EDBT 2026 Demo / reviewers in the wild / expert
Tianyue Luo
dblp:168/9558
· DBLP profile ↗
29ranked-venue papers
1as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 12 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Security and privacy · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VFCionX: Bridging Large and Small Models for Robust Vulnerability-Fixing Commit IdentificationabstractVulnerability-Fixing Commit Identification(VFCI) is a critical task in software security maintenance that aims to automatically identify code commits that patch security vulnerabilities. However, existing approaches face challenges in handling low-quality commit messages and entangled commits, which limit their identification performance. To address these issues, we propose VFCionX, a novel VFCI framework that integrates large and small language models in a collaborative architecture. VFCionX consists of three core modules: Message Classifier, Patch Classifier, and Ensemble Classifier. The Message Classifier employs a multi-source contextual augmentation strategy to enhance the quality of commit messages and fine-tunes the Qwen2.5-1.5B model, significantly improving classification performance in the textual modality. The Patch Classifier combines heuristic rules with a Qwen2.5-Coder-7B-driven file selector to filter noise from entangled commits, and incorporates a line-level feature extractor based on CodeBERT and CNN to capture local pattern differences between added and deleted code lines. The Ensemble Classifier integrates predictions from both channels using the AdaBoost algorithm, enhancing model robustness and generalization. Experimental results on five popular C/C++ repositories comprising 24,630 commits show that VFCionX achieves an F1-score of 81.47%, outperforming the best baseline by 9.42%. Ablation studies validate the effectiveness of each component, while sensitivity analysis reveals optimal parameter settings for balancing performance and noise resilience. This work provides a new and effective solution for robust vulnerability patch identification. Xing Cui, JingZheng Wu, Wenxiang Ou, Tianyue Luo, Xiang Ling 0001 |
AAAI | 4 |
| 2026 | SimFuzz: Similarity-guided Block-level Mutation for RISC-V Processor FuzzingabstractThe Instruction Set Architecture (ISA) defines processor operations and serves as the interface between hardware and software. As an open ISA, RISC-V lowers the barriers to processor design and encourages widespread adoption, but also exposes processors to security risks such as functional bugs. Processor fuzzing is a powerful technique for automatically detecting these bugs. However, existing fuzzing methods suffer from two main limitations. First, their emphasis on redundant test case generation causes them to overlook cross-processor corner cases. Second, they rely too heavily on coverage guidance. Current coverage metrics are biased and inefficient, and become ineffective once coverage growth plateaus.To overcome these limitations, we propose SimFuzz, a fuzzing framework that constructs a high-quality seed corpus from historical bug-triggering inputs and employs similarity-guided, block-level mutation to efficiently explore the processor input space. By introducing instruction similarity, SimFuzz expands the input space around seeds while preserving control-flow structure, enabling deeper exploration without relying on coverage feedback. We evaluate SimFuzz on three widely used open-source RISC-V processors: Rocket, BOOM, and XiangShan, and discover 17 bugs in total, including 14 previously unknown issues, 7 of which have been assigned CVE identifiers. These bugs affect the decode and memory units, cause instruction and data errors, and can lead to kernel instability or system crashes. Experimental results show that SimFuzz achieves up to 73.22% multiplexer coverage on the high-quality seed corpus. Our findings highlight critical security bugs in mainstream RISC-V processors and offer actionable insights for improving functional verification. Hao Lyu 0002, JingZheng Wu, Xiang Ling 0001, Yicheng Zhong, Tianyue Luo |
DATE | 6 |
| 2026 | Towards Graph-Based Code Generation: Competition-Level Coding Agents with Graph-Based Reasoning and Backtracking
Huidi Zhu, JingZheng Wu, Xiang Ling 0001, Tianyue Luo, Chen Zhao 0024 |
KSEM (3) | 4 |
| 2026 | Self-Supervised Learning for Pre-Training 3D Point Clouds: A SurveyabstractPoint cloud data have been extensively studied due to their compact form and flexibility in representing complex 3D geometries and structures. The ability of point cloud data to accurately capture and represent intricate 3D geometry makes it an ideal choice for a wide range of applications, including 3D computer graphics, autonomous driving, robotics, and augmented reality, all of which require an understanding of the underlying geometry and spatial structures. Given the challenges associated with annotating large-scale point clouds, self-supervised point cloud representation learning has attracted increasing attention in recent years. It aims to learn generic and useful point cloud representations from unlabeled data, circumventing the need for extensive manual annotation. In this paper, we present a comprehensive survey of self-supervised point cloud representation learning using DNNs. We begin by presenting the motivation and general trends in recent research, then briefly introduce commonly used datasets and evaluation metrics. Next, we extensively explore self supervised point cloud representation learning methods. Finally, we share our thoughts on some of the challenges and potential issues that future research into self supervised learning for pre-training 3D point clouds may encounter. Our curated bibliography can be found at https://github.com/EtronTech/Awesome_3DSSL. Ben Fei, Weidong Yang 0001, Qingyuan Zhou, Liwen Liu, Tianyue Luo, Ying He 0001 |
Comput. Vis. Media | 7 |
| 2026 | Generative Diffusion Prior for Unified Image and Video Restoration & EnhancementabstractAbstract Existing image restoration methods primarily rely on the posterior distribution of natural images but are often limited by their dependence on known degradations and supervised training. To this end, we propose Generative Diffusion Prior (GDP), an unsupervised sampling-based framework that effectively models posterior distributions for image and video restoration. GDP utilizes a single pre-trained denoising diffusion probabilistic model (DDPM) to solve a wide range of linear, non-linear, and blind inverse problems without explicit degradation assumptions. Specifically, GDP systematically explores a conditional guidance protocol, which proves more practical and effective than conventional methods of adding guidance. Furthermore, GDP incorporates a degradation model optimization mechanism during the denoising process, enabling blind image restoration. Besides, we introduce a patch-based strategy, allowing GDP to handle images of arbitrary resolution. We extensively evaluate GDP on multiple image and video restoration tasks, including super-resolution, deblurring, inpainting, and colorization, as well as more challenging applications such as low-light enhancement, HDR recovery, and LDR video enhancement. Experimental results demonstrate that GDP outperforms leading unsupervised methods across diverse benchmarks in both reconstruction accuracy and perceptual quality, while demonstrating robust generalization to images and videos of any size. Our project page at https://generativediffusionprior.github.io/. Ben Fei, Zhaoyang Lyu, Liang Pan, Junzhe Zhang 0002, Weidong Yang 0001, Tianyue Luo, Jinyi Wang, Bo Dai 0002, Ying He 0001, Wanli Ouyang |
Int. J. Comput. Vis. | 6 |
| 2026 | LibPass: An Entropy-Guided Black-Box Adversarial Attack Against Third-Party Library Detection Tools in the WildabstractTo mitigate the security and compliance risks posed by Android third-party libraries (TPLs), researchers have proposed many automated detection tools aimed at accurately identifying TPLs within Android applications (apps). The detection results serve as the foundation for practical downstream tasks such as software bill of materials (SBOM) generation, n-day vulnerability identification, compliance auditing, and software supply chain risk tracing, among others. However, despite the high levels of detection accuracy and efficiency achieved by state-of-the-art TPL detection tools, existing studies lack a systematic evaluation of these tools' robustness against potential malicious attacks. As a result, the risk of detection failure in real-world scenarios remains unmanageable. Tools with poor robustness may be rendered ineffective under attack, allowing unsafe and non-compliant TPLs within apps to evade scrutiny and analysis, thereby compromising user interests. To bridge this gap, we propose the first adversarial attack against TPL detection tools,LibPass. The core idea ofLibPassis to generate adversarial apps by crafting perturbations that go beyond the code transformations introduced by obfuscation techniques. These adversarial perturbations hinder the generalization capability of existing detection tools, which are primarily designed to counter code obfuscation, thereby enabling the evasion of TPL detection. To minimize attack overhead and enhance stealthiness,LibPassemploys an improved firefly algorithm to search for optimal adversarial apps. This work evaluates the effectiveness ofLibPassagainst five state-of-the-art TPL detection tools on three datasets of different types and benchmarks its performance against three baseline attack methods. Experimental results demonstrate thatLibPassachieves an average attack success rate of 61.33%, with a peak of 99.46%, underscoring the insufficient robustness of current TPL detection tools. Bolin Zhou, JingZheng Wu, Xiang Ling 0001, Jingkun Zhang, Tianyue Luo |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | Detecting Malicious Packages in PyPI and NPM by Clustering Installation ScriptsabstractSoftware repositories such as PyPI and npm are vital for software development but expose users to serious security risks from malicious packages. The malicious packages often execute their payloads immediately upon installation, leading to rapid system compromise. Existing detection methods are heavily dependent on difficult-to-obtain explicit knowledge, rendering them susceptible to overlooking emergent malicious packages.In this paper, we present a lightweight and effective method, namely EMPHunter, to detect malicious packages without requiring any explicit prior knowledge. EMPHunter is founded upon two fundamental and insightful observations. First, malicious packages are considerably rarer than benign ones, and second, the functionality of installation scripts for malicious packages diverges significantly from those of benign packages, with the latter frequently forming clusters. Consequently, EMPHunter utilizes the clustering technique to group the unique installation scripts of new-uploaded packages and identifies outliers as candidate malicious packages. It then ranks the outliers according to their deviate degrees and the distance between each of them and known malicious instances, effectively highlighting potential malicious packages.With EMPHunter, we successfully identified 122 previously unknown malicious packages from a pool of 267,009 newly-uploaded PyPI and npm packages, achieving an mAP (Mean Average Precision) of 0.813 and an exceptional recall of 0.992 when auditing the top-10 rankings. All detected packages have been officially confirmed as genuine malicious package by PyPI and npm. We assert that EMPHunter offers a valuable and advantageous supplement to existing detection tools, augmenting the arsenal of software supply chain security analysis. Wentao Liang, Xiang Ling 0001, Chen Zhao 0024, JingZheng Wu, Tianyue Luo |
IEEE Trans. Software Eng. | 5 |
| 2025 | PyReach: A Multi-Agent Framework for Vulnerability Reachability Analysis in PythonabstractModern Python applications heavily rely on third-party libraries (TPLs), which can introduce security risks when vulnerabilities in these libraries silently propagate into client code. Determining whether a known vulnerability in a third-party library (TPL) can potentially be triggered in a specific downstream application is a key aspect of vulnerability reachability analysis, a research area that remains a manual, error-prone task due to the dynamic nature of the Python language and its implicit coding patterns. We present PyReach, a multi-agent collaborative framework that automates vulnerability reachability analysis for Python programs. Instead of statically resolving all dynamic behavior, PyReach decomposes the reasoning process into three semantically guided agents: 1) Context Modeling Agent that extracts auxiliary semantic context by analyzing and summarizing the semantic context of external dependencies for each function in a call chain; 2) Reachability Analysis Agent that determines whether a function in a call chain alters a vulner-ability’s triggering conditions by analyzing its inside semantics; 3) Reachability Verification Agent that determines if execution paths from user-facing entry points can reach the vulnerable code under the right conditions. We evaluate PyReach on a custom-built dataset, the largest of its kind for Python vulnerability reachability analysis, consisting of 15 real-world Python CVEs and 45 corresponding client projects. Experimental results show that PyReach achieves $90 \%$ precision and $83.3 \%$ specificity, significantly outperforming a call-graph-based baseline, Jarvis. PyReach effectively distinguishes between truly affected and unaffected clients by reasoning over code semantics and trigger profiles. Our results highlight the value of combining modular semantic reasoning with constraint propagation for accurate and scalable vulnerability analysis in dynamic languages. Yueqin Wang, JingZheng Wu, Xiang Ling 0001, Tianyue Luo |
APSEC | 4 |
| 2025 | Version-level Third-Party Library Detection in Android Applications via Class Structural SimilarityabstractAndroid applications (apps) integrate reusable and well-tested third-party libraries (TPLs) to enhance functionality and shorten development cycles. However, recent research reveals that TPLs have become the largest attack surface for Android apps, where the use of insecure TPLs can compromise both developer and user interests. To mitigate such threats, researchers have proposed various tools to detect TPLs used by apps, supporting further security analyses such as vulnerable TPLs identification. Although existing tools achieve notable library-level TPL detection performance in the presence of obfuscation, they struggle with version-level TPL detection due to a lack of sensitivity to differences between versions. This limitation results in a high version-level false positive rate, significantly increasing the manual workload for security analysts. To resolve this issue, we propose SAD, a TPL detection tool with high version-level detection performance. SAD generates a candidate app class list for each TPL class based on the feature of nodes in class dependency graphs (CDGs). It then identifies the unique corresponding app class for each TPL class by performing class matching based on the similarity of their class summaries. Finally, SAD identifies TPL versions by evaluating the structural similarity of the sub-graph formed by matched classes within the CDGs of the TPL and the app. Extensive evaluation on three datasets demonstrates the effectiveness of SAD and its components. SAD achieves F1 scores of 97.64% and 84.82% for library-level and version-level detection on obfuscated apps, respectively, surpassing existing state-of-the-art tools. The version-level false positives reported by the best tool is 1.61 times that of SAD. We further evaluate the degree to which TPLs identified by detection tools correspond to actual TPL classes. Experimental results show that SAD achieves a class-level F1 score of 94.12%, 11% higher than the best tool, demonstrating the reliability of SAD and better supporting downstream tasks that rely on specific code. Bolin Zhou, JingZheng Wu, Xiang Ling 0001, Tianyue Luo, Jingkun Zhang |
EASE | 4 |
| 2025 | We Know What You're Looking For: Recommendation for Large-Scale Open Source SoftwareabstractBackground: In recent years, with the advancement of software engineering technologies and industry, Open Source Software (OSS) has become a mainstream model for software development and innovation. Increasingly, organizations and developers are adopting and customizing existing OSS to simplify and accelerate development processes. During OSS adoption, recommending suitable software based on user needs is crucial for enhancing development efficiency and addressing diverse requirements. However, the vast number and diversity of OSS make the recommendation task highly challenging. Despite progress in previous research, several issues remain, such as neglect of key software attributes, complexity in extracting multilingual features, and challenges of cold start and data sparsity. Aims: This paper presents AthenaRec, a large-scale OSS recommendation system comprising three core modules: Delphi, Argus, and Hestia. AthenaRec aims to recommend relevant and suitable software from a vast OSS based on user needs. Method: Specifically, Delphi first analyzes user queries to identify intention; Argus employs a heterogeneous ensemble recall approach to retrieve a large set of candidate software relevant to the identified intention; finally, Hestia adopts a two-stage deep ranking strategy. It performs coarse ranking by integrating multilingual modeling with contrastive learning, followed by fine ranking with a large language model, augmented by retrieval-augmented generation to incorporate external evidence. To evaluate the effectiveness of AthenaRec, we use a query dataset from real application scenarios. Results: Experimental results demonstrate that, on the test set of 7,500 queries, AthenaRec achieves superior recommendation performance, with Hits@20, MAP@20, NDCG@20, and MRR scores of$98.27 \%, 95.60 \%, 95.05 {\%}$, and 92.92%, respectively. On average, AthenaRec outperforms other top methods by 10.9% across all evaluation metrics. Conclusions: Additionally, we develop a Visual Studio Code (VSCode) plugin based on AthenaRec, which can be accessed via URL. We intend for this research to provide a reference for software developers, advancing the efficiency and accuracy of OSS recommendation. Xing Cui, JingZheng Wu, Xiang Ling 0001, Tianyue Luo |
ESEM | 4 |
| 2025 | The Seeds of the Future Sprout from History: Fuzzing for Unveiling Vulnerabilities in Prospective Deep-Learning LibrariesabstractThe widespread application of large language models (LLMs) underscores the importance of deep learning (DL) technologies that rely on foundational DL libraries such as PyTorch and TensorFlow. Despite their robust features, these libraries face challenges with scalability and adaptation to rapid advancements in the LLM community. In response, tech giants like Apple and Huawei are developing their own DL libraries to enhance performance, increase scalability, and safeguard intellectual property. Ensuring the security of these libraries is crucial, with fuzzing being a vital solution. However, existing fuzzing frameworks struggle with target flexibility, effectively testing bug-prone API sequences, and leveraging the limited available information in new libraries. To address these limitations, we propose FUTURE, the first universal fuzzing framework tailored for newly introduced and prospective DL libraries. FUTURE leverages historical bug information from existing libraries and fine-tunes LLMs for specialized code generation. This strategy helps identify bugs in new libraries and uses insights from these libraries to enhance security in existing ones, creating a cycle from history to future and back. To evaluate FUTURE's effectiveness, we conduct comprehensive evaluations on three newly introduced DL libraries. Evaluation results demonstrate that FUTURE significantly outperforms existing fuzzers in bug detection, success rate of bug reproduction, validity rate of code generation, and API coverage. Notably, FUTURE has detected 148 bugs across 452 targeted APIs, including 142 previously unknown bugs. Among these, 10 have been assigned CVE IDs. Additionally, FUTURE detects 7 bugs in PyTorch, demonstrating its ability to enhance security in existing libraries in reverse. JingZheng Wu, Xiang Ling 0001, Tianyue Luo, Zhiqing Rui |
ICSE | 4 |
| 2025 | RMGenie: An LLM-Based Agent Framework for Open Source Software README GenerationabstractOpen Source Software (OSS) plays a vital role in modern software ecosystems, with README files providing essential information on functionality, configuration, and usage. However, approximately 9.03% of OSS projects lack adequate README documentation, impacting developer efficiency and software maintainability. Recent advances in natural language processing and large language models (LLMs) have automated README generation to reduce manual effort. Yet, existing methods struggle with integrating real-time external knowledge, parsing complex software structures, and adapting to projectspecific requirements. To address these challenges, this paper introduces RMGenie, a framework that leverages LLM-based agents and external tool invocation for automated README generation. RMGenie constructs an agent-driven workflow enabling multi-round interactions with external tools to dynamically extract critical code insights, particularly for complex software structures. It employs a Tree of Actions model with an entropy-based scoring mechanism to optimize decision paths for README generation. Additionally, a reflexion Mechanism enhances accuracy by mitigating decision biases and tool invocation errors. Experimental results confirm that RMGenie significantly outperforms baseline methods in content completeness, instruction adherence, and factual accuracy. Furthermore, we develop a plugin that integrates RMGenie to automatically generate README files from GitHub repositories, improving developer productivity and standardizing OSS documentation. The plugin is available at VSCode Marketplace. Xing Cui, JingZheng Wu, Tianyue Luo, Xiang Ling 0001 |
ICSME | 4 |
| 2025 | OptionFuzz: Fuzzing SMT Solvers with Optimized Option Exploration via Large Language ModelsabstractSatisfiability Modulo Theory (SMT) solvers play a crucial role in various domains and applications. Therefore, ensuring their correctness and robustness becomes increasingly vital. Fuzzing is an efficient and effective method for validating the quality of SMT solvers, utilizing inputs that consist of solving formulas and configuration options. However, existing fuzzing methods focus solely on generating formulas or simply combining options and formulas, neglecting the complex interactions between options. Yet, randomly combining multiple options can lead to a combinatorial explosion and result in numerous invalid inputs. To overcome these limitations, we propose OptionFuzz, a fuzzer that optimizes option exploration by identifying relationships between solver options, reducing invalid inputs and mitigating combinatorial explosion. OptionFuzz identifies option relationships using large language models (LLMs), which analyze official documentation of options. These identified relationships are transformed to a relation graph, enabling efficient traversal to derive related option combinations and generate high-quality fuzz inputs. To evaluate OptionFuzz's effectiveness, we conduct comprehensive evaluations on two state-of-the-art SMT solvers, Z3 and CVC5. OptionFuzz demonstrates its effectiveness by accurately extracting option relationships with an accuracy of$\mathbf{9 5. 2 3 \%}$and a recall rate of$\mathbf{9 0. 1 0 \%}$. Leveraging these relationships, OptionFuzz reduces the number of options combinations to be tested by$\mathbf{7 0. 1 1 \%}$. Notably, OptionFuzz has detected$\mathbf{3 4}$unique bugs, 20 of which have been fixed by developers, and 5 have been assigned CVE IDs due to their severity. Yuhao Peng, JingZheng Wu, Xiang Ling 0001, Tianyue Luo |
ICSME | 5 |
| 2025 | Shrunk, Yet Complete: Code Shrinking-Resilient Android Third-Party Library DetectionabstractManaging third-party libraries is a costly and critical task for enterprises, essential for both vulnerability assessment and license compliance. Existing android software composition analysis tools focus on mitigating code obfuscation but neglect the impact of code optimization, which is deeply integrated into build pipelines and disrupts library structure.To tackle these challenges, we developed LibSleuth, a detection tool designed to be resilient to code shrinking and obfuscation. It is based on the observation that even after shrinking, the remaining code still retains functional completeness. LibSleuth adopts two novel strategies: (1) Method level functional module matching: We break down feature matching to method level and define a functional module as related methods that represent used functionality. This allows us to detect libraries based on functional module completeness to address code shrinking. (2) Context-enhanced multi-level filtering: To improve robustness against obfuscation and reduce the cost of pairing, LibSleuth leverages contextual relationships to enhance feature stability and adopts a coarse-to-fine progressive matching process.We evaluated LibSleuth on datasets containing obfuscated and optimized Android apps. LibSleuth outperforms state-of-the-art academic and commercial tools in both scenarios. Under combined code shrinking and obfuscation, LibSleuth achieves an average 27.74% higher version level F1-score. Moreover, our analysis of 10,000 real world Android apps shows that 20.35% still depend on vulnerable library, demonstrating the practical utility of LibSleuth for downstream tasks. Jingkun Zhang, JingZheng Wu, Xiang Ling 0001, Tianyue Luo, Bolin Zhou, Mutian Yang |
ASE | 4 |
| 2025 | CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge DeploymentabstractProprietary large language models (LLMs) exhibit strong generalization capabilities across diverse tasks and are increasingly deployed on edge devices for efficiency and privacy reasons. However, deploying proprietary LLMs at the edge without adequate protection introduces critical security threats. Attackers can extract model weights and architectures, enabling unauthorized copying and misuse. Even when protective measures prevent full extraction of model weights, attackers may still perform advanced attacks, such as fine-tuning, to further exploit the model. Existing defenses against these threats typically incur significant computational and communication overhead, making them impractical for edge deployment.
To safeguard the edge-deployed LLMs, we introduce CoreGuard, a computation- and communication-efficient protection method. CoreGuard employs an efficient protection protocol to reduce computational overhead and minimize communication overhead via a propagation protocol. Extensive experiments show that CoreGuard achieves upper-bound security protection with negligible overhead. Qinfeng Li, Tianyue Luo, Xuhong Zhang 0002, Yangfan Xie, Yier Jin, Hao Peng 0002, Xinkui Zhao, Xianwei Zhu, Jianwei Yin |
NeurIPS | 2 |
| 2025 | Point Patches Contrastive Learning for Enhanced Point Cloud CompletionabstractIn partial-to-complete point cloud completion, it is imperative that enabling every patch in the output point cloud faithfully represents the corresponding patch in partial input, ensuring similarity in terms of geometric content. To achieve this objective, we propose a straightforward method dubbed PPCL that aims to maximize the mutual information between two point patches from the encoder and decoder by leveraging a contrastive learning framework. Contrastive learning facilitates the mapping of two similar point patches to corresponding points in a learned feature space. Notably, we explore multi-layer point patches contrastive learning (MPPCL) instead of operating on the whole point cloud. The negatives are exploited within the input point cloud itself rather than the rest of the datasets. To fully leverage the local geometries present in the partial inputs and enhance the quality of point patches in the encoder, we introduce Multi-level Feature Learning (MFL) and Hierarchical Feature Fusion (HFF) modules. These modules are also able to facilitate the learning of various levels of features. Moreover, Spatial-Channel Transformer Point Up-sampling (SCT) is devised to guide the decoder to construct a complete and fine-grained point cloud by leveraging enhanced point patches from our point patches contrastive learning. Extensive experiments demonstrate that our PPCL can achieve better quantitive and qualitative performance over off-the-shelf methods across various datasets. Ben Fei, Liwen Liu, Tianyue Luo, Weidong Yang 0001, Lipeng Ma, Zhijun Li 0001, Wenming Chen 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Curriculumformer: Taming Curriculum Pre-Training for Enhanced 3-D Point Cloud UnderstandingabstractLearning universal representations of 3-D point clouds is essential for reducing the need for manual annotation of large-scale and irregular point cloud datasets. The current modus operandi for representative learning is self-supervised learning, which has shown great potential for improving point cloud understanding. Nevertheless, it remains an open problem how to employ auto-encoding for learning universal 3-D representations of irregularly structured point clouds, as previous methods focus on either global shapes or local geometries. To this end, we present a cascaded self-supervised point cloud representation learning framework, dubbed Curriculumformer, aiming to tame curriculum pre-training for enhanced point cloud understanding. Our main idea lies in devising a progressive pre-training strategy, which trains the Transformer in an easy-to-hard manner. Specifically, we first pre-train the Transformer using an upsampling strategy, which allows it to learn global information. Then, we follow up with a completion strategy, which enables the Transformer to gain insight into local geometries. Finally, we propose a Multi-Modal Multi-Modality Contrastive Learning (M4CL) strategy to enhance the ability of representation learning by enriching the Transformer with semantic information. In this way, the pre-trained Transformer can be easily transferred to a wide range of downstream applications. We demonstrate the superior performance of Curriculumformer on various discriminant and generative tasks, outperforming state-of-the-art methods. Moreover, Curriculumformer can also be integrated into other off-the-shelf methods to promote their performance. Our code is available at https://github.com/Fayeben/Curriculumformer. Ben Fei, Tianyue Luo, Weidong Yang 0001, Liwen Liu, Rui Zhang 0103, Ying He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | VulDL: Tree-based and Graph-based Neural Networks for Vulnerability Detection and LocalizationabstractWith the dramatic increase in the number and size of software in the industry, tremendous research has been studied to automatically detect vulnerabilities. However, existing detection methods have limitations in code semantic modeling and detection granularity, which makes them unable to meet the requirements of high accuracy and fine granularity at the same time. In this paper, we propose a general framework, namely VulDL, which can effectively identify whether a given code snippet has a vulnerability and locate the specific code line where the vulnerability resides. VulDL first represents the source code as a novel semantic data structure, namely the adapted code property graph. After that, tree-based and graph-based neural networks are designed, which learn features according to the hierarchies and neighborhoods, and further realize vulnerability identification and localization. Our evaluation shows that VulDL achieves F1-scores of 98.68% and 94.85% in the identification of buffer error and resource management error vulnerabilities and 97.73% on their combined vulnerabilities. On the FFmpeg+QEMU dataset, VulDL achieves an F1-score of 59.62%, which is more effective than existing methods. Besides, VulDL can locate vulnerabilities at the statement granularity with F1-scores of 97.88%, 98.31%, and 99.16% on the evaluated datasets. JingZheng Wu, Xiang Ling 0001, Xu Duan, Tianyue Luo, Mutian Yang |
EASE | 4 |
| 2024 | Towards Query-Efficient Decision-Based Adversarial Attacks Through Frequency DomainabstractDeep neural networks are vulnerable to adversarial examples, where decision-based attacks can generate adversarial examples based solely on the predicted labels. However, these attacks typically require excessive queries to attack one example. Considering this challenge, we propose FBA (Frequency based Boundary Attack), a decision-based attack against the limitation of query efficiency. FBA incorporates a novel search process, utilizing high-frequency based importance sampling for efficient gradient estimation. Empirical results confirm the superior query efficiency of our method. Specifically, FBA surpasses SOTA attacks by achieving a 54% average improvement in query efficiency, quantified by the reduction in perturbation size within the same number of queries. Jianhao Fu, Xiang Ling 0001, Yaguan Qian, Changjiang Li, Tianyue Luo, JingZheng Wu |
ICME | 5 |
| 2024 | APP-Miner: Detecting API Misuses via Automatically Mining API Path PatternsabstractExtracting API patterns from the source code has been extensively employed to detect API misuses. However, recent studies manually provide pattern templates as prerequisites, requiring prior software knowledge and limiting their extraction scope. This paper presents APP-Miner (API path pattern miner), a novel static analysis framework for extracting API path patterns via a frequent subgraph mining technique without pattern templates. The critical insight is that API patterns usually consist of APIs’ data-related operations and are commonplace. Therefore, we define API paths as the control flow graphs composed of APIs’ data-related operations, and thereby the maximum frequent subgraphs of the API paths are the probable API path patterns. We implemented APP-Miner and extensively evaluated it on four widely used open-source software: Linux kernel, OpenSSL, FFmpeg, and Apache httpd. We found 116, 35, 3, and 3 new API misuses from the above systems, respectively. Moreover, we gained 19 CVEs. Jiasheng Jiang, JingZheng Wu, Xiang Ling 0001, Tianyue Luo, Sheng Qu |
SP | 4 |
| 2024 | A Wolf in Sheep's Clothing: Practical Black-box Adversarial Attacks for Evading Learning-based Windows Malware Detection in the Wild
Xiang Ling 0001, Zhiyu Wu, Bin Wang 0062, JingZheng Wu, Shouling Ji, Tianyue Luo |
USENIX Security Symposium | 7 |
| 2023 | Generative Diffusion Prior for Unified Image Restoration and EnhancementabstractExisting image restoration methods mostly leverage the posterior distribution of natural images. However, they often assume known degradation and also require supervised training, which restricts their adaptation to complex real applications. In this work, we propose the Generative Diffusion Prior (GDP) to effectively model the posterior distributions in an unsupervised sampling manner. GDP utilizes a pre-train denoising diffusion generative model (DDPM) for solving linear inverse, non-linear, or blind problems. Specifically, GDP systematically explores a protocol of conditional guidance, which is verified more practical than the commonly used guidance way. Furthermore, GDP is strength at optimizing the parameters of degradation model during the denoising process, achieving blind image restoration. Besides, we devise hierarchical guidance and patch-based methods, enabling the GDP to generate images of arbitrary resolutions. Experimentally, we demonstrate GDP's versatility on several image datasets for linear problems, such as super-resolution, deblurring, inpainting, and colorization, as well as non-linear and blind issues, such as low-light enhancement and HDR image recovery. GDP outperforms the current leading unsupervised methods on the diverse benchmarks in reconstruction quality and perceptual quality. Moreover, GDP also generalizes well for natural images or synthesized images with arbitrary sizes from various tasks out of the distribution of the ImageNet training set. The project page is available at https://generativediffusionprior.github.io/ Ben Fei, Zhaoyang Lyu, Liang Pan, Junzhe Zhang 0002, Weidong Yang 0001, Tianyue Luo, Bo Zhang 0069, Bo Dai 0002 |
CVPR | 6 |
| 2023 | Automatic Program Repair via Learning Edits on Sequence Code Property GraphabstractIn recent years, deep learning has been widely applied in the research of automatic program repair, namely, learning-based program repair, which treats the program repair task as either a neural machine translation problem or a transformation toward the abstract syntax tree of the program. However, all existing learning-based program repair solutions have not adequately modeled the deep semantic information of the programs. To be specific, those solutions based on neural machine translation do not involve any semantic information of the programs at all, while other solutions based on abstract syntax tree transformations also do not involve information such as control flow and data flow. Additionally, the neural networks they use are also unable to sufficiently learn the semantic information. As a result, they cannot fix some bugs that involve semantic information.In this paper, we propose a new learning-based program repair approach called SGEPR. SGEPR uses a novel intermediate representation named sequence code property graph (SCPG) to model program semantic information, which can cover multiple types of semantic information in the program. SGEPR then utilizes a graph neural network based on the attention mechanism to learn the semantic information within SCPG. Finally, a well-designed prediction network is employed to generate the repair patches for the buggy program. To evaluate the performance of SGEPR, we conducted comparative experiments on a real-world dataset and the results show that SGEPR can achieve a top-3 accuracy of 56.85%, which is 15.69% higher than the state-of-the-art learning-based program repair method. JingZheng Wu, Xiang Ling 0001, Tianyue Luo |
ICPADS | 4 |
| 2023 | A Needle is an Outlier in a Haystack: Hunting Malicious PyPI Packages with Code ClusteringabstractAs the most popular Python software repository, PyPI has become an indispensable part of the Python ecosystem. Regrettably, the open nature of PyPI exposes end-users to substantial security risks stemming from malicious packages. Consequently, the timely and effective identification of malware within the vast number of newly-uploaded PyPI packages has emerged as a pressing concern. Existing detection methods are dependent on difficult-to-obtain explicit knowledge, such as taint sources, sinks, and malicious code patterns, rendering them susceptible to overlooking emergent malicious packages. In this paper, we present a lightweight and effective method, namely MPHunter, to detect malicious packages without requiring any explicit prior knowledge. MPHunter is founded upon two fundamental and insightful observations. First, malicious packages are considerably rarer than benign ones, and second, the functionality of installation scripts for malicious packages diverges significantly from those of benign packages, with the latter frequently forming clusters. Consequently, MPHunter utilizes clustering techniques to group the installation scripts of PyPI packages and identifies outliers. Subsequently, MPHunter ranks the outliers according to their outlierness and the distance between them and known malicious instances, thereby effectively highlighting potential evil packages. With MPHunter, we successfully identified 60 previously unknown malicious packages from a pool of 31,329 newly-uploaded packages over a two-month period. All of them have been confirmed by the PyPI official. Moreover, a manual analysis shows that MPHunter recognizes all potentially malicious installation scripts with a recall of 100% across all analyzed packages. We assert that MPHunter offers a valuable and advantageous supplement to existing detection techniques, augmenting the arsenal of software supply chain security analysis. Wentao Liang, Xiang Ling 0001, JingZheng Wu, Tianyue Luo |
ASE | 4 |
| 2023 | Adversarial attacks against Windows PE malware detection: A survey of the state-of-the-art
Xiang Ling 0001, Lingfei Wu 0001, Jiangyu Zhang, Zhenqing Qu, Xiang Chen 0017, Yaguan Qian, Chunming Wu 0001, Shouling Ji, Tianyue Luo, JingZheng Wu |
Comput. Secur. | 10 |
| 2022 | Cross Platform API Mappings based on API Documentation GraphsabstractAs different versions of the same application might be implemented based on different platforms/programming languages, it is significantly important to build an automated migration tool for the application programming interface (API) mapping relations between different platforms/programming languages. In this paper, we propose an approach to discover API mappings based on the API documentation. We first divide the information in the API documentation into different types of entities, relations, and attributes to construct their respective API Documentation Graphs (ADGs). Then, we encode nodes, edges and triplets of ADGs and input them to a new graph neural network (GNN) for entity alignment to obtain the API mappings between the two different platforms/programming languages. Taking HarmonyOS and Android as representative cases, we evaluate our approach based on their API documentation. The results show that our approach improves top-1, top-5, and top10 accuracies by 50.57%, 56.25%, and 52.66%, respectively, compared with documentation-based baselines. Yanjie Shao, Tianyue Luo, Xiang Ling 0001, Senwen Zheng |
QRS | 2 |
| 2019 | VulSniper: Focus Your Attention to Shoot Fine-Grained VulnerabilitiesabstractWith the explosive development of information technology, vulnerabilities have become one of the major threats to computer security. Most vulnerabilities with similar patterns can be detected effectively by static analysis methods. However, some vulnerable and non-vulnerable code is hardly distinguishable, resulting in low detection accuracy. In this paper, we define the accurate identification of vulnerabilities in similar code as a fine-grained vulnerability detection problem. We propose VulSniper which is designed to detect fine-grained vulnerabilities more effectively. In VulSniper, attention mechanism is used to capture the critical features of the vulnerabilities. Especially, we use bottom-up and top-down structures to learn the attention weights of different areas of the program. Moreover, in order to fully extract the semantic features of the program, we generate the code property graph, design a 144-dimensional vector to describe the relation between the nodes, and finally encode the program as a feature tensor. VulSniper achieves F1-scores of 80.6% and 73.3% on the two benchmark datasets, the SARD Buffer Error dataset and the SARD Resource Management Error dataset respectively, which are significantly higher than those of the state-of-the-art methods. Xu Duan, JingZheng Wu, Shouling Ji, Zhiqing Rui, Tianyue Luo, Mutian Yang |
IJCAI | 5 |
| 2015 | POSTER: PatchGen: Towards Automated Patch Detection and Generation for 1-Day VulnerabilitiesabstractA large fraction of source code in open-source systems such as Linux contain 1-day vulnerabilities. The command "patch" is used to apply the patches to source codes, and returns feedback information automatically. Unfortunately, this operation is not always successful when patching directly, and two typical error scenarios may occur as follows. 1. The patch may be applied in wrong place, meaning the fix location should be adjusted in patch. 2. The patch may be applied repeatedly, meaning a verification should be executed before applying. To resolve the above scenarios, we propose PatchGen, a new system to quickly detect and generate patches for 1-day vulnerabilities in OS distributions. Comparing with the previous works on 1-day vulnerabilities detection, PatchGen is able to solve the above two error scenarios and use a quick, syntax-based approach that scales to OS distribution-sized code base no matter the code written in what types of language. We implement the PatchGen prototype, and evaluate it by checking all codes from packages in Ubuntu Maverick/Oneiric, all SourceForge C and C++ projects, and the Linux kernel source. Specifically, it takes less than 10 minutes for PatcheGen to detect 175 1-day vulnerabilities and generate 140 patches for Linux Kernel. All of the results have been manually confirmed and tested in the real systems. Tianyue Luo, Chen Ni, Mutian Yang, JingZheng Wu |
CCS | 1 |
| 2015 | POSTER: biTheft: Stealing Your Secrets by Bidirectional Covert Channel Communication with Zero-Permission Android ApplicationabstractAndroid has 81.5% of the smartphone market now, and it is also suffering from the explosive growth of malicious applications (or apps). These apps steal users' secret data and transmit it out of the phones. By analyzing the required permissions and the abnormal behaviors, some malicious apps may be easily detected. However, in this paper, we present a bidirectional covert channel in Android, named biTheft, which steals secrets and privacies covertly without any permission. biTheft firstly collects secret data from a set of unprotected shared resources in Android system. Then, it analyzes and infers secrets from the data. With the Intent mechanism, biTheft transmits secrets by legally launching some activities of other apps without requiring any permission itself. biTheft also monitors the usages and statuses of the shared resources to receive commands from remote server. We implement a biTheft scenario, and demonstrate that some types of secrets can be stolen and transmitted out. With pre-agreement, biTheft dynamically adjusts according with the remote server commands. Comparing with the traditional covert channels, biTheft is more practical in the real world scenarios. JingZheng Wu, Mutian Yang, Zhifei Wu, Tianyue Luo, Yongji Wang 0002 |
CCS | 5 |