VLDB 2026 Research / reviewers in the wild / expert
Yue Duan
dblp:10/9994
· DBLP profile ↗
32ranked-venue papers
11as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 18 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 9 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decomposing and Composing: Towards Efficient Vision-Language Continual Learning via Rank-1 Expert Pool in a Single LoRAabstractContinual learning (CL) in vision-language models (VLMs) faces significant challenges in improving task adaptation and avoiding catastrophic forgetting. Existing methods usually have heavy inference burden or rely on external knowledge, while Low-Rank Adaptation (LoRA) has shown potential in reducing these issues by enabling parameter-efficient tuning. However, considering directly using LoRA to alleviate the catastrophic forgetting problem is non-trivial, we introduce a novel framework that restructures a single LoRA module as a decomposable Rank-1 Expert Pool. Our method learns to dynamically compose a sparse, task-specific update by selecting from this expert pool, guided by the semantics of the [CLS] token. In addition, we propose an Activation-Guided Orthogonal (AGO) loss that orthogonalizes critical parts of LoRA weights across tasks. This sparse composition and orthogonalization enable fewer parameter updates, resulting in domain-aware learning while minimizing inter-task interference and maintaining downstream task performance. Extensive experiments across multiple settings demonstrate state-of-the-art results in all metrics, surpassing zero-shot upper bounds in generalization. Notably, it reduces trainable parameters by 96.7% compared to the baseline method, eliminating reliance on external datasets or task-ID discriminators. The merged LoRAs retain less weights and incur no inference latency, making our method computationally lightweight. Zhan Fa, Yue Duan, Jian Zhang 0090, Lei Qi 0001, Wanqi Yang, Yinghuan Shi |
AAAI | 2 |
| 2026 | PDLogger: Automatic Multi-log Generation for Practical Software Development
Shengchen Duan, Yihua Xu, Yue Duan |
DSN | 5 |
| 2026 | Large-Scale Security Analysis of Multi-Token Smart Contracts: Uncovering Hidden Flaws in Batch Transfers
Ashok Kasthuri, Sajad Meisami, Lingxiao Jiang, Binghui Wang, Yue Duan |
DSN | 5 |
| 2025 | Divide-And-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-Supervised Continual Learning
Yue Duan, Taicai Chen, Lei Qi 0001, Yinghuan Shi |
ICCV | 1 |
| 2025 | CollisionRepair: First-Aid and Automated Patching for Storage Collision Vulnerabilities in Smart Contracts
Wanjing Han, Yue Duan, Mu Zhang 0001 |
USENIX Security Symposium | 3 |
| 2025 | SigScope: Detecting and Understanding Off-Chain Message Signing-related Vulnerabilities in Decentralized ApplicationsabstractIn Web 3.0, an emerging paradigm of building decentralized applications or DApps is off-chain message signing, which has advantages in performance, cost efficiency, and usability compared to conventional transaction-signing schemes. However, message signing burdens DApp developers with extra coding complexity and message designing, leading to new security risks. Sajad Meisami, Hugo Dabadie, Song Li 0006, Yuzhe Tang, Yue Duan |
WWW | 5 |
| 2025 | DeepVMUnProtect: Neural Network-Based Recovery of VM-Protected Android Apps for Semantics-Aware Malware DetectionabstractThe emerging virtual machine-based Android packers render existing unpacking techniques ineffective. The state-of-the-art unpacker falls short because it relies on unreliable heuristics and manually crafted semantic models. Hence, it cannot precisely recover app semantics necessary for malware detection. In this paper, we proposeDeepVMUnProtect, a deep learning-based approach to automatically and accurately capture the semantics of VM-packed code, so as to facilitate semantic-based Android malware classification. Experiments have shown thatDeepVMUnProtectoutperforms the state-of-the-art tool on recovering opcode semantics in Qihoo(58.3%), Baidu(47.5%) and NMMP (58.8%) respectively, and can enable semantics-aware malware detection which prior work fails to do. Mu Zhang 0001, Xiaopeng Ke, Yue Duan, Sheng Zhong 0002, Fengyuan Xu |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image ClusteringabstractRecently, some works integrate SSL techniques into deep clustering frameworks to enhance image clustering performance. However, they all need pretraining, clustering learning, or a trained clustering model as prerequisites, limiting the flexible and out-of-box application of SSL learners in the image clustering task. This work introduces ASD, an adaptor that enables the cold-start of SSL learners for deep image clustering without any prerequisites. Specifically, we first randomly sample pseudo-labeled data from all unlabeled data, and set an instance-level classifier to learn them with semantically aligned instance-level labels. With the ability of instance-level classification, we track the class transitions of predictions on unlabeled data to extract high-level similarities of instance-level classes, which can be utilized to assign cluster-level labels to pseudo-labeled data. Finally, we use the pseudo-labeled data with assigned cluster-level labels to trigger a general SSL learner trained on the unlabeled data for image clustering. We show the superior performance of ASD across various benchmarks against the latest deep image clustering approaches and very slight accuracy gaps compared to SSL methods using ground-truth, e.g., only 1.33% on CIFAR-10. Moreover, ASD can also further boost the performance of existing SSL-embedded deep image clustering methods. Yue Duan, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | PG-LBO: Enhancing High-Dimensional Bayesian Optimization with Pseudo-Label and Gaussian Process GuidanceabstractVariational Autoencoder based Bayesian Optimization (VAE-BO) has demonstrated its excellent performance in addressing high-dimensional structured optimization problems. However, current mainstream methods overlook the potential of utilizing a pool of unlabeled data to construct the latent space, while only concentrating on designing sophisticated models to leverage the labeled data. Despite their effective usage of labeled data, these methods often require extra network structures, additional procedure, resulting in computational inefficiency. To address this issue, we propose a novel method to effectively utilize unlabeled data with the guidance of labeled data. Specifically, we tailor the pseudo-labeling technique from semi-supervised learning to explicitly reveal the relative magnitudes of optimization objective values hidden within the unlabeled data. Based on this technique, we assign appropriate training weights to unlabeled data to enhance the construction of a discriminative latent space. Furthermore, we treat the VAE encoder and the Gaussian Process (GP) in Bayesian optimization as a unified deep kernel learning process, allowing the direct utilization of labeled data, which we term as Gaussian Process guidance. This directly and effectively integrates the goal of improving GP accuracy into the VAE training, thereby guiding the construction of the latent space. The extensive experiments demonstrate that our proposed method outperforms existing VAE-BO algorithms in various optimization scenarios. Our code will be published at https://github.com/TaicaiChen/PG-LBO. Taicai Chen, Yue Duan, Dong Li 0016, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 2 |
| 2024 | Roll with the Punches: Expansion and Shrinkage of Soft Label Selection for Semi-supervised Fine-Grained LearningabstractWhile semi-supervised learning (SSL) has yielded promising results, the more realistic SSL scenario remains to be explored, in which the unlabeled data exhibits extremely high recognition difficulty, e.g., fine-grained visual classification in the context of SSL (SS-FGVC). The increased recognition difficulty on fine-grained unlabeled data spells disaster for pseudo-labeling accuracy, resulting in poor performance of the SSL model. To tackle this challenge, we propose Soft Label Selection with Confidence-Aware Clustering based on Class Transition Tracking (SoC) by reconstructing the pseudo-label selection process by jointly optimizing Expansion Objective and Shrinkage Objective, which is based on a soft label manner. Respectively, the former objective encourages soft labels to absorb more candidate classes to ensure the attendance of ground-truth class, while the latter encourages soft labels to reject more noisy classes, which is theoretically proved to be equivalent to entropy minimization. In comparisons with various state-of-the-art methods, our approach demonstrates its superior performance in SS-FGVC. Checkpoints and source code are available at https://github.com/NJUyued/SoC4SS-FGVC. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
AAAI | 1 |
| 2024 | Marco: A Stochastic Asynchronous Concolic ExplorerabstractConcolic execution is a powerful program analysis technique for code path exploration. Despite recent advances that greatly improved the efficiency of concolic execution engines, path constraint solving remains a major bottleneck of concolic testing. An intelligent scheduler for inputs/branches becomes even more crucial. Our studies show that the previously under-studied branch-flipping policy adopted by state-of-the-art concolic execution engines has several limitations. We propose to assess each branch by its potential for new code coverage from a global view, concerning the path divergence probability at each branch. To validate this idea, we implemented a prototype Marco and evaluated it against the state-of-the-art concolic executor on 30 real-world programs from Google's Fuzzbench, Binutils, and UniBench. The result shows that Marco can outperform the baseline approach and make continuous progress after the baseline approach terminates. Jie Hu 0031, Yue Duan, Heng Yin 0001 |
ICSE | 2 |
| 2024 | PC2: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal RetrievalabstractIn the realm of cross-modal retrieval, seamlessly integrating diverse modalities within multimedia remains a formidable challenge, especially given the complexities introduced by noisy correspondence learning (NCL). Such noise often stems from mismatched data pairs, which is a significant obstacle distinct from traditional noisy labels. This paper introduces Pseudo-Classification based Pseudo-Captioning (PC$^2$) framework to address this challenge. PC$^2$ offers a threefold strategy: firstly, it establishes an auxiliary "pseudo-classification" task that interprets captions as categorical labels, steering the model to learn image-text semantic similarity through a non-contrastive mechanism. Secondly, unlike prevailing margin-based techniques, capitalizing on PC$^2$'s pseudo-classification capability, we generate pseudo-captions to provide more informative and tangible supervision for each mismatched pair. Thirdly, the oscillation of pseudo-classification is borrowed to assistant the correction of correspondence. In addition to technical contributions, we develop a realistic NCL dataset called Noise of Web (NoW), which could be a new powerful NCL benchmark where noise exists naturally. Empirical evaluations of PC$^2$ showcase marked improvements over existing state-of-the-art robust cross-modal retrieval techniques on both simulated and realistic datasets with various NCL settings. The contributed dataset and source code are released at https://github.com/alipay/PC2-NoiseofWeb. Yue Duan, Zhangxuan Gu, Zhenzhe Ying, Lei Qi 0001, Changhua Meng, Yinghuan Shi |
ACM Multimedia | 1 |
| 2024 | SigmaDiff: Semantics-Aware Deep Graph Matching for Pseudocode Diffing
Lian Gao, Yu Qu, Yue Duan, Heng Yin 0001 |
NDSS | 4 |
| 2024 | An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
Shenao Yan, Yue Duan, Hanbin Hong, Kiho Lee, Doowon Kim, Yuan Hong 0001 |
USENIX Security Symposium | 3 |
| 2024 | MutexMatch: Semi-Supervised Learning With Mutex-Based Consistency RegularizationabstractThe core issue in semi-supervised learning (SSL) lies in how to effectively leverage unlabeled data, whereas most existing methods tend to put a great emphasis on the utilization of high-confidence samples yet seldom fully explore the usage of low-confidence samples. In this article, we aim to utilize low-confidence samples in a novel way with our proposed mutex-based consistency regularization, namely MutexMatch. Specifically, the high-confidence samples are required to exactly predict "what it is" by the conventional true-positive classifier (TPC), while low-confidence samples are employed to achieve a simpler goal-to predict with ease "what it is not" by the true-negative classifier (TNC). In this sense, we not only mitigate the pseudo-labeling errors but also make full use of the low-confidence unlabeled data by the consistency of dissimilarity degree. MutexMatch achieves superior performance on multiple benchmark datasets, i.e., Canadian Institute for Advanced Research (CIFAR)-10, CIFAR-100, street view house numbers (SVHN), self-taught learning 10 (STL-10), and mini-ImageNet. More importantly, our method further shows superiority when the amount of labeled data is scarce, e.g., 92.23% accuracy with only 20 labeled data on CIFAR-10. Code has been released at https://github.com/NJUyued/MutexMatch4SSL. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Towards Semi-supervised Learning with Non-random Missing LabelsabstractSemi-supervised learning (SSL) tackles the label missing problem by enabling the effective usage of unlabeled data. While existing SSL methods focus on the traditional setting, a practical and challenging scenario called label Missing Not At Random (MNAR) is usually ignored. In MNAR, the labeled and unlabeled data fall into different class distributions resulting in biased label imputation, which deteriorates the performance of SSL models. In this work, class transition tracking based Pseudo-Rectifying Guidance (PRG) is devised for MNAR. We explore the class-level guidance information obtained by the Markov random walk, which is modeled on a dynamically created graph built over the class tracking matrix. PRG unifies the historical information of class distribution and class transitions caused by the pseudo-rectifying procedure to maintain the model’s unbiased enthusiasm towards assigning pseudo-labels to all classes, so as the quality of pseudo-labels on both popular classes and rare classes in MNAR could be improved. Finally, we show the superior performance of PRG across a variety of MNAR scenarios, outperforming the latest SSL approaches combining bias removal solutions by a large margin. Code and model weights are available at https://github.com/NJUyued/PRG4SSL-MNAR. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 1 |
| 2023 | Proxy Hunting: Understanding and Characterizing Proxy-based Upgradeable Smart Contracts in Blockchains
William Edward Bodell III, Sajad Meisami, Yue Duan |
USENIX Security Symposium | 3 |
| 2022 | Towards Automated Safety Vetting of Smart Contracts in Decentralized ApplicationsabstractWe propose VetSC, a novel UI-driven, program analysis guided model checking technique that can automatically extract contract semantics in DApps so as to enable targeted safety vetting. To facilitate model checking, we extract business model graphs from contract code that capture its intrinsic business and safety logic. To automatically determine what safety specifications to check, we retrieve textual semantics from DApp user interfaces. To exclude untrusted UI text, we also validate the UI-logic consistency and detect any discrepancies. We have implemented VetSC and applied it to 34 real-world DApps. Experiments have demonstrated that VetSC can accurately interpret smart contract code, enable autonomous safety vetting, and discover safety risks in real-world Dapps. Using our tool, we have successfully discovered 19 new safety risks in the wild, such as expired lottery tickets and double voting. Yue Duan, Shucheng Li, Minghao Li 0003, Fengyuan Xu, Mu Zhang 0001 |
CCS | 1 |
| 2022 | DC-SSL: Addressing Mismatched Class Distribution in Semi-supervised LearningabstractConsistency-based Semi-supervised learning (SSL) has achieved promising performance recently. However, the success largely depends on the assumption that the labeled and unlabeled data share an identical class distribution, which is hard to meet in real practice. The distribution mismatch between the labeled and unlabeled sets can cause severe bias in the pseudo-labels of SSL, resulting in significant performance degradation. To bridge this gap, we put forward a new SSL learning framework, named Distribution Consistency SSL (DC-SSL), which rectifies the pseudolabels from a distribution perspective. The basic idea is to directly estimate a reference class distribution (RCD), which is regarded as a surrogate of the ground truth class distribution about the unlabeled data, and then improve the pseudo-labels by encouraging the predicted class distribution (PCD) of the unlabeled data to approach RCD gradually. To this end, this paper revisits the Exponentially Moving Average (EMA) model and utilizes it to estimate RCD in an iteratively improved manner, which is achieved with a momentum-update scheme throughout the training procedure. On top of this, two strategies are proposed for RCD to rectify the pseudo-label prediction, respectively. They correspond to an efficient training-free scheme and a training-based alternative that generates more accurate and reliable predictions. DC-SSL is evaluated on multiple SSL benchmarks and demonstrates remarkable performance improvement over competitive methods under matched- and mismatched-distribution scenarios. Zhen Zhao 0001, Luping Zhou, Yue Duan, Lei Wang 0001, Lei Qi 0001, Yinghuan Shi |
CVPR | 3 |
| 2022 | RDA: Reciprocal Distribution Alignment for Robust Semi-supervised Learning
Yue Duan, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi |
ECCV (30) | 1 |
| 2022 | Probabilistic Path Prioritization for Hybrid FuzzingabstractHybrid fuzzing that combines fuzzing and concolic execution has become an advanced technique for software vulnerability detection. Based on the observation that fuzzing and concolic execution are complementary in nature, state-of-the-art hybrid fuzzing systems deploy “optimal concolic testing” and “demand launch” strategies. Although these ideas sound intriguing, we point out several fundamental limitations in them, due to unrealistic or oversimplified assumptions. Further, we propose a novel “discriminative dispatch” strategy and design a probabilistic hybrid fuzzing system to better utilize the capability of concolic execution. Specifically, we design a Monte Carlo-based probabilistic path prioritization model to quantify each path’s difficulty, and then prioritize them for concolic execution. Our model assigns the most difficult paths to concolic execution. We implement a prototype named${\sf DigFuzz}$and evaluate our system with two representative datasets and real-world programs. Results show that the concolic execution in${\sf DigFuzz}$outperforms than those in state-of-the-art hybrid fuzzing systems in every major aspect. In particular, the concolic execution in${\sf DigFuzz}$contributes to discovering more vulnerabilities (12 versus 5) and producing more code coverage (18.9 versus 3.8 percent) on the CQE dataset than the concolic execution in Driller. Lei Zhao 0012, Pengcheng Cao, Yue Duan, Heng Yin 0001, Jifeng Xuan |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2020 | DeepBinDiff: Learning Program-Wide Code Representations for Binary Diffing
Yue Duan, Xuezixiang Li, Heng Yin 0001 |
NDSS | 1 |
| 2019 | Send Hardest Problems My Way: Probabilistic Path Prioritization for Hybrid Fuzzing
Lei Zhao 0012, Yue Duan, Heng Yin 0001, Jifeng Xuan |
NDSS | 2 |
| 2019 | Automatic Generation of Non-intrusive Updates for Third-Party Libraries in Android Applications
Yue Duan, Lian Gao, Jie Hu 0031, Heng Yin 0001 |
RAID | 1 |
| 2019 | Be Sensitive and Collaborative: Analyzing Impact of Coverage Metrics in Greybox Fuzzing
Yue Duan, Heng Yin 0001, Chengyu Song |
RAID | 2 |
| 2018 | Evaluation of GF-3 Quad-Polarized SAR Imagery for Coastal Wetland ObservationabstractThe objective of this paper is to evaluate the potential of Gaofen (GF)-3 quad-polarized Synthetic Aperture Radar (SAR) for coastal wetland observation. The north coastal wetland of Zhejiang is selected as the test case, which is operated by on-site measure as reference results, matched up with three scenes of GF-3 Quad Polarized Strip I (QPSI) mode SAR imagery. Based on the collected datasets, three well-known methods of Pauli, Freeman and H/A/Alpha decomposition are performed to data processing, and the experimental results reveal that reliable observation capability of GF-3 quad-polarized SAR is verified. The promising preliminary results indicate that GF-3 is encouraging for operational implementation, especially for coastal wetland observation. Yun Shao 0001, Wei Tian 0006, Yue Duan, Kun Li 0002, Long Liu 0002 |
IGARSS | 4 |
| 2018 | Things You May Not Know About Android (Un)Packers: A Systematic Study based on Whole-System Emulation
Yue Duan, Mu Zhang 0001, Abhishek Vasisht Bhaskar, Heng Yin 0001, Xiaorui Pan, Tongxin Li 0002, Xueqiang Wang, XiaoFeng Wang 0001 |
NDSS | 1 |
| 2017 | Dark Hazard: Learning-based, Large-Scale Discovery of Hidden Sensitive Operations in Android Apps
Xiaorui Pan, Xueqiang Wang, Yue Duan, XiaoFeng Wang 0001, Heng Yin 0001 |
NDSS | 3 |
| 2017 | JSForce: A Forced Execution Engine for Malicious JavaScript Detection
Xunchao Hu, Yue Duan, Andrew Henderson, Heng Yin 0001 |
SecureComm | 3 |
| 2015 | Towards Automatic Generation of Security-Centric Descriptions for Android AppsabstractTo improve the security awareness of end users, Android markets directly present two classes of literal app information: 1) permission requests and 2) textual descriptions. Unfortunately, neither can serve the needs. A permission list is not only hard to understand but also inadequate; textual descriptions provided by developers are not security-centric and are significantly deviated from the permissions. To fill in this gap, we propose a novel technique to automatically generate security-centric app descriptions, based on program analysis. We implement a prototype system, DescribeME, and evaluate our system using both DroidBench and real-world Android apps. Experimental results demonstrate that DescribeME enables a promising technique which bridges the gap between descriptions and permissions. A further user study shows that automatically produced descriptions are not only readable but also effectively help users avoid malware and privacy-breaching apps. Mu Zhang 0001, Yue Duan, Heng Yin 0001 |
CCS | 2 |
| 2014 | Semantics-Aware Android Malware Classification Using Weighted Contextual API Dependency GraphsabstractThe drastic increase of Android malware has led to a strong interest in developing methods to automate the malware analysis process. Existing automated Android malware detection and classification methods fall into two general categories: 1) signature-based and 2) machine learning-based. Signature-based approaches can be easily evaded by bytecode-level transformation attacks. Prior learning-based works extract features from application syntax, rather than program semantics, and are also subject to evasion. In this paper, we propose a novel semantic-based approach that classifies Android malware via dependency graphs. To battle transformation attacks, we extract a weighted contextual API dependency graph as program semantics to construct feature sets. To fight against malware variants and zero-day malware, we introduce graph similarity metrics to uncover homogeneous application behaviors while tolerating minor implementation differences. We implement a prototype system, DroidSIFT, in 23 thousand lines of Java code. We evaluate our system using 2200 malware samples and 13500 benign samples. Experiments show that our signature detection can correctly label 93\% of malware instances; our anomaly detector is capable of detecting zero-day malware with a low false negative rate (2\%) and an acceptable false positive rate (5.15\%) for a vetting purpose. Mu Zhang 0001, Yue Duan, Heng Yin 0001, Zhiruo Zhao |
CCS | 2 |
| 2011 | Automatic Reputation Computation through Document Analysis: A Social Network ApproachabstractWe develop and study two social network-based algorithms for automatically computing authors' reputations from a collection of textual documents. First, given a set of documents, both algorithms examine keyword reference behaviors of the authors to construct a social network. This social network represents the relationship among the authors in terms of information reference behavior. With the resulting network, the first algorithm computes each author's reputation value considering only direct referential activities while the second considers indirect activities as well. We discuss the reputation values computed by the two algorithms and compare them with the reputation ratings given by a human domain expert. We also analyze the social network through a community detection algorithm. We observed several interesting phenomena including the network being scale-free and having negative assortativity. JooYoung Lee, Yue Duan, Jae C. Oh, Wenliang Du 0001, Howard Blair, Lusha Wang |
ASONAM | 2 |