VLDB 2026 Research / reviewers in the wild / expert
Lei Duan
dblp:04/545
· DBLP profile ↗
50ranked-venue papers in the field
10as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 26 (3 first)Data Mining & Knowledge Discovery · 16 (6 first)Information Retrieval & Web Search · 6Other / Interdisciplinary · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AURORA: An Adaptive Multi-granularity Graph Learning Framework for Drug Repositioning
Yidan Zhang 0001, Lei Duan, Huiru Zheng, Jiaxuan Xu 0001 |
DASFAA (3) | 2 |
| 2026 | SADD-RFCO: semi-supervised anomalous data detection based on random forest with co-training
Song Deng, Mengfei Sun, Lei Duan, Yi He 0007 |
Knowl. Inf. Syst. | 3 |
| 2026 | Dynamic Anchor-Based One-Step Hypergraph Ensemble Clustering
Jiaxuan Xu 0001, Lei Duan, Xinye Wang, Liang Du 0003, Yidan Zhang 0001, Zhen Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Perspective-Based Multi-task Learning for Outlier Interpretation
Zhuoling Li, Lili Guan, Xinye Wang, Zhengyong Pan, Lei Duan |
DASFAA (4) | 5 |
| 2025 | MVIC: Multi-view Information Collaborative Fusion for Drug-Drug Interaction Prediction
Xianxian Zhao, Chengxin He, Lei Duan |
DASFAA (2) | 3 |
| 2025 | mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUsabstractTransformer-based large language models (LLMs) have demonstrated outstanding performance across diverse domains, particularly in the emerging pretrain-then-finetune paradigm. LoRA, a parameter-efficient fine-tuning method, is commonly used to adapt a base LLM to multiple downstream tasks. Further, LLM platforms enable developers to fine-tune multiple models and develop various domain-specific applications simultaneously. However, existing model parallelism schemes suffer from high communication overhead and inefficient GPU utilization. In this paper, we present mLoRA, a parallelism-efficient fine-tuning system designed for training multiple LoRA across GPUs and machines. mLoRA introduces a novel LoRA-aware pipeline parallelism scheme that efficiently pipelines LoRA adapters and their distinct fine-tuning stages across GPUs and machines, along with a new LoRA-efficient operator to enhance GPU utilization. Our extensive evaluation shows that mLoRA can significantly reduce average fine-tuning task completion time, e.g., by 30%, compared to state-of-the-art methods like FSDP. More importantly, mLoRA enables simultaneous fine-tuning of larger models, e.g., two Llama-2-13B models on four NVIDIA RTX A6000 48GB GPUs, which is not feasible for FSDP due to high memory requirements. Hence, mLoRA not only increases fine-tuning efficiency but also makes it more accessible on cost-effective GPUs. Zhengmao Ye, Dengchun Li, Zetao Hu, Tingfeng Lan, Jian Sha, Shicong Zhang, Lei Duan, Jie Zuo, Hui Lu 0001, Yuanchun Zhou, MingJie Tang |
Proc. VLDB Endow. | 7 |
| 2024 | Community-Guided Contrastive Learning with Anomaly-Aware Reconstruction for Anomaly Detection on Attributed Networks
Xinye Wang, Chengxin He, Xiaocong Chen, Zhaohang Luo, Lei Duan, Jie Zuo |
DASFAA (7) | 6 |
| 2024 | An Efficient Adaptive Multi-Kernel Learning With Safe Screening Rule for Outlier DetectionabstractRecent advances in multi-kernel-based methods for outlier detection have positioned them as an attractive way to detect instances that are markedly different from the remaining data in a dataset. Currently, most outlier detection approaches based on multi-kernel learning are simply a convex combination of various kernels with handcrafted weights, meaning that these weights may not be suitable. Meanwhile, this combination of weights does not sufficiently consider the intrinsic correlations of instances when fusing different kernels. Thus, a key challenge is how to adaptively learn an appropriate combination of weights for capturing a new feature space in which outliers can be better detected than the original space. Simultaneously, it is still a burning issue to get the optimal combination of weights due to considerable computational cost and memory usage when the feature or instance size is large. In this paper, we propose a novel method forefficientadaptivemulti-kernel foroutlierdetection (EAMOD), which automatically learns the optimal weight for each training instance under different kernels using a non-negative function. In addition, we design a safe screening rule (SSR) for EAMOD to improve its training efficiency without any loss of accuracy. To the best of our knowledge, it is the first attempt to develop SSR for multi-kernel-based outlier detection methods. Extensive experiments show that EAMOD is effective and efficient. Xinye Wang, Lei Duan, Chengxin He, Yuanyuan Chen 0006, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Robust Multi-Kernel Nearest Neighborhood for Outlier DetectionabstractOutlier detection methods based on distance measure have been used in numerous applications due to their effectiveness and interpretability. However, distances among instances heavily depend on the feature space in which they reside. For an outlier, distances from it to the normal instances may be extremely close in one feature space, failing to separate them from each other, while this situation is reversed in another space. Meanwhile, the distance measure is sensitive to a few “marginal instances” (i.e., normal instances located very close to outliers in the feature space) during the estimation of whether a test instance is an outlier or not. In this paper, we propose a robust multi-kernel nearest neighborhood (RMKN) method for outlier detection. Specifically, in the training phase, we only consider normal instances and transform them into a Polynomial kernel function weighted digraph to capture their geometric relationships in the original feature space. Then, we develop an objective function based on the weighted digraph to find a latent feature space via multi-kernel learning such that distances among normal instances in this latent feature space are as close as possible while preserving their original distributions. In the detecting phase, we design an outlying score based on the two-stage multi-kernel k-nearest nearest neighbors to detect outliers. Extensive experiments with ten datasets show that RMKN is effective and robust Xinye Wang, Lei Duan, Zhenyang Yu, Chengxin He, Zhifeng Bao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Learning Enhanced Representations via Contrasting for Multi-view Outlier Detection
Xiaocong Chen, Xinye Wang, Lei Duan |
DASFAA (4) | 5 |
| 2023 | CHSR: Cross-view Learning from Heterogeneous Graph for Session-Based Recommendation
Junchen Wang, Lei Duan, Yidan Zhang 0001, Zhaohang Luo |
DASFAA (2) | 2 |
| 2023 | TUAF: Triple-Unit-Based Graph-Level Anomaly Detection with Adaptive Fusion Readout
Zhenyang Yu, Xinye Wang, Bingzhe Zhang, Zhaohang Luo, Lei Duan |
DASFAA (4) | 5 |
| 2023 | IFGDS: An Interactive Fraud Groups Detection System for Medicare Claims Data
Zhenyang Yu, Kaiming Zhan, Lei Duan |
DASFAA (4) | 6 |
| 2023 | Enhancing GNN-based Fraud Detector via Semantic Extraction and Max-Representation-MarginabstractFraud detection aims to identify fraudsters from normal users. In graph environments, both fraudsters and normal users are modeled as nodes, while edges represent the connections between them. However, fraudulent nodes in the real world often camouflage themselves by establishing numerous fake connections with normal nodes, making them challenging to be identified. Existing fraud detection methods struggle to address this issue, they utilize graph neural networks to aggregate normal informations from normal neighbors, which leads to the smoothing of the fraudulent information. Furthermore, these methods exhibit poor generalization performance as they are unable to detect new fraudsters which not present in the training process. To overcome these limitations, this paper proposes GFAN, a novel model based on Graph Feature enhAncement Network. Specifically, GFAN introduces a specific semantic extraction module to screen and delete fake connections by evaluating the confidence level of edge presence. Additionally, GFAN provides a representation enhanced co-training module that highlights camouflaged fraudulent representations by training the small sphere and large margin support vector data description. Experimental results show that GFAN outperforms other competitive graph-based fraud detectors on public datasets. The GFAN code is available at: https://github.com/scu-kdde/OAM-GFAN-2023. Bingzhe Zhang, Xinye Wang, Zhenyang Yu, Yuanhao Zhang, Chengxin He, Song Deng, Zhaohang Luo, Lei Duan |
ICDM | 8 |
| 2023 | Memory-Enhanced Transformer for Representation Learning on Temporal Heterogeneous GraphsabstractAbstract Temporal heterogeneous graphs can model lots of complex systems in the real world, such as social networks and e-commerce applications, which are naturally time-varying and heterogeneous. As most existing graph representation learning methods cannot efficiently handle both of these characteristics, we propose a Transformer-like representation learning model, named THAN, to learn low-dimensional node embeddings preserving the topological structure features, heterogeneous semantics, and dynamic patterns of temporal heterogeneous graphs, simultaneously. Specifically, THAN first samples heterogeneous neighbors with temporal constraints and projects node features into the same vector space, then encodes time information and aggregates the neighborhood influence in different weights via type-aware self-attention. To capture long-term dependencies and evolutionary patterns, we design an optional memory module for storing and evolving dynamic node representations. Experiments on three real-world datasets demonstrate that THAN outperforms the state-of-the-arts in terms of effectiveness with respect to the temporal link prediction task. Longhai Li, Lei Duan, Junchen Wang, Chengxin He, Guicai Xie, Song Deng, Zhaohang Luo |
Data Sci. Eng. | 2 |
| 2023 | Dynamic Ridesharing With Minimal Regret: Towards an Enhanced Engagement Among Three StakeholdersabstractIn dynamic ridesharing, the platform serves as the mediator by tailoring the assignment result between workers and riders with a focus on a certain objective. Existing studies generally focus on either one or two stakeholders when modelling the problem while the wellbeing of the other parties may be ignored or even undermined. For example, purely maximizing the total revenue of the ridesharing platform may cause the loss of riders and in turn lead to a low served rate, because those expensive orders will be processed in priority. In this paper, we for the first time study how to incorporate the willingness of all stakeholders (i.e., the platform, workers and riders). Given a set of workers and a set of rider requests, we aim to return the matchable worker-rider pairs in order to minimize theregret. Specifically, two types of regret are defined: (i) theserved rate regret, which refers to the rate of unserved requests, catering for the reputation and profit of the platform and workers; (ii) therevenue regret, which considers the portion of revenue loss from unserved riders, catering for the focus of workers and riders in the trip schedule. We prove the NP-hardness of this problem. To tackle this problem, we first propose a dynamic programming insertion algorithm to improve the efficiency of inserting a rider request into a trip schedule of a worker. Furthermore, two kinds of heuristic algorithms are devised to match rider requests with workers effectively. Comprehensive experiments on two real-world datasets verify the effectiveness, efficiency and scalability of our solutions in dealing with different supply-demand relationships in practice. Tingting Wang 0009, Hui Luo 0001, Zhifeng Bao, Lei Duan |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | An Efficient Method for Outlying Aspect Mining Based on Genetic Algorithm
Lei Duan, Xinye Wang |
ADMA (1) | 2 |
| 2022 | Effective Mining of Contrast Hybrid Patterns from Nominal-numerical Mixed Data
Lei Duan, Zhenyang Yu |
ADMA (1) | 2 |
| 2022 | Improving Text-based Similar Product Recommendation for Dynamic Product Advertising at YahooabstractRetrieving similar products is a critical functionality required by many e-commerce websites as well as dynamic product advertising systems. Retargeting and Prospecting are two major forms of dynamic product advertising. Typically, after a user interacts with a product on an advertiser website (e.g., Macy's), when the user later visits a website (e.g., yahoo.com) supported by a dynamic product advertising system, the same product may be shown to the user as a Retargeting product ad, while some similar products may be shown to the user as Prospecting product ads on the web page. Similar products can enrich users' ad experience based on users' intent on the Prospecting product ads through which the users interacted. These product ads can also serve as substitutes when Retargeting ad candidates are out of stock. However, it is challenging to retrieve similar products among billions of products in a product catalog efficiently. Deep Siamese models allow efficient retrieval but do not put enough emphasize on key product attributes. To improve the quality of the similar products, we propose to first use a Siamese Transformer-based model to retrieve similar products and then refine them with the attribute "product name" that indicates the type of a product (e.g., running shoes, engagement ring, etc.) for post filtering. We propose a novel product name generation model that fine tunes a pre-trained Transformer-based language model with a sequence to sequence objective. To the best of our knowledge, this is the first work using a generative approach for identifying product attributes. We introduce two applications of the proposed approach for the dynamic product advertising system of Yahoo for Retargeting and Prospecting respectively. Offline evaluation and online A/B testing shows that the proposed approach retrieves high quality similar products, leading to an increase of ad clicks and ad revenue. Xiao Bai 0002, Lei Duan, Richard Tang, Gaurav Batra, Ritesh Agrawal |
CIKM | 2 |
| 2022 | MORN: Molecular Property Prediction Based on Textual-Topological-Spatial Multi-View LearningabstractPredicting molecular properties has significant implications for the discovery and generation of drugs and further research in the domain of medicinal chemistry. Learning representations of molecules plays a central role in deep learning-driven property prediction. However, the diversity of molecular features (e.g., chemical system languages, structure notations) brings inconsistency in molecular representation. Moreover, the scarcity of labeled molecular data limits the accuracy of the molecular property prediction model. To address the above issues, we proposed a two-stage method, named MORN, for learning molecular representations for molecular property prediction from a multi-view perspective. In the first stage, textual-topological-spatial multi-views were proposed to learn the molecular representations, so as to capture both chemical system language and structure notation features simultaneously. In the second stage, an adaptive strategy was used to fuse molecular representations learned from multi-views to predict molecular properties. To alleviate the limitation of the scarcity of labeled molecular data, the label restriction was introduced in both multi-view representation learning and fusion stages. The performance of MORN was assessed by seven benchmark molecular datasets and one self-built molecular dataset. Experimental results demonstrated that MORN is effective in molecular property prediction. Yidan Zhang 0001, Xinye Wang, Zhenyang Yu, Lei Duan |
CIKM | 5 |
| 2022 | A Trace Ratio Maximization Method for Parameter Free Multiple Kernel Clustering
Yan Chen 0036, Liang Du 0003, Lei Duan |
DASFAA (2) | 4 |
| 2022 | AdCSE: An Adversarial Method for Contrastive Learning of Sentence Embeddings
Renhao Li, Lei Duan, Guicai Xie, Shan Xiao |
DASFAA (3) | 2 |
| 2021 | HMNet: Hybrid Matching Network for Few-Shot Link Prediction
Shan Xiao, Lei Duan, Guicai Xie, Renhao Li, Geng Deng, Jyrki Nummenmaa |
DASFAA (1) | 2 |
| 2021 | Efficient Mining of Outlying Sequential Behavior Patterns
Lei Duan, Guicai Xie, Longhai Li, Jyrki Nummenmaa |
DASFAA (2) | 2 |
| 2020 | EvsJSON: An Efficient Validator for Split JSON Documents
Bangjun He, Jie Zuo, Qiaoyan Feng, Guicai Xie, Ruiqi Qin 0001, Lei Duan |
DASFAA (3) | 7 |
| 2020 | Efficient Mining of Outlying Sequence Patterns for Analyzing Outlierness of Sequence DataabstractRecently, a lot of research work has been proposed in different domains to detect outliers and analyze the outlierness of outliers for relational data. However, while sequence data is ubiquitous in real life, analyzing the outlierness for sequence data has not received enough attention. In this article, we study the problem of mining outlying sequence patterns in sequence data addressing the question: given a query sequence s in a sequence dataset D , the objective is to discover sequence patterns that will indicate the most unusualness (i.e., outlierness) of s compared against other sequences. Technically, we use the rank defined by the average probabilistic strength ( aps ) of a sequence pattern in a sequence to measure the outlierness of the sequence. Then a minimal sequence pattern where the query sequence is ranked the highest is defined as an outlying sequence pattern. To address the above problem, we present OSPMiner, a heuristic method that computes aps by incorporating several pruning techniques. Our empirical study using both real and synthetic data demonstrates that OSPMiner is effective and efficient. Tingting Wang 0009, Lei Duan, Guozhu Dong, Zhifeng Bao |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | Discovering Relationship Patterns Among Associated Temporal Event Sequences
Lei Duan, Ruiqi Qin 0001, Jyrki Nummenmaa |
DASFAA (1) | 2 |
| 2018 | A Player Behavior Model for Predicting Win-Loss Outcome in MOBA Games
Xuan Lan, Lei Duan, Ruiqi Qin 0001, Timo Nummenmaa, Jyrki Nummenmaa |
ADMA | 2 |
| 2018 | Bus-OLAP: A Data Management Model for Non-on-Time Events Query Over Bus Journey DataabstractIncreasing the on-time rate of bus service can prompt the people’s willingness to travel by bus, which is an effective measure to mitigate the city traffic congestion. Performing queries on the bus arrival can be used to identify and analyze various kinds of non-on-time events that happened during the bus journey, which is helpful for detecting the factors of delaying events, and providing decision support for optimizing the bus schedules. We propose a data management model, called Bus-OLAP, for querying bus journey data, considering the characteristics of bus running and the scenarios of non-on-time analysis. While fulfilling typical requirements of bus journey data queries, Bus-OLAP not only provides a flexible way to manage the data and to implement multiple granularity data query and update, but it also supports distributed queries and computation. The experiments on real-world bus journey data verify that Bus-OLAP is effective and efficient. Lei Duan, Tinghai Pang, Jyrki Nummenmaa, Jie Zuo, Changjie Tang |
Data Sci. Eng. | 1 |
| 2017 | Mining Top-k Distinguishing Temporal Sequential Patterns from Event Sequences
Lei Duan, Guozhu Dong, Jyrki Nummenmaa |
DASFAA (2) | 1 |
| 2016 | Mining Distinguishing Customer Focus Sets for Online Shopping Decision Support
Lei Duan, Jyrki Nummenmaa, Guozhu Dong, Pan Qin |
ADMA | 2 |
| 2016 | Mining Top-k Distinguishing Sequential Patterns with Flexible Gap Constraints
Lei Duan, Guozhu Dong, Haiqing Zhang, Changjie Tang |
WAIM (1) | 2 |
| 2016 | Efficient discovery of contrast subspaces for object explanation and characterization
Lei Duan, Guanting Tang, Jian Pei 0001, James Bailey 0001, Guozhu Dong, Xuan Vinh Nguyen, Akiko Campbell, Changjie Tang |
Knowl. Inf. Syst. | 1 |
| 2015 | Mining Itemset-based Distinguishing Sequential Patterns with Gap Constraint
Lei Duan, Guozhu Dong, Jyrki Nummenmaa, Changjie Tang |
DASFAA (1) | 2 |
| 2015 | Mining outlying aspects on numeric data
Lei Duan, Guanting Tang, Jian Pei 0001, James Bailey 0001, Akiko Campbell, Changjie Tang |
Data Min. Knowl. Discov. | 1 |
| 2014 | Mining Frequent Closed Sequential Patterns with Non-user-defined Gap Constraints
Lei Duan, Jyrki Nummenmaa, Song Deng, Zhong-Qi Li, Changjie Tang |
ADMA | 2 |
| 2014 | Efficient Mining of Density-Aware Distinguishing Sequential Patterns with Gap Constraints
Xianming Wang, Lei Duan, Guozhu Dong, Zhonghua Yu, Changjie Tang |
DASFAA (1) | 2 |
| 2014 | YouRank: Let User Engagement Rank Microblog Search Results
Wenbo Wang 0002, Lei Duan, Anirudh Koul, Amit P. Sheth |
ICWSM | 2 |
| 2014 | Mining Contrast Subspaces
Lei Duan, Guanting Tang, Jian Pei 0001, James Bailey 0001, Guozhu Dong, Akiko Campbell, Changjie Tang |
PAKDD (1) | 1 |
| 2014 | Improving search relevance for short queries in community question answeringabstractRelevant question retrieval and ranking is a typical task in community question answering (CQA). Existing methods mainly focus on long and syntactically structured queries. However, when an input query is short, the task becomes challenging, due to a lack information regarding user intent. In this paper, we mine different types of user intent from various sources for short queries. With these intent signals, we propose a new intent-based language model. The model takes advantage of both state-of-the-art relevance models and the extra intent information mined from multiple sources. We further employ a state-of-the-art learning-to-rank approach to estimate parameters in the model from training data. Experiments show that by leveraging user intent prediction, our model significantly outperforms the state-of-the-art relevance models in question search. Haocheng Wu, Wei Wu 0014, Ming Zhou 0001, Enhong Chen, Lei Duan, Harry Shum |
WSDM | 5 |
| 2013 | Mining effective multi-segment sliding window for pathogen incidence rate prediction
Lei Duan, Changjie Tang, Guozhu Dong, Xianming Wang, Jie Zuo, Zhong-Qi Li |
Data Knowl. Eng. | 1 |
| 2011 | Mining Good Sliding Window for Positive Pathogens Prediction in Pathogenic Spectrum Analysis
Lei Duan, Changjie Tang, Chi Gou, Jie Zuo |
ADMA (2) | 1 |
| 2010 | Temporal query log profiling to improve web search rankingabstractTemporal information can be leveraged and incorporated to improve web search ranking. In this work, we propose a method to improve the ranking of search results by identifying the fundamental properties of temporal behavior of low-quality hosts and spam-prone queries in search logs and modeling those properties as quantifiable features. In particular, we introduce the concepts of host churn, a measure of changes in host visibility for user queries, and query volatility, a measure of semantic instability of query results, and propose the methods for construction of temporal profiles from search query logs that can be used for estimation of a set of features based on the introduced concepts. The utility of the proposed concepts has been experimentally demonstrated for two language-independent search tasks: the regression-based ranking of search results and a novel classification problem of detecting spam-prone queries introduced in this work. Alexander Kotov 0001, Pranam Kolari, Lei Duan, Yi Chang 0001 |
CIKM | 3 |
| 2010 | Mining Contrast Inequalities in Numeric Dataset
Lei Duan, Jie Zuo, Tianqing Zhang, Jing Peng 0002 |
WAIM | 1 |
| 2010 | An Efficient Approach for Mining Segment-Wise Intervention Rules in Time-Series Streams
Yue Wang 0014, Jie Zuo, Ning Yang 0001, Lei Duan |
WAIM | 4 |
| 2009 | Mining Class Contrast Functions by Gene Expression Programming
Lei Duan, Changjie Tang, Tianqing Zhang, Jie Zuo |
ADMA | 1 |
| 2009 | Improving web page classification by label-propagation over click graphsabstractIn this paper, we present a semi-supervised learning method for web page classification, leveraging click logs to augment training data by propagating class labels to unlabeled similar documents. Current state-of-the-art classifiers are supervised and require large amounts of manually labeled data. We hypothesize that unlabeled documents similar to our positive and negative labeled documents tend to be clicked through by the same user queries. Our proposed method leverages this hypothesis and augments our training set by modeling the similarity between documents in a click graph. We experiment with three different web page classifiers and show empirical evidence that our proposed approach outperforms state-of-the-art methods and reduces the amount of human effort to label training data. Patrick Pantel, Lei Duan, Scott Gaffney |
CIKM | 3 |
| 2009 | Threshold selection for web-page classification with highly skewed class distributionabstractWe propose a novel cost-efficient approach to threshold selection for binary web-page classification problems with imbalanced class distributions. In many binary-classification tasks the distribution of classes is highly skewed. In such problems, using uniform random sampling in constructing sample sets for threshold setting requires large sample sizes in order to include a statistically sufficient number of examples of the minority class. On the other hand, manually labeling examples is expensive and budgetary considerations require that the size of sample sets be limited. These conflicting requirements make threshold selection a challenging problem. Our method of sample-set construction is a novel approach based on stratified sampling, in which manually labeled examples are expanded to reflect the true class distribution of the web-page population. Our experimental results show that using false positive rate as the criterion for threshold setting results in lower-variance threshold estimates than using other widely used accuracy measures such as F1 and precision. Lei Duan, Yiping Zhou, Byron Dom |
WWW | 2 |
| 2007 | A Coding Hierarchy Computing Based Clustering Algorithm
Jing Peng 0002, Changjie Tang, Dongqing Yang, An-long Chen, Lei Duan |
ADMA | 5 |
| 2006 | Distance Guided Classification with Gene Expression Programming
Lei Duan, Changjie Tang, Tianqing Zhang, Dagang Wei |
ADMA | 1 |