VLDB 2026 Research / reviewers in the wild / expert
Wei Guan 0006
dblp:98/3846-6
· DBLP profile ↗
19ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-8979-6847ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEAR: LLM-Powered Sequential Recommendation via Fusion of Collaborative, Semantic, and Rating InformationabstractAs users' preferences evolve over time, personalized online services increasingly rely on sequential recommender systems to predict future interactions by modeling patterns in historical user behavior. However, existing methods for sequential recommendation (SR) face two key challenges: they struggle to simultaneously leverage collaborative, semantic, and rating information, and the use of hard labels during training provides limited supervision. In this paper, we introduce SEAR, an LLM-powered Sequential recommEndation framework via fusion of collAborative, semantic, and Rating information. The proposed deep model comprises an embedding layer and a sequence encoder. The embedding layer transforms user-item interactions into three types of embeddings: collaborative, semantic, and rating. The sequence encoder then integrates these embeddings and identifies sequential patterns to model user representations. To enhance the utilization of item semantics, we integrate a large language model (LLM) to extract LLM embeddings. These embeddings are then employed to initialize the semantic embedding layer, collaborative embedding layer, and item embeddings. To capture more nuanced user behavior patterns, we generate preference-weighted soft labels based on the next k interactions. Extensive experiments validate the effectiveness of SEAR, and ablation studies further highlight the distinct contributions of the collaborative, semantic, and rating information. Wei Guan 0006, Jian Cao 0001, Qiqi Cai, Jianqi Gao 0001, Jinyu Cai, See-Kiong Ng |
WWW | 1 |
| 2026 | DUAL: A Federated Unsupervised Anomaly Detection Framework for Collaborative Business ProcessesabstractDetecting anomalies in business processes is imperative for achieving operational success, especially as multiple participants increasingly engage in collaborative efforts to complete processes, i.e., collaborative business processes, amidst rapid economic development. However, existing business process anomaly detection approaches are typically designed for centralized training, making them impractical for collaborative business processes that require stringent privacy measures. In this paper, we introduce a feDerated Unsupervised AnomaLy detection framework for collaborative business processes, named DUAL. DUAL enables participants to collaboratively detect anomalies via a third-party coordinator, which constructs a global view from the hidden representations of participants' local sub-traces without exposing the sub-traces themselves. To further strengthen privacy protection, DUAL injects differential privacy noise into the transmitted hidden representations. To mitigate the global-local discrepancy, we introduce novel event execution errors which indicate whether the client who is supposed to execute the reconstructed event is consistent with the observed one. Extensive experiments demonstrate that DUAL effectively detects anomalies in collaborative processes while preserving privacy, achieving performance comparable to a centralized setting that requires access to participants' raw sub-traces. Wei Guan 0006, Jian Cao 0001, Haiyan Zhao 0002, Jianqi Gao 0001, Shiyou Qian |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Promoting Knowledge Base Question Answering by Directing LLMs to Generate Task-relevant Logical FormsabstractKnowledge base question answering (KBQA) refers to the system that produces answers to user queries by reasoning with a large-scale structured knowledge base. Advanced works have achieved great success either by generating logical forms (LF) or directly generating answers. Although the former typically yields better performance, these generated LF could be inaccurate, e.g., non-executable. In this regard, large language models (LLMs) have shown exciting potential for accurate generation. However, it is challenging to fine-tune LLMs to generate LF. This is because the context retrieved for prediction typically leads to an excessive number of reasoning paths. In this context, LLMs can generate numerous LF corresponding to these reasoning paths, but a few LF can result in correct answers. Thus, fine-tuning LLMs to generate answer-relevant LF would conflict with the prior knowledge of the LLMs. In this work, we propose a novel learning framework, FM-KBQA, to fine-tune LLMs using multi-task learning for KBQA. Specifically, we propose to fine-tune LLMs using an additional objective: generating the index of reasoning paths that lead to correct answers. This will direct LLMs to pay attention to answer-relevant paths among numerous reasoning paths by completing a simple task where the selected reasoning paths can be supplementary for non-executable LF. Directly generating answers can make LLMs pay attention to the answer-relevant reasoning paths, but it is much more challenging than generating the index of reasoning paths. To verify FM-KBQA's effectiveness, we conduct experiments on mainstream benchmarks, such as WebQuestionsSP (WQSP) and ComplexWebQuestions (CWQ). Extensive evaluations across two public benchmark datasets underscore the superiority of FM-KBQA over current state-of-the-art methods. Jianqi Gao 0001, Jian Cao 0001, Ranran Bu, Nengjun Zhu, Wei Guan 0006, Hang Yu 0006 |
AAAI | 5 |
| 2025 | DABL: Detecting Semantic Anomalies in Business Processes Using Large Language ModelsabstractDetecting anomalies in business processes is crucial for ensuring operational success. While many existing methods rely on statistical frequency to detect anomalies, it's important to note that infrequent behavior doesn't necessarily imply undesirability. To address this challenge, detecting anomalies from a semantic viewpoint proves to be a more effective approach. However, current semantic anomaly detection methods treat a trace (i.e., process instance) as multiple event pairs, disrupting long-distance dependencies. In this paper, we introduce DABL, a novel approach for detecting semantic anomalies in business processes using large language models (LLMs). We collect 143,137 real-world process models from various domains. By generating normal traces through the playout of these process models and simulating both ordering and exclusion anomalies, we fine-tune Llama 2 using the resulting log. Through extensive experiments, we demonstrate that DABL surpasses existing state-of-the-art semantic anomaly detection methods in terms of both generalization ability and learning of given processes. Users can directly apply DABL to detect semantic anomalies in their own datasets without the need for additional training. Furthermore, DABL offers the ability to interpret anomalies' causes in natural language, providing valuable insights into the detected anomalies. Wei Guan 0006, Jian Cao 0001, Jianqi Gao 0001, Haiyan Zhao 0002, Shiyou Qian |
AAAI | 1 |
| 2025 | Enhancing Graph-based Fraud Detection by Adversarial Confidence ReweightingabstractGraph-based fraud detection has emerged as a pivotal tool in risk management, leveraging the power of graph neural networks to enhance node representations by aggregating information from neighboring nodes. However, this aggregation process can sometimes introduce noise by incorporating neighbors from different categories, potentially diluting the central node’s representation. To tackle this issue, we introduce an innovative Adversarial Confidence Reweighting (ACR) technique designed to allocate discriminative weights to samples automatically. This approach effectively minimizes the impact of noisy neighbors on the representation of nodes. By introducing controlled adversarial perturbations to the nodes being classified, we can assess the extent of representation dilution. Furthermore, we employ an adaptive node re-weighted learning objective, which dynamically adjusts node weights based on a confidence measure derived from prediction accuracy. Our experimental evaluations across three public datasets demonstrate that our adversarial strategy significantly surpasses the baseline model in terms of detection performance. Jianqi Gao 0001, Jian Cao 0001, Shiyou Qian, Wei Guan 0006 |
ICASSP | 4 |
| 2025 | FeadSeq: A Personalized Federated Anomaly Detection Framework for Discrete Event SequencesabstractEvent sequence anomaly detection has garnered considerable attention in research, encompassing applications such as identifying anomalies in system logs, anomalous transaction users, and so on. Yet, prevailing anomaly detection methods often rely solely on local data for training, potentially leading to imperfect detection performance. In this article, we introduce a personalized Federated anomaly detection framework for discrete event Sequences, named FeadSeq. Specifically, we propose a separate architecture for sequence reconstruction networks (SEPRE) which partitions the network into two parts: a shared part and a standalone part, better suited for federated learning schemes. In tandem, we propose a novel partial shared federated learning scheme that employs a mask strategy to alleviate communication overhead and produce personalized local models to address the statistical heterogeneity of data among clients. This scheme dictates that a subset of weights is communicated between clients and servers for collaborative training, while the remaining weights are trained exclusively locally. To evaluate the effectiveness of FeadSeq, we conduct extensive experiments on both system logs and business process event logs. The results affirm the superiority of FeadSeq over existing personalized federated learning algorithms, showcasing not only improved performance but also reduced communication overhead. Wei Guan 0006, Jian Cao 0001, Haiyan Zhao 0002, Yang Gu 0002, Shiyou Qian |
ACM Trans. Knowl. Discov. Data | 1 |
| 2025 | Survey and Benchmark of Anomaly Detection in Business ProcessesabstractEffective management of business processes is crucial for organizational success. However, despite meticulous design and implementation, anomalies are inevitable and can result in inefficiencies, delays, or even significant financial losses. Numerous methods for detecting anomalies in business processes have been proposed recently. However, there is no comprehensive benchmark to evaluate these methods. Consequently, the relative merits of each method remain unclear due to differences in their experimental setup, choice of datasets and evaluation measures. In this paper, we present a systematic literature review and taxonomy of business process anomaly detection methods. Additionally, we select at least one method from each category, resulting in 16 methods that are cross-benchmarked against 32 synthetic logs and 19 real-life logs from different industry domains. Our analysis provides insights into the strengths and weaknesses of different anomaly detection methods. Ultimately, our findings can help researchers and practitioners in the field of process mining make informed decisions when selecting and applying anomaly detection methods to real-life business scenarios. Finally, some future directions are discussed in order to promote the evolution of business process anomaly detection. Wei Guan 0006, Jian Cao 0001, Haiyan Zhao 0002, Yang Gu 0002, Shiyou Qian |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | ProSwats: A Proxy-based Scientific Workflow Retrieval Approach by Bridging the Gap between Textual and Structural SemanticsabstractIt is time-consuming and knowledge-intensive for scientists to find practical workflows from the massive number of scientific workflow models. Currently, the retrieval approaches are mainly based on text matching between natural language queries and the descriptions of candidate workflows. Notably, the workflow structure also provides essential semantics, but the challenge lies in effectively matching these two pieces of heterogeneous information. To address this issue, we propose a Proxy-based Scientific workflow retrieval approach, ProSwats, which selects a workflow as the Proxy for each text query to bridge the gap between textual and structural semantics. ProSwats consists of two stages: workflow pre-selection based on text similarity and workflow ranking based on a matching degree prediction model. The textual and structural features are integrated by the proxy in this model, which is used to predict and rank the degree of semantic matching between the user query and candidate workflows. Crucially, ProSwats incorporates a confidence-aware learning mechanism to adapt to the varying reliability of proxies, enhancing generalizability. Extensive experimental results on two real-world datasets demonstrate that ProSwats outperforms state-of-the-art methods with statistical significance. Yang Gu 0002, Jian Cao 0001, Shiyou Qian, Nengjun Zhu, Wei Guan 0006 |
ICWS | 5 |
| 2024 | COMB: Interconnected Transformers-Based Autoencoder for Multi-Perspective Business Process Anomaly DetectionabstractIn business processes, anomalies are prevalent, arising from diverse factors, such as software malfunctions and operator errors. Detecting these anomalies is imperative, as it significantly influences not only the financial well-being of a business but also the dependability of event logs for subsequent analysis. However, existing deep business process anomaly detection approaches either lack effective temporal dependency modeling between events or encounter challenges associated with gradient vanishing. In this paper, we introduce COMB, which stands for an interConnected transfOrmers-based autoencoder for Multi-perspective Business process anomaly detection. Considering the interdependent nature of multi-perspectives, COMB leverages multiple parallel and interconnected transformers, facilitated by the inclusion of aggregation layers. These layers serve as integration points for information from various perspectives. Notably, considering how control flow influences other perspectives, COMB incorporates innovative mask adapters to enhance its detection performance. Furthermore, we propose a novel method for calculating anomaly scores, which effectively mitigates the influence of varying numbers of potential attribute values. Our extensive experimental evaluation encompasses both synthetic and real-life logs, and the results clearly demonstrate that COMB outperforms state-of-the-art methods in both trace-level and attribute-level anomaly detection. Wei Guan 0006, Jian Cao 0001, Yan Yao 0001, Yang Gu 0002, Shiyou Qian |
ICWS | 1 |
| 2024 | MANSOR: A module alignment method based on neighbor information for scientific workflowabstractSummary Finding similar scientific workflow modules that can substitute essential components of the privacy workflows from public repositories to create a personalized workflow is growing in popularity. This is a cost‐effective and error‐free strategy for the scientific community. Currently, module alignment approaches heavily depend on syntactic information such as modules' names and types. However, the contextual semantic information of modules, which encompasses inputs, outputs, and datalinks connecting modules, can also convey their functions. Unfortunately, this information is scarcely utilized during the module alignment process. In this work, we propose a module alignment method based on neighbor information for scientific workflow (MANSOR). Specifically, we present a rule‐based attribute similarity computation approach for calculating initial module‐pair similarity in two workflows. The relation similarity is then employed to iteratively fine‐tune the matching degree with uncertainty of module pairs by considering the contextual semantics until reaching a steady state. After pairwise comparison of workflows, the module alignment results are determined based on their module‐pair similarity. Experimental results on the human‐curated corpus of ratings indicate that MANSOR outperforms the existing state‐of‐the‐art approaches with statistical significance. Yang Gu 0002, Jian Cao 0001, Shiyou Qian, Nengjun Zhu, Wei Guan 0006 |
Concurr. Comput. Pract. Exp. | 5 |
| 2024 | GAMA: A multi-graph-based anomaly detection framework for business processes via graph neural networks
Wei Guan 0006, Jian Cao 0001, Yang Gu 0002, Shiyou Qian |
Inf. Syst. | 1 |
| 2024 | WAKE: A Weakly Supervised Business Process Anomaly Detection Framework via a Pre-Trained AutoencoderabstractThe ability to detect anomalies in business processes is crucial for achieving success in business operations. While unsupervised anomaly detection approaches have gained popularity in recent years due to their label-free nature, in some cases, a limited number of labelled anomalies can be provided and using them can improve the performance of anomaly detection. To address this issue, we propose a novel framework for anomaly detection that uses a pre-trained autoencoder to extract feature representations of traces. An anomaly score generator based on a multi-layer perceptron is utilized to evaluate the extracted features. The entire framework is trained using a joint loss that ensures the generated anomaly scores satisfy a specific distribution without compromising the autoencoder's ability to reconstruct normal traces. The feature encoder is fine-tuned to provide insights into the cause of anomalies. Additionally, we design a novel technique for calculating anomaly scores to mitigate the effects of varying numbers of potential attribute values. We conduct extensive experiments on both synthetic and real-life logs, and our results demonstrate that our proposed method, WAKE, outperforms state-of-the-art unsupervised deep business process anomaly detection methods by a significant margin. Additionally, it outperforms other weakly supervised anomaly detection methods as well Wei Guan 0006, Jian Cao 0001, Haiyan Zhao 0002, Yang Gu 0002, Shiyou Qian |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Remaining Time Prediction for Collaborative Business Processes with Privacy Preservation
Jian Cao 0001, Wei Guan 0006, Shiyou Qian, Haiyan Zhao 0002 |
ICSOC (2) | 3 |
| 2023 | Plan, Generate and Match: Scientific Workflow Recommendation with Large Language Models
Yang Gu 0002, Jian Cao 0001, Shiyou Qian, Wei Guan 0006 |
ICSOC (1) | 5 |
| 2023 | SWARM: A Scientific Workflow Fragments Recommendation Approach via Contrastive Learning and Semantic Matching
Yang Gu 0002, Jian Cao 0001, Jinghua Tang, Shiyou Qian, Wei Guan 0006 |
ICSOC (2) | 5 |
| 2023 | AIMED: An automatic and incremental approach for business process model repair under concept drift
Wei Guan 0006, Jian Cao 0001, Yang Gu 0002, Shiyou Qian |
Inf. Syst. | 1 |
| 2023 | When to Invoke a Prediction Service for Business Process Monitoring?abstractPredictive monitoring of business processes aims at forecasting the future information of a business process and has gained increasing attention in recent years. With the development of cloud computing, prediction models, including remaining time prediction models, can be provided as cloud services. Reducing the number of invocation times of a remaining time prediction service is necessary since a large quantity of business process instances may be initiated and monitored every day. However, most of the current research focuses on designing new algorithms to improve prediction accuracy but there is no approach available to decide when to invoke a prediction service for each business process instance. In this article, we propose a deep reinforcement learning based strategy that can learn the policies of selecting prediction points for the remaining time prediction. Specifically, the learned policies can dynamically decide the next prediction point at which the remaining time prediction service will be invoked for a business process instance. We performed extensive experiments on five real-world datasets. The experiment results show this strategy can reduce the number of prediction points significantly and still maintain the high prediction accuracy. Jian Cao 0001, Naixuan Wang, Shiyou Qian, Wei Guan 0006 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | SWORTS: A Scientific Workflow Retrieval Approach by Learning Textual and Structural SemanticsabstractFinding scientific workflow models that can be reused or repurposed from public repositories is becoming popular in the scientific community. Currently, the retrieval approaches for workflow models are mainly based on text matching between queries and the descriptions of models. However, the structure information of these models, which includes inputs, outputs, data processing modules, and datalinks connecting modules, expresses their functions in a more detailed manner, yet this information is not used in the model retrieval process. Therefore, we propose a two-stage framework for Scientific WOrkflow Retrieval by learning Textual and Structural semantics (SWORTS). The framework comprises a workflow pre-selection step and a workflow ranking step. Specifically, we use text similarity approaches to quickly identify candidate workflow models in the first step. Then, a hierarchy-based matching degree prediction model, which takes both textual and structural features into account, is trained to predict and rank the degree of semantic matching between requirement specifications and candidate workflows. The experiment results on real-world datasets demonstrate that SWORTS achieves the best performance with relatively balanced effectiveness and efficiency among state-of-the-art methods. Yang Gu 0002, Jian Cao 0001, Shiyou Qian, Wei Guan 0006 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | GRASPED: A GRU-AE Network Based Multi-Perspective Business Process Anomaly Detection ModelabstractIn process-aware information systems (PAISs), anomalies are ubiquitous, having a number of different underlying causes, such as software malfunctions or operator errors. The presence of anomalies not only has an enormous impact on the economic well-being of the business, but also interferes with our ability to mine useful information from event logs. In this article, we propose GRASPED, aGRU-AE Network based multi-perSpective business ProcEss anomalyDetection model. GRASPED can detect anomalies not only from the control flow but also from the data perspective of a business process. GRASPED is based on an autoencoder (AE) with gated recurrent units (GRU) which is trained in an unsupervised fashion (i.e., does not require any labeling of the data). In addition, GRASPED introduces the teacher forcing method as well as the attention mechanism to improve its detection performance. GRASPED does not require training on a clean log (i.e., it can be trained and perform anomaly detection directly on logs containing anomalies). We conduct extensive experiments on synthetic logs as well as real-life logs. The experiment results show that GRASPED outperforms the state-of-the-art methods for both trace-level and attribute-level anomaly detection. Wei Guan 0006, Jian Cao 0001, Yang Gu 0002, Shiyou Qian |
IEEE Trans. Serv. Comput. | 1 |