VLDB 2026 Research / reviewers in the wild / expert
Ziwei Chai
dblp:325/1758
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphLLM: Boosting Graph Reasoning Ability of Large Language ModelabstractThe advancement of Large Language Models (LLMs) has remarkably pushed the boundaries towards artificial general intelligence (AGI), with their exceptional ability on understanding diverse types of information, including but not limited to images and audio. Despite this progress, a critical gap remains in empowering LLMs to proficiently understand and reason on graph data, which is ubiquitous in Big Data applications such as social networks, knowledge graphs, and molecular databases. Recent studies underscore LLMs' underwhelming performance on fundamental graph reasoning tasks. In this paper, we endeavor to unearth the obstacles that impede LLMs in graph reasoning, pinpointing the common practice of converting graphs into natural language descriptions (Graph2Text) as a fundamental bot tleneck. To overcome this impediment, we introduce GraphLLM, a pioneering end-to-end approach that synergistically integrates graph learning models with LLMs through a novel Dynamic Task Configuration System. This system employs a Hierarchical Graph Processing Pipeline that combines Local Structure Analyzers for node-level features with Global Pattern Synthesizers for graph level understanding, enabling scalable processing of large-scale graph data. Our empirical evaluations across four fundamental graph reasoning tasks validate the effectiveness of GraphLLM. The results exhibit a substantial average accuracy enhancement of 54.44%, alongside a noteworthy context reduction of 96.45% across various graph reasoning tasks, demonstrating significant potential for Big Data graph analytics. Ziwei Chai, Tianjie Zhang, Kaiqiao Han, Xiaohai Hu, Xuanwen Huang, Yang Yang 0009 |
IEEE Trans. Big Data | 1 |
| 2025 | Cypher-RI: Reinforcement Learning for Integrating Schema Selection into Cypher GenerationabstractThe increasing utilization of graph databases across various fields stems from their capacity to represent intricate interconnections. Nonetheless, exploiting the full capabilities of graph databases continues to be a significant hurdle, largely because of the inherent difficulty in translating natural language into Cypher. Recognizing the critical role of schema selection in database query generation and drawing inspiration from recent progress in reasoning-augmented approaches trained through reinforcement learning to enhance inference capabilities and generalization, we introduce Cypher-RI, a specialized framework for the Text-to-Cypher task. Distinct from conventional approaches, our methodology seamlessly integrates schema selection within the Cypher generation pipeline, conceptualizing it as a critical element in the reasoning process. The schema selection mechanism is guided by textual context, with its outcomes recursively shaping subsequent inference processes. Impressively, our 7B-parameter model, trained through this RL paradigm, demonstrates superior performance compared to baselines, exhibiting a 9.41\% accuracy improvement over GPT-4o on CypherBench. These results underscore the effectiveness of our proposed reinforcement learning framework, which integrates schema selection to enhance both the accuracy and reasoning capabilities in Text-to-Cypher tasks. Hanchen Su, Xuyuan Li, Zhuoyi Lu, Ziwei Chai, Haozheng Wang |
NeurIPS | 5 |
| 2025 | Enhancing Cross-domain Link Prediction via Evolution Process ModelingabstractThis paper proposes CrossLink, a novel framework for cross-domain link prediction. CrossLink learns the evolution pattern of a specific downstream graph and subsequently makes pattern-specific link predictions. It employs a technique called conditioned link generation, which integrates both evolution and structure modeling to perform evolution-specific link prediction. This conditioned link generation is carried out by a transformer-decoder architecture, enabling efficient parallel training and inference. CrossLink is trained on extensive dynamic graphs across diverse domains, encompassing 6 million dynamic edges. Extensive experiments on eight untrained graphs demonstrate that CrossLink achieves state-of-the-art performance in cross-domain link prediction. Compared to advanced baselines under the same settings, CrossLink shows an average improvement of 11.40% in Average Precision across eight graphs. Impressively, it surpasses the fully supervised performance of 8 advanced baselines on 6 untrained graphs. Project Page is https://zjunet.github.io/CrossLink/ Xuanwen Huang, Wei Chow, Yize Zhu, Ziwei Chai, Chunping Wang 0001, Lei Chen 0082, Yang Yang 0009 |
WWW | 5 |
| 2024 | An Expert is Worth One Token: Synergizing Multiple Expert LLMs as Generalist via Expert Token RoutingabstractZiwei Chai, Guoyin Wang, Jing Su, Tianjie Zhang, Xuanwen Huang, Xuwu Wang, Jingjing Xu, Jianbo Yuan, Hongxia Yang, Fei Wu, Yang Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ziwei Chai, Guoyin Wang 0002, Jing Su 0005, Tianjie Zhang, Xuanwen Huang, Xuwu Wang, Hongxia Yang, Fei Wu 0001, Yang Yang 0009 |
ACL (1) | 1 |
| 2024 | InfiAgent-DABench: Evaluating Agents on Data Analysis TasksabstractIn this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. Agents need to solve these tasks end-to-end by interacting with an execution environment. This benchmark contains DAEval, a dataset consisting of 603 data analysis questions derived from 124 CSV files, and an agent framework which incorporates LLMs to serve as data analysis agents for both serving and evaluating. Since data analysis questions are often open-ended and hard to evaluate without human supervision, we adopt a format-prompting technique to convert each question into a closed-form format so that they can be automatically evaluated. Our extensive benchmarking of 34 LLMs uncovers the current challenges encountered in data analysis tasks. In addition, building upon our agent framework, we develop a specialized agent, DAAgent, which surpasses GPT-3.5 by 3.9% on DABench. Evaluation datasets and toolkits for InfiAgent-DABench are released at https://github.com/InfiAgent/InfiAgent. Xueyu Hu, Ziyu Zhao 0001, Ziwei Chai, Guoyin Wang 0002, Xuwu Wang, Jing Su 0005, Jiwei Li 0001, Kun Kuang 0001, Yang Yang 0009, Hongxia Yang, Fei Wu 0001 |
ICML | 4 |
| 2024 | Can GNN be Good Adapter for LLMs?abstractRecently, large language models (LLMs) have demonstrated superior capabilities in understanding and zero-shot learning on textual data, promising significant advances for many text-related domains. In the graph domain, various real-world scenarios also involve textual data, where tasks and node features can be described by text. These text-attributed graphs (TAGs) have broad applications in social media, recommendation systems, etc. Thus, this paper explores how to utilize LLMs to model TAGs. Previous methods for TAG modeling are based on million-scale LMs. When scaled up to billion-scale LLMs, they face huge challenges in computational costs. Additionally, they also ignore the zero-shot inference capabilities of LLMs. Therefore, we propose GraphAdapter, which uses a graph neural network (GNN) as an efficient adapter in collaboration with LLMs to tackle TAGs. In terms of efficiency, the GNN adapter introduces only a few trainable parameters and can be trained with low computation costs. The entire framework is trained using auto-regression on node text (next token prediction). Once trained, GraphAdapter can be seamlessly fine-tuned with task-specific prompts for various downstream tasks. Through extensive experiments across multiple real-world TAGs, GraphAdapter based on Llama 2 gains an average improvement of approximately 5% in terms of node classification. Furthermore, GraphAdapter can also adapt to other language models, including RoBERTa, GPT-2. The promising results demonstrate that GNNs can serve as effective adapters for LLMs in TAG modeling. Xuanwen Huang, Kaiqiao Han, Yang Yang 0009, Dezheng Bao, Quanjin Tao, Ziwei Chai, Qi Zhu 0008 |
WWW | 6 |
| 2023 | Towards Learning to Discover Money Laundering Sub-network in Massive Transaction NetworkabstractAnti-money laundering (AML) systems play a critical role in safeguarding global economy. As money laundering is considered as one of the top group crimes, there is a crucial need to discover money laundering sub-network behind a particular money laundering transaction for a robust AML system. However, existing rule-based methods for money laundering sub-network discovery is heavily based on domain knowledge and may lag behind the modus operandi of launderers. Therefore, in this work, we first address the money laundering sub-network discovery problem with a neural network based approach, and propose an AML framework AMAP equipped with an adaptive sub-network proposer. In particular, we design an adaptive sub-network proposer guided by a supervised contrastive loss to discriminate money laundering transactions from massive benign transactions. We conduct extensive experiments on real-word datasets in AliPay of Ant Group. The result demonstrates the effectiveness of our AMAP in both money laundering transaction detection and money laundering sub-network discovering. The learned framework which yields money laundering sub-network from massive transaction network leads to a more comprehensive risk coverage and a deeper insight to money laundering strategies. Ziwei Chai, Yang Yang 0009, Jiawang Dan, Changhua Meng, Weiqiang Wang 0002, Yifei Sun 0002 |
AAAI | 1 |
| 2023 | Time2Graph+: Bridging Time Series and Graph Representation Learning via Multiple AttentionsabstractTime series modeling has attracted great research interests in the last decades. Among the literature, shapelet-based models aim to extract representative subsequences, and could offer explanatory insights. In order to capture the shapelet dynamics and evolutions, we propose a novel framework of bridging time series representation learning and graph modeling, with two different implementations. We first formulate the process of extracting time-aware shapelets, then briefly introduce the key idea of transforming time series data into shapelet evolution graphs, to model the shapelet evolutionary patterns. A straightforward solution is to enumerate all possible shapelet transitions among adjacent time series segments, and apply a random-walk-based graph embedding algorithm to learn the time series representations (Time2Graph). We further extend Time2Graph by adopting graph attention mechanism to refine the procedure of modeling shapelet evolutions, namely Time2Graph+. Specifically, we transform each time series data into a unique and unweighted shapelet graph, and use GAT to automatically capture the correlations between shapelets. Experimental results show the significant improvements of Time2Graph+, and extensive observational analysis demonstrate the effectiveness and interpretability brought by attentions. Furthermore, the success of online deployment of Time2Graph+ model in State Grid of China validates the whole framework in the real-world application. Ziqiang Cheng, Yang Yang 0009, Wenjie Hu 0003, Zhangchi Ying, Ziwei Chai, Chunping Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Can Abnormality be Detected by Graph Neural Networks?abstractAnomaly detection in graphs has attracted considerable interests in both academia and industry due to its wide applications in numerous domains ranging from finance to biology. Meanwhile, graph neural networks (GNNs) is emerging as a powerful tool for modeling graph data. A natural and fundamental question that arises here is: can abnormality be detected by graph neural networks? In this paper, we aim to answer this question, which is nontrivial. As many existing works have explored, graph neural networks can be seen as filters for graph signals, with the favor of low frequency in graphs. In other words, GNN will smooth the signals of adjacent nodes. However, abnormality in a graph intuitively has the characteristic that it tends to be dissimilar to its neighbors, which are mostly normal samples. It thereby conflicts with the general assumption with traditional GNNs. To solve this, we propose a novel Adaptive Multi-frequency Graph Neural Network (AMNet), aiming to capture both low-frequency and high-frequency signals, and adaptively combine signals of different frequencies. Experimental results on real-world datasets demonstrate that our model achieves a significant improvement comparing with several state-of-the-art baseline methods. Ziwei Chai, Siqi You, Yang Yang 0009, Shiliang Pu, Jiarong Xu, Haoyang Cai |
IJCAI | 1 |