VLDB 2026 Research / reviewers in the wild / expert
Caigao Jiang
dblp:292/3817
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-9383-2479ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Learning paradigms · 43% Trustworthy machine learning · 31% Language models and text generation · 14% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 50% Data models and query languages · 50% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
continual learning |
1.5 | 2 | 2025 | Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning · ICLR 2025 Prompt-augmented Temporal Point Process for Streaming Event Sequence · NeurIPS 2023 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.9 | 1 | 2025 | Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning · ICLR 2025 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
function vectors |
0.9 | 1 | 2025 | Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning · ICLR 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning · ICLR 2025 |
Mathematical optimization
solver code generation |
0.9 | 1 | 2025 | LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch · ICLR 2025 |
Data mining › sequence analysis
event sequence modeling |
0.8 | 1 | 2024 | EasyTPP: Towards Open Benchmarking Temporal Point Processes · ICLR 2024 |
Data models and query languages › natural language interface
natural language interface to database |
0.8 | 1 | 2024 | Demonstration of DB-GPT: Next Generation Data Interaction System Empowered by Large Language Models · Proc. VLDB Endow. 2024 |
Data mining › probabilistic model
temporal point process |
0.8 | 1 | 2024 | EasyTPP: Towards Open Benchmarking Temporal Point Processes · ICLR 2024 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
0.8 | 1 | 2024 | Demonstration of DB-GPT: Next Generation Data Interaction System Empowered by Large Language Models · Proc. VLDB Endow. 2024 |
Performance modeling and evaluation
benchmarking |
0.8 | 1 | 2024 | EasyTPP: Towards Open Benchmarking Temporal Point Processes · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
temporal point process |
0.7 | 1 | 2023 | Prompt-augmented Temporal Point Process for Streaming Event Sequence · NeurIPS 2023 |
Natural language and speech › Language models and text generation
instruction tuning |
0.3 | 1 | 2025 | LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model reasoning
self-correction |
0.3 | 1 | 2025 | LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 3.3self-correction · 1.7multi-instruction tuning · 1.7model alignment · 1.7neural temporal point process · 1.5multi-agent framework · 1.5regularization · 0.9instruction tuning · 0.9prompt tuning · 0.7continuous-time retrieval prompt pool · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction TuningabstractCatastrophic forgetting (CF) poses a significant challenge in machine learning, where a model forgets previously learned information upon learning new tasks.
Despite the advanced capabilities of Large Language Models (LLMs), they continue to face challenges with CF during continual learning. The majority of existing research focuses on analyzing forgetting patterns through a singular training sequence, thereby overlooking the intricate effects that diverse tasks have on model behavior.
Our study explores CF across various settings, discovering that model forgetting is influenced by both the specific training tasks and the models themselves. To this end, we interpret forgetting by examining the function vector (FV), a compact representation of functions in LLMs, offering a model-dependent indicator for the occurrence of CF. Through theoretical and empirical analyses, we demonstrated that CF in LLMs primarily stems from biases in function activation rather than the overwriting of task processing functions.
Leveraging these insights, we propose a novel function vector guided training methodology, incorporating a regularization technique to stabilize the FV and mitigate forgetting. Empirical tests on four benchmarks confirm the effectiveness of our proposed training method, substantiating our theoretical framework concerning CF and model function dynamics. Gangwei Jiang, Caigao Jiang, Siqiao Xue, Jun Zhou 0011, Linqi Song, Defu Lian, Ying Wei 0001 |
ICLR | 2 |
| 2025 | LLMOPT: Learning to Define and Solve General Optimization Problems from ScratchabstractOptimization problems are prevalent across various scenarios. Formulating and then solving optimization problems described by natural language often requires highly specialized human expertise, which could block the widespread application of optimization-based decision making. To automate problem formulation and solving, leveraging large language models (LLMs) has emerged as a potential way. However, this kind of approach suffers from the issue of optimization generalization. Namely, the accuracy of most current LLM-based methods and the generality of optimization problem types that they can model are still limited. In this paper, we propose a unified learning-based framework called LLMOPT to boost optimization generalization. Starting from the natural language descriptions of optimization problems and a pre-trained LLM, LLMOPT constructs the introduced five-element formulation as a universal model for learning to define diverse optimization problem types. Then, LLMOPT employs the multi-instruction tuning to enhance both problem formalization and solver code generation accuracy and generality. After that, to prevent hallucinations in LLMs, such as sacrificing solving accuracy to avoid execution errors, the model alignment and self-correction mechanism are adopted in LLMOPT. We evaluate the optimization generalization ability of LLMOPT and compared methods across six real-world datasets covering roughly 20 fields such as health, environment, energy and manufacturing, etc. Extensive experiment results show that LLMOPT is able to model various optimization problem types such as linear/nonlinear programming, mixed integer programming, and combinatorial optimization, and achieves a notable 11.08% average solving accuracy improvement compared with the state-of-the-art methods. The code is available at https://github.com/caigaojiang/LLMOPT. Caigao Jiang, Xiang Shu, Hong Qian, Aimin Zhou |
ICLR | 1 |
| 2024 | Enhancing Event Sequence Modeling with Contrastive Relational InferenceabstractNeural temporal point processes(TPPs) have shown promise for modeling continuous-time event sequences. However, capturing the interactions between events is challenging yet critical for performing inference tasks like forecasting on event sequence data. Existing TPP models have focused on parameterizing the conditional distribution of future events but struggle to model event interactions. In this paper, we propose a novel approach that leverages Neural Relational Inference (NRI) to learn a relation graph that infers interactions while simultaneously learning the dynamics patterns from observational data. Our approach, the Contrastive Relational Inference-based Hawkes Process (CRIHP), reasons about event interactions under a variational inference framework. It utilizes intensity-based learning to search for prototype paths to contrast relationship constraints. Extensive experiments on three real-world datasets demonstrate the effectiveness of our model in capturing event interactions for event sequence modeling tasks. Yan Wang 0002, Zhixuan Chu, Caigao Jiang, Hongyan Hao, Minjie Zhu, Xindong Cai, Qing Cui, James Y. Zhang, Siqiao Xue, Jun Zhou 0011 |
ICASSP | 4 |
| 2024 | EasyTPP: Towards Open Benchmarking Temporal Point ProcessesabstractContinuous-time event sequences play a vital role in real-world domains such as healthcare, finance, online shopping, social networks, and so on. To model such data, temporal point processes (TPPs) have emerged as the most natural and competitive models, making a significant impact in both academic and application communities. Despite the emergence of many powerful models in recent years, there hasn't been a central benchmark for these models and future research endeavors. This lack of standardization impedes researchers and practitioners from comparing methods and reproducing results, potentially slowing down progress in this field.
In this paper, we present EasyTPP, the first central repository of research assets (e.g., data, models, evaluation programs, documentations) in the area of event sequence modeling. Our EasyTPP makes several unique contributions to this area: a unified interface of using existing datasets and adding new datasets; a wide range of evaluation programs that are easy to use and extend as well as facilitate reproducible research; implementations of popular neural TPPs, together with a rich library of modules by composing which one could quickly build complex models. We will actively maintain this benchmark and welcome contributions from other researchers and practitioners.
Our benchmark will help promote reproducible research in this field, thus accelerating research progress as well as making more significant real-world impacts. The code and data are available at \url{https://github.com/ant-research/EasyTemporalPointProcess}. Siqiao Xue, Xiaoming Shi 0001, Zhixuan Chu, Yan Wang 0002, Hongyan Hao, Fan Zhou 0012, Caigao Jiang, James Y. Zhang, Qingsong Wen, Jun Zhou 0011, Hongyuan Mei |
ICLR | 7 |
| 2024 | Demonstration of DB-GPT: Next Generation Data Interaction System Empowered by Large Language ModelsabstractThe recent breakthroughs in large language models (LLMs) are positioned to transition many areas of software. In this paper, we present DB-GPT, a revolutionary and product-ready Python library that integrates LLMs into traditional data interaction tasks to enhance user experience and accessibility. DB-GPT is designed to understand data interaction tasks described by natural language and provide context-aware responses powered by LLMs, making it an indispensable tool for users ranging from novice to expert. Its system design supports deployment across local, distributed, and cloud environments. Beyond handling basic data interaction tasks like Text-to-SQL with LLMs, it can handle complex tasks like generative data analysis through a Multi-Agents framework and the Agentic Workflow Expression Language (AWEL). The Service-oriented Multi-model Management Framework (SMMF) ensures data privacy and security, enabling users to employ DB-GPT with private LLMs. Additionally, DB-GPT offers a series of product-ready features designed to enable users to integrate DB-GPT within their product environments easily. The code of DB-GPT is available at Github. Siqiao Xue, Danrui Qi, Caigao Jiang, Fangyin Cheng, Keting Chen, Ganglin Wei, Wang Zhao 0005, Fan Zhou 0012, Shaodong Liu, Hongjun Yang, Faqiang Chen |
Proc. VLDB Endow. | 3 |
| 2023 | Continual Learning in Predictive AutoscalingabstractPredictive Autoscaling is used to forecast the workloads of servers and prepare the resources in advance to ensure service level objectives (SLOs) in dynamic cloud environments. However, in practice, its prediction task often suffers from performance degradation under abnormal traffics caused by external events (such as sales promotional activities and applications' re-configurations), for which a common solution is to re-train the model with data of a long historical period, but at the expense of high computational and storage costs. To better address this problem, we propose a replay-based continual learning method, i.e., Density-based Memory Selection and Hint-based Network Learning Model (DMSHM), using only a small part of the historical log to achieve accurate predictions. First, we discover the phenomenon of sample overlap when applying replay-based continual learning in prediction tasks. In order to surmount this challenge and effectively integrate new sample distribution, we propose a density-based sample selection strategy that utilizes kernel density estimation to calculate sample density as a reference to compute sample weight, and employs weight sampling to construct a new memory set. Then we implement hint-based network learning based on hint representation to optimize the parameters. Finally, we conduct experiments on public and industrial datasets to demonstrate that our proposed method outperforms state-of-the-art continual learning methods in terms of memory capacity and prediction accuracy. Furthermore, we demonstrate remarkable practicability of DMSHM in real industrial applications. Hongyan Hao, Zhixuan Chu, Shiyi Zhu, Gangwei Jiang, Yan Wang 0002, Caigao Jiang, James Y. Zhang, Siqiao Xue, Jun Zhou 0011 |
CIKM | 6 |
| 2023 | Prompt-augmented Temporal Point Process for Streaming Event SequenceabstractNeural Temporal Point Processes (TPPs) are the prevalent paradigm for modeling continuous-time event sequences, such as user activities on the web and financial transactions. In real world applications, the event data typically comes in a streaming manner, where the distribution of the patterns may shift over time. Under the privacy and memory constraints commonly seen in real scenarios, how to continuously monitor a TPP to learn the streaming event sequence is an important yet under-investigated problem. In this work, we approach this problem by adopting Continual Learning (CL), which aims to enable a model to continuously learn a sequence of tasks without catastrophic forgetting. While CL for event sequence is less well studied, we present a simple yet effective framework, PromptTPP, by integrating the base TPP with a continuous-time retrieval prompt pool. In our proposed framework, prompts are small learnable parameters, maintained in a memory space and jointly optimized with the base TPP so that the model is properly instructed to learn event streams arriving sequentially without buffering past examples or task-specific attributes. We formalize a novel and realistic experimental setup for modeling event streams, where PromptTPP consistently sets state-of-the-art performance across two real user behavior datasets. Siqiao Xue, Yan Wang 0002, Zhixuan Chu, Xiaoming Shi 0001, Caigao Jiang, Hongyan Hao, Gangwei Jiang, Xiaoyun Feng, James Zhang, Jun Zhou 0011 |
NeurIPS | 5 |