EDBT 2026 Demo / reviewers in the wild / expert
Defu Cao
dblp:274/1535
· DBLP profile ↗
21ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0003-0240-3818ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 15 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large ModelsabstractIn recent years, there has been increasing attention on the capabilities of large-scale models, particularly in handling complex tasks that small-scale models are unable to perform. Notably, large language models (LLMs) have demonstrated ``intelligent'' abilities such as complex reasoning and abstract language comprehension, reflecting cognitive-like behaviors. However, current research on emergent abilities in large models predominantly focuses on the relationship between model performance and size, leaving a significant gap in the systematic quantitative analysis of the internal structures and mechanisms driving these emergent abilities. Drawing inspiration from neuroscience research on brain network structure and self-organization, we propose (i) a general network representation of large models, (ii) a new analytical framework — *Neuron-based Multifractal Analysis (NeuroMFA)* - for structural analysis, and (iii) a novel structure-based metric as a proxy for emergent abilities of large models. By linking structural features to the capabilities of large models, *NeuroMFA* provides a quantitative framework for analyzing emergent phenomena in large models. Our experiments show that the proposed method yields a comprehensive measure of the network's evolving heterogeneity and organization, offering theoretical foundations and a new perspective for investigating emergence in large models. Xiongye Xiao, Heng Ping, Defu Cao, Yaxing Li, Yizhuo Zhou, Nikos Kanakaris, Paul Bogdan |
ICLR | 4 |
| 2025 | MultiverseAD: Enhancing spatial-temporal synchronous attention networks with causal knowledge for multivariate time series anomaly detection
Xudong Jia 0002, Defu Cao, Niangxi Zhuang, Wei Peng 0005, Baokang Zhao, Peng Xun, Chiran Shen |
Neural Networks | 2 |
| 2024 | GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series ForecastingabstractTime series forecasting is an essential area of machine learning with a wide range of real-world applications. Most of the previous forecasting models aim to capture dynamic characteristics from uni-modal numerical historical data. Although extra knowledge can boost the time series forecasting performance, it is hard to collect such information. In addition, how to fuse the multimodal information is non-trivial. In this paper, we first propose a general principle of collecting the corresponding textual information from different data sources with the help of modern large language models (LLM). Then, we propose a prompt-based LLM framework to utilize both the numerical data and the textual information simultaneously, named GPT4MTS. In practice, we propose a GDELT-based multimodal time series dataset for news impact forecasting, which provides a concise and well-structured version of time series dataset with textual information for further research in communication. Through extensive experiments, we demonstrate the effectiveness of our proposed method on forecasting tasks with extra-textual information. Furong Jia 0002, Yixiang Zheng, Defu Cao, Yan Liu 0002 |
AAAI | 4 |
| 2024 | TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series ForecastingabstractThe past decade has witnessed significant advances in time series modeling with deep learning. While achieving state-of-the-art results, the best-performing architectures vary highly across applications and domains. Meanwhile, for natural language processing, the Generative Pre-trained Transformer (GPT) has demonstrated impressive performance via training one general-purpose model across various textual datasets. It is intriguing to explore whether GPT-type architectures can be effective for time series, capturing the intrinsic dynamic attributes and leading to significant accuracy improvements. In this paper, we propose a novel framework, TEMPO, that can effectively learn time series representations. We focus on utilizing two essential inductive biases of the time series task for pre-trained models: (i) decomposition of the complex interaction between trend, seasonal and residual components; and (ii) introducing the design of prompts to facilitate distribution adaptation in different types of time series. TEMPO expands the capability for dynamically modeling real-world temporal phenomena from data within diverse domains. Our experiments demonstrate the superior performance of TEMPO over state-of-the-art methods on zero shot setting for a number of time series benchmark datasets. This performance gain is observed not only in scenarios involving previously unseen datasets but also in scenarios with multi-modal inputs. This compelling finding highlights TEMPO's potential to constitute a foundational model-building framework. Defu Cao, Furong Jia 0002, Sercan Ö. Arik, Tomas Pfister, Yixiang Zheng, Wen Ye 0001, Yan Liu 0002 |
ICLR | 1 |
| 2024 | Neuro-Inspired Information-Theoretic Hierarchical Perception for Multimodal LearningabstractIntegrating and processing information from various sources or modalities are critical for obtaining a comprehensive and accurate perception of the real world in autonomous systems and cyber-physical systems. Drawing inspiration from neuroscience, we develop the Information-Theoretic Hierarchical Perception (ITHP) model, which utilizes the concept of information bottleneck. Different from most traditional fusion models that incorporate all modalities identically in neural networks, our model designates a prime modality and regards the remaining modalities as detectors in the information pathway, serving to distill the flow of information. Our proposed perception model focuses on constructing an effective and compact information flow by achieving a balance between the minimization of mutual information between the latent state and the input modal state, and the maximization of mutual information between the latent states and the remaining modal states. This approach leads to compact latent state representations that retain relevant information while minimizing redundancy, thereby substantially enhancing the performance of multimodal representation learning. Experimental evaluations on the MUStARD, CMU-MOSI, and CMU-MOSEI datasets demonstrate that our model consistently distills crucial information in multimodal learning scenarios, outperforming state-of-the-art benchmarks. Remarkably, on the CMU-MOSI dataset, ITHP surpasses human-level performance in the multimodal sentiment binary classification task across all evaluation metrics (i.e., Binary Accuracy, F1 Score, Mean Absolute Error, and Pearson Correlation). Xiongye Xiao, Gengshuo Liu, Defu Cao, Yaxing Li, Tianqing Fang, Mingxi Cheng, Paul Bogdan |
ICLR | 4 |
| 2024 | An Empirical Examination of Balancing Strategy for Counterfactual Estimation on Time SeriesabstractCounterfactual estimation from observations represents a critical endeavor in numerous application fields, such as healthcare and finance, with the primary challenge being the mitigation of treatment bias. The balancing strategy aimed at reducing covariate disparities between different treatment groups serves as a universal solution. However, when it comes to the time series data, the effectiveness of balancing strategies remains an open question, with a thorough analysis of the robustness and applicability of balancing strategies still lacking. This paper revisits counterfactual estimation in the temporal setting and provides a brief overview of recent advancements in balancing strategies. More importantly, we conduct a critical empirical examination for the effectiveness of the balancing strategies within the realm of temporal counterfactual estimation in various settings on multiple datasets. Our findings could be of significant interest to researchers and practitioners and call for a reexamination of the balancing strategy in time series settings. Chuizheng Meng, Defu Cao, Biwei Huang, Yi Chang 0001, Yan Liu 0002 |
ICML | 3 |
| 2024 | Mixture of Projection Experts for Multivariate Long-Term Time Series ForecastingabstractMultivariate long-term time series forecasting (MLTSF), applicable across various domains, has gained increasing research attention. Channel-independent (CI) models, including Linear and Transformer-based architectures, have recently achieved state-of-the-art (SOTA) performance for MLTSF. Notably, Linear models can deliver satisfactory forecasting performance even with just a single linear projection layer. However, we identify a limitation in this architecture: a single linear projection struggles to adequately capture the inter- and intra-variate heterogeneity in temporal patterns. Similarly, any complex models like Transformer-based models that use a single projection layer to generate final predictions, may face capacity bottlenecks. To overcome this, we propose the Mixture of Projection Experts (MoPE), which replaces the single linear projection with multiple projection branches, and employs a gate network to dynamically assign weights to each branch based on the input data. We applied MoPE to multiple SOTA models and evaluated it on nine real-world datasets. Results show that MoPE boosts forecasting accuracy by an average of 9.59%, demonstrating its effectiveness in mitigating the limitations of a single projection layer. Additionally, our experiments demonstrate that integrating our proposal into CI models enhances their generalization to unseen variates. Interpretability analysis also reveals MoPE's ability to disentangle different temporal patterns. Overall, our paper establishes MoPE as an effective solution for MLTSF tasks. Hao Niu 0001, Guillaume Habault, Defu Cao, Roberto Legaspi, Huy Quang Ung, James Enouen, Shinya Wada, Chihiro Ono, Atsunori Minamikawa, Yan Liu 0002 |
ICMLA | 3 |
| 2024 | Collaborative Multi-Task Representation for Natural Language UnderstandingabstractMulti-task learning has shown large benefits in Natural Language Understanding (NLU). However, current state-of-the-arts (SOTAs) like MT-DNN and MMoE do not model task relationships explicitly and fail to obtain effective task alignment. In this paper, we propose a Collaborative Multi-Task Representation (CMTR) framework to tackle this problem. We capture instance-level task relations through a task interaction layer, which helps guide the fusion of task-oriented representations into the final representation. Moreover, tailored loss functions are proposed to facilitate the learning of task alignment. Specifically, we leverage knowledge distillation as an auxiliary loss to assist the adaptation layers in generating task-oriented representations. We also introduce a regularization loss to learn better gating functions for multi-task fusion. Empirically, CMTR outperforms SOTA multi-task learning frameworks on most natural language understanding tasks in the GLUE benchmark. Furthermore, it achieves better task alignment and demonstrates good interpretability. Yaming Yang 0001, Defu Cao, Ming Zeng 0009, Jing Yu 0007, Yunhai Tong, Yujing Wang 0002 |
IJCNN | 2 |
| 2024 | Active Sequential Posterior Estimation for Sample-Efficient Simulation-Based InferenceabstractComputer simulations have long presented the exciting possibility of scientific insight into complex real-world processes. Despite the power of modern computing, however, it remains challenging to systematically perform inference under simulation models. This has led to the rise of simulation-based inference (SBI), a class of machine learning-enabled techniques for approaching inverse problems with stochastic simulators. Many such methods, however, require large numbers of simulation samples and face difficulty scaling to high-dimensional settings, often making inference prohibitive under resource-intensive simulators. To mitigate these drawbacks, we introduce active sequential neural posterior estimation (ASNPE). ASNPE brings an active learning scheme into the inference loop to estimate the utility of simulation parameter candidates to the underlying probabilistic model. The proposed acquisition scheme is easily integrated into existing posterior estimation pipelines, allowing for improved sample efficiency with low computational overhead. We further demonstrate the effectiveness of the proposed method in the travel demand calibration setting, a high-dimensional inverse problem commonly requiring computationally expensive traffic simulators. Our method outperforms well-tuned benchmarks and state-of-the-art posterior estimation methods on a large-scale real-world traffic network, as well as demonstrates a performance advantage over non-active counterparts on a suite of SBI benchmark environments. Sam Griesemer, Defu Cao, Zijun Cui, Carolina Osorio, Yan Liu 0002 |
NeurIPS | 2 |
| 2024 | MuGSI: Distilling GNNs with Multi-Granularity Structural Information for Graph ClassificationabstractRecent works have introduced GNN-to-MLP knowledge distillation (KD) frameworks to combine both GNN's superior performance and MLP's fast inference speed. However, existing KD frameworks are primarily designed for node classification within single graphs, leaving their applicability to graph classification largely unexplored. Two main challenges arise when extending KD for node classification to graph classification: (1) The inherent sparsity of learning signals due to soft labels being generated at the graph level; (2) The limited expressiveness of student MLPs, especially in datasets with limited input feature spaces. To overcome these challenges, we introduce MuGSI, a novel KD framework that employs Multi-granularity Structural Information for graph classification. Specifically, we propose multi-granularity distillation loss in MuGSI to tackle the first challenge. This loss function is composed of three distinct components: graph-level distillation, subgraph-level distillation, and node-level distillation. Each component targets a specific granularity of the graph structure, ensuring a comprehensive transfer of structural knowledge from the teacher model to the student model. To tackle the second challenge, MuGSI proposes to incorporate a node feature augmentation component, thereby enhancing the expressiveness of the student MLPs and making them more capable learners. We perform extensive experiments across a variety of datasets and different teacher/student model architectures. The experiment results demonstrate the effectiveness, efficiency, and robustness of MuGSI. Codes are publicly available at: https://github.com/tianyao-aka/MuGSI. Tianjun Yao, Defu Cao, Kun Zhang 0001, Guangyi Chen 0002 |
WWW | 3 |
| 2023 | Estimating Treatment Effects from Irregular Time Series Observations with Hidden ConfoundersabstractCausal analysis for time series data, in particular estimating individualized treatment effect (ITE), is a key task in many real world applications, such as finance, retail, healthcare, etc. Real world time series, i.e., large-scale irregular or sparse and intermittent time series, raise significant challenges to existing work attempting to estimate treatment effects. Specifically, the existence of hidden confounders can lead to biased treatment estimates and complicate the causal inference process. In particular, anomaly hidden confounders which exceed the typical range can lead to high variance estimates. Moreover, in continuous time settings with irregular samples, it is challenging to directly handle the dynamics of causality. In this paper, we leverage recent advances in Lipschitz regularization and neural controlled differential equations (CDE) to develop an effective and scalable solution, namely LipCDE, to address the above challenges. LipCDE can directly model the dynamic causal relationships between historical data and outcomes with irregular samples by considering the boundary of hidden confounders given by Lipschitz constrained neural networks. Furthermore, we conduct extensive experiments on both synthetic and real world datasets to demonstrate the effectiveness and scalability of LipCDE. Defu Cao, James Enouen, Yujing Wang 0002, Xiangchen Song, Chuizheng Meng, Hao Niu 0001, Yan Liu 0002 |
AAAI | 1 |
| 2023 | SVGformer: Representation Learning for Continuous Vector Graphics using TransformersabstractAdvances in representation learning have led to great success in understanding and generating data in various domains. However, in modeling vector graphics data, the pure data-driven approach often yields unsatisfactory results in downstream tasks as existing deep learning methods often require the quantization of SVG parameters and cannot exploit the geometric properties explicitly. In this paper, we propose a transformer-based representation learning model (SVG-former) that directly operates on continuous input values and manipulates the geometric information of SVG to encode outline details and long-distance dependencies. SVGfomer can be used for various downstream tasks: reconstruction, classification, interpolation, retrieval, etc. We have conducted extensive experiments on vector font and icon datasets to show that our model can capture high-quality representation information and outperform the previous state-of-the-art on downstream tasks significantly. Defu Cao, Jose Echevarria, Yan Liu 0002 |
CVPR | 1 |
| 2023 | Coupled Multiwavelet Operator Learning for Coupled Differential Equations
Xiongye Xiao, Defu Cao, Ruochen Yang, Gengshuo Liu, Chenzhong Yin, Radu Balan, Paul Bogdan |
ICLR | 2 |
| 2023 | Time-delayed Multivariate Time Series PredictionsabstractA major issue with real-time monitoring is to collect complete data. Hardware or software failures, network issues or, more frequently, time delays can disrupt such a collection. This results in having two versions of the same information: one in real-time but with potentially missing data, and the another, albeit complete, is delayed. Many works have studied how to handle missing data for classification and prediction. However, to the best of our knowledge, they do not consider how to leverage the delayed complete data to assist in learning the representation of real-time available data with missing values. This is despite the fact that the delayed complete data contain all the information (e.g., periodicities and trends). In this paper, we propose a framework to enhance the representation learning of the real-time available data by aligning the representation of past real-time but with missing data to that of past delayed but complete data. We test both a distance metric and contrastive learning to achieve this alignment. We implement our framework on a Transformer-based model and experiment it on three datasets. The efficiency of our solution is evaluated against seven baselines and considering four distinct patterns of missing data. Our experiments show that this proposal has a significant improvement in prediction accuracy (5.21% on average) over the baselines. Hao Niu 0001, Guillaume Habault, Roberto Legaspi, Chuizheng Meng, Defu Cao, Shinya Wada, Chihiro Ono, Yan Liu 0002 |
SDM | 5 |
| 2022 | Enhancing Self-Attention with Knowledge-Assisted Attention MapsabstractJiangang Bai, Yujing Wang, Hong Sun, Ruonan Wu, Tianmeng Yang, Pengfei Tang, Defu Cao, Mingliang Zhang1, Yunhai Tong, Yaming Yang, Jing Bai, Ruofei Zhang, Hao Sun, Wei Shen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jiangang Bai, Yujing Wang 0002, Ruonan Wu, Tianmeng Yang, Defu Cao, Mingliang Zhang 0004, Yunhai Tong, Yaming Yang 0001, Jing Bai 0010, Ruofei Zhang, Hao Sun 0015 |
NAACL-HLT | 7 |
| 2022 | Counterfactual Neural Temporal Point Process for Estimating Causal Influence of Misinformation on Social MediaabstractRecent years have witnessed the rise of misinformation campaigns that spread specific narratives on social media to manipulate public opinions on different areas, such as politics and healthcare. Consequently, an effective and efficient automatic methodology to estimate the influence of the misinformation on user beliefs and activities is needed. However, existing works on misinformation impact estimation either rely on small-scale psychological experiments or can only discover the correlation between user behaviour and misinformation. To address these issues, in this paper, we build up a causal framework that model the causal effect of misinformation from the perspective of temporal point process. To adapt the large-scale data, we design an efficient yet precise way to estimate the \textbf{Individual Treatment Effect} (ITE) via neural temporal point process and gaussian mixture models. Extensive experiments on synthetic dataset verify the effectiveness and efficiency of our model. We further apply our model on a real-world dataset of social media posts and engagements about COVID-19 vaccines. The experimental results indicate that our model recognized identifiable causal effect of misinformation that hurts people's subjective emotions toward the vaccines. Defu Cao, Yan Liu 0002 |
NeurIPS | 2 |
| 2022 | Mu2ReST: Multi-resolution Recursive Spatio-Temporal Transformer for Long-Term Prediction
Hao Niu 0001, Chuizheng Meng, Defu Cao, Guillaume Habault, Roberto Legaspi, Shinya Wada, Chihiro Ono, Yan Liu 0002 |
PAKDD (1) | 3 |
| 2021 | Spectral Temporal Graph Neural Network for Trajectory PredictionabstractAn effective understanding of the contextual environment and accurate motion forecasting of surrounding agents is crucial for the development of autonomous vehicles and social mobile robots. This task is challenging since the behavior of an autonomous agent is not only affected by its own intention, but also by the static environment and surrounding dynamically interacting agents. Previous works focused on utilizing the spatial and temporal information in time domain while not sufficiently taking advantage of the cues in frequency domain. To this end, we propose a Spectral Temporal Graph Neural Network (SpecTGNN), which can capture inter-agent correlations and temporal dependency simultaneously in frequency domain in addition to time domain. SpecTGNN operates on both an agent graph with dynamic state information and an environment graph with the features extracted from context images in two streams. The model integrates graph Fourier transform, spectral graph convolution and temporal gated convolution to encode history information and forecast future trajectories. Moreover, we incorporate a multi-head spatio-temporal attention mechanism to mitigate the effect of error propagation in a long time horizon. We demonstrate the performance of SpecTGNN on two public trajectory prediction benchmark datasets, which achieves state-of-the-art performance in terms of prediction accuracy. Defu Cao, Jiachen Li 0001, Hengbo Ma, Masayoshi Tomizuka |
ICRA | 1 |
| 2020 | Multivariate Time-series Anomaly Detection via Graph Attention NetworkabstractAnomaly detection on multivariate time-series is of great importance in both data mining research and industrial applications. Recent approaches have achieved significant progress in this topic, but there is remaining limitations. One major limitation is that they do not capture the relationships between different time-series explicitly, resulting in inevitable false alarms. In this paper, we propose a novel self-supervised framework for multivariate time-series anomaly detection to address this issue. Our framework considers each univariate time-series as an individual feature and includes two graph attention layers in parallel to learn the complex dependencies of multivariate time-series in both temporal and feature dimensions. In addition, our approach jointly optimizes a forecasting-based model and a reconstruction-based model, obtaining better time-series representations through a combination of single-timestamp prediction and reconstruction of the entire time-series. We demonstrate the efficacy of our model through extensive experiments. The proposed method outperforms other state-of-the-art models on three real-world datasets. Further analysis shows that our method has good interpretability and is useful for anomaly diagnosis. Yujing Wang 0002, Juanyong Duan, Congrui Huang, Defu Cao, Yunhai Tong, Bixiong Xu, Jing Bai 0010, Jie Tong, Qi Zhang 0066 |
ICDM | 5 |
| 2020 | Spectral Temporal Graph Neural Network for Multivariate Time-series ForecastingabstractMultivariate time-series forecasting plays a crucial role in many real-world applications. It is a challenging problem as one needs to consider both intra-series temporal correlations and inter-series correlations simultaneously. Recently, there have been multiple works trying to capture both correlations, but most, if not all of them only capture temporal correlations in the time domain and resort to pre-defined priors as inter-series relationships. In this paper, we propose Spectral Temporal Graph Neural Network (StemGNN) to further improve the accuracy of multivariate time-series forecasting. StemGNN captures inter-series correlations and temporal dependencies jointly in the spectral domain. It combines Graph Fourier Transform (GFT) which models inter-series correlations and Discrete Fourier Transform (DFT) which models temporal dependencies in an end-to-end framework. After passing through GFT and DFT, the spectral representations hold clear patterns and can be predicted effectively by convolution and sequential learning modules. Moreover, StemGNN learns inter-series correlations automatically from the data without using pre-defined priors. We conduct extensive experiments on ten real-world datasets to demonstrate the effectiveness of StemGNN. Defu Cao, Yujing Wang 0002, Juanyong Duan, Ce Zhang 0001, Congrui Huang, Yunhai Tong, Bixiong Xu, Jing Bai 0010, Jie Tong, Qi Zhang 0066 |
NeurIPS | 1 |
| 2020 | FTCLNet: Convolutional LSTM with Fourier Transform for Vulnerability DetectionabstractAs software vulnerabilities become increasingly serious, it is necessary to detect them efficiently and accurately. However, vulnerabilities are diverse and context sensitive. Previous solutions either rely on features defined by experts, or use only recurrent neural networks on code sequence. It is difficult to extract complex features of vulnerabilities in traditional code space. This article proposes a deep convolutional LSTM neural network with Fourier transform for vulnerability detection. The discrete Fourier transform method convert code space into frequency domain, which significantly helps deep models learn remarkable patterns. This article combines convolutional neural network (CNN) with long short term memory (LSTM) network to extract local and global features in frequency domain, and utilize attention mechanism to decide the weight of each element in code space. Besides, this method rewrite the source code and convert them to vectors without guidance from the specified domain knowledge. Experiments on Buffer Error dataset (CWE-119) and Resource Management Error dataset (CWE-399) show that this new method achieves a significantly improved results. Defu Cao, Xianhua Liu 0001 |
TrustCom | 1 |