VLDB 2026 Research / reviewers in the wild / expert
Binwu Wang
dblp:262/4302
· DBLP profile ↗
39ranked-venue papers
8as first author
39since 2021 · last 2026
0000-0002-4638-0382ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 19 since 2021Databases, data management, data science and information retrieval · 14 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | U2B: Scale-unbiased Representation Converter for Graph Classification with Imbalanced and Balanced Scale DistributionsabstractGraph classification is a critical task in analyzing graph data, with applications across various domains. While graph neural networks (GNNs) have achieved remarkable results, their ability to generalize across graphs of varying scales remains a challenge. Conventional models often perform well on large-scale graphs but struggle with distributions that are skewed towards small scales. Conversely, models tailored to address scale imbalances frequently prioritize small-scale graphs, leading to diminished performance in more balanced scenarios. To overcome these limitations, we introduce a Unbalanced-Balanced Representation Converter (U2B), which exhibits no explicit bias toward graph scales. U2B employs a two-step workflow: a distillation phase to extract base features from both node-level and graph-level representations, followed by a refinement phase to generate unbiased representations for improved balance. In the distillation phase, a static constraint guides node-level adjustments, improving the representation of nodes in small graphs. Simultaneously, a dynamic constraint in the graph-level process mitigates biases toward features from large graphs. To ensure harmony between the representations, a consistency alignment loss is introduced, aligning node-level and graph-level features to create more cohesive and balanced graph representations. Extensive experiments on multiple datasets show that U2B achieves competitive performance. Jiaming Ma, Pengkun Wang 0001, Zhengyang Zhou, Binwu Wang, Yang Wang 0015 |
AAAI | 7 |
| 2026 | Augur: Modeling Covariate Causal Associations in Time Series via Large Language ModelsabstractZhiqing Cui, Binwu Wang, Qingxiang Liu, Yeqiang Wang, Zhengyang Zhou, Yuxuan Liang, Yang Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhiqing Cui, Binwu Wang, Qingxiang Liu 0004, Yeqiang Wang, Zhengyang Zhou, Yuxuan Liang 0002, Yang Wang 0015 |
ACL (1) | 2 |
| 2026 | We Need a More Robust Classifier: Dual Causal Learning Empowers Domain-Incremental Time Series ClassificationabstractThe World Wide Web thrives on intelligent services that rely on accurate time series classification, which has recently witnessed significant progress driven by advances in deep learning. However, existing studies face challenges in domain incremental learning. In this paper, we propose a lightweight and robust dual-causal disentanglement framework (DualCD) to enhance the robustness of models under domain incremental scenarios, which can be seamlessly integrated into time series classification models. Specifically, DualCD first introduces a temporal feature disentanglement module to capture class-causal features and spurious features. The causal features can offer sufficient predictive power to support the classifier in domain incremental learning settings. To accurately capture these causal features, we further design a dual-causal intervention mechanism to eliminate the influence of both intra-class and inter-class confounding features. This mechanism constructs variant samples by combining the current class's causal features with intra-class spurious features and with causal features from other classes. The causal intervention loss encourages the model to accurately predict the labels of these variant samples based solely on the causal features. Extensive experiments on multiple datasets and models demonstrate that DualCD effectively improves performance in domain incremental scenarios. We summarize our rich experiments into a comprehensive benchmark to facilitate research in domain incremental time series classification. Peibo Duan, Haodong Jing, Mingyang Geng, Jialu Xu, Bin Zhang 0001, Binwu Wang |
WWW | 9 |
| 2026 | QuiZSF: A Retrieval-Augmented Framework for Zero-Shot Time Series ForecastingabstractAccurate forecasting of sequential data streams is a cornerstone of modern Web services, supporting applications such as traffic management, user behavior modeling, and online anomaly prevention. However, in many Web environments, new domains emerge rapidly and labeled history data is scarce, which makes zero-shot forecasting particularly challenging. Existing time-series pre-trained models (TSPMs) show promise but they lack the ability to dynamically incorporate external knowledge, while conventional retrieval-augmented generation (RAG) methods are rarely extended beyond text. In this work, we present QuiZSF, a retrieval-augmented forecasting framework that integrates search and forecasting for time series data. The framework performs search by retrieving structurally similar sequences from a large-scale time-series database, and it performs forecasting by integrating the retrieved knowledge into the target sequence. Specifically, QuiZSF introduces a ChronoRAG Base, a hierarchical tree-structured database that enables scalable and domain-aware retrieval, a Multi-grained Series Interaction Learner that captures fine- and coarse-grained dependencies between target and retrieved sequences, and a Model Cooperation Coherer that adapts retrieved knowledge to TSPMs. This design teaches models to actively perform search, align auxiliary information across modalities, and leverage it for more accurate forecasting. Extensive experiments on five public benchmarks demonstrate that QuiZSF consistently outperforms strong baselines, ranking first in up to 87.5% of zero-shot forecasting settings while maintaining high efficiency. Zhengyang Zhou, Qihe Huang, Binwu Wang, Yang Wang 0015 |
WWW | 4 |
| 2026 | TimeFormer: Transformer with attention modulation empowered by temporal characteristics for time series forecastingabstractAlthough Transformers excel in natural language processing, their extension to time series forecasting remains challenging due to insufficient consideration of the differences between textual and temporal modalities. In this paper, we develop a novel Transformer architecture designed for time series data, aiming to maximize its representational capacity. We identify two key but often overlooked characteristics of time series: (1) unidirectional influence from the past to the future, and (2) the phenomenon of decaying influence over time. These characteristics are introduced to enhance the attention mechanism of Transformers. We propose TimeFormer, whose core innovation is a self-attention mechanism with two modulation terms (MoSA), designed to capture these temporal priors of time series under the constraints of the Hawkes process and causal masking. Additionally, TimeFormer introduces a framework based on multi-scale and subsequence analysis to capture semantic dependencies at different temporal scales, enriching the temporal dependencies. Extensive experiments conducted on multiple real-world datasets show that TimeFormer significantly outperforms state-of-the-art methods, achieving up to a 7.45% reduction in MSE compared to the best baseline and setting new benchmarks on 94.04% of evaluation metrics. Moreover, we demonstrate that the MoSA mechanism can be broadly applied to enhance the performance of other Transformer-based models. Peibo Duan, Baixin Li, Mingyang Geng, Changsheng Zhang 0001, Bin Zhang 0001, Binwu Wang |
Expert Syst. Appl. | 9 |
| 2026 | MADGCN: A Meteorology-Aware Spatio-Temporal Graph Convolution Network for Long-Term Air Pollution ForecastingabstractAir quality forecasting has attracted increasing attention as global air pollution worsens. Spatiotemporal graph neural networks have become a leading paradigm, thanks to their ability to capture complex spatial and temporal dynamics in Air Quality Index (AQI) data. However, existing methods remain limited by weak modeling of long-range temporal dependencies and insufficient integration of meteorological factors. Building on a publicly available nationwide air quality dataset spanning eight years, we propose MADGCN, a Meteorology-Aware Decoupled Spatio-Temporal Convolutional Network that jointly addresses long-horizon temporal modeling and meteorological context fusion. MADGCN includes a dynamic causality discovery module grounded in Granger causality, which captures time-varying causal relationships between meteorological conditions and AQI dynamics. The inferred causal structures further guide a causal graph convolution module and a PatchMixer module, enabling effective spatial interaction modeling and multiscale temporal dependency learning. Extensive experiments against 16 strong baselines show that MADGCN achieves competitive performance for long-horizon air pollution forecasting and generalizes well under high-pollution regimes. Binwu Wang, Zhiqing Cui, Guangjun Wang, Zhengyang Zhou, Fan Meng 0008, Jingjia Luo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Spatiotemporal Causal Decoupling Model for Air Quality ForecastingabstractDue to the profound impact of air pollution on human health, livelihoods, and economic development, air quality forecasting is of paramount significance. Initially, we employ the causal graph method to scrutinize the constraints of existing research in comprehensively modeling the causal relationships between the air quality index (AQI) and meteorological features. In order to enhance prediction accuracy, we introduce a novel air quality forecasting model, AirCade, which incorporates a causal decoupling approach. AirCade leverages a spatiotemporal module in conjunction with knowledge embedding techniques to capture the internal dynamics of AQI. Subsequently, a causal decoupling module is proposed to disentangle synchronous causality from past AQI and meteorological features, followed by the dissemination of acquired knowledge to future time steps to enhance performance. Additionally, we introduce a causal intervention mechanism to explicitly represent the uncertainty of future meteorological features, thereby bolstering the model’s robustness. Our evaluation of AirCade on an open-source air quality dataset demonstrates over 20% relative improvement over state-of-the-art models. Our source code is available at https://github.com/PoorOtterBob/AirCade. Jiaming Ma, Kuo Yang 0002, Binwu Wang, Pengkun Wang 0001, Yang Wang 0015 |
ICASSP | 5 |
| 2025 | Time-Space-Interlaced Spatiotemporal Graph Forecasting via Two-Stage Summarized AttentionabstractTypical spatiotemporal graph forecasting methods process graph-structured spatiotemporal data respectively from spatial and temporal perspectives with the idea of divide and conquer. Existing works are incapable of capturing long-term transdimensional correlations among different spatial points in different time planes, i.e., time-space-interlaced correlations. To tackle this issue, we propose a two-stage summarized attention network to establish transdimensional direct message passing routes between different data points in different time planes and spaces, thus enabling the extraction of time-space-interlaced long-term correlations. Specifically, a novel spatiotemporal embedding is proposed to implement time-space-interlaced learning by expanding orthogonal spatial and temporal dimensionalities into one-dimensionality, a series of temporal context fusion units are added to address the fluctuation dislocation insensitivity issue which is caused by time-space-dimension expansion, and an ingenious two-stage design can significantly reduce the computation complexity of such time-space-interlaced learning. Extensive experiments illustrate the superior performance of our proposed approach on real-world spatiotemporal datasets. Zhaoyang Sun, Yudong Zhang 0005, Kai Wang 0036, Binwu Wang, Yang Wang 0015, Xu Wang 0029 |
ICASSP | 5 |
| 2025 | Robust Spatio-Temporal Centralized Interaction for OOD LearningabstractRecently, spatiotemporal graph convolutional networks have achieved dominant performance in spatiotemporal prediction tasks. However, most models relying on node-to-node messaging interaction exhibit sensitivity to spatiotemporal shifts, encountering out-of-distribution (OOD) challenges. To address these issues, we introduce \textbf{\underline{S}}patio-\textbf{\underline{T}}emporal \textbf{\underline{O}}OD \textbf{\underline{P}}rocessor (STOP), which employs a centralized messaging mechanism along with a message perturbation mechanism to facilitate robust spatiotemporal interactions. Specifically, the centralized messaging mechanism integrates Context-Aware Units for coarse-grained spatiotemporal feature interactions with nodes, effectively blocking traditional node-to-node messages. We also implement a message perturbation mechanism to disrupt this messaging process, compelling the model to extract generalizable contextual features from generated variant environments. Finally, we customize a spatiotemporal distributionally robust optimization approach that exposes the model to challenging environments, thereby further enhancing its generalization capabilities. Compared with 14 baselines across six datasets, STOP achieves up to \textbf{17.01\%} improvement in generalization performance and \textbf{18.44\%} improvement in inductive learning performance. The code is available at https://github.com/PoorOtterBob/STOP. Jiaming Ma, Binwu Wang, Pengkun Wang 0001, Zhengyang Zhou, Xu Wang 0029, Yang Wang 0015 |
ICML | 2 |
| 2025 | Causal Learning Meet Covariates: Empowering Lightweight and Effective Nationwide Air Quality ForecastingabstractAir quality prediction plays a crucial role in the development of smart cities, garnering significant attention from both academia and industry. Current air quality prediction models encounter two major limitations: their high computational complexity limits scalability to nationwide datasets, and they often regard weather covariates as optional auxiliary information. In reality, weather covariates can have a substantial impact on air quality indices (AQI), exhibiting a significant causal association. In this paper, we first present a nationwide air quality dataset to address the lack of open-source, large-scale datasets in this field. Then we propose a causal learning model, CauAir, for air quality prediction that harnesses the powerful representation capabilities of the Transformer to explicitly model the causal association between weather covariates and AQI. To address the high complexity of traditional Transformers, we design CachLormer, which features two key innovations: a simplified architecture with redundant components removed, and a cache-attention mechanism that employs learnable embeddings for perceiving causal association between AQI and weather covariates in a coarsegrained perspective. We use information theory to illustrate the superiority of the proposed model. Finally, experimental results on three datasets with 28 as the baseline demonstrate that our model achieves competitive performance, while maintaining high training efficiency and low memory consumption. The source code is available at CauAir Official Repository. Jiaming Ma, Zhiqing Cui, Binwu Wang, Pengkun Wang 0001, Zhengyang Zhou, Zhe Zhao 0008, Yang Wang 0015 |
IJCAI | 3 |
| 2025 | DisMS-TS: Eliminating Redundant Multi-scale Features for Time Series ClassificationabstractReal-world time series typically exhibit complex temporal variations, making the time series classification task notably challenging. Recent advancements have demonstrated the potential of multi-scale analysis approaches, which provide an effective solution for capturing these complex temporal patterns. However, existing multi-scale analysis-based time series prediction methods fail to eliminate redundant scale-shared features across multi-scale time series, resulting in the model over- or under-focusing on scale-shared features. To address this issue, we propose a novel end-to-end Disentangled Multi-Scale framework for Time Series classification (DisMS-TS). The core idea of DisMS-TS is to eliminate redundant shared features in multi-scale time series, thereby improving prediction performance. Specifically, we propose a temporal disentanglement module to capture scale-shared and scale-specific temporal representations, respectively. Subsequently, to effectively learn both scale-shared and scale-specific temporal representations, we introduce two regularization terms that ensure the consistency of scale-shared representations and the disparity of scale-specific representations across all temporal scales. Extensive experiments conducted on multiple datasets validate the superiority of DisMS-TS over its competitive baselines, with the accuracy improvement up to 9.71%. Peibo Duan, Binwu Wang, Qi Chu 0012, Changsheng Zhang 0001, Bin Zhang 0001 |
ACM Multimedia | 3 |
| 2025 | Many Minds, One Goal: Time Series Forecasting via Sub-task Specialization and Inter-agent CooperationabstractTime series forecasting is a critical and complex task, characterized by diverse temporal patterns, varying statistical properties, and different prediction horizons across datasets and domains. Conventional approaches typically rely on a single, unified model architecture to handle all forecasting scenarios. However, such monolithic models struggle to generalize across dynamically evolving time series with shifting patterns. In reality, different types of time series may require distinct modeling strategies. Some benefit from homogeneous multi-scale forecasting awareness, while others rely on more complex and heterogeneous signal perception. Relying on a single model to capture all temporal diversity and structural variations leads to limited performance and poor interpretability. To address this challenge, we propose a Multi-Agent Forecasting System (MAFS) that abandons the one-size-fits-all paradigm. MAFS decomposes the forecasting task into multiple sub-tasks, each handled by a dedicated agent trained on specific temporal perspectives (e.g., different forecasting resolutions or signal characteristics). Furthermore, to achieve holistic forecasting, agents share and refine information through different communication topology, enabling cooperative reasoning across different temporal views. A lightweight voting aggregator then integrates their outputs into consistent final predictions. Extensive experiments across 11 benchmarks demonstrate that MAFS significantly outperforms traditional single-model approaches, yielding more robust and adaptable forecasts. Qihe Huang, Zhengyang Zhou, Yangze Li, Kuo Yang 0002, Binwu Wang, Yang Wang 0015 |
NeurIPS | 5 |
| 2025 | MoFo: Empowering Long-term Time Series Forecasting with Periodic Pattern ModelingabstractThe stable periodic patterns present in the time series data serve as the foundation for long-term forecasting. However, existing models suffer from limitations such as continuous and chaotic input partitioning, as well as weak inductive biases, which restrict their ability to capture such recurring structures. In this paper, we propose MoFo, which interprets periodicity as both the correlation of period-aligned time steps and the trend of period-offset time steps. We first design period-structured patches—2D tensors generated through discrete sampling—where each row contains only period-aligned time steps, enabling direct modeling of periodic correlations. Period-offset time steps within a period are aligned in columns. To capture trends across these offset time steps, we introduce a period-aware modulator. This modulator introduces an adaptive strong inductive bias through a regulated relaxation function, encouraging the model to generate attention coefficients that align with periodic trends. This function is end-to-end trainable, enabling the model to adaptively capture the distinct periodic patterns across diverse datasets. Extensive empirical results on widely used benchmark datasets demonstrate that MoFo achieves competitive performance while maintaining high memory efficiency and fast training speed. Jiaming Ma, Binwu Wang, Qihe Huang, Pengkun Wang 0001, Zhengyang Zhou, Yang Wang 0015 |
NeurIPS | 2 |
| 2025 | Less but More: Linear Adaptive Graph Learning Empowering Spatiotemporal ForecastingabstractThe effectiveness of Spatiotemporal Graph Neural Networks (STGNNs) critically hinges on the quality of the underlying graph topology. While end-to-end adaptive graph learning methods have demonstrated promising results in capturing latent spatiotemporal dependencies, they often suffer from high computational complexity and limited expressive capacity. In this paper, we propose MAGE for efficient spatiotemporal forecasting. We first conduct a theoretical analysis demonstrating that the ReLU activation function employed in existing methods amplifies edge-level noise during graph topology learning, thereby compromising the fidelity of the learned graph structures. To enhance model expressiveness, we introduce a sparse yet balanced mixture-of-experts strategy, where each expert perceives the unique underlying graph through kernel-based functions and operates with linear complexity relative to the number of nodes. The sparsity mechanism ensures that each node interacts exclusively with compatible experts, while the balancing mechanism promotes uniform activation across all experts, enabling diverse and adaptive graph representations. Furthermore, we theoretically establish that a single graph convolution using the learned graph in MAGE is mathematically equivalent to multiple convolutional steps under conventional graphs. We evaluate MAGE against advanced baselines on multiple real-world spatiotemporal datasets. MAGE achieves competitive performance while maintaining strong computational efficiency. Jiaming Ma, Binwu Wang, Kuo Yang 0002, Zhengyang Zhou, Pengkun Wang 0001, Xu Wang 0029, Yang Wang 0015 |
NeurIPS | 2 |
| 2025 | ComS2T: A Complementary Spatiotemporal Learning System for Data-Adaptive Model EvolutionabstractSpatiotemporal (ST) learning has become a crucial technique to enable smart cities and sustainable urban development. Current ST learning models capture the heterogeneity via various spatial convolution and temporal evolution blocks. However, rapid urbanization leads to fluctuating distributions in urban data and city structures, resulting in existing methods suffering generalization and data adaptation issues. Despite efforts, existing methods fail to deal with newly arrived observations, and the limitation of those methods with generalization capacity lies in the repeated training that leads to inconvenience, inefficiency and resource waste. Motivated by complementary learning in neuroscience, we introduce a prompt-based complementary spatiotemporal learning termed ComS2T, to empower the evolution of models for data adaptation. We first disentangle the neural architecture into two disjoint structures, a stable neocortex for consolidating historical memory, and a dynamic hippocampus for new knowledge update. Then we train the dynamic spatial and temporal prompts by characterizing distribution of main observations to enable prompts adaptive to new data. This data-adaptive prompt mechanism, combined with a two-stage training process, facilitates fine-tuning of the neural architecture conditioned on prompts, thereby enabling efficient adaptation during testing. Extensive experiments validate the efficacy of ComS2T in adapting various spatiotemporal out-of-distribution scenarios while maintaining effective inferences. Zhengyang Zhou, Qihe Huang, Binwu Wang, Jianpeng Hou, Kuo Yang 0002, Yuxuan Liang 0002, Yu Zheng 0004, Yang Wang 0015 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | BiST: A Lightweight and Efficient Bi-directional Model for Spatiotemporal PredictionabstractWhile existing spatiotemporal prediction models have shown promising performance, they often rely on the assumption of input-label spatiotemporal consistency, and their high complexity raises concerns about scalability. To enhance both efficiency and performance, we integrate label information into the learning process and propose a spatiotemporal dynamic theory that outlines a bi-directional learning paradigm. Building on this paradigm, we design BiST, a lightweight yet effective Bi -directional S patio -T emporal prediction model. BiST incorporates two key processes: a forward spatiotemporal learning process and a backward correction process. The forward process utilizes MLP layers exclusively to model input correlations and generate base prediction. In the backward process, we implement a spatiotemporal decoupling module, which can learn the residual modeling deviation between input and label representations from a decoupled perspective. After smoothing the residual with a diffusion module, we can obtain the correction term to correct the base predictions. This innovative design enables BiST to achieve competitive performance while remaining lightweight. We evaluate BiST against 26 baselines across 13 datasets, including a large-scale dataset with ten thousand nodes and a longrange dataset spanning 20 years. An impressive experimental result demonstrates that BiST achieves a 8.13% improvement in performance compared to state-of-the-art models while consuming only 1.86% of the training time and 7.36% of the memory usage. Jiaming Ma, Binwu Wang, Pengkun Wang 0001, Zhengyang Zhou, Xu Wang 0029, Yang Wang 0015 |
Proc. VLDB Endow. | 2 |
| 2025 | MobiMixer: A Multi-Scale Spatiotemporal Mixing Model for Mobile Traffic PredictionabstractUnderstanding mobile traffic data and predicting future trends are essential for wireless operators and service providers to allocate resources efficiently and manage energy effectively. Despite the strong performance of existing models, accurately forecasting mobile traffic remains a challenge due to limited spatial and temporal modeling capabilities and high computational complexity. This paper introduces MobiMixer, a lightweight and efficient multi-scale spatiotemporal mixing model. Its core concept is to integrate multi-scale information from both spatial and temporal dimensions to improve performance on mobile traffic data. We develop a hierarchical interaction module that incorporates super nodes to enable global high-level feature interactions among nodes with common patterns. Additionally, we employ a dynamic time warping strategy to decouple mobile traffic sequences into stable and seasonal components, which are then modeled at different scales using a multi-scale temporal mixing module. We conduct extensive experiments on mobile traffic datasets collected from four international cities. Compared with 21 state-of-the-art benchmark models, MobiMixer demonstrates highly competitive performance, achieving a maximum improvement of 48.49% on the Milan mobile dataset. The model achieves an improvement in training efficiency of up to 10.69 times and reduces memory usage by 33.01%. The source code is available athttps://github.com/PoorOtterBob/Submitted_Code. Jiaming Ma, Binwu Wang, Pengkun Wang 0001, Zhengyang Zhou, Yudong Zhang 0005, Xu Wang 0029, Yang Wang 0015 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Towards Dynamic Spatial-Temporal Graph Learning: A Decoupled PerspectiveabstractWith the progress of urban transportation systems, a significant amount of high-quality traffic data is continuously collected through streaming manners, which has propelled the prosperity of the field of spatial-temporal graph prediction. In this paper, rather than solely focusing on designing powerful models for static graphs, we shift our focus to spatial-temporal graph prediction in the dynamic scenario, which involves a continuously expanding and evolving underlying graph. To address inherent challenges, a decoupled learning framework (DLF) is proposed in this paper, which consists of a spatial-temporal graph learning network (DSTG) with a specialized decoupling training strategy. Incorporating inductive biases of time-series structures, DSTG can interpret time dependencies into latent trend and seasonal terms. To enable prompt adaptation to the evolving distribution of the dynamic graph, our decoupling training strategy is devised to iteratively update these two types of patterns. Specifically, for learning seasonal patterns, we conduct thorough training for the model using a long time series (e.g., three months of data). To enhance the learning ability of the model, we also introduce the masked auto-encoding mechanism. During this period, we frequently update trend patterns to expand new information from dynamic graphs. Considering both effectiveness and efficiency, we develop a subnet sampling strategy to select a few representative nodes for fine-tuning the weights of the model. These sampled nodes cover unseen patterns and previously learned patterns. Experiments on dynamic spatial-temporal graph datasets further demonstrate the competitive performance, superior efficiency, and strong scalability of the proposed framework. Binwu Wang, Pengkun Wang 0001, Yudong Zhang 0005, Xu Wang 0029, Zhengyang Zhou, Lei Bai 0001, Yang Wang 0015 |
AAAI | 1 |
| 2024 | Gradient Reactivation Enhanced Causal Attention for Out-Of-Distribution Generalizable Graph ClassificationabstractSeeking for generalizable graph representations becomes hot spot in the area of graph learning. Recently, causality theory has been applied for extracting the causal relations between graph data and labels, which are generalizable under distribution shift and result in better OOD generalization. In this paper, for more accurately capturing causal representation of graph data, we propose a gradient reactivation enhanced causal subgraph extraction method. The proposed model utilizes attention mechanism to extract the causal features and attenuates the confounding effect of shortcut features. For ensuring stability of extracted causal features, we propose a novel gradient reactivation method to filter features with greater effect on making prediction. Extensively experimental result proves the effectiveness of the proposed model. Xu Wang 0029, Pengfei Gu, Yudong Zhang 0005, Binwu Wang, Pengkun Wang 0001, Yang Wang 0015 |
ICASSP | 4 |
| 2024 | Graph Networks Stand Strong: Enhancing Robustness via Stability ConstraintsabstractGraph neural networks (GNNs) have achieved great success in graph classification tasks across many domains. However, the varying quality of real-world graph data leads to stability and reliability issues for real-world applications of graph neural networks (GNNs). Improving the robustness of GNNs would help enhance the quality and safety of GNNs in real-world applications. Recently, there have been studies that incorporate insights from information theory, causal theory, etc. into graph classification tasks to improve robustness. However, these strategies rely on extensive task-specific designs that increase model complexity and limit the scope of the methods. In this work, we leverage the interdependence between model stability and robustness by introducing stability constraints to graph neural network models through two different consistency regularization methods. To balance the trade-off between stability constraints and classification performance, we adaptively adjust the strength of the constraints dynamically using multi-objective optimization, making our method applicable to graph classification tasks of varying scales and domains. Extensive experiments on graph datasets from different domains demonstrate the superiority of our proposed method. Zhe Zhao 0008, Pengkun Wang 0001, Haibin Wen, Yudong Zhang 0005, Binwu Wang, Yang Wang 0015 |
ICASSP | 5 |
| 2024 | Kill Two Birds with One Stone: Rethinking Data Augmentation for Deep Long-tailed LearningabstractReal-world tasks are universally associated with training samples that exhibit a long-tailed class distribution, and traditional deep learning models are not suitable for fitting this distribution, thus resulting in a biased trained model. To surmount this dilemma, massive deep long-tailed learning studies have been proposed to achieve inter-class fairness models by designing sophisticated sampling strategies or improving existing model structures and loss functions. Habitually, these studies tend to apply data augmentation strategies to improve the generalization performance of their models. However, this augmentation strategy applied to balanced distributions may not be the best option for long-tailed distributions. For a profound understanding of data augmentation, we first theoretically analyze the gains of traditional augmentation strategies in long-tailed learning, and observe that augmentation methods cause the long-tailed distribution to be imbalanced again, resulting in an intertwined imbalance: inherent data-wise imbalance and extrinsic augmentation-wise imbalance, i.e., two 'birds' co-exist in long-tailed learning. Motivated by this observation, we propose an adaptive Dynamic Optional Data Augmentation (DODA) to address this intertwined imbalance, i.e., one 'stone' simultaneously 'kills' two 'birds', which allows each class to choose appropriate augmentation methods by maintaining a corresponding augmentation probability distribution for each class during training. Extensive experiments across mainstream long-tailed recognition benchmarks (e.g., CIFAR-100-LT, ImageNet-LT, and iNaturalist 2018) prove the effectiveness and flexibility of the DODA in overcoming the intertwined imbalance. Binwu Wang, Pengkun Wang 0001, Wei Xu 0055, Xu Wang 0029, Yudong Zhang 0005, Kun Wang 0056, Yang Wang 0015 |
ICLR | 1 |
| 2024 | Make Bricks with a Little Straw: Large-Scale Spatio-Temporal Graph Learning with Restricted GPU-Memory Capacity
Binwu Wang, Pengkun Wang 0001, Zhengyang Zhou, Zhe Zhao 0008, Wei Xu 0055, Yang Wang 0015 |
IJCAI | 1 |
| 2024 | STONE: A Spatio-temporal OOD Learning Framework Kills Both Spatial and Temporal ShiftsabstractTraffic prediction is a crucial task in the Intelligent Transportation System (ITS), receiving significant attention from both industry and academia. Numerous spatio-temporal graph convolutional networks have emerged for traffic prediction and achieved remarkable success. However, these models have limitations in terms of generalization and scalability when dealing with Out-of-Distribution (OOD) graph data with both structural and temporal shifts. To tackle the challenges of spatio-temporal shift, we propose a framework called STONE by learning invariable node dependencies, which achieve stable performance in variable environments. STONE initially employs gated-transformers to extract spatial and temporal semantic graphs. These two kinds of graphs represent spatial and temporal dependencies, respectively. Then we design three techniques to address spatio-temporal shifts. Firstly, we introduce a Fréchet embedding method that is insensitive to structural shifts, and this embedding space can integrate loose position dependencies of nodes within the graph. Secondly, we propose a graph intervention mechanism to generate multiple variant environments by perturbing two kinds of semantic graphs without any data augmentations, and STONE can explore invariant node representation from environments. Finally, we further introduce an explore-to-extrapolate risk objective to enhance the variety of generated environments. We conduct experiments on multiple traffic datasets, and the results demonstrate that our proposed model exhibits competitive performance in terms of generalization and scalability. Binwu Wang, Jiaming Ma, Pengkun Wang 0001, Xu Wang 0029, Yudong Zhang 0005, Zhengyang Zhou, Yang Wang 0015 |
KDD | 1 |
| 2024 | LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed ProblemsabstractThe long-tailed distribution is the underlying nature of real-world data, and it presents unprecedented challenges for training deep learning models. Existing long-tailed learning paradigms based on re-balancing or data augmentation have partially alleviated the long-tailed problem. However, they still have limitations, such as relying on manually designed augmentation strategies, having a limited search space, and using fixed augmentation strategies. To address these limitations, this paper proposes a novel LLM-based long-tailed data augmentation framework called LLM-AutoDA, which leverages large-scale pretrained models to automatically search for the optimal augmentation strategies suitable for long-tailed data distributions. In addition, it applies this strategy to the original imbalanced data to create an augmented dataset and fine-tune the underlying long-tailed learning model. The performance improvement on the validation set serves as a reward signal to update the generation model, enabling the generation of more effective augmentation strategies in the next iteration. We conducted extensive experiments on multiple mainstream long-tailed learning benchmarks. The results show that LLM-AutoDA outperforms state-of-the-art data augmentation methods and other re-balancing methods significantly. Pengkun Wang 0001, Zhe Zhao 0008, Haibin Wen, Fanfu Wang, Binwu Wang, Qingfu Zhang 0001, Yang Wang 0015 |
NeurIPS | 5 |
| 2024 | When Imbalance Meets Imbalance: Structure-driven Learning for Imbalanced Graph ClassificationabstractGraph Neural Networks (GNNs) can learn representative graph-level features to achieve efficient graph classification. But GNNs usually assume an environment where both class and structure distribution are balanced. Although previous works have considered the graph classification problem under the scenario of class imbalance or structure imbalance, they habitually ignored the obvious fact that class imbalance and structural imbalance are often intertwined in the real world. In this paper, we propose a carefully designed structure-driven learning framework called ImbGNN to address the potential intertwined class imbalance and structural imbalance in graph classification. Specifically, we find that feature-oriented augmentation (e.g., feature masking) and structure-oriented augmentation (e.g., edge perturbation) will have differential impacts when applied to different graphs. Therefore, we design optional augmentation based on the average degree distribution to alleviate structural imbalance. Furthermore, based on the imbalance of graph size distribution, we utilize a similarity-friendly graph random walk to extract a core subgraph to improve the accuracy of graph kernel similarity calculation, and then construct a more reasonable kernel-based graph of graphs, thereby alleviating the class imbalance and size imbalance. Extensive experiments on multiple benchmark datasets demonstrate that our proposed ImbGNN framework outperforms previous baselines on imbalanced graph classification tasks. The code of ImbGNN is available in~https://github.com/Xiaovy/ImbGNN. Wei Xu 0055, Pengkun Wang 0001, Zhe Zhao 0008, Binwu Wang, Xu Wang 0029, Yang Wang 0015 |
WWW | 4 |
| 2024 | Meta Koopman decomposition for time series forecasting under temporal distribution shifts
Yudong Zhang 0005, Xu Wang 0029, Zhaoyang Sun, Pengkun Wang 0001, Binwu Wang, Yang Wang 0015 |
Adv. Eng. Informatics | 5 |
| 2024 | Face Anti-Spoofing with Unknown Attacks: A Comprehensive Feature Extraction and Representation Perspective
Li-Min Li, Binwu Wang, Xu Wang 0029, Pengkun Wang 0001, Yudong Zhang 0005, Yang Wang 0015 |
J. Comput. Sci. Technol. | 2 |
| 2024 | Adaptive and Interactive Multi-Level Spatio-Temporal Network for Traffic ForecastingabstractTraffic forecasting is a challenging research topic due to the complex spatial and temporal dependencies among different roads. Though great efforts have been made on traffic forecasting, existing works still have the following shortcomings: i) Most methods only directly perform on the original road network topology which cannot accommodate the diverse traffic patterns and multi-granularity traffic forecasting requirements driven by the natural multi-level urban structure and layout, ii) The existing studies based on the spatio-temporal multi-granularity perspective ignore the interactions between the fine-grained information and coarse-grained information, resulting in the spatio-temporal correlation under multi-granularity inaccurately modeled. To solve the problems, we propose an Adaptive and Interactive Multi-level Spatio-Temporal network (AIMST) for traffic forecasting. Specifically, we first devise a learnable adaptive hierarchical clustering method to automatically generate more coarse-grained graphs from the initial road networks and the traffic data. Then, the spatio-temporal graph convolutional networks are executed on the constructed hierarchical traffic graph of each level correspondingly to capture the spatio-temporal patterns. Furthermore, a multi-level bidirectional interaction module is designed to emphasize the multi-grained interaction patterns among different levels. Extensive experiments on two real-world traffic datasets demonstrate that our framework is superior to several state-of-the-art baselines. Yudong Zhang 0005, Pengkun Wang 0001, Binwu Wang, Xu Wang 0029, Zhe Zhao 0008, Zhengyang Zhou, Lei Bai 0001, Yang Wang 0015 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Modeling Spatio-Temporal Mobility Across Data Silos via Personalized Federated LearningabstractSpatio-temporal mobility modeling plays a pivotal role in the advancement of mobile computing. Nowadays, data is frequently held by various distributed silos, which are isolated from each other and confront limitations on data sharing. Given this, there have been some attempts to introduce federated learning into spatio-temporal mobility modeling. Meanwhile, the distributional heterogeneity inherent in the spatio-temporal data also puts forward requirements for model personalization. However, the existing methods tackle personalization in a model-centric manner and fail to explore the data characteristics in various data silos, thus ignoring the fact that the fundamental cause of insufficient personalization in the model is the heterogeneous distribution of data. In this paper, we propose a novel distribution-oriented personalizedFederated learning framework forCross-siloSpatio-Temporal mobility modeling (namedFedCroST), that leverages learnable spatio-temporal prompts to implicitly represent the local data distribution patterns of data silos and guide the local models to learn the personalized information. Specifically, we focus on the potential characteristics within temporal distribution and devise a conditional diffusion module to generate temporal prompts that serve as guidance for the evolution of the time series. Simultaneously, we emphasize the structure distribution inherent in node neighborhoods and propose adaptive spatial structure partition to construct the spatial prompts, augmenting the spatial information representation. Furthermore, we introduce a denoising autoencoder to effectively harness the learned multi-view spatio-temporal features and obtain personalized representations adapted to local tasks. Our proposal highlights the significance of latent spatio-temporal data distributions in enabling personalized federated spatio-temporal learning, providing new insights into modeling spatio-temporal mobility in data silo scenarios. Extensive experiments conducted on real-world datasets demonstrate that FedCroST outperforms the advanced baselines by a large margin in diverse cross-silo spatio-temporal mobility modeling tasks. Yudong Zhang 0005, Xu Wang 0029, Pengkun Wang 0001, Binwu Wang, Zhengyang Zhou, Yang Wang 0015 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Predicting Collective Human Mobility via Countering Spatiotemporal HeterogeneityabstractHuman mobility forecasting is the key to energizing considerable mobile computing services. However, we find that the collective mobility suffers the spatiotemporal heterogeneity issue and therefore leads to inferior performances of conventional homogeneous aggregations. Given two fundamental factors, i.e., data and objectives in machine learning, we propose to counter such heterogeneity by improving data utilization and optimization objectives. 1) From data utilization perspective, we discover that such heterogeneity is inherently induced by mobility-related context factors and thus these factors can be exploited to learn heterogeneous mobility patterns. 2) From the optimization perspective, the dependencies among output elements, which give another prior to learning, can extract heterogeneous correlations within output sequences. Specifically, we propose a novel Context-Directional SpatioTemporal Graph Network (CD-STGNet), which tackles the above-mentioned heterogeneity, for achieving accurate mobility predictions. Firstly, we improve data utilization by inputting the encoded context-wise interactions to a direction field learner, which realizes directional spatial aggregations. Secondly, regarding series learning and optimization objectives, a context-trend highway is designed to enable context-aware temporal learning while two regularization objectives are proposed to keep the correlations among predicted elements consistent with the ground-truth. Experiments demonstrate that CD-STGNet surpasses competitive baselines by 13% to 22% and boosts the interpretability of context-directional learning. Zhengyang Zhou, Kuo Yang 0002, Yuxuan Liang 0002, Binwu Wang, Hongyang Chen 0001, Yang Wang 0015 |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Long-Tailed Time Series Classification via Feature Space Rebalancing
Pengkun Wang 0001, Xu Wang 0029, Binwu Wang, Yudong Zhang 0005, Lei Bai 0001, Yang Wang 0015 |
DASFAA (1) | 3 |
| 2023 | A Knowledge-Driven Memory System for Traffic Flow Prediction
Binwu Wang, Yudong Zhang 0005, Pengkun Wang 0001, Xu Wang 0029, Lei Bai 0001, Yang Wang 0015 |
DASFAA (4) | 1 |
| 2023 | Pattern Expansion and Consolidation on Evolving Graphs for Continual Traffic PredictionabstractRecently, spatiotemporal graph convolutional networks are becoming popular in the field of traffic flow prediction and significantly improve prediction accuracy. However, the majority of existing traffic flow prediction models are tailored to static traffic networks and fail to model the continuous evolution and expansion of traffic networks. In this work, we move to investigate the challenge of traffic flow prediction on an expanding traffic network. And we propose an efficient and effective continual learning framework to achieve continuous traffic flow prediction without the access to historical graph data, namely Pattern Expansion and Consolidation based on Pattern Matching based (PECPM). Specifically, we first design a pattern bank based on pattern matching to store representative patterns of the road network. With the expansion of the road network, the model configured with such a bank module can achieve continuous traffic prediction by effectively managing patterns stored in the bank. The core idea is to continuously update new patterns while consolidating learned ones. Specifically, we design a pattern expansion mechanism that can detect evolved and new patterns from the updated network, then these unknown patterns are expanded into the pattern bank to adapt to the updated road network. Additionally, we propose a pattern consolidation mechanism that includes both a bank preservation mechanism and a pattern traceability mechanism. This can effectively consolidate the learned patterns in the bank without requiring access to detailed historical graph data. We construct experiments on real-world traffic datasets to demonstrate the competitive performance, superior efficiency, and strong generalization ability of PECPM. Binwu Wang, Yudong Zhang 0005, Xu Wang 0029, Pengkun Wang 0001, Zhengyang Zhou, Lei Bai 0001, Yang Wang 0015 |
KDD | 1 |
| 2023 | An Observed Value Consistent Diffusion Model for Imputing Missing Values in Multivariate Time SeriesabstractMissing values, which are common in multivariate time series, is most important obstacle towards the utilization and interpretation of those data. Great efforts have been employed on how to accurately impute missing values in multivariate time series, and existing works either use deep learning networks to achieve deterministic imputations or aim at generating different plausible imputations by sampling multiple noises from a same distribution and then denoising them. However, these models either fall short of modeling the uncertainties of imputations due to their deterministic nature or perform poorly in terms of interpretability and imputation accuracy due to their ignorance of the correlations between the latent representations of both observed and missing values which are parts of samples from a same distribution. To this end, in this paper, we explicitly take the correlations between observed and missing values into account, and theoretically re-derive the Evidence Lower BOund (ELBO) of conditional diffusion model in the scenario of multivariate time series imputation. Based on the newly derived ELBO, we further propose a novel multivariate imputation diffusion model (MIDM) which is equipped with novel noise sampling, adding and denoising mechanisms for multivariate time series imputation, and the series of newly designed technologies jointly ensure the involving of the consistency between observed and missing values. Extensive experiments on both the tasks of multivariate time series imputation and forecasting witness the superiority of our proposed MIDM model on generating conditional estimations. Xu Wang 0029, Pengkun Wang 0001, Yudong Zhang 0005, Binwu Wang, Zhengyang Zhou, Yang Wang 0015 |
KDD | 5 |
| 2023 | CrossGNN: Confronting Noisy Multivariate Time Series Via Cross Interaction RefinementabstractRecently, multivariate time series (MTS) forecasting techniques have seen rapid development and widespread applications across various fields. Transformer-based and GNN-based methods have shown promising potential due to their strong ability to model interaction of time and variables. However, by conducting a comprehensive analysis of the real-world data, we observe that the temporal fluctuations and heterogeneity between variables are not well handled by existing methods. To address the above issues, we propose CrossGNN, a linear complexity GNN model to refine the cross-scale and cross-variable interaction for MTS. To deal with the unexpected noise in time dimension, an adaptive multi-scale identifier (AMSI) is leveraged to construct multi-scale time series with reduced noise. A Cross-Scale GNN is proposed to extract the scales with clearer trend and weaker noise. Cross-Variable GNN is proposed to utilize the homogeneity and heterogeneity between different variables. By simultaneously focusing on edges with higher saliency scores and constraining those edges with lower scores, the time and space complexity (i.e., $O(L)$) of CrossGNN can be linear with the input sequence length $L$. Extensive experimental results on 8 real-world MTS datasets demonstrate the effectiveness of CrossGNN compared with state-of-the-art methods. Qihe Huang, Shouhong Ding, Binwu Wang, Zhengyang Zhou, Yang Wang 0015 |
NeurIPS | 5 |
| 2023 | Towards Learning in Grey Spatiotemporal Systems: A Prophet to Non-consecutive Spatiotemporal DynamicsabstractSpatiotemporal forecasting is an imperative topic in data science due to its critical applications in smart cities. Existing works mostly perform consecutive predictions of following steps with observations continuously obtained, where nearest observations can be exploited as the key knowledge for status estimation. However, the practical issues of early activity planning and sensor failures elicit a new task, non-consecutive forecasting. In this paper, we define spatiotemporal learning systems with missing observations as Grey Spatiotemporal Systems (G2S) and propose a Factor-Decoupled learning framework for G2S to hierarchically decouple multi-level factors, and enable flexible aggregations with uncertainty estimations. We especially select representative sequences to capture periodicity and instantaneous variations, and infer the non-consecutive future statuses under expected exogenous factors, compensating the missing observations. Given the inherent incompleteness and critical applications of G2S, a DisEntangled Uncertainty Quantification is put forward, to identify two types of uncertainty for model interpretations and robustness promotions. Experiments demonstrate that our solution can promote the performance by at least 8.50% on early planning and 2.01%-18.00% on sensor failures. The appendix of this paper can be found at https://github.com/zzyy0929/SDM-G2S. Zhengyang Zhou, Kuo Yang 0002, Binwu Wang, Yunan Zong, Yang Wang 0015 |
SDM | 4 |
| 2023 | Knowledge Expansion and Consolidation for Continual Traffic Prediction With Expanding GraphsabstractAccurate traffic prediction plays a vital role in intelligent transport managements and applications. However, in the vast majority of existing works, the focus is mainly on modeling spatiotemporal correlations in static traffic networks. Thus, the continuous expansion and evolution of traffic networks are ignored. In this work, we study the problem of traffic prediction with expanding road network structures under the continual learning paradigm. Considering the model prediction performance, efficiency, and data accessibility, a SpatioTemporal Knowledge Expansion and Consolidation (STKEC) framework is proposed. This framework contains an influence-based knowledge expansion strategy to help the spatiotemporal learning model integrate new spatiotemporal traffic patterns and a memory-augmented knowledge consolidation mechanism to preserve the learned spatiotemporal patterns without accessing the data in previous graphs. Extensive experiments are conducted on a large-scale dataset and verify the superior performance of STKEC in continual traffic prediction. Binwu Wang, Yudong Zhang 0005, Pengkun Wang 0001, Xu Wang 0029, Lei Bai 0001, Yang Wang 0015 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Countering Modal Redundancy and Heterogeneity: A Self-Correcting Multimodal FusionabstractFusing multimodal heterogeneous data plays a vital role in recognition and prediction tasks in various fields, e.g., action recognition and traffic accident forecast. Yet, there remain some key challenges, such as heterogeneous feature interaction and feature redundancies, that significantly affect the performance of multimodal fusion. To tackle these challenges, we first devise a Unified Feature Interaction Module (UFIM) in which a novel orthogonal attention component is designed to obtain fine-grained inter-modal interaction information among heterogeneous features. Then, we propose a novel Self-Correcting Transformer Module (SCTM) which employs a modified transformer to obtain the one-to-many correlation information between the current modal feature and the merged features of other modalities to alleviate the redundancy problem. Extensive experiments on four cross-domain tasks demonstrate the effectiveness and generalization ability of our proposed method. Pengkun Wang 0001, Xu Wang 0029, Binwu Wang, Yudong Zhang 0005, Lei Bai 0001, Yang Wang 0015 |
ICDM | 3 |
| 2022 | CMT-Net: A Mutual Transition Aware Framework for Taxicab Pick-ups and Drop-offs Co-PredictionabstractWith increasing population of modern cities, accurate estimation of regional passenger demands is critical to online taxicab services as such platforms aim at a reformation of taxicab scheduling for a more efficient order dispatching. Though great efforts have been made on passenger demand predictions, existing works still have the following shortcomings: i) they mostly performed based on uniform grid partition, which results in the imbalance of demand volumes among regions and even non-vehicle regions in such partition, ii) none of previous demand forecasting efforts have highlighted the important mutual influences between pick-ups and drop-offs, which are of great significance for taxicab scheduling. To this end, we first devise a multi-kernel based clustering to achieve a taxicab-behavior and geographic-aware sub-region partition, hence a more balanced and compact regional division is obtained. Subsequently, we emphasize the essential factors with regard to mutual transition quantification in taxicab predictions, then propose a Transfer-LSTM and an Origin-Destination-based transition matrix to respectively capture the drop-to-pick and pick-to-drop spatiotemporal transition patterns. Hence, a novel mutual-transition-aware co-prediction framework is devised by capturing complex spatiotemporal interactions between pick-ups and drop-offs. Extensive experiments on two real-world taxicab datasets demonstrate our co-prediction framework is superior to state-of-the-art methods, thus providing novel perspectives to urban human mobility understanding and transition-based taxicab scheduling. Yudong Zhang 0005, Binwu Wang, Ziyang Shan, Zhengyang Zhou, Yang Wang 0015 |
WSDM | 2 |