Chengqing Yu

dblp:264/6276 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-8314-8251ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 APT: Affine Prototype-Timestamp for Time Series Forecasting Under Distribution Shift
abstract
Time series forecasting under distribution shift remains challenging, as existing deep learning models often rely on local statistical normalization (e.g., mean and variance) that fails to capture global distribution shift. Methods like RevIN and its variants attempt to decouple distribution and pattern but still struggle with missing values, noisy observations, and invalid channel-wise affine transformation. To address these limitations, we propose Affine Prototype-Timestamp(APT), a lightweight and flexible plug-in module that injects global distribution features into the normalization–forecasting pipeline. By leveraging timestamp-conditioned prototype learning, APT dynamically generates affine parameters that modulate both input and output series, enabling the backbone to learn from self-supervised, distribution-aware clustered instances. APT is compatible with arbitrary forecasting backbones and normalization strategies while introducing minimal computational overhead. Extensive experiments across six benchmark datasets and multiple backbone-normalization combinations demonstrate that APT significantly improves forecasting performance under distribution shift.
Yujie Li 0008, Zezhi Shao, Chengqing Yu, Yisong Fu, Tao Sun 0011, Yongjun Xu 0001, Fei Wang 0014
AAAI3
2026 A new feature reconstruction method and multilabel ensemble strategy for non-intrusive load recognition
Chengming Yu, Hui Liu 0023, Chengqing Yu, Guangxi Yan, Zijie Cao
Knowl. Based Syst.3
2025 Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual Recognition
abstract
Multi-teacher Knowledge Distillation (KD) transfers diverse knowledge from a teacher pool to a student network. The core problem of multi-teacher KD is how to balance distillation strengths among various teachers. Most existing methods often develop weighting strategies from an individual perspective of teacher performance or teacher-student gaps, lacking comprehensive information for guidance. This paper proposes Multi-Teacher Knowledge Distillation with Reinforcement Learning (MTKD-RL) to optimize multi-teacher weights. In this framework, we construct both teacher performance and teacher-student gaps as state information to an agent. The agent outputs the teacher weight and can be updated by the return reward from the student. MTKD-RL reinforces the interaction between the student and teacher using an agent in an RL-based decision mechanism, achieving better matching capability with more meaningful weights. Experimental results on visual recognition tasks, including image classification, object detection, and semantic segmentation tasks, demonstrate that MTKD-RL achieves state-of-the-art performance compared to the existing multi-teacher KD works.
Chuanguang Yang, Xinqiang Yu, Zhulin An, Chengqing Yu, Libo Huang 0001, Yongjun Xu 0001
AAAI5
2025 STA-GANN: A Valid and Generalizable Spatio-Temporal Kriging Approach
abstract
Spatio-temporal tasks often encounter incomplete data arising from missing or inaccessible sensors, making spatio-temporal kriging crucial for inferring the completely missing temporal information. However, current models struggle with ensuring the validity and generalizability of inferred spatio-temporal patterns, especially in capturing dynamic spatial dependencies and temporal shifts, and optimizing the generalizability of unknown sensors. To overcome these limitations, we propose Spatio-Temporal Aware Graph Adversarial Neural Network (STA-GANN), a novel GNN-based kriging framework that improves spatio-temporal pattern validity and generalization. STA-GANN integrates (i) Decoupled Phase Module that senses and adjusts for timestamp shifts. (ii) Dynamic Data-Driven Metadata Graph Modeling to update spatial relationships using temporal data and metadata; (iii) An adversarial transfer learning strategy to ensure generalizability. Extensive validation across nine datasets from four fields and theoretical evidence both demonstrate the superior performance of STA-GANN.
Yujie Li 0008, Zezhi Shao, Chengqing Yu, Tangwen Qian, Zhao Zhang 0011, Yifan Du 0004, Shaoming He, Fei Wang 0014, Yongjun Xu 0001
CIKM3
2025 BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models
abstract
The advent of universal time series forecasting models has revolutionized zero-shot forecasting across diverse domains, yet the critical role of data diversity in training these models remains underexplored. Existing large-scale time series datasets often suffer from inherent biases and imbalanced distributions, leading to suboptimal model performance and generalization. To address this gap, we introduce BLAST, a novel pre-training corpus designed to enhance data diversity through a balanced sampling strategy. First, BLAST incorporates 321 billion observations from publicly available datasets and employs a comprehensive suite of statistical metrics to characterize time series patterns. Then, to facilitate pattern-oriented sampling, the data is implicitly clustered using grid-based partitioning. Furthermore, by integrating grid sampling and grid mixup techniques, BLAST ensures a balanced and representative coverage of diverse patterns. Experimental results demonstrate that models pre-trained on BLAST achieve state-of-the-art performance with a fraction of the computational resources and training tokens required by existing methods. Our findings highlight the pivotal role of data diversity in improving both training efficiency and model performance for the universal forecasting task.
Zezhi Shao, Yujie Li 0008, Fei Wang 0014, Chengqing Yu, Yisong Fu, Tangwen Qian, Bin Xu 0019, Boyu Diao, Yongjun Xu 0001, Xueqi Cheng 0001
KDD (2)4
2025 Merlin: Multi-View Representation Learning for Robust Multivariate Time Series Forecasting with Unfixed Missing Rates
abstract
Multivariate Time Series Forecasting (MTSF) involves predicting future values of multiple interrelated time series. Recently, deep learning-based MTSF models have gained significant attention for their promising ability to mine semantics (global and local information) within MTS data. However, these models are pervasively susceptible to missing values caused by malfunctioning data collectors. These missing values not only disrupt the semantics of MTS, but their distribution also changes over time. Nevertheless, existing models lack robustness to such issues, leading to suboptimal forecasting performance. To this end, in this paper, we propose Multi-View Representation Learning (Merlin), which can help existing models achieve semantic alignment between incomplete observations with different missing rates and complete observations in MTS. Specifically, Merlin consists of two key modules: offline knowledge distillation and multi-view contrastive learning. The former utilizes a teacher model to guide a student model in mining semantics from incomplete observations, similar to those obtainable from complete observations. The latter improves the student model's robustness by learning from positive/negative data pairs constructed from incomplete observations with different missing rates, ensuring semantic alignment across different missing rates. Therefore, Merlin is capable of effectively enhancing the robustness of existing models against unfixed missing rates while preserving forecasting accuracy. Experiments on four real-world datasets demonstrate the superiority of Merlin.
Chengqing Yu, Fei Wang 0014, Chuanguang Yang, Zezhi Shao, Tao Sun 0011, Tangwen Qian, Wei Wei 0002, Zhulin An, Yongjun Xu 0001
KDD (2)1
2025 Selective Learning for Deep Time Series Forecasting
abstract
Benefiting from high capacity for capturing complex temporal patterns, deep learning (DL) has significantly advanced time series forecasting (TSF). However, deep models tend to suffer from severe overfitting due to the inherent vulnerability of time series to noise and anomalies. The prevailing DL paradigm uniformly optimizes all timesteps through the MSE loss and learns those uncertain and anomalous timesteps without difference, ultimately resulting in overfitting. To address this, we propose a novel selective learning strategy for deep TSF. Specifically, selective learning screens a subset of the whole timesteps to calculate the MSE loss in optimization, guiding the model to focus on generalizable timesteps while disregarding non-generalizable ones. Our framework introduces a dual-mask mechanism to target timesteps: (1) an uncertainty mask leveraging residual entropy to filter uncertain timesteps, and (2) an anomaly mask employing residual lower bound estimation to exclude anomalous timesteps. Extensive experiments across eight real-world datasets demonstrate that selective learning can significantly improve the predictive performance for typical state-of-the-art deep models, including 37.4% MSE reduction for Informer, 8.4% for TimesNet, and 6.5% for iTransformer.
Yisong Fu, Zezhi Shao, Chengqing Yu, Yujie Li 0008, Zhulin An, Cheems Wang, Yongjun Xu 0001, Fei Wang 0014
NeurIPS3
2025 On the Integration of Spatial-Temporal Knowledge: A Lightweight Approach to Atmospheric Time Series Forecasting
abstract
Transformers have gained attention in atmospheric time series forecasting (ATSF) for their ability to capture global spatial-temporal correlations. However, their complex architectures lead to excessive parameter counts and extended training times, limiting their scalability to large-scale forecasting. In this paper, we revisit ATSF from a theoretical perspective of atmospheric dynamics and uncover a key insight: spatial-temporal position embedding (STPE) can inherently model spatial-temporal correlations even without attention mechanisms. Its effectiveness arises from integrating geographical coordinates and temporal features, which are intrinsically linked to atmospheric dynamics. Based on this, we propose **STELLA**, a **S**patial-**T**emporal knowledge **E**mbedded **L**ightweight mode**L** for ASTF, utilizing only STPE and an MLP architecture in place of Transformer layers. With 10k parameters and one hour of training, STELLA achieves superior performance on five datasets compared to other advanced methods. The paper emphasizes the effectiveness of spatial-temporal knowledge integration over complex architectures, providing novel insights for ATSF.
Yisong Fu, Fei Wang 0014, Zezhi Shao, Boyu Diao, Lin Wu 0006, Zhulin An, Chengqing Yu, Yujie Li 0008, Yongjun Xu 0001
NeurIPS7
2025 Heterogeneity in Multivariate Time Series: Comprehensive Analysis and Adaptive Modeling
abstract
Multivariate time series (MTS) data are ubiquitous in complex dynamic systems such as meteorology, transportation, and energy.However, data heterogeneity caused by cross-domain variations has become a central bottleneck restricting model generalization and consistency in comparative studies.This paper systematically reviews recent MTS forecasting research, revealing that inconsistencies in experimental conclusions primarily arise from neglecting substantial differences in data distributions and characteristics.To address this issue, we introduce BasicTS, an fair and scalable benchmark designed to fairly quantify the impact of heterogeneity on model performance.Subsequently, to tackle generalization challenges posed by heterogeneity, this tutorial proposes two adaptive solutions: (i) developing BLAST, a balanced and diversity-enhanced pre-training corpus that explicitly models heterogeneity, significantly improving zero-shot general forecasting; and (ii) introducing ARIES, a relational assessment and model recommendation framework that leverages a statistical pattern-to-model matching mechanism to automatically select optimal forecasting models for specific real-world sequences.Through comprehensive experiments and case studies, we demonstrate that precisely characterizing and leveraging data heterogeneity, beyond mere model design, is crucial for improving the robustness of MTS forecasting.This research provides methodological guidance and practical insights for academia and industry to fully exploit the value of time series data and make data-driven decisions.
Zezhi Shao, Chengqing Yu, Fei Wang 0014
SSTD2
2025 Distributed Traffic Signal Control Model for Accurate Policy Learning Under Dynamic Traffic Flow: A Graph Forecast-State Vector Driven Deep Reinforcement Learning Framework
abstract
The increasing severity of traffic congestion has driven the application of deep reinforcement learning (DRL) in traffic signal control. It utilizes spatial states to interact with environments and develop control policy, which has been extensively studied and demonstrated to be efficient. However, the periodic fluctuations in traffic flow may produce states with the same spatial state but different trends. Relying solely on spatial states for interactions will mix up interactive experiences of these states, and bring instability to control policy learning for dynamic traffic flows. Moreover, many studies concentrate on vehicles as the main focus of traffic signal control, overlooking the demands of pedestrians, who are also vital participants in traffic. Therefore, this study proposes a distributed DRL traffic signal control method Graph Forecast-Vector Dueling Deep Q Network (GF-VDDQN) based on state vector. This method combines graph neural networks prediction and deep reinforcement learning, with the key idea being construction of state vector for agents to understand environments and learn policy accurately. Considering the principle of ‘people-oriented’, a comprehensive reward with green duration loss and vehicle waiting time is designed. A continuous duration control action set is introduced to enhance intersection utilization. Additionally, the distributed GF-VDDQN allows agents to collaborate with each other. To evaluate performances of GF-VDDQN, we conduct experiments in the SUMO traffic simulator, comparing it with advanced and traditional traffic signal control methods. Results show that state vector can assist agents in accurately understanding environments and avoiding experience mixture problem. GF-VDDQN performs superior than other methods.
Fuhao Yu, Xiwei Mi, Chengqing Yu, Yuanting Jiang
IEEE Trans. Intell. Transp. Syst.3
2025 Exploring Progress in Multivariate Time Series Forecasting: Comprehensive Benchmarking and Heterogeneity Analysis
abstract
Multivariate Time Series (MTS) analysis is crucial to understanding and managing complex systems, such as traffic and energy systems, and a variety of approaches to MTS forecasting have been proposed recently. However, we often observe inconsistent or seemingly contradictory performance findings across different studies. This hinders our understanding of the merits of different approaches and slows down progress. We address the need for means of assessing MTS forecasting proposals reliably and fairly, in turn enabling better exploitation of MTS as seen in different applications. Specifically, we first propose BasicTS+, a benchmark designed to enable fair, comprehensive, and reproducible comparison of MTS forecasting solutions. BasicTS+ establishes a unified training pipeline and reasonable settings, enabling an unbiased evaluation. Second, we identify the heterogeneity across different MTS as an important consideration and enable classification of MTS based on their temporal and spatial characteristics. Disregarding this heterogeneity is a prime reason for difficulties in selecting the most promising technical directions. Third, we apply BasicTS+ along with rich datasets to assess the capabilities of more than 30 MTS forecasting solutions. This provides readers with an overall picture of the cutting-edge research on MTS forecasting.
Zezhi Shao, Fei Wang 0014, Yongjun Xu 0001, Wei Wei 0002, Chengqing Yu, Zhao Zhang 0011, Di Yao 0001, Tao Sun 0011, Guangyin Jin, Xin Cao 0001, Gao Cong, Christian S. Jensen, Xueqi Cheng 0001
IEEE Trans. Knowl. Data Eng.5
2025 GinAR+: A Robust End-to-End Framework for Multivariate Time Series Forecasting With Missing Values
abstract
Spatial-Temporal Graph Neural Networks (STGNNs) have been widely utilized in multivariate time series forecasting (MTSF), but they rely on the assumption of data completeness. In practice, due to factors such as natural disaster, STGNNs frequently encounter the challenge of missing data resulting from numerous malfunctioning data collectors. In this case, on the one hand, due to the presence of missing values, STGNNs easily generate incorrect spatial correlations, leading to the performance degradation. On the other hand, STGNNs require separate training of models for different missing rates, limiting their robustness. To address these challenges, we first propose two important components (interpolation attention and adaptive graph convolution), which utilize normal values to recover missing values into reliable representations and reconstruct spatial correlations. Then, we replace the fully connected layers in simple recursive units with these two components and propose Graph Interpolation Attention Recursive Network (GinAR), aiming to recursively correct spatial correlations and achieve end-to-end MTSF with missing values. Finally, we use data with different missing rates as positive and negative data pairs. By employing contrastive learning to train GinAR, we propose GinAR+ and enhance its robustness to data with different missing rates. Experiments validate the superiority of GinAR+ and our motivation.
Chengqing Yu, Fei Wang 0014, Zezhi Shao, Tangwen Qian, Zhao Zhang 0011, Wei Wei 0002, Zhulin An, Qi Wang 0025, Yongjun Xu 0001
IEEE Trans. Knowl. Data Eng.1
2024 Online Policy Distillation with Decision-Attention
abstract
Policy Distillation (PD) has become an effective method to improve deep reinforcement learning tasks. The core idea of PD is to distill policy knowledge from a teacher agent to a student agent. However, the teacher-student framework requires a well-trained teacher model which is computationally expensive. In the light of online knowledge distillation, we study the knowledge transfer between different policies that can learn diverse knowledge from the same environment. In this work, we propose Online Policy Distillation (OPD) with Decision-Attention (DA), an online learning framework in which different policies operate in the same environment to learn different perspectives of the environment and transfer knowledge to each other to obtain better performance together. With the absence of a well-performance teacher policy, the group-derived targets play a key role in transferring group knowledge to each student policy. However, naive aggregation functions tend to cause student policies quickly homogenize. To address the challenge, we introduce the Decision-Attention module to the online policies distillation framework. The Decision-Attention module can generate a distinct set of weights for each policy to measure the importance of group members. We use the Atari platform for experiments with various reinforcement learning algorithms, including PPO and DQN. In different tasks, our method can perform better than an independent training policy on both PPO and DQN algorithms. This suggests that our OPD-DA can transfer knowledge between different policies well and help agents obtain more rewards.
Xinqiang Yu, Chuanguang Yang, Chengqing Yu, Libo Huang 0001, Zhulin An, Yongjun Xu 0001
IJCNN3
2024 GinAR: An End-To-End Multivariate Time Series Forecasting Model Suitable for Variable Missing
abstract
Multivariate time series forecasting (MTSF) is crucial for decision-making to precisely forecast the future values/trends, based on the complex relationships identified from historical observations of multiple sequences. Recently, Spatial-Temporal Graph Neural Networks (STGNNs) have gradually become the theme of MTSF model as their powerful capability in mining spatial-temporal dependencies, but almost of them heavily rely on the assumption of historical data integrity. In reality, due to factors such as data collector failures and time-consuming repairment, it is extremely challenging to collect the whole historical observations without missing any variable. In this case, STGNNs can only utilize a subset of normal variables and easily suffer from the incorrect spatial-temporal dependency modeling issue, resulting in the degradation of their forecasting performance. To address the problem, in this paper, we propose a novel Graph Interpolation Attention Recursive Network (named GinAR) to precisely model the spatial-temporal dependencies over the limited collected data for forecasting. In GinAR, it consists of two key components, that is, interpolation attention and adaptive graph convolution to take place of the fully connected layer of simple recursive units, and thus are capable of recovering all missing variables and reconstructing the correct spatial-temporal dependencies for recursively modeling of multivariate time series data, respectively. Extensive experiments conducted on five real-world datasets demonstrate that GinAR outperforms 11 SOTA baselines, and even when 90% of variables are missing, it can still accurately predict the future values of all variables.
Chengqing Yu, Fei Wang 0014, Zezhi Shao, Tangwen Qian, Zhao Zhang 0011, Wei Wei 0002, Yongjun Xu 0001
KDD1
2024 WGformer: A Weibull-Gaussian Informer based model for wind speed prediction
Ziyi Shi, Zheyuan Jiang, Chengqing Yu, Xiwei Mi
Eng. Appl. Artif. Intell.5
2024 Semi-supervised anomaly detection with contamination-resilience and incremental training
Liheng Yuan, Fanghua Ye 0001, Heng Li 0008, Cuiying Gao, Chengqing Yu, Wei Yuan 0001, Xinge You
Eng. Appl. Artif. Intell.6
2024 MRIformer: A multi-resolution interactive transformer for wind speed multi-step prediction
Chengqing Yu, Guangxi Yan, Chengming Yu, Xiwei Mi
Inf. Sci.1
2023 Clustering-property Matters: A Cluster-aware Network for Large Scale Multivariate Time Series Forecasting
abstract
Large-scale Multivariate Time Series(MTS) widely exist in various real-world systems, imposing significant demands on model efficiency. A recent work, STID, addressed the high complexity issue of popular Spatial-Temporal Graph Neural Networks(STGNNs). Despite its success, when applied to large-scale MTS data, the number of parameters of STID for modeling spatial dependencies increases substantially, leading to over-parameterization issues and suboptimal performance. These observations motivate us to explore new approaches for modeling spatial dependencies in a parameter-friendly manner. In this paper, we argue that the spatial properties of variables are essentially the superposition of multiple cluster centers. Accordingly, we propose a Cluster-Aware Network(CANet), which effectively captures spatial dependencies by mining the implicit cluster centers of variables. CANet solely optimizes the cluster centers instead of the spatial information of all nodes, thereby significantly reducing the parameter amount. Extensive experiments on two large-scale datasets validate our motivation and demonstrate the superiority of CANet.
Yuan Wang 0037, Zezhi Shao, Tao Sun 0011, Chengqing Yu, Yongjun Xu 0001, Fei Wang 0014
CIKM4
2023 DSformer: A Double Sampling Transformer for Multivariate Time Series Long-term Prediction
abstract
Multivariate time series long-term prediction, which aims to predict the change of data in a long time, can provide references for decision-making. Although transformer-based models have made progress in this field, they usually do not make full use of three features of multivariate time series: global information, local information, and variables correlation. To effectively mine the above three features and establish a high-precision prediction model, we propose a double sampling transformer (DSformer), which consists of the double sampling (DS) block and the temporal variable attention (TVA) block. Firstly, the DS block employs down sampling and piecewise sampling to transform the original series into feature vectors that focus on global information and local information respectively. Then, TVA block uses temporal attention and variable attention to mine these feature vectors from different dimensions and extract key information. Finally, based on a parallel structure, DSformer uses multiple TVA blocks to mine and integrate different features obtained from DS blocks respectively. The integrated feature information is passed to the generative decoder based on a multi-layer perceptron to realize multivariate time series long-term prediction. Experimental results on nine real-world datasets show that DSformer can outperform eight existing baselines.
Chengqing Yu, Fei Wang 0014, Zezhi Shao, Tao Sun 0011, Lin Wu 0006, Yongjun Xu 0001
CIKM1
2020 A novel axle temperature forecasting method based on decomposition, reinforcement learning optimization and neural network
Hui Liu 0023, Chengming Yu, Chengqing Yu, Haiping Wu
Adv. Eng. Informatics3