Chen Zhang 0007

dblp:94/4084-7 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0002-4767-9597ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 AgentAsk: Multi-Agent Systems Need to Ask
abstract
Bohan Lin, Kuo Yang, Zelin Tan, Yingchuan Lai, Chen Zhang, Guibin Zhang, Xinlei Yu, Miao Yu, Xu Wang, Yudong Zhang, Yang Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Bohan Lin, Kuo Yang 0002, Zelin Tan, Yingchuan Lai, Chen Zhang 0007, Guibin Zhang, Xu Wang 0029, Yudong Zhang 0001, Yang Wang 0015
ACL (1)5
2026 Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
abstract
Zelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang, Guancheng Wan, Yifan Zhou, Qiang He, Xiangyuan Xue, Heng Zhou, Yutao Fan, Zhong-Zhi Li, Zaibin Zhang, Guibin Zhang, Chen Zhang, Zhenfei Yin, Philip Torr, Lei Bai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang, Guancheng Wan, Xiangyuan Xue, Yutao Fan, Zhongzhi Li, Zaibin Zhang, Guibin Zhang, Chen Zhang 0007, Zhenfei Yin, Philip Torr 0001, Lei Bai 0001
ACL (1)14
2026 Adaptive Change Detection in Partially Observable Dynamic Networks
abstract
Sequential change detection in high-dimensional dynamic networks has attracted growing attention in modern applications. A major challenge is that as network scale increases, limited sensing resources make it difficult to fully observe links at each time point, resulting in partial observability and complicating detection. To tackle this, we propose a Latent Bi-Transit Network (LBTN) that learns unobserved edge formation, latent node states, and the evolution mechanisms of dynamic networks. Based on LBTN, we design a variational Bayesian method to infer sparse node-level changes and employ a likelihood ratio test as the detection statistic. By further formulating the statistic as the reward function of a combinatorial multi-armed bandit (CMAB) problem, we develop a Thompson sampling strategy to adaptively select edges for observation, balancing exploration and exploitation. Comprehensive simulations and real-world experiments show that our method consistently outperforms existing baselines and its variants across diverse scenarios.
Haijie Xu, Chen Zhang 0007
IEEE Trans. Knowl. Data Eng.3
2026 Addressing Glaucoma Structure-Function Relationship: A Multi-Task Learning Framework With Multi-Modal and Unpaired Data
abstract
Glaucoma, an irreversible neurodegenerative disorder, can lead to vision loss and blindness. Visual field (VF) tests are crucial for quantifying functional damage in glaucoma, but the tests are time-consuming and the results have high variations, influenced by subjective behaviors and psychological states of patients. Consequently, predicting VF test results using objective, non-invasive, reproducible optical coherence tomography (OCT) data coupled with deep learning modeling is promising in improving clinical care. However, existing methods only focus on predicting single VF indicators, such as threshold sensitivities or deviation maps, and have poor performance in severe glaucoma. Since different VF indicators are correlated, developing a joint prediction model is beneficial. Furthermore, the availability of more VF test data than corresponding OCTs in most datasets poses a challenge in utilizing unpaired VF test data. This study proposes a multi-modal, multi-task learning framework based on OCT data for VF prediction. We introduce a dynamic weighted loss function to improve prediction performance for eyes with severe glaucoma. Additionally, we construct a novel PairMatcher model for augmenting unpaired VF data. Extensive experiments demonstrate that our framework outperforms existing methods, showcasing its potential for VF prediction in glaucoma.
Xuming An 0002, Jacqueline Chua, Ruben Hemelings, Rahat Husain, Rachel Chong, Tina Wong, Tin Aung, Damon Wing Kee Wong, Chen Zhang 0007, Leopold Schmetterer
IEEE Trans. Medical Imaging10
2025 COFlowNet: Conservative Constraints on Flows Enable High-Quality Candidate Generation
abstract
Generative flow networks (GFlowNets) have been considered as powerful tools for generating candidates with desired properties. Given that evaluating the property of candidates can be complex and time-consuming, existing GFlowNets train proxy models for efficient online evaluation. However, the performance of proxy models is heavily dependent on the amount of data and is of considerable uncertainty. Therefore, it is of great interest that how to develop an offline GFlowNet that does not rely on online evaluation. Under the offline setting, the limited data results in an insufficient exploration of state space. The insufficient exploration means that offline GFlowNets can hardly generate satisfying candidates out of the distribution of training data. Therefore, it is critical to restrict the offline model to act in the distribution of training data. The distinctive training goal of GFlownets poses a unique challenge for making such restrictions. Tackling the challenge, we propose Conservative Offline GFlowNet (COFlowNet) in this paper. We define unsupported flow, edges containing unseen states in training data. Models can learn extremely little knowledge about unsupported flow from training data. By constraining the model from exploring unsupported flows, we restrict COFlowNet to explore as optimal trajectories on the training set as possible, thus generating better candidates. In order to improve the diversity of candidates, we further introduce a quantile version of unsupported flow restriction. Experimental results on several widely-used datasets validate the effectiveness of COFlowNet in generating high-scored and diverse candidates. All implementations are available at https://github.com/yuxuan9982/COflownet.
Yudong Zhang 0005, Xu Wang 0029, Zhaoyang Sun, Chen Zhang 0007, Pengkun Wang 0001, Yang Wang 0015
ICLR5
2025 TraffiDent: A Dataset for Understanding the Interplay Between Traffic Dynamics and Incidents
abstract
Long-separated research has been conducted on two highly correlated tracks: traffic and incidents. Traffic track witnesses complicating deep learning models, e.g., to push the prediction a few percent more accurate, and the incident track only studies the incidents alone, e.g., to infer the incident risk. We, for the first time, spatiotemporally aligned the two tracks in a large-scale region (16,972 traffic nodes) from year 2022 to 2024: our TraffiDent dataset includes traffic, i.e., time-series indexes on traffic flow, lane occupancy, and average vehicle speed, and incident, whose records are spatiotemporally aligned with traffic data, with seven different incident classes. Additionally, each node includes detailed physical and policy-level meta-attributes of lanes. Previous datasets typically contain only traffic or incident data in isolation, limiting research to general forecasting tasks. TraffiDent integrates both, enabling detailed analysis of traffic-incident interactions and causal relationships. To demonstrate its broad applicability, we design: (1) post-incident traffic forecasting to quantify the impact of different incidents on traffic indexes; (2) incident classification using traffic indexes to determine the incidents types for precautions measures; (3) global causal analysis among the traffic indexes, meta-attributes, and incidents to give high-level guidance of the interrelations of various factors; (4) local causal analysis within road nodes to examine how different incidents affect the road segments' relations. The dataset is available at https://xaitraffic.github.io.
Xiaochuan Gou, Ziyue Li 0002, Junpeng Lin, Zhishuai Li, Chen Zhang 0007, Di Wang 0015, Xiangliang Zhang 0001
NeurIPS7
2025 SPEVS-CC: Separated parameter estimation with variable selection based on canonical correlation analysis for multivariate functional regression
abstract
The rapid advancement of real-time data monitoring technologies has established functional data analysis as a crucial predictive tool across diverse applications. Despite progress in multivariate functional regression and principal component analysis, significant challenges persist in function-to-function regression and feature selection. These include the inaccurate selection of predictor variables from extensive predictors and flawed parameter estimation, which compromise the precision of function-based data predictions. This paper introduces the SPEVS-CC method, a novel canonical correlation-based feature selection technique specifically designed for function-on-function regression. By effectively decoupling variable selection from model fitting, the SPEVS-CC method enhances both adaptability and interpretability. Validated through rigorous experiments and real-world applications, including in the Shenzhen subway system, semiconductor manufacturing, and ocean climate, SPEVS-CC significantly reduces mean squared error, confirming its robustness and practical utility. This methodological breakthrough harmonizes variable selection with functional regression, providing unmatched interpretability and usability in industrial applications.
Xing Yang 0003, Haijie Xu, Yiming Shi, Chen Zhang 0007
Adv. Eng. Informatics4
2025 Multi-Scenario Cellular KPI Prediction Based on Spatiotemporal Graph Neural Network
abstract
With the increasing demand for high-quality telecommunication services, cellular KPI prediction becomes crucial for telecommunication network monitoring and management. In this work, we propose a novel framework for cellular KPI prediction, which considers its distribution discrepancy under different network operation scenarios. In particular, three specific predictors for normal, target alarm, and neighbor alarm scenarios are proposed based on spatiotemporal graph neural networks and unified through transfer learning. Temporal convolution and attention mechanism are embedded to model the impact of anomalies on KPIs and its propagation across neighboring cells according to the cellular network topology. An experiment on a real cellular KPI dataset shows the effectiveness of the proposed method compared to the state-of-the-arts. Note to Practitioners—Cellular network KPI prediction under scenarios of network alarms is crucial to evaluate the impact of alarms on network services and guides cellular network maintenance policies. This problem is similar to a general multivariate time series prediction problem with data multimodality. However, the first challenge in our case is that, under different scenarios, i.e., normal, target alarm, and neighbor alarm, the effective information and spatiotemporal dependencies among KPIs are different. The second challenge is the imbalanced or sparse sample size for specific scenarios, deteriorating the model performance. This paper proposes a cellular KPI prediction framework consisting of three scenario-specific predictors with similar but different modules to process different scenario-specific data. To address the dataset imbalance across scenarios, we adopt a transfer learning strategy to unify the training and prediction of three predictors. The experiment results on a real cellular KPI dataset demonstrate that the proposed framework is more feasible and effective than the state-of-the-art models for multivariate time series prediction. Future research can consider developing maintenance policies such that the cost caused by abnormal KPIs can be minimized.
Junpeng Lin, Dandan Miao, Huiru He, Jiantao Ye, Chen Zhang 0007, Yan-Fu Li
IEEE Trans Autom. Sci. Eng.8
2025 Nonlinear Causal Discovery via Dynamic Latent Variables
abstract
Distinguishing causality from mere correlation is a cornerstone in empirical research, as conflating the two can result in significant errors in decision-making, affecting policy formulation and the validity of scientific inferences. Traditional experimental designs, such as randomized trials, often fall short in complex systems where variables interact in a high-dimensional space with limited data. This paper aims to address these challenges by introducing an innovative causal discovery approach, extending beyond conventional methodologies by incorporating algorithmic advances in computational efficiency and design. We present a novel double Gaussian process state space causal model (GPSSCM) that contends with the multifaceted nature of causal inference, accounting for noisy observations and latent variables, which are commonly encountered in dynamic systems. Our methodological contribution includes the application of a Markov chain Monte Carlo technique for unraveling latent state dynamics and an expectation-maximization (EM) algorithm for robust parameter estimation. The acyclic nature of the causal graph is ensured through an integrated acyclic constraint within the EM framework, maintaining the integrity of the causal model. The efficacy of our proposed GPSSCM is evaluated through a series of tests on both synthetic data and empirical case studies from the industrial domain. The results highlight the model’s capacity to accurately infer complex nonlinear causal relationships, demonstrating its superiority over traditional structural equation modeling, especially when dealing with time series data and latent variables. This paper not only contributes a sophisticated tool for researchers and practitioners but also enriches the literature on causal discovery by offering a new perspective on the analysis of intricate systems, thereby facilitating more informed and ethical decision-making across various scientific fields. Note to Practitioners—Understanding the intricate web of causality is crucial for making informed decisions in various scientific and professional fields. Our study presents a Gaussian process state space causal model, which enhances the analysis of dynamic causal relationships in complex systems, particularly when dealing with noisy observations and latent variables. Leveraging a combination of Markov chain Monte Carlo and expectation maximization algorithms, the model ensures accurate estimation of parameters and causal structures. This paper is particularly relevant for those in fields such as economics, transportation, and biology, offering a sophisticated tool to support ethical decision-making and safeguard against errors in high-stakes environments. The practical implications of this research are underscored by its ability to inform targeted interventions and predict outcomes under new conditions, advancing the comprehension and application of causality in real-world scenarios.
Xing Yang 0003, Chen Zhang 0007
IEEE Trans Autom. Sci. Eng.4
2024 NondBREM: Nondeterministic Offline Reinforcement Learning for Large-Scale Order Dispatching
abstract
One of the most important tasks in ride-hailing is order dispatching, i.e., assigning unserved orders to available drivers. Recent order dispatching has achieved a significant improvement due to the advance of reinforcement learning, which has been approved to be able to effectively address sequential decision-making problems like order dispatching. However, most existing reinforcement learning methods require agents to learn the optimal policy by interacting with environments online, which is challenging or impractical for real-world deployment due to high costs or safety concerns. For example, due to the spatiotemporally unbalanced supply and demand, online reinforcement learning-based order dispatching may significantly impact the revenue of the ride-hailing platform and passenger experience during the policy learning period. Hence, in this work, we develop an offline deep reinforcement learning framework called NondBREM for large-scale order dispatching, which learns policy from only the accumulated logged data to avoid costly and unsafe interactions with the environment. In NondBREM, a Nondeterministic Batch-Constrained Q-learning (NondBCQ) module is developed to reduce the algorithm extrapolation error and a Random Ensemble Mixture (REM) module that integrates multiple value networks with multi-head networks is utilized to improve the model generalization and robustness. Extensive experiments on large-scale real-world ride-hailing datasets show the superiority of our design.
Guang Wang 0001, Xu Wang 0029, Zhengyang Zhou, Chen Zhang 0007, Zheng Dong 0002, Yang Wang 0015
AAAI5
2024 Fine-Grained Passenger Load Prediction inside Metro Network via Smart Card Data
abstract
Metro system serves as the backbone for urban public transportation. Accurate passenger load prediction for the metro system plays a crucial role in metro service quality improvement, such as helping operators schedule train timetables and passengers plan their trips. However, existing works can only predict low‐grained passenger flows of origin‐destination (O‐D) paths or inflows/outflows of each station but cannot predict passenger load distribution over the whole metro network. To this end, this paper proposes an end‐to‐end inference framework, PIPE, for passenger load prediction of every metro segment between two adjacent stations, by only utilizing smart card data. In particular, PIPE includes two modules. The first is the core. It formulates the travel time distribution of each metro segment as a truncated Gaussian distribution. Since there might be several possible routes for certain O‐D paths, the population‐level travel time distribution of these O‐D paths would be a mixture of travel times of different routes. Considering the route preference may change over time, a dynamic truncated Gaussian mixture model is proposed for parameter inference of each truncated Gaussian distribution of each metro segment. The second module serves as the supplement, which compiles a bunch of methods for predicting passenger flows of O‐D paths. Built upon them, PIPE is able to predict the travel time that future passengers of each O‐D path will take for passing each metro segment and consequently can predict the passenger load of each metro segment in the short future. Numerical studies from Singapore’s metro system demonstrate the efficacy of our method.
Xiancai Tian, Chen Zhang 0007, Baihua Zheng
Int. J. Intell. Syst.2
2024 Coupled Epidemic-Information Propagation With Stranding Mechanism on Multiplex Metapopulation Networks
abstract
Acknowledging the significance of information propagation and individual adaptive behavior has been regarded as an indispensable prerequisite for a complete understanding of epidemic spreading. Recent studies have widely considered the metapopulation model, where epidemics spread over a single layer of physical networks via individual mobility. However, these advances neglected the interventions of accompanied information and individual behavior response related to epidemics. In this article, we develop a coupled epidemic-information propagation model on multiplex metapopulation networks leveraging the microscopic Markov chain (MMC) approach, aiming to explore the spatiotemporal characteristics of epidemic spreading process. Taking the individual adaptive behavior into account, the stranding mechanism based on infection level and medical resources is introduced to capture the population size dynamics during individual mobility among different patches. Theoretical epidemic threshold is analytically derived under the improved framework. Extensive numerical simulations are performed to validate our theoretical analysis and further examine the impacts of information propagation and spreading parameters on epidemic threshold and steady-state prevalence. Our results indicate that both the scale of information diffusion and the specific configuration of spreading parameters can significantly suppress the epidemic prevalence. These findings shed a novel light on theoretical research and decision-making of coupled epidemic-information process in the spatiotemporal perspective.
Xuming An 0002, Chen Zhang 0007, Lin Hou 0003, Kaibo Wang
IEEE Trans. Comput. Soc. Syst.2
2024 Heterogeneous Multivariate Functional Time Series Modeling: A State Space Approach
abstract
Functional data have been gaining increasing popularity in the field of time series analysis. However, so far modeling heterogeneous multivariate functional time series remains a research gap. To fill it, this paper proposes a time-varying functional state space model (TV-FSSM). It uses functional decomposition to extract features of the functional observations, where the decomposition coefficients are regarded as latent states that evolve according to a tensor autoregressive model. This two-layer structure can on the one hand efficiently extract continuous functional features, and on the other provide a flexible and generalized description of data heterogeneity among different time points. An expectation maximization (EM) framework is developed for parameter estimation, where regularization and constraints are incorporated for better model interoperability. As the sample size grows, an incremental learning version of the EM algorithm is given to efficiently update the model parameters. Some model properties, including model identifiability conditions, convergence issues, time complexities, and bounds of its one-step-ahead prediction errors, are also presented. Extensive experiments on both real and synthetic datasets are performed to evaluate the predictive accuracy and efficiency of the proposed framework.
Junpeng Lin, Chen Zhang 0007
IEEE Trans. Knowl. Data Eng.3
2023 MM-DAG: Multi-task DAG Learning for Multi-modal Data - with Application for Traffic Congestion Analysis
abstract
This paper proposes to learn Multi-task, Multi-modal Direct Acyclic Graphs (MM-DAGs), which are commonly observed in complex systems, e.g., traffic, manufacturing, and weather systems, whose variables are multi-modal with scalars, vectors, and functions. This paper takes the traffic congestion analysis as a concrete case, where a traffic intersection is usually regarded as a DAG. In a road network of multiple intersections, different intersections can only have someoverlapping and distinct variables observed. For example, a signalized intersection has traffic light-related variables, whereas unsignalized ones do not. This encourages the multi-task design: with each DAG as a task, the MM-DAG tries to learn the multiple DAGs jointly so that their consensus and consistency are maximized. To this end, we innovatively propose a multi-modal regression for linear causal relationship description of different variables. Then we develop a novel Causality Difference (CD) measure and its differentiable approximator. Compared with existing SOTA measures, CD can penalize the causal structural difference among DAGs with distinct nodes and can better consider the uncertainty of causal orders. We rigidly prove our design's topological interpretation and consistency properties. We conduct thorough simulations and one case study to show the effectiveness of our MM-DAG. The code is available under https://github.com/Lantian72/MM-DAG.
Ziyue Li 0002, Zhishuai Li, Lei Bai 0001, Man Li 0003, Fugee Tsung, Wolfgang Ketter, Rui Zhao 0001, Chen Zhang 0007
KDD9
2023 Multi-view metro station clustering based on passenger flows: a functional data-edged network community detection approach
Chen Zhang 0007, Baihua Zheng, Fugee Tsung
Data Min. Knowl. Discov.1
2023 A Cluster-Oriented Bayesian Network Approach for Mixed-Type Event Prediction With Application in Order Logistics
abstract
Temporal representation and reasoning for probabilistic events describe temporal causal relationships between events. This has been widely used in several applications to predict events accurately. However, there are two challenges: the occurrence time points of events may have distinct types, distributed randomly or at several fixed time points; only limited historical data are available in some cases. This article presents a mixed-type event prediction algorithm based on a cluster-oriented Bayesian network (BN) model to address the highlighted challenges. The proposed model categorizes events as random events or timing events based on their temporal features. The similarity between events is measured according to event types and features. A clustering algorithm for events is further implemented to help reduce the model size and build a simpler and more accurate BN. The experimental results show that the proposed model significantly improves performance under small data sizes.
Xing Yang 0003, Chen Zhang 0007
IEEE Trans. Ind. Informatics3
2023 Deep Cascade-Learning Model via Recurrent Attention for Immunofixation Electrophoresis Image Analysis
abstract
Immunofixation Electrophoresis (IFE) analysis has been an indispensable prerequisite for the diagnosis of M-protein, which is an important criterion to recognize diversified plasma cell diseases. Existing intelligent methods of IFE diagnosis commonly employ a single unified classifier to directly classify whether M-protein exists and which isotype of M-protein is. However, this unified classification is not optimal because the two tasks have different characteristics and require different feature extraction techniques. Classifying the M-protein existence depends on the presence or absence of dense bands in IFE data, while classifying the M-protein isotype depends on the location of dense bands. Consequently, a cascading two-classifier framework suitable to the two tasks respectively may achieve better performance. In this paper, we propose a novel deep cascade-learning model, which sequentially integrates a positive-negative classifier based on deep collocative learning and an isotype classifier based on recurrent attention model to address these two tasks respectively. Specifically, the attention mechanism can mimic the visual perception of clinicians, where only the most informative local regions are extracted through sequential partial observations. This not only avoids the interference of redundant regions but also saves computational power. Further, domain knowledge about SP lane and heavy-light-chain lanes is also introduced to assist our attention location. Extensive numerical experiments show that our deep cascade-learning outperforms state-of-the-art methods on recognized evaluation metrics and can effectively capture the co-location of dense bands in different lanes.
Xuming An 0002, Pengchang Li, Chen Zhang 0007
IEEE Trans. Medical Imaging3
2023 A Hidden Markov Model for Condition Monitoring of Time Series Data in Complex Network Systems
abstract
Time series data are ubiquitous in complex network systems, where each component of the system is treated as a vertex and has sequentially collected autocorrelated observations for condition monitoring. In many cases, components transit between different states, such as “normal,” “degrading,” “failure,” etc., and under different states their distributions vary, resulting in the mixture marginal distribution of each vertex's time series data. Moreover, states of different components influence each other through the system's network topology, i.e., the state of a component itself and its neighbors jointly affect its data distribution. For efficient condition monitoring, it is important to capture the state evolution of all vertices as a whole over time. However, the state-switching may be unobservable, which complicates the modeling. Hence, this article proposes a hidden Markov model for networked time series data with mixture marginal distributions. The Markov transition rule for latent states captures the state-switching behavior and the AR model reveals the temporal dependence of vertices. The states of different vertices further influence each other through the network topology structure. The Baum–Welch algorithm is used to estimate the parameters of the proposed model, allowing for state inference and data fitting. Extensive numerical studies as well as real case studies demonstrate the effectiveness and applicability of the proposed model.
Wanshan Li, Chen Zhang 0007
IEEE Trans. Reliab.2
2023 Deep Reinforcement Learning for Dynamic Opportunistic Maintenance of Multi-Component Systems With Load Sharing
abstract
Opportunistic maintenance (OM), which shows its superiority on complex multi-component systems by integrating the maintenance activities of multiple components to reduce the maintenance cost, has been widely studied over the past decade. To our knowledge, most of the existing OM works are developed based on fixed maintenance thresholds without fully utilizing the health state of the multi-component system. This article presents an OM optimization problem of multi-component systems with load sharing, solved by a modified proximal policy optimization approach based on deep reinforcement learning algorithm. The load sharing effect is reflected in the hazard rate function, which further changes the failure probability of the components. Meanwhile, the health states can be recovered by executing imperfect maintenance and corrective maintenance. The optimization problem is formulated as an infinite-horizon MDP with mixed discrete and continuous state and action space to maximize the total discounted reward, taking into account the system reliability and the maintenance cost. The difficulty caused by the mixed action space is solved by designing a parameterized action space structure and multi-task reinforcement learning framework. The effectiveness of the proposed algorithm is tested on a four-component system and a real-world scenario configured with the high-pressure feedwater heater system in the nuclear power plant. The results show that the performance of the algorithm is stable when facing large-scale problems. The algorithm proposed in this study also contributes to the imperfect maintenance optimization with state-of-the-art optimization techniques.
Chen Zhang 0007, Yan-Fu Li, David W. Coit
IEEE Trans. Reliab.1
2022 GRELEN: Multivariate Time Series Anomaly Detection from the Perspective of Graph Relational Learning
abstract
System monitoring and anomaly detection is a crucial task in daily operation. With the rapid development of cyber-physical systems and IT systems, multiple sensors get involved to represent the system state from different perspectives, which inspires us to detect anomalies considering feature dependence relationship among sensors instead of focusing on individual sensor's behavior. In this paper, we propose a novel Graph Relational Learning Network (GReLeN) to detect multivariate time series anomalies from the perspective of between-sensor dependence relationship learning. Variational AutoEncoder (VAE) serves as the overall framework for feature extraction and system representation. Graph Neural Network (GNN) and stochastic graph relational learning strategy are also imposed to capture the between-sensor dependence. Then a composite anomaly metric is established with the learned dependence structure explicitly. The experiments on four real-world datasets show our superiority in detection accuracy, anomaly diagnosis, and model interpretation.
Chen Zhang 0007, Fugee Tsung
IJCAI2
2022 Individualized passenger travel pattern multi-clustering based on graph regularized tensor latent dirichlet allocation
abstract
Abstract Individual passenger travel patterns have significant value in understanding passenger’s behavior, such as learning the hidden clusters of locations, time, and passengers. The learned clusters further enable commercially beneficial actions such as customized services, promotions, data-driven urban-use planning, peak hour discovery, and so on. However, the individualized passenger modeling is very challenging for the following reasons: 1) The individual passenger travel data are multi-dimensional spatiotemporal big data, including at least the origin, destination, and time dimensions; 2) Moreover, individualized passenger travel patterns usually depend on the external environment, such as the distances and functions of locations, which are ignored in most current works. This work proposes a multi-clustering model to learn the latent clusters along the multiple dimensions of Origin, Destination, Time, and eventually, Passenger (ODT-P). We develop a graph-regularized tensor Latent Dirichlet Allocation (LDA) model by first extending the traditional LDA model into a tensor version and then applies to individual travel data. Then, the external information of stations is formulated as semantic graphs and incorporated as the Laplacian regularizations; Furthermore, to improve the model scalability when dealing with massive data, an online stochastic learning method based on tensorized variational Expectation-Maximization algorithm is developed. Finally, a case study based on passengers in the Hong Kong metro system is conducted and demonstrates that a better clustering performance is achieved compared to state-of-the-arts with the improvement in point-wise mutual information index and algorithm convergence speed by a factor of two.
Ziyue Li 0002, Chen Zhang 0007, Fugee Tsung
Data Min. Knowl. Discov.3
2022 A Data-Driven Method for Online Monitoring Tube Wall Thinning Process in Dynamic Noisy Environment
abstract
Tube internal erosion, which corresponds to its wall thinning process, is one of the major safety concerns for tubes. Many sensing technologies have been developed to detect a tube wall thinning process. Among them, fiber Bragg grating (FBG) sensors are the most popular ones due to their precise measurement properties. Most of the current works focus on how to design different types of FBG sensors according to certain physical laws and only test their sensors in controlled laboratory conditions. However, in practice, an industrial system usually suffers from harsh and dynamic environmental conditions, and FBG signals are affected by many unpredictable factors. Consequently, the FBG signals have more fluctuations and are polluted by noises. Hence, the signals no longer directly follow the assumed physical laws and their proposed thinning detection mechanisms no longer work. Targeting at this, this article develops a data-driven model for FBG signal feature extraction and tube wall thickness monitoring using data analytic techniques. In particular, we develop a spatiotemporal model to describe dynamic FBG signals and extract features related to thickness. By taking physical law as guideline, we trace the relationship between the extracted features and the tube wall thickness, based on which we construct an online statistical monitoring scheme for tube wall thinning process. We use both laboratory test and field trial experiment to demonstrate the efficacy and efficiency of the proposed scheme.Note to Practitioners—This article is motivated by the real industrial needs of inner erosion detection of tubes in harsh environment. Most of the current research works focus on designing various sensing apparatuses based on fiber Bragg grating (FBG) sensors for nondestructive erosion detection. These apparatuses prove to be able to collect signals reflecting tube wall thickness in static and controllable laboratory environment qualitatively. However, in reality, the industrial environment, which is impacted by many changing factors, is dynamic and uncontrollable. Consequently, the signals collected by these FBG apparatuses would have larger variations that mask the signals related to thickness. Furthermore, current methods have neither mentioned how to process their collected data to capture the unnoticeably slow but accumulative erosion information efficiently nor constructed online monitoring algorithms to detect the tube wall thinning process based on the collected signals quantitatively. Built upon their apparatuses but targeting at their unsolved challenges, we propose a novel data-driven approach for FBG signal analysis that can remove the environmental influence and extract features only related to tube wall thickness, and using the extracted features, we construct a statistical process control scheme to monitor tube wall thickness and detect erosion in real time efficiently.
Chen Zhang 0007, Jun Long Lim, Ouyang Liu, Aayush Madan, Yongwei Zhu, Shili Xiang, Kai Wu 0004, Rebecca Yen-Ni Wong, Eugene Phua Jiliang, Karan M. Sabnani, Keng Boon Siah, Emily Hao Jianzhong, Steven C. H. Hoi
IEEE Trans Autom. Sci. Eng.1
2022 Segment-Wise Time-Varying Dynamic Bayesian Network with Graph Regularization
abstract
Time-varying dynamic Bayesian network (TVDBN) is essential for describing time-evolving directed conditional dependence structures in complex multivariate systems. In this article, we construct a TVDBN model, together with a score-based method for its structure learning. The model adopts a vector autoregressive (VAR) model to describe inter-slice and intra-slice relations between variables. By allowing VAR parameters to change segment-wisely over time, the time-varying dynamics of the network structure can be described. Furthermore, considering some external information can provide additional similarity information of variables. Graph Laplacian is further imposed to regularize similar nodes to have similar network structures. The regularized maximum a posterior estimation in the Bayesian inference framework is used as a score function for TVDBN structure evaluation, and the alternating direction method of multipliers (ADMM) with L-BFGS-B algorithm is used for optimal structure learning. Thorough simulation studies and a real case study are carried out to verify our proposed method’s efficacy and efficiency.
Xing Yang 0003, Chen Zhang 0007, Baihua Zheng
ACM Trans. Knowl. Discov. Data2
2021 Holistic Prediction for Public Transport Crowd Flows: A Spatio Dynamic Graph Network Approach
Bingjie He, Chen Zhang 0007, Baihua Zheng, Fugee Tsung
ECML/PKDD (1)3
2020 Tensor Completion for Weakly-Dependent Data on Graph for Metro Passenger Flow Prediction
abstract
Low-rank tensor decomposition and completion have attracted significant interest from academia given the ubiquity of tensor data. However, low-rank structure is a global property, which will not be fulfilled when the data presents complex and weak dependencies given specific graph structures. One particular application that motivates this study is the spatiotemporal data analysis. As shown in the preliminary study, weakly dependencies can worsen the low-rank tensor completion performance. In this paper, we propose a novel low-rank CANDECOMP / PARAFAC (CP) tensor decomposition and completion framework by introducing the L1-norm penalty and Graph Laplacian penalty to model the weakly dependency on graph. We further propose an efficient optimization algorithm based on the Block Coordinate Descent for efficient estimation. A case study based on the metro passenger flow data in Hong Kong is conducted to demonstrate an improved performance over the regular tensor completion methods.
Ziyue Li 0002, Nurettin Sergin, Chen Zhang 0007, Fugee Tsung
AAAI4
2020 Time-Warped Sparse Non-negative Factorization for Functional Data Analysis
abstract
This article proposes a novel time-warped sparse non-negative factorization method for functional data analysis. The proposed method on the one hand guarantees the extracted basis functions and their coefficients to be positive and interpretable, and on the other hand is able to handle weakly correlated functions with different features. Furthermore, the method incorporates time warping into factorization and hence allows the extracted basis functions of different samples to have temporal deformations. An efficient framework of estimation algorithms is proposed based on a greedy variable selection approach. Numerical studies together with case studies on real-world data demonstrate the efficacy and applicability of the proposed methodology.
Chen Zhang 0007, Steven C. H. Hoi, Fugee Tsung
ACM Trans. Knowl. Discov. Data1
2019 Partially Observable Multi-Sensor Sequential Change Detection: A Combinatorial Multi-Armed Bandit Approach
abstract
This paper explores machine learning to address a problem of Partially Observable Multi-sensor Sequential Change Detection (POMSCD), where only a subset of sensors can be observed to monitor a target system for change-point detection at each online learning round. In contrast to traditional Multisensor Sequential Change Detection tasks where all the sensors are observable, POMSCD is much more challenging because the learner not only needs to detect on-the-fly whether a change occurs based on partially observed multi-sensor data streams, but also needs to cleverly choose a subset of informative sensors to be observed in the next learning round, in order to maximize the overall sequential change detection performance. In this paper, we present the first online learning study to tackle POMSCD in a systemic and rigorous way. Our approach has twofold novelties: (i) we attempt to detect changepoints from partial observations effectively by exploiting potential correlations between sensors, and (ii) we formulate the sensor subset selection task as a Multi-Armed Bandit (MAB) problem and develop an effective adaptive sampling strategy using MAB algorithms. We offer theoretical analysis for the proposed online learning solution, and further validate its empirical performance via an extensive set of numerical studies together with a case study on real-world data sets.
Chen Zhang 0007, Steven C. H. Hoi
AAAI1