Zhichao Chen 0001

dblp:16/5141-1 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0001-5785-0741ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Toward Intrinsically Calibrated Uncertainty Quantification in Industrial Data-Driven Models via Diffusion Sampler
abstract
In modern process industries, data-driven models are important tools for real-time monitoring when key performance indicators are difficult to measure directly. While accurate predictions are essential, reliable uncertainty quantification (UQ) is equally critical for safety, reliability, and decision-making, but remains a major challenge in current data-driven approaches. In this work, we introduce a diffusion-based posterior sampling framework that inherently produces well-calibrated predictive uncertainty via faithful posterior sampling, eliminating the need for post hoc calibration. In extensive evaluations on synthetic distributions, the Raman-based phenylacetic acid soft sensor benchmark, and a real ammonia synthesis case study, our method achieves practical improvements over existing UQ techniques in both uncertainty calibration and predictive accuracy. These results highlight diffusion samplers as a principled and scalable paradigm for advancing uncertainty-aware modeling in industrial applications.
Yiran Ma, Jerome Le Ny, Zhichao Chen 0001
IEEE Trans. Ind. Informatics3
2026 Slack More, Predict Better: Proximal Relaxation for Probabilistic Latent Variable Model-Based Soft Sensors
abstract
Nonlinear probabilistic latent variable models (NPLVMs) are a cornerstone of soft sensor modeling due to their capacity for uncertainty delineation. However, conventional NPLVMs are trained using amortized variational inference, where neural networks parameterize the variational posterior. While facilitating model implementation, this parameterization converts the distributional optimization problem within an infinite-dimensional function space to parameter optimization within a finite-dimensional parameter space, which introduces an approximation error gap, thereby degrading soft sensor modeling accuracy. To alleviate this issue, we introduce KProxNPLVM, a novel NPLVM that pivots to relaxing the objective itself and improving the NPLVM's performance. Specifically, we first prove the approximation error induced by the conventional approach. Based on this, we design the Wasserstein distance as the proximal operator to relax the learning objective, yielding a new variational inference strategy derived from solving this relaxed optimization problem. Based on this foundation, we provide a rigorous derivation of KProxNPLVM's optimization implementation, prove the convergence of our algorithm can finally sidestep the approximation error, and propose the KProxNPLVM by summarizing the abovementioned content. Finally, extensive experiments on synthetic and real-world industrial datasets are conducted to demonstrate the efficacy of the proposed KProxNPLVM.
Zehua Zou, Yiran Ma, Yulong Zhang 0005, Zhengnan Li, Jinhao Xie, Zhichao Chen 0001
IEEE Trans. Ind. Informatics8
2026 Blending Data and Knowledge for Process Industrial Modeling Under Riemannian Preconditioned Bayesian Framework
abstract
Integrating graph neural networks (GNNs) with variational inference (VI) provides a promising direction for blending structured prior knowledge with observational empirical data for data-driven industrial process modeling. However, this task requires inference of the normalized adjacency matrix (NAM), where each row is normalized to be non-negative and to sum to one, matching the support of Dirichlet distribution. This requirement presents two main technical challenges: 1) intractable Kullback-Leibler (KL) divergence optimization between Dirichlet distributions, and 2) constrained optimization for standard- gradient-descent-based neural network parameter optimization. To handle issue 1), we first formulate the inference of the NAM as a differential equation simulation problem and derive an easy-to-implement expression to iteratively improve the KL divergence without explicitly computing it. Based on this, to alleviate issue 2), we involve Riemannian optimization to precondition this simulation procedure, which ensures that the inferred NAM conforms to the row-normalization constraint. After that, we collectively designate these approaches for NAM inference as Preconditioned-Simulation-Induced Variational Inference ($\psi$-VI), and provide theoretical guarantees of convergence. On this foundation, we propose a new graph neural network architecture, the Preconditioned-Simulation-Induced-based Variational Graph Neural Network ($\psi$-VGNN) for industrial process modeling. Finally, we validate the efficacy of$\psi$-VGNN through comprehensive experiments on industrial modeling tasks.
Zhichao Chen 0001, Yulong Zhang 0005, Odin Zhang, Fangyikang Wang, Le Yao, Hao Wang 0049
IEEE Trans. Knowl. Data Eng.1
2026 Robust Missing Value Imputation With Proximal Optimal Transport for Low-Quality IIoT Data
abstract
Accurate imputation of missing data is crucial in the Industrial Internet-of-Things (IIoT), where operations are often compromised by noisy samples from harsh environments. Traditional imputation methods struggle with such noise due to their black-box nature or lack of adaptability. To address this issue, we recast data imputation as a distribution alignment challenge, utilizing the flexibility of optimal transport (OT) to handle noisy samples. Specifically, we first introduce the Proximal Optimal Transport (POT) problem, where the transportation cost is obtained by the network simplex approach with a selective matching mechanism, which renders it capable of matching distributions with noisy samples. Subsequently, we propose the POT-I framework, where the objective is to minimize the transport cost of POT. The produced gradient is used to refine the imputation value, which achieves missing data imputation (MDI) while getting robustness to noisy samples. Experiments on real-world IIoT datasets demonstrate the superiority of POT-I over state-of-the-art imputation methods.
Hao Wang 0049, Zhichao Chen 0001, Degui Yang, Dangjun Zhao, Buge Liang
IEEE Trans. Neural Networks Learn. Syst.2
2025 FreDF: Learning to Forecast in the Frequency Domain
abstract
Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation within historical data, the correlations among future labels are often overlooked. Specifically, modern forecasting models primarily adhere to the Direct Forecast (DF) paradigm, generating multi-step forecasts independently and disregarding label correlations over time. In this work, we demonstrate that the learning objective of DF is biased in the presence of label correlation. To address this issue, we propose the Frequency-enhanced Direct Forecast (FreDF), which mitigates label correlation by learning to forecast in the frequency domain, thereby reducing estimation bias. Our experiments show that FreDF significantly outperforms existing state-of-the-art methods and is compatible with a variety of forecast models. Code is available at https://github.com/Master-PLC/FreDF.
Hao Wang 0049, Lichen Pan, Zhichao Chen 0001, Degui Yang, Sen Zhang 0006, Xinggao Liu, Haoxuan Li 0001, Dacheng Tao
ICLR4
2025 Optimal Transport for Time Series Imputation
abstract
Missing data imputation through distribution alignment has demonstrated advantages for non-temporal datasets but exhibits suboptimal performance in time-series applications. The primary obstacle is crafting a discrepancy measure that simultaneously (1) captures temporal patterns—accounting for periodicity and temporal dependencies inherent in time-series—and (2) accommodates non-stationarity, ensuring robustness amidst multiple coexisting temporal patterns. In response to these challenges, we introduce the Proximal Spectrum Wasserstein (PSW) discrepancy, a novel discrepancy tailored for comparing two \textit{sets} of time-series based on optimal transport. It incorporates a pairwise spectral distance to encapsulate temporal patterns, and a selective matching regularization to accommodate non-stationarity. Subsequently, we develop the PSW for Imputation (PSW-I) framework, which iteratively refines imputation results by minimizing the PSW discrepancy. Extensive experiments demonstrate that PSW-I effectively accommodates temporal patterns and non-stationarity, outperforming prevailing time-series imputation methods. Code is available at https://github.com/FMLYD/PSW-I.
Hao Wang 0049, Zhengnan Li, Haoxuan Li 0001, Xu Chen 0017, Mingming Gong, BinChen, Zhichao Chen 0001
ICLR7
2025 Unbiased Recommender Learning from Implicit Feedback via Weakly Supervised Learning
abstract
Implicit feedback recommendation is challenged by the missing negative feedback essential for effective model training. Existing methods often resort to negative sampling, a technique that assumes unlabeled interactions as negative samples. This assumption risks misclassifying potential positive samples within the unlabeled data, thereby undermining model performance. To address this issue, we introduce PURL, a model-agnostic framework that reframes implicit feedback recommendation as a weakly supervised learning task, eliminating the need for negative samples. However, its unbiasedness hinges on the accurate estimation of the class prior. To address this challenge, we propose Progressive Proximal Transport (PPT), which estimates the class prior by minimizing the proximal transport cost between positive and unlabeled samples. Experiments on three real-world datasets validate the efficacy of PURL in terms of improved recommendation quality. Code is available at https://github.com/HowardZJU/weakrec.
Hao Wang 0049, Zhichao Chen 0001, Haotian Wang 0001, Yanchao Tan, Pan Li 0005, Tianqiao Liu, Xu Chen 0017, Haoxuan Li 0001, Zhouchen Lin
ICML2
2025 Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation
abstract
Heterogeneous treatment effect (HTE) estimation from observational data poses significant challenges due to treatment selection bias. Existing methods address this bias by minimizing distribution discrepancies between treatment groups in latent space, focusing on global alignment. However, the fruitful aspect of local proximity, where similar units exhibit similar outcomes, is often overlooked. In this study, we propose Proximity-enhanced CounterFactual Regression (CFR-Pro) to exploit proximity for enhancing representation balancing within the HTE estimation context. Specifically, we introduce a pair-wise proximity regularizer based on optimal transport to incorporate the local proximity in discrepancy calculation. However, the curse of dimensionality renders the proximity measure and discrepancy estimation ineffective-exacerbated by limited data availability for HTE estimation. To handle this problem, we further develop an informative subspace projector, which trades off minimal distance precision for improved sample complexity. Extensive experiments demonstrate that CFR-Pro accurately matches units across different treatment groups, effectively mitigates treatment selection bias, and significantly outperforms competitors. Code is available at https://github.com/HowardZJU/CFR-Pro.
Hao Wang 0049, Zhichao Chen 0001, Zhaoran Liu, Xu Chen 0017, Haoxuan Li 0001, Zhouchen Lin
KDD (2)2
2025 Time-o1: Time-Series Forecasting Needs Transformed Label Alignment
abstract
Training time-series forecasting models poses unique challenges in loss function design. Most existing approaches adopt temporal mean squared error, but this study reveals two critical limitations: (1) it ignores the presence of label autocorrelation, which biases it from the true label sequence likelihood; (2) it involves excessive number of tasks, which complicates optimization, especially for long-term forecasting. To address these issues, we introduce Time-o1, a transform-enhanced loss function for time-series forecasting. The central idea is to transform the label sequence into decorrelated components with discriminated significance. Models are then trained to align the most significant components, thereby effectively mitigating label autocorrelation and reducing task amount. Experiments demonstrate that Time-o1 achieves state-of-the-art performance and is compatible with various forecast models. Code is available at https://github.com/Master-PLC/Time-o1.
Hao Wang 0049, Pan Li 0005, Zhichao Chen 0001, Xu Chen 0017, Qingyang Dai, Lei Wang 0198, Haoxuan Li 0001, Zhouchen Lin
NeurIPS3
2025 Iterative Missing Data Imputation with Model Form Adaptation and Non-Missing Feature Supervision
abstract
Iterative imputation is a prevalent method for missing data imputation, where each feature is imputed iteratively by treating it as a target variable estimated from all other features. However, iterative imputation method suffers from two principal limitations: (1) it imposes a single parametric model form to impute all features, neglecting the potential for optimal models to vary among features, which risks model misspecification; and (2) it assumes every feature contains missing values, overlooking the potential presence of non-missing features, termed as oracle features, which are informative for imputation. To address these limitations, we propose kernel point imputation (KPI), a bi-level optimization framework for iterative missing data imputation. At the inner level, KPI adaptively learns the optimal model form for each feature within a reproducing kernel Hilbert space, addressing limitation (1). At the outer level, KPI utilizes oracle features as supervisory signals to iteratively refine the imputations, addressing limitation (2). Experiments demonstrate that KPI outperforms competitive imputation methods. Code is available at https://github.com/FMLYD/kpi.git.
Hao Wang 0049, Zhengnan Li, Zhichao Chen 0001, Xu Chen 0017, Shuting He, Haoxuan Li 0001, Zhouchen Lin
NeurIPS3
2025 TMoE-P: Toward the Pareto Optimum for Multivariate Soft Sensors
abstract
Multivariate soft sensors seek to provide accurate estimation of multiple quality variables through the analysis of measurable process variables, representing a significant advance over the traditional focus on single-quality variable sensors within industrial manufacturing. Current progress stays in applying parameter-sharing neural architectures while ignoring two fundamental issues: (1) catastrophic interference, where the indiscriminate sharing of parameters degrades performance due to the discrepancy of objectives; (2) seesaw optimization, where the optimizer overly focuses on one dominant yet simple objective at the expense of others. To address these issues, we reformulate multivariate soft sensors as a multi-objective optimization problem and propose the Task-aware Mixture-of-Experts framework for achieving the Pareto optimum (TMoE-P). Specifically, to handle issue(1), we propose an Objective-aware Mixture-of-Experts (OMoE) module, which consists of objective-specific and objective-shared experts to realize parameter sharing while accommodating the discrepancy between objectives. To handle issue(2), we devise a Pareto Objective Weighting (POW) module, which dynamically balances the weights of learning objectives to approximate the Pareto optimum among competing objectives. Our evaluations on a public soft sensor benchmark showcase TMoE-P’s superior performance, confirming its enhanced accuracy and robustness.Note to Practitioners—Addressing the burgeoning complexity of estimating multiple quality variables in industrial manufacturing processes, this study introduces a novel Task-aware Mixture-of-Experts framework aiming for the Pareto Optimum (TMoE-P). This framework mitigates the issues of catastrophic interference and seesaw optimization by achieving a delicate balance between parameter sharing and maintaining distinctness between objectives, guiding towards the Pareto optimum. The empirical findings affirm the framework’s capability to adeptly handle the complex interplay of variables, potentially enhancing the efficiency and reliability of industrial processes. While initially developed for multivariate data in industrial applications, the TMoE-P framework’s modular and adaptable nature facilitates its extension to a broad spectrum of data structures, optimizers, and neural architectures. Its adaptability offers practitioners a valuable tool to fulfill specific task requirements, making a modest contribution to industrial process control and monitoring.
Licheng Pan, Hao Wang 0049, Zhichao Chen 0001, Yuxin Huang 0007, Zhaoran Liu, Qunshan He, Xinggao Liu
IEEE Trans Autom. Sci. Eng.3
2025 Controllable Mixture-of-Experts for Multivariate Soft Sensors
abstract
Multivariate soft sensors are critical in industrial manufacturing for providing precise estimations of multiple quality variables and ensuring data reliability and completeness. Existing methods predominantly focus on parameter-sharing architectures while overlooking two fundamental issues: 1) catastrophic interference, where sharing parameters across all tasks leads to performance degradation; 2) uncontrollable optimization, where the optimizer lacks controllability over task priorities, misleading specific tasks converge to unexpected values. To handle these issues and enhance multivariate modeling, we propose a controllable Mixture-of-Experts (ControlMoE) framework for industrial soft sensors, consisting of a Mixture-of-Sequential-Experts (MoSE) module and a Proportional-Integral-Derivative Calibrating (PIDC) module. The MoSE module integrates sequential expert networks with task-specific gating networks, enabling parameter sharing while preserving task distinctions to mitigate catastrophic interference. The PIDC module employs an innovative PID controller to dynamically calibrate task learning weights, ensuring expected outcomes through feedback control and enhancing optimization controllability. Evaluations on two industrial datasets from Chinese nuclear monitoring stations demonstrate that ControlMoE achieves superior accuracy and controllability in multivariate soft sensor modeling. Note to Practitioners—This paper tackles the challenges in applying multivariate soft sensors within industrial manufacturing processes, particularly in environments such as nuclear monitoring where precision and control over multiple quality variables are critical. Traditional methods often suffer from performance degradation due to parameter sharing across all tasks-known as catastrophic interference–and lack control in optimizing multiple tasks simultaneously, leading to suboptimal outcomes. Our proposed framework, ControlMoE, introduces a novel architecture that combines sequential expert networks with a dynamic weight adjustment mechanism using PID controllers. This enables precise control over task-specific optimizations, ensuring that the desired importance and sequence of tasks are maintained, thereby enhancing both accuracy and operational control. We believe this framework could be widely applicable in various industrial settings where multivariate monitoring and control are necessary, improving not only performance but also the reliability of the data produced for critical decision-making.
Licheng Pan, Hao Wang 0049, Zhichao Chen 0001, Yuxin Huang 0007, Yunlong Niu, Zhaoran Liu, Qunshan He, Xinggao Liu
IEEE Trans Autom. Sci. Eng.3
2025 LSPT-D: Local Similarity Preserved Transport for Direct Industrial Data Imputation
abstract
Accurate imputation of missing data is pivotal in real-world industrial applications. Traditional direct imputers, which utilize basic statistics to replace missing elements, offer a practical solution but struggle to adapt to the complex patterns in industrial data, leaving a gap in the research landscape. This study explores the untapped potential of direct imputers, enhancing their adaptability and capacity to handle complex patterns in industrial data through optimal transport (OT) theory, with a focus on preserving local sample-wise similarity as an exemplar. To these ends, we construct a Local Similarity Preserved Transport (LSPT) problem, with a solution algorithm based on the Frank-Wolfe technique to compute transport cost. Subsequently, we propose the LSPT-D framework, which employs the transport cost of LSPT for distribution matching, directing the gradient flow to the missing data points to update the imputations directly. This strategy maintains local similarity throughout the imputation process thereby enhancing the overall imputation quality. Our experiments demonstrate that LSPT-D outperform various baselines in industrial missing data imputation. Note to Practitioners—Accurate missing data imputation is essential for enhancing the reliability of data analytics and reducing decision-making risks in industrial automation. This study introduces LSPT-D, a non-parametric imputation technique based on OT technology. Unique to LSPT-D is its ability to preserve local similarity during the imputation process, rendering it particularly advantageous for datasets with varying operational phases and load conditions. In industrial applications, LSPT-D not only significantly improves imputation quality compared to various baseline methods but also maintains modest running costs. Additionally, it serves as an exemplar for developing OT-based imputation strategies that capitalize on the inherent properties of data to improve imputation performance. However, LSPT-D operates under the independent and identically distributed assumption and is thus best applied in scenarios where temporal dependencies, such as trends and seasonality, are minimal or have been previously neutralized.
Hao Wang 0049, Xinggao Liu, Zhaoran Liu, Yilin Liao, Yuxin Huang 0007, Zhichao Chen 0001
IEEE Trans Autom. Sci. Eng.7
2025 Entire Space Counterfactual Learning for Reliable Content Recommendations
abstract
Post-click conversion rate (CVR) estimation is a fundamental task in developing effective recommender systems, yet it faces challenges from data sparsity and sample selection bias. To handle both challenges, the entire space multitask models are employed to decompose the user behavior track into a sequence of exposure$\rightarrow $click$\rightarrow $conversion, constructing surrogate learning tasks for CVR estimation. However, these methods suffer from two significant defects: (1) intrinsic estimation bias (IEB), where the CVR estimates are higher than the actual values; (2) false independence prior (FIP), where the causal relationship between clicks and subsequent conversions is potentially overlooked. To overcome these limitations, we develop a model-agnostic framework, namely Entire Space Counterfactual Multitask Model (ESCM2), which incorporates a counterfactual risk minimizer within the entire space multitask framework to regularize CVR estimation. Experiments conducted on large-scale industrial recommendation datasets and an online industrial recommendation service demonstrate that ESCM2 effectively mitigates IEB and FIP defects and substantially enhances recommendation performance.
Hao Wang 0049, Zhichao Chen 0001, Zhaoran Liu, Degui Yang, Xinggao Liu, Haoxuan Li 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Debiased Recommendation via Wasserstein Causal Balancing
abstract
Recommendation systems are pivotal in improving user experience on various digital platforms. However, observational training data in recommendation systems introduce selection bias, which leads to a distributional discrepancy between training data and real-world scenarios, resulting in suboptimal performance. Current causal debiasing methods such as inverse propensity score and doubly robust rely on accurately estimated propensity scores, typically optimized through negative log-likelihood (NLL) minimization. However, recent studies have highlighted the limitations of this approach, as perfect NLL minimization may not adequately correct for selection bias. To address this issue, we propose Wasserstein Balancing Metric (WBM), a novel metric that measures and enhances the balancing capacity of propensity scores in causal debiasing methods by minimizing the Wasserstein discrepancy between reweighted populations. On the basis, we introduce IPS-WBM and DR-WBM, incorporating WBM as a regularizer in standard inverse propensity score and doubly robust estimators, which enhances causal balancing capacity without introducing additional bias. Extensive experiments on three real-world recommendation datasets demonstrate that our methods improve the causal balancing capability of learned propensities and enhance debiasing performance.
Hao Wang 0049, Zhichao Chen 0001, Honglei Zhang 0002, Zhengnan Li, Licheng Pan, Haoxuan Li 0001, Mingming Gong
ACM Trans. Inf. Syst.2
2025 Improving Data-Driven Inferential Sensor Modeling by Industrial Knowledge: A Bayesian Perspective
abstract
Accurate quality variable inference by process variables is the core of industrial inferential sensor modeling, where recent advancements have seen deep learning (DL) models achieving remarkable success. However, integrating knowledge of unit operations is critical for improving inferential sensor performance, yet it has received little attention. The main challenge lies in the incompleteness and correctness of industrial knowledge due to its semi-empirical nature and inevitable engineering errors. Addressing this, this article introduces the gradient knowledge network based on the graph neural network’s message-passing mechanism within the variational Bayesian inference framework, which naturally copes with the abovementioned issues by fusing observational data. Initially, the prior knowledge about the process variables, which mirrors the graph in graph neural network, is parameterized as Dirichlet distribution based on the analysis of message-passing mechanism. However, the divergence computation and normalization constraints are challenging for model implementation. To navigate these challenges, the Bayesian inference problem is transformed into an optimization problem, subsequently recast as a simulation problem induced by the gradient field, ensuring compatibility with DL backends. Furthermore, a theoretical iteration equation is derived to maintain the normalization constraint. The architecture of the proposed model and its learning algorithm are then detailed. Finally, various experiments are conducted on two real industrial processes to demonstrate the model’s efficacy from the perspective of prediction accuracy, sensitivity analysis, and ablation study.
Zhichao Chen 0001, Hao Wang 0049, Zhiqiang Ge
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Rethinking the Diffusion Models for Missing Data Imputation: A Gradient Flow Perspective
abstract
Diffusion models have demonstrated competitive performance in missing data imputation (MDI) task. However, directly applying diffusion models to MDI produces suboptimal performance due to two primary defects. First, the sample diversity promoted by diffusion models hinders the accurate inference of missing values. Second, data masking reduces observable indices for model training, obstructing imputation performance. To address these challenges, we introduce $\underline{\text{N}}$egative $\underline{\text{E}}$ntropy-regularized $\underline{\text{W}}$asserstein gradient flow for $\underline{\text{Imp}}$utation (NewImp), enhancing diffusion models for MDI from a gradient flow perspective. To handle the first defect, we incorporate a negative entropy regularization term into the cost functional to suppress diversity and improve accuracy. To handle the second defect, we demonstrate that the imputation procedure of NewImp, induced by the conditional distribution-related cost functional, can equivalently be replaced by that induced by the joint distribution, thereby naturally eliminating the need for data masking. Extensive experiments validate the effectiveness of our method. Code is available at [https://github.com/JustusvLiebig/NewImp](https://github.com/JustusvLiebig/NewImp).
Zhichao Chen 0001, Haoxuan Li 0001, Fangyikang Wang, Odin Zhang, Hu Xu 0007, Hao Wang 0049
NeurIPS1
2024 Analyzing and Improving Supervised Nonlinear Dynamical Probabilistic Latent Variable Model for Inferential Sensors
abstract
Nonlinear dynamical probabilistic latent variable model (NDPLVM) and its variants, essential in industrial inferential sensors, face challenges in latent space inference and deep learning (DL) backend implementation. The first issue arises from the assumption that covariates directly infer the latent variable, potentially leading to inaccuracies. The second issue involves the discrepancy between the probabilistic distribution function form of NDPLVMs and data sample-based operation of DL backends. Addressing these, this study introduces the optimal control-NDPLVM (OC-NDPLVM), a model designed to enhance performance by analyzing NDPLVMs learning and tackling these issues. For the first problem, NDPLVMs' learning is reinterpreted as an optimization problem, solved by alternating direction method of multipliers, and selecting the inference network's input via studying optimal solution's structure. To address the second issue, OC-NDPLVM adapts mean and covariance equations for compatibility with DL backends. This model's effectiveness is validated through experiments on inferential sensor datasets.
Zhichao Chen 0001, Hao Wang 0049, Guofei Chen, Yiran Ma, Le Yao, Zhiqiang Ge
IEEE Trans. Ind. Informatics1
2024 Heat Equation Stein Variational Ensemble: Rethinking and Advancing Uncertainty-Aware Soft Sensor Modeling
abstract
Data-driven soft sensors have been prevalent in industrial key performance indicator prediction. However, rigorously quantifying the uncertainty of predictions has been consistently neglected. This neglect can result in industrial practitioners being ignorant of the reliability of the predictions, potentially leading to dangerous misjudgments. To bridge this gap, this study interprets uncertainty quantification (UQ) from a Bayesian perspective, which conceptualizes model uncertainty as the parameter distribution, thus distinguishing it from data uncertainty. Based on this interpretation, a novel UQ method, heat equation stein variational ensemble (HESVE), is proposed. HESVE introduces an efficient nonparametric approach called Stein variational gradient descent (SVGD) to approximate parameter distributions more precisely. A heat equation-based method is also adopted for adaptive hyperparameter tuning of SVGD. In addition, our method incorporates the variable noise estimator to capture heteroscedastic data noise, which contributes to uncertainty decomposition. Experiments demonstrate the superiority of HESVE in multiple aspects compared to other state-of-the-art methods.
Yiran Ma, Zhichao Chen 0001
IEEE Trans. Ind. Informatics2
2024 SPOT-I: Similarity Preserved Optimal Transport for Industrial IoT Data Imputation
abstract
Missing data imputation is a critical aspect of the Industrial Internet-of-Things (IIoT), which is uniquely challenged by local relationships within data due to different operational contexts and phases. Current imputation methods struggle to accommodate local relationships due to their black-box nature or limited capacity. To bridge this gap, we approach data imputation as a distribution alignment problem and leverage optimal transport to instantiate it for enhanced capacity. Specifically, we first introduce the similarity preserved optimal transport (SPOT) problem, with a conditional gradient solution to compute the transport cost. Subsequently, we propose the SPOT for imputation (SPOT-I) framework. It minimizes the transport cost of SPOT for distribution alignment and uses the gradient to update imputations, which maintains local similarity and refines imputation due to the characteristics of SPOT. Experiments on IIoT datasets showcase the superiority of SPOT-I over state-of-the-art imputation methods.
Hao Wang 0049, Zhichao Chen 0001, Zhaoran Liu, Licheng Pan, Hu Xu 0007, Yilin Liao, Xinggao Liu
IEEE Trans. Ind. Informatics2
2024 Variational Inference Over Graph: Knowledge Representation for Deep Process Data Analytics
abstract
With the advent of the industrial Big Data era, accurate estimation of product quality and monitoring of working conditions from historical data have become crucial in the process industry. However, the majority of data-driven approaches predominantly rely on observational data, overlooking the valuable empirical knowledge derived from experience or underlying mechanisms. In order to leverage this knowledge, researchers employ various graph neural network-based methods which introduce connections among process variables for feature extraction. Nevertheless, it is imperative to recognize that process knowledge undergoes changes due to internal or external concept drift. To address this challenge, we propose a novel deep learning module called “variational inference over graph” to effectively harness shifting knowledge. Building upon the self-attention mechanism, we design a probabilistic self-attention mechanism for encoding and reconciling prior knowledge. Instead of directly encoding the prior knowledge through graph neural network edges, we incorporate it as regularization term within the variational inference framework that accounts for knowledge shift. Furthermore, we introduce reparameterization estimator to control the variance resulting from knowledge uncertainty. To showcase the capability of our proposed method, we conduct various experiments on quality prediction task in real industrial processes.
Zhichao Chen 0001, Zhiqiang Ge
IEEE Trans. Knowl. Data Eng.1
2023 Monotonic Neural Ordinary Differential Equation: Time-series Forecasting for Cumulative Data
abstract
Time-Series Forecasting based on Cumulative Data (TSFCD) is a crucial problem in decision-making across various industrial scenarios. However, existing time-series forecasting methods often overlook two important characteristics of cumulative data, namely monotonicity and irregularity, which limit their practical applicability. To address this limitation, we propose a principled approach called Monotonic neural Ordinary Differential Equation (MODE) within the framework of neural ordinary differential equations. By leveraging MODE, we are able to effectively capture and represent the monotonicity and irregularity in practical cumulative data. Through extensive experiments conducted in a bonus allocation scenario, we demonstrate that MODE outperforms state-of-the-art methods, showcasing its ability to handle both monotonicity and irregularity in cumulative data and delivering superior forecasting performance.
Zhichao Chen 0001, Leilei Ding, Zhixuan Chu, Yucheng Qi, Jianmin Huang, Hao Wang 0049
CIKM1
2023 Unsupervised Anomaly Detection & Diagnosis: A Stein Variational Gradient Descent Approach
abstract
Detecting and diagnosing anomalies in observational data plays a crucial role in various real-world applications, such as e-commerce applet maintenance. Unsupervised machine learning techniques are typically employed for anomaly detection and diagnosis due to their convenience and independence from labeled data. Density estimation (DE), as one of the most widely used unsupervised machine learning techniques for anomaly detection, can be categorized into kernel density estimation (KDE)-based methods and normalizing flow (NF)-based methods. While KDE-based methods offer fast computation speed, they often ignore the complex manifold structure present in observational data. On the other hand, NF-based methods address the manifold issue but suffer from longer computation times. In this study, we propose a novel DE-based anomaly detection & diagnosis method using Stein Variational Gradient Descent (SVGD), aiming to leverage the strengths of KDE and NF approaches. Firstly, we rigorously derive the DE capability of SVGD through mathematical analysis. Subsequently, we demonstrate the ability of the SVGD method to perform anomaly diagnosis based on input feature attribution. Finally, to validate the effectiveness of our approach, we conduct experiments using synthetic, benchmark, and industrial datasets. The results demonstrate the superior performance and practical applicability of our proposed method.
Zhichao Chen 0001, Leilei Ding, Jianmin Huang, Zhixuan Chu, Qingyang Dai, Hao Wang 0049
CIKM1
2023 Optimal Transport for Treatment Effect Estimation
abstract
Estimating individual treatment effects from observational data is challenging due to treatment selection bias. Prevalent methods mainly mitigate this issue by aligning different treatment groups in the latent space, the core of which is the calculation of distribution discrepancy. However, two issues that are often overlooked can render these methods invalid: (1) mini-batch sampling effects (MSE), where the calculated discrepancy is erroneous in non-ideal mini-batches with outcome imbalance and outliers; (2) unobserved confounder effects (UCE), where the unobserved confounders are not considered in the discrepancy calculation. Both of these issues invalidate the calculated discrepancy, mislead the training of estimators, and thus impede the handling of treatment selection bias. To tackle these issues, we propose Entire Space CounterFactual Regression (ESCFR), which is a new take on optimal transport technology in the context of causality. Specifically, based on the canonical optimal transport framework, we propose a relaxed mass-preserving regularizer to address the MSE issue and design a proximal factual outcome regularizer to handle the UCE issue. Extensive experiments demonstrate that ESCFR estimates distribution discrepancy accurately, handles the treatment selection bias effectively, and outperforms prevalent competitors significantly.
Hao Wang 0049, Jiajun Fan, Zhichao Chen 0001, Haoxuan Li 0001, Weiming Liu 0005, Tianqiao Liu, Quanyu Dai, Yichao Wang 0002, Zhenhua Dong, Ruiming Tang
NeurIPS3
2022 ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate Estimation
abstract
Accurate estimation of post-click conversion rate is critical for building recommender systems, which has long been confronted with sample selection bias and data sparsity issues. Methods in the Entire Space Multi-task Model (ESMM) family leverage the sequential pattern of user actions, \ie $impression\rightarrow click \rightarrow conversion$ to address data sparsity issue. However, they still fail to ensure the unbiasedness of CVR estimates. In this paper, we theoretically demonstrate that ESMM suffers from the following two problems: (1) Inherent Estimation Bias (IEB) for CVR estimation, where the CVR estimate is inherently higher than the ground truth; (2) Potential Independence Priority (PIP) for CTCVR estimation, where ESMM might overlook the causality from click to conversion. To this end, we devise a principled approach named Entire Space Counterfactual Multi-task Modelling (ESCM$^2$), which employs a counterfactual risk miminizer as a regularizer in ESMM to address both IEB and PIP issues simultaneously. Extensive experiments on offline datasets and online environments demonstrate that our proposed ESCM$^2$ can largely mitigate the inherent IEB and PIP issues and achieve better performance than baseline models.
Hao Wang 0049, Tai-Wei Chang, Tianqiao Liu, Jianmin Huang, Zhichao Chen 0001, Ruopeng Li
SIGIR5
2022 Knowledge Automation Through Graph Mining, Convolution, and Explanation Framework: A Soft Sensor Practice
abstract
In industrial processes, data-driven soft sensors have played an important role for the effective process control, optimization, and monitoring. Deep learning technique has been widely used in soft sensor field in recent years for its excellent feature representation capability in spatial and temporal scales. However, the shortcomings for deep learning technique seriously hinder its application in industrial processes. For example, the knowledge cannot be added into the model, and the model prediction could not be well explained. To solve those problems, the graph mining, convolution, and explanation framework is proposed for knowledge automation in this article. Based on the equivalence analysis of the self-attention mechanism (SAM) and graph convolution (GC) operation, the spatial SAM is adopted for knowledge discovery from data directly. After that, the GC layer considering the relationship between process variables can utilize the knowledge for constructing soft sensor models. Besides, to explain which knowledge contributes to the final model prediction, the graph neural network explainer is designed for explaining the model output. Finally, the effectiveness and feasibility of the framework are evaluated on an industrial process, in which the knowledge discovered from the data is of great consistence with the prior knowledge, and the final explanation indicated that most of the knowledge is consistent with the prior knowledge contributed to the prediction.
Zhichao Chen 0001, Zhiqiang Ge
IEEE Trans. Ind. Informatics1