Zhong Chen 0003

dblp:70/2509-3 · DBLP profile ↗
← Back
17ranked-venue papers in the field
8as first author
15since 2021 · last 2026
0000-0002-7483-9699ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 10 (5 first)Big Data, Cloud & Distributed Data Systems · 5 (3 first)Database Systems & Data Management · 1Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2026 A survey on computational pathology foundation models: datasets, adaptation strategies, and evaluation tasks
abstract
Abstract Computational pathology foundation models (CPathFMs) have emerged as a powerful approach for analyzing histopathological data, leveraging self-supervised learning to extract robust feature representations from unlabeled whole-slide images. These models, categorized into uni-modal and multi-modal frameworks, have demonstrated promise in automating complex pathology tasks such as segmentation, classification, and biomarker discovery. However, the development of CPathFMs presents significant challenges, such as limited data accessibility, high variability across datasets, the necessity for domain-specific adaptation, and the lack of standardized evaluation benchmarks. This survey provides a comprehensive review of CPathFMs in computational pathology, focusing on pre-training datasets, adaptation strategies, and evaluation tasks. We analyze key techniques, such as contrastive learning, masked image modeling and multi-modal integration, and highlight existing gaps in current research. Finally, we explore future directions from four perspectives for advancing CPathFMs. This survey serves as a valuable resource for researchers, clinicians, and AI practitioners, guiding the advancement of CPathFMs toward robust and clinically applicable AI-driven pathology solutions.
Dong Li 0034, Guihong Wan, Xintao Wu, Yi He 0007, Zhong Chen 0003, Ajit Johnson Nirmal, Christine G. Lian, Peter K. Sorger, Yevgeniy R. Semenov, Chen Zhao 0010
Knowl. Inf. Syst.6
2025 ℓ1, ∞ Mixed Norm Promoted Row Sparsity for Fast Online CUR Decomposition Learning in Varying Feature Spaces
abstract
Online learning enables effective predictive modeling on complex data streams. To overcome the negative impact of possibly high-dimensional data, sparse online learning (SOL) has been proposed by imposing various sparse constraints to sheer the resultant model structure. However, most existing SOL studies focused on a fixed feature space, whereas in practice the steaming data observations may increment in both quantity and feature dimensions, leading to varying feature spaces. In this paper, we propose a novel ℓ1,∞-mixed norm-based row sparsity SOL algorithm (SOOFS) to handle data streams in varying feature spaces. We empower SOOFS with a tailored online CUR matrix decomposition method based on the promoted row sparsity to actively and adaptively select informative instances in the sliding windows, facilitating stable online performance over time. Empirical results on ten benchmark datasets substantiate the superiority of SOOFS over three state-of-the-art competitors in terms of classification accuracy and model sparsity.
Zhong Chen 0003, Yi He 0007, Di Wu 0056, Wenbin Zhang 0002, Zhiqiang Deng
SDM1
2025 Online Outlier Detection in Open Feature Spaces
abstract
Outlier detection is essential for data compliance, fraud prevention, and strategic decision-making. Finding outliers relies on study of feature space to find anomalous instances. As the feature dimension increases, it will inevitably complicate the process and hinder the models from finding genuine outliers. In this paper, we investigate an ever-more challenging task, online outlier detection (OOD) problem, where data points to be examined for outlier detection are characterized by two dynamic changes: (1) increasing volume instead of a static set; and (2) evolving feature space instead of a known set. Such instance and feature space dynamics impedes traditional OD techniques reliant on geometric data structure for distinguishing outliers. To aid, we propose a new approach coinedOnline Outlier Detection in Open Feature Spaces, which circumvents this limitation by learning a latent hypersphere representation, respectively positioning regular and anomalous data points inside and outside its boundary. The crux of our approach tailors a reconstruction loss, allowing each data point to be represented as anadditionof its pertinent feature embeddings. Each of these embeddings is updated non-intrusively, championing both efficient and incremental learning of the latent hypersphere. Extensive experiments on twelve benchmark datasets underscore the robustness and superior performance of our method against seven leading counterparts. Code is released inhttps://github.com/X1aoLian/OODOFS.git.
Heng Lian 0001, Yi He 0007, Di Wu 0056, Zhong Chen 0003, Xingquan Zhu 0001, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.4
2024 ℓ1, 2-Norm and CUR Decomposition based Sparse Online Active Learning for Data Streams with Streaming Features
abstract
Aiming at learning from a sequence of data instances over time, online learning has attracted increasing attention in the big data era. As two important variants, sparse online learning has been extensively explored by facilitating sparse constraints for online models such as truncated gradient, ℓ1-norm regularization, ℓ1-ball projection, and regularized dual averaging; while online active learning aims to build an online prediction model with a limited number of labeled instances, deploying the so called query strategies to select informative instances over time. However, most existing studies consider sparse online learning or online active learning with fixed feature spaces, whereby in real practice the features may be dynamically evolved over time. To the end, we propose a novel unified one-pass online learning framework named OASF for simultaneously online active learning and sparse online learning tailored for data streams described by open feature spaces, where new features can emerge constantly, and old features may be vanished over various time spans. Specifically, we technically develop an effective online CUR matrix decomposition based on the ℓ1,2mixed norm constraint for simultaneously selecting important up-to-date samples in a sliding window and facilitating stable and meaningful features in open feature spaces over time. If the loss function is simultaneously Lipschitz and convex, a sub-linear regret bound of our proposed algorithm is guaranteed with. Extensive experiments that are conducted with multiple streaming datasets have demonstrated the effectiveness of the proposed OASF compared with state-of-the-art online active learning and sparse online learning methods.
Zhong Chen 0003, Yi He 0007, Di Wu 0056, Liudong Zuo, Keren Li, Wenbin Zhang 0002, Zhiqiang Deng
IEEE Big Data1
2024 Accounting for Cancer Patients with Severe Outcomes: An Anomaly Detection Perspective
abstract
Health outcomes and radiation-induced toxicities for cancer patients undergoing radiation therapy are influenced by several factors, including the disease process, treatment plan, and various symptoms produced by both the disease and treatment. Accurate prediction and assessment of a patient’s health status are pivotal for precision and personalized healthcare, especially for patients suffering severe outcomes such as pain, depression, and sleep disorders. Motivated by the issue of extreme class imbalance, this study investigates a set of unsupervised anomaly detection approaches to classify patients with mild/intermediate and severe outcomes using patient-reported outcomes (PROs) datasets. We found that the HBOS method demonstrated superior performance for the minority class, representing cancer patients with severe health statuses. Moreover, IForest and KNN showcased good potential by effectively considering both majority and minority patients. This study may provide insightful guidelines for clinical practice.
Yang Yan 0003, Christopher Lominska, Gregory N. Gan, Zhong Chen 0003
IEEE Big Data5
2024 Optimal Scheduling Algorithms for Cost-Effective Bandwidth Reservation in HPNs
abstract
Vast amounts of data are continually being produced in various scientific fields. Once these large datasets are generated, they often require rapid transfer over long distances using bandwidth reservation services provided by high-performance networks (HPNs) dedicated to collaborative data storage and analysis. The primary goal of data transfer is typically to achieve the earliest completion time (ECT), but users may also seek to minimize financial costs associated with the transfer. Balancing these differing requirements can be challenging. In this paper, we explore the trade-off between ECT and cost in data transfers that use bandwidth reservation on variable paths with fixed bandwidth within dedicated HPNs. Our investigation focuses on two types of bandwidth reservation requests (BRRs) and their scheduling: (i) minimizing data transfer cost while meeting a data transfer deadline, and (ii) achieving ECT while meeting a specified maximum cost. To optimize the scheduling of both types of BRRs, we propose two novel algorithms and conduct extensive simulations to demonstrate their effectiveness and efficiency.
Liudong Zuo, Pan Lai, Zhong Chen 0003
IEEE Big Data3
2024 Advancing Graph Counterfactual Fairness Through Fair Representation Learning
Zichong Wang, Zhibo Chu, Ronald Blanco, Zhong Chen 0003, Shu-Ching Chen, Wenbin Zhang 0002
ECML/PKDD (7)4
2024 Individual Fairness with Group Awareness Under Uncertainty
Zichong Wang, Jocelyn Dzuong, Xiaoyong Yuan, Zhong Chen 0003, Yanzhao Wu 0001, Wenbin Zhang 0002
ECML/PKDD (5)4
2024 Robust Sparse Online Learning for Data Streams with Streaming Features
abstract
Sparse online learning has received extensive attention during the past few years. Most of existing algorithms that utilize ℓ1-norm regularization or ℓ1-ball projection assume that the feature space is fixed or changes by following explicit constraints. However, this assumption does not always hold in many real applications. Motivated by this observation, we propose a new online learning algorithm tailored for data streams described by open feature spaces, where new features can be occurred, and old features may be vanished over various time spans. Our algorithm named RSOL provides a strategy to adapt quickly to such feature dynamics by encouraging sparse model representation with an ℓ1- and ℓ2 -mixed regularizer. We leverage the proximal operator of the ℓ1,2 -mixed norm and show that our RSOL algorithm enjoys a closed-form solution at each iteration. A sub-linear regret bound of our proposed algorithm is guaranteed with a solid theoretical analysis. Empirical results benchmarked on nine streaming datasets validate the effectiveness of the proposed RSOL method over three state-of-the-art algorithms.
Zhong Chen 0003, Yi He 0007, Di Wu 0056, Huixin Zhan, Victor S. Sheng, Kun Zhang 0012
SDM1
2023 Simplex2vec Backward: From Vectors Back to Simplicial Complex
abstract
Simplicial neural networks (SNNs) were proposed to generate higher-order simplicial complex representations as vectors that encode not only pairwise relationships but also higher-order interactions between nodes. Although these vectors allowing us to consider richer data representations compared to typical graph convolution, most real-world graphs associated with molecule or human-related activities are often sensitive and might contain confidential information, e.g., molecular geometry or friend lists. However, little works investigate the potential threats for these simplicial complexes (higher-order interactions between nodes). We name this threat by Simplicial Complexes Reconstruction Attack (SCRA) and conduct this attack by studying whether the vectors can be inverted to (approximately) recover the simplicial complexes who used to generate them. Specifically, we first generate the vectors via a k-simplex2vec approach that extends the node2vec algorithm to simplices of higher dimensions to associate Euclidean vectors to simplicial complexes. We then present a Simplex2vec Backward algorithm to perform the SCRA on k-simplex2vec vectors by pointwise mutual information (PMI) matrix reconstruction.
Huixin Zhan, Kun Zhang 0012, Zhong Chen 0003, Victor S. Sheng
CIKM3
2023 Defending the Graph Reconstruction Attacks for Simplicial Neural Networks
abstract
Releasing the representations of nodes in real-world graphs associated with people or human-related activities, such as social and economic networks, gives adversaries a potential way to infer the sensitive information of edges. For example, graph convolutional layers initially aggregate node representations with their neighbors before passing them through non-linear activation functions. Hence, the released node representations may potentially breach edge privacy of the node neighbors. Thus, in this work, we study whether representations can be inverted to recover the graph used to generate them. We study three types of outputs that are trained on the graph, i.e., representations output from graph convolutional networks (GCNs), representations output from graph attention networks (GATs), and representations output from our proposed simplicial neural networks (SNNs). Unlike the first two types of representations that only encode pairwise relationships, the third type of representation, i.e., SNN outputs, encodes higher-order interactions (e.g., homological features) between nodes. We propose two graph reconstruction attacks (GRAs), i.e., Type-1 and Type-2 attacks, to recover a graph’s adjacency matrix from the three types of outputs trained on the graph. Specifically, our GRAs utilize a graph-decoder to minimize the reconstruction loss for the generated adjacency matrix via back-propagation. Our conclusions are two folds. First, our Type-2 attack achieves the best performance among all current GRAs. Second, we find that GCN outputs obtain the least precision and AUC on five datasets, followed by the GAT outputs, followed by the SNN outputs. Therefore, the SNN outputs reveal the lowest privacy-preserving ability to defend the GRAs. We further propose an unbiased multi-bit rectifier, by which the server can communicate with the nodes to privately collect their representations to defend the GRAs from potential adversaries.
Huixin Zhan, Liyuan Gao, Kun Zhang 0012, Zhong Chen 0003, Victor S. Sheng
DSAA4
2023 MMA: Multi-Metric-Autoencoder for Analyzing High-Dimensional and Incomplete Data
Cheng Liang 0003, Di Wu 0056, Yi He 0007, Teng Huang 0001, Zhong Chen 0003, Xin Luo 0001
ECML/PKDD (5)5
2023 An effective cost-sensitive sparse online learning framework for imbalanced streaming data classification and its application to online anomaly detection
Zhong Chen 0003, Victor S. Sheng, Andrea Edwards, Kun Zhang 0012
Knowl. Inf. Syst.1
2022 Proximal Cost-sensitive Sparse Group Online Learning
abstract
Effective streaming feature selection in dynamic on-line environments is essential in numerous applications. However, most existing methods evaluate high-dimensional features individually and ignore the potentially pertainable group structures of features. Moreover, the class imbalance underlying streaming data may further decrease the discriminative efficacy of the selected features, resulting in deteriorated classification performance. Motivated by this observation, we propose a proximal cost-sensitive sparse group online learning (PCSGOL) framework to handle imbalanced and high-dimensional streaming data. Specifically, we formulate this issue as a new cost-sensitive online optimization problem by leveraging the ℓ2-norm, ℓ1-norm, and group-wise sparsity constraints in the dual averaging regularization. The average weighted distance is also introduced in PCSGOL to achieve stable prediction results. We mathematically derive closed-form solutions to the optimization problems with four modified hinge loss functions, leading to four variants of PCSGOL. Extensive empirical studies on real-world streaming datasets demonstrate the effectiveness of our proposed method.
Zhong Chen 0003, Huixin Zhan, Victor S. Sheng, Andrea Edwards, Kun Zhang 0012
IEEE Big Data1
2022 Projection Dual Averaging Based Second-order Online Learning
abstract
Most existing online learning methods focus on mining ever-evolving streaming data based on the principle of first-order optimization. However, one drawback of these methods is the slow convergence rate in each iteration, resulting in sub-optimal solutions and deteriorated performance. Second-order methods, while are able to provide faster convergence, have been under-studied due to the high cost of computing the curvature information. To address this problem, in this paper, we develop a second-order projection dual averaging based online learning (SPDA) method to effectively handle high-throughput streaming data. By fully exploiting the regularized dual averaging optimization, the second-order information, and an optimal projection operator, SPDA converges fast with fairly optimal solutions. Two speed-up versions of SPDA, i.e., SPDA-diag and SPDA-sketch, are developed via the diagonal operator and Hessian sketch, respectively. Theoretical derivations on the regret bound of SPDA establish a solid convergence guarantee for this method. Extensive experiments demonstrate the efficacy of the proposed algorithms on large-scale online learning tasks, such as online binary and multi-class classification and online anomaly detection, shedding light on their potential wide applications.
Zhong Chen 0003, Huixin Zhan, Victor S. Sheng, Andrea Edwards, Kun Zhang 0012
ICDM1
2018 Online Density Estimation over Streaming Data: A Local Adaptive Solution
abstract
Accurate online density estimation is crucial to numerous applications that are prevalent with streaming data. Existing online approaches for density estimation somewhat lack prompt adaptability when facing drifting concepts, resulting in delayed or even deteriorated approximations. To alleviate this issue, in this work, we propose an adaptive local online density estimator, i.e. ALoKDE, for real-time density estimation on data streams. Two strategies, a statistical test for concept drift detection and an adaptive weighted local online density estimation when the drift occurs, are tightly integrated into ALoKDE. Specifically, using a weighted form, ALoKDE seeks to provide an unbiased estimation by factoring in the statistical hallmarks of the latest learned distribution and any potential distributional changes that could be introduced by each incoming instance. To ensure a high-precision estimate, ALoKDE integrates three key components: local sampling, optimal bandwidth selection at a temporal basis, and adaptive weighting factor determination. We further analyze the asymptotic properties of ALoKDE and derive its theoretical error bounds regarding bias, variance, MSE and MISE. Extensive comparative studies on various artificial and real-world streaming data demonstrate the efficacy of ALoKDE in online density estimation and real-time classification.
Zhong Chen 0003, Zhide Fang, Jiabin Zhao, Wei Fan 0001, Andrea Edwards, Kun Zhang 0012
IEEE BigData1
2017 CSTG: An Effective Framework for Cost-sensitive Sparse Online Learning
abstract
Sparse online learning and cost-sensitive learning are two important areas of machine learning and data mining research. Each has been well studied with many interesting algorithms developed. However, very limited published work addresses the joint study of these two fields. In this paper, to tackle the high-dimensional data streams with skewed distributions, we introduce a framework of cost-sensitive sparse online learning. Our proposed framework is a substantial extension of the influential Truncated Gradient (TG) method by formulating a new convex optimization problem, where the two mutual restraint factors, misclassification cost and sparsity, can be simultaneously and favorably balanced. We theoretically analyze the regret and cost bounds of the proposed algorithm, and pinpoint its theoretical merit compared to the existing related approaches. Large-scale empirical comparisons to five baseline methods on eight real-world streaming datasets demonstrate the encouraging performance of the developed method. Algorithm implementation and datasets are available upon request.
Zhong Chen 0003, Zhide Fang, Wei Fan 0001, Andrea Edwards, Kun Zhang 0012
SDM1