VLDB 2026 Research / reviewers in the wild / expert
Shujian Yu
dblp:154/5763
· DBLP profile ↗
78ranked-venue papers
20as first author
49since 2021 · last 2027
0000-0002-6385-1705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 16 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 19 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Cluster-aware continuous treatment effect estimation via Hellinger-based balanced representation learningabstractEstimating individualized treatment effects under continuous treatments (e.g., medication dosage) is challenging because treatment-response heterogeneity is often entangled with treatment-assignment bias in observational data. This paper proposes H-AdvCE, a Hellinger distance-based representation learning framework for cluster-aware continuous treatment effect estimation. The key idea is to learn a latent representation that preserves outcome-relevant heterogeneity while reducing dependence between the representation and treatment assignment. Theoretically, we derive a new counterfactual generalization error bound based on the Hellinger distance and show that it is tighter than commonly used Kullback–Leibler (KL) divergence-based alternatives. To optimize this objective, we introduce a variational Hellinger regularizer and formulate the learning problem as a min–max adversarial game. Empirically, H-AdvCE achieves strong performance on IHDP, News, and TCGA benchmarks for continuous dose-response estimation. Beyond prediction accuracy, the learned latent space reveals response-aware subgroups with distinct dose-response patterns, demonstrating its potential for clustering-driven interpretation in continuous treatment effect analysis. Our code is available at https://github.com/syzhao117/AdvCE . Yiran Zhuo, Hongchang Liu, Lincen Yang, Shujian Yu |
Signal Process. | 5 |
| 2026 | AnchorNet: A Clinically Grounded Causal Model for Personalized Medicine
Louk van Remmerden, Mark Hoogendoorn, Vincent François-Lavet, Shujian Yu, Hein van Hout |
AIME (1) | 4 |
| 2026 | MIMIR: Masked Image Modeling for Mutual Information-based Adversarial Robustness
Shujian Yu, Zhuoran Liu 0001, Stjepan Picek |
NDSS | 2 |
| 2025 | Cauchy-Schwarz Divergence Transfer EntropyabstractTransfer entropy (TE) is a powerful information-theoretic tool for analyzing causality in time series and complex systems. In this work, we propose a new formulation of TE using the Cauchy-Schwarz (CS) divergence. The resulting CS-TE offers a closed-form estimator and naturally extends to capture more complex causal relationships, such as indirect causation and synergistic effects, beyond just pairwise interactions. We also explore the feasibility of using a classifier, rather than regression models, to perform Granger tests in a supervised way. Lastly, we demonstrate the effectiveness of CS-TE on benchmark simulated data and stock indices from 14 stock markets. The code and supplementary material are available in our project repository: https://github.com/SJYuCNEL/Cauchy-Schwarz-Transfer-Entropy. Zhaozhao Ma, Shujian Yu |
ICASSP | 2 |
| 2025 | A Hierarchical Taxonomy For Deep State Space ModelsabstractModeling nonlinear dynamical systems is a challenging task in fields such as speech processing, music generation, and video prediction. This paper introduces a hierarchical framework for Deep State Space Models (DSSMs), categorizing them by their conditional independence properties and Markov assumptions and positioning existing models within this framework, including the Stochastic Recurrent Neural Network (SRNN), Variational Recurrent Neural Network (VRNN), and Recurrent State Space Model (RSSM). We discuss different options for the inference networks and demonstrate how integrating normalizing flows can enhance model flexibility by capturing complex distributions. Our work not only clarifies the relationships among existing models but also paves the way for the development of new, more effective approaches for modeling nonlinear dynamics. In particular, we propose the Autoregressive State Space Model (ArSSM) and evaluate its effectiveness in speech and polyphonic music modeling tasks. Shiqin Tang, Pengxing Feng, Shujian Yu, Yining Dong, S. Joe Qin |
ICASSP | 3 |
| 2025 | Deep Dynamic Probabilistic Canonical Correlation AnalysisabstractThis paper presents Deep Dynamic Probabilistic Canonical Correlation Analysis (D2PCCA), a model that integrates deep learning with probabilistic modeling to analyze nonlinear dynamical systems. Building on the probabilistic extensions of Canonical Correlation Analysis (CCA), D2PCCA captures nonlinear latent dynamics and supports enhancements such as KL annealing for improved convergence and normalizing flows for a more flexible posterior approximation. D2PCCA naturally extends to multiple observed variables, making it a versatile tool for encoding prior knowledge about sequential datasets and providing a probabilistic understanding of the system’s dynamics. Experimental validation on real financial datasets demonstrates the effectiveness of D2PCCA and its extensions in capturing latent dynamics. Shiqin Tang, Shujian Yu, Yining Dong, S. Joe Qin |
ICASSP | 2 |
| 2025 | Start Smart: Leveraging Gradients For Enhancing Mask-based XAI MethodsabstractMask-based explanation methods offer a powerful framework for interpreting deep learning model predictions across diverse data modalities, such as images and time series, in which the central idea is to identify an instance-dependent mask that minimizes the performance drop from the resulting masked input. Different objectives for learning such masks have been proposed, all of which, in our view, can be unified under an information-theoretic framework that balances performance degradation of the masked input with the complexity of the resulting masked representation. Typically, these methods initialize the masks either uniformly or as all-ones.
In this paper, we argue that an effective mask initialization strategy is as important as the development of novel learning objectives, particularly in light of the significant computational costs associated with existing mask-based explanation methods. To this end, we introduce a new gradient-based initialization technique called StartGrad, which is the first initialization method specifically designed for mask-based post-hoc explainability methods. Compared to commonly used strategies, StartGrad is provably superior at initialization in striking the aforementioned trade-off. Despite its simplicity, our experiments demonstrate that StartGrad enhances the optimization process of various state-of-the-art mask-explanation methods by reaching target metrics faster and, in some cases, boosting their overall performance. Buelent Uendes, Shujian Yu, Mark Hoogendoorn |
ICLR | 2 |
| 2025 | Aggregation of Dependent Expert Distributions in Multimodal Variational AutoencodersabstractMultimodal learning with variational autoencoders (VAEs) requires estimating joint distributions to evaluate the evidence lower bound (ELBO). Current methods, the product and mixture of experts, aggregate single-modality distributions assuming independence for simplicity, which is an overoptimistic assumption. This research introduces a novel methodology for aggregating single-modality distributions by exploiting the principle of *consensus of dependent experts* (CoDE), which circumvents the aforementioned assumption. Utilizing the CoDE method, we propose a novel ELBO that approximates the joint likelihood of the multimodal data by learning the contribution of each subset of modalities. The resulting CoDE-VAE model demonstrates better performance in terms of balancing the trade-off between generative coherence and generative quality, as well as generating more precise log-likelihood estimations. CoDE-VAE further minimizes the generative quality gap as the number of modalities increases. In certain cases, it reaches a generative quality similar to that of unimodal VAEs, which is a desirable property that is lacking in most current methods. Finally, the classification accuracy achieved by CoDE-VAE is comparable to that of state-of-the-art multimodal VAE models. Rogelio Andrade Mancisidor, Robert Jenssen, Shujian Yu, Michael Kampffmeyer |
ICML | 3 |
| 2025 | MvHo-IB: Multi-view Higher-Order Information Bottleneck for Brain Disorder Diagnosis
Kunyu Zhang, Qiang Li 0046, Shujian Yu |
MICCAI (15) | 3 |
| 2025 | Cross-Modal Retrieval with Cauchy-Schwarz DivergenceabstractEffective cross-modal retrieval requires robust alignment of heterogeneous data types. Most existing methods focus on bi-modal retrieval tasks and rely on distributional alignment techniques such as Kullback-Leibler divergence, Maximum Mean Discrepancy, and correlation alignment. However, these methods often suffer from critical limitations, including numerical instability, sensitivity to hyperparameters, and their inability to capture the full structure of the underlying distributions. In this paper, we introduce the Cauchy-Schwarz (CS) divergence, a hyperparameter-free measure that improves both training stability and retrieval performance. We further propose a novel Generalized CS (GCS) divergence inspired by Holder's inequality. This extension enables direct alignment of three or more modalities within a unified mathematical framework through a bidirectional circular comparison scheme, eliminating the need for exhaustive pairwise comparisons. Extensive experiments on six benchmark datasets demonstrate the effectiveness of our method in both bi-modal and tri-modal retrieval tasks. The code of our CS/GCS divergence is publicly available at https://github.com/JiahaoZhang666/CSD. Wenzhe Yin, Shujian Yu |
ACM Multimedia | 3 |
| 2025 | InfoDPCCA: Information-Theoretic Dynamic Probabilistic Canonical Correlation AnalysisabstractExtracting meaningful latent representations from high-dimensional sequential data is a crucial challenge in machine learning, with applications spanning natural science and engineering. We introduce InfoDPCCA, a dynamic probabilistic Canonical Correlation Analysis (CCA) framework designed to model two interdependent sequences of observations. InfoDPCCA leverages a novel information-theoretic objective to extract a shared latent representation that captures the mutual structure between the data streams and balances representation compression and predictive sufficiency while also learning separate latent components that encode information specific to each sequence. Unlike prior dynamic CCA models, such as DPCCA, our approach explicitly enforces the shared latent space to encode only the mutual information between the sequences, improving interpretability and robustness. We further introduce a two-step training scheme to bridge the gap between information-theoretic representation learning and generative modeling, along with a residual connection mechanism to enhance training stability. Through experiments on synthetic and medical fMRI data, we demonstrate that InfoDPCCA excels as a tool for representation learning. Code of InfoDPCCA is available at https://github.com/marcusstang/InfoDPCCA. Shiqin Tang, Shujian Yu |
UAI | 2 |
| 2025 | Generalized Cauchy-Schwarz divergence: Efficient estimation and applications in deep learning
Mingfei Lu, Shujian Yu, Robert Jenssen, Badong Chen |
Neurocomputing | 2 |
| 2025 | The Conditional Cauchy-Schwarz Divergence With Applications to Time-Series Data and Sequential Decision MakingabstractThe Cauchy-Schwarz (CS) divergence was developed by Príncipe et al. in 2000. In this paper, we extend the classic CS divergence to quantify the closeness between two conditional distributions and show that the developed conditional CS divergence can be elegantly estimated by a kernel density estimator from given samples. We illustrate the advantages (e.g., rigorous faithfulness guarantee, lower computational complexity, higher statistical power, and much more flexibility in a wide range of applications) of our conditional CS divergence over previous proposals, such as the conditional Kullback-Leibler divergence and the conditional maximum mean discrepancy. We also demonstrate the compelling performance of conditional CS divergence in two machine learning tasks related to time series data and sequential inference, namely time series clustering and uncertainty-guided exploration for sequential decision making. Shujian Yu, Sigurd Løkse, Robert Jenssen, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | An information bottleneck approach for feature selection
Qi Zhang 0089, Mingfei Lu, Shujian Yu, Jingmin Xin, Badong Chen |
Pattern Recognit. | 3 |
| 2025 | Guest Editorial: Special Issue on Information Theoretic Methods for the Generalization, Robustness, and Interpretability of Machine Learning
Badong Chen, Shujian Yu, Robert Jenssen, José C. Príncipe, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | BrainIB: Interpretable Brain Network-Based Psychiatric Diagnosis With Graph Information BottleneckabstractDeveloping new diagnostic models based on the underlying biological mechanisms rather than subjective symptoms for psychiatric disorders is an emerging consensus. Recently, machine learning (ML)-based classifiers using functional connectivity (FC) for psychiatric disorders and healthy controls (HCs) are developed to identify brain markers. However, existing ML-based diagnostic models are prone to overfitting (due to insufficient training samples) and perform poorly in new test environments. Furthermore, it is difficult to obtain explainable and reliable brain biomarkers elucidating the underlying diagnostic decisions. These issues hinder their possible clinical applications. In this work, we propose BrainIB, a new graph neural network (GNN) framework to analyze functional magnetic resonance images (fMRI), by leveraging the famed information bottleneck (IB) principle. BrainIB is able to identify the most informative edges in the brain (i.e., subgraph) and generalizes well to unseen data. We evaluate the performance of BrainIB against three baselines and seven state-of-the-art (SOTA) brain network classification methods on three psychiatric datasets and observe that our BrainIB always achieves the highest diagnosis accuracy. It also discovers the subgraph biomarkers that are consistent with clinical and neuroimaging findings. The source code and implementation details of BrainIB are freely available at the GitHub repository (https://github.com/SJYuCNEL/brain-and-Information-Bottleneck). Kaizhong Zheng, Shujian Yu, Baojuan Li, Robert Jenssen, Badong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | DIB-X: Formulating Explainability Principles for a Self-Explainable Model Through Information Theoretic LearningabstractThe recent development of self-explainable deep learning approaches has focused on integrating well-defined explainability principles into learning process, with the goal of achieving these principles through optimization. In this work, we propose DIB-X, a self-explainable deep learning approach for image data, which adheres to the principles of minimal, sufficient, and interactive explanations. The minimality and sufficiency principles are rooted from the trade-off relationship within the information bottleneck framework. Distinctly, DIB-X directly quantifies the minimality principle using the recently proposed matrix-based Rényi’s α-order entropy functional, circumventing the need for variational approximation and distributional assumption. The interactivity principle is realized by incorporating existing domain knowledge as prior explanations, fostering explanations that align with established domain understanding. Empirical results on MNIST and two marine environment monitoring datasets with different modalities reveal that our approach primarily provides improved explainability with the added advantage of enhanced classification performance. Changkyu Choi, Shujian Yu, Michael Kampffmeyer, Arnt-Børre Salberg, Nils Olav Handegard, Robert Jenssen |
ICASSP | 2 |
| 2024 | Rethinking Information-theoretic Generalization: Loss Entropy Induced PAC BoundsabstractInformation-theoretic generalization analysis has achieved astonishing success in characterizing the generalization capabilities of noisy and iterative learning algorithms. However, current advancements are mostly restricted to average-case scenarios and necessitate the stringent bounded loss assumption, leaving a gap with regard to computationally tractable PAC generalization analysis, especially for long-tailed loss distributions. In this paper, we bridge this gap by introducing a novel class of PAC bounds through leveraging loss entropies. These bounds simplify the computation of key information metrics in previous PAC information-theoretic bounds to one-dimensional variables, thereby enhancing computational tractability. Moreover, our data-independent bounds provide novel insights into the generalization behavior of the minimum error entropy criterion, while our data-dependent bounds improve over previous results by alleviating the bounded loss assumption under both leave-one-out and supersample settings. Extensive numerical studies indicate strong correlations between the generalization error and the induced loss entropy, showing that the presented bounds adeptly capture the patterns of the true generalization gap under various learning scenarios. Yuxin Dong 0003, Tieliang Gong, Hong Chen 0004, Shujian Yu, Chen Li 0011 |
ICLR | 4 |
| 2024 | Cauchy-Schwarz Divergence Information Bottleneck for RegressionabstractThe information bottleneck (IB) approach is popular to improve the generalization, robustness and explainability of deep neural networks. Essentially, it aims to find a minimum sufficient representation $\mathbf{t}$ by striking a trade-off between a compression term $I(\mathbf{x};\mathbf{t})$ and a prediction term $I(y;\mathbf{t})$, where $I(\cdot;\cdot)$ refers to the mutual information (MI). MI is for the IB for the most part expressed in terms of the Kullback-Leibler (KL) divergence, which in the regression case corresponds to prediction based on mean squared error (MSE) loss with Gaussian assumption and compression approximated by variational inference.
In this paper, we study the IB principle for the regression problem and develop a new way to parameterize the IB with deep neural networks by exploiting favorable properties of the Cauchy-Schwarz (CS) divergence. By doing so, we move away from MSE-based regression and ease estimation by avoiding variational approximations or distributional assumptions. We investigate the improved generalization ability of our proposed CS-IB and demonstrate strong adversarial robustness guarantees. We demonstrate its superior performance on six real-world regression tasks over other popular deep IB approaches. We additionally observe that the solutions discovered by CS-IB always achieve the best trade-off between prediction accuracy and compression ratio in the information plane. The code is available at \url{https://github.com/SJYuCNEL/Cauchy-Schwarz-Information-Bottleneck}. Shujian Yu, Sigurd Løkse, Robert Jenssen, José C. Príncipe |
ICLR | 1 |
| 2024 | Jacobian Regularizer-based Neural Granger CausalityabstractWith the advancement of neural networks, diverse methods for neural Granger causality have emerged, which demonstrate proficiency in handling complex data, and nonlinear relationships. However, the existing framework of neural Granger causality has several limitations. It requires the construction of separate predictive models for each target variable, and the relationship depends on the sparsity on the weights of the first layer, resulting in challenges in effectively modeling complex relationships between variables as well as unsatisfied estimation accuracy of Granger causality. Moreover, most of them cannot grasp full-time Granger causality. To address these drawbacks, we propose a **J**acobian **R**egularizer-based **N**eural **G**ranger **C**ausality (**JRNGC**) approach, a straightforward yet highly effective method for learning multivariate summary Granger causality and full-time Granger causality by constructing a single model for all target variables. Specifically, our method eliminates the sparsity constraints of weights by leveraging an input-output Jacobian matrix regularizer, which can be subsequently represented as the weighted causal matrix in the post-hoc analysis. Extensive experiments show that our proposed approach achieves competitive performance with the state-of-the-art methods for learning summary Granger causality and full-time Granger causality while maintaining lower model complexity and high scalability. Wanqi Zhou, Shuanghao Bai, Shujian Yu, Qibin Zhao, Badong Chen |
ICML | 3 |
| 2024 | BAN: Detecting Backdoors Activated by Adversarial Neuron NoiseabstractBackdoor attacks on deep learning represent a recent threat that has gained significant attention in the research community.
Backdoor defenses are mainly based on backdoor inversion, which has been shown to be generic, model-agnostic, and applicable to practical threat scenarios. State-of-the-art backdoor inversion recovers a mask in the feature space to locate prominent backdoor features, where benign and backdoor features can be disentangled. However, it suffers from high computational overhead, and we also find that it overly relies on prominent backdoor features that are highly distinguishable from benign features. To tackle these shortcomings, this paper improves backdoor feature inversion for backdoor detection by incorporating extra neuron activation information. In particular, we adversarially increase the loss of backdoored models with respect to weights to activate the backdoor effect, based on which we can easily differentiate backdoored and clean models. Experimental results demonstrate our defense, BAN, is 1.37$\times$ (on CIFAR-10) and 5.11$\times$ (on ImageNet200) more efficient with an average 9.99\% higher detect success rate than the state-of-the-art defense BTI DBF. Our code and trained models are publicly available at https://github.com/xiaoyunxxy/ban. Zhuoran Liu 0001, Stefanos Koffas, Shujian Yu, Stjepan Picek |
NeurIPS | 4 |
| 2024 | Domain Adaptation with Cauchy-Schwarz DivergenceabstractDomain adaptation aims to use training data from one or multiple source domains to learn a hypothesis that can be generalized to a different, but related, target domain. As such, having a reliable measure for evaluating the discrepancy of both marginal and conditional distributions is crucial. We introduce Cauchy-Schwarz (CS) divergence to the problem of unsupervised domain adaptation (UDA). The CS divergence offers a theoretically tighter generalization error bound than the popular Kullback-Leibler divergence. This holds for the general case of supervised learning, including multi-class classification and regression. Furthermore, we illustrate that the CS divergence enables a simple estimator on the discrepancy of both marginal and conditional distributions between source and target domains in the representation space, without requiring any distributional assumptions. We provide multiple examples to illustrate how the CS divergence can be conveniently used in both distance metric- or adversarial training-based UDA frameworks, resulting in compelling performance. The code of our paper is available at \url{https://github.com/ywzcode/CS-adv}. Wenzhe Yin, Shujian Yu, Yicong Lin, Jie Liu 0043, Jan-Jakob Sonke, Efstratios Gavves |
UAI | 2 |
| 2024 | R2-trans: Fine-grained visual categorization with redundancy reduction
Shuo Ye, Shujian Yu, Yu Wang 0106, Xinge You |
Image Vis. Comput. | 2 |
| 2024 | CI-GNN: A Granger causality-inspired graph neural network for interpretable brain network-based psychiatric diagnosis
Kaizhong Zheng, Shujian Yu, Badong Chen |
Neural Networks | 2 |
| 2023 | Robust and Fast Measure of Information via Low-Rank RepresentationabstractThe matrix-based Rényi's entropy allows us to directly quantify information measures from given data, without explicit estimation of the underlying probability distribution. This intriguing property makes it widely applied in statistical inference and machine learning tasks. However, this information theoretical quantity is not robust against noise in the data, and is computationally prohibitive in large-scale applications. To address these issues, we propose a novel measure of information, termed low-rank matrix-based Rényi's entropy, based on low-rank representations of infinitely divisible kernel matrices. The proposed entropy functional inherits the specialty of of the original definition to directly quantify information from data, but enjoys additional advantages including robustness and effective calculation. Specifically, our low-rank variant is more sensitive to informative perturbations induced by changes in underlying distributions, while being insensitive to uninformative ones caused by noises. Moreover, low-rank Rényi's entropy can be efficiently approximated by random projection and Lanczos iteration techniques, reducing the overall complexity from O(n³) to O(n²s) or even O(ns²), where n is the number of data samples and s ≪ n. We conduct large-scale experiments to evaluate the effectiveness of this new information measure, demonstrating superior results compared to matrix-based Rényi's entropy in terms of both performance and computational efficiency. Yuxin Dong 0003, Tieliang Gong, Shujian Yu, Hong Chen 0004, Chen Li 0011 |
AAAI | 3 |
| 2023 | Causal Recurrent Variational Autoencoder for Medical Time Series GenerationabstractWe propose causal recurrent variational autoencoder (CR-VAE), a novel generative model that is able to learn a Granger causal graph from a multivariate time series x and incorporates the underlying causal mechanism into its data generation process. Distinct to the classical recurrent VAEs, our CR-VAE uses a multi-head decoder, in which the p-th head is responsible for generating the p-th dimension of x (i.e., x^p). By imposing a sparsity-inducing penalty on the weights (of the decoder) and encouraging specific sets of weights to be zero, our CR-VAE learns a sparse adjacency matrix that encodes causal relations between all pairs of variables. Thanks to this causal matrix, our decoder strictly obeys the underlying principles of Granger causality, thereby making the data generating process transparent. We develop a two-stage approach to train the overall objective. Empirically, we evaluate the behavior of our model in synthetic data and two real-world human brain datasets involving, respectively, the electroencephalography (EEG) signals and the functional magnetic resonance imaging (fMRI) data. Our model consistently outperforms state-of-the-art time series generative models both qualitatively and quantitatively. Moreover, it also discovers a faithful causal graph with similar or improved accuracy over existing Granger causality-based causal inference methods. Code of CR-VAE is publicly available at https://github.com/hongmingli1995/CR-VAE. Shujian Yu, José C. Príncipe |
AAAI | 2 |
| 2023 | The Analysis of Deep Neural Networks by Information Theory: From Explainability to GeneralizationabstractDespite their great success in many artificial intelligence tasks, deep neural networks (DNNs) still suffer from a few limitations, such as poor generalization behavior for out-of-distribution (OOD) data and the "black-box" nature. Information theory offers fresh insights to solve these challenges. In this short paper, we briefly review the recent developments in this area, and highlight our contributions. Shujian Yu |
AAAI | 1 |
| 2023 | Revisiting the Robustness of the Minimum Error Entropy Criterion: A Transfer Learning Case StudyabstractCoping with distributional shifts is an important part of transfer learning methods in order to perform well in real-life tasks. However, most of the existing approaches in this area either focus on an ideal scenario in which the data does not contain noises or employ a complicated training paradigm or model design to deal with distributional shifts. In this paper, we revisit the robustness of the minimum error entropy (MEE) criterion, a widely used objective in statistical signal processing to deal with non-Gaussian noises, and investigate its feasibility and usefulness in real-life transfer learning regression tasks, where distributional shifts are common. Specifically, we put forward a new theoretical result showing the robustness of MEE against covariate shift. We also show that by simply replacing the mean squared error (MSE) loss with the MEE on basic transfer learning algorithms such as fine-tuning and linear probing, we can achieve competitive performance with respect to state-of-the-art transfer learning algorithms. We justify our arguments on both synthetic data and 5 real-world time-series data. Luis P. Silvestrin, Shujian Yu, Mark Hoogendoorn |
ECAI | 2 |
| 2023 | Towards a More Stable and General Subgraph Information BottleneckabstractGraph Neural Networks (GNNs) have been widely applied to graph-structured data. However, the lack of interpretability impedes its practical deployment especially in high-risk areas such as medical diagnosis. Recently, the Information Bottleneck (IB) principle has been extended to GNNs to identify a compact subgraph that is most informative to class labels, which significantly improves the interpretability on decision. However, existing Graph Information Bottleneck (GIB) models are either unstable during the training (due to the difficulty of mutual information estimation) or only focus on a special kind of graph (e.g., brain networks) that suffer from poor generalization to general graph datasets with varying graph sizes. In this work, we extend the recently developed Brain Information Bottleneck (BrainIB) to general graphs by introducing matrix-based Rényi’s α-order mutual information to stablize the training; and by designing a novel mask strategy to deal with varying graph sizes such that the new method can also be used for social networks, molecules, etc. Extensive experiments on different types of graph datasets demonstrate the superior stability and generality of our model. Kaizhong Zheng, Shujian Yu, Badong Chen |
ICASSP | 3 |
| 2023 | Sequential Invariant Information BottleneckabstractPrevious approaches to the problem of generalization for out-of-distribution (OOD) data usually assume that data from each environment is available simultaneously, which is unrealistic in real-world applications. In this paper, we develop a new framework termed the sequential invariant information bottleneck (seq-IIB) to improve the generalization ability of learning agents in sequential environments. Our main idea is to combine the merits of the famed Information Bottleneck (IB) principle with the Invariant Risk Minimization (IRM), such that the learning agent can gradually remove spurious features and remain invariant and compact task-relevant information in a sequential manner. Experimental results on three MNIST-like datasets show the effectiveness of our method. Shujian Yu, Badong Chen |
ICASSP | 2 |
| 2023 | Coping with change: Learning invariant and minimum sufficient representations for fine-grained visual categorization
Shuo Ye, Shujian Yu, Wenjin Hou, Yu Wang 0106, Xinge You |
Comput. Vis. Image Underst. | 2 |
| 2023 | Gated information bottleneck for generalization in sequential environments
Francesco Alesiani, Shujian Yu |
Knowl. Inf. Syst. | 2 |
| 2023 | Multiscale principle of relevant information for hyperspectral image classification
Yantao Wei, Shujian Yu, Luis Gonzalo Sánchez Giraldo, José C. Príncipe |
Mach. Learn. | 2 |
| 2023 | Optimal Randomized Approximations for Matrix-Based Rényi's EntropyabstractThe Matrix-based Rényi’s entropy enables us to directly measure information quantities from given data without the costly probability density estimation of underlying distributions, thus has been widely adopted in numerous statistical learning and inference tasks. However, exactly calculating this new information quantity requires access to the eigenspectrum of a semi-positive definite (SPD) matrix$A$which grows linearly with the number of samples$n$, resulting in a$O(n^{3})$time complexity that is prohibitive for large-scale applications. To address this issue, this paper takes advantage of stochastic trace approximations for matrix-based Rényi’s entropy with arbitrary$\alpha \in \mathbb {R}^{+}$orders, lowering the complexity by converting the entropy approximation to a matrix-vector multiplication problem. Specifically, we develop random approximations for integer-order$\alpha $cases and polynomial series approximations (Taylor and Chebyshev) for fractional$\alpha $cases, leading to a$O(n^{2}sm)$overall time complexity, where$s, m \ll n$denote the number of vector queries and the polynomial order respectively. We theoretically establish statistical guarantees for all approximation algorithms and give explicit order of$s$and$m$with respect to the approximation error$\epsilon $, showing optimal convergence rate for both parameters up to a logarithmic factor. Large-scale simulations and real-world applications validate the effectiveness of the developed approximations, demonstrating remarkable speedup with negligible loss in accuracy. Yuxin Dong 0003, Tieliang Gong, Shujian Yu, Chen Li 0011 |
IEEE Trans. Inf. Theory | 3 |
| 2023 | Selective Imputation for Multivariate Time Series Datasets With Missing ValuesabstractMultivariate time series often contain missing values for reasons such as failures in data collection mechanisms. Since these missing values can complicate the analysis of time series data, imputation techniques are typically used to deal with this issue. However, the quality of the imputation directly affects the performance of downstream tasks. In this paper, we propose a selective imputation method that identifies a subset of timesteps with missing values to impute in a multivariate time series dataset. This selection, which will result in shorter and simpler time series, is based on both reducing the uncertainty of the imputations and representing the original time series as good as possible. In particular, the method uses multi-objective optimization techniques to select the optimal set of points, and in this selection process, we leverage the beneficial properties of the Multi-task Gaussian Process (MGP). The method is applied to different datasets to analyze the quality of the imputations and the performance obtained in downstream tasks, such as classification or anomaly detection. The results show that much shorter and simpler time series are able to maintain or even improve both the quality of the imputations and the performance of the downstream tasks. Ane Blázquez-García, Kristoffer Wickstrøm, Shujian Yu, Karl Øyvind Mikalsen, Ahcène Boubekki, Angel Conde, Usue Mori, Robert Jenssen, José Antonio Lozano 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | A Componentwise Approach to Weakly Supervised Semantic Segmentation Using Dual-Feedback NetworkabstractRecent weakly supervised semantic segmentation methods generate pseudolabels to recover the lost position information in weak labels for training the segmentation network. Unfortunately, those pseudolabels often contain mislabeled regions and inaccurate boundaries due to the incomplete recovery of position information. It turns out that the result of semantic segmentation becomes determinate to a certain degree. In this article, we decompose the position information into two components: high-level semantic information and low-level physical information, and develop a componentwise approach to recover each component independently. Specifically, we propose a simple yet effective pseudolabels updating mechanism to iteratively correct mislabeled regions inside objects to precisely refine high-level semantic information. To reconstruct low-level physical information, we utilize a customized superpixel-based random walk mechanism to trim the boundaries. Finally, we design a novel network architecture, namely, a dual-feedback network (DFN), to integrate the two mechanisms into a unified model. Experiments on benchmark datasets show that DFN outperforms the existing state-of-the-art methods in terms of intersection-over-union (mIoU). Zhengqiang Zhang, Qinmu Peng, Sichao Fu, Yiu-Ming Cheung, Shujian Yu, Xinge You |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | Learning to Transfer with von Neumann Conditional DivergenceabstractThe similarity of feature representations plays a pivotal role in the success of problems related to domain adaptation. Feature similarity includes both the invariance of marginal distributions and the closeness of conditional distributions given the desired response y (e.g., class labels). Unfortunately, traditional methods always learn such features without fully taking into consideration the information in y, which in turn may lead to a mismatch of the conditional distributions or the mixup of discriminative structures underlying data distributions. In this work, we introduce the recently proposed von Neumann conditional divergence to improve the transferability across multiple domains. We show that this new divergence is differentiable and eligible to easily quantify the functional dependence between features and y. Given multiple source tasks, we integrate this divergence to capture discriminative information in y and design novel learning objectives assuming those source tasks are observed either simultaneously or sequentially. In both scenarios, we obtain favorable performance against state-of-the-art methods in terms of smaller generalization error on new tasks and less catastrophic forgetting on source tasks (in the sequential setup). Ammar Shaker, Shujian Yu, Daniel Oñoro-Rubio |
AAAI | 2 |
| 2022 | Deep Deterministic Independent Component Analysis for Hyperspectral UnmixingabstractWe develop a new neural network based independent component analysis (ICA) method by directly minimizing the dependence amongst all extracted components. Using the matrix-based Rényi’s α-order entropy functional, our network can be directly optimized by stochastic gradient descent (SGD), without any variational approximation or adversarial training. As a solid application, we evaluate our ICA in the problem of hyperspectral unmixing (HU) and refute a statement that "ICA does not play a role in unmixing hyperspectral data", which was initially suggested by [1]. Code and additional remarks of our DDICA is available at https://github.com/hongmingli1995/DDICA. Shujian Yu, José C. Príncipe |
ICASSP | 2 |
| 2022 | Multi-View Information Bottleneck Without Variational ApproximationabstractBy "intelligently" fuse the complementary information across different views, multi-view learning is able to improve the performance of classification task. In this work, we extend the information bottleneck principle to supervised multi-view learning scenario and use the recently proposed matrix-based Rényi’s α-order entropy functional to optimize the resulting objective directly, without the necessity of variational approximation or adversarial training. Empirical results in both synthetic and real-world datasets suggest that our method enjoys improved robustness to noise and redundant information in each view, especially given limited training samples. Code is available at https://github.com/archy666/MEIB. Qi Zhang 0089, Shujian Yu, Jingmin Xin, Badong Chen |
ICASSP | 2 |
| 2022 | Modular-Relatedness for Continual Learning
Ammar Shaker, Francesco Alesiani, Shujian Yu |
IDA | 3 |
| 2022 | Principle of relevant information for graph sparsificationabstractGraph sparsification aims to reduce the number of edges of a graph while maintaining its structural properties. In this paper, we propose the first general and effective information-theoretic formulation of graph sparsification, by taking inspiration from the Principle of Relevant Information (PRI). To this end, we extend the PRI from a standard scalar random variable setting to structured data (i.e., graphs). Our Graph-PRI objective is achieved by operating on the graph Laplacian, made possible by expressing the graph Laplacian of a subgraph in terms of a sparse edge selection vector w. We provide both theoretical and empirical justifications on the validity of our Graph-PRI approach. We also analyze its analytical solutions in a few special cases. We finally present three representative real-world applications, namely graph sparsification, graph regularized multi-task learning, and medical imaging-derived brain network classification, to demonstrate the effectiveness, the versatility and the enhanced interpretability of our approach over prevalent sparsification techniques. Code of Graph-PRI is available at https://github.com/SJYuCNEL/PRI-Graphs. Shujian Yu, Francesco Alesiani, Wenzhe Yin, Robert Jenssen, José C. Príncipe |
UAI | 1 |
| 2022 | Causality detection with matrix-based transfer entropy
Wanqi Zhou, Shujian Yu, Badong Chen |
Inf. Sci. | 2 |
| 2022 | Modularizing Deep Learning via Pairwise Learning With KernelsabstractBy redefining the conventional notions of layers, we present an alternative view on finitely wide, fully trainable deep neural networks as stacked linear models in feature spaces, leading to a kernel machine interpretation. Based on this construction, we then propose a provably optimal modular learning framework for classification that does not require between-module backpropagation. This modular approach brings new insights into the label requirement of deep learning (DL). It leverages only implicit pairwise labels (weak supervision) when learning the hidden modules. When training the output module, on the other hand, it requires full supervision but achieves high label efficiency, needing as few as ten randomly selected labeled examples (one from each class) to achieve 94.88% accuracy on CIFAR-10 using a ResNet-18 backbone. Moreover, modular training enables fully modularized DL workflows, which then simplify the design and implementation of pipelines and improve the maintainability and reusability of models. To showcase the advantages of such a modularized workflow, we describe a simple yet reliable method for estimating reusability of pretrained modules as well as task transferability in a transfer learning setting. At practically no computation overhead, it precisely described the task space structure of 15 binary classification tasks from CIFAR-10. Shiyu Duan, Shujian Yu, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Measuring Dependence with Matrix-based Entropy FunctionalabstractMeasuring the dependence of data plays a central role in statistics and machine learning. In this work, we summarize and generalize the main idea of existing information-theoretic dependence measures into a higher-level perspective by the Shearer's inequality. Based on our generalization, we then propose two measures, namely the matrix-based normalized total correlation and the matrix-based normalized dual total correlation, to quantify the dependence of multiple variables in arbitrary dimensional space, without explicit estimation of the underlying data distributions. We show that our measures are differentiable and statistically more powerful than prevalent ones. We also show the impact of our measures in four different machine learning problems, namely the gene regulatory network inference, the robust machine learning under covariate shift and non-Gaussian noises, the subspace outlier detection, and the understanding of the learning dynamics of convolutional neural networks, to demonstrate their utilities, advantages, as well as implications to those problems. Shujian Yu, Francesco Alesiani, Robert Jenssen, José C. Príncipe |
AAAI | 1 |
| 2021 | Deep Deterministic Information Bottleneck with Matrix-Based Entropy FunctionalabstractWe introduce the matrix-based Rényi’s α-order entropy functional to parameterize Tishby et al. information bottleneck (IB) principle [1] with a neural network. We term our methodology Deep Deterministic Information Bottleneck (DIB), as it avoids variational inference and distribution assumption. We show that deep neural networks trained with DIB outperform the variational objective counterpart and those that are trained with other forms of regularization, in terms of generalization performance and robustness to adversarial attack. Code available at https://github.com/yuxi120407/DIB. Shujian Yu, José C. Príncipe |
ICASSP | 2 |
| 2021 | Gated Information Bottleneck for Generalization in Sequential EnvironmentsabstractDeep neural networks suffer from poor generalization to unseen environments when the underlying data distribution is different from that in the training set. By learning minimum sufficient representations from training data, the information bottleneck (IB) approach has demonstrated its effectiveness to improve generalization in different AI applications. In this work, we propose a new neural network-based IB approach, termed gated information bottleneck (GIB), that dynamically drops spurious correlations and progressively selects the most task-relevant features across different environments by a trainable soft mask (on raw features). GIB enjoys a simple and tractable objective, without any variational approximation or distributional assumption. We empirically demonstrate the superiority of GIB over other popular neural network-based IB approaches in adversarial robustness and out-of-distribution (OOD) detection. Meanwhile, we also establish the connection between IB theory and invariant causal representation learning, and observed that GIB demonstrates appealing performance when different environments arrive sequentially, a more practical scenario where invariant risk minimization (IRM) fails. Francesco Alesiani, Shujian Yu |
ICDM | 2 |
| 2021 | Information-Theoretic Methods in Deep Neural Networks: Recent Advances and Emerging OpportunitiesabstractWe present a review on the recent advances and emerging opportunities around the theme of analyzing deep neural networks (DNNs) with information-theoretic methods. We first discuss popular information-theoretic quantities and their estimators. We then introduce recent developments on information-theoretic learning principles (e.g., loss functions, regularizers and objectives) and their parameterization with DNNs. We finally briefly review current usages of information-theoretic concepts in a few modern machine learning problems and list a few emerging opportunities. Shujian Yu, Luis Gonzalo Sánchez Giraldo, José C. Príncipe |
IJCAI | 1 |
| 2021 | Bilevel Continual LearningabstractContinual Learning (CL) studies the problem of learning a sequence of tasks, one at a time, such that the learning of each new task does not lead to the deterioration in performance on the previously seen ones while exploiting previously learned features. This paper presents Bilevel Continual Learning (BiCL), a general framework for continual learning that fuses bilevel optimization and recent advances in meta-learning for deep neural networks. BiCL is able to train both deep discriminative and generative models under the conservative setting of the online continual learning. Experimental results show that BiCL provides competitive performance in terms of accuracy for the current task while reducing the effect of catastrophic forgetting. Ammar Shaker, Francesco Alesiani, Shujian Yu, Wenzhe Yin |
IJCNN | 3 |
| 2021 | Understanding Convolutional Neural Networks With Information Theory: An Initial ExplorationabstractA novel functional estimator for Rényi's α -entropy and its multivariate extension was recently proposed in terms of the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the utility and possible applications of these new estimators are rather new and mostly unknown to practitioners. In this brief, we first show that this estimator enables straightforward measurement of information flow in realistic convolutional neural networks (CNNs) without any approximation. Then, we introduce the partial information decomposition (PID) framework and develop three quantities to analyze the synergy and redundancy in convolutional layer representations. Our results validate two fundamental data processing inequalities and reveal more inner properties concerning CNN training. Shujian Yu, Kristoffer Wickstrøm, Robert Jenssen, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Composite Dynamic Texture Synthesis Using Hierarchical Linear Dynamical SystemabstractWe demonstrate that a systematic inclusion of prior structural constraints on the states of a linear dynamical system significantly improves its ability to model complex multidimensional sequences. This constrained LDS, typically termed as the hierarchical linear dynamical system (HLDS), is a Kalman filter based topology that extracts relevant self-segmenting information from the input signal in an unsupervised manner by hierarchically constraining its information representing state subspaces thereby slowing down the signal dynamics. We highlight some of its practical advantages over the existing methods in real-world video applications. As a concrete application, we show that the HLDS, despite being a linear model trained in an unsupervised setting, is able to capture the dynamics of complex texture sequences consisting of multiple co-occurring textures. We compare its performance with a similarly trained LDS model in the reconstruction and synthesis of such signals. Rishabh Singh, Shujian Yu, José C. Príncipe |
ICASSP | 2 |
| 2020 | Measuring the Discrepancy between Conditional Distributions: Methods, Properties and ApplicationsabstractWe propose a simple yet powerful test statistic to quantify the discrepancy between two conditional distributions. The new statistic avoids the explicit estimation of the underlying distributions in high-dimensional space and it operates on the cone of symmetric positive semidefinite (SPS) matrix using the Bregman matrix divergence. Moreover, it inherits the merits of the correntropy function to explicitly incorporate high-order statistics in the data. We present the properties of our new statistic and illustrate its connections to prior art. We finally show the applications of our new statistic on three different machine learning problems, namely the multi-task learning over graphs, the concept drift detection, and the information-theoretic feature selection, to demonstrate its utility and advantage. Code of our statistic is available at https://bit.ly/BregmanCorrentropy. Shujian Yu, Ammar Shaker, Francesco Alesiani, José C. Príncipe |
IJCAI | 1 |
| 2020 | Online Meta-Forest for Regression Data StreamsabstractStream learning is essential when there is limited memory, time and computational power. However, existing streaming methods are mostly designed for classification with only a few exceptions for regression problems. Although being fast, the performance of these online regression methods is inadequate due to their dependence on merely linear models. Besides, only a few stream methods are based on meta-learning that aims at facilitating the dynamic choice of the right model. Nevertheless, these approaches are restricted to recommend learners on a window and not on the instance level. In this paper, we present a novel approach, named Online Meta-Forest, that incrementally induces an ensemble of meta-learners that selects the best set of predictors for each test example. Each meta-learner has the ability to find a non-linear mapping of the input space to the set of induced models. We conduct a series of experiments demonstrating that Online Meta-Forest outperforms related methods on 16 out of 25 evaluated benchmark and domain datasets in transportation. Ammar Shaker, Christoph Gärtner, Shujian Yu |
IJCNN | 4 |
| 2020 | Towards Interpretable Multi-task Learning Using Bilevel Programming
Francesco Alesiani, Shujian Yu, Ammar Shaker, Wenzhe Yin |
ECML/PKDD (2) | 2 |
| 2020 | Coarse-to-fine salient object detection with low-rank matrix recovery
Qi Zheng 0003, Shujian Yu, Xinge You |
Neurocomputing | 2 |
| 2020 | On Kernel Method-Based Connectionist Models and Supervised Deep Learning Without BackpropagationabstractWe propose a novel family of connectionist models based on kernel machines and consider the problem of learning layer by layer a compositional hypothesis class (i.e., a feedforward, multilayer architecture) in a supervised setting. In terms of the models, we present a principled method to “kernelize” (partly or completely) any neural network (NN). With this method, we obtain a counterpart of any given NN that is powered by kernel machines instead of neurons. In terms of learning, when learning a feedforward deep architecture in a supervised setting, one needs to train all the components simultaneously using backpropagation (BP) since there are no explicit targets for the hidden layers (Rumelhart, Hinton, & Williams, 1986 ). We consider without loss of generality the two-layer case and present a general framework that explicitly characterizes a target for the hidden layer that is optimal for minimizing the objective function of the network. This characterization then makes possible a purely greedy training scheme that learns one layer at a time, starting from the input layer. We provide instantiations of the abstract framework under certain architectures and objective functions. Based on these instantiations, we present a layer-wise training algorithm for an [Formula: see text]-layer feedforward network for classification, where [Formula: see text] can be arbitrary. This algorithm can be given an intuitive geometric interpretation that makes the learning dynamics transparent. Empirical results are provided to complement our theory. We show that the kernelized networks, trained layer-wise, compare favorably with classical kernel machines as well as other connectionist models trained by BP. We also visualize the inner workings of the greedy kernelized models to validate our claim on the transparency of the layer-wise algorithm. Shiyu Duan, Shujian Yu, Yunmei Chen, José C. Príncipe |
Neural Comput. | 2 |
| 2020 | Multivariate Extension of Matrix-Based Rényi's $\alpha$α-Order Entropy FunctionalabstractThe matrix-based Rényi's α-order entropy functional was recently introduced using the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the current theory in the matrix-based Rényi's α-order entropy functional only defines the entropy of a single variable or mutual information between two random variables. In information theory and machine learning communities, one is also frequently interested in multivariate information quantities, such as the multivariate joint entropy and different interactive quantities among multiple variables. In this paper, we first define the matrix-based Rényi's α-order joint entropy among multiple variables. We then show how this definition can ease the estimation of various information quantities that measure the interactions among multiple variables, such as interactive information and total correlation. We finally present an application to feature selection to show how our definition provides a simple yet powerful way to estimate a widely-acknowledged intractable quantity from data. A real example on hyperspectral image (HSI) band selection is also provided. Shujian Yu, Luis Gonzalo Sánchez Giraldo, Robert Jenssen, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Multiview Hybrid Embedding: A Divide-and-Conquer ApproachabstractWe present a novel cross-view classification algorithm where the gallery and probe data come from different views. A popular approach to tackle this problem is the multiview subspace learning (MvSL) that aims to learn a latent subspace shared by multiview data. Despite promising results obtained on some applications, the performance of existing methods deteriorates dramatically when the multiview data is sampled from nonlinear manifolds or suffers from heavy outliers. To circumvent this drawback, motivated by the Divide-and-Conquer strategy, we propose multiview hybrid embedding (MvHE), a unique method of dividing the problem of cross-view classification into three subproblems and building one model for each subproblem. Specifically, the first model is designed to remove view discrepancy, whereas the second and third models attempt to discover the intrinsic nonlinear structure and to increase the discriminability in intraview and interview samples, respectively. The kernel extension is conducted to further boost the representation power of MvHE. Extensive experiments are conducted on four benchmark datasets. Our methods demonstrate the overwhelming advantages against the state-of-the-art MvSL-based cross-view classification approaches in terms of classification accuracy and robustness. Jiamiao Xu, Shujian Yu, Xinge You, Mengjun Leng, Xiaoyuan Jing, C. L. Philip Chen |
IEEE Trans. Cybern. | 2 |
| 2019 | Understanding autoencoders with information theoretic concepts
Shujian Yu, José C. Príncipe |
Neural Networks | 1 |
| 2019 | Robust Visual Tracking Using Multi-Frame Multi-Feature Joint ModelingabstractIt remains a huge challenge to design effective and efficient trackers under complex scenarios, including occlusions, illumination changes and pose variations. To cope with this problem, a promising solution is to integrate the temporal consistency across consecutive frames and multiple feature cues in a unified model. Motivated by this idea, we propose a novel correlation filter-based tracker in this paper, in which the temporal relatedness is reconciled under a multi-task learning framework and the multiple feature cues are modeled using a multi-view learning approach. We demonstrate that the resulting regression model can be efficiently learned by exploiting the structure of blockwise diagonal matrix. A fast blockwise diagonal matrix inversion algorithm is developed thereafter for efficient online tracking. Meanwhile, we incorporate an adaptive scale estimation mechanism to strengthen the stability of scale variation tracking. We implement our tracker using two types of features and test it on two benchmark datasets. The experimental results demonstrate the superiority of our proposed approach when compared with the other state-of-the-art trackers. Peng Zhang 0040, Shujian Yu, Jiamiao Xu, Xinge You, Xiubao Jiang, Xiaoyuan Jing, Dacheng Tao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Request-and-Reverify: Hierarchical Hypothesis Testing for Concept Drift Detection with Expensive LabelsabstractOne important assumption underlying common classification models is the stationarity of the data. However, in real-world streaming applications, the data concept indicated by the joint distribution of feature and label is not stationary but drifting over time. Concept drift detection aims to detect such drifts and adapt the model so as to mitigate any deterioration in the model's predictive performance. Unfortunately, most existing concept drift detection methods rely on a strong and over-optimistic condition that the true labels are available immediately for all already classified instances. In this paper, a novel Hierarchical Hypothesis Testing framework with Request-and-Reverify strategy is developed to detect concept drifts by requesting labels only when necessary. Two methods, namely Hierarchical Hypothesis Testing with Classification Uncertainty (HHT-CU) and Hierarchical Hypothesis Testing with Attribute-wise "Goodness-of-fit" (HHT-AG), are proposed respectively under the novel framework. In experiments with benchmark datasets, our methods demonstrate overwhelming advantages over state-of-the-art unsupervised drift detectors. More importantly, our methods even outperform DDM (the widely used supervised drift detector) when we use significantly fewer labels. Shujian Yu, José C. Príncipe |
IJCAI | 1 |
| 2018 | Co-regularized multiview nonnegative matrix factorization with correlation constraint for representation learningabstractWith the increasing availability of multiview nonnegative data in real applications, multiview representation learning based on nonnegative matrix factorization (NMF) has attracted more and more attentions. However, existing NMF-based methods are sensitive to noises and are difficult to generate discriminative features with noisy views. To address these problems, we propose a co-regularized multiview nonnegative matrix factorization method with correlation constraint for nonnegative representation learning, which jointly exploits consistent and complementary information across different views. Different from previous works, we aim at integrating information from multiple views efficiently and making it more robust to the presence of noisy views. More specifically, we exploit the complementary information of multiple views through the co-regularization to accommodate the presence of the noisy views. Meanwhile, correlation constraint is imposed on the low-dimensional space to learn a common latent representation shared by different views. For the induced objective function, we derive an alternative algorithm to solve the optimization problem. The experimental results on four real datasets demonstrate the effectiveness and robustness of the proposed algorithm. Weihua Ou, Shujian Yu, Pengpeng Wang |
Multim. Tools Appl. | 4 |
| 2018 | Multi-view manifold learning with locality alignment
Xinge You, Shujian Yu, Chang Xu 0002, Wei Yuan 0001, Xiaoyuan Jing, Taiping Zhang, Dacheng Tao |
Pattern Recognit. | 3 |
| 2017 | Robust linear discriminant analysis with a Laplacian assumption on projection distributionabstractLinear discriminant analysis (LDA) is typically carried out using Fisher's method, which relies heavily on the estimation of sample mean vectors and covariance matrices. However, Fisher LDA is vulnerable to outliers as it happens to other multivariate statistical methods. In this paper, we analyzed the optimal discriminant design based on the criterion of minimizing total misclassification rate, assuming that the projected samples follow Laplacian distribution. The corresponding optimization objective can be approximated as a linear programming problem. We illustrated the relations of our proposed discriminant to Fisher LDA and minimax probability machine (MPM) from the perspective of projection-pursuit. Experiments on 6 real world benchmark dataset from UCI repository validate the effectiveness of our method. Shujian Yu, Xiubao Jiang |
ICASSP | 1 |
| 2017 | Autoencoders trained with relevant information: Blending Shannon and Wiener's perspectivesabstractIt is almost seventy years after the publication of Claude Shannon's “A Mathematical Theory of Communication” [1] and Norbert Wiener's “Extrapolation, Interpolation and Smoothing of Stationary Time Series” [2]. The pioneering works of Shannon and Wiener lay the foundation of communication, data storage, control, and other information technologies. This paper briefly reviews Shannon and Wiener's perspectives on the problem of message transmission over noisy channel and also experimentally evaluates the feasibility of integrating these two perspectives to train autoencoders close to the information limit. To this end, the principle of relevant information (PRI) is used and validated to optimally encode input imagery in the presence of noise. Shujian Yu, Matthew Emigh, Eder Santana, José C. Príncipe |
ICASSP | 1 |
| 2017 | Concept Drift Detection with Hierarchical Hypothesis TestingabstractWhen using statistical models (such as a classifier) in a streaming environment, there is often a need to detect and adapt to concept drifts to mitigate any deterioration in the model's predictive performance over time. Unfortunately, the ability of popular concept drift approaches in detecting these drifts in the relationship of the response and predictor variable is often dependent on the distribution characteristics of the data streams, as well as its sensitivity on parameter tuning. This paper presents Hierarchical Linear Four Rates (HLFR), a framework that detects concept drifts for different data stream distributions (including imbalanced data) by leveraging a hierarchical set of hypothesis tests in an online setting. The performance of HLFR is compared to benchmark approaches using both simulated and real-world datasets spanning the breadth of concept drift types. HLFR significantly outperforms benchmark approaches in terms of accuracy, G-mean, recall, delay in detection and adaptability across the various datasets. Shujian Yu, Zubin Abraham |
SDM | 1 |
| 2016 | Multiple adaptive kernel size KLMS for Beijing PM2.5 predictionabstractThe kernel least mean square (KLMS) algorithm is an efficient non-linear adaptive filter that operates in the reproducing kernel Hilbert space (RKHS). In realistic applications of system identification or time series prediction, there are usually multiple inputs that demand multiple kernels or kernel parameters. This paper proposes to use a tensor product kernel for KLMS that accommodates multiple inputs. Furthermore, instead of arbitrarily setting kernel parameters, appropriate kernel sizes can be chosen by a gradient descent based adaptive algorithm that minimizes the square of instant error, which helps KLMS to better capture the underlying system mechanism. Effectiveness of the proposed algorithm is shown by experiments conducted for both simulated dataset and an important real-world problem - Beijing PM2.5 prediction. Shujian Yu, Guibiao Xu, Badong Chen, José C. Príncipe |
IJCNN | 2 |
| 2016 | Multi-view non-negative matrix factorization by patch alignment framework with view consistency
Weihua Ou, Shujian Yu, Gai Li, Kesheng Zhang |
Neurocomputing | 2 |
| 2016 | STFT-like time frequency representations of nonstationary signal with arbitrary sampling schemes
Shujian Yu, Xinge You, Weihua Ou, Xiubao Jiang, Yi Mou |
Neurocomputing | 1 |
| 2016 | Dynamic texture modeling and synthesis using multi-kernel Gaussian process dynamic model
Xinge You, Shujian Yu, Jixin Zou, Haiquan Zhao 0001 |
Signal Process. | 3 |
| 2016 | Kernel Learning for Dynamic Texture SynthesisabstractDynamic textures (DTs) that represent moving scenes such as flames, smoke, and waves, exhibit fixed dynamics within a period of time and have been successfully modeled using linear dynamic systems (LDS). In this paper, we show that the widely used LDS model can be approximated using a principal component regression (PCR) model with the main advantage of simplicity. Furthermore, to capture the nonlinearity of training frames, we extend traditional PCR to its kernelized version and introduce kernel principal component regression (KPCR) to model and synthesize DTs. To ensure algorithm stability, we remove the standard state model and directly apply the quantized kernel least mean squares algorithm from signal processing domain to approximate the performance achieved with KPCR. We term this improvement kernel adaptive dynamic texture synthesis (KADTS), which also has the benefits of computational and memory efficiency. These advantages make KADTS ideally suited for real-world applications, since the majority of electronic devices, including cell phones and laptops, suffer from limited memory and real-time constraints. We demonstrate, via both theoretical and experimental analyses, the connections between DT synthesis using KPCR and KADTS with a regularization network theory. We also show the superiority of our proposed algorithms for DT synthesis compared with other dynamic system-based benchmarks. MATLAB code is available from our project homepage http://bmal.hust.edu.cn/project/dts.html. Xinge You, Weigang Guo, Shujian Yu, Kan Li 0002, José C. Príncipe, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2015 | Robust Discriminative Nonnegative Patch Alignment for Occluded Face Recognition
Weihua Ou, Gai Li, Shujian Yu, Fujia Ren, Yuan Yan Tang |
ICONIP (4) | 3 |
| 2015 | Webcam-Based Visual Gaze Estimation Under Desktop Environment
Shujian Yu, Weihua Ou, Xinge You, Xiubao Jiang, Yi Mou, Weigang Guo, Yuan Yan Tang, C. L. Philip Chen |
ICONIP (2) | 1 |
| 2015 | Generalized Kernel Normalized Mixed-Norm Algorithm: Analysis and Simulations
Shujian Yu, Xinge You, Xiubao Jiang, Weihua Ou, Yixiao Zhao, C. L. Philip Chen, Yuan Yan Tang |
ICONIP (2) | 1 |
| 2015 | Kernel normalized mixed-norm algorithm for system identificationabstractKernel methods provide an efficient nonparametric model to produce adaptive nonlinear filtering (ANF) algorithms. However, in practical applications, standard squared error based kernel methods suffer from two main issues: (1) a constant step size is used, which degrades the algorithm performance in non-stationary environment, and (2) additive noises are assumed to follow Gaussian distribution, while in practice the noises are generally non-Gaussian and follow other statistical distributions. To address these two issues simultaneously, this paper proposes a novel kernel normalized mixed-norm (KNMN) algorithm. Compared to the standard squared error based kernel methods, the KNMN algorithm extends the linear mixed-norm adaptive filtering algorithms to Reproducing Kernel Hilbert Space (RKHS) and introduces a normalized step size as well as adaptive mixing parameter. We also conduct the mean square convergence analysis and demonstrate the desirable performance of the KNMN algorithm in solving the system identification problem. Shujian Yu, Xinge You, Weihua Ou, Yuan Yan Tang |
IJCNN | 1 |
| 2015 | Dynamic Texture Synthesis via Image ReconstructionabstractThis paper addresses the problem of synthesizing continuous and infinitely varying stream of texture videos by doing operations on finite texture videos. Given an input texture video, such as flame, water, smoke, etc, we can synthesize a longer texture video holding the same texture appearance. Dynamic textures have been modeled as linear dynamic systems (LDS) by unfolding the video frames into column vectors and modeling their dynamic trajectory as time evolves. After the vectors are projected onto a lower dimensional space by Singular Value Decomposition (SVD), dynamic texture synthesis is achieved by driving the system with random noise. However, because of its over-simplified appearance model and under-constrained dynamic model. It is usually hard to synthesize long and visual pleasing texture video sequences. In this paper, we propose a new dynamic texture synthesis framework via creatively fitting the basic LDS with a newly developed patch reconstruction technique to efficiently enhance high quality texture details while maintaining the temporal coherence of the reconstructed texture patches. The patch reconstruction technique is inspired by locally linear embedding (LLE) and based on the assumption that small patches in the low-and high-quality images form manifolds with similar local geometry. The newly synthesized patches are finally stitched together by graph cuts to make up the output texture videos. Experiments on standard dynamic texture databases demonstrate that our method exhibits superior performance on synthesizing dynamic textures. Weigang Guo, Xinge You, Weiyong Xue, Shujian Yu, Xiubao Jiang |
SMC | 5 |
| 2015 | Human Heart Rate Estimation Using Ordinary Cameras under Natural MovementabstractNon-contact face-video based human heart rate (HR) estimation has attracted a lot of attentions in recent years. Almost all the state-of-the-art webcam or smartphone based HR estimation methods comprise three main steps: firstly, a region of interest (ROI) on the human face is detected in each video frame, then, the target signal is obtained by fusing multiple raw traces, which are extracted from the RGB channels across all the video frames, finally, HR is estimated by applying frequency analysis approach to the target signal. However, three major drawbacks impede the applicability of the current methods: (1) the performance of ROI detection is susceptible to head motion and facial expression, (2) there is still a lack of well-accepted method for fusing raw traces to form the target signal, and (3) the adopted frequency analysis approaches always provide estimation results with low resolution and high side lobes. To address these issues, we propose a novel HR estimation method which is applicable to ordinary cameras subject to natural head movement or facial expression. The proposed method features ROI detection via facial feature detection and tracking, target signal extraction via Independent Component Analysis (ICA) in the RGB channels, and HR estimation via real-valued iterative adaptive approach (RIAA). Experimental results validate the superiority of our proposed method. Shujian Yu, Xinge You, Xiubao Jiang, Yi Mou, Weihua Ou, Yuan Yan Tang, C. L. Philip Chen |
SMC | 1 |
| 2014 | Content-Adaptive Rain and Snow Removal Algorithms for Single Image
Shujian Yu, Yixiao Zhao, Yi Mou, Jinghui Wu, Xiaopeng Yang 0002, Baojun Zhao |
ISNN | 1 |
| 2014 | Data-Driven Bridge Detection in Compressed Domain from Panchromatic Satellite Imagery
Yixiao Zhao, Shujian Yu, Jinghui Wu, Zijing Chen, Xiaopeng Yang 0002, Baojun Zhao |
ISNN | 2 |