Rui Zhang 0012

dblp:60/2536-12 · DBLP profile ↗
← Back
47ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-8104-5432ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Rethinking Real Image Editing: Unleashing Diverse Editing Operators via Multi-Objective Optimization
abstract
Text-conditioned diffusion models have revolutionized the field of controllable real image editing, enabling high-fidelity and precise image manipulation. Recent methods target specific editing tasks, using internal representations from reconstruction to ensure consistency. Although effective for single tasks, they fail to balance precision and consistency across diverse image editing tasks. In this work, we propose a novel inference-time real-image editing framework that enables executing multiple editing tasks by tuning editing operators. Our key insight is to treat real image editing as a multi-objective optimization problem, optimizing editing operators for a Pareto optimal solution that balances editing accuracy and consistency at each denoising iteration. Additionally, we design a benchmark for operator-guided real-image editing that covers various local and global editing tasks. Extensive experimental evaluations demonstrate the method’s effectiveness in executing precise edits while preserving image fidelity across all tasks, thereby establishing it as the new state-of-the-art.
Xi Yang 0008, Huiru Shao, Rui Zhang 0012, Kaizhu Huang
WACV5
2026 Lena-TRNN: Exploring energy flow for time series prediction
Penglei Gao, Rui Zhang 0012, Xi Yang 0008, Zhuang Qian, Kaizhu Huang
Neural Networks2
2026 Diff-Oracle: Learning Styles and Contents to Augment Realistic Oracle Characters in Diffusion Model
abstract
Recognizing oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains because of the scarcity of oracle character images. To overcome this issue, we propose Diff-Oracle, a novel multi-modal conditional diffusion model that generates a diverse range of controllable oracle characters by inputting random combinations of references. Given the challenge of accurately describing oracle character styles using natural language, Diff-Oracle departs from traditional diffusion models that rely primarily on text prompts by introducing a style encoder. This encoder extracts style prompts from existing oracle character images, where style details are converted into a text embedding format via a pre-trained language-vision model. Additionally, given the lack of explicit content information for oracle characters, ensuring that generated characters accurately represent the intended glyphs is challenging. Therefore, we pre-generate pixel-level paired oracle character images (i.e., style and content images) by an image-to-image translation model, providing content information for the generation process. Meanwhile, Diff-Oracle integrates a content encoder designed to capture specific content details from content reference images. Extensive experiments on Oracle-241 and OBC306 datasets demonstrate that Diff-Oracle significantly outperforms existing generative methods in image quality and diversity. Moreover, Diff-Oracle substantially benefits downstream recognition tasks, outperforming all existing state-of-the-art methods by a large margin. In particular, on the challenging OBC306 dataset, Diff-Oracle achieves a 7.70% accuracy gain in the zero-shot setting and reaches 84.62% accuracy for unseen oracle characters, setting a new benchmark for oracle character recognition. The code is available at https://github.com/JJJingLi/Diff-Oracle .
Jing Li 0049, Qiufeng Wang 0001, Siyuan Wang 0017, Rui Zhang 0012, Kaizhu Huang, Erik Cambria
ACM Trans. Multim. Comput. Commun. Appl.4
2025 BFANet: Revisiting 3D Semantic Segmentation with Boundary Feature Analysis
abstract
3D semantic segmentation plays a fundamental and crucial role to understand 3D scenes. While contemporary state-of-the-art techniques predominantly concentrate on elevating the overall performance of 3D semantic segmentation based on general metrics (e.g. mIoU, mAcc, and oAcc), they unfortunately leave the exploration of challenging regions for segmentation mostly neglected. In this paper, we revisit 3D semantic segmentation through a more granular lens, shedding light on subtle complexities that are typically overshadowed by broader performance metrics. Concretely, we have delineated 3D semantic segmentation errors into four comprehensive categories as well as corresponding evaluation metrics tailored to each. Building upon this categorical framework, we introduce an innovative 3D semantic segmentation network called BFANet that incorporates detailed analysis of semantic boundary features. First, we design the boundary-semantic module to decouple point cloud features into semantic and boundary features, and fuse their query queue to enhance semantic features with attention. Second, we introduce a more concise and accelerated boundary pseudo-label calculation algorithm, which is 3.9 times faster than the state-of-the-art, offering compatibility with data augmentation and enabling efficient computation in training. Extensive experiments on benchmark data indicate the superiority of our BFANet model, confirming the significance of emphasizing the four uniquely designed metrics. Code is available at https://github.com/weiguangzhao/BFANet.
Weiguang Zhao, Rui Zhang 0012, Qiufeng Wang 0001, Kaizhu Huang
CVPR2
2025 From 2D Images to 3D Model: Weakly Supervised Multi-View Face Reconstruction with Deep Fusion
abstract
While weakly supervised multi-view face reconstruction (MVR) is garnering increased attention, one critical issue still remains open: how to effectively interact and fuse multiple image information to reconstruct high-precision 3D models. In this regard, we propose a novel pipeline called Deep Fusion MVR (DF-MVR) to explore the feature correspondences between multi-view images and reconstruct high-precision 3D faces. Specifically, we present a novel multi-view feature fusion backbone that utilizes face masks to align features from multiple encoders and integrates one multi-layer attention mechanism to enhance feature interaction and fusion, resulting in one unified facial representation. Additionally, we develop one concise face mask mechanism that facilitates multi-view feature fusion and facial reconstruction by identifying common areas and guiding the network’s focus on critical facial features (e.g., eyes, brows, nose, and mouth). Experiments on Pixel-Face and Bosphorus datasets indicate the superiority of the proposed method. Without the 3D annotation, DF-MVR achieves relative 5.2% and 3.0% RMSE improvement over the existing weakly supervised MVRs, respectively, on Pixel-Face and Bosphorus datasets. Our code is available at https://github.com/weiguangzhao/DF_MVR.
Weiguang Zhao, Chaolong Yang, Jianan Ye, Rui Zhang 0012, Yuyao Yan, Xi Yang 0008, Bin Dong 0003, Amir Hussain 0001, Kaizhu Huang
ICME4
2025 Towards Training-Free Open-World Classification with 3D Generative Models
abstract
3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring robust subsequent knowledge adaptation capabilities. While current approaches predominantly rely on 2D pre-trained models through 3D-to-2D projection, their performance degrades severely under arbitrary object orientations. Unlike these present efforts, this work makes a pioneering exploration of 3D generative models for 3D open-world classification-specifically, leverageing the accumulated prior knowledge from these models to provide anchors for novel categories, while integrating a rotation-invariant feature extractor. This innovative synergy endows our pipeline with the advantages of being training-free and pose-invariant, thus well suited to adapt novel categories in 3D open-world classification. Extensive experiments on benchmark datasets demonstrate the potential of this pipeline, achieving state-of-the-art performance on ModelNet10‡ and McGill‡ with 32.7% and 8.7% overall accuracy improvement, respectively. The code is available in the supplementary materials.
Xinzhe Xia, Weiguang Zhao, Yuyao Yan, Guanyu Yang 0002, Rui Zhang 0012, Kaizhu Huang, Xi Yang 0008
ACM Multimedia5
2025 Open-Pose 3D zero-shot learning: Benchmark and challenges
Weiguang Zhao, Guanyu Yang 0002, Rui Zhang 0012, Chenru Jiang, Chaolong Yang, Yuyao Yan, Amir Hussain 0001, Kaizhu Huang
Neural Networks3
2024 Instance-Specific Model Perturbation Improves Generalized Zero-Shot Learning
abstract
Zero-shot learning (ZSL) refers to the design of predictive functions on new classes (unseen classes) of data that have never been seen during training. In a more practical scenario, generalized zero-shot learning (GZSL) requires predicting both seen and unseen classes accurately. In the absence of target samples, many GZSL models may overfit training data and are inclined to predict individuals as categories that have been seen in training. To alleviate this problem, we develop a parameter-wise adversarial training process that promotes robust recognition of seen classes while designing during the test a novel model perturbation mechanism to ensure sufficient sensitivity to unseen classes. Concretely, adversarial perturbation is conducted on the model to obtain instance-specific parameters so that predictions can be biased to unseen classes in the test. Meanwhile, the robust training encourages the model robustness, leading to nearly unaffected prediction for seen classes. Moreover, perturbations in the parameter space, computed from multiple individuals simultaneously, can be used to avoid the effect of perturbations that are too extreme and ruin the predictions. Comparison results on four benchmark ZSL data sets show the effective improvement that the proposed framework made on zero-shot methods with learned metrics.
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, Xi Yang 0008
Neural Comput.3
2024 ES-GNN: Generalizing Graph Neural Networks Beyond Homophily With Edge Splitting
abstract
While Graph Neural Networks (GNNs) have achieved enormous success in multiple graph analytical tasks, modern variants mostly rely on the strong inductive bias of homophily. However, real-world networks typically exhibit both homophilic and heterophilic linking patterns, wherein adjacent nodes may share dissimilar attributes and distinct labels. Therefore, GNNs smoothing node proximity holistically may aggregate both task-relevant and irrelevant (even harmful) information, limiting their ability to generalize to heterophilic graphs and potentially causing non-robustness. In this work, we propose a novel Edge Splitting GNN (ES-GNN) framework to adaptively distinguish between graph edges either relevant or irrelevant to learning tasks. This essentially transfers the original graph into two subgraphs with the same node set but complementary edge sets dynamically. Given that, information propagation separately on these subgraphs and edge splitting are alternatively conducted, thus disentangling the task-relevant and irrelevant features. Theoretically, we show that our ES-GNN can be regarded as a solution to a disentangled graph denoising problem, which further illustrates our motivations and interprets the improved generalization beyond homophily. Extensive experiments over 11 benchmark and 1 synthetic datasets not only demonstrate the effective performance of ES-GNN but also highlight its robustness to adversarial graphs and mitigation of the over-smoothing problem.
Jingwei Guo 0001, Kaizhu Huang, Rui Zhang 0012, Xinping Yi
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 EgPDE-Net: Building Continuous Neural Networks for Time Series Prediction With Exogenous Variables
abstract
While exogenous variables have a major impact on performance improvement in time series analysis, interseries correlation and time dependence among them are rarely considered in the present continuous methods. The dynamical systems of multivariate time series could be modeled with complex unknown partial differential equations (PDEs) which play a prominent role in many disciplines of science and engineering. In this article, we propose a continuous-time model for arbitrary-step prediction to learn an unknown PDE system in multivariate time series whose governing equations are parameterized by self-attention and gated recurrent neural networks. The proposed model, exogenous-guided PDE network (EgPDE-Net), takes account of the relationships among the exogenous variables and their effects on the target series. Importantly, the model can be reduced into a regularized ordinary differential equation (ODE) problem with specially designed regularization guidance, which makes the PDE problem tractable to obtain numerical solutions and feasible to predict multiple future values of the target series at arbitrary time points. Extensive experiments demonstrate that our proposed model could achieve competitive accuracy over strong baselines: on average, it outperforms the best baseline by reducing 9.85% on RMSE and 13.98% on MAE for arbitrary-step prediction.
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Ping Guo 0002, John Yannis Goulermas, Kaizhu Huang
IEEE Trans. Cybern.3
2024 Learning Disentangled Graph Convolutional Networks Locally and Globally
abstract
Graph convolutional networks (GCNs) emerge as the most successful learning models for graph-structured data. Despite their success, existing GCNs usually ignore the entangled latent factors typically arising in real-world graphs, which results in nonexplainable node representations. Even worse, while the emphasis has been placed on local graph information, the global knowledge of the entire graph is lost to a certain extent. In this work, to address these issues, we propose a novel framework for GCNs, termed LGD-GCN, taking advantage of both local and global information for disentangling node representations in the latent space. Specifically, we propose to represent a disentangled latent continuous space with a statistical mixture model, by leveraging neighborhood routing mechanism locally. From the latent space, various new graphs can then be disentangled and learned, to overall reflect the hidden structures with respect to different factors. On the one hand, a novel regularizer is designed to encourage interfactor diversity for model expressivity in the latent space. On the other hand, the factor-specific information is encoded globally via employing a message passing along these new graphs, in order to strengthen intrafactor consistency. Extensive evaluations on both synthetic and five benchmark datasets show that LGD-GCN brings significant performance gains over the recent competitive models in both disentangling and node classification. Particularly, LGD-GCN is able to outperform averagely the disentangled state-of-the-arts by 7.4% on social network datasets.
Jingwei Guo 0001, Kaizhu Huang, Xinping Yi, Rui Zhang 0012
IEEE Trans. Neural Networks Learn. Syst.4
2024 Continuous Image Outpainting with Neural ODE
abstract
Generalised image outpainting is an important and active research topic in computer vision, which aims to extend appealing content all-side around a given image. Existing state-of-the-art outpainting methods often rely on discrete extrapolation to extend the feature map in the bottleneck. They thus suffer from content unsmoothness, especially in circumstances where the outlines of objects in the extrapolated regions are incoherent with the input sub-images. To mitigate this issue, we design a novel bottleneck with Neural ODEs to make continuous extrapolation in latent space, which could be a plug-in for many deep learning frameworks. Our ODE-based network continuously transforms the state and makes accurate predictions by learning the incremental relationship among latent points, leading to both smooth and structured feature representation. Experimental results on three real-world datasets both applied on transformer-based and CNN-based frameworks show that our methods could generate more realistic and coherent images against the state-of-the-art image outpainting approaches. Our code is available at https://github.com/PengleiGao/Continuous-Image-Outpainting-with-Neural-ODE .
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Kaizhu Huang
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Towards Better Robustness against Common Corruptions for Unsupervised Domain Adaptation
abstract
Recent studies have investigated how to achieve robustness for unsupervised domain adaptation (UDA). While most efforts focus on adversarial robustness, i.e. how the model performs against unseen malicious adversarial perturbations, robustness against benign common corruption (RaCC) surprisingly remains under-explored for UDA. Towards improving RaCC for UDA methods in an unsupervised manner, we propose a novel Distributionally and Discretely Adversarial Regularization (DDAR) framework in this paper. Formulated as a min-max optimization with a distribution distance, DDAR1is theoretically well-founded to ensure generalization over unknown common corruptions. Meanwhile, we show that our regularization scheme effectively reduces a surrogate of RaCC, i.e., the perceptual distance between natural data and common corruption. To enable a abetter adversarial regularization, the design of the optimization pipeline relies on an image discretization scheme that can transform "out-of-distribution" adversarial data into "in-distribution" data augmentation. Through extensive experiments, in terms of RaCC, our method is superior to conventional unsupervised regularization mechanisms, widely improves the robustness of existing UDA methods, and achieves state-of-the-art performance.
Kaizhu Huang, Rui Zhang 0012, Dawei Liu 0001, Jieming Ma
ICCV3
2023 Decoupled Learning for Long-Tailed Oracle Character Recognition
Jing Li 0049, Bin Dong 0003, Qiufeng Wang 0001, Lei Ding 0012, Rui Zhang 0012, Kaizhu Huang
ICDAR (4)5
2023 Graph Neural Networks with Diverse Spectral Filtering
abstract
Spectral Graph Neural Networks (GNNs) have achieved tremendous success in graph machine learning, with polynomial filters applied for graph convolutions, where all nodes share the identical filter weights to mine their local contexts. Despite the success, existing spectral GNNs usually fail to deal with complex networks (e.g., WWW) due to such homogeneous spectral filtering setting that ignores the regional heterogeneity as typically seen in real-world networks. To tackle this issue, we propose a novel diverse spectral filtering (DSF) framework, which automatically learns node-specific filter weights to exploit the varying local structure properly. Particularly, the diverse filter weights consist of two components — A global one shared among all nodes, and a local one that varies along network edges to reflect node difference arising from distinct graph parts — to balance between local and global information. As such, not only can the global graph characteristics be captured, but also the diverse local patterns can be mined with awareness of different node positions. Interestingly, we formulate a novel optimization problem to assist in learning diverse filters, which also enables us to enhance any spectral GNNs with our DSF framework. We showcase the proposed framework on three state-of-the-arts including GPR-GNN, BernNet, and JacobiConv. Extensive experiments over 10 benchmark datasets demonstrate that our framework can consistently boost model performance by up to 4.92% in node classification tasks, producing diverse filters with enhanced interpretability.
Jingwei Guo 0001, Kaizhu Huang, Xinping Yi, Rui Zhang 0012
WWW4
2023 Robust generative adversarial network
Shufei Zhang, Zhuang Qian, Kaizhu Huang, Rui Zhang 0012, Jimin Xiao, Canyi Lu
Mach. Learn.4
2023 Generalized image outpainting with U-transformer
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, John Yannis Goulermas, Yujie Geng, Yuyao Yan, Kaizhu Huang
Neural Networks3
2023 Towards better long-tailed oracle character recognition with adversarial data augmentation
abstract
Deciphering oracle bone script is of great significance to the study of ancient Chinese culture as well as archaeology. Although recent studies on oracle character recognition have made substantial progress, they still suffer from the long-tailed data situation that results in a noticeable performance drop on the tail classes. To mitigate this issue, we propose a generative adversarial framework to augment oracle characters in the problematic classes. In this framework, the generator produces synthetic data through convex combinations of all the available samples in the corresponding classes, and is further optimized through adversarial learning with the classifier and simultaneously the discriminator . Meanwhile, we introduce Repatch to generalize samples in the generator. Since tail classes do not have sufficient data for convex combinations , we propose the TailMix mechanism to generate suitable tail class samples from other classes. Experimental results show that our proposed algorithm obtains remarkable performance in oracle character recognition and achieves new state-of-the-art average (total) accuracy with 86.03% (89.46%), 86.54% (93.86%), 95.22% (96.17%) on the three datasets Oracle-AYNU, OBC306 and Oracle-20K, respectively.
Jing Li 0049, Qiufeng Wang 0001, Kaizhu Huang, Xi Yang 0008, Rui Zhang 0012, John Yannis Goulermas
Pattern Recognit.5
2023 Explainable Tensorized Neural Ordinary Differential Equations for Arbitrary-Step Time Series Prediction
abstract
In this work, we propose a continuous neural network architecture, referred to as Explainable Tensorized Neural - Ordinary Differential Equations (ETN-ODE) network for multi-step time series prediction at arbitrary time points. Unlike existing approaches which mainly handle univariate time series for multi-step prediction, or multivariate time series for single-step predictions, ETN-ODE is capable of handling multivariate time series with arbitrary-step predictions. An additional benefit is its tandem attention mechanism, with respect to temporal and variable attention, which enable it to greatly facilitate data interpretability. Specifically, the proposed model combines an explainable tensorized gated recurrent unit with ordinary differential equations, with the derivatives of the latent states parameterized through a neural network. We quantitatively and qualitatively demonstrate the effectiveness and interpretability of ETN-ODE on one arbitrary-step prediction task and five standard multi-step prediction tasks. Extensive experiments show that the proposed method achieves very accurate predictions at arbitrary time points while attaining very competitive performance against the baseline methods in standard multi-step time series prediction.
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Kaizhu Huang, John Yannis Goulermas
IEEE Trans. Knowl. Data Eng.3
2022 Outpainting by Queries
Penglei Gao, Xi Yang 0008, Jie Sun 0024, Rui Zhang 0012, Kaizhu Huang
ECCV (23)5
2022 Certifying Better Robust Generalization for Unsupervised Domain Adaptation
abstract
Recent studies explore how to obtain adversarial robustness for unsupervised domain adaptation (UDA). These efforts are however dedicated to achieving an optimal trade-off between accuracy and robustness on a given or seen target domain but ignore the robust generalization issue over unseen adversarial data. Consequently, degraded performance will be often observed when existing robust UDAs are applied to future adversarial data. In this work, we make a first attempt to address the robust generalization issue of UDA. We conjecture that the poor robust generalization of present robust UDAs may be caused by the large distribution gap among adversarial examples. We then provide an empirical and theoretical analysis showing that this large distribution gap is mainly owing to the discrepancy between feature-shift distributions. To reduce such discrepancy, a novel Anchored Feature-Shift Regularization (AFSR) method is designed with a certificated robust generalization bound. We conduct a series of experiments on benchmark UDA datasets. Experimental results validate the effectiveness of our proposed AFSR over many existing robust UDA methods.
Shufei Zhang, Kaizhu Huang, Qiufeng Wang 0001, Rui Zhang 0012, Chaoliang Zhong
ACM Multimedia5
2021 Mix-Up Augmentation for Oracle Character Recognition with Imbalanced Data Distribution
Jing Li 0049, Qiufeng Wang 0001, Rui Zhang 0012, Kaizhu Huang
ICDAR (1)3
2021 Towards Better Robust Generalization with Shift Consistency Regularization
abstract
While adversarial training becomes one of the most promising defending approaches against adversarial attacks for deep neural networks, the conventional wisdom through robust optimization may usually not guarantee good generalization for robustness. Concerning with robust generalization over unseen adversarial data, this paper investigates adversarial training from a novel perspective of shift consistency in latent space. We argue that the poor robust generalization of adversarial training is owing to the significantly dispersed latent representations generated by training and test adversarial data, as the adversarial perturbations push the latent features of natural examples in the same class towards diverse directions. This is underpinned by the theoretical analysis of the robust generalization gap, which is upper-bounded by the standard one over the natural data and a term of feature inconsistent shift caused by adversarial perturbation {–} a measure of latent dispersion. Towards better robust generalization, we propose a new regularization method {–} shift consistency regularization (SCR) {–} to steer the same-class latent features of both natural and adversarial data into a common direction during adversarial training. The effectiveness of SCR in adversarial training is evaluated through extensive experiments over different datasets, such as CIFAR-10, CIFAR-100, and SVHN, against several competitive methods.
Shufei Zhang, Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Rui Zhang 0012, Xinping Yi
ICML5
2021 Coarse-grained generalized zero-shot learning with efficient self-focus mechanism
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
Neurocomputing3
2021 Improving generative adversarial networks with simple latent distributions
Shufei Zhang, Kaizhu Huang, Zhuang Qian, Rui Zhang 0012, Amir Hussain 0001
Neural Comput. Appl.4
2020 Adversarial Rectification Network for Scene Text Regularization
Jing Li 0049, Qiufeng Wang 0001, Rui Zhang 0012, Kaizhu Huang
ICONIP (2)3
2020 Inductive Generalized Zero-Shot Learning with Adversarial Relation Network
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
ECML/PKDD (2)3
2020 Improving deep neural network performance by integrating kernelized Min-Max objective
Qiufeng Wang 0001, Rui Zhang 0012, Amir Hussain 0001, Kaizhu Huang
Neurocomputing3
2020 Novel deep neural network based pattern field classification architectures
Kaizhu Huang, Shufei Zhang, Rui Zhang 0012, Amir Hussain 0001
Neural Networks3
2020 Generative adversarial classifier for handwriting characters super-resolution
Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Jimin Xiao, Rui Zhang 0012
Pattern Recognit.5
2019 VSB-DVM: An End-to-End Bayesian Nonparametric Generalization of Deep Variational Mixture Model
abstract
Mixture of factor analyzers is a fundamental model in unsupervised learning, which is particularly useful for high dimensional data. Recent efforts on deep auto-encoding mixture models made a fruitful progress in clustering. However, in most cases, their performance depends highly on the results of pre-training. Moreover, they tend to ignore the prior information when making clustering assignment, leading to a less strict inference and consequently limiting the performance. In this paper, we propose an end-to-end Bayesian nonparametric generalization of deep mixture model with a Variational Auto-Encoder (VAE) framework. Specifically, we develop a novel model called VSB-DVM exploiting the Variational Stick-Breaking Process to design a Deep Variational Mixture Model. Distinct from the existing deep auto-encoding mixture models, this novel unsupervised deep generative model can learn low-dimensional representations and clustering simultaneously without pre-training. Importantly, a strict inference is proposed using weights of stick-breaking process in a variational way. Furthermore, able to capture the richer statistical structure of the data, VSB-DVM can also generate highly realistic samples for any specified cluster. A series of experiments are carried out, both qualitatively and quantitatively, on benchmark clustering and generation tasks. Comparative results show that the proposed model is able to generate diverse and high-quality samples of data, and also achieves encouraging clustering results outperforming the state-of-the-art algorithms on four real-world datasets.
Xi Yang 0008, Yuyao Yan, Kaizhu Huang, Rui Zhang 0012
ICDM4
2019 Generalized Adversarial Training in Riemannian Space
abstract
Adversarial examples, referred to as augmented data points generated by imperceptible perturbations of input samples, have recently drawn much attention. Well-crafted adversarial examples may even mislead state-of-the-art deep neural network (DNN) models to make wrong predictions easily. To alleviate this problem, many studies have focused on investigating how adversarial examples can be generated and/or effectively handled. All existing works tackle this problem in the Euclidean space. In this paper, we extend the learning of adversarial examples to the more general Riemannian space over DNNs. The proposed work is important in that (1) it is a generalized learning methodology since Riemmanian space will be degraded to the Euclidean space in a special case; (2) it is the first work to tackle the adversarial example problem tractably through the perspective of Riemannian geometry; (3) from the perspective of geometry, our method leads to the steepest direction of the loss function, by considering the second order information of the loss function. We also provide a theoretical study showing that our proposed method can truly find the descent direction for the loss function, with a comparable computational time against traditional adversarial methods. Finally, the proposed framework demonstrates superior performance over traditional counterpart methods, using benchmark data including MNIST, CIFAR-10 and SVHN.
Shufei Zhang, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICDM3
2018 W-Net: One-Shot Arbitrary-Style Chinese Character Generation with Deep Neural Networks
Haochuan Jiang, Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012
ICONIP (5)4
2018 Improving Deep Neural Network Performance with Kernelized Min-Max Objective
Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)3
2018 A new two-layer mixture of factor analyzers with joint factor loading model for the classification of small dataset problems
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
Neurocomputing3
2017 Field Support Vector Regression
Haochuan Jiang, Kaizhu Huang, Rui Zhang 0012
ICONIP (1)3
2017 Deep Mixtures of Factor Analyzers with Common Loadings: A Novel Deep Generative Approach to Clustering
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012
ICONIP (1)3
2017 Improve Deep Learning with Unsupervised Objective
Shufei Zhang, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)3
2017 Joint Learning of Unsupervised Dimensionality Reduction and Gaussian Mixture Model
Xi Yang 0008, Kaizhu Huang, John Yannis Goulermas, Rui Zhang 0012
Neural Process. Lett.4
2016 Learning Latent Features with Infinite Non-negative Binary Matrix Tri-factorization
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)3
2015 Two-layer Mixture of Factor Analyzers with Joint Factor Loading
abstract
Dimensionality Reduction (DR) is a fundamental yet active research topic in pattern recognition and machine learning. When used in classification, previous research usually performs DR separately, and then inputs the reduced features to other available models, e.g., Gaussian Mixture Model (GMM). Such independent learning could however significantly limit the classification performance, since the optimal subspace given by a particular DR approach may not be appropriate for the following classification model. More seriously, for high-dimensional data classification in the face of a limited number of samples (called small sample size or S3 problem), independent learning of DR and classification model may even deteriorate the classification accuracy. To solve this problem, we propose a joint learning model, called Two-layer Mixture of Factor Analyzers with Joint Factor Loading (2L-MJFA) for classification. More specifically, our proposed model enjoys a two-layer mixture structure, or a mixture of mixtures structure, with each component (representing each specific class) as another mixture model of Factor Analyzer (MFA). Importantly, all the involved factor analyzers are intentionally designed to share the same loading matrix. On one hand, such joint loading matrix can be considered as the dimensionality reduction matrix; on the other hand, a joint common matrix would largely reduce the parameters, making the proposed algorithm very suitable for S3 problems. We describe our model definition and propose a modified EM algorithm to optimize the model. A series of experiments demonstrates that our proposed model significantly outperforms the other three competitive algorithms on five data sets.
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas
IJCNN3
2015 Learning Imbalanced Classifiers Locally and Globally with One-Side Probability Machine
Kaizhu Huang, Rui Zhang 0012, Xu-Cheng Yin
Neural Process. Lett.2
2014 Unsupervised Dimensionality Reduction for Gaussian Mixture Model
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012
ICONIP (2)3
2014 A Novel Hybrid Approach for Combining Deep and Traditional Neural Networks
Rui Zhang 0012, Shufei Zhang, Kaizhu Huang
ICONIP (3)1
2013 One-Side Probability Machine: Learning Imbalanced Classifiers Locally and Globally
Rui Zhang 0012, Kaizhu Huang
ICONIP (2)1
2006 Combining Wavelet Analysis and Bayesian Networks for the Classification of Auditory Brainstem Response
abstract
The auditory brainstem response (ABR) has become a routine clinical tool for hearing and neurological assessment. In order to pick out the ABR from the background EEG activity that obscures it, stimulus-synchronized averaging of many repeated trials is necessary, typically requiring up to 2000 repetitions. This number of repetitions can be very difficult, time consuming and uncomfortable for some subjects. In this study, a method combining wavelet analysis and Bayesian networks is introduced to reduce the required number of repetitions, which could offer a great advantage in the clinical situation. 314 ABRs with 64 repetitions and 155 ABRs with 128 repetitions recorded from eight subjects are used here. A wavelet transform is applied to each of the ABRs, and the important features of the ABRs are extracted by thresholding and matching the wavelet coefficients. The significant wavelet coefficients that represent the extracted features of the ABRs are then used as the variables to build the Bayesian network for classification of the ABRs. In order to estimate the performance of this approach, stratified ten-fold cross-validation is used.
Rui Zhang 0012, Gerry McAllister, Bryan W. Scotney, Sally I. McClean, Glen Houston
IEEE Trans. Inf. Technol. Biomed.1
2005 Classification of the Auditory Brainstem Response (ABR) Using Wavelet Analysis and Bayesian Network
abstract
The auditory brainstem response (ABR) has become a routine clinical tool for hearing and neurological assessment. In order to pick out the ABR from the background EEG activity that obscures it, stimulus-synchronized averaging of many repeated trials is necessary and it typically requires up to 2000 repetitions. This number of repetitions can be very difficult, time consuming and uncomfortable for some subjects. In this study a method combining the wavelet analysis and the Bayesian network is introduced to reduce the required number of repetitions, which could offer a great advantage in the clinical situation. The important features of the ABR are extracted by thresholding and matching the wavelet coefficients. These extracted features are then used as the variables to build up the Bayesian network for classifying the ABR. 172 ABRs with 64 repetitions are applied in this study to learn the Bayesian network and estimate the conditional probability tables (CPTs). A further 142 ABRs with 64 repetitions are used to test the network. Moreover, this Bayesian network can also be applied to classify the ABRs with 128 repetitions.
Rui Zhang 0012, Gerry McAllister, Bryan W. Scotney, Sally I. McClean, Glen Houston
CBMS1