Xi Yang 0008

dblp:13/1520-8 · DBLP profile ↗
← Back
44ranked-venue papers
7as first author
37since 2021 · last 2026
0000-0002-8600-2570ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 7 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking Real Image Editing: Unleashing Diverse Editing Operators via Multi-Objective Optimization
abstract
Text-conditioned diffusion models have revolutionized the field of controllable real image editing, enabling high-fidelity and precise image manipulation. Recent methods target specific editing tasks, using internal representations from reconstruction to ensure consistency. Although effective for single tasks, they fail to balance precision and consistency across diverse image editing tasks. In this work, we propose a novel inference-time real-image editing framework that enables executing multiple editing tasks by tuning editing operators. Our key insight is to treat real image editing as a multi-objective optimization problem, optimizing editing operators for a Pareto optimal solution that balances editing accuracy and consistency at each denoising iteration. Additionally, we design a benchmark for operator-guided real-image editing that covers various local and global editing tasks. Extensive experimental evaluations demonstrate the method’s effectiveness in executing precise edits while preserving image fidelity across all tasks, thereby establishing it as the new state-of-the-art.
Xi Yang 0008, Huiru Shao, Rui Zhang 0012, Kaizhu Huang
WACV2
2026 Lena-TRNN: Exploring energy flow for time series prediction
Penglei Gao, Rui Zhang 0012, Xi Yang 0008, Zhuang Qian, Kaizhu Huang
Neural Networks3
2026 MedMAP: Promoting Incomplete Multi-Modal Brain Tumor Segmentation With Alignment
abstract
Brain tumor segmentation is often based on multiple magnetic resonance imaging (MRI). However, in clinical practice, certain modalities of MRI may be missing, which presents a more difficult scenario. To cope with this challenge, Knowledge Distillation, Domain Adaption, and Shared Latent Space have emerged as commonly promising strategies. However, recent efforts to address the missing modality problem in brain tumor segmentation typically overlook the modality gaps and thus fail to learn important invariant feature representations across different modalities. Such drawback consequently leads to limited performance for missing modality models. To ameliorate these problems, pre-trained models are used in natural visual segmentation tasks to minimize the gaps. However, promising pre-trained models are difficult to obtain in the brain tumor segmentation task due to the lack of sufficient data. Along this line, in this paper, we propose a novel paradigm that aligns latent features of involved modalities to a well-defined distribution anchor as the substitution of the pre-trained model. As a major contribution, we prove that our novel training paradigm ensures a tight evidence lower bound, thus theoretically certifying its effectiveness. Extensive experiments on different backbones validate that the proposed paradigm can enable invariant feature representations and produce models with narrowed modality gaps. Models with our alignment paradigm show their superior performance on both BraTS2018, BraTS2020 and Brain Metastasis datasets.
Zhaorui Tan, Muyin Chen, Xi Yang 0008, Haochuan Jiang, Kaizhu Huang
IEEE J. Biomed. Health Informatics4
2025 Disentangling Tabular Data Towards Better One-Class Anomaly Detection
abstract
Tabular anomaly detection under the one-class classification setting poses a significant challenge, as it involves accurately conceptualizing "normal" derived exclusively from a single category to discern anomalies from normal data variations. Capturing the intrinsic correlation among attributes within normal samples presents one promising method for learning the concept. To do so, the most recent effort relies on a learnable mask strategy with a reconstruction task. However, this wisdom may suffer from the risk of producing uniform masks, i.e., essentially nothing is masked, leading to less effective correlation learning. To address this issue, we presume that attributes related to others in normal samples can be divided into two non-overlapping and correlated subsets, defined as CorrSets, to capture the intrinsic correlation effectively. Accordingly, we introduce an innovative method that disentangles CorrSets from normal tabular data. To our knowledge, this is a pioneering effort to apply the concept of disentanglement for one-class anomaly detection on tabular data. Extensive experiments on 20 tabular datasets show that our method substantially outperforms the state-of-the-art methods and leads to an average performance improvement of 6.1% on AUC-PR and 2.1% on AUC-ROC.
Jianan Ye, Zhaorui Tan, Yijie Hu, Xi Yang 0008, Kaizhu Huang
AAAI4
2025 PO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly Detection
abstract
Point cloud anomaly detection under the anomaly-free setting poses significant challenges as it requires accurately capturing the features of 3D normal data to identify deviations indicative of anomalies. Current efforts focus on devising reconstruction tasks, such as acquiring normal data representations by restoring normal samples from altered, pseudo-anomalous counterparts. Our findings reveal that distributing attention equally across normal and pseudo-anomalous data tends to dilute the model’s focus on anomalous deviations. The challenge is further compounded by the inherently disordered and sparse nature of 3D point cloud data. In response to those predicaments, we introduce an innovative approach that emphasizes learning point offsets, targeting more informative pseudo-abnormal points, thus fostering more effective distillation of normal data representations. We also have crafted an augmentation technique that is steered by normal vectors, facilitating the creation of credible pseudo anomalies that enhance the efficiency of the training process. Our comprehensive experimental evaluation on the Anomaly-ShapeNet and Real3DAD datasets evidences that our proposed method outperforms existing state-of-the-art approaches, achieving an average enhancement of 9.0% and 1.4% in the AUC-ROC detection metric across these datasets, respectively. Code is available at https://github.com/yjnanan/PO3AD.
Jianan Ye, Weiguang Zhao, Xi Yang 0008, Kaizhu Huang
CVPR3
2025 Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant Representation
Zhaorui Tan, Xi Yang 0008, Tan Pan, Chen Jiang 0006, Xin Guo 0010, Qiufeng Wang 0001, Anh Nguyen 0003, Yuan Qi 0001, Kaizhu Huang
ICCV2
2025 From 2D Images to 3D Model: Weakly Supervised Multi-View Face Reconstruction with Deep Fusion
abstract
While weakly supervised multi-view face reconstruction (MVR) is garnering increased attention, one critical issue still remains open: how to effectively interact and fuse multiple image information to reconstruct high-precision 3D models. In this regard, we propose a novel pipeline called Deep Fusion MVR (DF-MVR) to explore the feature correspondences between multi-view images and reconstruct high-precision 3D faces. Specifically, we present a novel multi-view feature fusion backbone that utilizes face masks to align features from multiple encoders and integrates one multi-layer attention mechanism to enhance feature interaction and fusion, resulting in one unified facial representation. Additionally, we develop one concise face mask mechanism that facilitates multi-view feature fusion and facial reconstruction by identifying common areas and guiding the network’s focus on critical facial features (e.g., eyes, brows, nose, and mouth). Experiments on Pixel-Face and Bosphorus datasets indicate the superiority of the proposed method. Without the 3D annotation, DF-MVR achieves relative 5.2% and 3.0% RMSE improvement over the existing weakly supervised MVRs, respectively, on Pixel-Face and Bosphorus datasets. Our code is available at https://github.com/weiguangzhao/DF_MVR.
Weiguang Zhao, Chaolong Yang, Jianan Ye, Rui Zhang 0012, Yuyao Yan, Xi Yang 0008, Bin Dong 0003, Amir Hussain 0001, Kaizhu Huang
ICME6
2025 Stock Price Prediction with Attention-Based Framework by Integrating LLM-Generated Features
Yining Sun, Penglei Gao, Yuyao Yan, Xi Yang 0008
ICONIP (5)4
2025 Entropy-Guided Distillation for Medical Image Segmentation Under Missing Modalities
Yuyao Yan, Xi Yang 0008, Kaizhu Huang
ICONIP (5)3
2025 A Novel Bi-environmental Intuitionistic Fuzzy C-Means Clustering Algorithm
Yihao Zhang 0005, Youpeng Yang, Hao Lan Zhang 0001, Dongming Lu, Taoyu Wu, Xi Yang 0008
ICONIP (1)6
2025 Towards Training-Free Open-World Classification with 3D Generative Models
abstract
3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring robust subsequent knowledge adaptation capabilities. While current approaches predominantly rely on 2D pre-trained models through 3D-to-2D projection, their performance degrades severely under arbitrary object orientations. Unlike these present efforts, this work makes a pioneering exploration of 3D generative models for 3D open-world classification-specifically, leverageing the accumulated prior knowledge from these models to provide anchors for novel categories, while integrating a rotation-invariant feature extractor. This innovative synergy endows our pipeline with the advantages of being training-free and pose-invariant, thus well suited to adapt novel categories in 3D open-world classification. Extensive experiments on benchmark datasets demonstrate the potential of this pipeline, achieving state-of-the-art performance on ModelNet10‡ and McGill‡ with 32.7% and 8.7% overall accuracy improvement, respectively. The code is available in the supplementary materials.
Xinzhe Xia, Weiguang Zhao, Yuyao Yan, Guanyu Yang 0002, Rui Zhang 0012, Kaizhu Huang, Xi Yang 0008
ACM Multimedia7
2025 Revisiting 3D point cloud analysis with Markov process
Chenru Jiang, Wuwei Ma, Kaizhu Huang, Qiufeng Wang 0001, Xi Yang 0008, Weiguang Zhao, Junwei Wu 0001, Xinheng Wang 0001, Jimin Xiao, Zhenxing Niu
Pattern Recognit.5
2025 SCMix: Stochastic Compound Mixing for Open Compound Domain Adaptation in Semantic Segmentation
abstract
Open compound domain adaptation (OCDA) aims to transfer knowledge from a labeled source domain to a mix of unlabeled homogeneous compound target domains while generalizing to open unseen domains. Existing OCDA methods solve the intradomain gaps by a divide-and-conquer strategy, which decomposes the problem into several individual and parallel domain adaptation (DA) tasks. In this work, starting from the general DA theory, we establish a novel generalization bound for the setting of OCDA. Built upon this, we argue that conventional OCDA approaches may substantially underestimate the inherent variance inside the compound target domains for model generalization, constraining the model's performance. We subsequently present stochastic compound mixing (SCMix), an augmentation strategy with the primary objective of mitigating the divergence between the source and mixed target distributions. Theoretical analyses are conducted to substantiate the superiority of SCMix, proving that single-target mixing is a subgroup of our method. Extensive experiments show that our method attains a lower empirical risk on OCDA semantic segmentation tasks, thus supporting our theories. In particular, combining the transformer architecture, SCMix achieves a notable performance boost compared to SoTA results.
Zhaorui Tan, Zixian Su, Xi Yang 0008, Jie Sun 0024, Kaizhu Huang
IEEE Trans. Neural Networks Learn. Syst.4
2024 Unraveling Batch Normalization for Realistic Test-Time Adaptation
abstract
While recent test-time adaptations exhibit efficacy by adjusting batch normalization to narrow domain disparities, their effectiveness diminishes with realistic mini-batches due to inaccurate target estimation. As previous attempts merely introduce source statistics to mitigate this issue, the fundamental problem of inaccurate target estimation still persists, leaving the intrinsic test-time domain shifts unresolved. This paper delves into the problem of mini-batch degradation. By unraveling batch normalization, we discover that the inexact target statistics largely stem from the substantially reduced class diversity in batch. Drawing upon this insight, we introduce a straightforward tool, Test-time Exponential Moving Average (TEMA), to bridge the class diversity gap between training and testing batches. Importantly, our TEMA adaptively extends the scope of typical methods beyond the current batch to incorporate a diverse set of class information, which in turn boosts an accurate target estimation. Built upon this foundation, we further design a novel layer-wise rectification strategy to consistently promote test-time performance. Our proposed method enjoys a unique advantage as it requires neither training nor tuning parameters, offering a truly hassle-free solution. It significantly enhances model robustness against shifted domains and maintains resilience in diverse real-world scenarios with various batch sizes, achieving state-of-the-art performance on several major benchmarks. Code is available at https://github.com/kiwi12138/RealisticTTA.
Zixian Su, Jingwei Guo 0001, Xi Yang 0008, Qiufeng Wang 0001, Kaizhu Huang
AAAI4
2024 Semantic-Aware Data Augmentation for Text-to-Image Synthesis
abstract
Data augmentation has been recently leveraged as an effective regularizer in various vision-language deep neural networks. However, in text-to-image synthesis (T2Isyn), current augmentation wisdom still suffers from the semantic mismatch between augmented paired data. Even worse, semantic collapse may occur when generated images are less semantically constrained. In this paper, we develop a novel Semantic-aware Data Augmentation (SADA) framework dedicated to T2Isyn. In particular, we propose to augment texts in the semantic space via an Implicit Textual Semantic Preserving Augmentation, in conjunction with a specifically designed Image Semantic Regularization Loss as Generated Image Semantic Conservation, to cope well with semantic mismatch and collapse. As one major contribution, we theoretically show that Implicit Textual Semantic Preserving Augmentation can certify better text-image consistency while Image Semantic Regularization Loss regularizing the semantics of generated images would avoid semantic collapse and enhance image quality. Extensive experiments validate that SADA enhances text-image consistency and improves image quality significantly in T2Isyn models across various backbones. Especially, incorporating SADA during the tuning process of Stable Diffusion models also yields performance improvements.
Zhaorui Tan, Xi Yang 0008, Kaizhu Huang
AAAI2
2024 Mind the Gap: Promoting Missing Modality Brain Tumor Segmentation with Alignment
abstract
Brain tumor segmentation is often based on multiple magnetic resonance imaging (MRI). However, in clinical practice, certain modalities of MRI may be missing, which presents an even more difficult scenario. To cope with this challenge, knowledge distillation has emerged as one promising strategy. However, recent efforts typically overlook the modality gaps and thus fail to learn invariant feature representations across different modalities. Such drawback consequently leads to limited performance for both teachers and students. To ameliorate these problems, in this paper, we propose a novel paradigm that aligns latent features of involved modalities to a well-defined distribution anchor. As a major contribution, we prove that our novel training paradigm ensures a tight evidence lower bound, thus theoretically certifying its effectiveness. Extensive experiments on different backbones validate that the proposed paradigm can enable invariant feature representations and produce a teacher with narrowed modality gaps. This further offers superior guidance for missing modality students, achieving an average improvement of 1.75 on dice score.
Zhaorui Tan, Haochuan Jiang, Xi Yang 0008, Kaizhu Huang
BIBM4
2024 Rethinking Multi-Domain Generalization with A General Learning Objective
abstract
Multi-domain generalization$(mDG)$is universally aimed to minimize the discrepancy between training and testing distributions to enhance marginal-to-label distribution mapping. However, existing$mDG$literature lacks a general learning objective paradigm and often imposes constraints on static target marginal distributions. In this paper, we propose to leverage a Y-mapping to relax the constraint. We rethink the learning objective for$mDG$and design a new general learning objective to interpret and analyze most existing$mDG$wisdom. This general objective is bifurcated into two synergistic amis: learning domain-independent conditional features and maximizing a posterior. Explorations also extend to two effective regularization terms that incorporate prior information and suppress invalid causality, alleviating the issues that come with relaxed constraints. We theoretically contribute an upper bound for the domain alignment of domain-independent conditional features, disclosing that many previous$mDG$endeavors actually optimize partially the objective and thus lead to limited performance. As such, our study distills a general learning objective into four practical components, providing a general, robust, and flexible mechanism to handle complex domain shifts. Extensive empirical results indicate that the proposed objective with Y -mapping leads to substantially better$mDG$performance in various downstream tasks, including regression, segmentation, and classification. Code is available at htttps://github.com/zhaorui-t.an/GMDG/tree/main.
Zhaorui Tan, Xi Yang 0008, Kaizhu Huang
CVPR2
2024 Enhancing Semantic Segmentation in Open Compound Domain Adaptation Through Mixed Image and Epistemic Uncertainty
Yiqun Ma, Siyuan Wang 0017, Xi Yang 0008, Yuyao Yan
ICONIP (11)4
2024 Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
abstract
Vision models excel in image classification but struggle to generalize to unseen data, such as classifying images from unseen domains or discovering novel categories. In this paper, we explore the relationship between logical reasoning and deep learning generalization in visual classification. A logical regularization termed L-Reg is derived which bridges a logical analysis framework to image classification. Our work reveals that L-Reg reduces the complexity of the model in terms of the feature distribution and classifier weights. Specifically, we unveil the interpretability brought by L-Reg, as it enables the model to extract the salient features, such as faces to persons, for classification. Theoretical analysis and experiments demonstrate that L-Reg enhances generalization across various scenarios, including multi-domain generalization and generalized category discovery. In complex real-world scenarios where images span unknown classes and unseen domains, L-Reg consistently improves generalization, highlighting its practical efficacy.
Zhaorui Tan, Xi Yang 0008, Qiufeng Wang 0001, Anh Nguyen 0003, Kaizhu Huang
NeurIPS2
2024 Instance-Specific Model Perturbation Improves Generalized Zero-Shot Learning
abstract
Zero-shot learning (ZSL) refers to the design of predictive functions on new classes (unseen classes) of data that have never been seen during training. In a more practical scenario, generalized zero-shot learning (GZSL) requires predicting both seen and unseen classes accurately. In the absence of target samples, many GZSL models may overfit training data and are inclined to predict individuals as categories that have been seen in training. To alleviate this problem, we develop a parameter-wise adversarial training process that promotes robust recognition of seen classes while designing during the test a novel model perturbation mechanism to ensure sufficient sensitivity to unseen classes. Concretely, adversarial perturbation is conducted on the model to obtain instance-specific parameters so that predictions can be biased to unseen classes in the test. Meanwhile, the robust training encourages the model robustness, leading to nearly unaffected prediction for seen classes. Moreover, perturbations in the parameter space, computed from multiple individuals simultaneously, can be used to avoid the effect of perturbations that are too extreme and ruin the predictions. Comparison results on four benchmark ZSL data sets show the effective improvement that the proposed framework made on zero-shot methods with learned metrics.
Guanyu Yang 0002, Kaizhu Huang, Rui Zhang 0012, Xi Yang 0008
Neural Comput.4
2024 SaliencyCut: Augmenting plausible anomalies for anomaly detection
Jianan Ye, Yijie Hu, Xi Yang 0008, Qiufeng Wang 0001, Kaizhu Huang
Pattern Recognit.3
2024 EgPDE-Net: Building Continuous Neural Networks for Time Series Prediction With Exogenous Variables
abstract
While exogenous variables have a major impact on performance improvement in time series analysis, interseries correlation and time dependence among them are rarely considered in the present continuous methods. The dynamical systems of multivariate time series could be modeled with complex unknown partial differential equations (PDEs) which play a prominent role in many disciplines of science and engineering. In this article, we propose a continuous-time model for arbitrary-step prediction to learn an unknown PDE system in multivariate time series whose governing equations are parameterized by self-attention and gated recurrent neural networks. The proposed model, exogenous-guided PDE network (EgPDE-Net), takes account of the relationships among the exogenous variables and their effects on the target series. Importantly, the model can be reduced into a regularized ordinary differential equation (ODE) problem with specially designed regularization guidance, which makes the PDE problem tractable to obtain numerical solutions and feasible to predict multiple future values of the target series at arbitrary time points. Extensive experiments demonstrate that our proposed model could achieve competitive accuracy over strong baselines: on average, it outperforms the best baseline by reducing 9.85% on RMSE and 13.98% on MAE for arbitrary-step prediction.
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Ping Guo 0002, John Yannis Goulermas, Kaizhu Huang
IEEE Trans. Cybern.2
2024 Continuous Image Outpainting with Neural ODE
abstract
Generalised image outpainting is an important and active research topic in computer vision, which aims to extend appealing content all-side around a given image. Existing state-of-the-art outpainting methods often rely on discrete extrapolation to extend the feature map in the bottleneck. They thus suffer from content unsmoothness, especially in circumstances where the outlines of objects in the extrapolated regions are incoherent with the input sub-images. To mitigate this issue, we design a novel bottleneck with Neural ODEs to make continuous extrapolation in latent space, which could be a plug-in for many deep learning frameworks. Our ODE-based network continuously transforms the state and makes accurate predictions by learning the incremental relationship among latent points, leading to both smooth and structured feature representation. Experimental results on three real-world datasets both applied on transformer-based and CNN-based frameworks show that our methods could generate more realistic and coherent images against the state-of-the-art image outpainting approaches. Our code is available at https://github.com/PengleiGao/Continuous-Image-Outpainting-with-Neural-ODE .
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Kaizhu Huang
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation
abstract
Single-source domain generalization (SDG) in medical image segmentation is a challenging yet essential task as domain shifts are quite common among clinical image datasets. Previous attempts most conduct global-only/random augmentation. Their augmented samples are usually insufficient in diversity and informativeness, thus failing to cover the possible target domain distribution. In this paper, we rethink the data augmentation strategy for SDG in medical image segmentation. Motivated by the class-level representation invariance and style mutability of medical images, we hypothesize that unseen target data can be sampled from a linear combination of C (the class number) random variables, where each variable follows a location-scale distribution at the class level. Accordingly, data augmented can be readily made by sampling the random variables through a general form. On the empirical front, we implement such strategy with constrained Bezier transformation on both global and local (i.e. class-level) regions, which can largely increase the augmentation diversity. A Saliency-balancing Fusion mechanism is further proposed to enrich the informativeness by engaging the gradient information, guiding augmentation with proper orientation and magnitude. As an important contribution, we prove theoretically that our proposed augmentation can lead to an upper bound of the generalization risk on the unseen target domain, thus confirming our hypothesis. Combining the two strategies, our Saliency-balancing Location-scale Augmentation (SLAug) exceeds the state-of-the-art works by a large margin in two challenging SDG tasks. Code is available at https://github.com/Kaiseem/SLAug.
Zixian Su, Xi Yang 0008, Kaizhu Huang, Qiufeng Wang 0001, Jie Sun 0024
AAAI3
2023 Explore Epistemic Uncertainty in Domain Adaptive Semantic Segmentation
abstract
In domain adaptive segmentation, domain shift may cause erroneous high-confidence predictions on the target domain, resulting in poor self-training. To alleviate the potential error, most previous works mainly consider aleatoric uncertainty arising from the inherit data noise. This may however lead to overconfidence in incorrect predictions and thus limit the performance. In this paper, we take advantage of Deterministic Uncertainty Methods (DUM) to explore the epistemic uncertainty, which reflects accurately the domain gap depending on the model choice and parameter fitting trained on source domain. The epistemic uncertainty on target domain is evaluated on-the-fly to facilitate online reweighting and correction in the self-training process. Meanwhile, to tackle the class-wise quantity and learning difficulty imbalance problem, we introduce a novel data resampling strategy to promote simultaneous convergence across different categories. This strategy prevents the class-level over-fitting in source domain and further boosts the adaptation performance by better quantifying the uncertainty in target domain. We illustrate the superiority of our method compared with the state-of-the-art methods.
Zixian Su, Xi Yang 0008, Jie Sun 0024, Kaizhu Huang
CIKM3
2023 Divide and Conquer: 3D Point Cloud Instance Segmentation With Point-Wise Binarization
abstract
Instance segmentation on point clouds is crucially important for 3D scene understanding. Most SOTAs adopt distance clustering, which is typically effective but does not perform well in segmenting adjacent objects with the same semantic label (especially when they share neighboring points). Due to the uneven distribution of offset points, these existing methods can hardly cluster all instance points. To this end, we design a novel divide-and-conquer strategy named PBNet that binarizes each point and clusters them separately to segment instances. Our binary clustering divides offset instance points into two categories: high and low density points (HPs vs. LPs). Adjacent objects can be clearly separated by removing LPs, and then be completed and refined by assigning LPs via a neighbor voting method. To suppress potential over-segmentation, we propose to construct local scenes with the weight mask for each instance. As a plug-in, the proposed binary clustering can replace the traditional distance clustering and lead to consistent performance gains on many mainstream baselines. A series of experiments on ScanNetV2 and S3DIS datasets indicate the superiority of our model. In particular, PBNet ranks first on the ScanNetV2 official benchmark challenge, achieving the highest mAP. Code will be available publicly at https://github.com/weiguangzhao/PBNet.
Weiguang Zhao, Yuyao Yan, Chaolong Yang, Jianan Ye, Xi Yang 0008, Kaizhu Huang
ICCV5
2023 PAG: Protecting Artworks from Personalizing Image Generative Models
Zhaorui Tan, Siyuan Wang 0017, Xi Yang 0008, Kaizhu Huang
ICONIP (4)3
2023 Adversarial Example Detection with Latent Representation Dynamic Prototype
Taowen Wang, Zhuang Qian, Xi Yang 0008
ICONIP (4)3
2023 Towards Deeper and Better Multi-view Feature Fusion for 3D Semantic Segmentation
Chaolong Yang, Yuyao Yan, Weiguang Zhao, Jianan Ye, Xi Yang 0008, Amir Hussain 0001, Bin Dong 0003, Kaizhu Huang
ICONIP (15)5
2023 Generalized image outpainting with U-transformer
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, John Yannis Goulermas, Yujie Geng, Yuyao Yan, Kaizhu Huang
Neural Networks2
2023 Towards better long-tailed oracle character recognition with adversarial data augmentation
abstract
Deciphering oracle bone script is of great significance to the study of ancient Chinese culture as well as archaeology. Although recent studies on oracle character recognition have made substantial progress, they still suffer from the long-tailed data situation that results in a noticeable performance drop on the tail classes. To mitigate this issue, we propose a generative adversarial framework to augment oracle characters in the problematic classes. In this framework, the generator produces synthetic data through convex combinations of all the available samples in the corresponding classes, and is further optimized through adversarial learning with the classifier and simultaneously the discriminator . Meanwhile, we introduce Repatch to generalize samples in the generator. Since tail classes do not have sufficient data for convex combinations , we propose the TailMix mechanism to generate suitable tail class samples from other classes. Experimental results show that our proposed algorithm obtains remarkable performance in oracle character recognition and achieves new state-of-the-art average (total) accuracy with 86.03% (89.46%), 86.54% (93.86%), 95.22% (96.17%) on the three datasets Oracle-AYNU, OBC306 and Oracle-20K, respectively.
Jing Li 0049, Qiufeng Wang 0001, Kaizhu Huang, Xi Yang 0008, Rui Zhang 0012, John Yannis Goulermas
Pattern Recognit.4
2023 Semantic Similarity Distance: Towards better text-image consistency metric in text-to-image generation
Zhaorui Tan, Xi Yang 0008, Zihan Ye, Qiufeng Wang 0001, Yuyao Yan, Anh Nguyen 0003, Kaizhu Huang
Pattern Recognit.2
2023 Mind the Gap: Alleviating Local Imbalance for Unsupervised Cross-Modality Medical Image Segmentation
abstract
Unsupervised cross-modality medical image adaptation aims to alleviate the severe domain gap between different imaging modalities without using the target domain label. A key in this campaign relies upon aligning the distributions of source and target domain. One common attempt is to enforce the global alignment between two domains, which, however, ignores the fatal local-imbalance domain gap problem, i.e., some local features with larger domain gap are harder to transfer. Recently, some methods conduct alignment focusing on local regions to improve the efficiency of model learning. While this operation may cause a deficiency of critical information from contexts. To tackle this limitation, we propose a novel strategy to alleviate the domain gap imbalance considering the characteristics of medical images, namely Global-Local Union Alignment. Specifically, a feature-disentanglement style-transfer module first synthesizes the target-like source images to reduce the global domain gap. Then, a local feature mask is integrated to reduce the 'inter-gap' for local features by prioritizing those discriminative features with larger domain gap. This combination of global and local alignment can precisely localize the crucial regions in segmentation target while preserving the overall semantic consistency. We conduct a series of experiments with two cross-modality adaptation tasks, i,e. cardiac substructure and abdominal multi-organ segmentation. Experimental results indicate that our method achieves state-of-the-art performance in both tasks.
Zixian Su, Xi Yang 0008, Qiufeng Wang 0001, Yuyao Yan, Jie Sun 0024, Kaizhu Huang
IEEE J. Biomed. Health Informatics3
2023 Explainable Tensorized Neural Ordinary Differential Equations for Arbitrary-Step Time Series Prediction
abstract
In this work, we propose a continuous neural network architecture, referred to as Explainable Tensorized Neural - Ordinary Differential Equations (ETN-ODE) network for multi-step time series prediction at arbitrary time points. Unlike existing approaches which mainly handle univariate time series for multi-step prediction, or multivariate time series for single-step predictions, ETN-ODE is capable of handling multivariate time series with arbitrary-step predictions. An additional benefit is its tandem attention mechanism, with respect to temporal and variable attention, which enable it to greatly facilitate data interpretability. Specifically, the proposed model combines an explainable tensorized gated recurrent unit with ordinary differential equations, with the derivatives of the latent states parameterized through a neural network. We quantitatively and qualitatively demonstrate the effectiveness and interpretability of ETN-ODE on one arbitrary-step prediction task and five standard multi-step prediction tasks. Extensive experiments show that the proposed method achieves very accurate predictions at arbitrary time points while attaining very competitive performance against the baseline methods in standard multi-step time series prediction.
Penglei Gao, Xi Yang 0008, Rui Zhang 0012, Kaizhu Huang, John Yannis Goulermas
IEEE Trans. Knowl. Data Eng.2
2022 3D Random Occlusion and Multi-layer Projection for Deep Multi-camera Pedestrian Localization
Ming Xu 0011, Yuyao Yan, Jeremy S. Smith, Xi Yang 0008
ECCV (10)5
2022 Outpainting by Queries
Penglei Gao, Xi Yang 0008, Jie Sun 0024, Rui Zhang 0012, Kaizhu Huang
ECCV (23)3
2022 A Novel 3D Unsupervised Domain Adaptation Framework for Cross-Modality Medical Image Segmentation
abstract
We consider the problem of volumetric (3D) unsupervised domain adaptation (UDA) in cross-modality medical image segmentation, aiming to perform segmentation on the unannotated target domain (e.g. MRI) with the help of labeled source domain (e.g. CT). Previous UDA methods in medical image analysis usually suffer from two challenges: 1) they focus on processing and analyzing data at 2D level only, thus missing semantic information from the depth level; 2) one-to-one mapping is adopted during the style-transfer process, leading to insufficient alignment in the target domain. Different from the existing methods, in our work, we conduct a first of its kind investigation on multi-style image translation for complete image alignment to alleviate the domain shift problem, and also introduce 3D segmentation in domain adaptation tasks to maintain semantic consistency at the depth level. In particular, we develop an unsupervised domain adaptation framework incorporating a novel quartet self-attention module to efficiently enhance relationships between widely separated features in spatial regions on a higher dimension, leading to a substantial improvement in segmentation accuracy in the unlabeled target domain. In two challenging cross-modality tasks, specifically brain structures and multi-organ abdominal segmentation, our model is shown to outperform current state-of-the-art methods by a significant margin, demonstrating its potential as a benchmark resource for the biomedical and health informatics research community.
Zixian Su, Kaizhu Huang, Xi Yang 0008, Jie Sun 0024, Amir Hussain 0001, Frans Coenen
IEEE J. Biomed. Health Informatics4
2019 VSB-DVM: An End-to-End Bayesian Nonparametric Generalization of Deep Variational Mixture Model
abstract
Mixture of factor analyzers is a fundamental model in unsupervised learning, which is particularly useful for high dimensional data. Recent efforts on deep auto-encoding mixture models made a fruitful progress in clustering. However, in most cases, their performance depends highly on the results of pre-training. Moreover, they tend to ignore the prior information when making clustering assignment, leading to a less strict inference and consequently limiting the performance. In this paper, we propose an end-to-end Bayesian nonparametric generalization of deep mixture model with a Variational Auto-Encoder (VAE) framework. Specifically, we develop a novel model called VSB-DVM exploiting the Variational Stick-Breaking Process to design a Deep Variational Mixture Model. Distinct from the existing deep auto-encoding mixture models, this novel unsupervised deep generative model can learn low-dimensional representations and clustering simultaneously without pre-training. Importantly, a strict inference is proposed using weights of stick-breaking process in a variational way. Furthermore, able to capture the richer statistical structure of the data, VSB-DVM can also generate highly realistic samples for any specified cluster. A series of experiments are carried out, both qualitatively and quantitatively, on benchmark clustering and generation tasks. Comparative results show that the proposed model is able to generate diverse and high-quality samples of data, and also achieves encouraging clustering results outperforming the state-of-the-art algorithms on four real-world datasets.
Xi Yang 0008, Yuyao Yan, Kaizhu Huang, Rui Zhang 0012
ICDM1
2018 A new two-layer mixture of factor analyzers with joint factor loading model for the classification of small dataset problems
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas, Amir Hussain 0001
Neurocomputing1
2017 Deep Mixtures of Factor Analyzers with Common Loadings: A Novel Deep Generative Approach to Clustering
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012
ICONIP (1)1
2017 Joint Learning of Unsupervised Dimensionality Reduction and Gaussian Mixture Model
Xi Yang 0008, Kaizhu Huang, John Yannis Goulermas, Rui Zhang 0012
Neural Process. Lett.1
2016 Learning Latent Features with Infinite Non-negative Binary Matrix Tri-factorization
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, Amir Hussain 0001
ICONIP (1)1
2015 Two-layer Mixture of Factor Analyzers with Joint Factor Loading
abstract
Dimensionality Reduction (DR) is a fundamental yet active research topic in pattern recognition and machine learning. When used in classification, previous research usually performs DR separately, and then inputs the reduced features to other available models, e.g., Gaussian Mixture Model (GMM). Such independent learning could however significantly limit the classification performance, since the optimal subspace given by a particular DR approach may not be appropriate for the following classification model. More seriously, for high-dimensional data classification in the face of a limited number of samples (called small sample size or S3 problem), independent learning of DR and classification model may even deteriorate the classification accuracy. To solve this problem, we propose a joint learning model, called Two-layer Mixture of Factor Analyzers with Joint Factor Loading (2L-MJFA) for classification. More specifically, our proposed model enjoys a two-layer mixture structure, or a mixture of mixtures structure, with each component (representing each specific class) as another mixture model of Factor Analyzer (MFA). Importantly, all the involved factor analyzers are intentionally designed to share the same loading matrix. On one hand, such joint loading matrix can be considered as the dimensionality reduction matrix; on the other hand, a joint common matrix would largely reduce the parameters, making the proposed algorithm very suitable for S3 problems. We describe our model definition and propose a modified EM algorithm to optimize the model. A series of experiments demonstrates that our proposed model significantly outperforms the other three competitive algorithms on five data sets.
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012, John Yannis Goulermas
IJCNN1
2014 Unsupervised Dimensionality Reduction for Gaussian Mixture Model
Xi Yang 0008, Kaizhu Huang, Rui Zhang 0012
ICONIP (2)1