Ming Yang 0014

dblp:98/2604-14 · DBLP profile ↗
← Back
59ranked-venue papers
2as first author
30since 2021 · last 2026
0000-0001-8936-4270ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 14 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Cluster-aware contrastive learning for partially view-aligned clustering
abstract
Multi-view clustering aims to discover shared semantic information from different perspectives, but real-world time and space factors often lead to missing pairing relationships, resulting in partial data alignment between views and affecting clustering performance. Clustering on such data is referred to as the Partial View-aligned clustering Problem (PVP). Most existing PVP methods mainly capture shared semantic features in a common subspace to infer missing pairing relationships of samples between views. However, excessive reliance on shared feature representations between views could affect the enforcement of the clustering structures. To address this, we propose a novel Cluster-aware cOntrastive Learning method for partially view-Aligned clustering (COLA), by performing intra-view cluster contrastive learning to compact the clustering structure within a view and aligning the soft-label and distributions between views. In this way, the consistent clustering structure between views can be captured and then used to help recover the missing pairing relationships. Extensive experiments demonstrate that COLA outperforms state-of-the-art methods, improving NMI by 14.61% on average and improving ACC by 6.26% on average across eight datasets.
Wanqi Yang, Like Xin, Ming Yang 0014
Neurocomputing4
2026 Global-and-local guidance with synthesized view for unpaired multi-view clustering
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014
Inf. Process. Manag.4
2025 Hybrid Feature Collaborative Reconstruction Network for Few-Shot Fine-Grained Image Classification
abstract
Our research focuses on few-shot fine-grained image classification (FS-FGIC), which faces two main challenges: the similarity of fine-grained objects and a limited number of samples. Traditional feature reconstruction networks enhance key features through spatial reconstruction and error minimization but often fail to capture interclass differences with limited samples. We propose a Hybrid Feature Collaborative Reconstruction Network (HFCR-Net) with two key components: the Hybrid Feature Fusion Process (HFFP) and the Hybrid Feature Reconstruction Process (HFRP). In HFRP, dynamic weight adjustment is employed to enhance spatial dependencies and channel correlations, increasing inter-class differences. In HFRP, we introduce channel dimension reconstruction to improve the processes of support-to-query and query-to-support reconstruction, further increasing interclass and reducing intra-class differences. Extensive experiments on three widely used fine-grained datasets confirm the effectiveness and superiority of our approach.
Shulei Qiu, Wanqi Yang, Ming Yang 0014
ICASSP3
2025 Label-Semantics-Guided Multi-View Multi-Label Learning via High-Order Semantic Fusion
abstract
Incomplete multi-view multi-label learning faces significant challenges arising from semantic heterogeneity across modalities and incomplete modality availability. Traditional fusion approaches typically emphasize superficial feature alignment, neglecting high-order semantic interactions among modalities and labels, thus resulting in redundant or conflicting information integration. To address these limitations, we propose a novel Label Semantic Guided Adaptive Fusion framework. Specifically, we leverage pretrained language models to generate semantic embeddings for both multi-view data and associated labels, facilitating unified semantic understanding. Subsequently, we construct dual-domain hypergraphs separately within the modality and label semantic spaces to explicitly model complex high-order semantic correlations. Based on these hypergraphs, we employ hypergraph neural networks to mine intrinsic semantic relationships and dynamically assess semantic consistency between each modality and the label space. Finally, an adaptive weighting strategy guided by this semantic consistency measure is introduced to fuse modalities effectively, assigning high weights to modalities with greater semantic alignment. Extensive experiments demonstrate that our LSGMM improves fusion accuracy and robustness over state-of-the-art IMvML methods, confirming the effectiveness of integrating label semantics and high-order semantic relationships into adaptive multi-view fusion.
Kaixiang Wang 0001, Xiaojian Ding, Wanqi Yang, Ming Yang 0014
ACM Multimedia4
2025 E2MPL: An Enduring and Efficient Meta Prompt Learning Framework for Few-Shot Unsupervised Domain Adaptation
abstract
Few-shot unsupervised domain adaptation (FS-UDA) leverages a limited amount of labeled data from a source domain to enable accurate classification in an unlabeled target domain. Despite recent advancements, current approaches of FS-UDA continue to confront a major challenge: models often demonstrate instability when adapted to new FS-UDA tasks and necessitate considerable time investment. To address these challenges, we put forward a novel framework called Enduring and Efficient Meta-Prompt Learning (E2MPL) for FS-UDA. Within this framework, we utilize the pre-trained CLIP model as the backbone of feature learning. Firstly, we design domain-shared prompts, consisting of virtual tokens, which primarily capture meta-knowledge from a wide range of meta-tasks to mitigate the domain gaps. Secondly, we develop a task prompt learning network that adaptively learns task-specific prompts with the goal of achieving fast and stable task generalization. Thirdly, we formulate the meta-prompt learning process as a bilevel optimization problem, consisting of (outer) meta-prompt learner and (inner) task-specific classifier and domain adapter. Also, the inner objective of each meta-task has the closed-form solution, which enables efficient prompt learning and adaptation to new tasks in a single step. Extensive experimental studies demonstrate the promising performance of our framework in a domain adaptation benchmark dataset DomainNet. Compared with state-of-the-art methods, our approach has improved the average accuracy by at least 15 percentage points and reduces the average time by 64.67% in the 5-way 1-shot task; in the 5-way 5-shot task, it achieves at least a 9-percentage-point improvement in average accuracy and reduces the average time by 63.18%. Moreover, our method exhibits more enduring and stable performance than the other methods, i.e., reducing the average IQR value by over 40.80% and 25.35% in the 5-way 1-shot and 5-shot task, respectively.
Wanqi Yang, Lei Wang 0001, Ming Yang 0014, Yang Gao 0001
IEEE Trans. Image Process.6
2025 Selective Contrastive Learning for Unpaired Multi-View Clustering
abstract
In this article, we investigate a novel but insufficiently studied issue, unpaired multi-view clustering (UMC), where no paired observed samples exist in multi-view data, and the goal is to leverage the unpaired observed samples in all views for effective joint clustering. Existing methods in incomplete multi-view clustering usually utilize the sample pairing relationship between views to connect the views for joint clustering, but unfortunately, it is invalid for the UMC case. Therefore, we strive to mine a consistent cluster structure between views and propose an effective method, namely selective contrastive learning for UMC (scl-UMC), which needs to solve the following two challenging issues: 1) uncertain clustering structure under no supervision information and 2) uncertain pairing relationship between the clusters of views. Specifically, for the first one, we design an inner-view (IV) selective contrastive learning module to enhance the clustering structures and alleviate the uncertainty, which selects confident samples near the cluster centroids to perform contrastive learning in each view. For the second one, we design a cross-view (CV) selective contrastive learning module to first iteratively match the clusters between views and then tighten the matched clusters. Also, we utilize mutual information to further enhance the correlation of the matched clusters between views. Extensive experiments show the efficiency of our methods for UMC, compared with the state-of-the-art methods.
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.4
2025 Unpaired Multiview Clustering via Reliable View Guidance
abstract
This article focuses on unpaired multiview clustering (UMC), a challenging problem, where paired observed samples are unavailable across multiple views. The goal is to perform effective joint clustering using the unpaired observed samples in all views. In incomplete multiview clustering (IMC), existing methods typically rely on sample pairing between views to capture their complementary. However, this is not applicable in the case of UMC. Hence, we aim to extract the consistent cluster structure across views. In UMC, two challenging issues arise: the uncertain cluster structure due to the lack of labels and the uncertain pairing relationship due to the absence of paired samples. We assume that the view with a good cluster structure is the reliable view, which acts as a supervisor to guide the clustering of the other views. With the guidance of reliable views, a more certain cluster structure of these views is obtained while achieving alignment between the reliable views and the other views. Then, we propose reliable view guided UMC with one reliable view (RG-UMC) and reliable view guided UMC with multiple reliable views (RGs-UMC). Specifically, we design alignment modules with one reliable view and multiple reliable views, respectively, to adaptively guide the optimization process. Also, we utilize the compactness module to enhance the relationship of samples within the same cluster. Meanwhile, an orthogonal constraint is applied to the latent representation to obtain discriminate features. Extensive experiments show that both RG-UMC and RGs-UMC outperform the best state-of-the-art method by an average of 24.14% and 29.42% in normalized mutual information (NMI), respectively.
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.4
2025 Multilevel Reliable Guidance for Unpaired Multiview Clustering
abstract
In this article, we address the challenging problem of unpaired multiview clustering (UMC), which aims to achieve effective joint clustering using unpaired samples observed across multiple views. Traditional incomplete multiview clustering (IMC) methods typically rely on paired samples to capture complementary information between views. However, such strategies become impractical in the UMC due to the absence of paired samples. Although some researchers have attempted to address this issue by preserving consistent cluster structures across views, effectively mining such consistency remains challenging when the cluster structures with low confidence. Therefore, we propose a novel method, multilevel reliable guidance for UMC (MRG-UMC), which integrates multilevel clustering and reliable view guidance to learn consistent and confident cluster structures from three perspectives. Specifically, inner view multilevel clustering exploits high-confidence sample pairs across different levels to reduce the impact of boundary samples, resulting in more confident cluster structures. Synthesized-view alignment leverages a synthesized view to mitigate cross-view discrepancies and promote consistency. Cross-view guidance employs a reliable view guidance strategy to enhance the clustering confidence of poorly clustered views. These three modules are jointly optimized across multiple levels to achieve consistent and confident cluster structures. Furthermore, theoretical analyses verify the effectiveness of MRG-UMC in enhancing clustering confidence. Extensive experimental results show that MRG-UMC outperforms state-of-the-art UMC methods, achieving an average NMI improvement of 12.95% on multiview datasets. The source code is available at https://anonymous.4open.science/r/MRG-UMC-5E20.
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.4
2024 Cross-Modal Attention Alignment Network with Auxiliary Text Description for Zero-Shot Sketch-Based Image Retrieval
Hanwen Su, Jiyan Wang, Ming Yang 0014
ICANN (6)5
2024 Semi-Parametric Style Transfer with Multi-Perspective Feature Fusion and Information-Guided Alignment
abstract
The goal of style transfer is to render images with the attribute dependencies of style images while maintaining the original content structure.Some recent works mainly extract the statistical information of feature maps to match the target style, but the key challenge is that the captured feature representations are too homogeneous, and cross-domain mutual exclusion may occur in the alignment process, resulting in information loss, mismatch, planarization, and so on.To this end, we propose a semi-parametric framework based on multi-perspective feature fusion and guided alignment (SPMPFA), and a backtracking loss function for content maintenance.The SPMPFA and the backgracking loss work together to capture rich presentation information while maintaining structure, thus achieving style consistency between similar fine-grained semantics and global style hierarchy.Specifically, we first use adaptive aggregation and mapping Transformer (AMTransformer) to build a cross-domain graph carrying location information inside the module and use information based on weight aggregation as associated features to guide the style alignment trend.Then, we use the feature fusion strategy to adaptively fuse the heterogeneous representation information.Finally, the content structure is maintained to the maximum extent by using backtracking loss.Qualitative and quantitative experiments demonstrate the effectiveness of our work compared to other style transfer tasks.
Tianlong Zhang, Jing Lv, Ming Yang 0014
ICMR3
2024 Dual-view Pyramid Network for Video Frame Interpolation
abstract
Video frame interpolation is a critical component of video streaming, a vibrant research area dealing with requests of both service providers and users. However, existing methods cannot handle changing video resolutions while improving user perceptual quality. We aim to unleash the multifaceted knowledge yielded by the hierarchical views at multiple scales in a pyramid network. Specifically, we build a dual-view pyramid network by introducing pyramidal dual-view correspondence matching. It compels each scale to actively seek knowledge in view of both the current scale and a coarser scale, conducting robust correspondence matching by considering neighboring scales. Meanwhile, an auxiliary multi-scale collaborative supervision is devised to enforce the exchange of knowledge among scales and thus reduce error propagation from coarse to fine scales. Based on the robust capture of video dynamics via pyramidal dual-view correspondence matching, we further construct a pyramidal refinement module that formulates frame refinement as progressive latent representation generations by developing flow-guided cross-scale attention for feature fusion among frames. The proposed method is able to improve the perceptual quality on several benchmarks of varying video resolutions, while keeping low distortion and a compact model size.
Yao Luo, Ming Yang 0014, Jinhui Tang 0001
ACM Multimedia2
2024 Deep self-enhancement hashing for robust multi-label cross-modal retrieval
Hanwen Su, Fengyi Song, Ming Yang 0014
Pattern Recognit.5
2024 Deep Ranking Distribution Preserving Hashing for Robust Multi-Label Cross-Modal Retrieval
abstract
Deep supervised hashing techniques have exhibited remarkable efficiency in cross-modal retrieval tasks, because they enable the transformation of data from different modalities into compact binary codes that preserve semantic similarity structures. Nonetheless, existing methods often rely on pairwise or triplet relationships within known (or in-distribution) semantics during training, failing to capture the comprehensive ranking information inherent in web data that encompasses diverse concepts. In addition, these methods are vulnerable to out-of-distribution (OOD) semantic data when applied in realistic scenarios, resulting in suboptimal performance. In this paper, we propose ranking distribution preserving hashing (RDPH) to address these problems. We present a novel ranking loss, a differentiable surrogate that maximizes the NDCG metric for cross-modal retrieval. This loss incorporates two target ranking distributions derived from the ideal NDCG scores of samples and the cosine similarity of features. These distributions encourage RDPH to generate hash codes that approximate the desired inter-modal and intra-modal ranking distributions. To enhance the robustness of the hash codes against OOD data, RDPH leverages the CLIP paradigm to acquire OOD-resilient intermediate representations. Besides, we utilize the outlier exposure strategy to enhance the discriminative ability of OOD for hash codes under supervision by constructing auxiliary pseudo-OOD data from known data in feature space. Experiments on three datasets demonstrate that the proposed method achieves state-ofthe-art performance on regular retrieval tasks and good results on simulated real-world retrieval tasks.
Hanwen Su, Fengyi Song, Ming Yang 0014
IEEE Trans. Multim.5
2024 Iterative Multiview Subspace Learning for Unpaired Multiview Clustering
abstract
In real applications, several unpredictable or uncertain factors could result in unpaired multiview data, i.e., the observed samples between views cannot be matched. Since joint clustering among views is more effective than individual clustering in each view, we investigate unpaired multiview clustering (UMC), which is a valuable but insufficiently studied problem. Due to lack of matched samples between views, we could fail to build the connection between views. Therefore, we aim to learn the latent subspace shared by views. However, existing multiview subspace learning methods usually rely on the matched samples between views. To address this issue, we propose an iterative multiview subspace learning strategy [iterative unpaired multiview clustering (IUMC)], aiming to learn a complete and consistent subspace representation among views for UMC. Moreover, based on IUMC, we design two effective UMC methods: 1) Iterative unpaired multiview clustering via covariance matrix alignment (IUMC-CA) that further aligns the covariance matrix of subspace representations and then performs clustering on the subspace and 2) iterative unpaired multiview clustering via one-stage clustering assignments (IUMC-CY) that performs one-stage multiview clustering (MVC) by replacing the subspace representations with clustering assignments. Extensive experiments show the excellent performance of our methods for UMC, compared with the state-of-the-art methods. Also, the clustering performance of observed samples in each view can be considerably improved by those observed samples from the other views. In addition, our methods have good applicability in incomplete MVC.
Wanqi Yang, Like Xin, Lei Wang 0001, Ming Yang 0014, Wenzhu Yan, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 High-Level Semantic Feature Matters Few-Shot Unsupervised Domain Adaptation
abstract
In few-shot unsupervised domain adaptation (FS-UDA), most existing methods followed the few-shot learning (FSL) methods to leverage the low-level local features (learned from conventional convolutional models, e.g., ResNet) for classification. However, the goal of FS-UDA and FSL are relevant yet distinct, since FS-UDA aims to classify the samples in target domain rather than source domain. We found that the local features are insufficient to FS-UDA, which could introduce noise or bias against classification, and not be used to effectively align the domains. To address the above issues, we aim to refine the local features to be more discriminative and relevant to classification. Thus, we propose a novel task-specific semantic feature learning method (TSECS) for FS-UDA. TSECS learns high-level semantic features for image-to-class similarity measurement. Based on the high-level features, we design a cross-domain self-training strategy to leverage the few labeled samples in source domain to build the classifier in target domain. In addition, we minimize the KL divergence of the high-level feature distributions between source and target domains to shorten the distance of the samples between the two domains. Extensive experiments on DomainNet show that the proposed method significantly outperforms SOTA methods in FS-UDA by a large margin (i.e., ~10%).
Wanqi Yang, Shengqi Huang, Lei Wang 0001, Ming Yang 0014
AAAI5
2023 Cross Attention with Deep Local Features for Few-Shot Image Classification
Tengfei Chu, Jing Lv, Ming Yang 0014
ICANN (10)4
2023 Aggregate Distillation for Top-K Recommender System
Tengfei Chu, Yulou Yang, Ming Yang 0014
ICANN (7)4
2023 SDRNet: Shape Decoupled Regression Network for 3d face Reconstruction
abstract
In the field of computer vision, 3D face reconstruction from single-view images is a long-standing and challenging problem. Following the popular 3DMM-based reconstruction framework, recent works show great concerns about exploring discriminative information of identity and expression for shape regression commonly in coupling ways. Actually, identity and expression information may contribute differently in explaining the intrinsic shape of faces, and the former is inferior to the latter in explaining the great facial shape variations caused by extreme expression. In this paper, we propose a Shape Decoupled Regression Network (SDRNet) consisting of identity-focused branch and expression-focused branch with focused criteria for representation learning, which interact with the union branch to achieve the final 3DMM parameters regression for improved shape reconstruction. In SDRNet, the focused criteria estimate the 3D vertex prediction loss, while the predicted 3D shape is reconstructed only using the predicted parameter of identity or expression and introducing the ground-truth parameters of the left two. Extensive experiments on the challenging AFLW2000-3D and AFLW datasets demonstrate advanced performance in 3D face reconstruction and face alignment.
Shikun Zhang, Fengyi Song, Ming Yang 0014
ICASSP4
2023 Deep continual hashing for real-world multi-label image retrieval
Hanwen Su, Fengyi Song, Ming Yang 0014
Comput. Vis. Image Underst.5
2023 Deep continual hashing with gradient-aware memory for cross-modal retrieval
Xiaoyang Tan, Ming Yang 0014
Pattern Recognit.3
2023 Towards deeper match for multi-view oriented multiple kernel learning
Wenzhu Yan, Yanmeng Li, Ming Yang 0014
Pattern Recognit.3
2023 Robust Low Rank and Sparse Representation for Multiple Kernel Dimensionality Reduction
abstract
In the fields of pattern recognition and data mining, two problems need to be addressed. First, the curse of dimensionality degrades the performance of many practical data processing techniques. Second, due to the existence of noise and outliers, feature extraction on corrupted data cannot be effectively achieved. Recently, some representation based methods have produced promising results. However, these methods cannot handle the case in which nonlinear similarity exists and have failed to provide the quantized interpretability for the importance of features. In this paper, we propose a novel low rank and sparse representation method to realize dimensionality reduction and robustly extract latent low dimensional discriminative features. Specifically, we first adopt multiple kernel learning to map the original data into an embedded reproducing kernel Hilbert space (RKHS) and then kernel based similarity discriminative projection is learned to explore the within-class and between-class variability. Notably, this low dimensional feature learning strategy is definitely integrated into the low rank matrix recovery of the kernel matrix. Next, we introduce the regularization of$l_{2,1}$norm on error matrix to eliminate noise and on projection matrix to lead the selected features to be more compact and interpretable. The non-convex optimization problem is effectively solved by the alternating direction method of multipliers (ADMM) methods. Extensive experiments on seven benchmark datasets are conducted to demonstrate the effectiveness of our method.
Wenzhu Yan, Ming Yang 0014, Yanmeng Li
IEEE Trans. Circuits Syst. Video Technol.2
2022 A Lightweight Network with Multi-Stage Feature Fusion Module for Single-View 3d Face Reconstruction
abstract
3D face reconstruction has attracted great attentions of researchers from both academic and industry for its potential application in many scenarios such as face alignment and recognition across large poses. 3D Morphable Model which reconstructs a 3D face through basis coefficients prediction, is usually adopted as the typical parametric framework for 3D face and is suitable to combine with deep learning. Existing cascade regression method predicts coefficients by multiple iterations, which is time-consuming. In this paper, we propose an efficient and end-to-end method for single-view 3D face reconstruction. We build a lightweight network based on mobile blocks with faster speed for parameter extraction and smaller model size. Especially, a multi-stage feature fusion module is designed for enhancing the end-to-end learning. To match the setting of input image size, we updated the pose label of images under various sizes in training dataset before training. Extensive experiments on challenging datasets validate the efficiency of our method for both 3D face reconstruction and face alignment.
Shikun Zhang, Fengyi Song, Ming Yang 0014
ICIP5
2022 Exploring Occlusion-Sensitive Deep Network for Single-View 3D Face Reconstruction
abstract
Recovering 3D geometry from a single-view 2D face image is an ill-posed task full of various challenges, especially under occlusion conditions commonly seen with large poses, while partial facial information missing makes the burden much heavier. Although existing methods could solve this problem in the end-to-end fashion, they still have limited performance in the occluded scenes. It is intuitively for many methods to depress the influence of those occluded regions for reconstruction. But we propose an occlusion-sensitive weighting mechanism for balancing the contributions among occluded and non-occluded regions. Meanwhile, considering no dataset contains various occlusions for learning, the data augmentation technique is exploited to expand the training dataset, which further facilitates the learning of the occlusion-sensitive deep network. Extensive experiments on two challenging datasets validate the advanced performance of our method for both 3D face reconstruction and face alignment.1
Shikun Zhang, Fengyi Song, Ming Yang 0014
ICIP5
2022 Relation-Guided Network for Image-Text Retrieval
abstract
Image-text retrieval has made great progress, but it remains challenging due to heterogeneity between images and text. Enhancing the interaction by exploring the relationship between the image and text can reduce this problem, to some extent. How to explore and use the relationship between image and text to enhance the interaction between them is a critical problem. In this paper, we design an asymmetric structure network (RGN) to represent image and text. First, we mine the relationship between image and text, and extract the specific text information. Then we exploit this relationship to guide the generation of text embeddings, which can capture the rich and representative embeddings. Results on two datasets, Flickr30K dataset and MSCOCO dataset, show that our model can achieve competitive results.
Yulou Yang, Ming Yang 0014
ICIP3
2022 Few-Shot Unsupervised Domain Adaptation via Meta Learning
abstract
Unsupervised domain adaptation (UDA) has raised a lot of interests in recent years. However, current UDA methods are still not capable enough in dealing with two issues: 1) the scarcity of labeled data in source domain and 2) the need of a general model that can quickly adapt to solve new UDA tasks. To address this situation, we investigate available but rarely-studied setting called few-shot unsupervised domain adaptation (FS-UDA), in which the data of source domain is few-shot per category and the data of target domain remains unlabeled. To realize effective adaptation for FS-UDA tasks in the same source and target domains, we propose a novel meta learning method namely meta-FUDA, which leverages meta learning to perform task-level transfer and domain-level transfer jointly. Extensive experiments demonstrate the promising performance of our method on multiple benchmark data sets.
Wanqi Yang, Chengmei Yang, Shengqi Huang, Lei Wang 0001, Ming Yang 0014
ICME5
2022 Dual-scale correlation analysis for robust multi-label classification
Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Lei Wang 0001
Appl. Intell.2
2021 Few-shot Unsupervised Domain Adaptation with Image-to-Class Sparse Similarity Encoding
abstract
This paper investigates a valuable setting called few-shot unsupervised domain adaptation (FS-UDA), which has not been sufficiently studied in the literature. In this setting, the source domain data are labelled, but with few-shot per category, while the target domain data are unlabelled. To address the FS-UDA setting, we develop a general UDA model to solve the following two key issues: the few-shot labeled data per category and the domain adaptation between support and query sets. Our model is general in that once trained it will be able to be applied to various FS-UDA tasks from the same source and target domains. Inspired by the recent local descriptor based few-shot learning (FSL), our general UDA model is fully built upon local descriptors (LDs) for image classification and domain adaptation. By proposing a novel concept called similarity patterns (SPs), our model not only effectively considers the spatial relationship of LDs that was ignored in previous FSL methods, but also makes the learned image similarity better serve the required domain alignment. Specifically, we propose a novel IMage-to-class sparse Similarity Encoding (IMSE) method. It learns SPs to extract the local discriminative information for classification and meanwhile aligns the covariance matrix of the SPs for domain adaptation. Also, domain adversarial training and multi-scale local feature matching are performed upon LDs. Extensive experiments conducted on a multi-domain benchmark dataset DomainNet demonstrates the state-of-the-art performance of our IMSE for the novel setting of FS-UDA. In addition, for FSL, our IMSE can also show better performance than most of recent FSL methods on miniImageNet.
Shengqi Huang, Wanqi Yang, Lei Wang 0001, Luping Zhou, Ming Yang 0014
ACM Multimedia5
2021 Deep robust multilevel semantic hashing for multi-label cross-modal retrieval
Xiaoyang Tan, Jun Zhao 0007, Ming Yang 0014
Pattern Recognit.4
2021 SA-LuT-Nets: Learning Sample-Adaptive Intensity Lookup Tables for Brain Tumor Segmentation
abstract
In clinics, the information about the appearance and location of brain tumors is essential to assist doctors in diagnosis and treatment. Automatic brain tumor segmentation on the images acquired by magnetic resonance imaging (MRI) is a common way to attain this information. However, MR images are not quantitative and can exhibit significant variation in signal depending on a range of factors, which increases the difficulty of training an automatic segmentation network and applying it to new MR images. To deal with this issue, this paper proposes to learn a sample-adaptive intensity lookup table (LuT) that dynamically transforms the intensity contrast of each input MR image to adapt to the following segmentation task. Specifically, the proposed deep SA-LuT-Net framework consists of a LuT module and a segmentation module, trained in an end-to-end manner: the LuT module learns a sample-specific nonlinear intensity mapping function through communication with the segmentation module, aiming at improving the final segmentation performance. In order to make the LuT learning sample-adaptive, we parameterize the intensity mapping function by exploring two families of non-linear functions (i.e., piece-wise linear and power functions) and predict the function parameters for each given sample. These sample-specific parameters make the intensity mapping adaptive to samples. We develop our SA-LuT-Nets separately based on two backbone networks for segmentation, i.e., DMFNet and the modified 3D Unet, and validate them on BRATS2018 and BRATS2019 datasets for brain tumor segmentation. Our experimental results clearly demonstrate the superior performance of the proposed SA-LuT-Nets using either single or multiple MR modalities. It not only significantly improves the two baselines (DMFNet and the modified 3D Unet), but also wins a set of state-of-the-art segmentation methods. Moreover, we show that, the LuTs learnt using one segmentation model could also be applied to improving the performance of another segmentation model, indicating the general segmentation information captured by LuTs.
Biting Yu, Luping Zhou, Lei Wang 0001, Wanqi Yang, Ming Yang 0014, Pierrick Bourgeat, Jurgen Fripp
IEEE Trans. Medical Imaging5
2020 Learning Sample-Adaptive Intensity Lookup Table for Brain Tumor Segmentation
Biting Yu, Luping Zhou, Lei Wang 0001, Wanqi Yang, Ming Yang 0014, Pierrick Bourgeat, Jurgen Fripp
MICCAI (4)5
2020 A novel spectral-spatial based adaptive minimum spanning forest for hyperspectral image classification
Jing Lv, Ming Yang 0014, Wanqi Yang
GeoInformatica3
2020 An Effective MR-Guided CT Network Training for Segmenting Prostate in CT Images
abstract
Segmentation of prostate in medical imaging data (e.g., CT, MRI, TRUS) is often considered as a critical yet challenging task for radiotherapy treatment. It is relatively easier to segment prostate from MR images than from CT images, due to better soft tissue contrast of the MR images. For segmenting prostate from CT images, most previous methods mainly used CT alone, and thus their performances are often limited by low tissue contrast in the CT images. In this article, we explore the possibility of using indirect guidance from MR images for improving prostate segmentation in the CT images. In particular, we propose a novel deep transfer learning approach, i.e., MR-guided CT network training (namely MICS-NET), which can employ MR images to help better learning of features in CT images for prostate segmentation. In MICS-NET, the guidance from MRI consists of two steps: (1) learning informative and transferable features from MRI and then transferring them to CT images in a cascade manner, and (2) adaptively transferring the prostate likelihood of MRI model (i.e., well-trained convnet by purely using MR images) with a view consistency constraint. To illustrate the effectiveness of our approach, we evaluate MICS-NET on a real CT prostate image set, with the manual delineations available as the ground truth for evaluation. Our methods generate promising segmentation results which achieve (1) six percentages higher Dice Ratio than the CT model purely using CT images and (2) comparable performance with the MRI model purely using MR images.
Wanqi Yang, Yinghuan Shi, Sanghyun Park 0004, Ming Yang 0014, Yang Gao 0001, Dinggang Shen
IEEE J. Biomed. Health Informatics4
2019 Multi-view Locality Preserving Embedding with View Consistent Constraint for Dimension Reduction
Weiling Cai, Ming Yang 0014, Fengyi Song
KSEM (1)3
2019 Multi-Constraints-Based Enhanced Class-Specific Dictionary Learning for Image Classification
Ze Tian, Ming Yang 0014
PAKDD (3)2
2019 Sparse and low-rank representation for multi-label classification
Zhifen He, Ming Yang 0014
Appl. Intell.2
2019 Joint multi-label classification and label correlations with missing labels and feature selection
Zhifen He, Ming Yang 0014, Yang Gao 0001, Hui-Dong Liu, Yilong Yin
Knowl. Based Syst.2
2019 Calibrated Multi-label Classification with Label Correlations
Zhifen He, Ming Yang 0014, Hui-Dong Liu, Lei Wang 0001
Neural Process. Lett.2
2018 Attributes Consistent Faces Generation Under Arbitrary Poses
Fengyi Song, Jinhui Tang 0001, Ming Yang 0014, Weiling Cai, Wanqi Yang
ACCV (2)3
2018 Deep Correlation Structure Preserved Label Space Embedding for Multi-label Classification
abstract
Label embedding is an effective and efficient method which can jointly extract the information of all labels for better performance of multi-label classification. However, most existing embedding methods ignore information of feature space or intrinsic structure of previous label space, such that their learned latent space will not have strong predictability and discriminant ability. We propose a novel deep neural network (DNN) based model, namely Deep Correlation Structure Preserved Label Space Embedding (DCSPE). Specifically, DCSPE derives a deep latent space by performing feature-aware label space embedding with deep canonical correlation analysis (DCCA) and preserving the intrinsic structure of the previous label space with proposed deep multidimensional scaling (DMDS). Our DCSPE is achieved by integrating the DNN architectures of the two DNN based models and can learn a feature-aware structure preserved deep latent space. Furthermore, extensive experimental results on datasets with many labels demonstrate that our proposed approach is significantly better than the existing label embedding algorithms.
Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Yilong Yin
ACML2
2018 Deep Cross-View Label Embedding with Correlation and Structure Preserved for Multi-Label Classification
abstract
Label embedding is an important family of multi-label classification algorithms which can jointly extract the information of all labels for better performance. However, few works have been done on label embedding methods which consider the structure information of original feature and label space simultaneously. We propose a novel deep neural network (DNN) based model for learning an effective deep latent space, namely Deep Cross-view label space Embedding with Correlation and Structure preserved (DCECS). In DCECS, the latent space correlates with feature and label spaces closely by virtue of the deep cross-view embedding. Meanwhile, the latent space is also learned under the guidance of label correlation and local structure of feature space which are exploited by hypergraph and graph regularizations. The overall framework achieves the complementarity and correspondence between information of feature and label space, therefore the feature-aware deep latent space we learned has strong predictability and discriminant ability. Extensive experimental results on datasets with many labels demonstrate that our proposed approach is significantly better than the existing label embedding algorithms.
Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Yilong Yin
ICTAI2
2018 Feature Integration with Adaptive Importance Maps for Visual Tracking
abstract
Discriminative correlation filters have recently achieved excellent performance for visual object tracking. The key to success is to make full use of dense sampling and specific properties of circulant matrices in the Fourier domain. However, previous studies don't take into consideration the importance and complementary information of different features, simply concatenating them. This paper investigates an effective method of feature integration for correlation filters, which jointly learns filters, as well as importance maps in each frame. These importance maps borrow the advantages of different features, aiming to achieve complementary traits and improve robustness. Moreover, for each feature, an importance map is shared by its all channels to avoid overfitting. In addition, we introduce a regularization term for the importance maps and use the penalty factor to control the significance of features. Based on handcrafted and CNN features, we implement two trackers, which achieve a competitive performance compared with several state-of-the-art trackers.
Aishi Li, Ming Yang 0014, Wanqi Yang
IJCAI2
2018 Online multi-view subspace learning via group structure analysis for visual object tracking
Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Ming Yang 0014
Distributed Parallel Databases4
2018 Image filtering method using trimmed statistics and edge preserving
abstract
Image filtering is to retain the details of the image as much as possible and meanwhile suppress the noise pollution to great extent. This study presents an image filtering using the truncated statistics and edge preserving. In the first step of our method, the alpha‐trimmed filter is utilized to remove a variety of types of noises; in the second step, taking the image after alpha‐trimmed filtering as a guide image, the local linear model between the guide image and the target image is established; in the third step, the obtained local linear model is further simplified to reduce the time complexity; and finally, using the relationship between image local variance and the global variance, the local linear model is modified to enhance the details of the image and meanwhile remove halo phenomenon. This method has three advantages: (i) it is flexible to deal with the images stained by various types of high‐intensity noise; (ii) it is effective to keep the image details and profile information, and remove the halo phenomenon; and (iii) it runs in time linear in the image size, thus its computation complexity is low. Experimental results show that the proposed filter is robust and efficient.
Weiling Cai, Ming Yang 0014, Fengyi Song
IET Image Process.2
2018 Incomplete-Data Oriented Multiview Dimension Reduction via Sparse Low-Rank Representation
abstract
For dimension reduction on multiview data, most of the previous studies implicitly take an assumption that all samples are completed in all views. Nevertheless, this assumption could often be violated in real applications due to the presence of noise, limited access to data, equipment malfunction, and so on. Most of the previous methods will cease to work when missing values in one or multiple views occur, thus an incomplete-data oriented dimension reduction becomes an important issue. To this end, we mathematically formulate the above-mentioned issue as sparse low-rank representation through multiview subspace (SRRS) learning to impute missing values, by jointly measuring intraview relations (via sparse low-rank representation) and interview relations (through common subspace representation). Moreover, by exploiting various subspace priors in the proposed SRRS formulation, we develop three novel dimension reduction methods for incomplete multiview data: 1) multiview subspace learning via graph embedding; 2) multiview subspace learning via structured sparsity; and 3) sparse multiview feature selection via rank minimization. For each of them, the objective function and the algorithm to solve the resulting optimization problem are elaborated, respectively. We perform extensive experiments to investigate their performance on three types of tasks including data recovery, clustering, and classification. Both two toy examples (i.e., Swiss roll and -curve) and four real-world data sets (i.e., face images, multisource news, multicamera activity, and multimodality neuroimaging data) are systematically tested. As demonstrated, our methods achieve the performance superior to that of the state-of-the-art comparable methods. Also, the results clearly show the advantage of integrating the sparsity and low-rankness over using each of them separately.
Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Ming Yang 0014
IEEE Trans. Neural Networks Learn. Syst.5
2017 Infomax principle based pooling of deep convolutional activations for image retrieval
abstract
Neural activations produced by deep convolutional networks have recently become state-of-the-art representation for image retrieval. To obtain a global image representation, sum-pooling has been frequently used to aggregate activations of convolutional feature maps. This work first presents an understanding on the effectiveness of sum-pooling via probabilistic interpretation, by proving that sum-pooling is an upper bound of the probability that a visual pattern is present in an image. To further answer the optimality of sum-pooling, a quantitative analysis based on the Infomax principle in neural networks is provided. It shows that sum-pooling aligns well with the leading eigenvector of principal component analysis (PCA) applied to the activations of a feature map. Moreover, considering the 2D matrix structure of feature maps, a two-directional 2DPCA-based pooling scheme is proposed to aggregate the convolutional activations. Experiments on multiple benchmark image retrieval datasets demonstrate the above analysis and the superiority of the proposed pooling scheme.
Zhimin Gao, Lei Wang 0001, Luping Zhou, Ming Yang 0014
ICME4
2017 Cost Sensitive Matrix Factorization for Face Recognition
Jianwu Wan, Ming Yang 0014
IDEAL2
2017 Cost Sensitive Semi-Supervised Canonical Correlation Analysis for Multi-view Dimensionality Reduction
Jianwu Wan, Ming Yang 0014
Neural Process. Lett.3
2016 An Improved Recommender Model by Joint Learning of Both Similarity and Latent Feature Space
Yunxiang Tao, Ming Yang 0014
IDEAL2
2015 Discriminative cost sensitive Laplacian score for face recognition
Jianwu Wan, Ming Yang 0014, Yin-Juan Chen
Neurocomputing2
2014 Linear Regression Fisher Discrimination Dictionary Learning for Hyperspectral Image Classification
Ming Yang 0014, Hujun Yin
IDEAL2
2014 mPadal: a joint local-and-global multi-view feature selection method for activity recognition
Wanqi Yang, Yang Gao 0001, Longbing Cao, Ming Yang 0014, Yinghuan Shi
Appl. Intell.4
2014 Local histogram specification for face recognition under varying lighting conditions
Hui-Dong Liu, Ming Yang 0014, Yang Gao 0001, Chunyan Cui
Image Vis. Comput.2
2014 Bilinear discriminative dictionary learning for face recognition
Hui-Dong Liu, Ming Yang 0014, Yang Gao 0001, Yilong Yin
Pattern Recognit.2
2014 Fast Local Histogram Specification
abstract
Local histogram specification (LHS) is a useful technique for image processing. However, LHS faces a critical computational challenge when it is applied to high-resolution high-precision images. The calculation of the values in the cumulative distribution function (CDF) and the mapped value for the central pixel in each sliding window is time consuming with the computational complexity O(s + L) of the state-of-theart techniques, where s is the side length of the square window and L is the number of gray levels. In this paper, we propose a fast algorithm for LHS, called fast local histogram specification (FLHS). FLHS reduces the complexity of calculating the CDF value for the central pixel in each sliding window to O(s + √L), and the time complexity for the mapping procedure in each window to O(log L). This results in the overall time complexity of LHS reduced from O(s+L) to O(s+√L) in each sliding window. Theoretical analysis shows that the newly developed algorithm is efficient. Experimental results on the 8-bit and high-resolution high-precision (16-bit) images demonstrate the efficiency of our proposed algorithm.
Hui-Dong Liu, Ming Yang 0014, Yang Gao 0001, Longbing Cao
IEEE Trans. Circuits Syst. Video Technol.2
2014 Pairwise Costs in Semisupervised Discriminant Analysis for Face Recognition
abstract
In recent years, face recognition is being recognized as a cost-sensitive learning problem. Many cost-sensitive classifiers have been proposed. However, no sufficient attention is paid to the research on cost-sensitive dimensionality reduction, especially on the cost-sensitive semisupervised dimensionality reduction. To the best of our knowledge, cost sensitive semisupervised discriminant analysis (CS3DA) may be the first work. CS3DA first uses the sparse representation to infer a soft label for unlabeled sample and then learns the projection direction by incorporating misclassification costs into both labeled and unlabeled data. Although CS3DA reduces the loss of misclassification, it has two major drawbacks: 1) the sparsity is not a feature of face recognition, and therefore sparse approximations may not deliver the robustness or performance desired and 2) CS3DA is not proven to satisfy the minimal misclassification loss criterion. In this paper, we embed pairwise costs in semisupervised discriminant analysis (PCSDA) for face recognition. PCSDA first uses a simple l2approach to predict the label of unlabeled data, and then learns the projection direction by embedding pairwise costs in both labeled and unlabeled data. Compared with CS3DA, PCSDA has three major advantages: 1) l2approach is more accurate and robust than sparse representation for face recognition; 2) we prove that CS3DA approximates the pairwise Bayesian risk only when the classes are balanced and without outliers in face data sets; and 3) PCSDA approximates the pairwise Bayesian risk considering the class imbalance problem and outliers in face recognition. Hence, the projection direction obtained by using PCSDA can be more discriminative, immunes to outliers and class imbalance problem. The experimental results on AR, PIE, ORL, and extended Yale B data sets demonstrate the effectiveness of PCSDA.
Jianwu Wan, Ming Yang 0014, Yang Gao 0001, Yin-Juan Chen
IEEE Trans. Inf. Forensics Secur.2
2012 Local histogram specification using learned histograms for face recognition
abstract
In the field of face recognition, most existing preprocessing methods only try to filter the low frequency part of the spectrum of face images to eliminate illumination variations. In this paper, we introduce the Local Histogram Specification (LHS) to preprocess face images using learned histograms. Each local histogram to be specified is learned by estimating the distribution of gray values in the corresponding local region of all normal lighting images in the training set. The proposed method is able to alleviate both the low and high frequency parts of illumination on face images as well as enhance face features lying in the low frequency part. Reasonable window size is also empirically studied. Experimental results on two standard illumination variation datasets demonstrate the effectiveness and stability of our proposed method.
Hui-Dong Liu, Ming Yang 0014
ICIP2
2010 A novel hypothesis-margin based approach for feature selection with side pairwise constraints
Ming Yang 0014
Neurocomputing1
2008 A novel condensing tree structure for rough set feature selection
Ming Yang 0014
Neurocomputing1