Zengmao Wang

dblp:168/4719 · DBLP profile ↗
← Back
58ranked-venue papers
17as first author
43since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Calibrating graph neural networks for trustworthy graph-level predictions with confidence backward-readout
Junchao Qiu, Guojia Wan, Zengmao Wang, Bo Du 0001
Expert Syst. Appl.4
2026 SDLK-Net: Enhanced squeezed directional large kernel multi-scale multi-modal fusion network for salient object detection
Lingyu Yan, Rong Gao 0001, Zengmao Wang, Zhiwei Ye, Xinyun Wu
Appl. Intell.4
2026 CAFL: Conditional Attention Federated Learning for Image Emotion Analysis
abstract
Abstract The rapid proliferation of images on online platforms has made emotion analysis a task of paramount significance. However, these images are often privacy-sensitive, making Federated Learning (FL) a compelling paradigm over traditional centralized methods. A critical yet largely unaddressed challenge in applying FL to this domain is the severe concept drift stemming from the subjective and culturally diverse nature of emotional expression, which causes conventional FL algorithms to fail. In this paper, we propose CAFL (Conditional Attention Federated Learning) to fill this gap. CAFL empowers clients to learn collaboratively yet personally. It intelligently routes information through an adaptive gate that separates features into a personalized stream and a global stream. These streams are then processed by dedicated local and global prediction heads. Crucially, collaboration is guided by a conditional attention mechanism, where the server computes a personalized reference model for each client based on an attention-weighted aggregation of peer models, promoting knowledge sharing among kindred clients. Extensive experiments on various lightweight foundation models show that CAFL consistently outperforms existing FL methods, demonstrating its robustness and superior performance as a solution for distributed, privacy-sensitive image emotion analysis.
Chang Liu 0046, Zengmao Wang, Yongchao Xu, Bo Du 0001
Data Sci. Eng.2
2026 Wavelet Spectral-Spatial Mamba Network for Hyperspectral Image Classification
abstract
Spectral–spatial feature modeling plays a crucial role in hyperspectral image (HSI) classification. However, existing models based on convolutional neural networks (CNNs) and Transformers still face a trade-off between feature modeling capability and computational efficiency. Although recent wavelet-based HSI classification methods have demonstrated the advantages of frequency-domain analysis, they typically rely on a single wavelet basis, which limits their ability to capture diverse spectral–spatial patterns across different frequency bands. To address these issues, we propose Wavelet Spectral-Spatial Mamba (WSSMamba) by combining wavelet transform with state space modeling for HSI classification. WSSMamba introduces an Adaptive Wavelet Fusion Module (AWFM) to perform multi-scale frequency domain decomposition using multiple wavelet bases. This allows the model to extract both low-frequency global structure and high-frequency local details. A Wavelet Feature Enhancement (WFE) module is also designed to improve feature discriminability by applying channel and spatial attention mechanisms. Furthermore, we propose a Spectral-Spatial Cross-Fusion Strategy (SSCFS), which uses multi-directional state modeling to dynamically integrate high-frequency information. Extensive experiments on benchmark datasets demonstrate that WSSMamba outperforms state-of-the-art methods in classification performance.
Yongchao Song, Zhaowei Liu 0001, Weiqing Yan, Zengmao Wang, Xuan Wang 0021
IEEE Trans. Circuits Syst. Video Technol.5
2025 Dynamic Parallel Tree Search for Efficient LLM Reasoning
abstract
Tree of Thoughts (ToT) enhances Large Language Model (LLM) reasoning by structuring problem-solving as a spanning tree. However, recent methods focus on search accuracy while overlooking computational efficiency. The challenges of accelerating the ToT lie in the frequent switching of reasoning focus, and the redundant exploration of suboptimal solutions. To alleviate this dilemma, we propose Dynamic Parallel Tree Search (DPTS), a novel parallelism framework that aims to dynamically optimize the reasoning path in inference. It includes the Parallelism Streamline in the generation phase to build up a flexible and adaptive parallelism with arbitrary paths by cache management and alignment. Meanwhile, the Search and Transition Mechanism filters potential candidates to dynamically maintain the reasoning focus on more possible solutions with less redundancy. Experiments on Qwen-2.5 and Llama-3 on math and code datasets show that DPTS significantly improves efficiency by 2-4\times on average while maintaining or even surpassing existing reasoning algorithms in accuracy, making ToT-based reasoning more scalable and computationally efficient. Codes are released at: https://github.com/yifu-ding/DPTS.
Yifu Ding 0001, Shunyu Liu 0001, Yongcheng Jing, Zengmao Wang, Ziwei Liu 0002, Bo Du 0001, Xianglong Liu 0001, Dacheng Tao
ACL (1)8
2025 Enhancing Multimodal Chain-of-Thought Reasoning with Tree-Searched Self-Training
abstract
Despite impressive performance on general visual benchmarks, Multimodal Large Language Models (MLLMs) still struggle with generating consistent and accurate reasoning processes for complex visual reasoning tasks. The limited availability of multimodal reasoning datasets further complicates fine-tuning efforts to improve reasoning capabilities. To address these challenges, we introduce Tree-Searched Self-Training (TSST), a novel framework that enhances multimodal reasoning through self-improvement without relying on extensive manual annotations. TSST introduces a hierarchical tree-search mechanism that combines stepwise rationales generation with value-guided path selection, enabling the model to explore and identify high-quality reasoning trajectories autonomously. Training on these self-generated data significantly improves the model’s reasoning ability, boosting consistency and accuracy. Extensive experiments on ScienceQA and VCR benchmarks show that TSST outperforms existing fine-tuning baselines by 6.0% in accuracy, demonstrating its effectiveness in enhancing multimodal reasoning capabilities. The code is publicly available at: https://github.com/laaambs/tsst.
Yiwen Luo, Yong Luo 0002, Zengmao Wang
ICME4
2025 When Data-Free Knowledge Distillation Meets Non-Transferable Teacher: Escaping Out-of-Distribution Trap is All You Need
abstract
Data-free knowledge distillation (DFKD) transfers knowledge from a teacher to a student without access the real in-distribution (ID) data. Its common solution is to use a generator to synthesize fake data and use them as a substitute for real ID data. However, existing works typically assume teachers are trustworthy, leaving the robustness and security of DFKD from untrusted teachers largely unexplored. In this work, we conduct the first investigation into distilling non-transferable learning (NTL) teachers using DFKD, where the transferability from an ID domain to an out-of-distribution (OOD) domain is prohibited. We find that NTL teachers fool DFKD through divert the generator’s attention from the useful ID knowledge to the misleading OOD knowledge. This hinders ID knowledge transfer but prioritizes OOD knowledge transfer. To mitigate this issue, we propose Adversarial Trap Escaping (ATEsc) to benefit DFKD by identifying and filtering out OOD-like synthetic samples. Specifically, inspired by the evidence that NTL teachers show stronger adversarial robustness on OOD samples than ID samples, we split synthetic samples into two groups according to their robustness. The fragile group is treated as ID-like data and used for normal knowledge distillation, while the robust group is seen as OOD-like data and utilized for forgetting OOD knowledge. Extensive experiments demonstrate the effectiveness of ATEsc for improving DFKD against NTL teachers.
Ziming Hong, Runnan Chen, Zengmao Wang, Bo Han 0003, Bo Du 0001, Tongliang Liu
ICML3
2025 Learning to Exploit Leg Odometry Enables Terrain-Aware Quadrupedal Locomotion
abstract
The geometry of terrain is crucial for developing terrain-aware locomotion policies. Recent advancements in quadrupedal locomotion based on learning rely on depth information obtained from LiDARs and depth cameras. Despite the capabilities of these locomotion policies on terrains, they pose challenges in processing high-dimensional data in real time with onboard hardware. In this study, we develop a lightweight framework that utilizes only the intrinsic sensors of a quadrupedal robot to facilitate terrain-aware locomotion. We introduce a learning-based leg odometry, integrated with a locomotion policy trained through reinforcement learning. Utilizing blind localization from leg odometry alongside a pre-constructed height map enables the robot to navigate steps and stairs without incident.We assess the efficacy of our framework through simulations, where our results indicate that the robot achieves up to a 17% improvement in successful traversal rates and requires fewer point samples. By compensating for slippage during locomotion, our learning-based leg odometry surpasses traditional inertialleg odometry. Lastly, we validate the practical applicability of our models on a real robot, confirming their effectiveness in real-world settings.
Jiawei Jiang 0001, Bo Du 0001, Zengmao Wang
IROS4
2025 DM-PCL: Text-Driven Dual-Modal Prototype Consistency Learning for Weakly-Supervised Few-Shot Part Segmentation
Mengya Han, Yong Luo 0002, Han Hu 0003, Zengmao Wang, Lefei Zhang, Bo Du 0001, Ling-Yu Duan, Dacheng Tao
Int. J. Comput. Vis.4
2025 Detect Changes Like Humans: Incorporating Semantic Priors for Improved Change Detection
abstract
When given two similar images, humans identify their differences by comparing the appearance (e.g., color, texture) with the help of semantics (e.g., objects, relations). However, mainstream binary change detection models adopt a supervised training paradigm, where the annotated binary change map is the main constraint. Thus, such methods primarily emphasize difference-aware features between bi-temporal images, and the semantic understanding of changed landscapes is undermined, resulting in limited accuracy in the face of noise and illumination variations. To this end, this paper explores incorporating semantic priors from visual foundation models to improve the ability to detect changes. Firstly, we propose a Semantic-Aware Change Detection network (SA-CDNet), which transfers the knowledge of visual foundation models (i.e., FastSAM) to change detection. Inspired by the human visual paradigm, a novel dual-stream feature decoder is derived to distinguish changes by combining semantic-aware features and difference-aware features. Secondly, we explore a single-temporal pre-training strategy for better adaptation of visual foundation models. With pseudo-change data constructed from single-temporal segmentation datasets, we employ an extra branch of proxy semantic segmentation task for pre-training. We explore various settings like dataset combinations and landscape types, thus providing valuable insights. Experimental results on five challenging benchmarks demonstrate the superiority of our method over the existing state-of- the-art methods. The code is available at SA-CD.
Yuhang Gan, Wenjie Xuan, Zhiming Luo, Zengmao Wang, Juhua Liu, Bo Du 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Efficient Prompt Tuning of Large Vision-Language Model for Fine-Grained Ship Classification
abstract
Remote-sensing fine-grained ship classification (RS-FGSC) poses a significant challenge due to the high similarity between classes and the limited availability of labeled data, limiting the effectiveness of traditional supervised classification methods. Recent advancements in large pretrained vision-language models (VLMs) have demonstrated impressive capabilities in few-shot or zero-shot learning, particularly in understanding image content. This study delves into harnessing the potential of VLMs to enhance classification accuracy for unseen ship categories, which holds considerable significance in scenarios with restricted data due to cost or privacy constraints. Directly fine-tuning VLMs for RS-FGSC often encounters the challenge of overfitting the seen classes, resulting in suboptimal generalization to unseen classes, which highlights the difficulty in differentiating complex backgrounds and capturing distinct ship features. To address these issues, we introduce a novel prompt tuning technique that employs a hierarchical, multigranularity prompt design. Our approach integrates remote sensing ship priors through bias terms, learned from a small trainable network. This strategy enhances the model’s generalization capabilities while improving its ability to discern intricate backgrounds and learn discriminative ship features. Furthermore, we contribute to the field by introducing a comprehensive dataset, FGSCM-52, significantly expanding existing datasets with more extensive data and detailed annotations for less common ship classes. Extensive experimental evaluations demonstrate the superiority of our proposed method over current state-of-the-art techniques. The source code will be made publicly available.
Long Lan, Fengxiang Wang 0004, Xiangtao Zheng, Zengmao Wang, Xinwang Liu 0002
IEEE Trans. Geosci. Remote. Sens.4
2025 Multi-Modal Correction Network for Recommendation
abstract
Multi-modal contents have proven to be the powerful knowledge for recommendation tasks. Most state-of-the-art multi-modal recommendation methods mainly focus on aligning the semantic spaces of different modalities to enhance the item representations and do not pay much attention on the relevant knowledge in the multi-modalities for recommendation, resulting in that the positive effects of the relevant knowledge is reduced and the improvement of recommendation performance is limited. In this paper, we propose a multi-modal correction network termed MMCN to enhance the item representation with the important semantic knowledge in each modality by a residual structure with attention mechanisms and a hierarchical contrastive learning framework. The residual information is obtained through self-attention and cross-attention, which can learn the relevant knowledge across different modalities effectively. While hierarchical contrastive learning further captures the relevant knowledge not only at the feature level but also at the element-wise level with a matrix. Extensive experiments on three large-scale real-world datasets show the superiority of MMCN over state-of-the-art multi-modal recommendation methods.
Zengmao Wang, Yunzhen Feng, Xin Zhang 0091, Renjie Yang, Bo Du 0001
IEEE Trans. Knowl. Data Eng.1
2025 Medical Transformer With Mix Mask Generation for Thorax Disease Classification
abstract
Chest X-ray images have been highly involved in clinical diagnosis and treatment planning for thoracic disease. The process of medical images has attracted great attention in the machine learning community. However, the labeled medical images are limited and the regions of lesions are usually much smaller in the image. Most of the existing methods are prone to learning the spurious correlation for classification, resulting in poor generalization. In this paper, we propose a medical generation transformer network based on self-supervised learning and the adversarial strategy to capture the discriminative label-relevant regions with lesions in the images by extending the Chest X-ray images. In the proposed method, we first localize the label-relevant regions in each transformer layer. Then we keep the label-relevant regions to mask the image and construct the masked image with self-supervised learning. Thus we can generate more images to fine-tune the classification network with masked images that keep the label-relevant regions. Since the generated images are usually noisy to fine-tune the classification network, we adopt the adversarial probabilities to weight the importance of each generated image for training. Experimental results on two large-scale and popular chest X-ray datasets show that the proposed method can efficiently leverage the location of lesions to improve the performance of classification.
Ziyi Liu 0010, Zengmao Wang, Bo Du 0001
IEEE Trans. Multim.2
2025 Learning Intrinsic Invariance Within Intra-Class for Domain Generalization
abstract
Deep learning methods often struggle with the domain shift problem, leading to poor generalization on out-of-domain (OOD) data. To address the problem, domain generalization (DG) has been proposed to leverage the source domains to train a model that can generalize to OOD data. Existing domain generalization methods primarily focus on learning domain invariance, but they fail to ensure proximity among samples within the same category when domains are aligned for domain-invariant learning. Consequently, their generalization performance remains suboptimal. In this paper, we propose a novel approach to address this issue by iteratively approximating the category domain-invariant distribution from all domains. Our method involves an iterative loop where we initially estimate the domain-invariant distribution for each category by averaging the statistical characteristics across all domains. Then the adversarial perturbation alignment is adopted to keep each sample close to its corresponding category domain-invariant distribution. With the iterative loop, the deep network is optimized for robust domain invariance learning. Extensive experiments demonstrate that our proposed method consistently outperforms state-of-the-art approaches across various scenarios.
Chaoyang Zhou, Zengmao Wang, Bo Du 0001
IEEE Trans. Multim.2
2025 Controllable fashion rendering via brownian bridge diffusion model with latent sketch encoding
Zengmao Wang, Jituo Li, Wei Gao 0014
Vis. Comput.1
2024 Cycle Self-Refinement for Multi-Source Domain Adaptation
abstract
Multi-source domain adaptation (MSDA) aims to transfer knowledge from multiple source domains to the unlabeled target domain. In this paper, we propose a cycle self-refinement domain adaptation method, which progressively attempts to learn the dominant transferable knowledge in each source domain in a cycle manner. Specifically, several source-specific networks and a domain-ensemble network are adopted in the proposed method. The source-specific networks are adopted to provide the dominant transferable knowledge in each source domain for instance-level ensemble on predictions of the samples in target domain. Then these samples with high-confidence ensemble predictions are adopted to refine the domain-ensemble network. Meanwhile, to guide each source-specific network to learn more dominant transferable knowledge, we force the features of the target domain from the domain-ensemble network and the features of each source domain from the corresponding source-specific network to be aligned with their predictions from the corresponding networks. Thus the adaptation ability of source-specific networks and the domain-ensemble network can be improved progressively. Extensive experiments on Office-31, Office-Home and DomainNet show that the proposed method outperforms the state-of-the-art methods for most tasks.
Chaoyang Zhou, Zengmao Wang, Bo Du 0001, Yong Luo 0002
AAAI2
2024 Improving Generalized Zero-Shot Learning by Exploring the Diverse Semantics from External Class Names
abstract
Generalized Zero-Shot Learning (GZSL) methods often assume that the unseen classes are similar to seen classes, and thus perform poor when unseen classes are dissimilar to seen classes. Although some existing GZSL approaches can alleviate this issue by leveraging additional semantic information from test unseen classes, their generalization ability to dissimilar unseen classes is still unsatisfactory. This motivates us to study GZSL in the more practical setting, where unseen classes can be either similar or dissimilar to seen classes. In this paper, we propose a simple yet effective GZSL framework by exploring diverse semantics from external class names (DSECN), which is simultaneously robust on the similar and dissimilar unseen classes. This is achieved by introducing diverse semantics from external class names and aligning the introduced semantics to visual space using the classification head of pretrained network. Furthermore, we show that the design idea of DSECN can easily be integrate into other advanced GZSL approaches, such as the generative-based ones, and enhance their robustness for dissimilar unseen classes. Extensive experiments in the practical setting including both similar and dissimilar unseen classes show that our method significantly outperforms the state-of-the-art approaches on all datasets and can be trained very efficiently.
Yong Luo 0002, Zengmao Wang, Bo Du 0001
CVPR3
2024 LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation
Jing Zhang 0037, Di Wang 0023, Qiming Zhang 0001, Zengmao Wang, Bo Du 0001
IJCAI5
2024 Temporal Uplift Modeling for Online Marketing
abstract
In recent years, uplift modeling, also known as individual treatment effect (ITE) estimation, has seen wide applications in online marketing, such as delivering one-time issuance of coupons or discounts to motivate users' purchases. However, complex yet more realistic scenarios involving multiple interventions over time on users are still rarely explored. The challenges include handling the bias from time-varying confounders, determining optimal treatment timing, and selecting among numerous treatments. In this paper, to tackle the aforementioned challenges, we present a temporal point process-based uplift model (TPPUM) that utilizes users' temporal event sequences to estimate treatment effects via counterfactual analysis and temporal point processes. In this model, marketing actions are considered as treatments, user purchases as outcome events, and how treatments alter the future conditional intensity function of generating outcome events as the uplift. Empirical evaluations demonstrate that our method outperforms existing baselines on both real-world and synthetic datasets. In the online experiment conducted in a discounted bundle recommendation scenario involving an average of 3 to 4 interventions per day and hundreds of treatment candidates, we demonstrate how our model outperforms current state-of-the-art methods in selecting the appropriate treatment and timing of treatment, resulting in a 3.6% increase in application-level revenue.
Xin Zhang 0091, Kai Wang 0064, Zengmao Wang, Bo Du 0001, Runze Wu 0001, Tangjie Lv, Changjie Fan
KDD3
2024 SAT3D: Image-driven Semantic Attribute Transfer in 3D
abstract
GAN-based image editing task aims at manipulating image attributes in the latent space of generative models. Most of the previous 2D and 3D-aware approaches mainly focus on editing attributes in images with ambiguous semantics or regions from a reference image, which fail to achieve photographic semantic attribute transfer, such as the beard from a photo of a man. In this paper, we propose an image-driven Semantic Attribute Transfer method in 3D (SAT3D) by editing semantic attributes from a reference image. For the proposed method, the exploration is conducted in the style space of a pre-trained 3D-aware StyleGAN-based generator by learning the correlations between semantic attributes and style code channels. For guidance, we associate each attribute with a set of phrase-based descriptor groups, and develop a Quantitative Measurement Module (QMM) to quantitatively describe the attribute characteristics in images based on descriptor groups, which leverages the image-text comprehension capability of CLIP. During the training process, the QMM is incorporated into attribute losses to calculate attribute similarity between images, guiding target semantic transferring and irrelevant semantics preserving. We present our 3D-aware attribute transfer results across multiple domains and also conduct comparisons with classical 2D image editing methods, demonstrating the effectiveness and customizability of our SAT3D.
Zhijun Zhai, Zengmao Wang, Xiaoxiao Long, Kaixuan Zhou, Bo Du 0001
ACM Multimedia2
2024 What If the Input is Expanded in OOD Detection?
abstract
Out-of-distribution (OOD) detection aims to identify OOD inputs from unknown classes, which is important for the reliable deployment of machine learning models in the open world. Various scoring functions are proposed to distinguish it from in-distribution (ID) data. However, existing methods generally focus on excavating the discriminative information from a single input, which implicitly limits its representation dimension. In this work, we introduce a novel perspective, i.e., employing different common corruptions on the input space, to expand that. We reveal an interesting phenomenon termed *confidence mutation*, where the confidence of OOD data can decrease significantly under the corruptions, while the ID data shows a higher confidence expectation considering the resistance of semantic features. Based on that, we formalize a new scoring method, namely, *Confidence aVerage* (CoVer), which can capture the dynamic differences by simply averaging the scores obtained from different corrupted inputs and the original ones, making the OOD and ID distributions more separable in detection tasks. Extensive experiments and analyses have been conducted to understand and verify the effectiveness of CoVer.
Boxuan Zhang 0001, Jianing Zhu, Zengmao Wang, Tongliang Liu, Bo Du 0001, Bo Han 0003
NeurIPS3
2024 Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?
abstract
This paper investigates an under-explored challenge in large language models (LLMs): chain-of-thought prompting with noisy rationales, which include irrelevant or inaccurate reasoning thoughts within examples used for in-context learning. We construct NoRa dataset that is tailored to evaluate the robustness of reasoning in the presence of noisy rationales. Our findings on NoRa dataset reveal a prevalent vulnerability to such noise among current LLMs, with existing robust methods like self-correction and self-consistency showing limited efficacy. Notably, compared to prompting with clean rationales, base LLM drops by 1.4%-19.8% in accuracy with irrelevant thoughts and more drastically by 2.2%-40.4% with inaccurate thoughts. Addressing this challenge necessitates external supervision that should be accessible in practice. Here, we propose the method of contrastive denoising with noisy chain-of-thought (CD-CoT). It enhances LLMs' denoising-reasoning capabilities by contrasting noisy rationales with only one clean rationale, which can be the minimal requirement for denoising-purpose prompting. This method follows a principle of exploration and exploitation: (1) rephrasing and selecting rationales in the input space to achieve explicit denoising and (2) exploring diverse reasoning paths and voting on answers in the output space. Empirically, CD-CoT demonstrates an average improvement of 17.8% in accuracy over the base model and shows significantly stronger denoising capabilities than baseline methods. The source code is publicly available at: https://github.com/tmlr-group/NoisyRationales.
Zhanke Zhou, Rong Tao, Jianing Zhu, Yiwen Luo, Zengmao Wang, Bo Han 0003
NeurIPS5
2024 Boosting Semi-Supervised Object Detection in Remote Sensing Images With Active Teaching
abstract
The lack of object-level annotations poses a significant challenge for object detection in remote sensing images. To address this issue, active learning and semi-supervised learning techniques have been proposed to enhance the quality and quantity of annotations. Active learning focuses on selecting the most informative samples for annotation, while semi-supervised learning leverages the knowledge from unlabeled samples. In this paper, we propose a novel active learning method to boost semi-supervised object detection for remote sensing images with a teacher-student network, called SSOD-AT. The proposed method incorporates a RoI Comparison module (RoICM) to generate high-confidence pseudo-labels for Regions of Interest (RoIs). Meanwhile, the RoICM is utilized to identify the top-K uncertain images. To reduce redundancy in the top-K uncertain images for human labeling, a diversity criterion is introduced based on object-level prototypes of different categories using both labeled and pseudo-labeled images. Extensive experiments on DOTA and DIOR two popular datasets demonstrate that our proposed method outperforms state-of-the-art methods for object detection in remote sensing images. Compared with the best performance in the SOTA methods, the proposed method achieves 1% improvement at most cases in the whole active learning.
Boxuan Zhang 0001, Zengmao Wang, Bo Du 0001
IEEE Geosci. Remote. Sens. Lett.2
2024 Leveraging spatial residual attention and temporal Markov networks for video action understanding
Yangyang Xu 0001, Zengmao Wang
Neural Networks2
2024 MambaHSI: Spatial-Spectral Mamba for Hyperspectral Image Classification
abstract
Transformer has been extensively explored for hyperspectral image (HSI) classification. However, transformer poses challenges in terms of speed and memory usage because of its quadratic computational complexity. Recently, the Mamba model has emerged as a promising approach, which has strong long-distance modeling capabilities while maintaining a linear computational complexity. However, representing the HSI is challenging for the Mamba due to the requirement for an integrated spatial and spectral understanding. To remedy these drawbacks, we propose a novel HSI classification model based on a Mamba model, named MambaHSI, which can simultaneously model long-range interaction of the whole image and integrate spatial and spectral information in an adaptive manner. Specifically, we design a spatial Mamba block (SpaMB) to model the long-range interaction of the whole image at the pixel-level. Then, we propose a spectral Mamba block (SpeMB) to split the spectral vector into multiple groups, mine the relations across different spectral groups, and extract spectral features. Finally, we propose a spatial-spectral fusion module (SSFM) to adaptively integrate spatial and spectral features of a HSI. To our best knowledge, this is the first image-level HSI classification model based on the Mamba. We conduct extensive experiments on four diverse HSI datasets. The results demonstrate the effectiveness and superiority of the proposed model for HSI classification. This reveals the great potential of Mamba to be the next-generation backbone for HSI models. Codes are available athttps://github.com/li-yapeng/MambaHSI.
Yong Luo 0002, Lefei Zhang, Zengmao Wang, Bo Du 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Deep Session Heterogeneity-Aware Network for Click Through Rate Prediction
abstract
CTR (Click-Through Rate) prediction plays an essential role in online advertising systems. Most existing works attempt to capture users’ interests from sessions by assuming that behaviors within a session are homogeneous. However, user interest may change frequently. Thus it is hard to guarantee that behaviors in a session are homogeneous, resulting in users’ interests extracted from sessions being biased. In this paper, we propose a model named Deep Session Heterogeneity-aware Network (DSHN) by learning the relationships of behaviors within sessions and the relevance between the session and target item to alleviate the influence of irrelevant or heterogeneous sessions. We design a heterogeneity-aware mechanism to learn the heterogeneity of items within a session. Then we further design two modules: the Session Heterogeneity Learning module and the Relevance Inference module. The Session Heterogeneity Learning module weighs each session by summarizing the variation of session interest with and without any behavior. The relevance Inference module learns the relevance between the target item and each session in a similar way by learning session interest with and without the target item. Extensive experiments on four datasets demonstrate that our proposed DSHN achieves better results compared to the state-of-the-art.
Xin Zhang 0091, Zengmao Wang, Bo Du 0001, Jia Wu 0001, Erli Meng
IEEE Trans. Knowl. Data Eng.2
2024 Domain Complementary Adaptation by Leveraging Diversity and Discriminability From Multiple Sources
abstract
Due to the lack of labeled data in many real-world applications, unsupervised domain adaptation has attracted a great deal of attention in the machine learning community through its use of labeled data from source domains. However, how to make full use of the discriminative information from different sources remains a challenge due to various domain gaps. In this article, we propose a domain complementary adaptation method by leveraging the diversity between sources and the discriminability of each source with contrastive learning. In the proposed method, we adopt several branch networks, denoted as domain branch networks, to learn different views of discriminative domain-invariant features from each source. Moreover, an ensemble classification network trained with domain-invariant features from all domain branch networks is adopted to guide the domain branch networks in providing diverse knowledge. We design a domain mutual contrastive loss by forcing the domain branch networks to be different from one another and be consistent with the ensemble classification network to learn diverse domain-invariant features. To further improve the discriminability of domain branch networks, a domain structure-oriented contrastive loss is proposed to learn the discriminative intrinsic neighborhood structure across each source and target domain. Extensive experiments on the Office-31, Office-Home and DomainNet datasets show that the proposed method outperforms state-of-the-art methods.
Chaoyang Zhou, Zengmao Wang, Bo Du 0001
IEEE Trans. Multim.2
2023 Unified active and semi-supervised learning for hyperspectral image classification
Zengmao Wang, Bo Du 0001
GeoInformatica1
2023 KE-X: Towards subgraph explanations of knowledge graph embedding based on knowledge information gain
Guojia Wan, Yibing Zhan, Zengmao Wang, Liang Ding 0006, Zhigao Zheng 0001, Bo Du 0001
Knowl. Based Syst.4
2023 PAENL: personalized attraction enhanced network learning for recommendation
Zengmao Wang, Jedi S. Shang
Neural Comput. Appl.2
2023 Coherence-aware context aggregator for fast video object segmentation
Meng Lan, Zengmao Wang
Pattern Recognit.3
2023 Active Learning With Co-Auxiliary Learning and Multi-Level Diversity for Image Classification
abstract
Due to the fact that it is expensive and time-consuming to annotate a large amount of data, the available labeled data to train a deep neural network is usually scarce, resulting in the poor performance of the deep classification network in most cases. In this paper, we propose a novel active learning method to train a competitive deep classification neural network by labeling a limited amount of diverse images. The proposed method advances most of the active learning methods in two aspects. First, active learning is enhanced with a co-auxiliary learning strategy. We use an auxiliary network to provide diverse pseudo-labels for the primary network. The auxiliary network is also adopted to remove a part of redundancy information from the candidate pool of active learning when image predictions of the primary network and the auxiliary network are the same. Meanwhile, the primary network can also provide pseudo-labels to improve the performance of the auxiliary network. Second, we further remove the redundancy within the query samples with multi-level diversity selection. The multi-level diversity strategy not only considers the feature-level diversity but also the class-level diversity with the predictions of the primary network, and it can select both representative and uncertain samples effectively and efficiently from the large-scale data. The extensive experiments on four large-scale datasets show that the proposed method outperforms the state-of-the-art methods.
Zengmao Wang, Bo Du 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Graph-aware collaborative reasoning for click-through rate prediction
Xin Zhang 0091, Zengmao Wang, Bo Du 0001
World Wide Web (WWW)2
2022 Multi-marginal Contrastive Learning for Multilabel Subcellular Protein Localization
abstract
Protein subcellular localization(PSL) is an important task to study human cell functions and cancer pathogenesis. It has attracted great attention in the computer vision community. However, the huge size of immune histochemical (IHC) images, the disorganized location distribution in different tissue images and the limited training images are always the challenges for the PSL to learn a strong generalization model with deep learning. In this paper, we propose a deep protein subcellular localization method with multi-marginal contrastive learning to perceive the same PSLs in different tissue images and different PSLs within the same tissue image. In the proposed method, we learn the representation of an IHC image by fusing the global features from the downsampled images and local features from the selected patches with the activation map to tackle the oversize of an IHC image. Then a multi-marginal attention mechanism is proposed to generate contrastive pairs with different margins and improve the discriminative features of PSL patterns effectively. Finally, the ensemble prediction of each IHC image is obtained with different patches. The results on the benchmark datasets show that the proposed method achieves significant improvements for the PSL task.
Ziyi Liu 0010, Zengmao Wang, Bo Du 0001
CVPR2
2022 Self-Guided Network for Fine-Grained Object Localization Using Weakly Supervised Learning
abstract
Weakly supervised object localization (WSOL) aims at pre-dicting the location of objects with image-level labels. Fine-grained WSOL task has its characteristic challenge compared with generic object localization. The structural information of the objects in fine-grained benchmark has little relevance to class. Previous WSOL works mainly focus on learning the most class discriminative parts recursively, leading to se-rious structural feature missing issue. In this paper, we pro-pose a self-guided network (SGN) which consists of two branch deep classification networks. It adopts a coarse-to-fine strategy to detect the structural information of the ob-ject. First, we devise a self-adaptive method (SAM) to de-tect the most body structure of the object by directly leveraging the feature recognition ability of the first classifier. Then, an object structure generation (OSG) method is proposed in the fine localization phase. OSG helps the second classifier to learn the boundary feature of the object with less back-ground noise. Extensive experiments on four well-known fine-grained benchmarks, including CUB, FGVC Aircraft, Stanford Dogs, and Stanford Cars show that the proposed SGN outperforms the state-of-the-art WSOL methods.
Xiangyu Qu, Zengmao Wang
ICME2
2022 Self-paced Supervision for Multi-source Domain Adaptation
abstract
Multi-source domain adaptation has attracted great attention in machine learning community. Most of these methods focus on weighting the predictions produced by the adaptation networks of different domains. Thus the domain shifts between certain of domains and target domain are not effectively relieved, resulting in that these domains are not fully exploited and even may have a negative influence on multi-source domain adaptation task. To address such challenge, we propose a multi-source domain adaptation method to gradually improve the adaptation ability of each source domain by producing more high-confident pseudo-labels with self-paced learning for conditional distribution alignment. The proposed method first trains several separate domain branch networks with single domains and an ensemble branch network with all domains. Then we obtain some high-confident pseudo-labels with the branch networks and learn the branch specific pseudo-labels with self-paced learning. Each branch network reduces the domain gap by aligning the conditional distribution with its branch specific pseudo-labels and the pseudo-labels provided by all branch networks. Experiments on Office31, Office-Home and DomainNet show that the proposed method outperforms the state-of-the-art methods.
Zengmao Wang, Chaoyang Zhou, Bo Du 0001, Fengxiang He
IJCAI1
2022 CFDA: Collaborative Filtering with Dual Autoencoder for Recommender System
abstract
In recent years, deep neural networks have been widely used in recommender systems. Neural collaborative filtering is a popular work to model complex interactions between users and items with deep learning. However, methods that are based on collaborative filtering usually focus on learning embedding with the factorization of pairwise interactions, thereby causing embedding to be insufficient in capturing the complex relationships between users and items. To alleviate the above problem, in this paper, we propose a novel recommendation method based on collaborative filtering with dual autoencoder (CFDA). In the proposed method, we use dual autoencoder to learn hidden representations of users and items simultaneously, and we minimize the deviation of the training data by learning the user and item representations. Extensive experiments on several datasets demonstrate that the proposed method outperforms the baseline methods that are based on neural collaborative filtering.
Zengmao Wang
IJCNN2
2022 Unsupervised Domain Adaptation with Implicit Pseudo Supervision for Semantic Segmentation
abstract
Pseudo-labelling is a popular technique in unsupervised domain adaptation for semantic segmentation. However, pseudo labels are noisy and inevitably have confirmation bias due to the discrepancy between source and target domains and training process. In this paper, we train the model by the pseudo labels which are implicitly produced by itself to learn new complementary knowledge about target domain. Specifically, we propose a tri-learning architecture, where every two branches produce the pseudo labels to train the third one. And we align the pseudo labels based on the similarity of the probability distributions for each two branches. To further implicitly utilize the pseudo labels, we maximize the distances of features for different classes and minimize the distances for the same classes by triplet loss. Extensive experiments on GTA5 to Cityscapes and SYNTHIA to Cityscapes tasks show that the proposed method has considerable improvements.
Wanyu Xu, Zengmao Wang
IJCNN2
2022 Deep Dynamic Interest Learning With Session Local and Global Consistency for Click-Through Rate Predictions
abstract
Click-through rate (CTR) prediction is the core task in an online advertising system. How to capture the users’ dynamic interests through the behavior sequences to predict CTR is a challenging problem in the real-world applications. To address this challenge, we propose a deep dynamic interest learning by learning the local sessions and global sessions within sequences for CTR prediction. The local sessions and global sessions are used to capture the short-term dynamic interests and the long-term interests of users, respectively. To explore the heterogeneous behaviors within the global session efficiently, the interests in global sessions are forced to be consistent with that in local sessions. Then, the bi-LSTM network is adopted to capture the dynamic interests with the behaviors across global sessions effectively. Experimental results on various datasets show that the proposed method outperforms several state-of-the-art CTR prediction methods.
Xin Zhang 0091, Zengmao Wang, Bo Du 0001
IEEE Trans. Ind. Informatics2
2021 Multi-subband and Multi-subepoch Time Series Feature Learning for EEG-based Sleep Stage Classification
abstract
EEG plays an important role in the analysis and recognition of brain activity, and which has great potential in the field of biometrics, while EEG-based time series classification is complicated and difficult due to the nonstationary characteristics and individual difference. In this paper, we investigate the EEG signal classification problem and propose a multi-subband and multi-subepoch time series feature learning (MMTSFL) method for automatic sleep stage classification. Specifically, MMTSFL first decomposes multiple subbands with various frequency from raw EEG signals and partitions the obtained subbands in-to multiple consecutive subepochs, and then employs time series feature learning to obtain effective discriminant features. Moreover, amplitude-time based signal features are extracted from each subepoch to represent dynamic variation of EEG signals, and MMTSFL conduct further multipurpose feature learning for specific features, consistent features and temporal features simultaneously. Experiment results on three classification tasks of sleep quality evaluation, fatigue detection and sleep disease diagnosis demonstrate the superiority of the proposed method.
Panfeng An, Jianhui Zhao 0001, Zengmao Wang, Bo Du 0001
IJCB5
2021 Multistage reaction-diffusion equation network for image super-resolution
abstract
Abstract Deep learning‐based models have progressed considerably in single‐image super‐resolution. A high‐resolution pattern generation task is performed at the end of convolution neural networks (CNNs) with some convolution‐based operations in these models. However, this process may be difficult because all the work is done through the remarkable learning ability of CNN without any specific learning target. Reaction‐diffusion equation (RDE) is a mechanism involved in the pattern generation process that can serve as a guide for super‐resolution. It is proposed to embed RDE into a super‐resolution network by designing a reaction‐diffusion process block (RDPB) in this study. The proposed RDPB uses Euler method for iteratively solving one particular RDE, which is determined by the parameter generated through CNN. Accordingly, this module guides and leads the CNN in generating patterns for image super‐resolution. Moreover, a multistage framework is constructed to guide each network module further. On the basis of these two designs, the multistage reaction‐diffusion equation network is proposed for image super‐resolution. Experimental results demonstrated that the proposed model can obtain findings consistent with the conclusions of state‐of‐the‐art methods with a relatively shallow structure and small model size.
Xiaofeng Pu, Zengmao Wang
IET Image Process.2
2021 Joint representation learning with ratings and reviews for recommendation
Zengmao Wang, Haifeng Xia, Gang Chun
Neurocomputing1
2021 Incorporating Distribution Matching into Uncertainty for Multiple Kernel Active Learning
abstract
Due to the lack of the labeled data and the complex structures of various data, it is very hard to learn the uncertainty and representativeness accurately in active learning. In this paper, we propose a multiple kernel active learning framework that incorporates a group regularizer of distribution information into the estimation of uncertainty. The proposed method takes the advantage of multiple kernel learning to learn the kernel space in which the complex structures can be well captured by kernel weights. Meanwhile, we have developed an efficient optimization algorithm to solve the proposed method. Experimental results on twelve UCI benchmark data sets and eight subsets of ImageNet show that the proposed method outperforms several state-of-the-art active learning methods. Moreover, we also have applied the proposed method to multiple feature scenario on Caltech101, and the promising results are also obtained compared with single feature scenario.
Zengmao Wang, Bo Du 0001, Weiping Tu, Lefei Zhang, Dacheng Tao
IEEE Trans. Knowl. Data Eng.1
2020 Bi-adapting kernel learning for unsupervised domain adaptation
Zengmao Wang, Weiping Tu, Bo Du 0001, Yanxiang Cheng
Neurocomputing1
2020 Domain Adaptation With Neural Embedding Matching
abstract
Domain adaptation aims to exploit the supervision knowledge in a source domain for learning prediction models in a target domain. In this article, we propose a novel representation learning-based domain adaptation method, i.e., neural embedding matching (NEM) method, to transfer information from the source domain to the target domain where labeled data is scarce. The proposed approach induces an intermediate common representation space for both domains with a neural network model while matching the embedding of data from the two domains in this common representation space. The embedding matching is based on the fundamental assumptions that a cross-domain pair of instances will be close to each other in the embedding space if they belong to the same class category, and the local geometry property of the data can be maintained in the embedding space. The assumptions are encoded via objectives of metric learning and graph embedding techniques to regularize and learn the semisupervised neural embedding model. We also provide a generalization bound analysis for the proposed domain adaptation method. Meanwhile, a progressive learning strategy is proposed and it improves the generalization ability of the neural network gradually. Experiments are conducted on a number of benchmark data sets and the results demonstrate that the proposed method outperforms several state-of-the-art domain adaptation methods and the progressive learning strategy is promising.
Zengmao Wang, Bo Du 0001, Yuhong Guo
IEEE Trans. Neural Networks Learn. Syst.1
2019 Leveraging Ratings and Reviews with Gating Mechanism for Recommendation
abstract
Recommender system plays an important role to provide people with personalized information based on their history records. However, it is still a challenge to capture the preference of users accurately due to the sparsity of rating data and the heterogeneity of review data. In this paper, we propose a hybrid deep collaborative filtering model that jointly learns latent representations from ratings and reviews. Specifically, the model learns the rating feature and textual feature based on ratings and reviews simultaneously. Two embedding layers are employed to learn rating feature for users and items based on the user and item interactions, and two attention-based GRU networks learn context-aware representation from user and item reviews. Then a gating mechanism is used to leverage contributions from rating feature and textual feature. Experimental results on six real-world datasets demonstrate the superior performance of the proposed method over several state-of-the-art methods. Moreover, the keywords in reviews can be highlighted to interpret the predictions with the attention mechanism.
Haifeng Xia, Zengmao Wang, Bo Du 0001, Lefei Zhang, Gang Chun
CIKM2
2019 On combining active and transfer learning for medical data classification
abstract
This study presents a novel algorithm which combines active learning (AL) and transfer learning for medical data classification. The main idea of the proposed algorithm is iteratively querying a small number of informative unlabelled target samples, and, at the same time, removing the source samples which do not fit with the posterior probability distributions in the target domain, so as to combine the basic idea of AL with transfer learning. The experimental results obtained in the classification of the datasets from the University of California Irvine (UCI) Machine Learning Repository and The Cancer Imaging Archive (TCIA) confirm the effectiveness of the proposed algorithm.
Bo Du 0001, Zengmao Wang, Lefei Zhang
IET Comput. Vis.4
2019 Domain Adaptation With Discriminative Distribution and Manifold Embedding for Hyperspectral Image Classification
abstract
Hyperspectral remote sensing image classification has drawn a great attention in recent years due to the development of remote sensing technology. To build a high confident classifier, the large number of labeled data is very important, e.g., the success of deep learning technique. Indeed, the acquisition of labeled data is usually very expensive, especially for the remote sensing images, which usually needs to survey outside. To address this problem, in this letter, we propose a domain adaptation method by learning the manifold embedding and matching the discriminative distribution in source domain with neural networks for hyperspectral image classification. Specifically, we use the discriminative information of source image to train the classifier for the source and target images. To make the classifier can work well on both domains, we minimize the distribution shift between the two domains in an embedding space with prior class distribution in the source domain. Meanwhile, to avoid the distortion mapping of the target domain in the embedding space, we try to keep the manifold relation of the samples in the embedding space. Then, we learn the embedding on source domain and target domain by minimizing the three criteria simultaneously based on a neural network. The experimental results on two hyperspectral remote sensing images have shown that our proposed method can outperform several baseline methods.
Zengmao Wang, Bo Du 0001, Qian Shi 0001, Weiping Tu
IEEE Geosci. Remote. Sens. Lett.1
2019 Robust Graph-Based Semisupervised Learning for Noisy Labeled Data via Maximum Correntropy Criterion
abstract
Semisupervised learning (SSL) methods have been proved to be effective at solving the labeled samples shortage problem by using a large number of unlabeled samples together with a small number of labeled samples. However, many traditional SSL methods may not be robust with too much labeling noisy data. To address this issue, in this paper, we propose a robust graph-based SSL method based on maximum correntropy criterion to learn a robust and strong generalization model. In detail, the graph-based SSL framework is improved by imposing supervised information on the regularizer, which can strengthen the constraint on labels, thus ensuring that the predicted labels of each cluster are close to the true labels. Furthermore, the maximum correntropy criterion is introduced into the graph-based SSL framework to suppress labeling noise. Extensive image classification experiments prove the generalization and robustness of the proposed SSL method.
Bo Du 0001, Zengmao Wang, Lefei Zhang, Dacheng Tao
IEEE Trans. Cybern.3
2018 Matrix completion with Preference Ranking for Top-N Recommendation
abstract
Matrix completion has become a popular method for top-N recommendation due to the low rank nature of sparse rating matrices. However, many existing methods produce top-N recommendations by recovering a user-item matrix solely based on a low rank function or its relaxations, while ignoring other important intrinsic characteristics of the top-N recommendation tasks such as preference ranking over the items. In this paper, we propose a novel matrix completion method that integrates the low rank and preference ranking characteristics of recommendation matrix under a self-recovery model for top-N recommendation. The proposed method is formulated as a joint minimization problem and solved using an ADMM algorithm. We conduct experiments on E-commerce datasets. The experimental results show the proposed approach outperforms several state-of-the-art methods.
Zengmao Wang, Yuhong Guo, Bo Du 0001
IJCAI1
2017 On Gleaning Knowledge from Multiple Domains for Active Learning
abstract
How can a doctor diagnose new diseases with little historical knowledge, which are emerging over time? Active learning is a promising way to address the problem by querying the most informative samples. Since the diagnosed cases for new disease are very limited, gleaning knowledge from other domains (classical prescriptions) to prevent the bias of active leaning would be vital for accurate diagnosis. In this paper, a framework that attempts to glean knowledge from multiple domains for active learning by querying the most uncertain and representative samples from the target domain and calculating the importance weights for re-weighting the source data in a single unified formulation is proposed. The weights are optimized by both a supervised classifier and distribution matching between the source domain and target domain with maximum mean discrepancy. Besides, a multiple domains active learning method is designed based on the proposed framework as an example. The proposed method is verified with newsgroups and handwritten digits data recognition tasks, where it outperforms the state-of-the-art methods.
Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Ruimin Hu, Dacheng Tao
IJCAI1
2017 Multi-class active learning: A hybrid informative and representative criterion inspired approach
abstract
Labeling each instance in a large-scale data set is extremely labor- and time-consuming. One way to alleviate this problem is active learning, which aims to discover the most valuable instances for labeling to construct a powerful classifier with low generalization error. Considering both informativeness and representativeness provides a promising way to design a practical active learning. However, most existing active learning methods select instances favoring either informativeness or representativeness. Meanwhile, many are designed based on the binary class, so that they may present suboptimal solutions on the data sets with multiple classes. In this paper, a hybrid informative and representative criterion based multi-class active learning approach is proposed. We combine the informativeness and representativeness into one formulation, which can be solved under a unified framework. The informativeness is measured by the margin minimum while the representative information is measured by the maximum mean discrepancy. By minimizing the loss risk, we generalize the loss risk minimization principle to the multi-class active learning setting. Hence, the proposed method is not only suitable to the binary class but also the multiple classes. We conduct our experiments on twelve benchmark UCI data sets, and the experimental results demonstrate that the proposed method performs better than some state-of-the-art methods.
Zengmao Wang, Bo Du 0001, Lefei Zhang
IJCNN1
2017 Exploring Representativeness and Informativeness for Active Learning
abstract
How can we find a general way to choose the most suitable samples for training a classifier? Even with very limited prior information? Active learning, which can be regarded as an iterative optimization procedure, plays a key role to construct a refined training set to improve the classification performance in a variety of applications, such as text analysis, image recognition, social network modeling, etc. Although combining representativeness and informativeness of samples has been proven promising for active sampling, state-of-the-art methods perform well under certain data structures. Then can we find a way to fuse the two active sampling criteria without any assumption on data? This paper proposes a general active learning framework that effectively fuses the two criteria. Inspired by a two-sample discrepancy problem, triple measures are elaborately designed to guarantee that the query samples not only possess the representativeness of the unlabeled data but also reveal the diversity of the labeled data. Any appropriate similarity measure can be employed to construct the triple measures. Meanwhile, an uncertain measure is leveraged to generate the informativeness criterion, which can be carried out in different ways. Rooted in this framework, a practical active learning algorithm is proposed, which exploits a radial basis function together with the estimated probabilities to construct the triple measures and a modified best-versus-second-best strategy to construct the uncertain measure, respectively. Experimental results on benchmark datasets demonstrate that our algorithm consistently achieves superior performance over the state-of-the-art active learning algorithms.
Bo Du 0001, Zengmao Wang, Lefei Zhang, Liangpei Zhang 0001, Wei Liu 0005, Jialie Shen 0001, Dacheng Tao
IEEE Trans. Cybern.2
2017 A Novel Semisupervised Active-Learning Algorithm for Hyperspectral Image Classification
abstract
Less training samples are a challenging problem in hyperspectral image classification. Active learning and semisupervised learning are two promising techniques to address the problem. Active learning solves the problem by improving the quality of the training samples, while semisupervised learning solves the problem by increasing the quantity of the training samples. However, they pay too much attention to the discriminative information in the unlabeled data, leading to information bias to train supervised models, and much more effort to label samples. Therefore, a method to discover representativeness and discriminativeness by semisupervised active learning is proposed. It takes advantages of both active learning and semisupervised learning. The representativeness and discriminativeness are discovered with a labeling process based on a supervised clustering technique and classification results. Specifically, the supervised clustering results can discover important structural information in the unlabeled data, and the classification results are also highly confidential in the active-learning process. With these clustering results and classification results, we can assign pseudolabels to the unlabeled data. Meanwhile, the unlabeled samples that cannot be assigned with pseudolabels with high confidence at each iteration are regarded as candidates in active learning. The methodology is validated on four hyperspectral data sets. Significant improvements in classification accuracy are achieved by the proposed method with respect to the state-of-the-art methods.
Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Xiuping Jia
IEEE Trans. Geosci. Remote. Sens.1
2017 Robust and Discriminative Labeling for Multi-Label Active Learning Based on Maximum Correntropy Criterion
abstract
Multi-label learning draws great interests in many real world applications. It is a highly costly task to assign many labels by the oracle for one instance. Meanwhile, it is also hard to build a good model without diagnosing discriminative labels. Can we reduce the label costs and improve the ability to train a good model for multi-label learning simultaneously? Active learning addresses the less training samples problem by querying the most valuable samples to achieve a better performance with little costs. In multi-label active learning, some researches have been done for querying the relevant labels with less training samples or querying all labels without diagnosing the discriminative information. They all cannot effectively handle the outlier labels for the measurement of uncertainty. Since maximum correntropy criterion (MCC) provides a robust analysis for outliers in many machine learning and data mining algorithms, in this paper, we derive a robust multi-label active learning algorithm based on an MCC by merging uncertainty and representativeness, and propose an efficient alternating optimization method to solve it. With MCC, our method can eliminate the influence of outlier labels that are not discriminative to measure the uncertainty. To make further improvement on the ability of information measurement, we merge uncertainty and representativeness with the prediction labels of unknown data. It cannot only enhance the uncertainty but also improve the similarity measurement of multi-label data with labels information. Experiments on benchmark multi-label data sets have shown a superior performance than the state-of-the-art methods.
Bo Du 0001, Zengmao Wang, Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao
IEEE Trans. Image Process.2
2016 Multi-label Active Learning Based on Maximum Correntropy Criterion: Towards Robust and Discriminative Labeling
Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001, Dacheng Tao
ECCV (3)1
2016 A batch-mode active learning framework by querying discriminative and representative samples for hyperspectral image classification
Zengmao Wang, Bo Du 0001, Lefei Zhang, Liangpei Zhang 0001
Neurocomputing1
2015 Batch Mode Active Learning for Geographical Image Classification
Zengmao Wang, Bo Du 0001, Lefei Zhang, Wenbin Hu 0001, Dacheng Tao, Liangpei Zhang 0001
APWeb1