Zhihao Wu 0002

dblp:27/8792-2 · DBLP profile ↗
← Back
30ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0002-2704-0614ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease Diagnosis
abstract
Multimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention mechanisms - treat cross-modal interactions indiscriminately, lacking adaptive modulation of modality-specific contributions under varying clinical scenarios. We propose an adaptive fusion framework that dynamically routes and refines multimodal signals for enhancing disease recognition. The framework comprises two key components: 1) Dynamic Cross-Modal Expert Routing (CMER), which selectively activates convolutional neural network (CNN) experts from one modality based on contextual guidance from the other, ensuring only the most relevant feature extractors contribute to fusion; and 2) Top-K Expert-Guided Wavelet Fusion (TEWF), which performs discrete wavelet transform (DWT) to decompose selected features into low- and high-frequency subbands. Cross-modal attention is then applied specifically to high-frequency components, where lesion-specific microstructures reside, enabling frequency-aware fusion. Finally, inverse DWT (IDWT) reconstructs the fused representation, weighted by CMER-derived importance scores to amplify informative modality cues while suppressing redundancy. Experimental validation on two multimodal retinal datasets demonstrates that our method achieves state-of-the-art performance, outperforming existing fusion strategies by significant margins in disease classification accuracy and robustness.
Haoran Li 0024, Haoyu Cao 0002, Yongting Hu, Qihao Xu, Chengliang Liu 0003, Xiaoling Luo 0001, Zhihao Wu 0002, Yong Xu 0001, Wei Wang 0169
AAAI8
2026 Incomplete Multi-view Diabetic Retinopathy Grading via Self-Supervised Inter- and Intra-View Restoration
abstract
Multi-view diabetic retinopathy (DR) grading has achieved remarkable performance by capturing more comprehensive pathological features than single-view methods. However, complete multi-view fundus images are often difficult to obtain in clinical practice, and the performance degrades significantly when fewer views are available. To overcome this limitation, we propose the first incomplete multi-view DR grading framework, aiming to provide accurate diagnosis regardless of the number of available views. It introduces two novel modules. First, cross-view spatial correlation attention (CSCA) captures region correlations across views, automatically identifying and fusing diagnostically relevant spatial features to improve feature representation. Second, self-supervised mask consistency learning (SMCL) formulates a novel pretext task of missing-view information reconstruction by strategically masking inter- and intra-view regions, enabling the model to infer complete features from incomplete views. Benefiting from CSCA and SMCL, our method enhances structural feature consistency across views and effectively compensates for missing information during DR grading. Extensive experiments demonstrate that our method achieves state-of-the-art grading performance, particularly under realistic conditions where some views are unavailable.
Zhihao Wu 0002, Jie Wen 0001, Wuzhen Shi, LinLin Shen
AAAI1
2026 Weakly Supervised Salient Object Detection with Text Supervision
Zhihao Wu 0002, Jie Wen 0001, LinLin Shen, Xiaopeng Fan 0001, Yong Xu 0001, Jian Yang 0003, David Zhang 0001
Int. J. Comput. Vis.1
2026 Zero-shot referring expression comprehension via guidance of Multimodal Large Language Models
Rouyi Li, Shiyi Zheng, Zhihao Wu 0002, LinLin Shen
Pattern Recognit.4
2026 DeCenter: Density-Center Guided Perception Enhancement for UAV Object Detection
abstract
Unmanned aerial vehicle (UAV) object detection is essential for applications such as surveillance, agriculture, and disaster response. However, UAV imagery often contains small, dense, and occluded objects, posing challenges for existing methods. To address these challenges, we propose DeCenter, a novel Density-Center Guided Perception Enhancement framework for UAV object detection. DeCenter is composed of two key modules that jointly enhance the perception of small and crowded objects. First, the Density-Guided Object Center Heatmap Generator (DOCHG) adaptively generates Gaussian kernel-based heatmaps according to local density information, guiding the model to emphasize central neighborhoods of objects in crowded regions. This mechanism reduces overlaps between adjacent instances and alleviates missed detections under occlusion. Second, the Density-Center Feature Enhancement module (DCFE) integrates complementary cues from density features and object centers, adaptively balancing region-level object distribution with fine-grained localization. By fusing these signals, DCFE enhances the quality of feature representations, making them more discriminative for dense small objects while suppressing background noise. Experimental results on VisDrone and UAVDT datasets show that DeCenter achieves competitive overall accuracy with clear improvements in detecting dense small objects, offering an effective solution for UAV object detection. The code will be available at https://github.com/bluuzzz/decenter.
Zhiqing Shi, Zhihao Wu 0002, Jie Wen 0001, Mu Li 0005, Xiaopeng Fan 0001, Yaowei Wang 0001, LinLin Shen
IEEE Trans. Circuits Syst. Video Technol.2
2026 Adaptive Fine-Grained Fusion Network for Multimodal UAV Object Detection
abstract
Multimodal perception and fusion play a vital role in uncrewed aerial vehicle (UAV) object detection. Existing methods typically adopt global fusion strategies across modalities. However, due to illumination variation, the effectiveness of RGB and infrared modalities may differ across local regions within the same image, particularly in UAV perspectives where occlusions and dense small objects are prevalent, leading to suboptimal performance of global fusion methods. To address this issue, we propose an adaptive fine-grained fusion network for multimodal UAV object detection. First, we design a local feature consistency-based modality fusion module, which adaptively assigns local fusion weights according to the structural consistency of high-response regions across modalities, thereby enabling more effective aggregation of object-relevant features. Second, we introduce a mutual information-guided feature contrastive loss to encourage the preservation of modality-specific information during the early training phase. Experimental results demonstrate that the proposed method effectively addresses the issue of object occlusion in UAV perspectives, achieving state-of-the-art performance on multimodal UAV object detection benchmarks. Code will be available at https://github.com/lingf5877/AFFNet.
Zhanyan Tang, Zhihao Wu 0002, Mu Li 0005, Jie Wen 0001, Bob Zhang 0001, Yong Xu 0001, Jianqiang Li 0001
IEEE Trans. Image Process.2
2025 Mixture of Experts as Representation Learner for Deep Multi-View Clustering
abstract
Multi-view clustering (MVC) aims to integrate information from diverse data sources to facilitate the clustering process, which has achieved considerable success in various real-world applications. However, previous MVC methods typically employ one of two strategies: (1) designing separate feature extraction pipelines for each view, which restricts their ability to fully exploit collaborative potential; or (2) employing a single shared representation module, which hinders the capture of diverse, view-specific representations. To tackle these challenges, we introduce Deep Multi-View Clustering via Collaborative Experts (DMVC-CE), a novel MVC approach that employs the Mixture of Experts (MoE) framework. DMVC-CE incorporates a gating network that dynamically selects multiple experts for handling each data sample, capturing diverse and complementary information from different views. Additionally, to ensure balanced expert utilization and maintain their diversity, we introduce an equilibrium loss and a multi-expert distinctiveness enhancer. The equilibrium loss prevents excessive reliance on specific experts, while the distinctiveness enhancer encourages each expert to specialize in different aspects of the data, thereby promoting diversity in learned representations. Comprehensive experiments on various multi-view benchmark datasets demonstrate the superiority of DMVC-CE compared to state-of-the-art MVC baselines.
Yunhe Zhang 0001, Jinyu Cai, Zhihao Wu 0002, Pengyang Wang, See-Kiong Ng
AAAI3
2025 Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus Diseases
abstract
With the advancement of computer vision, numerous models have been proposed for screening of fundus diseases. However, the recognition of multiple fundus diseases is often hampered by the simultaneous presence of multiple disease types and the confluence of lesion types in fundus images. This paper addresses these challenges by conceptualizing them as multi-level feature fusion and self-supervised disease-indicative feature learning problems. We decode fundus images at various levels of granularity to delineate scenarios wherein multiple diseases and lesions co-occur. To effectively integrate these features, we introduce a hierarchical vision transformer (HVT) that adeptly captures both inter-level and intra-level dependencies. A novel forward-attention module is proposed to enhance the integration of lower-level semantic information into higher semantic layers, thereby enriching the representation of complex features. Additionally, we introduce a novel self-supervised mask-consistent feature learner (MCFL). Unlike traditional mask-autoencoders that reconstruct original images using encoder-decoder structures, MCFL utilizes a teacher-student framework to reconstruct mask-consistent feature maps. In this setup, exponential moving averaging is employed to derive classification-guided features, serving as labels for reconstruction rather than merely reconstructing the original images. This innovative approach facilitates the extraction of disease-indicative features. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art models.
Wei Wang 0169, Xiaoling Luo 0001, Zhihao Wu 0002, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001
AAAI4
2025 Weakly Supervised Salient Object Detection With Oversize Bounding Box
Zhihao Wu 0002, Yong Xu 0001, Jian Yang 0003, David Zhang 0001
Int. J. Comput. Vis.1
2025 Multi-view diabetic retinopathy grading via cross-view spatial alignment and adaptive vessel reinforcing
Xiaoyan Dou, Xiaoling Luo 0001, Zhihao Wu 0002, Chengliang Liu 0003, Tianyi Luo, Jie Wen 0001, Bingo Wing-Kuen Ling, Yong Xu 0001, Wei Wang 0169
Pattern Recognit.4
2025 Foregroundness-Aware Task Disentanglement and Self-Paced Curriculum Learning for Domain Adaptive Object Detection
abstract
Unsupervised domain adaptive object detection (UDA-OD) is a challenging problem since it needs to locate and recognize objects while maintaining the generalization ability across domains. Most existing UDA-OD methods directly integrate the adaptive modules into the detectors. This integration procedure can significantly sacrifice the detection performances, though it enhances the generalization ability. To solve this problem, we propose an effective framework, named foregroundness-aware task disentanglement and self-paced curriculum adaptation (FA-TDCA), to disentangle the UDA-OD task into four independent subtasks of source detector pretraining, classification adaptation, location adaptation, and target detector training. The disentanglement can transfer the knowledge effectively while maintaining the detection performance of our model. In addition, we propose a new metric, i.e., foregroundness, and use it to evaluate the confidence of the location result. We use both foregroundness and classification confidence to assess the label quality of the proposals. For effective knowledge transfer across domains, we utilize a self-paced curriculum learning paradigm to train adaptors and gradually improve the quality of the pseudolabels associated with the target samples. Experiment results indicate that our method achieves state-of-the-art results on four cross-domain object detection tasks.
Linhui Xiao, Chengliang Liu 0003, Zhihao Wu 0002, Yong Xu 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Spatial Continuity and Nonequal Importance in Salient Object Detection With Image-Category Supervision
abstract
Due to the inefficiency of pixel-level annotations, weakly supervised salient object detection with image-category labels (WSSOD) has been receiving increasing attention. Previous works usually endeavor to generate high-quality pseudolabels to train the detectors in a fully supervised manner. However, we find that the detection performance is often limited by two types of noise contained in pseudolabels: 1) holes inside the object or at the edge and outliers in the background and 2) missing object portions and redundant surrounding regions. To mitigate the adverse effects caused by them, we propose local pixel correction (LPC) and key pixel attention (KPA), respectively, based on two key properties of desirable pseudolabels: 1) spatial continuity, meaning an object region consists of a cluster of adjacent points; and 2) nonequal importance, meaning pixels have different importance for training. Specifically, LPC fills holes and filters out outliers based on summary statistics of the neighborhood as well as its size. KPA directs the focus of training toward ambiguous pixels in multiple pseudolabels to discover more accurate saliency cues. To evaluate the effectiveness of our method, we design a simple yet strong baseline we call weakly supervised saliency detector with Transformer (WSSDT) and unify the proposed modules into WSSDT. Extensive experiments on five datasets demonstrate that our method significantly improves the baseline and outperforms all existing congeneric methods. Moreover, we establish the first benchmark to evaluate WSSOD robustness. The results show that our method can improve detection robustness as well. The code and robustness benchmark are available at https://github.com/Horatio9702/SCNI.
Zhihao Wu 0002, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Biomedically Informed ECG Synthesis: Customizing Cardiac Cycle Phases with Diffusion Model
abstract
Cardiovascular diseases are a major global health challenge, with electrocardiography (ECG) being critical for diagnosis and monitoring. As artificial intelligence and automated ECG diagnostic technologies rapidly advance, the demand for large-scale ECG databases continues to grow. Generative ECG has become a mainstream method to enhance database size and diversity. However, existing methods typically generate ECG randomly or focus on limited physiological categories, lacking the ability to synthesize ECG with varying physiological features and cardiac cycles, which is crucial for various practical applications. In response to this need, we propose a novel approach introducing a diffusion model called DIFF-ECG to generate precisely customized ECG that accurately reflect diverse cardiac conditions. Segmentation-based quality assessments confirmed that the synthesized ECG accurately followed the specified cardiac cycle information, with our model significantly outperforming baseline diffusion and GAN-based methods. Therefore, our approach addresses the critical need for generating clinically relevant and customizable ECG, contributing significantly to the field of automated cardiac disease diagnosis. By enabling fine-tuning of cardiac cycle phases, our method significantly expands the application range of generative ECG, potentially improving the diagnostic accuracy for rare diseases and advancing personalized medicine.
Wei Wang 0169, Zhihao Wu 0002, Suyu Dong, Gongning Luo, Kuanquan Wang
BIBM4
2024 Kernel Readout for Graph Neural Networks
Jiajun Yu, Zhihao Wu 0002, Jinyu Cai, Adele Lu Jia, Jicong Fan 0001
IJCAI2
2024 Misclassification in Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) aims to train detectors using only image-category labels. Current methods typically first generate dense class-agnostic proposals and then select objects based on the classification scores of these proposals. These methods mainly focus on selecting the proposal having high Intersection-over-Union with the true object location, while ignoring the problem of misclassification, which occurs when some proposals exhibit semantic similarities with objects from other categories due to viewing perspective and background interference. We observe that the positive class that is misclassified typically has the following two characteristics: 1) It is usually misclassified as one or a few specific negative classes, and the scores of these negative classes are high; 2) Compared to other negative classes, the score of the positive class is relatively high. Based on these two characteristics, we propose misclassification correction (MCC) and misclassification tolerance (MCT) respectively. In MCC, we establish a misclassification memory bank to record and summarize the class-pairs with high frequencies of potential misclassifications in the early stage of training, that is, cases where the score of a negative class is significantly higher than that of the positive class. In the later stage of training, when such cases occur and correspond to the summarized class-pairs, we select the top-scoring negative class proposal as the positive training example. In MCT, we decrease the loss weights of misclassified classes in the later stage of training to avoid them dominating training and causing misclassification of objects from other classes that are semantically similar to them during inference. Extensive experiments on the PASCAL VOC and MS COCO demonstrate our method can alleviate the problem of misclassification and achieve the state-of-the-art results.
Zhihao Wu 0002, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001
IEEE Trans. Image Process.1
2024 Information Recovery-Driven Deep Incomplete Multiview Clustering Network
abstract
Incomplete multiview clustering (IMC) is a hot and emerging topic. It is well known that unavoidable data incompleteness greatly weakens the effective information of multiview data. To date, existing IMC methods usually bypass unavailable views according to prior missing information, which is considered a second-best scheme based on evasion. Other methods that attempt to recover missing information are mostly applicable to specific two-view datasets. To handle these problems, in this article, we propose an information-recovery-driven-deep IMC network, termed as RecFormer. Concretely, a two-stage autoencoder network with self-attention structure is built to synchronously extract high-level semantic representations of multiple views and recover the missing data. Besides, we develop a recurrent graph reconstruction mechanism that cleverly leverages the restored views to promote representation learning and further data reconstruction. Visualization of recovery results are given and sufficient experimental results confirm that our RecFormer has obvious advantages over other top methods.
Chengliang Liu 0003, Jie Wen 0001, Zhihao Wu 0002, Xiaoling Luo 0001, Chao Huang 0008, Yong Xu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Enhanced Spatial Feature Learning for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) has become an effective paradigm, which requires only class labels to train object detectors. However, WSOD detectors are prone to learn highly discriminative features corresponding to local objects rather than complete objects, resulting in imprecise object localization. To address the issue, designing backbones specifically for WSOD is a feasible solution. However, the redesigned backbone generally needs to be pretrained on large-scale ImageNet or trained from scratch, both of which require much more time and computational costs than fine-tuning. In this article, we explore to optimize the backbone without losing the availability of the original pretrained model. Since the pooling layer summarizes neighborhood features, it is crucial to spatial feature learning. In addition, it has no learnable parameters, so its modification will not change the pretrained model. Based on the above analysis, we further propose enhanced spatial feature learning (ESFL) for WSOD, which first takes full advantage of multiple kernels in a single pooling layer to handle multiscale objects and then enhances above-average activations within the rectangular neighborhood to alleviate the problem of ignoring unsalient object parts. The experimental results on the PASCAL VOC and the MS COCO benchmarks demonstrate that ESFL can bring significant performance improvement for the WSOD method and achieve state-of-the-art results.
Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 DICNet: Deep Instance-Level Contrastive Network for Double Incomplete Multi-View Multi-Label Classification
abstract
In recent years, multi-view multi-label learning has aroused extensive research enthusiasm. However, multi-view multi-label data in the real world is commonly incomplete due to the uncertain factors of data collection and manual annotation, which means that not only multi-view features are often missing, and label completeness is also difficult to be satisfied. To deal with the double incomplete multi-view multi-label classification problem, we propose a deep instance-level contrastive network, namely DICNet. Different from conventional methods, our DICNet focuses on leveraging deep neural network to exploit the high-level semantic representations of samples rather than shallow-level features. First, we utilize the stacked autoencoders to build an end-to-end multi-view feature extraction framework to learn the view-specific representations of samples. Furthermore, in order to improve the consensus representation ability, we introduce an incomplete instance-level contrastive learning scheme to guide the encoders to better extract the consensus information of multiple views and use a multi-view weighted fusion module to enhance the discrimination of semantic features. Overall, our DICNet is adept in capturing consistent discriminative representations of multi-view multi-label data and avoiding the negative effects of missing views and missing labels. Extensive experiments performed on five datasets validate that our method outperforms other state-of-the-art methods.
Chengliang Liu 0003, Jie Wen 0001, Xiaoling Luo 0001, Chao Huang 0008, Zhihao Wu 0002, Yong Xu 0001
AAAI5
2023 Highly Confident Local Structure Based Consensus Graph Learning for Incomplete Multi-view Clustering
abstract
Graph-based multi-view clustering has attracted extensive attention because of the powerful clustering-structure representation ability and noise robustness. Considering the reality of a large amount of incomplete data, in this paper, we propose a simple but effective method for incomplete multi-view clustering based on consensus graph learning, termed as HCLS_CGL. Unlike existing methods that utilize graph constructed from raw data to aid in the learning of consistent representation, our method directly learns a consensus graph across views for clustering. Specifically, we design a novel confidence graph and embed it to form a confidence structure driven consensus graph learning model. Our confidence graph is based on an intuitive similar-nearest-neighbor hypothesis, which does not require any additional information and can help the model to obtain a high-quality consensus graph for better clustering. Numerous experiments are performed to confirm the effectiveness of our method.
Jie Wen 0001, Chengliang Liu 0003, Gehui Xu, Zhihao Wu 0002, Chao Huang 0008, Lunke Fei, Yong Xu 0001
CVPR4
2023 Masked Two-channel Decoupling Framework for Incomplete Multi-view Weak Multi-label Learning
abstract
Multi-view learning has become a popular research topic in recent years, but research on the cross-application of classic multi-label classification and multi-view learning is still in its early stages. In this paper, we focus on the complex yet highly realistic task of incomplete multi-view weak multi-label learning and propose a masked two-channel decoupling framework based on deep neural networks to solve this problem. The core innovation of our method lies in decoupling the single-channel view-level representation, which is common in deep multi-view learning methods, into a shared representation and a view-proprietary representation. We also design a cross-channel contrastive loss to enhance the semantic property of the two channels. Additionally, we exploit supervised information to design a label-guided graph regularization loss, helping the extracted embedding features preserve the geometric structure among samples. Inspired by the success of masking mechanisms in image and text analysis, we develop a random fragment masking strategy for vector features to improve the learning ability of encoders. Finally, it is important to emphasize that our model is fully adaptable to arbitrary view and label absences while also performing well on the ideal full data. We have conducted sufficient and convincing experiments to confirm the effectiveness and advancement of our model.
Chengliang Liu 0003, Jie Wen 0001, Chao Huang 0008, Zhihao Wu 0002, Xiaoling Luo 0001, Yong Xu 0001
NeurIPS5
2023 Graph Convolutional Kernel Machine versus Graph Convolutional Networks
abstract
Graph convolutional networks (GCN) with one or two hidden layers have been widely used in handling graph data that are prevalent in various disciplines. Many studies showed that the gain of making GCNs deeper is tiny or even negative. This implies that the complexity of graph data is often limited and shallow models are often sufficient to extract expressive features for various tasks such as node classification. Therefore, in this work, we present a framework called graph convolutional kernel machine (GCKM) for graph-based machine learning. GCKMs are built upon kernel functions integrated with graph convolution. An example is the graph convolutional kernel support vector machine (GCKSVM) for node classification, for which we analyze the generalization error bound and discuss the impact of the graph structure. Compared to GCNs, GCKMs require much less effort in architecture design, hyperparameter tuning, and optimization. More importantly, GCKMs are guaranteed to obtain globally optimal solutions and have strong generalization ability and high interpretability. GCKMs are composable, can be extended to large-scale data, and are applicable to various tasks (e.g., node or graph classification, clustering, feature extraction, dimensionality reduction). The numerical results on benchmark datasets show that, besides the aforementioned advantages, GCKMs have at least competitive accuracy compared to GCNs.
Zhihao Wu 0002, Zhao Zhang 0001, Jicong Fan 0001
NeurIPS1
2023 Selecting High-Quality Proposals for Weakly Supervised Object Detection With Bottom-Up Aggregated Attention and Phase-Aware Loss
abstract
Weakly supervised object detection (WSOD) has received widespread attention since it requires only image-category annotations for detector training. Many advanced approaches solve this problem by a two-phase learning framework, that is, instance mining that classifies generated proposals via multiple instance learning, and instance refinement that iteratively refines bounding boxes using the supervision produced by the preceding stage. In this paper, we observe that the detection performance is usually limited by imprecise supervision, including part domination and untight boxes. To mitigate their adverse effects, we focus on selecting high-quality proposals as the supervision for WSOD. To be specific, for the issue of part domination, we propose bottom-up aggregated attention which incorporates low-level features from shallow layers to improve location representation of top-level features. In this manner, the proposals corresponding to entire objects can get high scores. Its advantage is that it can be flexibly plugged into the WSOD framework since there is no need to attach learnable parameters or learning branches. As regards the problem of untight boxes, we propose a phase-aware loss, which is the first work to measure supervision quality by the loss in the instance mining phase, to highlight correct boxes and suppress untight ones. In this work, we unify the proposed two modules into the framework of online instance classifier refinement. Extensive experiments on the PASCAL VOC and the MS COCO demonstrate that our method can significantly improve the performance of WSOD and achieve the state-of-the-art results. The code is available at https://github.com/Horatio9702/BUAA_PALoss.
Zhihao Wu 0002, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001
IEEE Trans. Image Process.1
2023 Localized Sparse Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering, which aims to solve the clustering problem on the incomplete multi-view data with partial view missing, has received more and more attention in recent years. Although numerous methods have been developed, most of the methods either cannot flexibly handle the incomplete multi-view data with arbitrary missing views or do not consider the negative factor of information imbalance among views. Moreover, some methods do not fully explore the local structure of all incomplete views. To tackle these problems, this paper proposes a simple but effective method, named localized sparse incomplete multi-view clustering (LSIMVC). Different from the existing methods, LSIMVC intends to learn a sparse and structured consensus latent representation from the incomplete multi-view data by optimizing a sparse regularized and novel graph embedded multi-view matrix factorization model. Specifically, in such a novel model based on the matrix factorization, a norm based sparse constraint is introduced to obtain the sparse low-dimensional individual representations and the sparse consensus representation. Moreover, a novel local graph embedding term is introduced to learn the structured consensus representation. Different from the existing works, our local graph embedding term aggregates the graph embedding task and consensus representation learning task into a concise term. Furthermore, to reduce the imbalance factor of incomplete multi-view learning, an adaptive weighted learning scheme is introduced to LSIMVC. Comprehensive experimental results performed on six incomplete multi-view databases verify that the performance of our LSIMVC is superior to the state-of-the-art IMC approaches.
Chengliang Liu 0003, Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Chao Huang 0008
IEEE Trans. Multim.2
2023 Multiple Instance Detection Networks With Adaptive Instance Refinement
abstract
Weakly supervised object detection (WSOD) aims to train object detectors by using only image-level annotations. Many recent works on WSOD adopt multiple instance detection networks (MIDN), which usually generate a certain number of proposals and regard proposal classification as a latent model learning within image classification. However, these methods tend to detect salient object, salient object parts and clustered objects due to lack of instance-level annotations during training. Thus a core issue is how to guarantee that the network learn as many objects with precise bounding boxes as possible. In this paper, we address this issue by exploiting the potential of proposal scores during training. We propose an adaptive instance refinement (AIR) framework with three novel designs, which can be integrated with MIDN into a single network. Specifically, adaptive instance mining attempts to discover all positive instances according to the score distribution of proposals and their spatial similarity. Adaptive score modulation dynamically adjusts proposal scores to make the network focus more on instances with different difficulties in different training iterations. Adaptive knowledge refinement distills important information from all previous stages by the weighted average of proposal scores. The experimental results on the PASCAL VOC 2007 and 2012 benchmarks and the MS COCO benchmark demonstrate that AIR significantly improves the performance of the original MIDN and achieves the state-of-the-art results.
Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Jian Yang 0003, David Zhang 0001
IEEE Trans. Multim.1
2022 Deep Object Detection with Example Attribute Based Prediction Modulation
abstract
Deep object detectors suffer from the gradient contribution imbalance during training. In this paper, we point out that such imbalance can be ascribed to the imbalance in example attributes, e.g., difficulty and shape variation degree. We further propose example attribute based prediction modulation (EAPM) to address it. In EAPM, first, the attribute of an example is defined by the prediction and the corresponding ground truth. Then, a modulating factor w.r.t the example attribute is introduced to modulate the prediction error. Finally, the new prediction and the ground-truth are input into the loss function. Essentially, we adjust the gradients of examples with specific attributes to reweight their contribution on the global gradients. We apply EAPM with focal loss and balanced L1 loss to simultaneously solve the imbalance in classification and localization. The experimental results on MS COCO demonstrate that EAPM can bring substantial improvement for deep object detectors.
Zhihao Wu 0002, Chengliang Liu 0003, Chao Huang 0008, Jie Wen 0001, Yong Xu 0001
ICASSP1
2022 Pixel-Level Anomaly Detection via Uncertainty-aware Prototypical Transformer
abstract
Pixel-level visual anomaly detection, which aims to recognize the abnormal areas from images, plays an important role in industrial fault detection and medical diagnosis. However, it is a challenging task due to the following reasons: i) the large variation of anomalies; and ii) the ambiguous boundary between anomalies and their normal surroundings. In this work, we present an uncertainty-aware prototypical transformer (UPformer), which takes into account both the diversity and uncertainty of anomaly to achieve accurate pixel-level visual anomaly detection. To this end, we first design a memory-guided prototype learning transformer encoder to learn and memorize the prototypical representations of anomalies for enabling the model to capture the diversity of anomalies. Additionally, an anomaly detection uncertainty quantizer is designed to learn the distributions of anomaly detection for measuring the anomaly detection uncertainty. Furthermore, an uncertainty-aware transformer decoder is proposed to leverage the detection uncertainties to guide the model to focus on the uncertain areas and generate the final detection results. As a result, our method achieves more accurate anomaly detection by combining the benefits of prototype learning and uncertainty estimation. Experimental results on five datasets indicate that our method achieves state-of-the-art anomaly detection performance.
Chao Huang 0008, Chengliang Liu 0003, Zheng Zhang 0006, Zhihao Wu 0002, Jie Wen 0001, Qiuping Jiang, Yong Xu 0001
ACM Multimedia4
2022 Abnormal Event Detection Using Deep Contrastive Learning for Intelligent Video Surveillance System
abstract
The continuous developments of urban and industrial environments have increased the demand for intelligent video surveillance. Deep learning has achieved remarkable performance for anomaly detection in surveillance videos. Previous approaches achieve anomaly detection with a single-pretext task (image reconstruction or prediction) and detect anomalies by larger reconstruction error or poor prediction. However, they cannot fully exploit the discriminative semantics and temporal context information. Moreover, tackling anomaly detection with a single pretext task is suboptimal due to the nonalignment between the pretext task and anomaly detection. In this article, we propose a temporal-aware contrastive network (TAC-Net) to address the abovementioned problems of anomaly detection for intelligence video surveillance. TAC-Net is an unsupervised method that utilizes deep contrastive self-supervised learning to capture the high-level semantic features and tackles anomaly detection with multiple self-supervised tasks. During inference phase, the multiple task losses and contrastive similarity are utilized to calculate the anomaly score. Experimental results show that our method is superior to state-of-the-art approaches on three benchmarks, which demonstrates the validity and advancement of TAC-Net.
Chao Huang 0008, Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Qiuping Jiang, Yaowei Wang 0001
IEEE Trans. Ind. Informatics2
2021 Structural Deep Incomplete Multi-view Clustering Network
abstract
In recent years, incomplete multi-view clustering has drawn increasing attention due to the existence of large amounts of unlabeled incomplete data whose views are not fully observed in the practical applications. Although many traditional methods have been extended to address the incomplete learning problem, most of them exploit the shallow models and ignore the geometric structure. To address these issues, we proposed a structural deep incomplete multi-view clustering network. Specifically, the proposed method can simultaneously explore the high-level features and high-order geometric structure information of data with several view-specific graph convolutional encoder networks and can directly obtain the optimal clustering indicator matrix in one stage. Experimental results on several datasets with the comparison of state-of-the-art methods validate the superiority of the proposed method.
Jie Wen 0001, Zhihao Wu 0002, Zheng Zhang 0006, Lunke Fei, Bob Zhang 0001, Yong Xu 0001
CIKM2
2020 DIMC-net: Deep Incomplete Multi-view Clustering Network
abstract
In this paper, a new deep incomplete multi-view clustering network, called DIMC-net, is proposed to address the challenge of multi-view clustering on missing views. In particular, DIMC-net designs several view-specific encoders to extract the high-level information of multiple views and introduces a fusion graph based constraint to explore the local geometric information of data. To reduce the negative influence of missing views, a weighted fusion layer is introduced to obtain the consensus representation shared by all views. Moreover, a clustering layer is introduced to guarantee that the obtained consensus representation is the best one for the clustering task. Compared with the existing deep learning based approaches, DIMC-net is more flexible and efficient since it can handle all kinds of incomplete cases and directly produce the clustering results. Experimental results show that DIMC-net achieves significant improvement over state-of-the-art incomplete multi-view clustering methods.
Jie Wen 0001, Zheng Zhang 0006, Zhao Zhang 0001, Zhihao Wu 0002, Lunke Fei, Yong Xu 0001, Bob Zhang 0001
ACM Multimedia4
2020 Lightweight image super-resolution with enhanced CNN
Chunwei Tian, Ruibin Zhuge, Zhihao Wu 0002, Yong Xu 0001, Wangmeng Zuo, Chen Chen 0001, Chia-Wen Lin
Knowl. Based Syst.3