VLDB 2026 Research / reviewers in the wild / expert
Haofeng Zhang 0001
dblp:17/4297 · also Hao-Feng Zhang 0001, Hao-feng Zhang 0001
· DBLP profile ↗
112ranked-venue papers
13as first author
84since 2021 · last 2026
0000-0002-4039-7618ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 8 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 5 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 13 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 7 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Foundation-Adaptive Integrated Refinement for Generalized Category DiscoveryabstractThe potential of Generalized Category Discovery (GCD) lies in its ability to identify previously undiscovered patterns in both labeled and unlabeled data by leveraging insights from partially labeled training samples. However, interference can arise due to the model's dual focus on discovering both novel and known categories, often leading to conflicts that obscure true patterns in the dataset. This paper presents a divide-and-conquer framework, Foundation-Adaptive Integrated Refinement (FAIR), which fine-tunes pretrained foundational weights for various purposes, divided into Foundation (pretrained weights), Adaptive (weights fine-tuned with a variance-preserving loss), and Integrated (weights adjusted for both labeled and unlabeled data). The Adaptive utilizes a newly proposed adaptive contrastive loss that introduces variances within classes to preserve the individuality of representations. The Integrated addresses inherent estimation errors while dynamically estimating the number of categories, incorporating a cosine-based perturbation mechanism as a relaxed margin to accommodate potential ground-truth deviations, rather than relying on biased estimates. Extensive experiments on six benchmark datasets demonstrate our method's effectiveness, outperforming state-of-the-art algorithms, especially on fine-grained datasets. Yuwei Bian, Yazhou Yao, Haofeng Zhang 0001 |
AAAI | 4 |
| 2026 | Learning a Fix and Explore Framework for Continuous Generalized Category DiscoveryabstractTo address the limitations of transductive learning in evolving real-world scenarios where unknown categories may continuously emerge, Continual Generalized Category Discovery (C-GCD) presents a novel paradigm that extends conventional category discovery frameworks. Unlike traditional static learning environments, C-GCD requires models to incrementally discover novel categories across multiple operational phases while maintaining discrimination capabilities for previously learned classes, posing significant challenges in balancing stability and plasticity. Prior approaches typically employ parameter-level knowledge distillation from historical models to alleviate catastrophic forgetting, which effectively preserves prior knowledge and optimizes computational efficiency. However, our analysis reveals that the persistent availability of samples from previous stages enables more sophisticated knowledge preservation strategies. Specifically, we present a Fix and Explore strategy that employs distinct learning methodologies for different types of potential data, aiming to preserve the features of old categories as much as possible and gradually exploring the potential distribution of new class latent spaces, we can enhance the model's ability to discover novel categories. This paper investigates this effect and introduces a novel heuristic paradigm to solve the C-GCD problem, called Fix and Explore (FaE), which aims to provide sufficient imaginative space for new classes while preserving the classification ability for old tasks. We conducted experiments across multiple datasets and performed detailed comparisons. The results demonstrate that our method achieves state-of-the-art performance at each stage across all datasets. Chunming Li, Haofeng Zhang 0001 |
AAAI | 3 |
| 2026 | Bridging Granularity Gaps: Hierarchical Semantic Learning for Cross-domain Few-shot SegmentationabstractCross-domain Few-shot Segmentation (CD-FSS) aims to segment novel classes from target domains that are not involved in training and have significantly different data distributions from the source domain, using only a few annotated samples, and recent years have witnessed significant progress on this task. However, existing CD-FSS methods primarily focus on style gaps between source and target domains while ignoring segmentation granularity gaps, resulting in insufficient semantic discriminability for novel classes in target domains. Therefore, we propose a Hierarchical Semantic Learning (HSL) framework to tackle this problem. Specifically, we introduce a Dual Style Randomization (DSR) module and a Hierarchical Semantic Mining (HSM) module to learn hierarchical semantic features, thereby enhancing the model's ability to recognize semantics at varying granularities. DSR simulates target domain data with diverse foreground-background style differences and overall style variations through foreground and global style randomization respectively, while HSM leverages multi-scale superpixels to guide the model to mine intra-class consistency and inter-class distinction at different granularities. Additionally, we also propose a Prototype Confidence-modulated Thresholding (PCMT) module to mitigate segmentation ambiguity when foreground and background are excessively similar. Extensive experiments are conducted on four popular target domain datasets, and the results demonstrate that our method achieves state-of-the-art performance. Sujun Sun, Haowen Gu, Yanxu Ren, Mingwu Ren, Haofeng Zhang 0001 |
AAAI | 6 |
| 2026 | Beyond Quadratic: Linear-Time Change Detection with RWKVabstractExisting paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this conflict. By building upon the Receptance Weighted Key Value (RWKV) framework, our ChangeRWKV uniquely combines the parallelizable training of Transformers with the linear-time inference of RNNs. Our approach core features two key innovations: a hierarchical RWKV encoder that builds multi-resolution feature representation, and a novel Spatial-Temporal Fusion Module (STFM) engineered to resolve spatial misalignments across scales while distilling fine-grained temporal discrepancies. ChangeRWKV not only achieves state-of-the-art performance on the LEVIR-CD benchmark, with an 85.46% IoU and 92.16% F1 score, but does so while drastically reducing parameters and FLOPs compared to previous leading methods. This work demonstrates a new, efficient, and powerful paradigm for operational-scale change detection. Gensheng Pei, Tao Chen 0012, Xia Yuan, Haofeng Zhang 0001, Xiangbo Shu, Yazhou Yao |
AAAI | 5 |
| 2026 | Few-shot Medical Image Segmentation via Boundary-extended Prototypes and Momentum Inference
Yazhou Zhu 0001, Yang Long 0001, Haofeng Zhang 0001 |
Comput. Vis. Image Underst. | 5 |
| 2026 | Text-vision fusion and semantic adaptive labeling for compositional zero-shot learning
Run Shi, Chenyi Jiang, Chunyan Xu, Haofeng Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | CoTeach-CLIP: Cross-modal collaborative teachers for zero-shot point cloud recognition
Jiabao Zuo, Haofeng Zhang 0001, Quanchen Zhou, Huan Wang 0013, Mingwu Ren |
Neurocomputing | 2 |
| 2026 | Diving into the Details: Holistic and partial feature fusion network for few-shot object counting
Haofeng Zhang 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | Learning clique-based inter-class affinity for compositional zero-shot learning
Chenyi Jiang, Qiaolin Ye, Zebin Wu 0001, Haofeng Zhang 0001 |
Pattern Recognit. | 5 |
| 2026 | Learning to transport for open set domain generalization
Chunming Li, Yang Long 0001, Haofeng Zhang 0001 |
Pattern Recognit. | 4 |
| 2026 | Leveraging language to generalize natural images to few-shot medical image segmentation
Feifan Song 0004, Yuntian Bo, Yang Long 0001, Haofeng Zhang 0001 |
Pattern Recognit. Lett. | 5 |
| 2026 | Bridging the semantic gap for zero-shot object counting with language-augmented pseudo-exemplars
Yanxu Ren, Haofeng Zhang 0001 |
Pattern Recognit. Lett. | 4 |
| 2026 | Contrastive Graph Modeling for Cross-Domain Few-Shot Medical Image SegmentationabstractCross-domain few-shot medical image segmentation (CD-FSMIS) offers a promising and data-efficient solution for medical applications where annotations are severely scarce and multimodal analysis is required. However, existing methods typically filter out domain-specific information to improve generalization, which inadvertently limits cross-domain performance and degrades source-domain accuracy. To address this, we present Contrastive Graph Modeling (C-Graph), a framework that leverages the structural consistency of medical images as a reliable domain-transferable prior. We represent image features as graphs, with pixels as nodes and semantic affinities as edges. A Structural Prior Graph (SPG) layer is proposed to capture and transfer target-category node dependencies and enable global structure modeling through explicit node interactions. Building upon SPG layers, we introduce a Subgraph Matching Decoding (SMD) mechanism that exploits semantic relations among nodes to guide prediction. Furthermore, we design a Confusion-minimizing Node Contrast (CNC) loss to mitigate node ambiguity and subgraph heterogeneity by contrastively enhancing node discriminability in the graph space. Our method significantly outperforms prior CD-FSMIS approaches across multiple cross-domain benchmarks, achieving state-of-the-art performance while simultaneously preserving strong segmentation accuracy on the source domain. Our code is available at https://github.com/primebo1/C-Graph. Yuntian Bo, Tao Zhou 0002, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | Uncertainty-Guided Prototype Reliability Enhancement Network for Few-Shot Medical Image SegmentationabstractFew-Shot Learning (FSL) has garnered increasing attention for data-scarce scenarios, particularly in medical segmentation tasks where only a few labeled data points are available. Existing few-shot segmentation methods typically learn prototypes from support images and employ nearest-neighbor searching to segment query images. Despite notable progress, effectively learning prototypes for each class remains a challenging task to achieve promising results. In this paper, we propose an Uncertainty-guided Prototype Reliability Enhancement Network (UPRE-Net) for few-shot medical image segmentation. Specifically, we present a dual-support branch to maximize the extraction of information from support images through augmentation techniques. To enhance the reliability of prototypes, we propose an Uncertainty-guided Prototype Generation (UPG) module. Within the UPG module, we first extract both global and local prototypes for each class and then apply uncertainty measures to select the most informative prototypes. Additionally, to effectively combine the prediction results from the dual-support branch, we present a Reliable Dynamic Fusion (RDF) module. This module dynamically integrates the two prediction results to generate a more reliable output. Furthermore, we present an Uncertainty-induced Weighted Loss (UWL) to ensure that the model pays more attention to these regions with high uncertainty. Experiments on four benchmark medical image datasets demonstrate that our proposed model significantly outperforms state-of-the-art methods. The code will be released at https://github.com/taozh2017/UPRENet. Tao Zhou 0002, Kaiwen Huang 0002, Yi Zhou 0007, Haofeng Zhang 0001, Boqiang Fan, Huazhu Fu |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Adversarial Prototypical Perturbation for Cross-domain Few-shot Medical Image SegmentationabstractFew-shot medical image segmentation (FSMIS) has become one of the potential solutions for limited annotated medical image analysis. However, the realistic multiple domains of medical data demands the FSMIS models generalizing across domains. Thus, the cross-domain few-shot medical image segmentation (CD-FSMIS) is introduced, and we propose the Adversarial Prototypical Perturbation (APP) model which employs the adversarial learning strategy in the prototype learning process for gaining the domain robust prototypes. Specifically, the method consists of two components: adversarial signal formulation (ASF) and interactive prototypical attack (IPA). The ASF module collects perturbations from the gained gradients from both of the intra-class variation measurement loss and the inter-class variation measurement loss, and the IPA module aims to impose gained perturbations on the prototypical representation construction process with the two stages of interactive attacking manner. Additionally, a local-imbalance aware whitening loss is designed to resist the shift-sensitive local components for further enforcing to learn the domain robust prototypical representation in the IPA module. Extensive experiments are conducted on three cross-domain medical imaging datasets, and the results demonstrate that our model outperforms the state-of-the-art few-shot medical image segmentation methods. The code is available at https://github.com/YazhouZhu19/APP . Yazhou Zhu 0001, Tong Xin 0002, Haofeng Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | FAMNet: Frequency-aware Matching Network for Cross-domain Few-shot Medical Image SegmentationabstractExisting few-shot medical image segmentation (FSMIS) models fail to address a practical issue in medical imaging: the domain shift caused by different imaging techniques, which limits the applicability to current FSMIS tasks. To overcome this limitation, we focus on the cross-domain few-shot medical image segmentation (CD-FSMIS) task, aiming to develop a generalized model capable of adapting to a broader range of medical image segmentation scenarios with limited labeled data from the novel target domain. Inspired by the characteristics of frequency domain similarity across different domains, we propose a Frequency-aware Matching Network (FAMNet), which includes two key components: a Frequency-aware Matching (FAM) module and a Multi-Spectral Fusion (MSF) module. The FAM module tackles two problems during the meta-learning phase: 1) intra-domain variance caused by the inherent support-query bias, due to the different appearances of organs and lesions, and 2) inter-domain variance caused by different medical imaging techniques. Additionally, we design an MSF module to integrate the different frequency features decoupled by the FAM module, and further mitigate the impact of inter-domain variance on the model's segmentation performance. Combining these two modules, our FAMNet surpasses existing FSMIS models and Cross-domain Few-shot Semantic Segmentation models on three cross-domain datasets, achieving state-of-the-art performance in the CD-FSMIS task. Yuntian Bo, Yazhou Zhu 0001, Lunbo Li, Haofeng Zhang 0001 |
AAAI | 4 |
| 2025 | Concentrate on Weakness: Mining Hard Prototypes for Few-Shot Medical Image SegmentationabstractFew-Shot Medical Image Segmentation (FSMIS) has been widely used to train a model that can perform segmentation from only a few annotated images. However, most existing prototype-based FSMIS methods generate multiple prototypes from the support image solely by random sampling or local averaging, which can cause particularly severe boundary blurring due to the tendency for normal features accounting for the majority of features of a specific category. Consequently, we propose to focus more attention to those weaker features that are crucial for clear segmentation boundary. Specifically, we design a Support Self-Prediction (SSP) module to identify such weak features by comparing true support mask with one predicted by global support prototype. Then, a Hard Prototypes Generation (HPG) module is employed to generate multiple hard prototypes based on these weak features. Subsequently, a Multiple Similarity Maps Fusion (MSMF) module is devised to generate final segmenting mask in a dual-path fashion to mitigate the imbalance between foreground and background in medical images. Furthermore, we introduce a boundary loss to further constraint the edge of segmentation. Extensive experiments on three publicly available medical image datasets demonstrate that our method achieves state-of-the-art performance. Code is available at https://github.com/jcjiang99/CoW. Jianchao Jiang, Haofeng Zhang 0001 |
IJCAI | 2 |
| 2025 | Few-shot Novel Category DiscoveryabstractThe recently proposed Novel Category Discovery (NCD) adapt paradigm of transductive learning hinders its application in more real-world scenarios. In fact, few labeled data in part of new categories can well alleviate this burden, which coincides with the ease that people can label few of new category data. Therefore, this paper presents a new setting in which a trained agent is able to flexibly switch between the tasks of identifying examples of known (labelled) classes and clustering novel (completely unlabeled) classes as the number of query examples increases by leveraging knowledge learned from only a few (handful) support examples. Drawing inspiration from the discovery of novel categories using prior-based clustering algorithms, we introduce a novel framework that further relaxes its assumptions to the real-world open set level by unifying the concept of model adaptability in few-shot learning. We refer to this setting as Few-Shot Novel Category Discovery (FSNCD) and propose Semi-supervised Hierarchical Clustering (SHC) and Uncertainty-aware K-means Clustering (UKC) to examine the model's reasoning capabilities. Extensive experiments and detailed analysis on five commonly used datasets demonstrate that our methods can achieve leading performance levels across different task settings and scenarios. Code is available at: https://github.com/Ashengl/FSNCD. Chunming Li, Haofeng Zhang 0001 |
IJCAI | 3 |
| 2025 | MAUP: Training-Free Multi-center Adaptive Uncertainty-Aware Prompting for Cross-Domain Few-Shot Medical Image Segmentation
Yazhou Zhu 0001, Haofeng Zhang 0001 |
MICCAI (7) | 2 |
| 2025 | Test-time Filtering Boosts Training-free Zero-shot Composed Image RetrievalabstractComposed Image Retrieval (CIR) is an image retrieval task where users provide a reference image along with modification text to retrieve a target image. Zero-shot CIR (ZS-CIR) attracts significant research interest owing to its strong generalization capability and independence from labeled training data. Most ZS-CIR methods employ late fusion and textual inversion, but these approaches fail to precisely convert the reference image and modification text into a target-aligned query. Consequently, the final query may retain redundant information—such as elements present in the reference image but absent in the target image. To address these limitations, a training-free ZS-CIR method called Test-time Filtering (TTF) is proposed in this paper, leveraging Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Specifically, an MLLM generates captions for reference images, which, along with the modification text, are fed into an LLM to produce target image descriptions. This converts the final query into plain text. Using these results, an MLLM filters the selected candidate images to select positive and negative samples. Finally, the similarity scores of candidate images matching positive samples are amplified, while those of negative samples are suppressed. The proposed method is evaluated on three ZS-CIR datasets—CIRR, FashionIQ, and CIRCO—with experimental results demonstrating superior performance over prior ZS-CIR approaches. The source code is available at https://github.com/After-lifes/TTF. Haoyue Chong, Lunbo Li, Haofeng Zhang 0001 |
MMAsia | 3 |
| 2025 | Leveraging Pseudo-triplet and Flexible Prompt for Zero-shot Composed Image RetrievalabstractCompared to supervised composed image retrieval (CIR), which requires a large number of manually labeled triplets (reference image, modification text, target image) for training, Zero-shot Composed Image Retrieval (ZS-CIR) only needs easily available image-caption pairs for training. Prior works for ZS-CIR have primarily used textual inversion to convert images from image-caption pairs into pseudo-words using a pre-trained Vision-language Model (VLM). These pseudo-words are then concatenated to a template, such as “a photo of $”. We find that this fixed template representation hinders the generalization ability of ZS-CIR, and that converting an image into a single pseudo-token fails to accurately represent the information contained in the image. Additionally, given the high performance demonstrated by traditional supervised CIR after triplet training, we aim to achieve a similar structure in ZS-CIR. In this paper, we propose a novel ZS-CIR paradigm based on pseudo-triplets and flexible prompts, rather than fixed templates. Specifically, we fine-tune the vision encoder and split the caption to generate pseudo-triplets using random patch masking. The obtained masked image tokens and corresponding split captions are fed into a Q-former, which is guided by the learned queries and generates queries containing both image and caption information. Finally, the learned queries are concatenated with the split caption and fed into the text encoder to obtain a composed query. Extensive experiments on commonly used ZS-CIR benchmarks such as CIRR, FashionIQ, and CIRCO demonstrate that our method outperforms previous ZS-CIR approaches. The source code is available at https://github.com/After-lifes/PFP. Haoyue Chong, Lunbo Li, Haofeng Zhang 0001 |
MMAsia | 3 |
| 2025 | RobustEMD: Domain robust matching for cross-domain few-shot medical image segmentation
Yazhou Zhu 0001, Minxian Li, Qiaolin Ye, Tong Xin 0002, Haofeng Zhang 0001 |
Artif. Intell. Medicine | 6 |
| 2025 | Synthesizing Spreading-out features for generative zero-shot image classification
Jingren Liu, Zheng Zhang 0006, Yang Long 0001, Wankou Yang, Yunyang Yan, Haofeng Zhang 0001 |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Imbuing, Enrichment and Calibration: Leveraging Language for Unseen Domain Extension
Chenyi Jiang, Jianqin Zhao, Jingjing Deng 0001, Zechao Li, Haofeng Zhang 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | A generic class-agnostic object counting network with adaptive offset deformable convolution
Yuanwu Xu, Haofeng Zhang 0001 |
Neurocomputing | 3 |
| 2025 | CREAM: Few-shot Object Counting with Cross REfinement and Adaptive density Map
Yuanwu Xu, Minxian Li, Qiaolin Ye, Lunbo Li, Haofeng Zhang 0001 |
Image Vis. Comput. | 6 |
| 2025 | Multi-domain feature-enhanced attribute updater for generalized zero-shot learning
Yuyan Shi, Chenyi Jiang, Feifan Song 0004, Qiaolin Ye, Yang Long 0001, Haofeng Zhang 0001 |
Neural Comput. Appl. | 6 |
| 2025 | Imaginary-Connected Embedding in Complex Space for Unseen Attribute-Object DiscriminationabstractCompositional Zero-Shot Learning (CZSL) aims to recognize novel compositions of seen primitives. Prior studies have attempted to either learn primitives individually (non-connected) or establish dependencies among them in the composition (fully-connected). In contrast, human comprehension of composition diverges from the aforementioned methods as humans possess the ability to make composition-aware adaptation for these primitives, instead of inferring them rigidly through the aforementioned methods. However, developing a comprehension of compositions akin to human cognition proves challenging within the confines of real space. This arises from the limitation of real-space-based methods, which often categorize attributes, objects, and compositions using three independent measures, without establishing a direct dynamic connection. To tackle this challenge, we expand the CZSL distance metric scheme to encompass complex spaces to unify the independent measures, and we establish an imaginary-connected embedding in complex space to model human understanding of attributes. To achieve this representation, we introduce an innovative visual bias-based attribute extraction module that selectively extracts attributes based on object prototypes. As a result, we are able to incorporate phase information in training and inference, serving as a metric for attribute-object dependencies while preserving the independent acquisition of primitives. We evaluate the effectiveness of our proposed approach on three benchmark datasets, illustrating its superiority compared to baseline methods. Chenyi Jiang, Yang Long 0001, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Exploring the Effectiveness of Open-Source Donation Platform: An Empirical Study on OpencollectiveabstractABSTRACT In recent years, with the development of the open‐source community, various open‐source donation platforms have emerged. These platforms effectively alleviate the financial pressures faced by open‐source projects through diversified funding sources and flexible donation methods. As one of the most representative open‐source donation platforms, Opencollective has garnered widespread attention from both the open‐source community and academia. Although Opencollective claims to provide more funding opportunities for open‐source projects, the extent to which it effectively addresses the financial challenges faced by these projects remains unclear. While there have been studies on the effectiveness of traditional donation models, research on the effectiveness of emerging donation platforms such as Opencollective is still limited. Given that a large number of open‐source projects are urgently seeking donations, understanding the effectiveness of donations through Opencollective is crucial for these projects. To address this gap, we have made an early step in this direction. This paper conducts a comprehensive study on the effectiveness of donations through the Opencollective, employing a combination of quantitative and qualitative analysis and identifies the following key findings: (1) Opencollective attracts a diverse group of participants, including individual donors, sponsors, contributors, and project managers, with individual donors constituting the largest group. Most donations are concentrated in the range of $5 to $10, indicating that the platform largely relies on small but frequent donations from individuals. (2) Only about 26.61% of open‐source projects receive donations through Opencollective, with approximately 64.38% of these projects receiving a total donation amount of less than $50,000. The likelihood of receiving donations increases with project scale, maturity and the number of stars. Among projects that have received donations, larger projects with stronger social media promotion, greater attention and more issues are more likely to receive additional donations. (3) The positive impact of donations on project development and spend activities is significant only in the short term, with no notable long‐term effects. In contrast, donations do not have a significant short‐term impact on community engagement. Although the long‐term effect is slightly positive, it is not statistically significant. (4) The main shortcomings of Opencollective include insufficient project management and collaboration features, inadequate user experience and interface design, high transaction fees, and a lack of transparency in fund allocation and usage. Our findings provide significant theoretical support and practical recommendations for the effectiveness of emerging donation platforms and the sustainable development of open‐source projects. Shuoxiao Zhang, Enyi Tang, Zhekai Zhang, Yixiao Shan, Haofeng Zhang 0001, Xuandong Li |
J. Softw. Evol. Process. | 6 |
| 2025 | Cross-Domain Few-Shot Medical Image Segmentation via Dynamic Semantic MatchingabstractCross-domain few-shot medical image segmentation (CDFSMIS) presents the fundamental challenge of segmenting novel anatomical or tissue structures on unfamiliar medical imaging domains with limited annotated data. In this paper, we conduct an in-depth investigation of CDFSMIS and identify two critical observations: 1) the conventional matching mechanisms from existing few-shot models are particularly vulnerable to discrepancies in local characteristics between different domains and 2) the semantic representations learned from source domains often lack robustness when generalizing to unfamiliar target domains. Motivated by these insights, we propose a novel Dynamic Semantic Matching (DSM) framework that addresses these challenges through a three-component approach. First, we design a support-query feature re-weighting (SFR) mechanism that leverages multilevel hidden features to suppress domain-specific contents. Second, we introduce a dynamic semantic information selection (DSIS) strategy that adaptively identifies and combines domain-robust channels to construct generalizable representations. Third, we develop a dual-perspective semantic center calculation method to address the inherent texture imbalance in medical images. Extensive experiments on four unfamiliar target domains (MS-CMR, PI-PMR, Chest-X-Ray and ISIC2018) demonstrate that our approach significantly outperforms state-of-the-art few-shot segmentation and cross-domain few-shot segmentation models, validating the effectiveness of DSM in simultaneously addressing domain generalization and semantic matching challenges in medical image segmentation. The source code is available at https://github.com/YazhouZhu19/DSM. Yazhou Zhu 0001, Tao Zhou 0002, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Few-Shot Medical Image Segmentation With High-Confidence Prior MaskabstractLabeling large amounts of medical data is travailing, leading to the blooming of few-shot medical image segmentation, which aims to segment the foreground of a query image given a labeled support set. Almost all current models adopt the cosine distance to measure the similarity between prototypes and query features. However, the limitation of the cosine distance is exacerbated by intra-class differences and inter-class imbalances in medical image scenarios, where angle-only evaluation can induce misclassification to under- and over-segmentation. Motivated by this, we propose a High-Confidence Prior Mask-guided Network (HCPMNet), comprising a High-Confidence Mask Generator (HCPMG), a Target Region Mining (TRM) module, and a Prototype-Oriented Expansion Match (POEM) module. Our HCPMNet offers key advantages: 1) HCPMG is the first to combinatively evaluate angle and magnitude similarity, generating high-confidence priori masks that accurately and completely localize target regions. 2) TRM mines and aggregates target class information under the guidance of priori masks. 3) POEM, based on both similarity metrics, correctly matches prototypes with query features. Extensive experiments on three general medical datasets show that our HCPMNet achieves a new SoTA with great superiority. Ziming Cheng, Jianqin Zhao, Jingjing Deng 0001, Haofeng Zhang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Dual Interspersion and Flexible Deployment for Few-Shot Medical Image SegmentationabstractAcquiring a large volume of annotated medical data is impractical due to time, financial, and legal constraints. Consequently, few-shot medical image segmentation is increasingly emerging as a prominent research direction. Nowadays, Medical scenarios pose two major challenges: 1) intra-class variation caused by diversity among support and query sets; 2) inter-class extreme imbalance resulting from background heterogeneity. However, existing prototypical networks struggle to tackle these obstacles effectively. To this end, we propose a Dual Interspersion and Flexible Deployment (DIFD) model. Drawing inspiration from military interspersion tactics, we design the dual Interspersion module to generate representative basis prototypes from support features. These basis prototypes are then deeply interacted with query features. Furthermore, we introduce a fusion factor to fuse and refine the basis prototypes. Ultimately, we seamlessly integrate and flexibly deploy the basis prototypes to facilitate correct matching between the query features and basis prototypes, thus conducive to improving the segmentation accuracy of the model. Extensive experiments on three publicly available medical image datasets demonstrate that our model significantly outshines other SoTAs (2.78% higher dice score on average across all datasets), achieving a new level of performance. The code is available at: https://github.com/zmcheng9/DIFD. Ziming Cheng, Yang Long 0001, Tao Zhou 0002, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Contextual Interaction via Primitive-based Adversarial Training for Compositional Zero-shot LearningabstractCompositional Zero-shot Learning (CZSL) aims to identify novel compositions via known attribute–object pairs. The primary challenge in CZSL tasks lies in the significant discrepancies introduced by the complex interaction between the visual primitives of attribute and object, consequently decreasing the classification performance toward novel compositions. Previous remarkable works primarily addressed this issue by focusing on disentangling strategy or utilizing object-based conditional probabilities to constrain the selection space of attributes. Unfortunately, few studies have explored the problem from the perspective of modeling the mechanism of visual primitive interactions. Inspired by the success of vanilla adversarial learning in Cross-Domain Few-shot Learning, we take a step further and devise a model-agnostic and Primitive-based Adversarial Training (PBadv) method to deal with this problem. Besides, the latest studies highlight the weakness of the perception of hard compositions even under data-balanced conditions. To this end, we propose a novel over-sampling strategy with object-similarity guidance to augment target compositional training data. We performed detailed quantitative analysis and retrieval experiments on well-established datasets, such as UT-Zappos50K, MIT-States, and C-GQA, to validate the effectiveness of our proposed method, and the State-of-the-Art (SOTA) performance demonstrates the superiority of our approach. The code is available at https://github.com/lisuyi/PBadv_czsl . Suyi Li 0005, Chenyi Jiang, Yang Long 0001, Zheng Zhang 0006, Haofeng Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Revealing the Proximate Long-Tail Distribution in Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to transfer knowledge from seen state-object pairs to novel unseen pairs. In this process, visual bias caused by the diverse interrelationship of state-object combinations blurs their visual features, hindering the learning of distinguishable class prototypes. Prevailing methods concentrate on disentangling states and objects directly from visual features, disregarding potential enhancements that could arise from a data viewpoint. Experimentally, we unveil the results caused by the above problem closely approximate the long-tailed distribution. As a solution, we transform CZSL into a proximate class imbalance problem. We mathematically deduce the role of class prior within the long-tailed distribution in CZSL. Building upon this insight, we incorporate visual bias caused by compositions into the classifier's training and inference by estimating it as a proximate class prior. This enhancement encourages the classifier to acquire more discernible class prototypes for each composition, thereby achieving more balanced predictions. Experimental results demonstrate that our approach elevates the model's performance to the state-of-the-art level, without introducing additional parameters. Chenyi Jiang, Haofeng Zhang 0001 |
AAAI | 2 |
| 2024 | Do They Share the Same Tail? Learning Individual Compositional Attribute Prototype for Generalized Zero-Shot Learning
Yuyan Shi, Chenyi Jiang, Run Shi, Haofeng Zhang 0001 |
ACCV (3) | 4 |
| 2024 | MTPNet: Learning Multiple Twin-support Prototypes for Few-shot Medical Image SegmentationabstractGiven the high annotation costs and ethical considerations associated with medical images, leveraging a limited number of annotated samples for Few-Shot Medical Image Segmentation (FSMIS) has become increasingly prevalent. However, existing models tend to focus on visible foreground support information, often overlooking extreme foreground-background imbalances. In addition, query images sometimes have slight different appearance compared to support images of the same category due to the differences in size as well as slicing angle, thus employing only support images to generate prototypes inevitably leads to matching bias. To address these challenges, we present an innovative approach through learning a Multiple Twin-support Prototypes Network (MTPNet). Our approach includes the design of the Scale Consistent Sampling (SCS) module, which adaptively adjusts the foreground and background points within the support set, thereby balancing the influence of various structural elements in the image. Additionally, the Twin-support Prototypes Extraction (TPE) module facilitates the critical interaction between query and support features to extract twin-support prototypes. This module incorporates a Backtrace Interaction Filter (BIF) to eliminate erroneous interaction prototypes. Extensive experimental validation on three widely used medical image datasets demonstrates that our method surpasses current State-of-the-arts, showcasing its potential to address key limitations in FSMIS. The code is available at https://github.com/FeifanSong/MTPNet. Feifan Song 0004, Ziming Cheng, Lunbo Li, Haofeng Zhang 0001 |
BIBM | 4 |
| 2024 | Frequency-aware Adaptive Filtering Network for Few-Shot Medical Image SegmentationabstractManual annotation of massive data in the biomedical field necessitates huge costs, expertise, and privacy considerations. However, Few-Shot Medical Image Segmentation (FSMIS) offers the possibility of learning a segmentation model with excellent performance from limited medical data. In this paper, we propose a Frequency-aware Adaptive Filtering Network (FAF-Net) for FSMIS, which is the first to calibrate the correlations between support prototype and query features from a frequency perspective, addressing the problems of FSMIS. The FAF-Net mines the commonalities between the support prototype and query features to reduce the intra-class differences through an Intra-class Commonality Miner (ICM), and introduces an Adaptive Fourier Filter (AFF) to adaptively filter the spectrum of the attention map, thus balancing the foreground and background classes. Additionally, an Adaptive Threshold Learner (ATL) is incorporated to learn an adaptive threshold for each spatial location of the predicted mask from the support features, overcoming the limitations of single thresholding. Extensive experiments on four generic medical datasets showcase that our model significantly outperforms SoTA methods (exceeding SoTA by 2.13% on average). Ziming Cheng, Tong Xin 0002, Haofeng Zhang 0001 |
BIBM | 4 |
| 2024 | Evolutionary Generalized Zero-Shot Learning
Dubing Chen, Chenyi Jiang, Haofeng Zhang 0001 |
IJCAI | 3 |
| 2024 | Learning Spatial Similarity Distribution for Few-shot Object Counting
Yuanwu Xu, Feifan Song 0004, Haofeng Zhang 0001 |
IJCAI | 3 |
| 2024 | Fusing spatial and frequency features for compositional zero-shot image classification
Suyi Li 0005, Chenyi Jiang, Qiaolin Ye, Wankou Yang, Haofeng Zhang 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Learning to segment complex vessel-like structures with spectral transformer
Huajun Liu, Jing Yang 0053, Hui Kong 0001, Haofeng Zhang 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Estimation of Near-Instance-Level Attribute Bottleneck for Zero-Shot Learning
Chenyi Jiang, Yuming Shen, Dubing Chen, Haofeng Zhang 0001, Ling Shao 0001, Philip Torr 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Extrinsic Calibration of Camera and LiDAR Systems With Three-Dimensional Towered CheckerboardsabstractWith the increasing utilization of cameras and three‐dimensional Light Detection and Ranging (LiDAR) systems in perception tasks, the fusion of these two sensor modalities has emerged as a prominent research focus in the fields of robotics and unmanned systems. While various extrinsic calibration methods have been developed, they often suffer from limited accuracy when using low‐resolution LiDAR sensors and require the placement of calibration targets at multiple locations. This paper introduces a novel calibration target known as the Three‐Dimensional Towered Checkerboard (3TC), along with a precise and straightforward extrinsic calibration approach for camera‐LiDAR systems. The 3TC consists of stacked cubes adorned with planar or 2D checkerboards, which provide the known positions of checkerboard corner points in three‐dimensional space. Leveraging the Iterative Closest Point (ICP) algorithm, the proposed method calculates the spatial relationship between LiDAR point cloud data and the 3TC model to infer the positions of checkerboard corner points in the LiDAR coordinate system. Subsequently, the Perspective‐n‐Point (PnP) algorithm is employed to establish the correlation between corner positions in the LiDAR coordinate system and the camera image, given the intrinsic parameters of the camera. By ensuring an adequate number of cubes and 2D checkerboards on a specific 3TC, along with accurately estimated corner point positions in LiDAR, a single frame of data from both the camera and LiDAR facilitates their extrinsic calibration. Experimental validations conducted across diverse camera and LiDAR systems, achieving minimal error close to the theoretical limit of the devices, attest to the robustness and precision of the 3TC and the proposed calibration methodology. Dexin Ren, Mingwu Ren, Haofeng Zhang 0001 |
Int. J. Intell. Syst. | 3 |
| 2024 | Unsupervised cross domain semantic segmentation with mutual refinement and information distillation
Dexin Ren, Zheng Zhang 0006, Wankou Yang, Mingwu Ren, Haofeng Zhang 0001 |
Neurocomputing | 6 |
| 2024 | SAFENet: Semantic-Aware Feature Enhancement Network for unsupervised cross-domain road scene segmentation
Dexin Ren, Minxian Li, Mingwu Ren, Haofeng Zhang 0001 |
Image Vis. Comput. | 5 |
| 2024 | Mutual Balancing in State-Object Components for Compositional Zero-Shot Learning
Chenyi Jiang, Qiaolin Ye, Yuming Shen, Zheng Zhang 0006, Haofeng Zhang 0001 |
Pattern Recognit. | 6 |
| 2024 | Learning De-biased prototypes for Few-shot Medical Image Segmentation
Yazhou Zhu 0001, Ziming Cheng, Haofeng Zhang 0001 |
Pattern Recognit. Lett. | 4 |
| 2024 | Few-Shot Medical Image Segmentation via Generating Multiple Representative DescriptorsabstractAutomatic medical image segmentation has witnessed significant development with the success of large models on massive datasets. However, acquiring and annotating vast medical image datasets often proves to be impractical due to the time consumption, specialized expertise requirements, and compliance with patient privacy standards, etc. As a result, Few-shot Medical Image Segmentation (FSMIS) has become an increasingly compelling research direction. Conventional FSMIS methods usually learn prototypes from support images and apply nearest-neighbor searching to segment the query images. However, only a single prototype cannot well represent the distribution of each class, thus leading to restricted performance. To address this problem, we propose to Generate Multiple Representative Descriptors (GMRD), which can comprehensively represent the commonality within the corresponding class distribution. In addition, we design a Multiple Affinity Maps based Prediction (MAMP) module to fuse the multiple affinity maps generated by the aforementioned descriptors. Furthermore, to address intra-class variation and enhance the representativeness of descriptors, we introduce two novel losses. Notably, our model is structured as a dual-path design to achieve a balance between foreground and background differences in medical images. Extensive experiments on four publicly available medical image datasets demonstrate that our method outperforms the state-of-the-art methods, and the detailed analysis also verifies the effectiveness of our designed module. Ziming Cheng, Tong Xin 0002, Tao Zhou 0002, Haofeng Zhang 0001, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Deconstructed Generation-Based Zero-Shot ModelabstractRecent research on Generalized Zero-Shot Learning (GZSL) has focused primarily on generation-based methods. However, current literature has overlooked the fundamental principles of these methods and has made limited progress in a complex manner. In this paper, we aim to deconstruct the generator-classifier framework and provide guidance for its improvement and extension. We begin by breaking down the generator-learned unseen class distribution into class-level and instance-level distributions. Through our analysis of the role of these two types of distributions in solving the GZSL problem, we generalize the focus of the generation-based approach, emphasizing the importance of (i) attribute generalization in generator learning and (ii) independent classifier learning with partially biased data. We present a simple method based on this analysis that outperforms SotAs on four public GZSL datasets, demonstrating the validity of our deconstruction. Furthermore, our proposed method remains effective even without a generative model, representing a step towards simplifying the generator-classifier structure. Our code is available at https://github.com/cdb342/DGZ. Dubing Chen, Yuming Shen, Haofeng Zhang 0001, Philip Torr 0001 |
AAAI | 3 |
| 2023 | Few-Shot Medical Image Segmentation via a Region-Enhanced Prototypical Transformer
Yazhou Zhu 0001, Tong Xin 0002, Haofeng Zhang 0001 |
MICCAI (4) | 4 |
| 2023 | Hashing One With AllabstractThe recent trend in unsupervised hashing requires not only a discrete representation space but also the ability to mine the similarities between data points. Determining and maintaining the relations between each datum and all the others maximally utilize the semantic diversity of the training set, but can we explore this on the holistic dataset using a single network end-to-end? In this paper, we take a step towards this vision by proposing Overview Hashing (OH). OH unifies the two ultimate goals of unsupervised hashing, i.e., (1) encoding compact features and (2) learning data similarities on a large scale end-to-end, into one model. In particular, we split the top of an encoder into a binary hash head and a continuous one. For an arbitrary datum, its similarities to all the others in the dataset are reflected in the Hamming distances of their hash heads. The distances then act as the weights to aggregate the continuous heads, shaping the final representation of this datum for loss computation. Hence, training with this representation simultaneously tunes the similarities of this datum to the whole dataset. In the context of a contrastive learning framework, we theoretically endorse our design by linking it to knowledge distillation and the attention mechanisms. Our experiments on the benchmarked datasets show the superiority of OH over the state-of-the-art hashing methods. Code is available at \hrefhttps://github.com/RosieYuu/OH \textcolorred https://github.com/RosieYuu/OH. Jiaguo Yu, Yuming Shen, Haofeng Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | Multi-Scale Superpoint Network for 3D Point Cloud Semantic Segmentationabstract3D point cloud semantic segmentation is a fundamental task for 3D scene understanding. However, most existing pipelines usually use k-NN or ball query operation to form hard neighborhoods, which may cross different semantic objects, resulting low-quality local features. To address this issue, we propose a multi-scale superpoint network that gradually generates multi-scale soft neighborhoods to extract geometric local features, thereby boosting the 3D semantic segmentation performance. Specifically, we present a simple yet efficient superpoint merging module that merge small-scale superpoints to obtain large-scale superpoint by considering the feature similarity of superpoints, so that we can obtain multi-scale geometric features of point clouds. We also develop a superpoint upsampling module that adopt inverse mapping function to propagate multi-scale features from low-resolution point cloud to high-resolution point cloud. By integrating our multi-scale superpoint network into a simple point based semantic segmentation network, our method can obtain SOTA results on S3DIS Area 5 and 6-fold, and competitive results on ScanNet v2. Ft Zheng, Le Hui, Jin Xie 0001, Haofeng Zhang 0001 |
MMAsia | 4 |
| 2023 | Data driven recurrent generative adversarial network for generalized zero shot image classification
Jie Zhang 0005, Shengbin Liao, Haofeng Zhang 0001, Yang Long 0001, Zheng Zhang 0006, Li Liu 0004 |
Inf. Sci. | 3 |
| 2023 | Combining Pixel-Level and Structure-Level Adaptation for Semantic Segmentation
Xiwen Bi, Dubing Chen, Haofeng Zhang 0001 |
Neural Process. Lett. | 5 |
| 2023 | Dynamic Unary Convolution in TransformersabstractIt is uncertain whether the power of transformer architectures can complement existing convolutional neural networks. A few recent attempts have combined convolution with transformer design through a range of structures in series, where the main contribution of this paper is to explore a parallel design approach. While previous transformed-based approaches need to segment the image into patch-wise tokens, we observe that the multi-head self-attention conducted on convolutional features is mainly sensitive to global correlations and that the performance degrades when these correlations are not exhibited. We propose two parallel modules along with multi-head self-attention to enhance the transformer. For local information, a dynamic local enhancement module leverages convolution to dynamically and explicitly enhance positive local patches and suppress the response to less informative ones. For mid-level structure, a novel unary co-occurrence excitation module utilizes convolution to actively search the local co-occurrence between patches. The parallel-designed Dynamic Unary Convolution in Transformer (DUCT) blocks are aggregated into a deep architecture, which is comprehensively evaluated across essential computer vision tasks in image-based classification, segmentation, retrieval and density estimation. Both qualitative and quantitative results show our parallel convolutional-transformer approach with dynamic and unary convolution outperforms existing series-designed structures. Haoran Duan 0001, Yang Long 0001, Haofeng Zhang 0001, Chris G. Willcocks, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Exploiting spatial relationships for visual tracking
Lunbo Li, Jianhui Guo, Haofeng Zhang 0001 |
Pattern Recognit. Lett. | 5 |
| 2023 | Generating diverse augmented attributes for generalized zero shot learning
Yuming Shen, Haofeng Zhang 0001 |
Pattern Recognit. Lett. | 4 |
| 2022 | Boosting Generative Zero-Shot Learning by Synthesizing Diverse Features with Attribute AugmentationabstractThe recent advance in deep generative models outlines a promising perspective in the realm of Zero-Shot Learning (ZSL). Most generative ZSL methods use category semantic attributes plus a Gaussian noise to generate visual features. After generating unseen samples, this family of approaches effectively transforms the ZSL problem into a supervised classification scheme. However, the existing models use a single semantic attribute, which contains the complete attribute information of the category. The generated data also carry the complete attribute information, but in reality, visual samples usually have limited attributes. Therefore, the generated data from attribute could have incomplete semantics. Based on this fact, we propose a novel framework to boost ZSL by synthesizing diverse features. This method uses augmented semantic attributes to train the generative model, so as to simulate the real distribution of visual features. We evaluate the proposed model on four benchmark datasets, observing significant performance improvement against the state-of-the-art. Yuming Shen, Haofeng Zhang 0001 |
AAAI | 4 |
| 2022 | Weighted Contrastive Hashing
Jiaguo Yu, Huming Qiu, Dubing Chen, Haofeng Zhang 0001 |
ACCV (5) | 4 |
| 2022 | Learning Internal Semantics with Expanded Categories for Generative Zero-Shot Learning
Haofeng Zhang 0001 |
ACCV (7) | 3 |
| 2022 | From Sparse to Dense: Semantic Graph Evolutionary Hashing for Unsupervised Cross-Modal Retrieval
Jiaguo Yu, Shengbin Liao, Zheng Zhang 0006, Haofeng Zhang 0001 |
ACCV (4) | 5 |
| 2022 | Class Concentration with Twin Variational Autoencoders for Unsupervised Cross-Modal Hashing
Yazhou Zhu 0001, Shengbin Liao, Qiaolin Ye, Haofeng Zhang 0001 |
ACCV (6) | 5 |
| 2022 | Zero-Shot Logit AdjustmentabstractSemantic-descriptor-based Generalized Zero-Shot Learning (GZSL) poses challenges in recognizing novel classes in the test phase. The development of generative models enables current GZSL techniques to probe further into the semantic-visual link, culminating in a two-stage form that includes a generator and a classifier. However, existing generation-based methods focus on enhancing the generator's effect while neglecting the improvement of the classifier. In this paper, we first analyze of two properties of the generated pseudo unseen samples: bias and homogeneity. Then, we perform variational Bayesian inference to back-derive the evaluation metrics, which reflects the balance of the seen and unseen classes. As a consequence of our derivation, the aforementioned two properties are incorporated into the classifier training as seen-unseen priors via logit adjustment. The Zero-Shot Logit Adjustment further puts semantic-based classifiers into effect in generation-based GZSL. Our experiments demonstrate that the proposed technique achieves state-of-the-art when combined with the basic generator, and it can improve various generative Zero-Shot Learning frameworks. Our codes are available on https://github.com/cdb342/IJCAI-2022-ZLA. Dubing Chen, Yuming Shen, Haofeng Zhang 0001, Philip Torr 0001 |
IJCAI | 3 |
| 2022 | Learning to Hash Naturally SortsabstractLearning to hash pictures a list-wise sorting problem. Its testing metrics, e.g., mean-average precision, count on a sorted candidate list ordered by pair-wise code similarity. However, scarcely does one train a deep hashing model with the sorted results end-to-end because of the non-differentiable nature of the sorting operation. This inconsistency in the objectives of training and test may lead to sub-optimal performance since the training loss often fails to reflect the actual retrieval metric. In this paper, we tackle this problem by introducing Naturally-Sorted Hashing (NSH). We sort the Hamming distances of samples' hash codes and accordingly gather their latent representations for self-supervised training. Thanks to the recent advances in differentiable sorting approximations, the hash head receives gradients from the sorter so that the hash encoder can be optimized along with the training procedure. Additionally, we describe a novel Sorted Noise-Contrastive Estimation (SortedNCE) loss that selectively picks positive and negative samples for contrastive learning, which allows NSH to mine data semantic relations during training in an unsupervised manner. Our extensive experiments show the proposed NSH model significantly outperforms the existing unsupervised hashing methods on three benchmarked datasets. Jiaguo Yu, Yuming Shen, Haofeng Zhang 0001, Philip Torr 0001 |
IJCAI | 4 |
| 2022 | Entropy-weighted reconstruction adversary and curriculum pseudo labeling for domain adaptation in semantic segmentation
Xiwen Bi, Xiaohong Zhang 0009, Haofeng Zhang 0001 |
Neurocomputing | 4 |
| 2022 | Semi-supervised cross-modal hashing with multi-view graph representation
Haofeng Zhang 0001, Lunbo Li, Wankou Yang, Li Liu 0004 |
Inf. Sci. | 2 |
| 2022 | Learning discriminative and representative feature with cascade GAN for generalized zero-shot learning
Jingren Liu, Liyong Fu, Haofeng Zhang 0001, Qiaolin Ye, Wankou Yang, Li Liu 0004 |
Knowl. Based Syst. | 3 |
| 2022 | Edge detection with attention: From global view to local focus
Huajun Liu, Zuyuan Yang, Haofeng Zhang 0001, Cailing Wang |
Pattern Recognit. Lett. | 3 |
| 2022 | From Less to More: Progressive Generalized Zero-Shot Detection With Curriculum LearningabstractObject detection, as one of the most important environment perception tasks for traffic safety in intelligent transportation systems, has been widely investigated recently. However, most of the researches focus on the fully supervised scenario, and inevitably lead to model failure. With the continuous development of Zero-Shot Learning (ZSL) models, Generalized Zero-Shot Detection (GZSD) has attracted great attention due to its ability of detecting unseen objects. Many researchers tend to map the detected visual features to semantic attributes and then separate seen and unseen domains during inference. But they have ignore that the generative methods generally have higher performance than these visual-semantic mapping methods, and they have been confirmed from previous GZSL methods. In order to make up for the vacancy of GZSD in the generative methods, we propose an idea of using curriculum learning to generate more precise unseen visual features. And with the excellent performance of WGAN-based method in sample synthesis, we realize the function of using semantics to generate visual features for unseen domains. In addition, we also adopt part of the idea of meta-learning to progressively correct the capability of the generator for better mitigating domain shift problem during the generation process. Through the above ideas, we can detect both seen and unseen bounding boxes and classify them accurately, by combining with the excellent detection ability of Faster-RCNN. Extensive experimental results on two popular datasets, i.e., MSCOCO and KITTI, show that our proposed method can outperform the state-of-the-art methods. Jingren Liu, Yi Chen 0023, Huajun Liu, Haofeng Zhang 0001, Yudong Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Confidence-and-Refinement Adaptation Model for Cross-Domain Semantic SegmentationabstractWith the rapid development of convolutional neural networks (CNNs), significant progress has been achieved in semantic segmentation. Despite the great success, such deep learning approaches require large scale real-world datasets with pixel-level annotations. However, considering that pixel-level labeling of semantics is extremely laborious, many researchers turn to utilize synthetic data with free annotations. But due to the clear domain gap, the segmentation model trained with the synthetic images tends to perform poorly on the real-world datasets. Unsupervised domain adaptation (UDA) for semantic segmentation recently gains an increasing research attention, which aims at alleviating the domain discrepancy. Existing methods in this scope either simply align features or the outputs across the source and target domains or have to deal with the complex image processing and post-processing problems. In this work, we propose a novel multi-level UDA model named Confidence-and-Refinement Adaptation Model (CRAM), which contains a confidence-aware entropy alignment (CEA) module and a style feature alignment (SFA) module. Through CEA, the adaptation is done locally via adversarial learning in the output space, making the segmentation model pay attention to the high-confident predictions. Furthermore, to enhance the model transfer in the shallow feature space, the SFA module is applied to minimize the appearance gap across domains. Experiments on two challenging UDA benchmarks “GTA5-to-Cityscapes” and “SYNTHIA-to-Cityscapes” demonstrate the effectiveness of CRAM. We achieve comparable performance with the existing state-of-the-art works with advantages in simplicity and convergence speed. Xiaohong Zhang 0009, Yi Chen 0023, Ziyi Shen, Yuming Shen, Haofeng Zhang 0001, Yudong Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Exploiting Web Images for Fine-Grained Visual Recognition via Dynamic Loss Correction and Global Sample SelectionabstractTo distinguish subtle differences among fine-grained categories, a large amount of well-labeled images are typically required. However, acquiring manual annotations for fine-grained categories is an extremely difficult task as it usually has a high demand for professional knowledge. To this end, directly leveraging web images for learning fine-grained models becomes a natural choice. Nevertheless, due to the existence of label noise, this learning paradigm tends to have a poor performance. In this work, we propose an end-to-end approach by combining dynamic loss correction and global sample selection to alleviate the problem of label noise. Specifically, we leverage the network to predict all samples, record the predictions of recent several epochs, and calculate the uncertainly-based dynamic loss for global sample selection. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed approach. The source code of our approach has been released on the website:https://github.com/NUST-Machine-Intelligence-Laboratory/dlc. Huafeng Liu 0004, Haofeng Zhang 0001, Jianfeng Lu 0003, Zhenmin Tang |
IEEE Trans. Multim. | 2 |
| 2021 | Dual Prototype Relaxation for Generalized Zero Shot LearningabstractGeneralized Zero Shot Learning (GZSL) is proposed to solve the training data missing problem by transferring the knowledge learned in seen classes to unseen classes. Many methods project the visual features into semantic space and find their nearest neighbours among the pre-defined attributes, which has achieved significant success. However, there are two problems involved in this type of methods, one is that the projection is a many-to-one mapping, which cannot maintain the diversity of features in semantic space, the other is that searching within all classes cannot well utilize the knowledge learned in the seen classes. In this paper, we propose a novel method named Dual Prototype Relaxation (DPR) by relaxing the projection from many-to-one to many-to-many. Specifically, we add noise to the semantic prototype in response to the projection of multiple features within a class, and reconstruct the visual features with the same relaxed prototype. Besides, in order to make better use of the knowledge learned in the seen classes, an Out-of-Domain (OoD) based method is employed to first classify the feature to seen or unseen domains, and then the same DPR model is applied to recognize its category within each domain. Extensive experiments on four popular datasets are conducted and the results show that our method can outperform many linear and deep state-of-the-art methods although our method is a linear one. Jie Zhang 0005, Haofeng Zhang 0001, Bingzhang Hu |
ICME | 2 |
| 2021 | Near-Real Feature Generative Network for Generalized Zero-Shot LearningabstractDue to the powerful feature synthesis ability, Generative Adversarial Networks (GAN) is well adapted to the Generalized Zero-Shot Learning (GZSL) task and has achieved great success. Most GAN models for GZSL usually employ random noise with normal distribution to synthesize unseen samples. However, the generated samples often have the same normal distribution as the input noise, which is unrealistic in most circumstances. Therefore, in this paper, we consider that the distribution of unseen classes should be follow that of seen classes and propose a near-real feature generative network (NereNet), which utilizes the most semantically similar seen samples to generate the noise for the unseen classes. Specifically, we first calculate the most similar seen classes for the unseen classes, and then train an encoder network to generate the corresponding noise, which is subsequently combined with the unseen classes attributes to generate unseen samples with GAN. Extensive experiments are conducted on four datasets, and the results demonstrate the effectiveness of our proposed method. Jingren Liu, Haoyue Bai 0003, Haofeng Zhang 0001, Li Liu 0004 |
ICME | 3 |
| 2021 | Attention-Guided Semantic Hashing for Unsupervised Cross-Modal RetrievalabstractRecently, due to the low storage consumption and high search efficiency of hashing methods and the powerful feature extraction capability of deep neural networks, deep cross-modal hashing has received extensive attention in the field of multi-media retrieval. However, existing methods tend to ignore the latent relationships between heterogeneous data when learning a common semantic subspace, and cannot retain more important semantic information when mining deep correlations. In this paper, an attention mechanism which focuses on the characteristics of the associated features is employed to propose an attention-aware semantic fusion matrix that integrates important information from different modalities. We introduce a novel network that can pass the extracted features through the attention module to efficiently encode rich and relevant features, and can also generate hash codes under the self-supervision of the proposed attention-aware semantic fusion matrix. Our experimental results and detailed analysis prove that our method can achieve better retrieval performance on the three popular datasets, compared with the recent unsupervised cross-modal hashing methods. Haofeng Zhang 0001, Lunbo Li, Li Liu 0004 |
ICME | 2 |
| 2021 | Learning Homogeneous and Heterogeneous Co-Occurrences for Unsupervised Cross-Modal RetrievalabstractImage-text retrieval, which focuses on unifying both visual and textual representations, is one of the major tasks of cross-modal information processing. With attention mechanism, previous methods performing well take advantage of not only the correspondence in image-text level but also the semantic alignment between the regions in images and corresponding words. However, few of them comprehend the importance of combing the semantic relationship between multimodalities and semantic correspondences in one modality at the same time. Inspired by the heterogeneous information learning of heterogeneous graph network, we propose a novel method called Homogeneous and Heterogeneous Co-Occurrences (H2CO) which mainly consists of two modules to achieve modal co-occurrence in a query of corporations. Specifically, Homogeneous Co-Occurrence Module captures correspondences with neighbors from single modal of regions and words respectively, while Heterogeneous Co-Occurrence Module aims to learn the relations about neighbors across modalities. Finally, the proposed method can aggregate the neighborhood features from both intra modality and inter modality at the same time, thus performs better on image-text matching for considering much more semantic information. Extensive experimental results on MS COCO and Flickr30K show the superior performance of our proposed modal over the state-of-the-arts. Haofeng Zhang 0001, Bingzhang Hu |
ICME | 3 |
| 2021 | Target-targeted Domain Adaptation for Unsupervised Semantic SegmentationabstractSemantic segmentation has attracted increasing attention due to its important role in self-driving, and it is often realized by supervised learning with large number of well labeled maps. However, the labeled images are hard to be obtained in most circumstances, and the common way for unsupervised semantic segmentation is usually implemented by transferring the knowledge from source supervised domain to target unsupervised domain. Most researches focus on encouraging target predictions to be closer to the source ones through a weight-sharing network, and achieve certain performance. However, these methods often suffer from the domain shift problem that the networks are often trained towards the source domain and lead to performance degradation. In this paper, we propose a target-targeted domain adaptation approach by focusing the training on target domain. Our model consists of two components: the Image-to-image Translation (IIT) module to translate the source image to target domain and the Target-targeted Segmentation Adaptation (TSA) module to focus the semantic segmentation on target domain. The IIT module deals with image space alignment while the TSA module bridges the domain gap at the segmentation map level. In addition, we design a closed-loop learning to promote each other by employing feedback from TSA to IIT. Extensive experiments on GTA5 and SYNTHIA to Cityscapes demonstrate the effectiveness of our method in domain adaptation of unsupervised semantic segmentation. Xiaohong Zhang 0009, Haofeng Zhang 0001, Jianfeng Lu 0003, Ling Shao 0001, Jing-Yu Yang 0001 |
ICRA | 2 |
| 2021 | Multi-Scale Spatial Transformer Network for LiDAR-Camera 3D Object DetectionabstractAccurate 3D object detection has recently aroused interest in the context of emerging autonomous driving technologies. Existing approaches predominantly use LiDAR-Camera fusion method to fulfill this challenging task, while neglecting the fact that LiDAR and camera data are spatially correlated, and cannot well retain the edge information. To solve these problems, in this paper, we propose a novel LiDAR-Camera 3D object detection method, namely the Multi-scale Spatial Transformer Network (MST-Net). The proposed method exploits an innovative spatial alignment scheme based on the projection transformer network (PTN) to mitigate the effects of the perspective view caused by sensors. In the process of generating 3D bounding boxes, the Atrous Spatial Pyramid Pool (ASPP) is applied to spatially aligned fusion features in order to preserve edge information to the greatest extent. Extensive experiments are conducted on the popular dataset KITTI, and the results can demonstrate the superiority of the proposed method. In addition, the effectiveness of these two strategies has been illustrated in ablation studies. Zhifan Wang, Xiaohong Zhang 0009, Tong Xin 0002, Haofeng Zhang 0001, Jianfeng Lu 0003 |
IJCNN | 5 |
| 2021 | Clustering-driven Deep Adversarial Hashing for scalable unsupervised cross-modal retrieval
Haofeng Zhang 0001, Lunbo Li, Zheng Zhang 0006, Debao Chen, Li Liu 0004 |
Neurocomputing | 2 |
| 2021 | Sparse graph based self-supervised hashing for scalable image retrieval
Haofeng Zhang 0001, Zheng Zhang 0006, Li Liu 0004, Ling Shao 0001 |
Inf. Sci. | 2 |
| 2021 | Modality independent adversarial network for generalized zero shot image classification
Haofeng Zhang 0001, Yinduo Wang, Yang Long 0001, Longzhi Yang, Ling Shao 0001 |
Neural Networks | 1 |
| 2021 | A plug-in attribute correction module for generalized zero-shot learning
Haofeng Zhang 0001, Haoyue Bai 0003, Yang Long 0001, Li Liu 0004, Ling Shao 0001 |
Pattern Recognit. | 1 |
| 2021 | Built-in Depth-Semantic Coupled Encoding for Scene Parsing, Vehicle Detection, and Road SegmentationabstractRecent representative scene parsing methods based on Convolutional Neural Networks (CNN) have greatly improved spatial resolution of pixel-wise labelling by exploiting multi-scale features and refined boundaries. However, the vast majority of previous works only utilize the color or textural information of images, without considering the depth information, which is beneficial for semantic reasoning. In this paper, we take advantages of the mutual benefit and strong correlation between depth information and semantic information in scene parsing by introducing the Built-in Depth-Semantic Coupled Encoding (BDSCE) module, which adaptively fuses RGB and depth features, and selectively highlights the depth-discriminative features. The proposed BDSCE module is compatible with existing CNN based methods, and can greatly improve scene parsing performance, particularly in those categories that have clear depth distinction and might be misclassified with RGB-only features. Furthermore, we also extend our proposed module to other urban scene semantic reasoning tasks such as vehicle detection and road segmentation, which are implemented by effectively learning and exploiting the encoded depth-semantics and transferring the learned representations with fine-tuning. The extensive experiments on the popular datasets Cityscapes and KITTI demonstrate that our method performs quite well and can significantly improve the state-of-the-art methods. Haofeng Zhang 0001, Ling Shao 0001, Jing-Yu Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | When Visual Disparity Generation Meets Semantic Segmentation: A Mutual Encouragement ApproachabstractSemantic segmentation and depth estimation play important roles in the field of autonomous driving. In recent years, the advantages of Convolutional Neural Networks (CNNs) have allowed these two topics to flourish. However, people always solve these two tasks separately and rarely solve them in a united model. In this paper, we propose a Mutual Encouragement Network (MENet), which includes a semantic segmentation branch and a disparity regression branch, and simultaneously generates semantic map and visual disparity. In the cost volume construction phase, the depth information is embedded in the semantic segmentation branch to increase contextual understanding. Similarly, the semantic information is also included in the disparity regression branch to generate more accurate disparity. Two branches mutually promote each other during training phase and inference phase. We conducted our method on the popular dataset KITTI, and the experimental results show that our method can outperform the state-of-the-art methods on both visual disparity generation and semantic segmentation. In addition, extensive ablation studies also demonstrate that the two tasks in our method can facilitate each other significantly with the proposed approach. Xiaohong Zhang 0009, Yi Chen 0023, Haofeng Zhang 0001, Shuihua Wang, Jianfeng Lu 0003, Jing-Yu Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Deep Unsupervised Self-Evolutionary Hashing for Image RetrievalabstractHashing methods have proven to be effective in the field of large-scale image retrieval. In recent years, the performance of hashing algorithms based on deep learning has greatly exceeded that of non-deep methods. However, most of the outstanding hashing methods are supervised models that heavily rely on annotated labels. In order to circumvent the huge overhead of labeling large-scale datasets, some unsupervised hashing algorithms have been proposed, such as pseudo labels and pseudo pairs. Since the image labels are strictly unavailable, some hyper-parameters in these methods are difficult to be selected, e.g., the final result is very sensitive to the picked number of categories or the chosen threshold of similarity for pairs. In addition, the calculation of pseudo-labels in high-dimensional space is not only computationally complex, but also has low precision. Therefore, in order to alleviate these issues in this paper, we propose a simple but effective Deep Unsupervised Self-evolutionary Hashing (DUSH) algorithm, which utilizes a curriculum learning strategy to iteratively select pseudo pairs from easy to hard in low dimensional Hamming space. Extensive experiments are conducted on four popular datasets, including two single-label datasets and two multi-label datasets, and the results show that our method can significantly outperform the state-of-the-art methods. Haofeng Zhang 0001, Yazhou Yao, Zheng Zhang 0006, Li Liu 0004, Jian Zhang 0002, Ling Shao 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Set and Rebase: Determining the Semantic Graph Connectivity for Unsupervised Cross-Modal HashingabstractThe label-free nature of unsupervised cross-modal hashing hinders models from exploiting the exact semantic data similarity. Existing research typically simulates the semantics by a heuristic geometric prior in the original feature space. However, this introduces heavy bias into the model as the original features are not fully representing the underlying multi-view data relations. To address the problem above, in this paper, we propose a novel unsupervised hashing method called Semantic-Rebased Cross-modal Hashing (SRCH). A novel ‘Set-and-Rebase’ process is defined to initialize and update the cross-modal similarity graph of training data. In particular, we set the graph according to the intra-modal feature geometric basis and then alternately rebase it to update the edges within according to the hashing results. We develop an alternating optimization routine to rebase the graph and train the hashing auto-encoders with closed-form solutions so that the overall framework is efficiently trained. Our experimental results on benchmarked datasets demonstrate the superiority of our model against state-of-the-art algorithms. Yuming Shen, Haofeng Zhang 0001, Yazhou Yao, Li Liu 0004 |
IJCAI | 3 |
| 2020 | Semantic combined network for zero-shot scene parsingabstractRecently, image‐based scene parsing has attracted increasing attention due to its wide application. However, conventional models can only be valid on images with the same domain of the training set and are typically trained using discrete and meaningless labels. Inspired by the traditional zero‐shot learning methods which employ auxiliary side information to bridge the source and target domains, the authors propose a novel framework called semantic combined network (SCN), which aims at learning a scene parsing model only from the images of the seen classes while targeting on the unseen ones. In addition, with the assistance of semantic embeddings of classes, the proposed SCN can further improve the performances of traditional fully supervised scene parsing methods. Extensive experiments are conducted on the data set Cityscapes, and the results show that the proposed SCN can perform well on both zero‐shot scene parsing (ZSSP) and generalised ZSSP settings based on several state‐of‐the‐art scenes parsing architectures. Furthermore, the authors test the proposed model under the traditional fully supervised setting and the results show that the proposed SCN can also significantly improve the performances of the original network models. Yinduo Wang, Haofeng Zhang 0001, Yang Long 0001, Longzhi Yang |
IET Image Process. | 2 |
| 2020 | Semantic-rebased cross-modal hashing for scalable unsupervised text-visual retrieval
Yuming Shen, Haofeng Zhang 0001, Li Liu 0004 |
Inf. Process. Manag. | 3 |
| 2020 | Learning discriminative domain-invariant prototypes for generalized zero shot learning
Yinduo Wang, Haofeng Zhang 0001, Zheng Zhang 0006, Yang Long 0001, Ling Shao 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Unsupervised deep triplet hashing with pseudo triplets for scalable image retrieval
Haofeng Zhang 0001, Zheng Zhang 0006, Qiaolin Ye |
Multim. Tools Appl. | 2 |
| 2020 | Asymmetric graph based zero shot learning
Yinduo Wang, Haofeng Zhang 0001, Zheng Zhang 0006, Yang Long 0001 |
Multim. Tools Appl. | 2 |
| 2020 | Deep transductive network for generalized zero shot learning
Haofeng Zhang 0001, Li Liu 0004, Yang Long 0001, Zheng Zhang 0006, Ling Shao 0001 |
Pattern Recognit. | 1 |
| 2020 | Pseudo distribution on unseen classes for generalized zero shot learning
Haofeng Zhang 0001, Jingren Liu, Yazhou Yao, Yang Long 0001 |
Pattern Recognit. Lett. | 1 |
| 2020 | A Probabilistic Zero-Shot Learning Method via Latent Nonnegative Prototype Synthesis of Unseen ClassesabstractZero-shot learning (ZSL), a type of structured multioutput learning, has attracted much attention due to its requirement of no training data for target classes. Conventional ZSL methods usually project visual features into semantic space and assign labels by finding their nearest prototypes. However, this type of nearest neighbor search (NNS)-based method often suffers from great performance degradation because of the nonuniform variances between different categories. In this article, we propose a probabilistic framework by taking covariance into account to deal with the above-mentioned problem. In this framework, we define a new latent space, which has two characteristics. The first is that the features in this space should gather within the classes and scatter between the classes, which is implemented by triplet learning; the second is that the prototypes of unseen classes are synthesized with nonnegative coefficients, which are generated by nonnegative matrix factorization (NMF) of relations between the seen classes and the unseen classes in attribute space. During training, the learned parameters are the projection model for triplet network and the nonnegative coefficients between the unseen classes and the seen classes. In the testing phase, visual features are projected into latent space and assigned with the labels that have the maximum probability among unseen classes for classic ZSL or within all classes for generalized ZSL. Extensive experiments are conducted on four popular data sets, and the results show that the proposed method can outperform the state-of-the-art methods in most circumstances. Haofeng Zhang 0001, Huaqi Mao, Yang Long 0001, Wankou Yang, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | A General Transductive Regularizer for Zero-Shot Learning
Huaqi Mao, Haofeng Zhang 0001, Yang Long 0001, Longzhi Yang |
BMVC | 2 |
| 2019 | Clustering-driven unsupervised deep hashing for image retrieval
Haofeng Zhang 0001, Yazhou Yao, Wankou Yang, Li Liu 0004 |
Neurocomputing | 3 |
| 2019 | Adversarial unseen visual feature synthesis for Zero-shot Learning
Haofeng Zhang 0001, Yang Long 0001, Li Liu 0004, Ling Shao 0001 |
Neurocomputing | 1 |
| 2019 | Dual-verification network for zero-shot learning
Haofeng Zhang 0001, Yang Long 0001, Wankou Yang, Ling Shao 0001 |
Inf. Sci. | 1 |
| 2019 | Zero-shot leaning and hashing with binary visual similes
Haofeng Zhang 0001, Yang Long 0001, Ling Shao 0001 |
Multim. Tools Appl. | 1 |
| 2019 | Zero-shot Hashing with orthogonal projection for image retrieval
Haofeng Zhang 0001, Yang Long 0001, Ling Shao 0001 |
Pattern Recognit. Lett. | 1 |
| 2019 | Triple Verification Network for Generalized Zero-Shot LearningabstractConventional Zero-shot Learning approaches often suffer from severe performance degradation in the Generalised Zero-shot Learning (GZSL) scenario, i.e. to recognise test images that are from both seen and unseen classes. This paper studies the Class-level Over-fitting (CO) and empirically shows its effects to GZSL. We then address ZSL as a Triple Verification problem and propose a unified optimisation of regression and compatibility functions, i.e. two main streams of existing ZSL approaches. The complementary losses mutually regularise the same model to mitigate the CO problem. Furthermore, we implement a deep extension paradigm to linear models and significantly outperforms state-of-the-art methods in both GZSL and ZSL scenarios on the four standard benchmarks. Haofeng Zhang 0001, Yang Long 0001, Yu Guan 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Depth Embedded Recurrent Predictive Parsing Network for Video ScenesabstractSemantic segmentation-based scene parsing plays an important role in automatic driving and autonomous navigation. However, most of the previous models only consider static images, and fail to parse sequential images because they do not take the spatial-temporal continuity between consecutive frames in a video into account. In this paper, we propose a depth embedded recurrent predictive parsing network (RPPNet), which analyzes preceding consecutive stereo pairs for parsing result. In this way, RPPNet effectively learns the dynamic information from historical stereo pairs, so as to correctly predict the representations of the next frame. The other contribution of this paper is to systematically study the video scene parsing (VSP) task, in which we use the RPPNet to facilitate conventional image paring features by adding spatial-temporal information. The experimental results show that our proposed method RPPNet can achieve fine predictive parsing results on cityscapes and the predictive features of RPPNet can significantly improve conventional image parsing networks in VSP task. Lingli Zhou, Haofeng Zhang 0001, Yang Long 0001, Ling Shao 0001, Jing-Yu Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | ParallelNet: A Depth-Guided Parallel Convolutional Network for Scene Segmentation
Haofeng Zhang 0001 |
PRICAI (1) | 2 |
| 2018 | 3SP-Net: Semantic Segmentation Network with Stereo Image Pairs for Urban Scene Parsing
Lingli Zhou, Haofeng Zhang 0001 |
PRICAI (1) | 2 |
| 2018 | Unsupervised Deep Hashing With Pseudo Labels for Scalable Image RetrievalabstractIn order to achieve efficient similarity searching, hash functions are designed to encode images into low-dimensional binary codes with the constraint that similar features will have a short distance in the projected Hamming space. Recently, deep learning-based methods have become more popular, and outperform traditional non-deep methods. However, without label information, most state-of-the-art unsupervised deep hashing (DH) algorithms suffer from severe performance degradation for unsupervised scenarios. One of the main reasons is that the ad-hoc encoding process cannot properly capture the visual feature distribution. In this paper, we propose a novel unsupervised framework that has two main contributions: 1) we convert the unsupervised DH model into supervised by discovering pseudo labels; 2) the framework unifies likelihood maximization, mutual information maximization, and quantization error minimization so that the pseudo labels can maximumly preserve the distribution of visual features. Extensive experiments on three popular data sets demonstrate the advantages of the proposed method, which leads to significant performance improvement over the state-of-the-art unsupervised hashing algorithms. Haofeng Zhang 0001, Li Liu 0004, Yang Long 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Fast orthogonal linear discriminant analysis with applications to image classificationabstractOrthogonalized variant of Linear Discriminant Analysisis (LDA) is an effective statistical learning tool for dimension reduction. However, existing orthogonalized LDA algorithms suffer from various drawbacks, including the requirement for expensive computing time. This paper develops an efficient algorithm for dimension reduction, referred to as Fast Orthogonal Linear Discriminant Analysis (FOLDA), which adopts an iterative procedure to extract the orthogonal projection vectors. Different from previous efforts, this new approach applies QR decomposition and regression to solve for a new projection vector in each time of iterations, leading to the by far cheaper computational cost. FOLDA can achieve comparable recognition rates to existing orthogonal LDA algorithms. Experimental results on image databases, such as MNIST, COIL20, MEPG-7, and OUTEX, show the effectiveness and efficiency of FOLDA. Qiaolin Ye, Ning Ye 0001, Haofeng Zhang 0001, Chunxia Zhao |
IJCNN | 3 |
| 2014 | Flexible orthogonal semisupervised learning for dimension reduction with image classification
Qiaolin Ye, Ning Ye 0001, Chunxia Zhao, Tongming Yin, Haofeng Zhang 0001 |
Neurocomputing | 5 |
| 2012 | Recursive "concave-convex" Fisher Linear Discriminant with applications to face, handwritten digit and terrain recognition
Qiaolin Ye, Chunxia Zhao, Haofeng Zhang 0001, Xiaobo Chen 0001 |
Pattern Recognit. | 3 |
| 2011 | Distance difference and linear programming nonparallel plane classifier
Qiaolin Ye, Chunxia Zhao, Haofeng Zhang 0001, Ning Ye 0001 |
Expert Syst. Appl. | 3 |
| 2011 | An improved cooperative particle swarm optimization and its application
Debao Chen, Chunxia Zhao, Haofeng Zhang 0001 |
Neural Comput. Appl. | 3 |
| 2008 | Particle swarm based stereo algorithm and disparity map evaluationabstractIn this paper, a new particle swarm based stereo algorithm is presented. Our motivation is to improve the accuracy of the disparity map by removing the mismatches caused by both occlusions and false targets. In our approach, the stereo matching problem is divided into two steps, including partial matching of segmented image and particle swarm optimization of the rest. The algorithm first takes advantage of SAD and Dynamic Programming to remove the mismatches mainly caused by visibility problems; after the first step, the algorithm selects all the rest image segmented regions, takes them as a particle and uses particle swarm to optimization it. In the second step, the cost function is defined on the pixel level, as well as on the segmented level, while the pixel level measures the data similarity based the current disparity map, the segmented level incorporates a smooth term. Results obtained for benchmark indicate that the proposed method is able to get rather accurate disparity maps. Haofeng Zhang 0001, Chunxia Zhao, Zhenmin Tang, Jing-Yu Yang 0001 |
ICARCV | 1 |
| 2008 | Curvature diffusion evolution in image filteringabstractThe neighborhood structure of a pixel in an image can be described more accurately by its two principal curvatures than its gradient or mean curvature-based estimation. Based on this idea, we propose a novel method -- minimum principal curvature-driven diffusion, in which the two principal curvatures are used in a curvature-driven diffusion equation for image filtering. The main advantage of the proposed method over the existing methods is that it preserves not only conventional structures, such as edges, but also some fine structures such as ridges or thin lines. Hong-nan Wang, Chunxia Zhao, Haofeng Zhang 0001, Yong Hu 0004 |
ICARCV | 3 |
| 2008 | Road-surface abstraction using ladar sensingabstractPropose a road-surface abstraction algorithm which suitable for structured and semi-structured road environments. Algorithm uses fuzzy cluster method which based on maximum entropy theory to cluster ladar points that belong to a scan line. After fitting clustered data linearly, one can abstract straight lines that belong to road-surface by their location and slope angle. We can acquire a current referenced horizontal by comparing several continuous ladar scan lines and then the algorithm abstracts obstacles on road-surface area. Experiments show our algorithm works well in spite of the road-boundary's shape is regular or not, and free from the impact of complex texture or irregular illumination of the road. Xia Yuan, Chunxia Zhao, Yun-fei Cai, Haofeng Zhang 0001, Debao Chen |
ICARCV | 4 |