EDBT 2026 Demo / reviewers in the wild / expert
Bingfeng Zhang
dblp:00/1201
· DBLP profile ↗
39ranked-venue papers
10as first author
37since 2021 · last 2026
0000-0001-5751-2873ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 10 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 22 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Modelsabstract3D understanding has drawn significant attention recently, leveraging Vision-Language Models (VLMs) to enable multi-modal reasoning between point cloud and text data. Current 3D-VLMs directly embed the 3D point clouds into 3D tokens, following large 2D-VLMs with powerful reasoning capabilities. However, this framework has a great computational cost limiting its application, where we identify that the bottleneck lies in processing all 3D tokens in the Large Language Model (LLM) part. This raises the question: how can we reduce the computational overhead introduced by 3D tokens while preserving the integrity of their essential information? To address this question, we introduce Hierarchical Compensatory Compression (HCC-3D) to efficiently compress 3D tokens while maintaining critical detail retention. Specifically, we first propose a global structure compression (GSC), in which we design global queries to compress all 3D tokens into a few key tokens while keeping overall structural information. Then, to compensate for the information loss in GSC, we further propose an adaptive detail mining (ADM) module that selectively recompresses salient but under-attended features through complementary scoring. Extensive experiments demonstrate that HCC-3D not only achieves extreme compression ratios (approximately 98%) compared to previous 3D VLMs, but also achieves new state-of-the-art performance, showing the great improvements on both efficiency and performance. Liheng Zhang, Bingfeng Zhang, Weifeng Liu 0001 |
AAAI | 4 |
| 2026 | Balancing semantic and structural decoding for fMRI-to-image reconstruction
Wanqi He, Hanyang Chi, Bingfeng Zhang |
Expert Syst. Appl. | 5 |
| 2026 | SCM: Semantic Segmentation with Dual-stream Semantic Synergy under Adverse Weather Conditions
Shuochen Tian, Jian Pang, Bingfeng Zhang, Weifeng Liu 0001 |
Multim. Syst. | 4 |
| 2026 | Unleashing the power of optimal head in CLIP and DINO for weakly supervised semantic segmentation
Xianglin Qiu, Siyue Yu, Bingfeng Zhang, Tammam Tillo, Jimin Xiao |
Pattern Recognit. | 3 |
| 2026 | Adaptive sparse contrastive learning for unsupervised object re-identification
Dingyuan Zheng, Yang Liu 0195, Jimin Xiao, Bingfeng Zhang |
Pattern Recognit. | 5 |
| 2026 | Cross-Hierarchical Decoding With SAM for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation (SSMIS) mainly leverages valuable information from unlabeled data to complement the limited labeled guidance. Incorporating SAM, with its excellent generalization capabilities, enhances the learning process from unlabeled data, as demonstrated by existing methods. However, most current SAM-based methods in SSMIS focus on unique prompt design, while the prompts generated for unlabeled data through pseudo-labels unavoidably introduce noise, limiting the following decoding process. In this paper, we propose a Cross-Hierarchical Decoding (CHD) process for SAM, which removes explicit prompts (e.g., point or box) and thus mitigates the influence of inaccurate pseudo labels. Specifically, our CHD is a two-stage decoder. The first stage uses the original decoder in SAM to generate probability masks, which are combined with a learnable mask interaction module in the second stage to achieve more fine-grained segmentation. Meanwhile, to remove the restriction that the original SAM can only segment foreground-background categories, we design a cross-class correlation module in CHD to capture class-wise interrelationships between different classes, thus achieving multi-class segmentation. Extensive experiments show that CHD achieves new state-of-the-art performance for SSMIS, significantly improving different baselines. Hanyang Chi, Xuru Gao, Guixun Luo, Bingfeng Zhang, Weifeng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | CO3+: Improved Collaborative Consortium of Foundation Models for Open-World Few-Shot LearningabstractOpen-World Few-Shot Learning (OFSL) is a critical research domain focused on accurately identifying target samples under conditions where data is scarce and labels are unreliable. This field is highly relevant to real-world scenarios, holding significant practical implications. Currently, the field has only a few solutions, primarily relying on conventional methods such as metric learning and feature aggregation. However, these methods often struggle in more complex scenarios. Recent breakthroughs in foundation models such as CLIP and DINO have demonstrated their strong representational capabilities, even in resource-limited environments. These advancements have led to a shift from “training model from scratch” towards “exploiting the extensive capabilities and expertise of these pre-trained foundation models for OFSL”. Inspired by this shift, we introduce the Improved Collaborative Consortium of Foundation Models (CO+3), an extension of CO3, first presented in AAAI 2024. CO+3significantly improves the accuracy of OFSL by integrating the strengths of four foundational models. It includes three decoupled blocks: (1) The Label Correction Block (LC-Block) rectifies unreliable labels, (2) the Data Augmentation Block (DA-Block) enriches the available data, and (3) the Text-guided Fusion Adapter (TeFu-Adapter) merges various features and reduces the impact of noisy labels through semantic constraints. We evaluate CO+3across eleven benchmark datasets, comparing it against recent state-of-the-art methods. Our thorough evaluations demonstrate that the proposed CO+3consistently surpasses existing methods by a substantial margin, particularly in high-noise scenarios. Shuai Shao 0006, Rui Xu 0012, Bingfeng Zhang, Baodi Liu, Weifeng Liu 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Unbiased Semantic Decoding With Vision Foundation Models for Few-Shot SegmentationabstractFew-shot segmentation (FSS) has garnered significant attention. Many recent approaches attempt to introduce the segment anything model (SAM) to handle this task. With the strong generalization ability and rich object-specific extraction ability of the SAM model, such a solution shows great potential in FSS. However, the decoding process of SAM highly relies on accurate and explicit prompts, making previous approaches mainly focus on extracting prompts from the support set, which is insufficient to activate the generalization ability of SAM, and this design is easy to result in a biased decoding process when adapting to the unknown classes. In this work, we propose an unbiased semantic decoding (USD) strategy integrated with SAM, which extracts target information from both the support and query set simultaneously to perform consistent predictions guided by the semantics of the contrastive language-image pretraining (CLIP) model. Specifically, to enhance the unbiased semantic discrimination of SAM, we design two feature enhancement strategies that leverage the semantic alignment capability of CLIP to enrich the original SAM features, mainly including a global supplement at the image level to provide a generalize category indicate with support image and a local guidance at the pixel level to provide a useful target location with query image. Besides, to generate target-focused prompt embeddings, a learnable visual-text target prompt generator (VTPG) is proposed by interacting target text embeddings and clip visual features. Without requiring retraining of the vision foundation models, the features with semantic discrimination draw attention to the target region through the guidance of prompt with rich target information. Experiments on both the PASCAL- $5^{i}$ and COCO- $20^{i}$ show that our proposed method outperforms the existing approaches by a clear margin and achieves new state-of-the-art performances. Bingfeng Zhang, Jian Pang, Weifeng Liu 0001, Baodi Liu, Honglong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | A Training-free Synthetic Data Selection Method for Semantic SegmentationabstractTraining semantic segmenter with synthetic data has been attracting great attention due to its easy accessibility and huge quantities. Most previous methods focused on producing large-scale synthetic image-annotation samples and then training the segmenter with all of them. However, such a solution remains a main challenge in that the poor-quality samples are unavoidable, and using them to train the model will damage the training process. In this paper, we propose a training-free Synthetic Data Selection (SDS) strategy with CLIP to select high-quality samples for building a reliable synthetic dataset. Specifically, given massive synthetic image-annotation pairs, we first design a Perturbation-based CLIP Similarity (PCS) to measure the reliability of synthetic image, thus removing samples with low-quality images. Then we propose a class-balance Annotation Similarity Filter (ASF) by comparing the synthetic annotation with the response of CLIP to remove the samples related to low-quality annotations. The experimental results show that using our method significantly reduces the data size by half, while the trained segmenter achieves higher performance. Siyue Yu, Jian Pang, Bingfeng Zhang |
AAAI | 4 |
| 2025 | POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) leverages Class Activation Maps (CAMs) to extract spatial information from image-level labels. However, CAMs primarily highlight the most discriminative foreground regions, leading to incomplete results. Prototype-based methods attempt to address this limitation by employing prototype CAMs instead of classifier CAMs. Nevertheless, existing prototype-based methods typically use a single prototype for each class, which is insufficient to capture all attributes of the foreground features due to the significant intra-class variations across different images. Consequently, these methods still struggle with incomplete CAM predictions. In this paper, we propose a novel framework called Prototypical Optimal Transport (POT) for WSSS. POT enhances CAM predictions by dividing features into multiple clusters and activating each cluster using its prototype. In this process, a similarity-aware optimal transport is employed to assign features to the most probable clusters. This similarity-aware strategy ensures the prioritization of significant cluster prototypes, thereby improving the accuracy of feature assignment. Additionally, we introduce an adaptive OT-based consistency loss to refine feature representations. This framework effectively overcomes the limitations of single-prototype methods, providing more complete and accurate CAM predictions. Extensive experimental results on standard WSSS benchmarks (PASCAL VOC and MS COCO) demonstrate that our method significantly improves the quality of CAMs and achieves state-of-the-art performances. The source code will be released https://github.com/jianwang91/POT. Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
CVPR | 3 |
| 2025 | Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic Segmentation
Siyue Yu, Bingfeng Zhang, Mingjie Sun, Yi Dong 0002, Jimin Xiao |
ICCV | 3 |
| 2025 | Class Token as Proxy: Optimal Transport-Assisted Proxy Learning for Weakly Supervised Semantic Segmentation
Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
ICCV | 3 |
| 2025 | Simulating Confidence Intervals for Conditional Value-at-Risk via Least-Squares MetamodelsabstractMetamodeling techniques have been applied to approximate portfolio loss as a function of financial risk factors, thus producing point estimates of various measures of portfolio risk based on Monte Carlo samples. Rather than point estimates, this paper focuses on the construction of confidence intervals (CIs) for a widely used risk measure, the so-called conditional value-at-risk (CVaR), when the least-squares method (LSM) is employed as a metamodel in the point estimation. To do so, we first develop lower and upper bounds of CVaR and construct CIs for these bounds. Then, the lower end of the CI for the lower bound and the upper end of the CI for the upper bound together form a CI of CVaR with justifiable statistical guarantee, which accounts for both the metamodel error and the noises of Monte Carlo samples. The proposed CI procedure reuses the samples simulated for LSM point estimation, thus requiring no additional simulation budget. We demonstrate via numerical examples that the proposed procedure may lead to a CI with the desired coverage probability and a much smaller width than that of an existing CI in the literature. History: Accepted by Bruno Tuffin, Area Editor for Simulation. Funding: This research was supported by the National Natural Science Foundation of China (NNSFC) [Grants 72101260 and 72471232], the Research Grants Council of Hong Kong (RGC-HK) [General Research Fund Project 11508620], InnoHK Initiative, the Government of the HKSAR, and Laboratory for AI-Powered Financial Technologies, and NNSFC/RGC-HK Joint Research Scheme [Project N_CityU 105/21]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0394 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0394 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Qidong Lai, Guangwu Liu, Bingfeng Zhang, Kun Zhang 0036 |
INFORMS J. Comput. | 3 |
| 2025 | Frozen CLIP-DINO: A Strong Backbone for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has witnessed great achievements with image-level labels. Several recent approaches use the CLIP model to generate pseudo labels for training an individual segmentation model, while there is no attempt to apply the CLIP model as the backbone to directly segment objects with image-level labels. In this paper, we propose WeCLIP and its advanced version WeCLIP+, to build the single-stage pipeline for weakly supervised semantic segmentation. For WeCLIP, the frozen CLIP model is applied as the backbone for semantic feature extraction, and a new light decoder is designed to interpret extracted semantic features for final prediction. Meanwhile, we utilize the above frozen backbone to generate pseudo labels for training the decoder. Such labels are fixed during training. We then propose a refinement module (RFM) to optimize them dynamically. For WeCLIP+, we introduce the frozen DINO model to achieve more comprehensive semantic feature extraction. The frozen DINO is combined with the frozen CLIP as the backbone, followed by a shared decoder to make predictions with less training cost. Moreover, a strengthened refinement module (RFM+) is designed to revise online pseudo labels with extra guidance from DINO features. Extensive experiments show that both WeCLIP and WeCLIP+ significantly outperform other approaches with less training cost. Particularly, WeCLIP+ gets mIoU of 83.9% on VOC 2012 test set and 56.3% on COCO val set. Additionally, these two approaches also obtain promising results for fully supervised settings. Bingfeng Zhang, Siyue Yu, Jimin Xiao, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Dynamic feature regularized loss for weakly supervised semantic segmentation
Bingfeng Zhang, Jimin Xiao, Yao Zhao 0001 |
Pattern Recognit. | 1 |
| 2025 | See Degraded Objects: A Physics-Guided Approach for Object Detection in Adverse EnvironmentsabstractIn adverse environments, the detector often fails to detect degraded objects because they are almost invisible and their features are weakened by the environment. Common approaches involve image enhancement to support detection, but they inevitably introduce human-invisible noise that negatively impacts the detector. In this work, we propose a physics-guided approach for object detection in adverse environments, which gives a straightforward solution that injects the physical priors into the detector, enabling it to detect poorly visible objects. The physical priors, derived from the imaging mechanism and image property, include environment prior and frequency prior. The environment prior is generated from the physical model, e.g., the atmospheric model, which reflects the density of environmental noise. The frequency prior is explored based on an observation that the amplitude spectrum could highlight object regions from the background. The proposed two priors are complementary in principle. Furthermore, we present a physics-guided loss that incorporates a novel weight item, which is estimated by applying the membership function on physical priors and could capture the extent of degradation. By backpropagating the physics-guided loss, physics knowledge is injected into the detector to aid in locating degraded objects. We conduct experiments in synthetic foggy environment, real foggy environment, and real underwater scenario. The results demonstrate that our method is effective and achieves state-of-the-art performance. The code is available at https://github.com/PangJian123/See-Degraded-Objects. Weifeng Liu 0001, Jian Pang, Bingfeng Zhang, Baodi Liu, Dapeng Tao |
IEEE Trans. Image Process. | 3 |
| 2024 | Adaptive Bidirectional Displacement for Semi-Supervised Medical Image SegmentationabstractConsistency learning is a central strategy to tackle unlabeled data in semi-supervised medical image segmentation (SSMIS), which enforces the model to produce consistent predictions under the perturbation. However, most current approaches solely focus on utilizing a specific single perturbation, which can only cope with limited cases, while employing multiple perturbations simultaneously is hard to guarantee the quality of consistency learning. In this paper, we propose an Adaptive Bidirectional Displacement (ABD) approach to solve the above challenge. Specifically, we first design a bidirectional patch displacement based on reliable prediction confidence for unlabeled data to generate new samples, which can effectively suppress uncontrollable regions and still retain the influence of input perturbations. Meanwhile, to enforce the model to learn the potentially uncontrollable content, a bidirectional displacement operation with inverse confidence is proposed for the labeled images, which generates samples with more unreliable information to facilitate model learning. Extensive experiments show that ABD achieves new state-of-the-art performances for SSMIS, significantly improving different base-lines. Source code is available at https://github.com/chy-upclABD. Hanyang Chi, Jian Pang, Bingfeng Zhang, Weifeng Liu 0001 |
CVPR | 3 |
| 2024 | Rethinking Prior Information Generation with CLIP for Few-Shot SegmentationabstractFew-shot segmentation remains challenging due to the limitations of its labeling information for unseen classes. Most previous approaches rely on extracting high-level fea-ture maps from the frozen visual encoder to compute the pixel- wise similarity as a key prior guidance for the decoder. However, such a prior representation suffers from coarse granularity and poor generalization to new classes since these high-level feature maps have obvious category bias. In this work, we propose to replace the visual prior representation with the visual-text alignment capacity to capture more reliable guidance and enhance the model generalization. Specifically, we design two kinds of trainingfree prior information generation strategy that attempts to utilize the semantic alignment capability of the Contrastive Language-Image Pre-training model (CLIP) to locate the target class. Besides, to acquire more accurate prior guidance, we build a high-order relationship of attention maps and utilize it to refine the initial prior information. Experiments on both the PASCAL-5i and COCO-20i datasets show that our method obtains a clearly substantial improvement and reaches the new state-of-the-art performance. The code is available on the project website11https://github.com/vangjin/PI-CLIP. Bingfeng Zhang, Jian Pang, Honglong Chen, Weifeng Liu 0001 |
CVPR | 2 |
| 2024 | Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has witnessed great achievements with image-level labels. Several recent approaches use the CLIP model to generate pseudo labels for training an individual segmentation model, while there is no attempt to apply the CLIP model as the backbone to directly segment objects with image-level labels. In this paper, we propose WeCLIP, a CLIP-based single-stage pipeline, for weakly supervised semantic segmentation. Specifically, the frozen CLIP model is applied as the backbone for semantic feature extraction, and a new decoder is designed to interpret extracted semantic features for final prediction. Meanwhile, we utilize the above frozen backbone to generate pseudo labels for training the decoder. Such labels cannot be optimized during training. We then propose a refinement module (RFM) to rectify them dynamically. Our architecture enforces the proposed decoder and RFM to benefit from each other to boost the final performance. Extensive experiments show that our approach significantly outperforms other approaches with less training cost. Additionally, our WeCLIP also obtains promising results for fully supervised settings. The code is available at https://github.com/zbf1991/WeCLIP. Bingfeng Zhang, Siyue Yu, Yunchao Wei, Yao Zhao 0001, Jimin Xiao |
CVPR | 1 |
| 2024 | PSDPM: Prototype-based Secondary Discriminative Pixels Mining for Weakly Supervised Semantic SegmentationabstractImage-level Weakly Supervised Semantic Segmentation (WSSS) has received increasing attention due to its low an-notation cost. Class Activation Mapping (CAM) generated through classifier weights in WSSS inevitably ignores cer-tain useful cues, while the CAM generated through class prototypes can alleviate that. However, because of the dif-ferent goals of image classification and semantic segmentation, the class prototypes still focus on activating primary discriminative pixels learned from classification loss, leading to incomplete CAM. In this paper, we propose a plug-and-play Prototype-based Secondary Discriminative Pixels Mining (PSDPM) framework for enabling class prototypes to activate more secondary discriminative pixels, thus gen-erating a more complete CAM. Specifically, we introduce a Foreground Pixel Estimation Module (FPEM) for esti-mating potential foreground pixels based on the correlations between primary and secondary discriminative pix-els and the semantic segmentation results of baseline meth-ods. Then, we enable WSSS model to learn discriminative features from secondary discriminative pixels through a consistency loss calculated between FPEM result and class-prototype CAM. Experimental results show that our PSDPM improves various baseline methods significantly and achieves new state-of-the-art performances on WSSS benchmarks. Codes are available at https://github.com/xinqiaozhao/PSDPM. Xinqiao Zhao, Ziqian Yang, Tianhong Dai, Bingfeng Zhang, Jimin Xiao |
CVPR | 4 |
| 2024 | Adversarial Erasing Transformer for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has attracted a lot of attention recently. Previous methods can be divided into two types, which are single-stage training and multi-stage training. In this paper, we focus on multi-stage training for image-level weakly supervised semantic segmentation. Many recent methods have tried to use transformer architecture as the backbone for CAM generation since it can capture global relationships to refine CAM accurately. However, we observe that such a backbone still fails to generate complete and smooth CAM. We argue that this is because the attention mechanism in the transformer can only pay attention to the most discriminative relationships. It is difficult to capture semantic-level long-range pair-wise relationships under image-level supervision. Thus, we propose an adversarial erasing transformer network called AETN, where an erasing attention mechanism is designed to establish more extensive pair-wise relationships. To cope with erasing, more target features will be forced to activate. Thus, better feature representation can be obtained for more accurate CAM generation. Besides, to further help our network learn better feature representation, we propose a self-consistent learning mechanism based on different augmentations. In this way, our AETN outperforms recent methods. Our AETN achieves 73.0 mIoU on the PASCAL VOC 2012 val set and 73.9 mIoU on the PASCAL VOC 2012 test set. Code is available a https://github.com/siyueyu/AETN. Bingfeng Zhang, Siyue Yu, Xuru Gao, Mingjie Sun, Eng Gee Lim, Jimin Xiao |
ECAI | 1 |
| 2024 | HTPSeg: A Semantic Segmentation Database for House-Tree-Person Psychological TestabstractThe House Tree Person (HTP) test is widely recommended for clinical application of mental illness. Traditionally, therapists assess a patient's mental state by analyzing the content of their HTP drawings manually, which is time-consuming and susceptible to the therapist's subjective influences. Recently, using intelligent diagnostic models to tackle HTP test has attracted much attention. However, most existing models attempt to make a binary classification for HTP drawings, i.e., positive or negative psychological states, making details that can reflect the psychological state of the patient lost. To address the above challenge, in this paper, we introduce the semantic segmentation task into the HTP test for the first time to generate pixel-level semantic information. We first construct a semantic segmentation dataset about HTP psychological diagnosis named HTPSeg. Subsequently, to overcome the domain gap between HTP drawings and natural images, we propose to use the Low-Rank Adaptation (LoRA) fine-tuning strategy to adapt the Segment Anything Model (SAM) to the task of HTP drawing analysis. Specifically, we integrate the rank decomposition matrices into the projection layer of the transformer block in SAM's image encoder for fine-tuning. Additionally, we freeze the Mask Decoder and fine-tune the Prompt Encoder using default embeddings. Extensive experiments indicate the effectiveness and efficiency of the proposed method. The dataset and code will be available at https://github.com/clown06/HTPSeg. Bingfeng Zhang, Weifeng Liu 0001 |
ICTAI | 4 |
| 2024 | MCNet: Magnitude consistency network for domain adaptive object detection under inclement environments
Jian Pang, Weifeng Liu 0001, Bingfeng Zhang, Xinghao Yang, Baodi Liu, Dapeng Tao |
Pattern Recognit. | 3 |
| 2024 | Cross-frame feature-saliency mutual reinforcing for weakly supervised video salient object detection
Jian Wang 0122, Siyue Yu, Bingfeng Zhang, Xinqiao Zhao, Ángel F. García-Fernández, Eng Gee Lim, Jimin Xiao |
Pattern Recognit. | 3 |
| 2024 | Enhanced online CAM: Single-stage weakly supervised semantic segmentation via collaborative guidance
Bingfeng Zhang, Xuru Gao, Siyue Yu, Weifeng Liu 0001 |
Pattern Recognit. | 1 |
| 2023 | Hunting Sparsity: Density-Guided Contrastive Learning for Semi-Supervised Semantic SegmentationabstractRecent semi-supervised semantic segmentation methods combine pseudo labeling and consistency regularization to enhance model generalization from perturbation-invariant training. In this work, we argue that adequate supervision can be extracted directly from the geometry of feature space. Inspired by density-based unsupervised clustering, we propose to leverage feature density to locate sparse regions within feature clusters defined by label and pseudo labels. The hypothesis is that lower-density features tend to be under-trained compared with those densely gathered. Therefore, we propose to apply regularization on the structure of the cluster by tackling the sparsity to increase intra-class compactness in feature space. With this goal, we present a Density-Guided Contrastive Learning (DGCL) strategy to push anchor features in sparse regions toward cluster centers approximated by high-density positive keys. The heart of our method is to estimate feature density which is defined as neighbor compactness. We design a multi-scale density estimation module to obtain the density from multiple nearest-neighbor graphs for robust density modeling. Moreover, a unified training framework is proposed to combine label-guided self-training and density-guided geometry regularization to form complementary supervision on unlabeled data. Experimental results on PAS-CAL VOC and Cityscapes under various semi-supervised settings demonstrate that our proposed method achieves state-of-the-art performances. The project is available at https://github.com/Gavinwxy/DGCL. Xiaoyang Wang 0007, Bingfeng Zhang, Limin Yu, Jimin Xiao |
CVPR | 2 |
| 2023 | Self-Compensating Learning for Few-Shot SegmentationabstractFew-shot segmentation (FSS) has witnessed rapid development. Most existing approaches extract prototypes from support images to segment query images. However, the integrity and validity of these support prototypes cannot be guaranteed. To solve the above drawbacks, we propose a self-compensating strategy, aiming to provide query-aware support information, to build more effective matching between support information and query images. Specifically, we design a prototype compensating module to mine useful information from the query prediction, to update original support prototypes as new query-aware support prototypes. Then the updated prototypes are utilized to perform the second matching with query features. In addition, we also compensate the information of original prior masks on the second matching phase, to improve the quality of prior masks. With improved prototype representations and prior knowledge, our approach can directly improve the performance of different approaches with new state-of-the-art performances. Bingfeng Zhang, Weifeng Liu 0001, Baodi Liu, Siyue Yu |
ICIP | 2 |
| 2023 | Credible Dual-Expert Learning for Weakly Supervised Semantic Segmentation
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Yao Zhao 0001 |
Int. J. Comput. Vis. | 1 |
| 2023 | Weight-guided class complementing for long-tailed image recognition
Xinqiao Zhao, Jimin Xiao, Siyue Yu, Hui Li 0085, Bingfeng Zhang |
Pattern Recognit. | 5 |
| 2023 | Weight-guided loss for long-tailed object detection and instance segmentation
Xinqiao Zhao, Jimin Xiao, Bingfeng Zhang, Waleed Al-Nuaimy |
Signal Process. Image Commun. | 3 |
| 2022 | Democracy Does Matter: Comprehensive Feature Mining for Co-Salient Object DetectionabstractCo-salient object detection, with the target of detecting co-existed salient objects among a group of images, is gaining popularity. Recent works use the attention mechanism or extra information to aggregate common co-salient features, leading to incomplete even incorrect responses for target objects. In this paper, we aim to mine comprehensive co-salient features with democracy and reduce background interference without introducing any extra information. To achieve this, we design a democratic prototype generation module to generate democratic response maps, covering sufficient co-salient regions and thereby involving more shared attributes of co-salient objects. Then a comprehensive prototype based on the response maps can be generated as a guide for final prediction. To suppress the noisy background information in the prototype, we propose a self-contrastive learning module, where both positive and negative pairs are formed without relying on additional classification information. Besides, we also design a democratic feature enhancement module to further strengthen the co-salient features by readjusting attention values. Extensive experiments show that our model obtains better performance than previous state-of-the-art methods, especially on challenging real-world cases (e.g., for CoCA, we obtain a gain of 2.0% for MAE, 5.4% for maximum F-measure, 2.3% for maximum E-measure, and 3.7% for S-measure) under the same settings. Source code is available at https://github.com/siyueyu/DCFM. Siyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee Lim |
CVPR | 3 |
| 2022 | CARD: Semi-supervised Semantic Segmentation via Class-agnostic Relation based DenoisingabstractRecent semi-supervised semantic segmentation methods focus on mining extra supervision from unlabeled data by generating pseudo labels. However, noisy labels are inevitable in this process which prevent effective self-supervision. This paper proposes that noisy labels can be corrected based on semantic connections among features. Since a segmentation classifier produces both high and low-quality predictions, we can trace back to feature encoder to investigate how a feature in a noisy group is related to those in the confident groups. Discarding the weak predictions from the classifier, rectified predictions are assigned to the wrongly predicted features through the feature relations. The key to such an idea lies in mining reliable feature connections. With this goal, we propose a class-agnostic relation network to precisely capture semantic connections among features while ignoring their semantic categories. The feature relations enable us to perform effective noisy label corrections to boost self-training performance. Extensive experiments on PASCAL VOC and Cityscapes demonstrate the state-of-the-art performances of the proposed methods under various semi-supervised settings. Xiaoyang Wang 0007, Jimin Xiao, Bingfeng Zhang, Limin Yu |
IJCAI | 3 |
| 2022 | Affinity Attention Graph Neural Network for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation is receiving great attention due to its low human annotation cost. In this paper, we aim to tackle bounding box supervised semantic segmentation, i.e., training accurate semantic segmentation models using bounding box annotations as supervision. To this end, we propose affinity attention graph neural network ($A^2$A2GNN). Following previous practices, we first generate pseudo semantic-aware seeds, which are then formed into semantic graphs based on our newly proposed affinity Convolutional Neural Network (CNN). Then the built graphs are input to our$A^2$A2GNN, in which an affinity attention layer is designed to acquire the short- and long- distance information from soft graph edges to accurately propagate semantic labels from the confident seeds to the unlabeled pixels. However, to guarantee the precision of the seeds, we only adopt a limited number of confident pixel seed labels for$A^2$A2GNN, which may lead to insufficient supervision for training. To alleviate this issue, we further introduce a new loss function and a consistency-checking mechanism to leverage the bounding box constraint, so that more reliable guidance can be included for the model optimization. Experiments show that our approach achieves new state-of-the-art performances on Pascal VOC 2012 datasets (val: 76.5 percent,test: 75.2 percent). More importantly, our approach can be readily applied to bounding box supervised instance segmentation task or other weakly supervised semantic segmentation tasks, with state-of-the-art or comparable performance among almot all weakly supervised tasks on PASCAL VOC or COCO dataset. Our source code will be available athttps://github.com/zbf1991/A2GNN. Bingfeng Zhang, Jimin Xiao, Jianbo Jiao, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | End-to-end weakly supervised semantic segmentation with reliable region mining
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Kaizhu Huang, Shan Luo 0001, Yao Zhao 0001 |
Pattern Recognit. | 1 |
| 2021 | Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency CoherenceabstractSparse labels have been attracting much attention in recent years. However, the performance gap between weakly supervised and fully supervised salient object detection methods is huge, and most previous weakly supervised works adopt complex training methods with many bells and whistles. In this work, we propose a one-round end-to-end training approach for weakly supervised salient object detection via scribble annotations without pre/post-processing operations or extra supervision data. Since scribble labels fail to offer detailed salient regions, we propose a local coherence loss to propagate the labels to unlabeled regions based on image features and pixel distance, so as to predict integral salient regions with complete object structures. We design a saliency structure consistency loss as self-consistent mechanism to ensure consistent saliency maps are predicted with different scales of the same image as input, which could be viewed as a regularization technique to enhance the model generalization ability. Additionally, we design an aggregation module (AGGM) to better integrate high-level features, low-level features and global context information for the decoder to aggregate various information. Extensive experiments show that our method achieves a new state-of-the-art performance on six benchmarks (e.g. for the ECSSD dataset: Fβ = 0.8995, Eξ = 0.9079 and MAE = 0.0489), with an average gain of 4.60% for F-measure, 2.05% for E-measure and 1.88% for MAE over the previous best performing method on this task. Source code is available at http://github.com/siyueyu/SCWSSOD. Siyue Yu, Bingfeng Zhang, Jimin Xiao, Eng Gee Lim |
AAAI | 2 |
| 2021 | Self-Guided and Cross-Guided Learning for Few-Shot SegmentationabstractFew-shot segmentation has been attracting a lot of attention due to its effectiveness to segment unseen object classes with a few annotated samples. Most existing approaches use masked Global Average Pooling (GAP) to encode an annotated support image to a feature vector to facilitate query image segmentation. However, this pipeline unavoidably loses some discriminative information due to the average operation. In this paper, we propose a simple but effective self-guided learning approach, where the lost critical information is mined. Specifically, through making an initial prediction for the annotated support image, the covered and uncovered foreground regions are encoded to the primary and auxiliary support vectors using masked GAP, respectively. By aggregating both primary and auxiliary support vectors, better segmentation performances are obtained on query images. Enlightened by our self-guided module for 1-shot segmentation, we propose a cross-guided module for multiple shot segmentation, where the final mask is fused using predictions from multiple annotated samples with high-quality support vectors contributing more and vice versa. This module improves the final prediction in the inference stage without re-training. Extensive experiments show that our approach achieves new state-of-the-art performances on both PASCAL-5iand COCO-20idatasets. Source code is available at https://github.com/zbf1991/SCL. Bingfeng Zhang, Jimin Xiao, Terry Qin |
CVPR | 1 |
| 2021 | Fast pixel-matching for video object segmentation
Siyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee Lim, Yao Zhao 0001 |
Signal Process. Image Commun. | 3 |
| 2020 | Reliability Does Matter: An End-to-End Weakly Supervised Semantic Segmentation ApproachabstractWeakly supervised semantic segmentation is a challenging task as it only takes image-level information as supervision for training but produces pixel-level predictions for testing. To address such a challenging task, most recent state-of-the-art approaches propose to adopt two-step solutions, i.e. 1) learn to generate pseudo pixel-level masks, and 2) engage FCNs to train the semantic segmentation networks with the pseudo masks. However, the two-step solutions usually employ many bells and whistles in producing high-quality pseudo masks, making this kind of methods complicated and inelegant. In this work, we harness the image-level labels to produce reliable pixel-level annotations and design a fully end-to-end network to learn to predict segmentation maps. Concretely, we firstly leverage an image classification branch to generate class activation maps for the annotated categories, which are further pruned into confident yet tiny object/background regions. Such reliable regions are then directly served as ground-truth labels for the parallel segmentation branch, where a newly designed dense energy loss function is adopted for optimization. Despite its apparent simplicity, our one-step solution achieves competitive mIoU scores (val: 62.6, test: 62.9) on Pascal VOC compared with those two-step state-of-the-arts. By extending our one-step method to two-step, we get a new state-of-the-art performance on the Pascal VOC (val: 66.3, test: 66.5). Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, Kaizhu Huang |
AAAI | 1 |
| 2020 | Fast Template Matching and Update for Video Object Tracking and SegmentationabstractIn this paper, the main task we aim to tackle is the multi-instance semi-supervised video object segmentation across a sequence of frames where only the first-frame box-level ground-truth is provided. Detection-based algorithms are widely adopted to handle this task, and the challenges lie in the selection of the matching method to predict the result as well as to decide whether to update the target template using the newly predicted result. The existing methods, however, make these selections in a rough and inflexible way, compromising their performance. To overcome this limitation, we propose a novel approach which utilizes reinforcement learning to make these two decisions at the same time. Specifically, the reinforcement learning agent learns to decide whether to update the target template according to the quality of the predicted result. The choice of the matching method will be determined at the same time, based on the action history of the reinforcement learning agent. Experiments show that our method is almost 10 times faster than the previous state-of-the-art method with even higher accuracy (region similarity of 69.1% on DAVIS 2017 dataset). Mingjie Sun, Jimin Xiao, Eng Gee Lim, Bingfeng Zhang, Yao Zhao 0001 |
CVPR | 4 |