VLDB 2026 Research / reviewers in the wild / expert
Junsong Fan
dblp:150/4094
· DBLP profile ↗
27ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0001-6989-2711ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAGD: Boundary-Enhanced Segment Anything in 3D Gaussian via Gaussian Decompositionabstract3D Gaussian Splatting has emerged as an alternative 3D representation for novel view synthesis, benefiting from its high-quality rendering results and real-time rendering speed. However, the 3D Gaussians learned by 3D-GS have ambiguous structures without any geometry constraints. This inherent issue in 3D-GS leads to a rough boundary when segmenting individual objects. To remedy these problems, we propose SAGD, a conceptually simple yet effective boundary-enhanced segmentation pipeline for 3D-GS to improve segmentation accuracy while preserving segmentation speed. Specifically, we introduce a Gaussian Decomposition scheme, which ingeniously utilizes the special structure of 3D Gaussians, finds out, and then decomposes the boundary Gaussians. Moreover, to achieve fast interactive 3D segmentation, we introduce a novel training-free pipeline by lifting a 2D foundation model to 3D-GS. Extensive experiments demonstrate that our approach achieves high-quality 3D segmentation without rough boundary issues, which can be easily applied to other scene editing tasks. Our code is publicly available at https://github.com/XuHu0529/SAGS. Yuxi Wang 0001, Lue Fan, Chuanchen Luo, Junsong Fan, Zhen Lei 0001, Qing Li 0001, Junran Peng, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Using Unreliable Pseudo-Labels for Label-Efficient Semantic Segmentation
Yujun Shen, Junsong Fan, Yuxi Wang 0001, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Bootstrap Masked Visual Modeling via Hard Patch MiningabstractMasked visual modeling has attracted much attention due to its promising potential in learning generalizable representations. Typical approaches urge models to predict specific contents of masked tokens, which can be intuitively considered as teaching a student (the model) to solve given problems (predicting masked contents). Under such settings, the performance is highly correlated with mask strategies (the difficulty of provided problems). We argue that it is equally important for the model to stand in the shoes of a teacher to produce challenging problems by itself. Intuitively, patches with high values of reconstruction loss can be regarded as hard samples, and masking those hard patches naturally becomes a demanding reconstruction task. To empower the model as a teacher, we propose Hard Patch Mining (HPM), predicting patch-wise losses and subsequently determining where to mask. Technically, we introduce an auxiliary loss predictor, which is trained with a relative objective to prevent overfitting to exact loss values. To gradually guide the training procedure, we propose an easy-to-hard mask strategy. Empirically, HPM brings significant improvements under both image and video benchmarks. Interestingly, solely incorporating the extra loss prediction objective leads to better representations, verifying the efficacy of determining where is hard to reconstruct. Junsong Fan, Yuxi Wang 0001, Kaiyou Song, Tiancai Wang, Xiangyu Zhang 0005, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Fully Data-Driven Pseudo Label Estimation for Pointly-Supervised Panoptic SegmentationabstractThe core of pointly-supervised panoptic segmentation is estimating accurate dense pseudo labels from sparse point labels to train the panoptic head. Previous works generate pseudo labels mainly based on hand-crafted rules, such as connecting multiple points into polygon masks, or assigning the label information of labeled pixels to unlabeled pixels based on the artificially defined traversing distance. The accuracy of pseudo labels is limited by the quality of the hand-crafted rules (polygon masks are rough at object contour regions, and the traversing distance error will result in wrong pseudo labels). To overcome the limitation of hand-crafted rules, we estimate pseudo labels with a fully data-driven pseudo label branch, which is optimized by point labels end-to-end and predicts more accurate pseudo labels than previous methods. We also train an auxiliary semantic branch with point labels, it assists the training of the pseudo label branch by transferring semantic segmentation knowledge through shared parameters. Experiments on Pascal VOC and MS COCO demonstrate that our approach is effective and shows state-of-the-art performance compared with related works. Codes are available at https://github.com/BraveGroup/FDD. Jing Li 0112, Junsong Fan, Yuran Yang, Shuqi Mei, Jun Xiao 0005, Zhaoxiang Zhang 0001 |
AAAI | 2 |
| 2024 | Continual Forgetting for Pre-Trained Vision ModelsabstractFor privacy and security concerns, the need to erase un-wanted information from pre-trained vision models is becoming evident nowadays. In real-world scenar-ios, erasure requests originate at any time from both users and model owners. These requests usually form a sequence. Therefore, under such a setting, selective information is expected to be continuously removed from a pre-trained model while maintaining the rest. We define this problem as continual forgetting and identify two key challenges. (i) For unwanted knowledge, efficient and effective deleting is crucial. (ii) For remaining knowledge, the impact brought by the forgetting procedure should be minimal. To address them, we propose Group Sparse LoRA (GS-LoRA). Specifically, towards (i), we use LoRA modules to fine-tune the FFN layers in Transformer blocks for each forgetting task independently, and towards (ii), a simple group sparse regularization is adopted, enabling automatic selection of specific LoRA groups and zeroing out the others. GS-LoRA is effective, parameter-efficient, data-efficient, and easy to implement. We conduct extensive experiments on face recognition, object detection and image classification and demonstrate that GS-LoRA manages to forget specific classes with minimal impact on other classes. Codes will be released on https://github.com/bjzhb666/GS-LoRA. Hongbo Zhao 0006, Bolin Ni, Junsong Fan, Yuxi Wang 0001, Yuntao Chen, Gaofeng Meng, Zhaoxiang Zhang 0001 |
CVPR | 3 |
| 2024 | Point-Supervised Panoptic Segmentation via Estimating Pseudo Labels from Learnable Distance
Jing Li 0112, Junsong Fan, Zhaoxiang Zhang 0001 |
ECCV (16) | 2 |
| 2024 | General Geometry-Aware Weakly Supervised 3D Object Detection
Guowen Zhang, Junsong Fan, Liyi Chen 0002, Zhaoxiang Zhang 0001, Zhen Lei 0001, Lei Zhang 0006 |
ECCV (51) | 2 |
| 2024 | Enhancing Sound Source Localization via False Negative EliminationabstractSound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling negatives in prior arts can lead to the false negative issue, where the sounds semantically similar to visual instance are sampled as negatives and incorrectly pushed away from the visual anchor/query. As a result, this misalignment of audio and visual features could yield inferior performance. To address this issue, we propose a novel audio-visual learning framework which is instantiated with two individual learning schemes: self-supervised predictive learning (SSPL) and semantic-aware contrastive learning (SACL). SSPL explores image-audio positive pairs alone to discover semantically coherent similarities between audio and visual features, while a predictive coding module for feature alignment is introduced to facilitate the positive-only learning. In this regard SSPL acts as a negative-free method to eliminate false negatives. By contrast, SACL is designed to compact visual features and remove false negatives, providing reliable visual anchor and audio negatives for contrast. Different from SSPL, SACL releases the potential of audio-visual contrastive learning, offering an effective alternative to achieve the same goal. Comprehensive experiments demonstrate the superiority of our approach over the state-of-the-arts. Furthermore, we highlight the versatility of the learned representation by extending the approach to audio-visual event classification and object detection tasks. Zengjie Song, Jiangshe Zhang 0001, Yuxi Wang 0001, Junsong Fan, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Hard Patches Mining for Masked Image ModelingabstractMasked image modeling (MIM) has attracted much research attention due to its promising potential for learning scalable visual representations. In typical approaches, models usually focus on predicting specific contents of masked patches, and their performances are highly related to pre-defined mask strategies. Intuitively, this procedure can be considered as training a student (the model) on solving given problems (predict masked patches). However, we argue that the model should not only focus on solving given problems, but also stand in the shoes of a teacher to produce a more challenging problem by itself. To this end, we propose Hard Patches Mining (HPM), a brand-new framework for MIM pre-training. We observe that the reconstruction loss can naturally be the metric of the difficulty of the pretraining task. Therefore, we introduce an auxiliary loss predictor, predicting patch-wise losses first and deciding where to mask next. It adopts a relative relationship learning strategy to prevent overfitting to exact reconstruction loss values. Experiments under various settings demonstrate the effectiveness of HPM in constructing masked images. Furthermore, we empirically find that solely introducing the loss prediction objective leads to powerful representations, verifying the efficacy of the ability to be aware of where is hard to reconstruct.11Code: https://github.com/Haochen-wang409/HPM Kaiyou Song, Junsong Fan, Yuxi Wang 0001, Zhaoxiang Zhang 0001 |
CVPR | 3 |
| 2023 | DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) is a practical yet challenging task. Due to large-scale datasets, most existing methods use a network pretrained in other datasets to extract features, which are not suitable enough for WTAL. To address this problem, researchers design several modules for feature enhancement, which improve the performance of the localization module, especially modeling the temporal relationship between snippets. However, all of them omit that ambiguous snippets deliver contradictory information, which would reduce the discriminability of linked snippets. Considering this phenomenon, we propose Discriminability-Driven Graph Network (DDG-Net), which explicitly models ambiguous snippets and discriminative snippets with well-designed connections, preventing the transmission of ambiguous information and enhancing the discriminability of snippet-level representations. Additionally, we propose feature consistency loss to prevent the assimilation of features and drive the graph convolution network to generate more discriminative representations. Extensive experiments on THUMOS14 and ActivityNet1.2 benchmarks demonstrate the effectiveness of DDG-Net, establishing new state-of-the-art results on both datasets. Source code is available at https://github.com/XiaojunTang22/ICCV2023-DDGNet. Junsong Fan, Chuanchen Luo, Zhaoxiang Zhang 0001, Man Zhang 0005, Zongyuan Yang |
ICCV | 2 |
| 2023 | DropPos: Pre-Training Vision Transformers by Reconstructing Dropped PositionsabstractAs it is empirically observed that Vision Transformers (ViTs) are quite insensitive to the order of input tokens, the need for an appropriate self-supervised pretext task that enhances the location awareness of ViTs is becoming evident. To address this, we present DropPos, a novel pretext task designed to reconstruct Dropped Positions. The formulation of DropPos is simple: we first drop a large random subset of positional embeddings and then the model classifies the actual position for each non-overlapping patch among all possible positions solely based on their visual appearance. To avoid trivial solutions, we increase the difficulty of this task by keeping only a subset of patches visible. Additionally, considering there may be different patches with similar visual appearances, we propose position smoothing and attentive reconstruction strategies to relax this classification problem, since it is not necessary to reconstruct their exact positions in these cases. Empirical evaluations of DropPos show strong capabilities. DropPos outperforms supervised pre-training and achieves competitive results compared with state-of-the-art self-supervised alternatives on a wide range of downstream benchmarks. This suggests that explicitly encouraging spatial reasoning abilities, as DropPos does, indeed contributes to the improved location awareness of ViTs. The code is publicly available at https://github.com/Haochen-Wang409/DropPos. Junsong Fan, Yuxi Wang 0001, Kaiyou Song, Zhaoxiang Zhang 0001 |
NeurIPS | 2 |
| 2023 | Toward Practical Weakly Supervised Semantic Segmentation via Point-Level Supervision
Junsong Fan, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 1 |
| 2023 | Memory-Based Cross-Image Contexts for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) trains segmentation models by only weak labels, aiming to save the burden of expensive pixel-level annotations. This paper tackles the WSSS problem of utilizing image-level labels as the weak supervision. Previous approaches address this problem by focusing on generating better pseudo-masks from weak labels to train the segmentation model. However, they generally only consider every single image and overlook the potential cross-image contexts. We emphasize that the cross-image contexts among a group of images can provide complementary information for each other to obtain better pseudo-masks. To effectively employ cross-image contexts, we develop an end-to-end cross-image context module containing a memory bank mechanism and a transformer-based cross-image attention module. The former extracts cross-image contexts online from the feature encodings of input images and stores them as the memory. The latter mines useful information from the memorized contexts to provide the original queries with additional information for better pseudo-mask generation. We conduct detailed experiments on the Pascal VOC 2012 and the COCO dataset to demonstrate the advantage of utilizing cross-image contexts. Besides, state-of-the-art performance is also achieved. Codes are available at https://github.com/js-fan/MCIC.git. Junsong Fan, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | MMT: Cross Domain Few-Shot Learning via Meta-Memory TransferabstractFew-shot learning aims to recognize novel categories solely relying on a few labeled samples, with existing few-shot methods primarily focusing on the categories sampled from the same distribution. Nevertheless, this assumption cannot always be ensured, and the actual domain shift problem significantly reduces the performance of few-shot learning. To remedy this problem, we investigate an interesting and challenging cross-domain few-shot learning task, where the training and testing tasks employ different domains. Specifically, we propose a Meta-Memory scheme to bridge the domain gap between source and target domains, leveraging style-memory and content-memory components. The former stores intra-domain style information from source domain instances and provides a richer feature distribution. The latter stores semantic information through exploration of knowledge of different categories. Under the contrastive learning strategy, our model effectively alleviates the cross-domain problem in few-shot learning. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on cross-domain few-shot semantic segmentation tasks on the COCO-20$^{i}$, PASCAL-5$^{i}$, FSS-1000, and SUIM datasets and positively affects few-shot classification tasks on Meta-Dataset. Wenjian Wang 0002, Lijuan Duan, Yuxi Wang 0001, Junsong Fan, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Joint Power and 3D Trajectory Optimization for UAV-Enabled Wireless Powered Communication Networks With ObstaclesabstractUnmanned aerial vehicle (UAV)-enabled wireless powered communication networks (WPCNs) are promising technologies in 5G/6G wireless communications, while there are several challenges about UAV power allocation and scheduling to enhance the energy utilization efficiency, considering the existence of obstacles. In this work, we consider a UAV-enabled WPCN scenario that a UAV needs to cover the ground wireless devices (WDs). During the coverage process, the UAV needs to collect data from the WDs and charge them simultaneously. To this end, we formulate a joint-UAV power and three-dimensional (3D) trajectory optimization problem (JUPTTOP) to simultaneously increase the total number of the covered WDs, increase the time efficiency, and reduce the total flying distance of UAV so as to improve the energy utilization efficiency in the network. Due to the difficulties and complexities, we decompose it into two sub optimization problems, which are the UAV power allocation optimization problem (UPAOP) and UAV 3D trajectory optimization problem (UTTOP), respectively. Then, we propose an improved non-dominated sorting genetic algorithm-II with$K$-means initialization operator and Variable dimension mechanism (NSGA-II-KV) for solving the UPAOP. For UTTOP, we first introduce a pretreatment method, and then use an improved particle swarm optimization with Normal distribution initialization, Genetic mechanism, Differential mechanism and Pursuit operator (PSO-NGDP) to deal with this sub optimization problem. Simulation results verify the effectiveness of the proposed strategies under different scales and settings of the networks. Hongyang Pan, Yanheng Liu 0001, Geng Sun 0001, Junsong Fan, Shuang Liang 0003, Chau Yuen |
IEEE Trans. Commun. | 4 |
| 2023 | Coarse Mask Guided Interactive Object SegmentationabstractInteractive object segmentation aims to produce object masks with user interactions, such as clicks, bounding boxes, and scribbles. Click point is the most popular interactive cue for its efficiency, and related deep learning methods have attracted lots of interest in recent years. Most works encode click points as gaussian maps and concatenate them with images as the model's input. However, the spatial and semantic information of gaussian maps would be noised through multiple convolution layers and won't be fully exploited by top layers for mask prediction. To pass click information to top layers exactly and efficiently, we propose a coarse mask guided model (CMG) which predicts coarse masks with a coarse module to guide the object mask prediction. Specifically, the coarse module encodes user clicks as query features and enriches their semantic information with backbone features through transformer layers, coarse masks are generated based on the enriched query feature and fed into CMG's decoder. Benefiting from the efficiency of transformer, CMG's coarse module and decoder module are lightweight and computationally efficient, making the interaction process more smooth. Experiments on several segmentation benchmarks demonstrate the effectiveness of our method, and we get new state-of-the-art results compared with previous works. Jing Li 0112, Junsong Fan, Yuxi Wang 0001, Yuran Yang, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Towards Noiseless Object Contours for Weakly Supervised Semantic SegmentationabstractImage-level label based weakly supervised semantic segmentation has attracted much attention since image labels are very easy to obtain. Existing methods usually generate pseudo labels from class activation map (CAM) and then train a segmentation model. CAM usually highlights partial objects and produce incomplete pseudo labels. Some methods explore object contour by training a contour model with CAM seed label supervision and then propagate CAM score from discriminative regions to nondiscriminative regions with contour guidance. The propagation process suffers from the noisy intra-object contours, and inadequate propagation results produce incomplete pseudo labels. This is because the coarse CAM seed label lacks sufficient precise semantic information to suppress contour noise. In this paper, we train a SANCE model which utilizes an auxiliary segmentation module to supplement high-level semantic information for contour training by backbone feature sharing and online label supervision. The auxiliary segmentation module also provides more accurate localization map than CAM for pseudo label generation. We evaluate our approach on Pascal VOC 2012 and MS COCO 2014 benchmarks and achieve stateof- the-art performance, demonstrating the effectiveness of our method. The source code can be found at https://github.com/BraveGroup/SANCE Jing Li 0112, Junsong Fan, Zhaoxiang Zhang 0001 |
CVPR | 2 |
| 2022 | Remember the Difference: Cross-Domain Few-Shot Semantic Segmentation via Meta-Memory TransferabstractFew-shot semantic segmentation intends to predict pixel-level categories using only a few labeled samples. Existing few-shot methods focus primarily on the categories sampled from the same distribution. Nevertheless, this assumption cannot always be ensured. The actual domain shift problem significantly reduces the performance of few-shot learning. To remedy this problem, we propose an interesting and challenging cross-domain few-shot semantic segmentation task, where the training and test tasks perform on different domains. Specifically, we first propose a meta-memory bank to improve the generalization of the segmentation network by bridging the domain gap between source and target domains. The meta-memory stores the intra-domain style information from source domain instances and transfers it to target samples. Subsequently, we adopt a new contrastive learning strategy to explore the knowledge of different categories during the training stage. The negative and positive pairs are obtained from the proposed memory-based style augmentation. Comprehensive experiments demon-strate that our proposed method achieves promising results on cross-domain few-shot semantic segmentation tasks on COCO-20i, PASCAL-Si, FSS-1000, and SUIM datasets. Wenjian Wang 0002, Lijuan Duan, Yuxi Wang 0001, Qing En, Junsong Fan, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2022 | Pointly-Supervised Panoptic Segmentation
Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan |
ECCV (30) | 1 |
| 2022 | Interact with Open Scenes: A Life-long Evolution Framework for Interactive Segmentation ModelsabstractExisting interactive segmentation methods mainly focus on optimizing user interacting strategies, as well as making better use of clicks provided by users. However, the intention of the interactive segmentation model is to obtain high-quality masks with limited user interactions, which are supposed to be applied to unlabeled new images. But most existing methods overlooked the generalization ability of their models when witnessing new target scenes. To overcome this problem, we propose a life-long evolution framework for interactive models in this paper, which provides a possible solution for dealing with dynamic target scenes with one single model. Given several target scenes and an initial model trained with labels on the limited closed dataset, our framework arranges sequentially evolution steps on each target set. Specifically, we propose an interactive-prototype module to generate and refine pseudo masks, and apply a feature alignment module in order to adapt the model to a new target scene and keep the performance on previous images at the same time. All evolution steps above do not require ground truth labels as supervision. We conduct thorough experiments on PASCAL VOC, Cityscapes, and COCO datasets, demonstrating the effectiveness of our framework in solving new target datasets and maintaining performance on previous scenes at the same time. Ruitong Gan, Junsong Fan, Yuxi Wang 0001, Zhaoxiang Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | 3D Position Scheduling of UAV Secure Communications with Multiple ConstraintsabstractUnmanned aerial vehicle (UAV) communication is a promising technology in 5G/6G wireless communications. However, there are several challenges for ensuring secure communications in practical scenarios. In this paper, we consider a UAV-enabled communication scenario that a UAV needs to maintain secure communication with the ground communication nodes (GCNs), subject to the known ground eavesdropping nodes (GENs). UAV needs to select optimal communication positions and avoid obstacles. We formulate a UAV secrecy scheduling optimization problem (USSOP) to maximize the average secrecy rate and the minimum secrecy rate jointly. Then, we propose a particle swarm optimization with $\underline {normal}$ distribution initialization, $\underline {differential}$ mechanism and $\underline {avoiding}$ obstacles operator (PSONDA) to solve the USSOP. Simulation results show that this method performs better than other comparison algorithms. Junsong Fan, Yanheng Liu 0001, Geng Sun 0001, Hongyang Pan, Aimin Wang 0001, Shuang Liang 0003 |
SMC | 1 |
| 2022 | Toward few-shot domain adaptation with perturbation-invariant representation and transferable prototypes
Junsong Fan, Yuxi Wang 0001, He Guan, Chunfeng Song, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 1 |
| 2022 | Multimodal channel-wise attention transformer inspired by multisensory integration mechanisms of the brain
Qianqian Shi 0001, Junsong Fan, Zuoren Wang, Zhaoxiang Zhang 0001 |
Pattern Recognit. | 2 |
| 2020 | CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation with only image-level labels saves large human effort to annotate pixel-level labels. Cutting-edge approaches rely on various innovative constraints and heuristic rules to generate the masks for every single image. Although great progress has been achieved by these methods, they treat each image independently and do not take account of the relationships across different images. In this paper, however, we argue that the cross-image relationship is vital for weakly supervised segmentation. Because it connects related regions across images, where supplementary representations can be propagated to obtain more consistent and integral regions. To leverage this information, we propose an end-to-end cross-image affinity module, which exploits pixel-level cross-image relationships with only image-level labels. By means of this, our approach achieves 64.3% and 65.3% mIoU on Pascal VOC 2012 validation and test set respectively, which is a new state-of-the-art result by only using image-level labels for weakly supervised semantic segmentation, demonstrating the superiority of our approach. Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan, Chunfeng Song, Jun Xiao 0005 |
AAAI | 1 |
| 2020 | Learning Integral Objects With Intra-Class Discriminator for Weakly-Supervised Semantic SegmentationabstractImage-level weakly-supervised semantic segmentation (WSSS) aims at learning semantic segmentation by adopting only image class labels. Existing approaches generally rely on class activation maps (CAM) to generate pseudo-masks and then train segmentation models. The main difficulty is that the CAM estimate only covers partial foreground objects. In this paper, we argue that the critical factor preventing to obtain the full object mask is the classification boundary mismatch problem in applying the CAM to WSSS. Because the CAM is optimized by the classification task, it focuses on the discrimination across different image-level classes. However, the WSSS requires to distinguish pixels sharing the same image-level class to separate them into the foreground and the background. To alleviate this contradiction, we propose an efficient end-to-end Intra-Class Discriminator (ICD) framework, which learns intra-class boundaries to help separate the foreground and the background within each image-level class. Without bells and whistles, our approach achieves the state-of-the-art performance of image label based WSSS, with mIoU 68.0% on the VOC 2012 semantic segmentation benchmark, demonstrating the effectiveness of the proposed approach. Junsong Fan, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
CVPR | 1 |
| 2020 | Employing Multi-estimations for Weakly-Supervised Semantic Segmentation
Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan |
ECCV (17) | 1 |
| 2014 | Improved Biogeography-Based Optimization approach to secondary protein predictionabstractIn recent years, many bio-inspired computation algorithms have been proposed to solve constraint problems. Biogeography-Based Optimization (BBO) is one of these newly proposed optimization algorithms. As a new way to solve complicated optimization problems, BBO has a quick convergence. In this paper, we proposed an improved BBO for solving protein structure prediction problems. Comparative experiments with standard BBO and differential evolution algorithm (DE) are also conducted, and the results demonstrate this improved BBO approach performs better in solving these complicated protein prediction problems. Junsong Fan, Haibin Duan, Guangming Xie |
IJCNN | 1 |