EDBT 2026 Demo / reviewers in the wild / expert
Yupei Wang
dblp:208/4142
· DBLP profile ↗
23ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-9771-6229ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Earth-Adapter: Bridge the Geospatial Domain Gaps with a Frequency-Guided Mixture of AdaptersabstractVision Foundation Models (VFMs), while powerful, often struggle in Remote Sensing (RS) segmentation tasks when combined with existing Parameter-Efficient Fine-Tuning (PEFT) methods. We observe that this limitation primarily arises from their inability to effectively handle the pervasive artifacts in RS imagery. To address this, we introduce Earth-Adapter, the first PEFT method specifically designed for RS artifact mitigation. Earth-Adapter introduces a novel Frequency-Guided Mixture of Adapters (MoA) approach, structured around a ''divide and conquer" strategy. It first utilizes Discrete Fourier Transformation (DFT) to "divide" features into distinct frequency components, thereby effectively isolating artifact-related information from semantic signals. Subsequently, to ''conquer" these artifact, MoA independently optimizes features within different subspaces and dynamically assigns weights via a router to aggregate the refined representations. This enables adaptive refinement of the VFM’s representation space to mitigate the impact of artifacts. This simple yet highly effective PEFT method demonstrably mitigates artifacts and significantly enhances VFMs performance on RS segmentation tasks. Extensive experiments demonstrate Earth-Adapter's effectiveness on in-domain semantic segmentation (SS), as well as Domain Adaptive (DA) and Domain Generalized (DG) semantic segmentation tasks. Compared with the baseline Rein, Earth-Adapter significantly improves mIoU by 1.2% in SS, 9.0% in DA, and 3.1% in DG benchmarks. Our code and weights will be released soon. Xiaoxing Hu, Ziyang Gong, Yupei Wang, Yuru Jia, Fei Lin 0005, Dexiang Gao, Jianhong Han, Gen Luo, Xue Yang 0005 |
AAAI | 3 |
| 2026 | Style-adaptive detection transformer for single-source domain generalized object detection
Jianhong Han, Yupei Wang, Liang Chen 0004 |
Neurocomputing | 2 |
| 2026 | DFANet: Prototype-driven dynamic feature alignment for natural-to-remote-sensing cross-domain few-shot semantic segmentation
Yupei Wang |
Pattern Recognit. Lett. | 3 |
| 2025 | Decoupled Global-Local Alignment for Improving Compositional UnderstandingabstractContrastive Language-Image Pre-training (CLIP) has achieved success on multiple downstream tasks by aligning image and text modalities. However, the nature of global contrastive learning limits CLIP's ability to comprehend compositional concepts, such as relations and attributes. Although recent studies employ global hard negative samples to improve compositional understanding, these methods significantly compromise the model's inherent general capabilities by forcibly distancing textual negative samples from images in the embedding space. To overcome this limitation, we introduce a Decoupled Global-Local Alignment (DeGLA) framework that improves compositional understanding while substantially mitigating losses in general capabilities. To optimize the retention of the model's inherent capabilities, we incorporate a self-distillation mechanism within the global alignment process, aligning the learnable image-text encoder with a frozen teacher model derived from an exponential moving average. Under the constraint of self-distillation, it effectively mitigates the catastrophic forgetting of pretrained knowledge during fine-tuning. To improve compositional understanding, we first leverage the in-context learning capability of Large Language Models (LLMs) to construct about 2M high-quality negative captions across five types. Subsequently, we propose the Image-Grounded Contrast (IGC) loss and Text-Grounded Contrast (TGC) loss to enhance vision-language compositionally. Experimental results across both general and compositional reasoning tasks validate the effectiveness of the DeGLA framework. Our code is released at https://github.com/xiaoxing2001/DeGLA. Xiaoxing Hu, Kaicheng Yang 0002, Ziyong Feng, Yupei Wang |
ACM Multimedia | 6 |
| 2025 | Boosting Domain Generalization in Remote Sensing Image Segmentation via Style Mapping and General Prototypical Contrast
Yupei Wang, Xiaoxing Hu, Yongkang Hu, Shanghang Zhang, Liang Chen 0004 |
Int. J. Comput. Vis. | 1 |
| 2025 | Soft-Guided Open-Vocabulary Semantic Segmentation of Remote Sensing ImagesabstractOpen-vocabulary remote sensing semantic segmentation strives to assign both seen and unseen class labels to individual pixels in remote sensing images. Existing models follow the “fine-tune” paradigm based on Vision-Language Models (VLMs). However, as VLMs are predominantly tailored to natural scenes, these directly fine-tuned models often collapse into the seen categories and show insensitivity in perceiving remote sensing semantic cues. This critical issue of model collapse is closely related with the miss-alignment between image and text, making them struggle with the unique challenges of remote sensing images, such as complex and diverse scenes, and objects with significant scale differences. To this end, we propose a soft-guided open-vocabulary remote sensing semantic segmentation framework, which is the first to explore how to softly adapt VLMs to the downstream task of semantic segmentation for remote sensing images. Concretely, instead of directly fine-tuning, we introduce a generalization compensation strategy, which employs an additional frozen VLM encoder to provide implicit semantic guidance for dynamic optimization of visual representation. By introducing prior knowledge from frozen encoder, this soft strategy compensates potential losses incurred during fine-tuning, thus enhancing the model’s pixel-level perceptual alignment while avoiding model collapse. Afterwards, to optimize the sensitivity of VLMs’ textual and visual embeddings to remote sensing semantic information, bias-guided image-text collaborative optimization is presented to achieve a bilateral interaction of semantic information with the guidance of RS-Bias. Finally, an improved upsampling decoder is employed to obtain the progressive refinement and calibration of cost map through the integration of multi-scale information and textual embeddings. Extensive experiments demonstrate that our method achieves state-of-the-art performance on widely used challenging benchmarks. Code is available at https://github.com/H1NATA111/SGSeg.git. Yupei Wang, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | DATR: Unsupervised Domain Adaptive Detection Transformer With Dataset-Level Adaptation and Prototypical AlignmentabstractWith the success of the DEtection TRansformer (DETR), numerous researchers have explored its effectiveness in addressing unsupervised domain adaptation tasks. Existing methods leverage carefully designed feature alignment techniques to align the backbone or encoder, yielding promising results. However, effectively aligning instance-level features within the unique decoder structure of the detector has largely been neglected. Related techniques primarily align instance-level features in a class-agnostic manner, overlooking distinctions between features from different categories, which results in only limited improvements. Furthermore, the scope of current alignment modules in the decoder is often restricted to a limited batch of images, failing to capture the dataset-level cues, thereby severely constraining the detector's generalization ability to the target domain. To this end, we introduce a strong DETR-based detector named Domain Adaptive detection TRansformer (DATR) for unsupervised domain adaptation of object detection. First, we propose the Class-wise Prototypes Alignment (CPA) module, which effectively aligns cross-domain features in a class-aware manner by bridging the gap between the object detection task and the domain adaptation task. Then, the designed Dataset-level Alignment Scheme (DAS) explicitly guides the detector to achieve global representation and enhance inter-class distinguishability of instance-level features across the entire dataset, which spans both domains, by leveraging contrastive learning. Moreover, DATR incorporates a mean-teacher-based self-training framework, utilizing pseudo-labels generated by the teacher model to further mitigate domain bias. Extensive experimental results demonstrate superior performance and generalization capabilities of our proposed DATR in multiple domain adaptation scenarios. Code is released at https://github.com/h751410234/DATR. Liang Chen 0004, Jianhong Han, Yupei Wang |
IEEE Trans. Image Process. | 3 |
| 2024 | Frequency Spectrum Features Modeling for Real-Time Tiny Object Detection in Remote Sensing ImageabstractRecently, object detection in remote sensing images has achieved rapid advancement. However, due to critical issues, such as low spatial resolution and complex background noises, it is still difficult to achieve satisfactory object detection performance for remote sensing images. For current widely used object detection methods, the feature resolution of the backbone network is decreased gradually with successive pooling operations. In this way, object spatial details are largely lost for the deeper feature layers, resulting in the difficulty of accurate object detection, especially for tiny objects. However, current methods fail to eliminate the adverse effects due to the loss of object details. To this end, considering that high-frequency information is more likely to be overlooked and high-frequency object details may be beneficial for detecting tiny objects, we propose to improve the previous spatial feature modeling pipeline with the learned features in the frequency domain. Specifically, discrete cosine transform (DCT) is first used to transform the original image into the frequency domain, obtaining the corresponding frequency spectrum features. We then utilize a dual-domain feature extraction (DFE) network based on a lightweight attention mechanism to align the features in two different domains. Finally, a domain synergy fusion (DSF) module is further employed to match and fuse the features in the spatial domain and the obtained features in the frequency domain. Extensive experimental results are obtained on the challenging remote sensing datasets, DIOR and DOTA. Experimental results show that our method can increase at least 2.9%${\mathbf {AP}}_{\mathbf {s}}^{\mathbf {50}}$in DIOR and 3.5%${\mathbf {AP}}_{\mathbf {s}}^{\mathbf {50}}$in DOTA compared to the new state-of-the-art methods, which effectively demonstrates the superiority of our proposed method. Zhaoyi Luo, Yupei Wang, Liang Chen 0004, Wenying Yang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Encouraging the Mutual Interact Between Dataset-Level and Image-Level Context for Semantic Segmentation of Remote Sensing ImageabstractRecently, semantic segmentation of remote sensing images has witnessed rapid advancement with the adoption of deep neural networks. Contextual cues, referring to the long-range correlation between pixels, are crucial for achieving accurate segmentation results, particularly for objects with less discriminative characteristics in these images. Currently, most studies are centered on incorporating contextual cues by aggregating context information at the dataset-level or image-level. However, current research often treats contextual cue modeling at the dataset-level and image-level as independent procedures, neglecting the intrinsic correlation between these two feature levels. Consequently, the obtained contextual cues are sub-optimal. This issue is particularly critical in the semantic segmentation of remote sensing images. To address this, we propose to encourage mutual interaction between dataset-level and image-level contextual cues. Firstly, we propose an interactive dataset-image context aggregation scheme to obtain complementary and consistent multi-level contextual cues. Additionally, we introduce a parallel feature interaction network that progressively extracts and fuses features across multiple layers, enabling effective integration of multi-level contexts. Furthermore, we introduce an enhanced shifted window-based cross-attention mechanism to augment model efficiency. Extensive experimental results on the widely used Vaihingen, GaoFen-2 and iSAID datasets effectively demonstrate the superiority of our proposed method over state-of-art methods. Yupei Wang, Liang Chen 0004 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Remote Sensing Teacher: Cross-Domain Detection Transformer With Learnable Frequency-Enhanced Feature Alignment in Remote Sensing ImageryabstractUnsupervised Domain Adaptation (UDA) is critical for remote sensing object detection in real applications, aiming to address the significant performance degradation issue caused by the domain gap between source and target domain. This method achieves cross-domain alignment by leveraging the unlabeled target domain data, thus avoiding the expensive annotation cost. However, existing works mainly cope with CNN-based object detectors, which are characterized by complex adversarial learning architecture and fail to accurately align the features in remote sensing images with sparsely allocated objects and inevitable background noise. Compared to CNN-based methods, the DEtection TRansformer (DETR) largely simplifies the object detection pipeline and demonstrates the great potential by its intrinsic characteristics of global relation modeling between any pixels. On this basis, we propose the first strong DETR-based baseline, Remote Sensing Teacher, for unsupervised domain adaptation in remote sensing object detection. Specifically, the Remote Sensing Teacher introduces an innovative Learnable Frequency-enhanced feature Alignment (LFA) module. Within this module, we initially transform the features into frequency space to simplify the attention solver and effectively capture domain-specific information. Subsequently, the module significantly enhances the global feature representations of sparsely allocated objects by using a lightweight attention mechanism. Following this, the module incorporates learnable filters with a gated mechanism, enabling selective alignment of features in noisy backgrounds. Additionally, the Remote Sensing Teacher employs a Self-adaptive Pseudo label Assigner (SPA) that can automatically adjust the class-wise confidence threshold according to the model’s learning status, thereby enabling the generation of high-quality pseudo-labels in scenarios with a long-tailed distribution. Leveraging these pseudo-labels further mitigates the domain bias of the detector by establishing alignment at the label level. Extensive experimental results demonstrate superior performance and generalization capabilities of our proposed Remote Sensing Teacher in multiple remote sensing adaptation scenarios. Code is released at https://github.com/h751410234/RemoteSensingTeacher. Jianhong Han, Yupei Wang, Liang Chen 0004, Zhaoyi Luo |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Advancing Realistic Precipitation Nowcasting With a Spatiotemporal Transformer-Based Denoising Diffusion ModelabstractRecent advances in deep learning have significantly improved the quality of precipitation nowcasting. Current approaches are either based on deterministic or generative models. Deterministic models perceive nowcasting as a spatiotemporal prediction task, relying on distance functions like L2-norm loss for training. While improving meteorological evaluation metrics, they inevitably produce blurry predictions with no reference value. In contrast, generative models aim to capture realistic precipitation distributions and generate nowcasting products by sampling within these distributions. However, designing a generative model that produces realistic samples satisfying meteorological evaluation indexes in real-time remains challenging, given the triple dilemma of generative learning: achieving high sample quality, mode coverage, and fast sampling simultaneously. Recently, diffusion models exhibit impressive sample quality but suffer from time-consuming sampling, severely hindering their application in nowcasting. Moreover, samples generated by the U-Net denoiser of current denoising diffusion model are prone to yield poor meteorological evaluation metrics such as CSI. To this end, we propose a spatiotemporal Transformer-based conditional diffusion model with rapid diffusion strategy. Concretely, we incorporate an adversarial mapping-based rapid diffusion strategy to overcome the time-consuming sampling process for standard diffusion models, enabling timely nowcasting. Additionally, a meticulously designed spatiotemporal Transformer-based denoiser is incorporated into diffusion models, remedying the defects in U-Net denoisers by estimating diffusion scores and improving nowcasting skill scores. Case studies of typical weather events such as thunderstorms, as well as quantitative indicators, demonstrate the effectiveness of the proposed method in generating sharper and more precise precipitation forecasts while maintaining satisfied meteorological evaluation metrics. Zewei Zhao, Xichao Dong, Yupei Wang, Cheng Hu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MDTNet: Multiscale Deformable Transformer Network With Fourier Space Losses Toward Fine-Scale Spatiotemporal Precipitation NowcastingabstractDeep learning (DL)-based precipitation nowcasting algorithms have garnered significant attention in recent years. However, the presence of variable spatial scales in precipitation patterns poses challenges for methods that solely focus on capturing spatiotemporal correlations at a single scale. Moreover, current DL-based algorithms tend to model short-term (e.g., 10-min time span) rainfall locally neglecting long-term, global (e.g., 2-h time span) life-cycle evolution. Furthermore, widely used pixel-wise losses are prone to produce low effective-spatial-resolution predictions. To this end, we introduce a multiscale deformable transformer network to leverage echo contexts from image patches of varying spatial scales. Meanwhile, a multihead deformable self-attention mechanism is introduced for capturing precipitation spatiotemporal dynamics in a global manner. Moreover, to improve the spatial resolution of predictions, the Fourier space regularization and adversarial losses are proposed by narrowing the discrepancy of the Fourier spectra of predictions and references. Thanks to the introduced loss function, our model generates highly effective spatial-resolution predictions with abundant details. Extensive experiments on two real datasets show the substantial superiority of our method in terms of critical success index (CSI) compared to recent competitive approaches. At the same time, our predictions have more realistic precipitation details and significantly better fidelity. For example, on a vertically integrated liquid (VIL) product dataset, compared to baseline methods, our approach reduces the Fréchet inception distance (FID) value by a factor of$2\sim 4$while improves the CSI score by 3%~5% approximately. Zewei Zhao, Xichao Dong, Yupei Wang, Jianping Wang 0003, Yubao Chen, Cheng Hu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Geometric Boundary Guided Feature Fusion and Spatial-Semantic Context Aggregation for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images aims to achieve pixel-level semantic category assignment for input images. This task has achieved significant advances with the rapid development of deep neural network. Most current methods mainly focus on effectively fusing the low-level spatial details and high-level semantic cues. Other methods also propose to incorporate the boundary guidance to obtain boundary preserving segmentation. However, current methods treat the multi-level feature fusion and the boundary guidance as two separate tasks, resulting in sub-optimal solutions. Moreover, due to the large inter-class difference and small intra-class consistency within remote sensing images, current methods often fail to accurately aggregate the long-range contextual cues. These critical issues make current methods fail to achieve satisfactory segmentation predictions, which severely hinder downstream applications. To this end, we first propose a novel boundary guided multi-level feature fusion module to seamlessly incorporate the boundary guidance into the multi-level feature fusion operations. Meanwhile, in order to further enforce the boundary guidance effectively, we employ a geometric-similarity-based boundary loss function. In this way, under the explicit guidance of boundary constraint, the multi-level features are effectively combined. In addition, a channel-wise correlation guided spatial-semantic context aggregation module is presented to effectively aggregate the contextual cues. In this way, subtle but meaningful contextual cues about pixel-wise spatial context and channel-wise semantic correlation are effectively aggregated, leading to spatial-semantic context aggregation. Extensive qualitative and quantitative experimental results on ISPRS Vaihingen and GaoFen-2 datasets demonstrate the effectiveness of the proposed method. Yupei Wang, Yongkang Hu, Xiaoxing Hu, Liang Chen 0004, Shanqing Hu |
IEEE Trans. Image Process. | 1 |
| 2022 | SAR-to-Optical Image Translating Through Generate-Validate Adversarial NetworksabstractSynthetic aperture radar (SAR) has the advantages of high resolution in all-weather and all-day. However, SAR images are hard to be understood, due to their unique imaging mechanism. The SAR to optical image translation can assist in interpreting and has become a topic of growing interest in the field of remote sensing. In this letter, a SAR to optical image translation network is proposed, called generate-validate adversarial networks (GVANs). More specifically, there are two Pix2Pix networks form the cyclic structure. The validate module is employed to increase the training process and improve the edge retention ability. In order to improve multidomain images adaptability, the embedded layer is proposed. Additionally, the dilation convolution layer is employed in the generator, which is more suitable for the characteristics of SAR images. The proposed method has experimented on the SEN1-2 dataset. The result demonstrates the superiority of the proposed method over state-of-the-art methods. Hao Shi 0006, Bocheng Zhang, Yupei Wang, Zihan Cui, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Dual-Path Sparse Hierarchical Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images aims to label every pixel with the correct semantic category. The core challenge of the current deep convolutional network (ConvNet)-based methods lies in the difficulty of effectively aggregating high-level categorical semantics and low-level local details along the hierarchy of backbone. Most current approaches consider only fusing adjacent feature layers gradually with short-range feature connections, which lack the diversity of feature interactions, such as long-range cross-scale connections. To this end, we propose a novel dual-path sparse hierarchical network that is characterized by rich cross-scale feature interactions. Multiscale features are first sparsely grouped with a predefined interval, which is then aggregated via both long-range and short-range cross-scale connections in a hierarchical manner. Moreover, in order to further enrich the diversity of feature interactions, we also introduce another fusion path in parallel but with different sparsity for feature grouping, forming a dual-path network. In this way, our model is able to effectively aggregate multilevel features by incorporating both long-range and short-range feature interactions in both parallel and hierarchical manner. Meanwhile, the semantic and resolution gap between multilevel features can also be bridged. Yupei Wang, Hao Shi 0006, Shan Dong, Yin Zhuang, Liang Chen 0004 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Motion-Guided Global-Local Aggregation Transformer Network for Precipitation NowcastingabstractNowadays deep learning based weather radar echo extrapolation methods have competently improved nowcasting quality. Current pure convolutional or convolutional recurrent neural network based extrapolation pipelines inherently struggle in capturing both global and local spatiotemporal interactions simultaneously, thereby limiting nowcasting performances, e.g., they not only tend to underestimate heavy rainfalls’ spatial coverage and intensity but also fail to precisely predict non-linear motion patterns. Furthermore, the usually adopted pixel-wise objective functions lead to blurry predictions. To this end, we propose a novel motion-guided global-local aggregation Transformer network for effectively combining spatiotemporal cues at different time scales, thereby strengthening global-local spatiotemporal aggregation urgently required by the extrapolation task. First, we divide existing observations into both short and long term sequences to represent echo dynamics at different time scales. Then, to introduce reasonable motion guidance to Transformer, we customize an end-to-end module for jointly extracting Motion Representation of Short and Long term echo sequences (MRS, MRL), while estimating optical flow. Subsequently, based on Transformer architecture, MRS is used as queries to retrospect the most useful information from MRL for an effective aggregation of global long-term and local short-term cues. Finally, the fused feature is employed for future echo prediction. Additionally, for the blurry prediction problem, predictions from our model trained with an adversarial regularization achieve superior performances not only in nowcasting skill scores but also in precipitation details and image clarity over existing methods. Extensive experiments on two challenging radar echo datasets demonstrate the effectiveness of our proposed method. Xichao Dong, Zewei Zhao, Yupei Wang, Jianping Wang 0003, Cheng Hu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | FMCW Radar-Based Hand Gesture Recognition Using Spatiotemporal Deformable and Context-Aware Convolutional 5-D Feature RepresentationabstractRecently, frequency-modulated continuous-wave (FMCW) radar-based hand gesture recognition (HGR) using deep learning has achieved favorable performance. However, many existing methods use extracted features separately, i.e., using one of the range, Doppler, azimuth, or elevation angle information, or a combination of any two, to train convolutional neural networks (CNNs), which ignore the interrelation among the 5-D time-varying-range-Doppler-azimuth-elevation feature space. Although there have been methods using the 5-D information, their mining of the interrelation among the 5-D feature space is not sufficient, and there is still room for improvements. This article proposes a new processing scheme of HGR based on 5-D feature cubes that are jointly encoded by a 3-D fast Fourier transform (3-D-FFT)-based method. Then, a CNN is proposed by building two novel blocks, i.e., the spatiotemporal deformable convolution (STDC) block and the adaptive spatiotemporal context-aware convolution (ASTCAC) block. Concretely, STDC is designed to cope with hand gestures’ large spatiotemporal geometric transformations in the 5-D feature space. Moreover, ASTCAC is designed for modeling long-distance global relationships, e.g., relationships between pixels of the feature at the upper left corner and lower right corner, and exploring the global spatiotemporal context, in order to enhance the target feature representation and suppress interference. Finally, our presented method is verified on a large radar dataset, including 19 760 sets of 16 common hand gestures, collected by 19 subjects. Our method obtains a recognition rate of 99.53% on the validation dataset and that of 97.22% on the test dataset, which is significantly better than state-of-the-art methods. Xichao Dong, Zewei Zhao, Yupei Wang, Tao Zeng 0001, Jianping Wang 0003, Yi Sui 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | SSAP: Single-Shot Instance Segmentation With Affinity PyramidabstractProposal-free instance segmentation methods mainly generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, previous methods mostly employ separate modules for these two sub-tasks and require multiple passes for inference. In addition to the lack of efficiency, previous methods also failed to perform as well as proposal-based approaches. To this end, this work proposes a single-shot proposal-free instance segmentation method that requires only one single pass for prediction. Our method is based on learning an affinity pyramid, which computes the probability that two pixels belong to the same instance in a hierarchical manner. Moreover, incorporating with the learned affinity pyramid, a novel cascaded graph partition (CGP) module is presented to fuse the two predictions and segment instances efficiently. As an additional contribution, we conduct an experiment to demonstrate the benefits of proposal-free methods in capturing detailed structures from finely annotated training examples. Our approach is evaluated on the Cityscapes and COCO datasets and achieves state-of-the-art performance. Naiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 0012, Kaiqi Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | SSAP: Single-Shot Instance Segmentation With Affinity PyramidabstractRecently, proposal-free instance segmentation has received increasing attention due to its concise and efficient pipeline. Generally, proposal-free methods generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, previous methods mostly employ separate modules for these two sub-tasks and require multiple passes for inference. We argue that treating these two sub-tasks separately is suboptimal. In fact, employing multiple separate modules significantly reduces the potential for application. The mutual benefits between the two complementary sub-tasks are also unexplored. To this end, this work proposes a single-shot proposal-free instance segmentation method that requires only one single pass for prediction. Our method is based on a pixel-pair affinity pyramid, which computes the probability that two pixels belong to the same instance in a hierarchical manner. The affinity pyramid can also be jointly learned with the semantic class labeling and achieve mutual benefits. Moreover, incorporating with the learned affinity pyramid, a novel cascaded graph partition module is presented to sequentially generate instances from coarse to fine. Unlike previous time-consuming graph partition methods, this module achieves 5× speedup and 9% relative improvement on Average-Precision (AP). Our approach achieves new state of the art on the challenging Cityscapes dataset. Naiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 0012, Yinan Yu, Ming Yang 0007, Kaiqi Huang |
ICCV | 3 |
| 2019 | Focal Boundary Guided Salient Object DetectionabstractThe performance of salient object segmentation has been significantly advanced by using deep convolutional networks. However, these networks often produce blob-like saliency maps without accurate object boundaries. This is caused by the limited spatial resolution of their feature maps after multiple pooling operations, and might hinder downstream applications that require precise object shapes. To address this issue, we propose a novel deep model-Focal Boundary Guided (Focal- BG) network. Our model is designed to jointly learn to segment salient object masks and detect salient object boundaries. Our key idea is that additional knowledge about object boundaries can help to precisely identify the shape of the object. Moreover, our model incorporates a refinement pathway to refine the mask prediction, and makes use of the focal loss to facilitate the learning of the hard boundary pixels. To evaluate our model, we conduct extensive experiments. Our Focal-BG network consistently outperforms state-of-the-art methods on five major benchmarks. We provide a detailed analysis of these results and demonstrate that our joint modeling of salient object boundary and mask helps to better capture shape details, especially in the vicinity of object boundaries. Yupei Wang, Xin Zhao 0012, Xuecai Hu, Yin Li 0003, Kaiqi Huang |
IEEE Trans. Image Process. | 1 |
| 2019 | Deep Crisp Boundaries: From Boundaries to Higher-Level TasksabstractEdge detection has made significant progress with the help of deep convolutional networks (ConvNet). These ConvNet-based edge detectors have approached human level performance on standard benchmarks. We provide a systematical study of these detectors' outputs. We show that the detection results did not accurately localize edge pixels, which can be adversarial for tasks that require crisp edge inputs. As a remedy, we propose a novel refinement architecture to address the challenging problem of learning a crisp edge detector using ConvNet. Our method leverages a top-down backward refinement pathway, and progressively increases the resolution of feature maps to generate crisp edges. Our results achieve superior performance, surpassing human accuracy when using standard criteria on BSDS500, and largely outperforming the state-of-the-art methods when using more strict criteria. More importantly, we demonstrate the benefit of crisp edge maps for several important applications in computer vision, including optical flow estimation, object proposal generation, and semantic segmentation. Yupei Wang, Xin Zhao 0012, Yin Li 0003, Kaiqi Huang |
IEEE Trans. Image Process. | 1 |
| 2018 | Densely Cascaded Shadow Detection Network via Deeply Supervised Parallel FusionabstractShadow detection is an important and challenging problem in computer vision. Recently, single image shadow detection had achieved major progress with the development of deep convolutional networks. However, existing methods are still vulnerable to background clutters, and often fail to capture the global context of an input image. These global contextual and semantic cues are essential for accurately localizing the shadow regions. Moreover, rich spatial details are required to segment shadow regions with precise shape. To this end, this paper presents a novel model characterized by a deeply supervised parallel fusion (DSPF) network and a densely cascaded learning scheme. The DSPF network achieves a comprehensive fusion of global semantic cues and local spatial details by multiple stacked parallel fusion branches, which are learned in a deeply supervised manner. Moreover, the densely cascaded learning scheme is employed to refine the spatial details. Our method is evaluated on two widely used shadow detection benchmarks. Experimental results show that our method outperforms state-of-the-arts by a large margin. Yupei Wang, Xin Zhao 0012, Yin Li 0003, Xuecai Hu, Kaiqi Huang |
IJCAI | 1 |
| 2017 | Deep Crisp BoundariesabstractEdge detection had made significant progress with the help of deep Convolutional Networks (ConvNet). ConvNet based edge detectors approached human level performance on standard benchmarks. We provide a systematical study of these detector outputs, and show that they failed to accurately localize edges, which can be adversarial for tasks that require crisp edge inputs. In addition, we propose a novel refinement architecture to address the challenging problem of learning a crisp edge detector using ConvNet. Our method leverages a top-down backward refinement pathway, and progressively increases the resolution of feature maps to generate crisp edges. Our results achieve promising performance on BSDS500, surpassing human accuracy when using standard criteria, and largely outperforming state-of-the-art methods when using more strict criteria. We further demonstrate the benefit of crisp edge maps for estimating optical flow and generating object proposals. Yupei Wang, Xin Zhao 0012, Kaiqi Huang |
CVPR | 1 |