VLDB 2026 Research / reviewers in the wild / expert
Wen Yang 0001
dblp:42/3814-1
· DBLP profile ↗
117ranked-venue papers
9as first author
51since 2021 · last 2026
0000-0002-3263-8768ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 72 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 29 · 1 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 10 since 2021Systems, architecture and hardware · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniCalib: Targetless LiDAR-camera Calibration via Probabilistic Flow on Unified Depth RepresentationsabstractOnline targetless extrinsic LiDAR-camera calibration is essential for robust perception in computer vision applications such as autonomous driving. However, existing methods struggle with the significant modality gap between heterogeneous sensors and fail to handle unreliable correspondences arising from real-world challenges like occlusions and dynamic objects. To address these issues, we introduce UniCalib, a novel method that performs calibration by estimating a probabilistic flow on unified depth representations. UniCalib first bridges the modality gap by converting both the camera images and the sparse LiDAR points into unified, dense depth maps, enabling a unified encoder to learn consistent features. Subsequently, it learns a probabilistic flow field that captures the correspondence uncertainty to improve robustness. This probabilistic approach is reinforced by a reliability map and a perceptually weighted sparse flow loss, which guide the model to suppress the influence of unreliable regions. Experimental results on three datasets validate the accuracy and generalization of UniCalib. In particular, it achieves a mean translation error of 0.550cm and a rotation error of 0.044° on the KITTI dataset. The code is available at https://github.com/han-15/UniCalib. Xubo Zhu, Ji Wu 0012, Ximeng Cai, Wen Yang 0001, Huai Yu, Gui-Song Xia |
WACV | 5 |
| 2026 | Change-aware multi-temporal cloud removal
Guochu You, Runmin Dong, Wen Yang 0001, Gui-Song Xia |
Sci. China Inf. Sci. | 5 |
| 2026 | Oriented Tiny Object Detection: A Dataset, Benchmark, and Dynamic Unbiased LearningabstractDetecting oriented tiny objects, which are limited in appearance information yet prevalent in real-world applications, remains an intricate and under-explored problem. To address this, we systematically introduce a new dataset, a benchmark, and a dynamic coarse-to-fine learning scheme in this study. Our proposed dataset, AI-TOD-R, features the smallest object sizes among all oriented object detection datasets. Based on AI-TOD-R, we present a benchmark spanning a broad range of detection paradigms, including both fully-supervised and label-efficient approaches. Through investigation, we identify a learning bias presents across various learning pipelines: confident objects become increasingly confident, while vulnerable oriented tiny objects are further marginalized, hindering their detection performance. To mitigate this issue, we propose a Dynamic Coarse-to-Fine Learning (DCFL) scheme towards unbiased learning. DCFL dynamically updates prior positions to better align with the limited areas of oriented tiny objects, and it assigns samples in a way that balances both quantity and quality across different object shapes, thus mitigating biases in prior settings and sample selection. Extensive experiments across 10 challenging object detection datasets demonstrate that DCFL achieves state-of-the-art accuracy, high efficiency, and remarkable versatility. Chang Xu 0027, Ruixiang Zhang, Wen Yang 0001, Jian Ding 0001, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | DN-TOD: Robust tiny object detection amidst label noise
Chang Xu 0027, Wen Yang 0001, Ruixiang Zhang, Yan Zhang 0115, Gui-Song Xia |
Pattern Recognit. | 3 |
| 2026 | QuadricsReg: Large-Scale Point Cloud Registration Using Semantic Quadric Primitives
Ji Wu 0012, Huai Yu, Ximeng Cai, Mingfeng Wang, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Robotics | 6 |
| 2025 | LiDAR-enhanced 3D Gaussian Splatting MappingabstractThis paper introduces LiGSM, a novel LiDARenhanced 3D Gaussian Splatting (3DGS) mapping framework that improves the accuracy and robustness of 3D scene mapping by integrating LiDAR data. LiGSM constructs joint loss from images and LiDAR point clouds to estimate the poses and optimize their extrinsic parameters, enabling dynamic adaptation to variations in sensor alignment. Furthermore, it leverages LiDAR point clouds to initialize 3DGS, providing a denser and more reliable starting points compared to sparse SfM points. In scene rendering, the framework augments standard image-based supervision with depth maps generated from LiDAR projections, ensuring an accurate scene representation in both geometry and photometry. Experiments on public and self-collected datasets demonstrate that LiGSM outperforms comparative methods in pose tracking and scene rendering. Huai Yu, Ji Wu 0012, Wen Yang 0001, Gui-Song Xia |
ICRA | 4 |
| 2025 | Self-supervised Shutter Unrolling with Events
Mingyuan Lin, Yangguang Wang, Xiang Zhang 0022, Boxin Shi, Wen Yang 0001, Chu He, Gui-Song Xia, Lei Yu 0006 |
Int. J. Comput. Vis. | 5 |
| 2025 | ReFocal: Addressing Learning Imbalances for Accurate Tiny Object Detection in Aerial ImageryabstractTiny objects in aerial imagery usually exhibit an extremely limited number of pixels, significantly affecting the object detection model’s learning process. While existing research has attempted to improve tiny objects’ positive sample quantity for scale-balanced learning, the primary focus lies on the object level. We argue that mitigating learning imbalance requires a comprehensive consideration encompassing object-level, sample-level, and feature-level improvements. To this end, we propose ReFocal, a learning strategy comprised of ReFocal Loss and ReFocal feature pyramid network (FPN), to mitigate imbalances across these three levels. ReFocal Loss utilizes a magnitude factor to regulate the learning magnitude of objects with varying sample counts and a novel focal rate adjuster to differentiate sample quality at the sample level, enabling the detector to prioritize high-quality samples within each object. ReFocal FPN employs a refocusing mechanism to dynamically enhance detailed information in high-level feature maps without introducing additional computational cost, thus addressing the feature-level imbalance. Extensive experiments on AI-TOD-v2 and TinyPerson datasets demonstrate the superiority of our proposed method over previous single-stage methods, particularly for very tiny objects. Zijuan Chen, Chang Xu 0027, Wen Yang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Detecting Every Object From EventsabstractObject detection is critical in autonomous driving, and it is more practical yet challenging to localize objects of unknown categories: an endeavour known as Class-Agnostic Object Detection (CAOD). Existing studies on CAOD predominantly rely on RGB cameras, but these frame-based sensors usually have high latency and limited dynamic range, leading to safety risks under extreme conditions like fast-moving objects, overexposure, and darkness. In this study, we turn to the event-based vision, featured by its sub-millisecond latency and high dynamic range, for robust CAOD. We propose Detecting Every Object in Events (DEOE), an approach aimed at achieving high-speed, class-agnostic object detection in event-based vision. Built upon the fast event-based backbone: recurrent vision transformer, we jointly consider the spatial and temporal consistencies to identify potential objects. The discovered potential objects are assimilated as soft positive samples to avoid being suppressed as backgrounds. Moreover, we introduce a disentangled objectness head to separate the foreground-background classification and novel object discovery tasks, enhancing the model's generalization in localizing novel objects while maintaining a strong ability to filter out the background. Extensive experiments confirm the superiority of our proposed DEOE in both open-set and closed-set settings, outperforming strong baseline methods. Haitian Zhang, Chang Xu 0027, Xinya Wang, Bingde Liu, Guang Hua 0001, Lei Yu 0006, Wen Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | STAR-CD: Style-Aligned Remote Sensing Change Detection With Appearance-Relation ModelingabstractRemote sensing change detection plays a crucial part in monitoring the evolution of land cover. Despite remarkable progress in recent years, existing change detection models still struggle with two major challenges that lead to inaccurate change predictions. For one, current models lack sufficient global relation modeling, resulting in coarse boundary prediction and false identification of unwanted variations. For another, bitemporal images often display divergent imaging styles due to lighting, sensor, or seasonal differences, and this cross-temporal style inconsistency would amplify the pseudo-changes caused by non-semantic variations. In this paper, we present the style-alignment and appearance-relation modeling change detection model (STAR-CD) to tackle the above challenges. First, we propose an appearance-relation modeling block (ARMB) to jointly explore semantic differences from both local textures and global structures of remote sensing images, effectively reducing false alarms and refining change boundary predictions. Second, we design a Fourier-inspired style alignment module (FSAM), which suppresses style-induced noise and reveals true changes by texture-style decoupling in the frequency domain. Extensive experiments on the WHU-CD, LEVIR-CD, and CDD datasets demonstrate the superiority of our STAR-CD over state-of-the-art methods. In particular, STAR-CD surpasses previous methods by 2.40% and 4.34% in F1 and IoU on the CDD datasets, respectively. Visual comparisons further validate its strength in predicting precise change boundaries and robustly suppressing pseudo-changes. Yan Zhang 0115, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Unsupervised Multiview UAV Image Geolocalization via Iterative RenderingabstractUnmanned Aerial Vehicle (UAV) Cross-View Geo-Localization (CVGL) poses significant challenges due to the substantial view discrepancies between oblique UAV images and overhead satellite images. Existing methods heavily rely on supervised learning with labeled datasets to extract viewpoint-invariant features for cross-view retrieval. However, these approaches are computationally expensive, prone to overfitting region-specific cues, and exhibit limited generalizability to new regions. To overcome this issue, we propose an unsupervised solution that lifts the scene representation to 3D space from UAV observations for satellite image generation, providing a robust representation against view distortion. By generating orthogonal images that closely resemble satellite views, our method reduces view discrepancies in feature representation and mitigates shortcuts in region-specific image pairing. To further align the perspective of the rendered image with the real one, we design an iterative camera pose updating mechanism that progressively modulates the rendered query image with potential satellite targets, eliminating spatial offsets relative to the reference images. Additionally, this iterative refinement strategy enhances cross-view feature invariance through view-consistent fusion across iterations. As such, our unsupervised paradigm naturally avoids the problem of region-specific overfitting, enabling generic CVGL for UAV images without feature fine-tuning or data-driven training. Experiments on the University-1652 and SUES-200 datasets demonstrate that our approach significantly improves geo-localization accuracy while maintaining robustness across diverse regions. Notably, without model fine-tuning or paired training, our method achieves competitive performance with recent supervised methods. Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Li Mi, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Minimizing Sample Redundancy for Label-Efficient Object Detection in Aerial ImagesabstractObjects in aerial images tend to be densely scattered and appear in arbitrary orientations, making the annotation process quite costly. To reduce the annotation cost, existing methods propose randomly annotating a proportion of images or objects for aerial object detection with fewer label usage. These approaches, however, can lead to redundancy in labels and inherit the biases associated with the imbalance in datasets. To minimize sample redundancy and alleviate data imbalance, we propose a novel labeling pattern that acquires heterogeneous object labels in a class-orthogonal manner, preserving a broader diversity of samples for each category with less annotation effort. To improve data utility, we design a Dynamic Multi-View Learning (DML) strategy to overcome the sample quantity-quality dilemma in current pseudo-labeling methods—a high pseudo-label threshold reduces sample quantity, while low thresholds compromise sample quality. First, DML separates model predictions into multiple hierarchies for finer screening, mitigating the suppression of unlabeled objects in binary pseudo-label strategies. With this separation, DML learns to construct a new view by injecting high-quality samples and masking low-quality regions in this view, simultaneously expanding sample quantity while ensuring sample quality. Unlike previous methods that mine pseudo labels solely from unlabelled regions, DML releases this constraint by learning to expand high-quality samples with a dynamic view. Extensive experiments on five benchmark datasets validate our method’s state-of-the-art accuracy and label efficiency. Notably, with approximately 5% DOTA-v2.0 annotations, DML achieves nearly 90% of the fully supervised performance. The codes will be available at https://github.com/ZhangRuixiang-WHU/ALOD_DML/. Ruixiang Zhang, Chang Xu 0027, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | ConGeo: Robust Cross-View Geo-Localization Across Ground View Variations
Li Mi, Chang Xu 0027, Javiera Castillo-Navarro, Syrielle Montariol, Wen Yang 0001, Antoine Bosselut, Devis Tuia |
ECCV (14) | 5 |
| 2024 | Poisoning-Free Defense Against Black-Box Model ExtractionabstractRecent research has shown that an adversary can use a surrogate model to steal the functionality of a target deep learning model even under the black-box condition and without data curation, while the existing defense mainly relies on API poisoning to disturb the surrogate training. Unfortunately, due to poisoning, the defense is achieved at the price of fidelity loss, sacrificing the interests of honest users. To solve this problem, we propose an Adversarial Fine-Tuning (AdvFT) framework, incorporating the generative adversarial network (GAN) structure that disturbs the feature representations of out-of-distribution (OOD) queries while preserving those of in-distribution (ID) ones, circumventing the need for OOD sample collection and API poisoning. Extensive experiments verify the effectiveness of the proposed framework. Code is available at github.com/Hatins/AdvFT. Haitian Zhang, Guang Hua 0001, Wen Yang 0001 |
ICASSP | 3 |
| 2024 | FE-DeTr: Keypoint Detection and Tracking in Low-quality Image Frames with EventsabstractKeypoint detection and tracking in traditional image frames are often compromised by image quality issues such as motion blur and extreme lighting conditions. Event cameras offer potential solutions to these challenges by virtue of their high temporal resolution and high dynamic range. However, they have limited performance in practical applications due to their inherent noise in event data. This paper advocates fusing the complementary information from image frames and event streams to achieve more robust keypoint detection and tracking. Specifically, we propose a novel keypoint detection network that fuses the textural and structural information from image frames with the high-temporal-resolution motion information from event streams, namely FE-DeTr. The network leverages a temporal response consistency for supervision, ensuring stable and efficient keypoint detection. Moreover, we use a spatio-temporal nearest-neighbor search strategy for robust keypoint tracking. Extensive experiments are conducted on a new dataset featuring both image frames and event data captured under extreme conditions. The experimental results confirm the superior performance of our method over both existing frame-based and event-based methods. Our code, pre-trained models, and dataset are available at https://github.com/yuyangpoi/FE-DeTr. Xiangyuan Wang, Kuangyi Chen, Wen Yang 0001, Lei Yu 0006, Yannan Xing, Huai Yu |
ICRA | 3 |
| 2024 | QuadricsNet: Learning Concise Representation for Geometric Primitives in Point CloudsabstractThis paper presents a novel framework to learn a concise geometric primitive representation for 3D point clouds. Different from representing each type of primitive individually, we focus on the challenging problem of how to achieve a concise and uniform representation robustly. We employ quadrics to represent diverse primitives with only 10 parameters and propose the first end-to-end learning-based framework, namely QuadricsNet, to parse quadrics in point clouds. The relationships between quadrics mathematical formulation and geometric attributes, including the type, scale and pose, are insightfully integrated for effective supervision of QuaidricsNet. Besides, a novel pattern-comprehensive dataset with quadrics segments and objects is collected for training and evaluation. Experiments demonstrate the effectiveness of our concise representation and the robustness of QuadricsNet. Our code is available at https://github.com/MichaelWu99-lab/QuadricsNet. Ji Wu 0012, Huai Yu, Wen Yang 0001, Gui-Song Xia |
ICRA | 3 |
| 2024 | AODet: Anti-Occlusion for Enhanced Small Object Detection in Drone-Based RGBT ImageryabstractDrone-based RGBT person detection promotes various applications such as search and rescue due to its maneuverability. While existing research predominantly concentrates on refining fusion strategies and bolstering learning mechanisms for small objects, the pervasive yet unique occlusion challenge in drone-based RGBT settings remains inadequately addressed. In this work, we address the unique challenge of occlusion in the context of RGBT small object detection, particularly emphasizing its vulnerability and the distinct characteristics it exhibits across different modalities. We propose AODet, a novel Anti-Occlusion Detector meticulously crafted to tackle the challenges posed by occlusion in drone-based RGBT object detection. Our proposed approach significantly improves the detection performance of RGBT small objects, surpassing strong baselines on two large-scale datasets, VTUAV-det and RGBTDronePerson, by 1.30 points and 2.24 points in mAPsand ${\text{mAP}}_{50}^{{\text{tiny}}}$, respectively. Ziming Gui, Yan Zhang 0115, Xu Lei 0002, Ruixiang Zhang, Wen Yang 0001 |
IGARSS | 5 |
| 2024 | Decoupling Representation for Nighttime Aerial TrackingabstractNighttime aerial tracking is an indispensable step towards around-the-clock real-world applications. However, RGBbased tracking algorithms face significant challenges at night due to their vulnerability to illumination. Observing that different feature channels have varying sensitivity to illumination, we propose to decouple the representation for illuminationsensitive and illumination-insensitive embeddings. We devise a Nighttime aerial tracking scheme via Decouple Representations, termed NiDR, where the Illumination-Invariant Embedding (IIE) module and the Illumination-Sensitive Embedding (ISE) module are designed to decouple representations. We achieve this semantic decoupling by utilizing a pair of normlight and low-light images and regulating the reconstruction and consistency relations between features. Experiments on UAVDark135 exhibit the remarkable performance of NiDR under challenging nighttime scenarios, surpassing the secondbest competitor by a large margin of 3.1% on precision. Xu Lei 0002, Yan Zhang 0115, Chang Xu 0027, Wen Yang 0001, Wensheng Cheng |
IGARSS | 4 |
| 2024 | Ship Removal from High-Resolution Satellite Images: Dataset and ChallengesabstractThe removal of sensitive objects is critical for the secure sharing of remote sensing resources. Manual object removal is labor-intensive, and automated removal falls short, posing challenges in achieving the necessary quality for sensitive object removal. In this paper, we propose an integrated approach, combining state-of-the-art object detection models with advanced image inpainting techniques in a two-stage processing model tailored for detecting and removing sensitive targets in remote sensing images. We validate the pipeline’s performance using ship targets in remote sensing mapping. Specifically, we first curate a ship removal dataset, dubbed RSInpainting-Ships. We then introduce a refined random-mask strategy, guiding the network towards filling the background within the targeted region for removal. Initial experiments reveal three primary factors influencing performance: size, location, and wake. Size relates to the quantity of pixels to be filled and the complexity of the contextual information, location influences gradient changes in background structure, and wake brings in additional texture and determines the performance of ship removal. Yida Pan, Lanxin Zeng, Wen Yang 0001 |
IGARSS | 5 |
| 2024 | Detecting Line Segments in Motion-Blurred Images With EventsabstractMaking line segment detectors more reliable under motion blurs is one of the most important challenges for practical applications, such as visual SLAM and 3D line mapping. Existing line segment detection methods face severe performance degradation for accurately detecting and locating line segments when motion blur occurs. While event data shows strong complementary characteristics to images for minimal blur and edge awareness at high-temporal resolution, potentially beneficial for reliable line segment recognition. To robustly detect line segments over motion blurs, we propose to leverage the complementary information of images and events. Specifically, we first design a general frame-event feature fusion network to extract and fuse the detailed image textures and low-latency event edges, which consists of a channel-attention-based shallow fusion module and a self-attention-based dual hourglass module. We then utilize the state-of-the-art wireframe parsing networks to detect line segments on the fused feature map. Moreover, due to the lack of line segment detection datasets with pairwise motion-blurred images and events, we contribute two datasets, i.e., synthetic FE-Wireframe and realistic FE-Blurframe, for network training and evaluation. Extensive analyses on the component configurations demonstrate the design effectiveness of our fusion network. When compared to the state-of-the-arts, the proposed approach achieves the highest detection accuracy while maintaining comparable real-time performance. In addition to being robust to motion blur, our method also exhibits superior performance for line detection under high dynamic range scenes. Huai Yu, Hao Li 0114, Wen Yang 0001, Lei Yu 0006, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | CrossZoom: Simultaneous Motion Deblurring and Event Super-ResolvingabstractEven though the collaboration between traditional and neuromorphic event cameras brings prosperity to frame-event based vision applications, the performance is still confined by the resolution gap crossing two modalities in both spatial and temporal domains. This paper is devoted to bridging the gap by increasing the temporal resolution for images, i.e., motion deblurring, and the spatial resolution for events, i.e., event super-resolving, respectively. To this end, we introduce CrossZoom, a novel unified neural Network (CZ-Net) to jointly recover sharp latent sequences within the exposure period of a blurry input and the corresponding High-Resolution (HR) events. Specifically, we present a multi-scale blur-event fusion architecture that leverages the scale-variant properties and effectively fuses cross-modal information to achieve cross-enhancement. Attention-based adaptive enhancement and cross-interaction prediction modules are devised to alleviate the distortions inherent in Low-Resolution (LR) events and enhance the final results through the prior blur-event complementary information. Furthermore, we propose a new dataset containing HR sharp-blurry images and the corresponding HR-LR event streams to facilitate future research. Extensive qualitative and quantitative experiments on synthetic and real-world datasets demonstrate the effectiveness and robustness of the proposed method. Chi Zhang 0027, Xiang Zhang 0022, Mingyuan Lin, Cheng Li 0023, Chu He, Wen Yang 0001, Gui-Song Xia, Lei Yu 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Beyond Dehazing: Learning Intrinsic Hazy Robustness for Aerial Object DetectionabstractAccurate object detection in aerial imagery is crucial across numerous applications. However, haze can significantly degrade the performance of normal detectors, presenting a substantial obstacle in real-world scenarios. Previous solutions often resort to image dehazing as a pre-processing step to enhance image quality for subsequent detection. Despite being logically intuitive, their performance is limited due to the inherent objective mismatch between low-level image restoration tasks and high-level object detection tasks. In this article, we present haze-robust aerial object detection (HRAOD) to directly enhance detection robustness under hazy conditions. HRAOD constructs a clean-to-hazy distillation framework, enabling the detector to “see through haze,” without relying on the explicit image dehazing process. To address the challenge of extracting informative hazy features from blurry and low-contrast hazy images, we introduce a gradient-guided feature imitation method to emphasize the desired objects. Moreover, recognizing that different regions suffer from varying degradation degrees and pose distinct detection difficulties, we further propose a degradation-weighted response distillation method to mimic the normal predictions according to the degradation pattern adaptively. Due to the scarcity of hazy aerial data, we curate two remote sensing hazy aerial datasets, namely DOTA-Haze and SODA-A-Haze, and one drone hazy aerial dataset, DroneVehicle-Haze, for simulation. Extensive experimental results demonstrate the superiority of our method. Specifically, our HRAOD outperforms the state-of-the-art “dehaze + detect” method by 13.1 points in mAP on the DOTA-Haze dataset without incurring additional inference costs. HRAOD also performs favorably against other methods on real-world hazy scenes. Yan Zhang 0115, Ruixiang Zhang, Wen Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | NiDR: Nighttime Aerial Tracking via Decoupled RepresentationsabstractVanilla aerial trackers exhibit sensitivity to low-light conditions (e.g., nighttime aerial tracking scenario). To mitigate this, existing methods incorporate the light enhancement method as a preprocessing for aerial tracking. Despite the advancements, these approaches are restricted to the disparity in task objectives between the enhancer and tracker. Motivated by the observation that feature channels exhibit varying sensitivity to illumination, we propose to decouple the feature representation into two distinct parts: 1) illumination-invariant feature embedding and 2) illumination-sensitive feature embedding. The former, realized by the illumination invariant embedding (IIE) module, enhances features that remain invariant to illumination changes. Meanwhile, the latter, facilitated by the illumination sensitive embedding (ISE) module, aims to mitigate the negative impact of illumination-sensitive features on tracking performance. Building upon this decoupling strategy, we introduce NiDR, a simple yet effective nighttime aerial tracker. The proposed NiDR exhibits strong performance on three nighttime aerial tracking benchmarks (i.e., UAVDark135, NAT2021, and DarkTrack2021). Notably, it outperforms previous competitors by large margins, e.g., 3.1 points on the UAVDark135 and 2.0 points on the Darktrack2021 in terms of precision for nighttime scenarios. Xu Lei 0002, Yan Zhang 0115, Chang Xu 0027, Wensheng Cheng, Wen Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Learning Cross-View Visual Geo-Localization Without Ground TruthabstractCross-view geo-localization (CVGL) involves determining the geographical location of a query image by matching it with a corresponding GPS-tagged reference image. Current state-of-the-art methods predominantly rely on training models with labeled paired images, incurring substantial annotation costs and training burdens. In this study, we investigate the adaptation of frozen models for CVGL without requiring ground-truth pair labels. We observe that training on unlabeled cross-view images presents significant challenges, including establishing relationships within unlabeled data and reconciling view discrepancies between uncertain queries and references. To address these challenges, we propose a self-supervised learning framework to train a learnable adapter for a frozen foundation model (FM). This adapter is designed to map feature distributions from diverse views into a uniform space using unlabeled data exclusively. To establish relationships within unlabeled data, we introduce an expectation-maximization (EM)-based pseudolabeling module, which iteratively estimates matching between cross-view features and optimizes the adapter. To maintain the robustness of the FM’s representation, we incorporate an information consistency module with a reconstruction loss, ensuring that adapted features retain strong discriminative ability across views. Experimental results demonstrate that our proposed method achieves significant improvements over vanilla FMs and competitive accuracy compared to supervised methods while necessitating fewer training parameters and relying solely on unlabeled data. Evaluation of our adaptation for task-specific models further highlights its broad applicability. Particularly, on the University-1652 dataset, our method outperforms the FM baseline by a substantial margin, achieving about 39 points improvement in Recall@1 and more than 34 points increase in average precision (AP). The project is available athttps://collebt.github.io/EM-CVGL. Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Extracting Building Footprints in SAR Images via Distilling Boundary Information From Optical ImagesabstractBuildings represent pivotal entities in remote sensing imagery for various applications like urban planning and land resource management. Predominantly, methods for building footprint extraction in the literature focus on optical imagery with visual attributes that faithfully mirror the physical world. Nevertheless, the acquisition of high-quality optical images presents formidable challenges due to the susceptibility to illumination conditions and scene visibility. In contrast, synthetic aperture radar (SAR) images can be acquired in all-weather and all-time situations, unburdened by the aforementioned constraints. However, the coherent imaging mechanism engenders intricate complexities for building footprint extraction SAR images. To address this issue, this paper introduces the Boundary Information Distillation Network (BIDNet) to improve the prediction accuracy in SAR images by distilling knowledge from optical images. The proposed approach adopts a teacher-student framework, featuring two customized components: the Explicit Distillation Module (EDM) and the Latent Distillation Module (LDM). Different from the conventional practice of directly aligning feature maps, BIDNet focuses on leveraging the more conspicuous boundary information in optical images. The EDM operates by simultaneously yielding a boundary map to emphasize the boundary area and assimilating the explicit low-level features of two modalities. The LDM represents the structural attributes within the high-level latent feature space and aligns the representations of the two modalities. Within this module, intrinsic self-correlations among features originating from boundary regions are encoded, and so are the cross-correlations established between features from boundary regions and alternative areas. The two modules also serve as the conduit for knowledge distillation from the teacher network to the student network, enabling the utilization of optical imagery for enhancing the building footprint extraction in SAR imagery. Extensive experiments demonstrate that our BIDNet achieves state-of-the-art performance on the Multi-Sensor All Weather Mapping (MSAW) dataset, outperforming the strong baseline by 4.3-7.2 points in f1-score and 4.9-8.0 points in IoU. The source code and trained models will be publicly available. Lanxin Zeng, Wen Yang 0001, Jian Kang 0005, Huai Yu, Mihai Datcu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multimodal and Multiresolution Data Fusion for High-Resolution Cloud Removal: A Novel Baseline and BenchmarkabstractCloud removal is a significant and challenging problem in remote sensing, and in recent years, there have been notable advancements in this area. However, two major issues remain hindering the development of cloud removal: the unavailability of high-resolution imagery for existing datasets and the absence of evaluation regarding the semantic meaningfulness of the generated structures. In this paper, we introduce M3R-CR, a benchmark dataset for high-resolution Cloud Removal with Multi-Modal and Multi-Resolution data fusion. M3R-CR is the first public dataset for cloud removal to feature globally sampled high-resolution optical observations, paired with radar measurements and pixel-level land cover annotations. With this dataset, we consider the problem of cloud removal in high-resolution optical remote sensing imagery by integrating multi-modal and multi-resolution information. In this context, we have to take into account the alignment errors caused by the multi-resolution nature, along with the more pronounced misalignment issues in high-resolution images due to inherent imaging mechanism differences and other factors. Existing multi-modal data fusion based methods, which assume the image pairs are aligned accurately at pixel-level, are thus not appropriate for this problem. To this end, we design a new baseline named Align-CR to perform the low-resolution SAR image guided high-resolution optical image cloud removal. It gradually warps and fuses the features of the multi-modal and multi-resolution data during the reconstruction process, effectively mitigating concerns associated with misalignment. In the experiments, we evaluate the performance of cloud removal by analyzing the quality of visually pleasing textures using image reconstruction metrics and further analyze the generation of semantically meaningful structures using a well-established semantic segmentation task. The proposed Align-CR method is superior to other baseline methods in both areas. The project is available at https://github.com/zhu-xlab/M3R-CR. Yilei Shi, Patrick Ebel 0002, Wen Yang 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Learning Cross-Modality High-Resolution Representation for Thermal Small-Object DetectionabstractThermal infrared (TIR) object detection plays a crucial role in diverse around-the-clock applications, such as search and rescue operations and wildlife protection. Achieving rapid and robust detection of small objects from an aerial perspective is particularly significant in these scenarios. However, the task is compounded by two interrelated challenges, rendering it even more tricky. For one, small objects only occupy a few pixels and contain limited information. For another, TIR sensors are typically low-resolution (LR) due to inherent challenges associated with the imaging mechanism of the TIR spectrum. In contrast, high-resolution (HR) RGB sensors are readily available due to their cost-effectiveness and widespread application. Recognizing the importance of HR information, especially in the context of small object detection, we propose a cross-modality high-resolution knowledge distillation framework (CMHRD), which leverages knowledge from the HR-RGB modality and provides a novel strategy for TIR small object detection. The proposed framework introduces three key components: a super-resolution generative distillation loss for cross-modal high-resolution representation learning, a cross-modality affinity distillation loss to extract scene-level cross-modality information, and a response distillation loss aimed at mimicking the HR prediction. To facilitate research on small object detection with HR-RGB and LR-TIR data, we have curated and annotated two datasets, namely NOAA-Seal and VTUAV-det-small. Experimental results on the NOAA-Seal demonstrate that CMHRD yields significant improvements, achieving a remarkable 6.39 mAP50 increase over a strong baseline without introducing additional computational cost during inference. Experiments on single-category dataset VTUAV-det-small and multi-category dataset RTDOD also show consistent improvements brought by CMHRD. The project is available at https://github.com/NNNNerd/CMHRD. Yan Zhang 0115, Xu Lei 0002, Chang Xu 0027, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Dynamic Coarse-to-Fine Learning for Oriented Tiny Object DetectionabstractDetecting arbitrarily oriented tiny objects poses intense challenges to existing detectors, especially for label assignment. Despite the exploration of adaptive label assignment in recent oriented object detectors, the extreme geometry shape and limited feature of oriented tiny objects still induce severe mismatch and imbalance issues. Specifically, the position prior, positive sample feature, and instance are mismatched, and the learning of extreme-shaped objects is biased and unbalanced due to little proper feature supervision. To tackle these issues, we propose a dynamic prior along with the coarse-to-fine assigner, dubbed DCFL. For one thing, we model the prior, label assignment, and object representation all in a dynamic manner to alleviate the mismatch issue. For another, we leverage the coarse prior matching and finer posterior constraint to dynamically assign labels, providing appropriate and relatively balanced supervision for diverse instances. Extensive experiments on six datasets show substantial improvements to the baseline. Notably, we obtain the state-of-the-art performance for one-stage detectors on the DOTA-v1.5, DOTA-v2.0, and DIOR-R datasets under single-scale training and testing. Codes are available at https://github.com/Chasel-Tsui/mmrotate-dcfl. Chang Xu 0027, Jian Ding 0001, Jinwang Wang, Wen Yang 0001, Huai Yu, Lei Yu 0006, Gui-Song Xia |
CVPR | 4 |
| 2023 | Generalizing Event-Based Motion Deblurring in Real-World ScenariosabstractEvent-based motion deblurring has shown promising results by exploiting low-latency events. However, current approaches are limited in their practical usage, as they assume the same spatial resolution of inputs and specific blurriness distributions. This work addresses these limitations and aims to generalize the performance of event-based de-blurring in real-world scenarios. We propose a scale-aware network that allows flexible input spatial scales and enables learning from different temporal scales of motion blur. A two-stage self-supervised learning scheme is then developed to fit real-world data distribution. By utilizing the relativity of blurriness, our approach efficiently ensures the restored brightness and structure of latent images and further generalizes deblurring performance to handle varying spatial and temporal scales of motion blur in a self-distillation manner. Our method is extensively evaluated, demonstrating remarkable performance, and we also introduce a real-world dataset consisting of multi-scale blurry frames and events to facilitate research in event-based deblurring. Xiang Zhang 0022, Lei Yu 0006, Wen Yang 0001, Jianzhuang Liu, Gui-Song Xia |
ICCV | 3 |
| 2023 | Self-Supervised Dense Depth Estimation with Panoramic Image and Sparse LidarabstractThe 360-depth estimation with spherical images and LiDAR data has recently become increasingly popular in autonomous driving and scene reconstruction. Compared with perspective images, spherical images have omnidirectional FoV, which exceedingly matches LiDAR data. However, the spherical distortion makes the 360-depth estimation a great challenge. To address this problem, we propose a self-supervised 360 depth estimation network in this paper. The network consists of a spherical convolution branch to extract panoramic image features and a ResNet branch to extract LiDAR features. Then an attention-based decoder is designed to estimate the depth. The reprojection error is used to self-supervise the network training. Experiments on the KITTI-360 dataset demonstrate the effectiveness of the proposed method. Chenwei Lyu, Huai Yu, Zhipeng Zhao 0001, Pengliang Ji, Xiangli Yang, Wen Yang 0001 |
IGARSS | 6 |
| 2023 | Multi-Modal Multi-Task Learning for Semantic Segmentation of Land Cover Under Cloudy ConditionsabstractThe majority of the Earth’s surface is covered by clouds, causing optical images to suffer serious degradation of ground information. Synthetic Aperture Radar (SAR) images with the cloud-penetration capability could provide supplementary information to optical images. Thus, the fusion of optical and SAR image can remarkably improve the interpretation accuracy under cloudy conditions. In this paper, we propose to exploit related cloud removal task for accurate multi-modal semantic segmentation of land cover. Towards this goal, we develop an end-to-end learnable architecture which solves the tasks of cloud removal and semantic segmentation of land cover jointly. The cloud removal task encourages to learn knowledgeable features to overcome negative effects of semantic ambiguity. Our experiments show that the proposed algorithm can effectively improve the semantic segmentation accuracy under cloudy conditions. Yilei Shi, Wen Yang 0001, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2023 | Cross-Modal 2D-3D Localization with Single-Modal QueryabstractGlobal visual localization is an important task in geoscience with a plethora of applications such as SLAM and autonomous navigation. Current place recognition approaches restrict the modality of the query data which relies on the database data modality. However, real-world robots are equipped with different sensors in different application scenarios and it is difficult for data from a single fixed modality to accommodate all challenging environments. To overcome this limitation, we propose to build a generalized model that allows spherical images and point clouds to be retrieved under any single-modal query. Our 2D-3D dataset is created based on the KITTI360 dataset with spherical images and corresponding point clouds for training and evaluation. Extensive experimental results demonstrate the effectiveness of our proposed approach. Zhipeng Zhao 0001, Huai Yu, Chenwei Lyu, Pengliang Ji, Xiangli Yang, Wen Yang 0001 |
IGARSS | 6 |
| 2023 | Multi-level Feature Interaction and Efficient Non-Local Information Enhanced Channel Attention for image dehazing
Bohui Li, Zhiping Dan, Bo Du 0001, Wen Yang 0001, Jun Wan 0005 |
Neural Networks | 6 |
| 2023 | Learning to Extract Building Footprints From Off-Nadir Aerial ImagesabstractExtracting building footprints from aerial images is essential for precise urban mapping with photogrammetric computer vision technologies. Existing approaches mainly assume that the roof and footprint of a building are well overlapped, which may not hold in off-nadir aerial images as there is often a big offset between them. In this paper, we propose an offset vector learning scheme, which turns the building footprint extraction problem in off-nadir images into an instance-level joint prediction problem of the building roof and its corresponding “roof to footprint” offset vector. Thus the footprint can be estimated by translating the predicted roof mask according to the predicted offset vector. We further propose a simple but effective feature-level offset augmentation module, which can significantly refine the offset vector prediction by introducing little extra cost. Moreover, a new dataset, Buildings in Off-Nadir Aerial Images (BONAI), is created and released in this paper. It contains 268,958 building instances across 3,300 aerial images with fully annotated instance-level roof, footprint, and corresponding offset vector for each building. Experiments on the BONAI dataset demonstrate that our method achieves the state-of-the-art, outperforming other competitors by 3.37 to 7.39 points in F1-score. The codes, datasets, and trained models are available athttps://github.com/jwwangchn/BONAI.git. Jinwang Wang, Lingxuan Meng, Wen Yang 0001, Lei Yu 0006, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Learning to Super-Resolve Blurry Images With EventsabstractSuper-Resolution from a single motion Blurred image (SRB) is a severely ill-posed problem due to the joint degradation of motion blurs and low spatial resolution. In this article, we employ events to alleviate the burden of SRB and propose an Event-enhanced SRB (E-SRB) algorithm, which can generate a sequence of sharp and clear images with High Resolution (HR) from a single blurry image with Low Resolution (LR). To achieve this end, we formulate an event-enhanced degeneration model to consider the low spatial resolution, motion blurs, and event noises simultaneously. We then build an event-enhanced Sparse Learning Network (eSL-Net++) upon a dual sparse learning scheme where both events and intensity frames are modeled with sparse representations. Furthermore, we propose an event shuffle-and-merge scheme to extend the single-frame SRB to the sequence-frame SRB without any additional training process. Experimental results on synthetic and real-world datasets show that the proposed eSL-Net++ outperforms state-of-the-art methods by a large margin. Datasets, codes, and more results are available at https://github.com/ShinyWang33/eSL-Net-Plusplus. Lei Yu 0006, Bishan Wang, Xiang Zhang 0022, Wen Yang 0001, Jianzhuang Liu, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Learning to See Through With EventsabstractAlthough synthetic aperture imaging (SAI) can achieve the seeing-through effect by blurring out off-focus foreground occlusions while recovering in-focus occluded scenes from multi-view images, its performance is often deteriorated by dense occlusions and extreme lighting conditions. To address the problem, this paper presents an Event-based SAI (E-SAI) method by relying on the asynchronous events with extremely low latency and high dynamic range acquired by an event camera. Specifically, the collected events are first refocused by a Refocus-Net module to align in-focus events while scattering out off-focus ones. Following that, a hybrid network composed of spiking neural networks (SNNs) and convolutional neural networks (CNNs) is proposed to encode the spatio-temporal information from the refocused events and reconstruct a visual image of the occluded targets. Extensive experiments demonstrate that our proposed E-SAI method can achieve remarkable performance in dealing with very dense occlusions and extreme lighting conditions and produce high-quality images from pure events. Codes and datasets are available at https://dvs-whu.cn/projects/esai/. Lei Yu 0006, Xiang Zhang 0022, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | A3Track: Achieving Precise Target Tracking in Aerial Images With Receptive Field AlignmentabstractTracking arbitrary objects in aerial images presents formidable challenges to existing trackers. Among these challenges, the large scale variation and arbitrary geometry shape of visual targets are pronounced, resulting in two-fold mismatch issues between the feature receptive field and the tracking target. For one, there is a mismatch between the prior receptive field center and arbitrary-shaped targets. For another, the single receptive field mismatches the significantly scale-varied targets in the aerial imagery. To handle these challenges, we propose to Achieve precise Aerial tracking with receptive field Alignment, dubbed A3Track. The proposed A3Track is comprised of two modules: a Receptive Field Alignment (RFA) module and a Pyramid Receptive Field (PRF) module. First of all, we transform and update the receptive field center progressively, which drives the feature sampling location onto the targets’ main body, thus gradually yielding precise feature representation for arbitrary-shaped targets. We term this progressively updating process as the Receptive Field Alignment. Moreover, the PRF module constructs a set of pyramid features for the target, providing a multi-scale receptive field to handle the large scale variation of tracking objects. On four benchmarks, the new tracker A3Track achieves leading performance compared with existing methods and shows consistent improvements over baselines. The project is available at: https://chnleixu.github.io/A3Track-web/. Xu Lei 0002, Chang Xu 0027, Wensheng Cheng, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Partial Siamese With Multiscale Bi-Codec Networks for Remote Sensing Image Haze RemovalabstractRecently, the U-Shaped networks has been widely explored in remote sensing image dehazing and obtained promising performance. However, most of the existing dehazing methods based on U-Shaped framework lack the reconstruction constraints of haze areas, which is particularly important to restore haze-free images. Moreover, their encoding and decoding layers cannot effectively fuse multi-scale features, resulting in deviations in the color and texture of the dehazing image. To address these issues, in this paper, we propose a Partial Siamese with Multiscale Bi-codec Dehazing Network (PSMB-Net) which is mainly composed of a Partial Siamese Framework (PSF) and a Multiscale Bi-codec Information Fusion (MBIF) module. Specifically, the PSF is proposed to create dehazing prior information to guide the network to build Siamese constraints and achieve improved dehazing results. Furthermore, we design a MBIF module which can enhance feature extraction, and the multi-scale information is used to improve the reconstruction ability of the network for the color and texture of the dehazing image. Experimental results on challenging benchmark datasets demonstrate the superiority of our PSMB-Net over state-of-the-art image dehazing methods. The source code is available at https://github.com/thislzm/PSMB-Net. Zhiming Luo, Bo Du 0001, Wen Yang 0001, Jun Wan 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Categorical Inference Poisoning: Verifiable Defense Against Black-Box DNN Model Stealing Without Constraining Surrogate Data and Query TimesabstractDeep Neural Network (DNN) models have offered powerful solutions for a wide range of tasks, but the cost to develop such models is nontrivial, which calls for effective model protection. Although black-box distribution can mitigate some threats, model functionality can still be stolen via black-box surrogate attacks. Recent studies have shown that surrogate attacks can be launched in several ways, while the existing defense methods commonly assume attackers with insufficient in-distribution (ID) data and restricted attacking strategies. In this paper, we relax these constraints and assume a practical threat model in which the adversary not only has sufficient ID data and query times but also can adjust the surrogate training data labeled by the victim model. Then, we propose a two-step categorical inference poisoning (CIP) framework, featuring both poisoning for performance degradation (PPD) and poisoning for backdooring (PBD). In the first poisoning step, incoming queries are classified into ID and (out-of-distribution) OOD ones using an energy score (ES) based OOD detector, and the latter are further classified into high ES and low ES ones, which are subsequently passed to a strong and a weak PPD process, respectively. In the second poisoning step, difficult ID queries are detected by a proposed reliability score (RS) measurement and are passed to PBD. In doing so, the first step OOD poisoning leads to substantial performance degradation in surrogate models, the second step ID poisoning further embeds backdoors in them, while both can preserve model fidelity. Extensive experiments confirm that CIP can not only achieve promising performance against state-of-the-art black-box surrogate attacks like KnockoffNets and data-free model extraction (DFME) but also work well against stronger attacks with sufficient ID and deceptive data, better than the existing dynamic adversarial watermarking (DAWN) and deceptive perturbation defense methods. PyTorch code is available athttps://github.com/Hatins/CIP_master.git. Haitian Zhang, Guang Hua 0001, Xinya Wang, Hao Jiang 0010, Wen Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Single Image Deraining With Continuous Rain Density EstimationabstractSingle image deraining (SIDR) often suffers from over/under deraining due to the nonuniformity of rain densities and the variety of raindrop scales. In this paper, we propose acontinuousdensity-guided network (CODE-Net) for SIDR. Particularly, it is composed of a rain streak extractor and a denoiser, where the convolutional sparse coding (CSC) is exploited to filter out noises from the extracted rain streaks. Inspired by the reweighted iterative soft-threshold (ISTA) for CSC, we address the problem of continuous rain density estimation by learning the weights with channel attention blocks from sparse codes. We further develop a multiscale strategy to depict rain streaks appearing at different scales. Experiments on synthetic and real-world data demonstrate the superiority of our methods over recent state-of-the-arts, in terms of both quantitative and qualitative results. Additionally, instead of quantizing rain density with several levels, our CODE-Net can provide continuous-valued estimations of rain densities, which is more desirable in real applications. Lei Yu 0006, Bishan Wang, Jingwei He, Gui-Song Xia, Wen Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Synthetic Aperture Imaging with Events and FramesabstractThe Event-based Synthetic Aperture Imaging (E-SAI) has recently been proposed to see through extremely dense occlusions. However, the performance of E-SAI is not consistent under sparse occlusions due to the dramatic de-crease of signal events. This paper addresses this problem by leveraging the merits of both events and frames, leading to a fusion-based SAl (EF-SAI) that performs consistently under the different densities of occlusions. In particular, we first extract the feature from events and frames via multi-modal feature encoders and then apply a multi-stage fusion network for cross-modal enhancement and density-aware feature selection. Finally, a CNN decoder is employed to generate occlusion-free visual images from selected features. Extensive experiments show that our method effectively tackles varying densities of occlusions and achieves superior performance to the state-of-the-art SAl methods. Codes and datasets are available at https://github.com/smjsc/EF-SAI Xiang Zhang 0022, Lei Yu 0006, Shijie Lin, Wen Yang 0001 |
CVPR | 5 |
| 2022 | RFLA: Gaussian Receptive Field Based Label Assignment for Tiny Object Detection
Chang Xu 0027, Jinwang Wang, Wen Yang 0001, Huai Yu, Lei Yu 0006, Gui-Song Xia |
ECCV (9) | 3 |
| 2022 | Towards Robust Visual-Inertial Odometry with Multiple Non-Overlapping Monocular CamerasabstractWe present a Visual-Inertial Odometry (VIO) algorithm with multiple non-overlapping monocular cameras aiming at improving the robustness of the VIO algorithm. An initialization scheme and tightly-coupled bundle adjustment for multiple non-overlapping monocular cameras are proposed. With more stable features captured by multiple cameras, VIO can maintain stable state estimation, especially when one of the cameras tracked unstable or limited features. We also address the high CPU usage rate brought by multiple cameras by proposing a GPU-accelerated frontend. Finally, we use our pedestrian carried system to evaluate the robustness of the VIO algorithm in several challenging environments. The results show that the multi-camera setup yields significantly higher estimation robustness than a monocular system while not increasing the CPU usage rate (reducing the CPU resource usage rate and computational latency by 40.4% and 50.6% on each camera). A demo video can be found at https://youtu.be/r7QvPth1m10. Huai Yu, Wen Yang 0001, Sebastian A. Scherer |
IROS | 3 |
| 2022 | Object Detection in Aerial Images: A Large-Scale Benchmark and ChallengesabstractIn he past decade, object detection has achieved significant progress in natural images but not in aerial images, due to the massive variations in the scale and orientation of objects caused by the bird's-eye view of aerial images. More importantly, the lack of large-scale benchmarks has become a major obstacle to the development of object detection in aerial images (ODAI). In this paper, we present a large-scale Dataset of Object deTection in Aerial images (DOTA) and comprehensive baselines for ODAI. The proposed DOTA dataset contains 1,793,658 object instances of 18 categories of oriented-bounding-box annotations collected from 11,268 aerial images. Based on this large-scale and well-annotated dataset, we build baselines covering 10 state-of-the-art algorithms with over 70 configurations, where the speed and accuracy performances of each model have been evaluated. Furthermore, we provide a code library for ODAI and build a website for evaluating different algorithms. Previous challenges run on DOTA have attracted more than 1300 teams worldwide. We believe that the expanded large-scale DOTA dataset, the extensive baselines, the code library and the challenges can facilitate the designs of robust algorithms and reproducible research on the problem of object detection in aerial images. Jian Ding 0001, Nan Xue 0001, Gui-Song Xia, Xiang Bai, Wen Yang 0001, Michael Ying Yang, Serge J. Belongie, Jiebo Luo 0001, Mihai Datcu, Marcello Pelillo, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Asymmetric Siamese Networks for Semantic Change Detection in Aerial ImagesabstractGiven two multitemporal aerial images, semantic change detection (SCD) aims to locate the land-cover variations and identify their change types with pixelwise boundaries. This problem is vital in many earth vision-related tasks, such as precise urban planning and natural resource management. Existing state-of-the-art algorithms mainly identify the changed pixels by applying homogeneous operations on each input image and comparing the extracted features. However, in changed regions, totally different land-cover distributions often require heterogeneous feature extraction procedures for images acquired at different times. In this article, we present an asymmetric Siamese network (ASN) to locate and identify semantic changes through feature pairs obtained from modules of widely different structures, which involves areas of various sizes and applies different quantities of parameters to factor in the discrepancy across land-cover distributions during different times. To better train and evaluate our model, we create a large-scale well-annotated SEmantic Change detectiON Dataset (SECOND), while an adaptive threshold learning (ATL) module and a separated kappa (SeK) coefficient are proposed to alleviate the influences of label imbalance in model training and evaluation. The experimental results demonstrate that the proposed model can stably outperform the state-of-the-art algorithms with different encoder backbones. Kunping Yang, Gui-Song Xia, Zicheng Liu 0003, Bo Du 0001, Wen Yang 0001, Marcello Pelillo, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Instance Switching-Based Contrastive Learning for Fine-Grained Airplane DetectionabstractDetecting airplanes from high-resolution remote sensing images has a variety of applications. The characteristics of clear details, rich spatial and texture information of objects in high-resolution remote sensing images make it possible to identify different types of airplanes from backgrounds. However, airplanes usually exhibit slight inter-class discrepancy and unbalanced class distribution, which pose significant challenges to fine-grained detection of airplanes. In this paper, we propose the ISCL, an Instance Switching-based Contrastive Learning method for fine-grained airplane detection. Specifically, we introduce a Contrastive Learning-based Module (CLM) to widen the inter-class distance while narrowing the intra-class distance by optimizing feature space distribution with the InfoNCE+loss, which is built on a serial head in a cascaded way. Then, we design a Refined Instance Switching (ReIS) module to alleviate the class imbalance problem. To take full advantage of the CLM and ReIS, we further introduce an optimization strategy which is an organic combination of the two modules to widen the distances of different airplane categories that are easily confused. In addition, we contribute a fine-grained attribute-assisted dataset, dubbed GF-RarePlanes Dataset (GRD), to help the detectors better learn the subtle differences between the airplanes. Extensive experiments on two datasets (i.e., GF and FAIR1M) demonstrate that our proposed method can significantly improve the accuracy of fine-grained airplane detection under both HBB and OBB scenarios. Dataset and codes will be available at https://lanxin1011.github.io/ISCL/. Lanxin Zeng, Haowen Guo, Wen Yang 0001, Huai Yu, Lei Yu 0006, Tongyuan Zou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Optical-Enhanced Oil Tank Detection in High-Resolution SAR ImagesabstractIn recent years, object detection in high-resolution SAR images has made significant progress, especially after the introduction of deep learning. However, objects like dense oil tanks, which are compactly arranged in SAR images, are still challenging to recognize due to the unique imaging mechanism of SAR. Inspired by human learning from comparison, we propose a multi-stage framework for oil tank detection in SAR images using optical image enhancement. Specifically, in the training stage, we build a teacher-student network to align the semantic information between the two modalities, where the optical features are used to guide the corresponding SAR feature learning. While in the inference stage, the learned network detects oil tanks using only SAR images as input. Besides, a pre-training stage before training is applied to further improve the network’s ability for SAR feature extraction, which is realized by the proposed paired optical-SAR self-supervised learning. To verify the effectiveness of the proposed method, we perform experiments on our newly built SpaceNet6-OTD dataset. Extensive experiments demonstrate that the proposed method can effectively improve the accuracy of detecting oil tanks in SAR images. Datasets, codes, and more results will be released at: https://EIS-VIPG.github.io/SpaceNet6-OTD/. Ruixiang Zhang, Haowen Guo, Wen Yang 0001, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Event-Based Synthetic Aperture Imaging With a Hybrid NetworkabstractSynthetic aperture imaging (SAI) is able to achieve the see through effect by blurring out the off-focus foreground occlusions and reconstructing the in-focus occluded targets from multi-view images. However, very dense occlusions and extreme lighting conditions may bring significant disturbances to the SAI based on conventional frame-based cameras, leading to performance degeneration. To address these problems, we propose a novel SAI system based on the event camera which can produce asynchronous events with extremely low latency and high dynamic range. Thus, it can eliminate the interference of dense occlusions by measuring with almost continuous views, and simultaneously tackle the over/under exposure problems. To reconstruct the occluded targets, we propose a hybrid encoder-decoder network composed of spiking neural networks (SNNs) and convolutional neural networks (CNNs). In the hybrid network, the spatio-temporal information of the collected events is first encoded by SNN layers, and then transformed to the visual image of the occluded targets by a style-transfer CNN decoder. Through experiments, the proposed method shows remarkable performance in dealing with very dense occlusions and extreme lighting conditions, and high quality visual images can be reconstructed using pure event data. Xiang Zhang 0022, Lei Yu 0006, Wen Yang 0001, Gui-Song Xia |
CVPR | 4 |
| 2021 | Motion Deblurring with Real EventsabstractIn this paper, we propose an end-to-end learning framework for event-based motion deblurring in a self-supervised manner, where real-world events are exploited to alleviate the performance degradation caused by data inconsistency. To achieve this end, optical flows are predicted from events, with which the blurry consistency and photometric consistency are exploited to enable self-supervision on the deblurring network with real-world data. Furthermore, a piecewise linear motion model is proposed to take into account motion non-linearities and thus leads to an accurate model for the physical formation of motion blurs in the real-world scenario. Extensive evaluation on both synthetic and real motion blur datasets demonstrates that the proposed algorithm bridges the gap between simulated and real-world motion blurs and shows remarkable performance for eventbased motion deblurring in real-world scenarios. Lei Yu 0006, Bishan Wang, Wen Yang 0001, Gui-Song Xia, Xu Jia 0012, Zhendong Qiao, Jianzhuang Liu |
ICCV | 4 |
| 2021 | Toward Dataset Construction for Remote Sensing Image InterpretationabstractWith the rapid advancement of remote sensing (RS) technology, RS image interpretation has made great progress and been widely used in broad applications, in which the constructed benchmark datasets for developing and testing intelligent interpretation algorithms have been playing an increasingly critical role. Motivated by the essential prerequisites of dataset in the development of RS image interpretation algorithms, this manuscript provides a discussion on dataset construction for RS image interpretation. Specifically, we first analyze the current challenges of developing algorithms for RS image interpretation and a review on the widespread RS image datasets is conducted through the bibliometric analysis. We then propose some principles and discuss the methodology on constructing benchmark datasets. An implementation on creating the RS scene classification dataset demonstrates the practicability of our proposed framework and the experimental results show that our constructed dataset can serve as a promising benchmark for RS image scene interpretation. Yang Long 0002, Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 3 |
| 2021 | Learning Center Probability Map for Detecting Objects in Aerial ImagesabstractOne fundamental problem in Earth Vision is to accurately find the locations and identify the categories of the interesting objects in the aerial images, for which oriented bounding boxes (OBBs) are usually employed to depict better the objects emerging with arbitrary orientations. However, the regression of the OBBs always suffers from the ambiguous problem in the definition of the regression targets, which often reduces the convergency efficiency and decreases the detection accuracy. Although there are some methods like the binary segmentation map that can handle this problem, it brings a new problem of ambiguous background pixels in the OBBs. In this article, we propose to cast the OBB regression as a center-probability-map (CenterMap)-prediction problem, thus largely eliminating the ambiguities on the target definitions and the background pixels. The predicted CenterMaps are then used to generate the OBBs. The CenterMap OBB representation is simple, yet effective. Furthermore, to distinguish better the interesting objects from the cluttered background, a weighted pseudosegmentation-guided attention network is adopted to provide the object-level features for predicting the horizontal bounding boxes and the OBBs. The experimental results on three widely used data sets, i.e., DOTA, HRSC2016, and UCAS-AOD, demonstrate the effectiveness of our proposed method. Jinwang Wang, Wen Yang 0001, Heng-Chao Li 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Event Enhanced High-Quality Image Recovery
Bishan Wang, Jingwei He, Lei Yu 0006, Gui-Song Xia, Wen Yang 0001 |
ECCV (13) | 5 |
| 2020 | Image De-Raining Via RDL: When Reweighted Convolutional Sparse Coding Meets Deep LearningabstractOver the past few decades, image de-raining has witnessed substantial progress due to the development of priors and deep learning based methods. However, few studies combine the merits of both. In this paper, we argue that domain expertise of conventional convolutional sparse coding (CSC) is still valuable, and it can be combined with the key ingredients of deep learning to achieve further improved results. Specifically, motivated by the success of reweighting algorithms, we propose solving the CSC model by learning weighted iterative soft thresholding algorithm (LwISTA) in a convolutional manner where the reweighted ℓ1-norm is introduced. Based on this, we present a novel framework for single image de-raining, in which the channel attention is employed to learn the weight. Extensive experiments demonstrate the superiority of our method over recent state-of-the-art image de-raining methods, in terms of both quantitative and qualitative results. Jingwei He, Lei Yu 0006, Wen Yang 0001 |
ICASSP | 3 |
| 2020 | Robust Intensity Image Reconstruciton Based On Event CamerasabstractThe event camera is a novel sensor that records brightness change in the form of asynchronous events with high temporal resolution, and simultaneously outputs intensity images with a lower frame rate. Events recorded by sensors have a lot of noise and the intensity images captured often suffer from motion blur and noise effects. Therefore, to reconstruct high quality images is of great significance for the application of event camera in computer vision. However, the existing reconstruction methods only addressed the motion blur issue without considering the influence of noise. In this paper, we propose a variational model by using spatial smooth constraint regularization to recover clean image frames from blurry and noisy camera images and events at any frame rate. We present experimental results on synthetic dataset as well as real dataset with high speed and high dynamic range to demonstrate that the proposed algorithm is superior to the other reconstruction algorithms. Bishan Wang, Lei Yu 0006, Wen Yang 0001 |
ICIP | 5 |
| 2020 | Event-Based High Frame-Rate Video Reconstruction With A Novel Cycle-Event NetworkabstractsEvent-to-image translation is a popular problem where the goal is to obtain a mapping from an input event stream to an output intensity image using a set of aligned image pairs for training. However, due to the high temporal resolution of the event camera, the alignment of the ground truth to the events is difficult to acquire. In this paper, we firstly propose an enhanced Cycle-Consistency Generative Adversarial Networks (enhanced Cycle-GAN), called Cycle-Event Network, where paired data is not required for the training phase. Besides, noises from event cameras can severely contaminate the data quality and makes the reconstruction an ill-posed problem. In order to generate high frame-rate video from events with less noisy background and richer texture details, a novel attention mechanism (Residual Channel-wise Attention Gate) is then proposed to reweight the feature of the generator in Cycle-Event Network. The qualitative results are presented on several datasets, and the quantitative comparisons clearly demonstrate the effectiveness of our proposed Cycle-Event Networks. Binyi Su, Lei Yu 0006, Wen Yang 0001 |
ICIP | 3 |
| 2020 | Tiny Object Detection in Aerial ImagesabstractObject detection in Earth Vision has achieved great progress in recent years. However, tiny object detection in aerial images remains a very challenging problem since the tiny objects contain a small number of pixels and are easily confused with the background. To advance tiny object detection research in aerial images, we present a new dataset for Tiny Object Detection in Aerial Images (AI-TOD). Specifically, AI-TOD comes with 700,621 object instances for eight categories across 28,036 aerial images. Compared to existing object detection datasets in aerial images, the mean size of objects in AI-TOD is about 12.8 pixels, which is much smaller than others. To build a benchmark for tiny object detection in aerial images, we evaluate the state-of-the-art object detectors on our AI-TOD dataset. Experimental results show that direct application of these approaches on AI-TOD produces suboptimal object detection results, thus new specialized detectors for tiny object detection need to be designed. Therefore, we propose a multiple center points based learning network (M-CenterNet) to improve the localization performance of tiny object detection, and experimental results show the significant performance gain over the competitors. Jinwang Wang, Wen Yang 0001, Haowen Guo, Ruixiang Zhang, Gui-Song Xia |
ICPR | 2 |
| 2020 | Change Detection of Polarimetric SAR Images Using Minkowski Log-Ratio DistanceabstractThe Minkowski log-ratio (MLR) distance admits closed-form formula for mixture model of exponential families (Gaussian family, Wishart family, etc.), and is suitable for measuring the dissimilarity of polarimetric SAR (PolSAR) data. In this paper, MLR distance is introduced for PolSAR image change detection. Specifically, the PolSAR images are estimated by Wishart mixture models and over-segmented into superpixels first. After that, the statistical distribution differences between two corresponding superpixels are measured by MLR distance. Finally, the change detection map is obtained by the Kittler-Illingworth thresholding method based on the generated difference map. Qualitative and quantitative experimental analysis shows the MLR distance can provide a desirable difference map which facilitates the following thresholding stage. Shuailin Chen, Xiangli Yang, Tongyuan Zou, Dong Peng, Wen Yang 0001, Heng-Chao Li 0001 |
IGARSS | 5 |
| 2020 | Look at the Big Picture: Building Area Extraction with Global Density MapabstractThe automatic extraction of building areas from high-resolution satellite imagery has become an important and challenging research issue. Many recent studies have explored different deep learning-based semantic segmentation methods for better accuracy. However, the deep network usually takes sliding window cropped satellite images as inputs, which loses the global information and causes a high false positive rate. In this paper, we propose a density map guided attention mechanism for building area extraction to make the network look at the big picture. We exploit an FCN-based building density prediction network to generate a density heatmap from large satellite images. The density factors in heatmap control the classifier's threshold of building area extraction network that optimize the FP and recall rates. Furthermore, we propose a test-time overlap augmentation mechanism to improve the segmentation results. Our method outperforms state-of-the-art approaches and increases mIoU by about 3.08% to 93.31%, and decreases FP rate to 0.91%. Haowen Guo, Wensheng Cheng, Wen Yang 0001, Gui-Song Xia |
IGARSS | 3 |
| 2020 | Instance Segmentation with Oriented Proposals for Aerial ImagesabstractInstance segmentation is a challenging issue in remote sensing. The existing state-of-the-art methods use horizontal bounding box (HBB) to infer the instance mask of object. However, the objects in the aerial images have the characteristics of being distributed in arbitrary orientation, densely packed, and so on. In this case, the HBB usually contains a lot of background information and several neighboring objects, leading to coarse and inaccurate mask prediction. To solve the aforementioned problems, we propose a new instance segmentation method, ISOP, by inferring the mask on oriented bounding box (OBB) instead of HBB. We show that the proposed method leading to more accurate mask predictions, especially for densely packed objects. We evaluate our method in the iSAID dataset, and compared to the baseline, the ISOP has achieved around 17% improvement in terms of mAP and 11% for densely packed objects. Jian Ding 0001, Jinwang Wang, Wen Yang 0001, Gui-Song Xia |
IGARSS | 4 |
| 2020 | Edge-Driven Object Matching for UAV Images and Satellite SAR ImagesabstractThe task of matching between images acquired by terminal equipment and satellites is important and challenging due to the dramatic viewpoint changes and unknown orientations, especially with different imaging sensors. In this paper, we firstly present the task to match the optical/infrared images acquired by UAVs with satellite SAR images. Many previous works mainly focused on matching with the images of the same modality, and may not perform well to our task. To overcome the difficulties caused by the diversity among the three modalities of data, we mine the common features of them and propose a novel edge-driven matching framework to find the correspondence between the UAV images and SAR images. Experimental results demonstrate the effectiveness and superiority of our method. Ruixiang Zhang, Huai Yu, Wen Yang 0001, Heng-Chao Li 0001 |
IGARSS | 4 |
| 2020 | Monocular Camera Localization in Prior LiDAR Maps with 2D-3D Line CorrespondencesabstractLight-weight camera localization in existing maps is essential for vision-based navigation. Currently, visual and visual-inertial odometry (VO&VIO) techniques are well-developed for state estimation but with inevitable accumulated drifts and pose jumps upon loop closure. To overcome these problems, we propose an efficient monocular camera localization method in prior LiDAR maps using direct 2D-3D line correspondences. To handle the appearance differences and modality gaps between LiDAR point clouds and images, geometric 3D lines are extracted offline from LiDAR maps while robust 2D lines are extracted online from video sequences. With the pose prediction from VIO, we can efficiently obtain coarse 2D-3D line correspondences. Then the camera poses and 2D-3D correspondences are iteratively optimized by minimizing the projection error of correspondences and rejecting outliers. Experimental results on the EurocMav dataset and our collected dataset demonstrate that the proposed method can efficiently estimate camera poses without accumulated drifts or pose jumps in structured environments. Huai Yu, Weikun Zhen, Wen Yang 0001, Ji Zhang 0003, Sebastian A. Scherer |
IROS | 3 |
| 2020 | Mining Deep Semantic Representations for Scene Classification of High-Resolution Remote Sensing ImageryabstractScene classification is one of the most fundamental task in interpretation of high-resolution remote sensing (HRRS) images. Many recent works show that the probabilistic topic models which are capable of mining latent semantics of images can be effectively applied to HRRS scene classification. However, the existing approaches based on topic models simply utilize low-level hand-crafted features to form semantic features, which severely limit the representative capability of the semantic features derived from topic models. To alleviate this problem, this paper propose to build powerful semantic features using the probabilistic latent semantic analysis (pLSA) model, by employing the pre-trained deep convolutional neural networks (CNNs) as feature extractors rather than relying on the hand-crafted features. Specifically, we develop two methods to generate semantic features, called multi-scale deep semantic representation (MSDS) and multi-level deep semantic representation (MLDS), by extracting CNN features from different layers: (1) in MSDS, the final semantic features are learned by the pLSA with multi-scale features extracted from the convolutional layer of a pre-trained CNN; (2) in MLDS, we extract CNN features for densely sampled image patches at different size level from the fully-connected layer of a pre-trained CNN, and concatenate the sematic features learned by the pLSA at each level. We comprehensively evaluate the two methods on two public HRRS scene datasets, and achieve significant performance improvement over the state-of-the-art. The outstanding results demonstrate that the pLSA model is capable of discovering considerably discriminative semantic features from the deep CNN features. Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001 |
IEEE Trans. Big Data | 3 |
| 2020 | Mental Retrieval of Remote Sensing Images via Adversarial Sketch-Image Feature LearningabstractSearching the targets of interest in large-scale remote sensing images is a fundamental problem, which becomes a very challenging issue when there is no relevant example at hand but a mental picture in mind. Hand-drawn sketch as a precise and convenient expression of the mental picture makes sketch-based remote sensing image retrieval (SBRSIR) an ideal choice to cope with this issue. However, the accuracy of SBRSIR algorithm, which is critical to effective retrieval, is still far behind classic query-by-example image retrieval. Two central limiting factors for this performance gap are: 1) the lack of effective cross-domain representations for bridging the domain gap between sketches and remote sensing images and 2) the absence of large sketch/remote sensing image data sets for developing, evaluating, and comparing the SBRSIR approaches. In this article, we first develop a novel SBRSIR model to learn a deep joint embedding space with discriminative losses, where adversarial training is used for the embedding space to learn domain-invariant representations. Then, we contribute a sketch/remote sensing image data set specifically for SBRSIR and provide a benchmark for subsequent researchers. Extensive experiments on the data set and large scene images demonstrate the effectiveness and superiority of the method for both the seen and unseen categories. Wen Yang 0001, Tianbi Jiang, Shijie Lin, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Context Aggregation Network for Semantic Labeling in Aerial ImagesabstractMulti-scale object recognition and accurate object localization are two major problems for semantic segmentation in high resolution aerial images. To handle these problems, we design a Context Fuse Module to aggregate multi-scale features and propose an Attention Mix Module to combine different level features for higher localization accuracy. We further employ a Residual Convolutional Module to refine features in all levels. Based on these modules, we construct a new end-to-end network for semantic labeling in aerial images. Experiments demonstrate that our network outperforms other state-of-the-art models on the large-scale ISPRS Vaihingen 2D Semantic Labeling Challenge dataset. The model implementation code is made publicly available1. Wensheng Cheng, Wen Yang 0001, Youqi Pan, Haowen Guo |
ICIP | 2 |
| 2019 | Detecting and Positioning of Wind Turbine Blade Tips for UAV-Based Automatic InspectionabstractWind turbine detecting and blade tip positioning are essential parts of UAV-based automatic inspection. Previous methods which adopt LiDAR or ultrasonic sensor may cause cost increasement and system instability. We propose a vision-based approach for blade tip detecting and positioning. More precisely, we first detect each structure of wind turbine by combining Mask R-CNN detector and shape constraints. The pixel coordinates of blade tips can be extracted and then be used to solve a PnP problem. Finally we get GPS coordinates of blade tips which are useful for achieving path planning during the inspection. We evaluate the method on a well-annotated wind turbine datasets. The experimental results show the effectiveness and feasibility of the proposed method. Haowen Guo, Qiangqiang Cui, Jinwang Wang, Wen Yang 0001, Zhengrong Li |
IGARSS | 5 |
| 2019 | Mental Retrieval of Large-Scale Satellite Images Via Learned Sketch-Image Deep FeaturesabstractSearching targets of interest in large-scale satellite images is an imperative task, which becomes a challenging issue when the targets reside only in the mind of the user as a set of subjective visual patterns. In this paper, we take the advantage of hand-drawn sketches' strong intuition of describing mental target to address the problem of no available exemplar query. We introduce a multi-level-of-detail model to learn a cross-domain representation for bridging the gap between sketches and satellite images. To train the model, we propose a novel method of generating satellite images with corresponding level of details based on generative adversarial network. Experiments on both large-scale satellite images and commonly used RS datasets demonstrate the effectiveness and superiority of our method. Ruixiang Zhang, Wen Yang 0001, Gui-Song Xia |
IGARSS | 3 |
| 2019 | Combined Convolutional and Structured Features for Power Line Detection in UAV ImagesabstractPower line detection plays an important role in automated UAV inspection system, which is crucial for real-time motion planning and navigation along power lines. Previous methods which adopt traditional filters and gradients may fail to capture complete power lines due to noisy background. To overcome this, we develop an accurate power line detection method using rich convolutional and structured features. The proposed method fully exploits multiscale and structured prior information to conduct both accurate and efficient detection. We evaluate the method on two well-annotated power line datasets and achieve state-of-the-art performance compared with previous methods. Heng Zhang 0011, Wen Yang 0001, Huai Yu |
IGARSS | 2 |
| 2019 | Order-aware Embedding Neural Network for CTR PredictionabstractProduct based models, which represent multi-field categorical data as embedding vectors of features, then model feature interactions in terms of vector product of shared embedding, have been extensively studied and have become one of the most popular techniques for CTR prediction. However, if the shared embedding is applied: (1) the angles of feature interactions of different orders may conflict with each other, (2) the gradients of feature interactions of high-orders may vanish, which result in learned feature interactions less effective. To solve these problems, we propose a novel technique named Order-aware Embedding (i.e., multi-embeddings are learned for each feature, and different embeddings are applied for feature interactions of different orders), which can be applied to various models and generates feature interactions more effectively. We further propose a novel order-aware embedding neural network (OENN) based on this embedding technique for CTR prediction. Extensive experiments on three publicly available datasets demonstrate the effectiveness of Order-aware Embedding and show that our OENN outperforms the state-of-the-art models. Wei Guo 0006, Ruiming Tang, Huifeng Guo, Jianhua Han, Wen Yang 0001 |
SIGIR | 5 |
| 2019 | Unsupervised Change Detection of SAR Images Based on Variational Multivariate Gaussian Mixture Model and Shannon EntropyabstractIn this letter, we propose an unsupervised change detection method for synthetic aperture radar (SAR) images based on variational multivariate Gaussian mixture model (MGMM) and Shannon entropy. First, the difference features are generated from the Gabor wavelet transform of two SAR images. In variational inference framework, the variational MGMM is first introduced to implement accurate modeling for the data distribution of difference features and to output responsibilities. Subsequently, spatial information is explored on the responsibilities to yield thecontextual responsibilitiesfor improving the accuracy and reliability of change detection. Then,a posterioriprobabilities of the changed and unchanged classes are derived from thecontextual responsibilities, and Shannon entropy, being directly related to the classification error rate, is proposed to determine the optimal index integer. Finally, the binary change mask is achieved by separating the pixels into the changed and unchanged classes. The experiments on three pairs of SAR images for describing urban sprawl and water bodies demonstrate the effectiveness of the proposed method. Gang Yang 0006, Heng-Chao Li 0001, Wen Yang 0001, Kun Fu 0001, Yong-Jian Sun, William J. Emery |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Unsupervised Change Detection Based on a Unified Framework for Weighted Collaborative Representation With RDDL and Fuzzy ClusteringabstractIn this paper, we propose a novel unsupervised change detection method of remote sensing (RS) images based on a unified framework for weighted collaborative representation (WCR) with robust deep dictionary learning (RDDL) and fuzzy clustering. Specifically, WCR is employed to collaboratively represent neighborhood features with lower computational complexity, for which the RDDL model is built to learn more effective and representative overcomplete dictionary and enhance the robustness against the noise and outliers. Meanwhile, in order to make the resulting collaborative coefficients more beneficial for clustering, the unified framework for WCR with RDDL and fuzzy clustering is designed. By doing so, our framework not only precludes the utilization of third-party clustering algorithm, but also achieves better detection performance. Subsequently, the spatial constraint is enforced on the membership matrix to yield the updated one for further improving the accuracy of change detection. Finally, a binary change mask (CM) is achieved by assigning the pixels into the changed and unchanged classes. Experiments are performed on five pairs of RS images, and experimental results demonstrate the effectiveness of the proposed method. Gang Yang 0006, Heng-Chao Li 0001, Wei-Ye Wang, Wen Yang 0001, William J. Emery |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Bottle Detection in the Wild Using Low-Altitude Unmanned Aerial VehiclesabstractIn this paper, we propose a new dataset and benchmark for low altitude UAV object detection, aiming to find and localize waste plastic bottles in the wild, as well as to inspire the development of object detection models to be capable of detecting small and transparent objects. To this end, we collect 25, 407 UAV images of bottles with various kinds of backgrounds. Unlike traditional horizontal bounding box based annotation methods, we use the oriented bounding box to accurately and compactly annotate the bottles, which provides more detailed information for subsequent robotic grasping. The fully annotated images contain 34, 791 bottles, each of which is annotated by an arbitrary (5 d.o.f.) quadrilateral. To build a baseline for bottle detection, we evaluate several state-of-the-art object detection algorithms on our UAV-Bottle Dataset (UAV-BD), such as Faster R-CNN, SSD, YOLOv2 and RRPN. We also present an analysis of the dataset along with baseline approaches. Both the dataset and benchmark are made publicly available to the vision community on our website to advance research in the area of object detection from UAVs. Jinwang Wang, Wei Guo 0006, Huai Yu, Lin Duan, Wen Yang 0001 |
FUSION | 6 |
| 2018 | Accurate Registration of Multitemporal UAV Images Based on Detection of Major ChangesabstractAccurate registration of multitemporal images captured by UAV usually involves affine transformation and complicated non-rigid transformation, which makes it very difficult to achieve satisfying results. Pixel-wise correspondence is effective for handling images with complicated non-rigidity. However, objects with changes in the scene deform severely because there should be no pixel-wise correspondence but the algorithm erroneously matches the pixels. In this paper, we propose a coarse-to-fine registration method for multitemporal UAV images. First, a projective model is used to eliminate large scale changes as well as perspective distortion. Then the major changes of different temporal UAV images are most detected, which is used to mask the dense matches in changed areas. Finally, the optical flow field method is used to handle complicated non-rigid changes by matching dense SIFT feature. Experimental results on a challenging set of multitemporal UAV images demonstrate the effectiveness of our approach. Huai Yu, Jinwang Wang, Wen Yang 0001 |
FUSION | 4 |
| 2018 | Analysis of P-Band Repeat-Pass SAR Tomography Under Changing Weather ConditionsabstractIn this paper, the impact of changing weather conditions on repeat pass SAR tomography is addressed to support the upcoming spaceborne mission BIOMASS. In recent years it has been demonstrated that forest biomass retrieval can be improved by using P-band SAR tomography in tropical forest. Yet, these results were obtained by using campaign data acquired in a single day, while the revisit time of BIOMASS mission will be 3-4 days. To fill this gap, we simulate BIO-MASS repeat pass tomography using ground-based TropiS-CAT data with revisit time of 3 days and rainy days included. It is observed that the backscattered power within canopy layer, which is significantly correlated to the forest biomass, stays stable under changing weather conditions. The backscattered power variation of canopy layer are within 1.5 dB. For this forest site, this error is translated into an AGB error of about 50-80 t/ha, which is 20% or less of forest AGB. Yu Bai 0007, Stefano Tebaldini, Ho Tong Minh Dinh, Wen Yang 0001 |
IGARSS | 4 |
| 2018 | Recent Advances and Opportunities in Scene Classification of Aerial Images with Deep ModelsabstractScene classification is a fundamental task in interpretation of remote sensing images, and has become an active research topic in remote sensing community due to its important role in a wide range of applications. Over the past years, tremendous efforts have been made for developing powerful approaches for scene classification of remote sensing images, evolving from the traditional bag-of-visual-words model to the new generation deep convolutional neural networks (CNNs). The deep CNN based methods have exhibited remarkable breakthrough on performance, dramatically outperforming previous methods which strongly rely on hand-crafted features. However, performance with deep CNNs has gradually plateaued on existing public scene datasets, due to the notable drawbacks of these datasets, such as the small scale and low-diversity of training samples. Therefore, to promote the development of new methods and move the scene classification task a step further, we deeply discuss the existing problems in scene classification task, and accordingly present three open directions. We believe these potential directions will be instructive for the researchers in this field. Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2018 | Deep Semi-Nonnegative Matrix Factorization Based Unsupervised Change Detection of Remote Sensing ImagesabstractIn the paper, an unsupervised change detection method for remote sensing (RS) images based on deep semi-nonnegative matrix factorization (semi-NMF) is proposed. Firstly, the difference image is generated in different ways, depending on the types of input images. Then principal component analysis (PCA) is applied on the difference image to form the feature matrix X for improving the capability against various noise. In order to exploit more useful information from the resulting feature matrix, deep semi-NMF is introduced to factorize X into L+1 factors consisting of L nonrestricted matrices {Fl}l=1Land nonnegative cluster indicator matrix GL. Finally, the binary change mask (CM) is generated by assigning the pixels into changed and unchanged classes according to maximum criterion. The experimental results on two pairs of multitemporal RS images demonstrate the effectiveness of the proposed method. Gang Yang 0006, Heng-Chao Li 0001, Wen Yang 0001, William J. Emery |
IGARSS | 3 |
| 2017 | Accurate object matching for UAV imagery using multi-scale best-buddies similarityabstractThis paper presents a multi-scale Best-Buddies Similarity (BBS) method for matching objects in the UAV imagery. More precisely, we first extract proposed regions from the target image by Selective Search, then take a number of proposed regions with the highest confidence scores as templates and finally apply templates to the original BBS method and select the coordinate of the region with the highest overlap rate as the exact position of the query object in the target image. Experimental results on real data show the effectiveness of the proposed method. Xu Lei 0002, Jinwang Wang, Kaimin Fu, Huai Yu, Wen Yang 0001 |
IGARSS | 5 |
| 2017 | Stable feature point extraction for accurate multi-temporal SAR image registrationabstractFeature extraction is an important issue for image interpretation, many valuable feature extraction methods have been proposed to address synthetic aperture radar (SAR) image registration. However, the current methods care little about changes over multi-temporal SAR images, which result in unstable output features. The unstable property is a latent factor to affect the accuracy of SAR image registration. To overcome this problem, we propose a new method to detect stable features by intersecting Coherent Scatters (CS) and SAR-FAST corners. Thus the stable features are not only corners but also located at the time-invariant area. Then, the stable features are described by SIFT descriptors and the coarse registration is achieved by matching these stable points. Finally, the Powell algorithm is used to search the optimal Mutual Information location for precise registration. Experimental results demonstrate the effectiveness of the proposed method. Huai Yu, Wen Yang 0001, Mingsheng Liao |
IGARSS | 4 |
| 2017 | A UAV-based crack inspection system for concrete bridge monitoringabstractCrack inspection is one of the most important tasks for concrete bridge monitoring. Traditionally, human-based inspection systems using people's eyes are unsafe and costly. To overcome this problem, a novel unmanned aerial vehicle (UAV) based inspection system is developed to detect cracks on the lateral sides and underside of bridges. Applying obstacle avoidance modules, UAV can reach bridges' lateral sides and underside within a safe distance. Thus the onboard camera can capture images of bridge surface in a high resolution. To locate the positions of cracks and solve the problem of limited horizon of single image, a fast feature-based stitching algorithm is developed. Then an edge detection method using structured forests is utilized to detect cracks on the large panorama. With image distortion corrected, the positions of cracks can be located on the large panorama. Experimental results validate the applicability and efficiency of the proposed UAV-based crack inspection system. Huai Yu, Wen Yang 0001, Heng Zhang 0011, Wanjun He |
IGARSS | 2 |
| 2017 | Accurate extraction of cracks on the underside of concrete bridgesabstractCrack detection is crucial for maintaining the structural health and safety of concrete bridges. Previous studies mainly focus on the detection of thick and conspicuous cracks on bridge deck using high resolution images. However, it's difficult to obtain accurate crack detection results on the underside of bridges due to the fainter and thinner appearance and lower contrast of cracks to their background. To overcome this problem, we propose a modified Beamlet Tree-based crack detection method which searches for possible edges by hierarchically constructing difference filters. The proposed method is evaluated with real crack images on the underside of bridges collected by an unmanned aerial vehicle (UAV). Compared with competitors, our proposed method achieves better performance in terms of accuracy and completeness. Heng Zhang 0011, Huai Yu, Wen Yang 0001 |
IGARSS | 4 |
| 2017 | Unsupervised Classification of Polarimetric SAR Images via Riemannian Sparse CodingabstractUnsupervised classification plays an important role in understanding polarimetric synthetic aperture radar (PolSAR) images. One of the typical representations of PolSAR data is in the form of Hermitian positive definite (HPD) covariance matrices. Most algorithms for unsupervised classification using this representation either use statistical distribution models or adopt polarimetric target decompositions. In this paper, we propose an unsupervised classification method by introducing a sparsity-based similarity measure on HPD matrices. Specifically, we first use a novel Riemannian sparse coding scheme for representing each HPD covariance matrix as sparse linear combinations of other HPD matrices, where the sparse reconstruction loss is defined by the Riemannian geodesic distance between HPD matrices. The coefficient vectors generated by this step reflect the neighborhood structure of HPD matrices embedded in the Euclidean space and hence can be used to define a similarity measure. We apply the scheme for PolSAR data, in which we first oversegment the images into superpixels, followed by representing each superpixel by an HPD matrix. These HPD matrices are then sparse coded, and the resulting sparse coefficient vectors are then clustered by spectral clustering using the neighborhood matrix generated by our similarity measure. The experimental results on different fully PolSAR images demonstrate the superior performance of the proposed classification approach against the state-of-the-art approaches. Neng Zhong, Wen Yang 0001, Anoop Cherian, Xiangli Yang, Gui-Song Xia, Mingsheng Liao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Automatic extraction of linear arranged targets from polarimetric SAR imageryabstractIn this paper we propose a new extraction scheme for linear arranged targets in polarimetric SAR images based on a contrario theory. In this scheme, to reduce the influence of speckle, firstly a polarimetric whitening filter is applied to combine four images from different channels into a single channel image. Then a Cell-Averaging Constant False Alarm Rate detector with Weibull clutter background and a post-processing operator is used to detect point-like targets in the image. Finally, a searching approach based on contrario theory is applied to extract the linear arranged targets from other point-like targets. The experimental results varify the effectiveness in linear arranged targets extraction. Wei Guo 0006, Kaimin Fu, Pingping Huang, Wen Yang 0001 |
IGARSS | 4 |
| 2016 | Polarimetric SAR speckle filtering based on stochastic samplingabstractSpeckle noise is an inherent problem in synthetic aperture radar (SAR) imaging system. For further imagery analysis and interpretation, it demands better and more efficient polarimetric SAR (PolSAR) speckle-filtering algorithms. Inspired by the great success of stochastic denoising, our goal in this paper is to reduce speckle noise of PolSAR image with random walk model. Taking spatial distance into consideration, limited steps of random walks over small pixel neighbourhoods provides good stochastic samples, then the final estimation of denoised image based on transition probabilities is obtained. In addition, since polarimetric distance is used to measure the dissimilarities between pixels, the proposed method exploits both the spatial and polarimetric information. Despeckling results on simulated SAR data and real SAR data demonstrate that our method provides efficient speckle reduction as well as spatial resolution and polarimetric information preservation. Tianheng Yan, Xueke Yin, Wen Yang 0001, Carlos López-Martínez |
IGARSS | 3 |
| 2016 | Fusion of intensity/coherent information using region covariance features for unsupervised classification of SAR imageryabstractUnsupervised classification of synthetic aperture radar (SAR) imagery is an essential step in SAR image interpretation. There is a growing demand for an efficient way to fuse multi-information of SAR imagery. This paper presents an intensity/coherent information fusion algorithm by using region covariance features for unsupervised classification. More precisely, we firstly extract the intensity properties and coherent characteristics from each pixel of SAR imagery, then use the region covariance descriptor to fuse the intensity and coherent features, and finally exploit the K-means algorithm to obtain the final unsupervised classification map. Experimental results on SAR imagery demonstrate the effectiveness of the proposed fusion scheme. Xiangli Yang, Shangtan Tu, Yu Bai 0007, Wen Yang 0001 |
IGARSS | 4 |
| 2016 | Riemannian sparse coding for classification of PolSAR imagesabstractHermitian positive definite (HPD) covariance matrices form one of the most widely-used data representations in PolSAR applications. However, most of these applications either use statistical distribution models on the PolSAR covariance matrices or polarimetric target decomposition. In this paper, we study HPD matrices for PolSAR image classification in the context of sparse coding. More specifically, the PolSAR HPD matrices are first represented as sparse linear combinations of elements from a dictionary, where each element itself is an HPD matrix and the representation loss is measured by the affine-invariant Riemannian metric. We then introduce a sparsity induced similarity measure between two HPD matrices. Finally, we propose a supervised classification scheme using support vector machines on the Riemannian sparse codes and an unsupervised classification scheme encompassing a sparsity induced similarity measure followed by spectral clustering. The proposed methods are validated on the NASA/JPL AIRSAR fully PolSAR data. The experimental results demonstrate the effectiveness of our methods. Wen Yang 0001, Neng Zhong, Xiangli Yang, Anoop Cherian |
IGARSS | 1 |
| 2016 | Unsupervised Classification of PolSAR Imagery via Kernel Sparse Subspace ClusteringabstractUnsupervised classification is very important for the fully polarimetric synthetic aperture radar (PolSAR) image interpretation. The PolSAR covariance matrices, as one of the most widely used representations for PolSAR data, are Hermitian positive definite (HPD) and form a Riemannian manifold when endowed with an appropriate metric. Considering their geometric properties, we propose a new clustering algorithm by embedding the HPD matrices into Hilbert space and introduce sparse subspace clustering in the newly formed highly dimensional space to recover the latent cluster structure. Moreover, an improved scalable scheme is presented to classify large-scale PolSAR images, which involves dictionary learning and spatially reinforced joint coding for robustness against the speckle noise. Experimental results on real fully PolSAR data sets demonstrate the effectiveness of the proposed method. Wen Yang 0001, Neng Zhong, Xin Xu 0005 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Meaningful Object Segmentation From SAR Images via a Multiscale Nonlocal Active Contour ModelabstractThe segmentation of synthetic aperture radar (SAR) images is a long-standing yet challenging task, not only because of the presence of speckle but also due to the variations of surface backscattering properties in the images. Tremendous investigations have been made to suppress the speckle effects for the segmentation of SAR images, whereas few works are devoted to dealing with the variations of backscattering intensities in the images. To overcome the two difficulties, this paper presents a novel SAR image segmentation method by exploiting a multiscale active contour model based on the nonlocal processing principle. More precisely, we first formulize the SAR segmentation problem with an active contour model by integrating the nonlocal interactions between pairs of patches inside and outside the segmented regions. Second, a multiscale strategy is proposed to speed up the nonlocal active contour segmentation procedure and to avoid falling into a local minimum for achieving more accurate segmentation results. Experimental results on simulated and real SAR images demonstrate the efficiency and feasibility of the proposed method: It can not only achieve precise segmentations for images with heavy speckle and nonlocal intensity variations but also be used for SAR images from different types of sensors. Gui-Song Xia, Gang Liu 0013, Wen Yang 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Region-Based Change Detection for Polarimetric SAR Images Using Wishart Mixture ModelsabstractThe change detection of polarimetric synthetic aperture radar (PolSAR) images is a longstanding and challenging task, not only because of the speckle issue but also due to the complex texture, which generally appears highly heterogeneous. There are two widely used approaches for the change detection of PolSAR images: one is the post classification comparison algorithm, and the other is the directly unsupervised change detection algorithm. In this paper, we focus on the latter and propose a region-based change detection method for PolSAR images by means of Wishart mixture models (WMMs). The WMMs fit the distribution of PolSAR images with less errors both in the homogeneous and the extremely heterogeneous area. More precisely, two PolSAR images are first segmented into compact local regions using the customized simple-linear-iterative-clustering algorithm, while the WMMs are used to model each local region. To generate a difference map, statistical distribution differences measured by information theoretic divergence are then computed for corresponding local region pairs. The Cauchy-Schwarz divergence is adopted as its analytic expression can be derived for WMMs. Finally, the change detection results are obtained by the Kittler-Illingworth thresholding method with Markov random field-based smoothing. The proposed scheme is tested on different PolSAR data sets. Qualitative and quantitative evaluations show its superior performance comparing to the traditional pixel-level approach. Wen Yang 0001, Xiangli Yang, Tianheng Yan, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | A novel polarimetric-texture-structure descriptor for high-resolution PolSAR image classificationabstractA novel Polarimetric-Texture-Structure descriptor for high-resolution PolSAR image is presented in this paper. More precisely, a PolSAR image is represented by a tree of shapes, each of which is associated with several polarimetric and texture attributes. We first extract the texture properties and polarimetric characteristics from each shape, then use the shape co-occurrence patterns (SCOPs) to characterize the shape relationships, and finally use the resulting SCOPs distributions as features for PolSAR image classification. The proposed method not only has the strong ability to depict the texture and polarimetric properties, but also encodes the shape relationships on the tree. We compare the proposed method with the cluster based statistical feature (CSF) and the scattering mechanism based statistical feature (SMSF). Experimental results on high-resolution PolSAR sample dataset and a large scene for classification demonstrate the effectiveness of the proposed method. Yu Bai 0007, Wen Yang 0001, Gui-Song Xia, Mingsheng Liao |
IGARSS | 2 |
| 2015 | Dissimilarity measurements for processing and analyzing PolSAR data: A surveyabstractMeasuring the pairwise similarity/dissimilarity of data is of central importance for processing and analyzing PolSAR images. In the literature, a large variety of measurements have been used, however, it is still not clear how to choose appropriate similarity measures for a given task. This paper presents a brief summary and discussion of the dissimilarity measurements used for interpreting PolSAR Images. Wen Yang 0001, Gui-Song Xia, Carlos López-Martínez |
IGARSS | 1 |
| 2015 | Learning High-level Features for Satellite Image Classification With Limited Labeled SamplesabstractThis paper presents a novel method addressing the classification task of satellite images when limited labeled data is available together with a large amount of unlabeled data. Instead of using semi-supervised classifiers, we solve the problem by learning a high-level features, called semisupervised ensemble projection (SSEP). More precisely, we propose to represent an image by projecting it onto an ensemble of weak training (WT) sets sampled from a Gaussian approximation of multiple feature spaces. Given a set of images with limited labeled ones, we first extract preliminary features, e.g., color and textures, to form a low-level image description. We then propose a new semisupervised sampling algorithm to build an ensemble of informative WT sets by exploiting these feature spaces with a Gaussian normal affinity, which ensures both the reliability and diversity of the ensemble. Discriminative functions are subsequently learned from the resulting WT sets, and each image is represented by concatenating its projected values onto such WT sets for final classification. Moreover, we consider that the potential redundant information existed in SSEP and use sparse coding to reduce it. Experiments on high-resolution remote sensing data demonstrate the efficiency of the proposed method. Wen Yang 0001, Xiaoshuang Yin, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Texture Analysis with Shape Co-occurrence PatternsabstractThis paper presents a flexible shape-based texture analysis method by investigating the co-occurrence patterns of shapes. More precisely, a texture image is represented by a tree of shapes, each of which is associated with several attributes. The modeling of texture is thus converted to characterize the tree of shapes. To this aim, we first learn a set of co-occurrence patterns of shapes from texture images, then establish a bag-of-words model on the learned shape co-occurrence patterns (SCOPs), and finally use the resulting SCOPs distributions as features for texture analysis. In contrast with existing work, the proposed method not only inherits the strong ability to depict geometrical aspects of textures and the high robustness to variations of imaging conditions from the shape-based texture analysis method, but also provides a more flexible way to model shape relationships (high-order statistics) on the tree. To our knowledge, this is the first time to use co-occurrence patterns of explicit shapes as a tool for texture analysis. Experiments of texture retrieval and classification on various databases report state-of-the-art results and demonstrate the efficiency of the proposed method. Gang Liu 0013, Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001 |
ICPR | 3 |
| 2014 | Hierarchical segmentation of polarimetric SAR image via Non-Parametric Graph EntropyabstractPolSAR image segmentation has long been an important problem in the PolSAR remote sensing community. Many segmentation algorithms describe images in terms of a hierarchy of regions has attracted particular attention in recent years. However, they often contain more data than is required for an efficient description. In this paper, we propose an effective measure to extract hierarchical semantic structures from PolSAR images. First, we construct the Binary partition tree (BPT) which is a multi-scale image representation to obtain a hierarchy of regions. Once the tree has been constructed, every hierarchy can be considered as a region adjacency graph (RAG). Second, we use a Non-Parametric Graph Entropy as a measure of graph complexity to identify semantic structures within BPT hierarchies. Experimental results on NASA/JPL AIRSAR and DLR E-SAR images demonstrate the effectiveness of the proposed approach. Yu Bai 0007, Lixia Dong, Wen Yang 0001, Mingsheng Liao |
IGARSS | 4 |
| 2014 | Saliency detection based on distance between patches in polarimetric SAR imagesabstractIn this paper, we propose a saliency detection model for polarimetric SAR images based on inter-patch distances. The model is biology-based as it takes the human visual properties into account. Our model consists of local and global saliency detection, which obtained by multi-scale information extraction. What's more, inspired by the distance measures of different image patches, the model takes full use of the coherency matrix which includes the statistical information of pixels and precisely measures the similarity between patches. The experimental results demonstrate the effectiveness of our method. Pingping Huang, Lixia Dong, Wen Yang 0001 |
IGARSS | 5 |
| 2014 | Unsupervised classification of PolSAR data using large scale spectral clusteringabstractIn this paper, a spectral clustering based unsupervised classification scheme is proposed for processing large scale polarimetric synthetic aperture radar (PolSAR) data. Due to its high computational complexity, spectral clustering can hardly handle large PolSAR image. To overcome this bottleneck, a representative points based scheme is introduced. Instead of building pairwise affinity graph on the whole data set, we first build a bipartite graph between data points and a small set of selected representative points. Then an approximate large graph is constructed based on this bipartite graph. After that, spectral analysis on the approximate graph is solved efficiently by singular value decomposition (SVD). To integral context information, Markov random fields (MRF) model based smoothing is also performed to get the final clusters. We test the proposed approach on DLR ESAR data set. Experimental results demonstrate its effectiveness and efficiency. Li-Qi Lin, Pingping Huang, Wen Yang 0001, Xin Xu 0005 |
IGARSS | 4 |
| 2014 | SAR image segmentation via non-local active contoursabstractThis paper presents a method for SAR image segmentation by relying on active contour model with the non-local processing principle [1]. The idea is to partition a SAR image via computing the patch similarity in the SAR image non-locally, and formulize the segmentation problem with an active contour model. More precisely, after computing the statistical features of SAR images, non-local comparisons between feature patches are used to calculate the active contour energy, which is defined by integrating the interactions between pairs of patches inside and outside the segmented region. A level set method is finally used to minimize the non-local energy. Compared with existing approaches for SAR image segmentation, the only requirement of this method is a local similarity between patches, and it is less sensitive to initial segmentation. The experimental results show the effectiveness and feasibility of the proposed method. Gang Liu 0013, Gui-Song Xia, Wen Yang 0001, Nan Xue 0001 |
IGARSS | 3 |
| 2014 | Semi-supervised feature learning for remote sensing image classificationabstractThis paper presents a semi-supervised method for learning informative image representations, which is a crucial but challenging step for remote sensing image classification. More precisely, we propose to represent an image by projecting it onto an ensemble of prototype sets sampled from a Gaussian approximation of multiple feature spaces. Given a set of images with a few labeled ones, we first extract preliminary features, e.g. color and textures, to form a low-level image description. We then build an ensemble of informative prototype sets by exploiting these feature spaces with a Gaussian normal affinity. Discriminative functions are subsequently learned from the resulting prototype sets, and each image is represented by concatenating their projected values onto such prototypes for final classification. Experiments on two high-resolution remote sensing image sets demonstrate the efficiency of the proposed method on remote sensing image classification with different classifiers. Xiaoshuang Yin, Wen Yang 0001, Gui-Song Xia, Lixia Dong |
IGARSS | 2 |
| 2013 | A Hierarchical Scheme of Multiple Feature Fusion for High-Resolution Satellite Scene Categorization
Wen Shao, Wen Yang 0001, Gui-Song Xia, Gang Liu 0013 |
ICVS | 2 |
| 2013 | The algorithm of building area extraction based on boundary prior and conditional random field for SAR imageabstractIn this paper, an algorithm applied for building area extraction on SAR image is proposed, which is based on conditional random model, then a boundary prior relation is introduced to strengthen the description of prior item around the edge of building area, aiming at improving the classification performance nearby the boundary lines encompass building area. Firstly, pre-segmentation and boundary lines extraction can be accomplished respectively rely on mean shift algorithm and ratio of average edge detection. After that a combination term of the distances between the boundary lines and pixels around them and the pixels' label information can help to improve the prior item in CRF and build the boundary prior-CRF model. Finally, several experimental results on TerraSAR-X images prove that the proposed approach significantly improves the extraction accuracy and classification performance when compared to CRF. Chu He, Yu Zhang 0019, Xin Su 0003, Wen Yang 0001, Xin Xu 0005 |
IGARSS | 5 |
| 2013 | A perception-inspired building index for automatic built-up area detection in high-resolution satellite imagesabstractThis paper addresses the problem of automatic extraction of built-up areas from high-resolution remote sensing images. We propose a new building presence index from the point view of perception. We argue that built-up areas usually result in significant corners and junctions in high-resolution satellite images, due to the man-made structures and occlusion, and thus can be measured by the geometrical structures they contained. More precisely, we first detect corners and junctions by relying on a perception-inspired corner detector, called an a-contrario junction detector. Each detected corner is associated with a perceptual significance, which measures the structural saliency of the corner in the image and is independent of the contrast and scale. All these detected corners together with their significance are then used to compute the building index. The proposed approach is evaluated on a high-resolution satellite image set, including 15 big images from GeoEye-1, QuickBird and IKONOS. The results demonstrated that our method achieves the state-of-the-art results and can be used in practical applications. Gang Liu 0013, Gui-Song Xia, Xin Huang 0002, Wen Yang 0001, Liangpei Zhang 0001 |
IGARSS | 4 |
| 2013 | Change detection in multi-temporal TerraSAR-X SAR images using a hierarchical Markov model on regionsabstractThis paper addresses the problem of change detection in high-resolution multi-temporal synthetic aperture radar (SAR) images (e.g. TerraSAR-X SAR images). Given two images, the proposed method first computes a difference map between them, by taking into account both the spatial and temporal correlations. Change detection is then formulated as a binary (changed/unchanged) segmentation problem of the difference map. A hierarchical Markov model (HMM) is defined on the multi-scale over-segmented regions of the difference map. The change map is finally inferred by relying on the hierarchical marginal posterior mode (HMPM) of the HMM. Experimental results on multi-temporal TerraSAR-X SAR images demonstrate the effectiveness and the reliability of the proposed approach. Wen Yang 0001, Gui-Song Xia, Mingsheng Liao |
IGARSS | 2 |
| 2013 | Unsupervised PolSAR image classification based on ensemble partitioningabstractThis work introduces an unsupervised classification framework based on ensemble partitioning for polarimetric synthetic aperture radar (PolSAR) data, which can automatically determine the number of categories. First, the PolSAR image is divided into patches by an over-segmentation method. Second, ensemble partitioning is performed on the patch based dataset to obtain an ensemble similarity matrix. Third, a self-tuning spectral clustering method is adopted to automatically find the number of categories and the classification results, which is finally smoothed by a Markov random field based method. The experimental results on PolSAR image show the effectiveness of this unsupervised classification method. Xiaoshuang Yin, Wen Yang 0001, Chu He, Xin Xu 0005 |
IGARSS | 3 |
| 2013 | Unsupervised Satellite Image Classification Using Markov Field Topic ModelabstractRecently, the combination of topic models and random fields has been frequently and successfully applied to image classification due to their complementary effect. However, the number of classes is usually needed to be assigned manually. This letter presents an efficient unsupervised semantic classification method for high-resolution satellite images. We add label cost, which can penalize a solution based on a set of labels that appear in it by optimization of energy, to the random fields of latent topics, and an iterative algorithm is thereby proposed to make the number of classes finally be converged to an appropriate level. Compared with other mentioned classification algorithms, our method not only can obtain accurate semantic segmentation results by larger scale structures but also can automatically assign the number of segments. The experimental results on several scenes have demonstrated its effectiveness and robustness. Kan Xu, Wen Yang 0001, Gang Liu 0013 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2012 | Car detection from high-resolution aerial imagery using multiple featuresabstractDetecting cars in high-resolution aerial images has attracted particular attention in recent years. However, scene complexity, large illumination change and occlusions make the task very challenging. In this paper, we propose a robust and effective framework for car detection from high-resolution aerial imagery. More specifically, we first incorporate multiple diverse and complementary image descriptors, Histogram of Oriented Gradients (HOG), Local Binary Pattern (LBP) and Opponent Histogram. Subsequently taking computational efficiency and runtime complexity into account, we adopt an interactive bootstrapping approach to collect hard negatives for training an intersection kernel support vector machine (IKSVM). After training, detection is performed by exhaustive search. Finally for post-processing, we employ a greedy procedure for eliminating repetitive detections via non-maximum suppression. Furthermore, contextual information is utilized to refine the detections. Experimental results on Vaihingen dataset have demonstrated that the proposed method can achieve state-of-the-art performance in various real scenes. Wen Shao, Wen Yang 0001, Gang Liu 0013 |
IGARSS | 2 |
| 2012 | Laplacian Eigenmaps-Based Polarimetric Dimensionality Reduction for SAR Image ClassificationabstractIn this paper, we propose a novel scheme of polarimetric synthetic aperture radar (PolSAR) image classification. We apply Laplacian eigenmaps (LE), a nonlinear dimensionality reduction (NDR) technique, to a high-dimensional polarimetric feature representation for PolSAR land-cover classification. A wide variety of polarimetric signatures are chosen to construct a high-dimensional polarimetric manifold which can be mapped into the most compact low-dimensional structure by manifold-based dimensionality reduction techniques. This NDR technique is employed to obtain a low-dimensional intrinsic feature vector by the LE algorithm, which is beneficial to PolSAR land-cover classification owing to its local preserving property. The effectiveness of our PolSAR land-cover classification scheme with LE intrinsic feature vector is demonstrated with the RadarSat-2 C-band PolSAR data set and the 38th Research Institute of China Electronics Technology Group Corporation X-band PolInSAR data set. The performance of our method is measured by the separability in the feature space and the accuracy of classification. Comparisons on the feature space show that the LE intrinsic feature vector is more separable than different original feature vectors. Our LE intrinsic feature vector also improves the classification accuracy. Shangtan Tu, Wen Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2012 | SAR-Based Terrain Classification Using Weakly Supervised Hierarchical Markov Aspect ModelsabstractWe introduce the hierarchical Markov aspect model (HMAM), a computationally efficient graphical model for densely labeling large remote sensing images with their underlying terrain classes. HMAM resolves local ambiguities efficiently by combining the benefits of quadtree representations and aspect models-the former incorporate multiscale visual features and hierarchical smoothing to provide improved local label consistency, while the latter sharpen the labelings by focusing them on the classes that are most relevant for the broader local image context. The full HMAM model takes a grid of local hierarchical Markov quadtrees over image patches and augments it by incorporating a probabilistic latent semantic analysis aspect model over a larger local image tile at each level of the quadtree forest. Bag-of-word visual features are extracted for each level and patch, and given these, the parent-child transition probabilities from the quadtree and the label probabilities from the tile-level aspect models, an efficient forwards-backwards inference pass allows local posteriors for the class labels to be obtained for each patch. Variational expectation-maximization is then used to train the complete model from either pixel-level or tile-keyword-level labelings. Experiments on a complete TerraSAR-X synthetic aperture radar terrain map with pixel-level ground truth show that HMAM is both accurate and efficient, providing significantly better results than comparable single-scale aspect models with only a modest increase in training and test complexity. Keyword-level training greatly reduces the cost of providing training data with little loss of accuracy relative to pixel-level training. Wen Yang 0001, Dengxin Dai, Bill Triggs, Gui-Song Xia |
IEEE Trans. Image Process. | 1 |
| 2011 | Polsar scene classification based on fast approximate neareast neighbours searchabstractIn this paper, we attempt to solve the efficiency problem of PolSAR scene classification with non-parametric classifier. We employ the tree-structure based search strategy to perform fast approximate nearest neighbour search by introducing the multiple randomized kd-tree and hierarchical kmeans-tree into ONBNN classifier. The experimental results on RadarSat-2 PolSAR dataset demonstrate that our method can greatly improve the efficiency with a guaranteed accuracy. Wen Yang 0001 |
IGARSS | 2 |
| 2011 | Satellite Image Classification via Two-Layer Sparse Coding With Biased Image RepresentationabstractThis letter presents a method for satellite image classification aiming at the following two objectives: 1) involving visual attention into the satellite image classification; biologically inspired saliency information is exploited in the phase of the image representation, making our method more concentrated on the interesting objects and structures, and 2) handling the satellite image classification without the learning phase. A two-layer sparse coding (TSC) model is designed to discover the “true” neighbors of the images and bypass the intensive learning phase of the satellite image classification. The underlying philosophy of the TSC is that an image can be more sparsely reconstructed via the images (sparse I) belonging to the same category (sparse II). The images are classified according to a newly defined “image-to-category” similarity based on the coding coefficients. Requiring no training phase, our method achieves very promising results. The experimental comparisons are shown on a real satellite image database. Dengxin Dai, Wen Yang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2011 | Multilevel Local Pattern Histogram for SAR Image ClassificationabstractIn this letter, we propose a theoretically and computationally simple feature for synthetic aperture radar (SAR) image classification, the multilevel local pattern histogram (MLPH). The MLPH describes the size distributions of bright, dark, and homogenous patterns appearing in a moving window at various contrasts; these patterns are the elementary properties of SAR image texture. The MLPH is a very powerful descriptor of SAR images because it captures both local and global structural information. Additionally, it is robust to speckle noise. Experiments on a TerraSAR-X data set demonstrate that MLPH significantly outperforms four other widely used features in SAR image classification. Dengxin Dai, Wen Yang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2010 | Active Contours with Fitting Term Driven by Region and Edge Information
Gang Chen 0016, Wen Yang 0001, Chu He, Tai Hu, Shuai Fang |
ICIP | 3 |
| 2010 | Topographic gray level multiscale analysis and its application to histogram modificationabstractThis paper describes a framework for multi-scale gray level analysis of images. It defines scales based on gray levels and organizes the basic “atoms” with a topographic map. The aim of this approach is to separate a large number of pixels concentrating in a narrow range of gray values. The main advantage of the methodology is that it allows manipulating pixels according to gray levels and spatial relations simultaneously. We apply it to histogram modification of Synthetic Aperture Radar (SAR) images. The experiments on displaying and classification prove the superiority of the approach. Chu He, Xinping Deng, Gui-Song Xia, Wen Yang 0001 |
ICIP | 4 |
| 2010 | Fast semantic scene segmentation with conditional random fieldabstractIn this paper, we present a fast approach to obtain semantic scene segmentation with high precision. We employ a two-stage classifier to label all image pixels. First, we use the regularized logistic regression to combine different appearance-based features and the improved spatial layout of labeling information. In the second stage, we incorporate the local, regional and global cues into a conditional random field model to provide a final segmentation, and a fast max-margin training method is employed to learn the parameters of the model quickly. The comparison experiments on four multi-class image segmentation databases show that our approach can achieve comparable semantic segmentation results and work faster than that of the state-of-the-art approaches. Wen Yang 0001, Dengxin Dai, Bill Triggs, Gui-Song Xia, Chu He |
ICIP | 1 |
| 2010 | Active Contours with Thresholding Value for Image SegmentationabstractIn this paper, we propose an active contour with threshold value to detect objects and at the same time get rid of unimportant parts rather than extract all information. The basic ideal of our model is to introduce a weight matrix into region-based active contours, which can enhance the weight for the main parts while filter the weak intensity, such as shadows, illumination and so on. Moreover, we can choose threshold value to set weight matrix manually for accurate image segmentation. Thus, the proposed method can extract objects of interest in practice. Coupled partial differential equations are used to implement this method with level set algorithms. Experimental results show the advantages of our method in terms of accuracy for image segmentation. Gang Chen 0016, Iron Chen, Wen Yang 0001 |
ICPR | 4 |
| 2010 | Three-layer Spatial Sparse Coding for Image ClassificationabstractIn this paper, we propose a three-layer spatial sparse coding (TSSC) for image classification, aiming at three objectives: naturally recognizing image categories without learning phase, naturally involving spatial configurations of images, and naturally counteracting the intra-class variances. The method begins by representing the test images in a spatial pyramid as the to-be-recovered signals, and taking all sampled image patches at multiple scales from the labeled images as the bases. Then, three sets of coefficients are involved into the cardinal sparse coding to get the TSSC, one to penalize spatial inconsistencies of the pyramid cells and the corresponding selected bases, one to guarantee the sparsity of selected images, and the other to guarantee the sparsity of selected categories. Finally, the test images are classified according to a simple image-to-category similarity defined on the coding coefficients. In experiments, we test our method on two publicly available datasets and achieve significantly more accurate results than the conventional sparse coding with only a modest increase in computational complexity. Dengxin Dai, Wen Yang 0001, Tianfu Wu 0001 |
ICPR | 2 |
| 2010 | Semantic segmentation of Polarimetric SAR imagery using Conditional Random FieldsabstractThe paper proposes a fast and accurate semantic segmentation approach for a large Polarimetric SAR (PolSAR) image using Conditional Random Fields (CRFs). It efficiently incorporates the polarimetric signatures, texture and intensity features into a unite CRFs model, and employs a fast max-margin training method for parameters learning. Experiments on RadarSat-2 PolSAR data in Flevoland test site demonstrate that our approach achieves precise segmentation results with a few well-selected training samples. Wen Yang 0001 |
IGARSS | 1 |
| 2006 | Classification of Polarimetric SAR Data Based on Multidimensional Watershed Clustering
Wen Yang 0001, Yongfeng Cao |
ADMA | 1 |
| 2005 | Extended multi-level logistic Model and SAR image segmentation
Yongfeng Cao, Chuanzhao Han, Wen Yang 0001 |
IGARSS | 4 |
| 2005 | Registration of high resolution SAR and optical images based on multiple features
Wen Yang 0001, Chuanzhao Han, Yongfeng Cao |
IGARSS | 1 |