VLDB 2026 Research / reviewers in the wild / expert
Gui-Song Xia
dblp:97/594 · also Guisong Xia
· DBLP profile ↗
196ranked-venue papers
13as first author
110since 2021 · last 2026
0000-0001-7660-6090ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 98 · 7 first-author · 65 since 2021Graphics, computer vision, multimedia, augmented reality and games · 78 · 7 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 74 · 3 first-author · 36 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven ConsistencyabstractSuper-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applications due to significant domain discrepancies between simulated degradations and real-world imaging conditions. To bridge this synthetic-to-real gap, we propose a novel Self-supervised Event-based SRB (SE-SRB) framework that leverages neuromorphic event streams as physical priors and adopts a lightweight neural architecture tailored for effective domain adaptation. Specifically, the proposed SE-SRB introduces a self-supervised learning paradigm based on asymmetric integral driven consistency, which enforces temporal coherence between predictions derived from RGB and asynchronous event streams at different time points. Extensive experiments validate that SE-SRB consistently outperforms state-of-the-art methods on both synthetic and real-world datasets. Built upon a lightweight parallel two-stream architecture, SE-SRB achieves high computational efficiency, featuring reduced parameter count, lower FLOPs, and real-time inference capability (40 FPS). Chi Zhang 0027, Xiang Zhang 0022, Lei Yu 0006, Gui-Song Xia, Yuming Fang 0001, Wenhan Yang |
AAAI | 4 |
| 2026 | UniCalib: Targetless LiDAR-camera Calibration via Probabilistic Flow on Unified Depth RepresentationsabstractOnline targetless extrinsic LiDAR-camera calibration is essential for robust perception in computer vision applications such as autonomous driving. However, existing methods struggle with the significant modality gap between heterogeneous sensors and fail to handle unreliable correspondences arising from real-world challenges like occlusions and dynamic objects. To address these issues, we introduce UniCalib, a novel method that performs calibration by estimating a probabilistic flow on unified depth representations. UniCalib first bridges the modality gap by converting both the camera images and the sparse LiDAR points into unified, dense depth maps, enabling a unified encoder to learn consistent features. Subsequently, it learns a probabilistic flow field that captures the correspondence uncertainty to improve robustness. This probabilistic approach is reinforced by a reliability map and a perceptually weighted sparse flow loss, which guide the model to suppress the influence of unreliable regions. Experimental results on three datasets validate the accuracy and generalization of UniCalib. In particular, it achieves a mean translation error of 0.550cm and a rotation error of 0.044° on the KITTI dataset. The code is available at https://github.com/han-15/UniCalib. Xubo Zhu, Ji Wu 0012, Ximeng Cai, Wen Yang 0001, Huai Yu, Gui-Song Xia |
WACV | 7 |
| 2026 | Change-aware multi-temporal cloud removal
Guochu You, Runmin Dong, Wen Yang 0001, Gui-Song Xia |
Sci. China Inf. Sci. | 6 |
| 2026 | Interacted Planes Reveal 3D Line Mappingabstract3D line mapping from multi-view RGB images provides a compact and structured visual representation of scenes. We study the problem from a physical and topological perspective: a 3D line most naturally emerges as the edge of a finite 3D planar patch. We present LiP-Map, a line-plane joint optimization framework that explicitly models learnable line and planar primitives. This coupling enables accurate and detailed 3D line mapping while maintaining strong efficiency (typically completing a reconstruction in 3 to 5 minutes per scene). LiP-Map pioneers the integration of planar topology into 3D line mapping, not by imposing pairwise coplanarity constraints but by explicitly constructing interactions between plane and line primitives, thus offering a principled route toward structured reconstruction in man-made environments. On more than 100 scenes from ScanNetV2, ScanNet++, Hypersim, 7Scenes, and Tanks&Temple, LiP-Map improves both accuracy and completeness over state-of-the-art methods. Beyond line mapping quality, LiP-Map significantly advances line-assisted visual localization, establishing strong performance on 7Scenes. Zeran Ke, Bin Tan 0002, Gui-Song Xia, Yujun Shen, Nan Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Seeing Through Satellite Images at Street ViewsabstractThis paper studies the task of SatStreet-view synthesis, which aims to render photorealistic street-view panorama images and videos given a satellite image and specified camera positions or trajectories. Our approach involves learning a satellite image conditioned neural radiance field from paired images captured from both satellite and street viewpoints, which comes to be a challenging learning problem due to the sparse-view nature and the extremely large viewpoint changes between satellite and street-view images. We tackle the challenges based on a task-specific observation that street-view specific elements, including the sky and illumination effects, are only visible in street-view panoramas, and present a novel approach, Sat2Density++, to accomplish the goal of photo-realistic street-view panorama rendering by modeling these street-view specific elements in neural networks. In the experiments, our method is evaluated on both urban and suburban scene datasets, demonstrating that Sat2Density++ is capable of rendering photorealistic street-view panoramas that are consistent across multiple views and faithful to the satellite image. Ming Qian, Bin Tan 0002, Qiuyu Wang, Xianwei Zheng, Hanjiang Xiong, Gui-Song Xia, Yujun Shen, Nan Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Understanding Data Influence With Differential ApproximationabstractData plays a pivotal role in the groundbreaking advancements in artificial intelligence. The quantitative analysis of data significantly contributes to model training, enhancing both the efficiency and quality of data utilization. However, existing data analysis tools often lag in accuracy. For instance, many of these tools even assume that the loss function of neural networks is convex. These limitations make it challenging to implement current methods effectively. In this paper, we introduce a new formulation to approximate a sample's influence by accumulating the differences in influence between consecutive learning steps, which we term Diff-In. Specifically, we formulate the sample-wise influence as the cumulative sum of its changes/differences across successive training iterations. By employing second-order approximations, we approximate these difference terms with high accuracy while eliminating the need for model convexity required by existing methods. Despite being a second-order method, Diff-In maintains computational complexity comparable to that of first-order methods and remains scalable. This efficiency is achieved by computing the product of the Hessian and gradient, which can be efficiently approximated using finite differences of first-order gradients. We assess the approximation accuracy of Diff-In both theoretically and empirically. Our theoretical analysis demonstrates that Diff-In achieves significantly lower approximation error compared to existing influence estimators. Extensive experiments further confirm its superior performance across multiple benchmark datasets in three data-centric tasks: data cleaning, data deletion, and coreset selection. Notably, our experiments on data pruning for large-scale vision-language pre-training show that Diff-In can scale to millions of data points and outperforms strong baselines. Haoru Tan, Sitong Wu, Xiuzhe Wu, Wang Wang, Zeke Xie, Gui-Song Xia, Xiaojuan Qi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Oriented Tiny Object Detection: A Dataset, Benchmark, and Dynamic Unbiased LearningabstractDetecting oriented tiny objects, which are limited in appearance information yet prevalent in real-world applications, remains an intricate and under-explored problem. To address this, we systematically introduce a new dataset, a benchmark, and a dynamic coarse-to-fine learning scheme in this study. Our proposed dataset, AI-TOD-R, features the smallest object sizes among all oriented object detection datasets. Based on AI-TOD-R, we present a benchmark spanning a broad range of detection paradigms, including both fully-supervised and label-efficient approaches. Through investigation, we identify a learning bias presents across various learning pipelines: confident objects become increasingly confident, while vulnerable oriented tiny objects are further marginalized, hindering their detection performance. To mitigate this issue, we propose a Dynamic Coarse-to-Fine Learning (DCFL) scheme towards unbiased learning. DCFL dynamically updates prior positions to better align with the limited areas of oriented tiny objects, and it assigns samples in a way that balances both quantity and quality across different object shapes, thus mitigating biases in prior settings and sample selection. Extensive experiments across 10 challenging object detection datasets demonstrate that DCFL achieves state-of-the-art accuracy, high efficiency, and remarkable versatility. Chang Xu 0027, Ruixiang Zhang, Wen Yang 0001, Jian Ding 0001, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | DN-TOD: Robust tiny object detection amidst label noise
Chang Xu 0027, Wen Yang 0001, Ruixiang Zhang, Yan Zhang 0115, Gui-Song Xia |
Pattern Recognit. | 6 |
| 2026 | Revisiting Fine-Grained Image Analysis by Semantic-Part AlignmentabstractFine-grained image analysis is widely recognized as highly challenging, since distinguishing individual differences within a certain category, species, or type often depends on tiny, subtle patterns. However, learning fine-grained semantic categories from these subtle part patterns is inherently fragile, as they can easily be overwhelmed by the dominant patterns resting in the coarse-category information. Therefore, how to enhance the relation between the fine-grained semantics and these subtle patterns is the key. To push this frontier, a novel semantic-part alignment (SPA) learning scheme is proposed in this paper. Its general idea is to firstly measure the relevance of each part to the fine-grained semantics, and then regularize the fine-grained visual representation learning. Specifically, it consists of three key components, namely, joint semantic-part modeling, semantic-part set modeling, and optimal semantic-part transport. The joint semantic-part modeling associates each part in an image with the fine-grained semantics in a latent space. Then, the optimal semantic-part transport component is devised to enhance the relation between fine-grained semantic embeddings and the discriminative part embeddings. Notably, the proposed SPA is plug-in-and-play, easy-to-implement, and insensitive to the latent embedding dimension and loss weight. Experiments show the proposed method can substantially boost performance on multiple fine-grained image analysis tasks across various baselines. Qi Bi, Jingjun Yi, Haolan Zhan, Wei Ji 0011, Gui-Song Xia |
IEEE Trans. Image Process. | 5 |
| 2026 | QuadricsReg: Large-Scale Point Cloud Registration Using Semantic Quadric Primitives
Ji Wu 0012, Huai Yu, Ximeng Cai, Mingfeng Wang, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Robotics | 7 |
| 2025 | Learning Fine-grained Domain Generalization via Hyperbolic State Space HallucinationabstractFine-grained domain generalization (FGDG) aims to learn a fine-grained representation that can be well generalized to unseen target domains when only trained on the source domain data. Compared with generic domain generalization, FGDG is particularly challenging in that the fine-grained category can be only discerned by some subtle and tiny patterns. Such patterns are particularly fragile under the cross-domain style shifts caused by illumination, color and etc. To push this frontier, this paper presents a novel Hyperbolic State Space Hallucination (HSSH) method. It consists of two key components, namely, state space hallucination (SSH) and hyperbolic manifold consistency (HMC). SSH enriches the style diversity for the state embeddings by firstly extrapolating and then hallucinating the source images. Then, the pre- and post- style hallucinate state embeddings are projected into the hyperbolic manifold. The hyperbolic state space models the high-order statistics, and allows a better discernment of the fine-grained patterns. Finally, the hyperbolic distance is minimized, so that the impact of style variation on fine-grained patterns can be eliminated. Experiments on three FGDG benchmarks demonstrate its state-of-the-art performance. Qi Bi, Jingjun Yi, Haolan Zhan, Wei Ji 0011, Gui-Song Xia |
AAAI | 5 |
| 2025 | VHM: Versatile and Honest Vision Language Model for Remote Sensing Image AnalysisabstractThis paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset comprising both factual and deceptive questions (HnstD). Unlike prevailing remote sensing image-text datasets, in which image captions focus on a few prominent objects and their relationships, VersaD captions provide detailed information about image properties, object attributes, and the overall scene. This comprehensive captioning enables VHM to thoroughly understand remote sensing images and perform diverse remote sensing tasks. Moreover, different from existing remote sensing instruction datasets that only include factual questions, HnstD contains additional deceptive questions stemming from the non-existence of objects. This feature prevents VHM from producing affirmative answers to nonsense queries, thereby ensuring its honesty. In our experiments, VHM significantly outperforms various vision language models on common tasks of scene classification, visual question answering, and visual grounding. Additionally, VHM achieves competent performance on several unexplored tasks, such as building vectorizing, multi-label classification and honest question answering. Chao Pang 0001, Xingxing Weng, Jiang Wu 0003, Yi Liu 0028, Jiaxing Sun 0001, Litong Feng, Gui-Song Xia, Conghui He |
AAAI | 10 |
| 2025 | Bypass Back-propagation: Optimization-based Structural Pruning for Large Language Models via Policy GradientabstractRecent Large-Language Models (LLMs) pruning methods typically operate at the posttraining phase without the expensive weight finetuning, however, their pruning criteria often rely on heuristically hand-crafted metrics, potentially leading to suboptimal performance.We instead propose a novel optimizationbased structural pruning that learns the pruning masks in a probabilistic space directly by optimizing the loss of the pruned model.To preserve efficiency, our method eliminates the back-propagation through the LLM per se during optimization, requiring only the forward pass of the LLM.We achieve this by learning an underlying Bernoulli distribution to sample binary pruning masks, where we decouple the Bernoulli parameters from LLM loss, facilitating efficient optimization via policy gradient estimator without back-propagation.Thus, our method can 1) support global and heterogeneous pruning (i.e., automatically determine different redundancy for different layers), and 2) optionally initialize with a metric-based method (for our Bernoulli distributions).Extensive experiments conducted on LLaMA, LLaMA-2, LLaMA-3, Vicuna, and Mistral models using the C4 and WikiText2 datasets demonstrate the promising performance of our method in efficiency and effectiveness.Code is available at https://github.com/ ethanygao/backprop-free_LLM_pruning. Yuan Gao 0015, Zujing Liu, Bo Du 0001, Gui-Song Xia |
ACL (1) | 5 |
| 2025 | UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image GenerationabstractRecently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descriptions. However, these models often struggle with achieving precise control over pixel-level layouts, object appearances, and global styles when using text prompts alone. To mitigate this issue, previous works introduce conditional images as auxiliary inputs for image generation, enhancing control but typically necessitating specialized models tailored to different types of reference inputs. In this paper, we explore a new approach to unify controllable generation within a single framework. Specifically, we propose the unified image-instruction adapter (UNIC-Adapter) built on the Multi-Modal-Diffusion Transformer architecture, to enable flexible and controllable generation across diverse conditions without the need for multiple specialized models. Our UNIC-Adapter effectively extracts multi-modal instruction information by incorporating both conditional images and task instructions, injecting this information into the image generation process through a cross-attention mechanism enhanced by Rotary Position Embedding. Experimental results across a variety of tasks, including pixel-level spatial control, subject-driven image generation, and style-image-based image synthesis, demonstrate the effectiveness of our UNIC-Adapter in unified controllable image generation. Lunhao Duan, Shanshan Zhao 0001, Yinglun Li, Weihua Luo, Kaifu Zhang, Mingming Gong, Gui-Song Xia |
CVPR | 10 |
| 2025 | Exploring Scene Affinity for Semi-Supervised LiDAR Semantic SegmentationabstractThis paper explores scene affinity (AIScene), namely intra-scene consistency and inter-scene correlation, for semi-supervised LiDAR semantic segmentation in driving scenes. Adopting teacher-student training, AIScene employs a teacher network to generate pseudo-labeled scenes from unlabeled data, which then supervise the student network’s learning. Unlike most methods that include all points in pseudo-labeled scenes for forward propagation but only pseudo-labeled points for backpropagation, AIScene removes points without pseudo-labels, ensuring consistency in both forward and backward propagation within the scene. This simple point erasure strategy effectively prevents unsupervised, semantically ambiguous points (excluded in backpropagation) from affecting the learning of pseudo-labeled points. Moreover, AIScene incorporates patch-based data augmentation, mixing multiple scenes at both scene and instance levels. Compared to existing augmentation techniques that typically perform scene-level mixing between two scenes, our method enhances the semantic diversity of labeled (or pseudo-labeled) scenes, thereby improving the semi-supervised performance of segmentation models. Experiments show that AIScene outperforms previous methods on two popular benchmarks across four settings, achieving notable improvements of 1.9% and 2.1% in the most challenging 1% labeled data. The code will be released at https://github.com/azhuantou/AIScene. Chuandong Liu, Xingxing Weng, Shuguo Jiang, Pengcheng Li 0017, Lei Yu 0006, Gui-Song Xia |
CVPR | 6 |
| 2025 | AdaDCP: Learning an Adapter with Discrete Cosine Prior for Clear-to-Adverse Domain Generalization
Qi Bi, Yixian Shen, Jingjun Yi, Gui-Song Xia |
ICCV | 4 |
| 2025 | Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation
Guopeng Li 0004, Shouhong Ding, Yuan Gao 0015, Gui-Song Xia |
ICCV | 6 |
| 2025 | Beyond Pixel Uncertainty: Bounding the OoD Objects in Road Scenes
Huachao Zhu, Zelong Liu, Zhichao Sun 0004, Yuda Zou, Gui-Song Xia, Yongchao Xu |
ICCV | 5 |
| 2025 | LiDAR-enhanced 3D Gaussian Splatting MappingabstractThis paper introduces LiGSM, a novel LiDARenhanced 3D Gaussian Splatting (3DGS) mapping framework that improves the accuracy and robustness of 3D scene mapping by integrating LiDAR data. LiGSM constructs joint loss from images and LiDAR point clouds to estimate the poses and optimize their extrinsic parameters, enabling dynamic adaptation to variations in sensor alignment. Furthermore, it leverages LiDAR point clouds to initialize 3DGS, providing a denser and more reliable starting points compared to sparse SfM points. In scene rendering, the framework augments standard image-based supervision with depth maps generated from LiDAR projections, ensuring an accurate scene representation in both geometry and photometry. Experiments on public and self-collected datasets demonstrate that LiGSM outperforms comparative methods in pose tracking and scene rendering. Huai Yu, Ji Wu 0012, Wen Yang 0001, Gui-Song Xia |
ICRA | 5 |
| 2025 | Holistic Large-Scale Scene Reconstruction via Mixed Gaussian SplattingabstractRecent advances in 3D Gaussian Splatting have shown remarkable potential for novel view synthesis. However, most existing large-scale scene reconstruction methods rely on the divide-and-conquer paradigm, which often leads to the loss of global scene information and requires complex parameter tuning due to scene partitioning and local optimization. To address these limitations, we propose MixGS, a novel holistic optimization framework for large-scale 3D scene reconstruction. MixGS models the entire scene holistically by integrating camera pose and Gaussian attributes into a view-aware representation, which is decoded into fine-detailed Gaussians. Furthermore, a novel mixing operation combines decoded and original Gaussians to jointly preserve global coherence and local fidelity. Extensive experiments on large-scale scenes demonstrate that MixGS achieves state-of-the-art rendering quality and competitive speed, while significantly reducing computational requirements, enabling large-scale scene reconstruction training on a single 24GB VRAM GPU. Chuandong Liu, Huijiao Wang, Lei Yu 0006, Gui-Song Xia |
NeurIPS | 4 |
| 2025 | Mitigating representation bias for class-incremental semantic segmentation of remote sensing images
Xiaoqian Sun, Xingxing Weng, Chao Pang 0001, Gui-Song Xia |
Sci. China Inf. Sci. | 4 |
| 2025 | Self-supervised Shutter Unrolling with Events
Mingyuan Lin, Yangguang Wang, Xiang Zhang 0022, Boxin Shi, Wen Yang 0001, Chu He, Gui-Song Xia, Lei Yu 0006 |
Int. J. Comput. Vis. | 7 |
| 2025 | High-Quality Pseudo-Labeling for Point Cloud Segmentation With Scene-Level AnnotationabstractThis paper investigates indoor point cloud semantic segmentation under scene-level annotation, which is less explored compared to methods relying on sparse point-level labels. In the absence of precise point-level labels, current methods first generate point-level pseudo-labels, which are then used to train segmentation models. However, generating accurate pseudo-labels for each point solely based on scene-level annotations poses a considerable challenge, substantially affecting segmentation performance. Consequently, to enhance accuracy, this paper proposes a high-quality pseudo-label generation framework by exploring contemporary multi-modal information and region-point semantic consistency. Specifically, with a cross-modal feature guidance module, our method utilizes 2D-3D correspondences to align point cloud features with corresponding 2D image pixels, thereby assisting point cloud feature learning. To further alleviate the challenge presented by the scene-level annotation, we introduce a region-point semantic consistency module. It produces regional semantics through a region-voting strategy derived from point-level semantics, which are subsequently employed to guide the point-level semantic predictions. Leveraging the aforementioned modules, our method can rectify inaccurate point-level semantic predictions during training and obtain high-quality pseudo-labels. Significant improvements over previous works on ScanNet v2 and S3DIS datasets under scene-level annotation can demonstrate the effectiveness. Additionally, comprehensive ablation studies validate the contributions of our approach's individual components. Lunhao Duan, Shanshan Zhao 0001, Xingxing Weng, Jing Zhang 0037, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Towards Human-Level 3D Relative Pose Estimation: Generalizable, Training-Free, With Single ReferenceabstractHumans can easily deduce the relative pose of a previously unseen object, without labeling or training, given only a single query-reference image pair. This is arguably achieved by incorporating i) 3D/2.5D shape perception from a single image, ii) render-and-compare simulation, and iii) rich semantic cue awareness to furnish (coarse) reference-query correspondence. Motivated by this, we propose a novel 3D generalizable relative pose estimation method by elaborating 3D/2.5D shape perception with a 2.5D shape from an RGB-D reference, fulfilling the render-and-compare paradigm with an off-the-shelf differentiable renderer, and leveraging the semantic cues from a pretrained model like DINOv2. Specifically, our differentiable renderer takes the 2.5D rotatable mesh textured by the RGB and the semantic maps (obtained by DINOv2 from the RGB input), then renders new RGB and semantic maps (with back-surface culling) under a novel rotated view. The refinement loss comes from comparing the rendered RGB and semantic maps with the query ones, back-propagating the gradients through the differentiable renderer to refine the 3D relative pose. As a result, our method can be readily applied to unseen objects, given only a single RGB-D reference, without labeling or training. Extensive experiments on LineMOD, LM-O, and YCB-V show that our training-free method significantly outperforms the state-of-the-art supervised methods, especially under the rigorous Acc@5/10/15$^\circ$∘ metrics and the challenging cross-dataset settings. Yuan Gao 0015, Yajing Luo, Kui Jia, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Partial Distribution Matching via Partial Wasserstein Adversarial NetworksabstractThis paper studies the problem of distribution matching (DM), which is a fundamental machine learning problem seeking to robustly align two probability distributions. Our approach is established on a relaxed formulation, called partial distribution matching (PDM), which seeks to match a fraction of the distributions instead of matching them completely. We theoretically derive the Kantorovich-Rubinstein duality for the partial Wasserstein-1 (PW) discrepancy, and develop a partial Wasserstein adversarial network (PWAN) that efficiently approximates the PW discrepancy based on this dual form. Partial matching can then be achieved by optimizing the network using gradient descent. Two practical tasks, point set registration and partial domain adaptation are investigated, where the goals are to partially match distributions in 3D space and high-dimensional feature space respectively. The experiment results confirm that the proposed PWAN effectively produces highly robust matching results, performing better or on par with the state-of-the-art methods. Nan Xue 0001, Rebecka Jörnsten, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | On the Robustness of Object Detection Models on Aerial ImagesabstractThe robustness of object detection models is a major concern when applied to real-world scenarios. The performance of most models tends to degrade when confronted with images affected by corruptions, since they are usually trained and evaluated on clean datasets. While numerous studies have explored the robustness of object detection models on natural images, there is a paucity of research focused on models applied to aerial images, which feature complex backgrounds, substantial variations in scales, and orientations of objects. This article addresses the challenge of assessing the robustness of object detection models on aerial images, with a specific emphasis on scenarios where images are affected by clouds. In this study, we introduce two novel benchmarks based on DOTA-v1.0. The first benchmark encompasses 19 prevalent corruptions, while the second focuses on the cloud-corrupted condition—a phenomenon uncommon in natural images yet frequent in aerial photography. We systematically evaluate the robustness of mainstream object detection models and perform necessary ablation experiments. Through our investigations, we find that rotation-invariant modeling and enhanced backbone architectures can improve the robustness of models. Furthermore, increasing the capacity of Transformer-based backbones can strengthen their robustness. The benchmarks we propose and our comprehensive experimental analyses can facilitate research on robust object detection on aerial images. The codes and datasets are available at:https://github.com/hehaodong530/DOTA-C. Haodong He, Jian Ding 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | STAR-CD: Style-Aligned Remote Sensing Change Detection With Appearance-Relation ModelingabstractRemote sensing change detection plays a crucial part in monitoring the evolution of land cover. Despite remarkable progress in recent years, existing change detection models still struggle with two major challenges that lead to inaccurate change predictions. For one, current models lack sufficient global relation modeling, resulting in coarse boundary prediction and false identification of unwanted variations. For another, bitemporal images often display divergent imaging styles due to lighting, sensor, or seasonal differences, and this cross-temporal style inconsistency would amplify the pseudo-changes caused by non-semantic variations. In this paper, we present the style-alignment and appearance-relation modeling change detection model (STAR-CD) to tackle the above challenges. First, we propose an appearance-relation modeling block (ARMB) to jointly explore semantic differences from both local textures and global structures of remote sensing images, effectively reducing false alarms and refining change boundary predictions. Second, we design a Fourier-inspired style alignment module (FSAM), which suppresses style-induced noise and reveals true changes by texture-style decoupling in the frequency domain. Extensive experiments on the WHU-CD, LEVIR-CD, and CDD datasets demonstrate the superiority of our STAR-CD over state-of-the-art methods. In particular, STAR-CD surpasses previous methods by 2.40% and 4.34% in F1 and IoU on the CDD datasets, respectively. Visual comparisons further validate its strength in predicting precise change boundaries and robustly suppressing pseudo-changes. Yan Zhang 0115, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Unsupervised Multiview UAV Image Geolocalization via Iterative RenderingabstractUnmanned Aerial Vehicle (UAV) Cross-View Geo-Localization (CVGL) poses significant challenges due to the substantial view discrepancies between oblique UAV images and overhead satellite images. Existing methods heavily rely on supervised learning with labeled datasets to extract viewpoint-invariant features for cross-view retrieval. However, these approaches are computationally expensive, prone to overfitting region-specific cues, and exhibit limited generalizability to new regions. To overcome this issue, we propose an unsupervised solution that lifts the scene representation to 3D space from UAV observations for satellite image generation, providing a robust representation against view distortion. By generating orthogonal images that closely resemble satellite views, our method reduces view discrepancies in feature representation and mitigates shortcuts in region-specific image pairing. To further align the perspective of the rendered image with the real one, we design an iterative camera pose updating mechanism that progressively modulates the rendered query image with potential satellite targets, eliminating spatial offsets relative to the reference images. Additionally, this iterative refinement strategy enhances cross-view feature invariance through view-consistent fusion across iterations. As such, our unsupervised paradigm naturally avoids the problem of region-specific overfitting, enabling generic CVGL for UAV images without feature fine-tuning or data-driven training. Experiments on the University-1652 and SUES-200 datasets demonstrate that our approach significantly improves geo-localization accuracy while maintaining robustness across diverse regions. Notably, without model fine-tuning or paired training, our method achieves competitive performance with recent supervised methods. Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Li Mi, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Learning Modality-Invariant Feature for Multimodal Image Matching via Knowledge DistillationabstractMultimodal remote sensing image matching is essential for multi-source information fusion. Recently, learning-based feature matching networks have significantly enhanced the performance of unimodal image matching tasks through data-driven approaches. However, progress in applying these learning-based methods to multimodal image matching has been slower. A major obstacle is the substantial nonlinear radiometric differences between modalities, which require networks to learn modality-invariant features from large amounts of paired data. To address this, we propose EMINet, an efficient method for learning modality-invariant features from limited data to improve matching performance. Our approach constructs a high-performance teacher network by combining the DINOv2 foundational model, the keypoint and descriptor extraction network SuperPoint, and the feature matching network SuperGlue. Leveraging the strong semantic representation capability of DINOv2, the teacher network achieves excellent cross-modality matching ability. To meet low-latency requirements in practical applications, we introduce two novel knowledge distillation strategies: Semantic Window Relation Distillation (SWRD) and Cross-Triplet Descriptor Distillation (CTDD). SWRD improves the discriminative power of the student network’s descriptors by learning patch-level distributions from DINOv2, while CTDD enforces cross-modality triplet constraints to enhance modality invariance of the student network. Experimental results demonstrate that EMINet outperforms several state-of-the-art methods on various datasets, including Optical-SAR, Optical-NIR, and Optical-IR datasets. Yepeng Liu 0002, Wenpeng Lai, Yuliang Gu, Gui-Song Xia, Bo Du 0001, Yongchao Xu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Holistic Response Lifting for Weakly Supervised Land-Cover Classification
Qiyuan Ma, Xianwei Zheng, Linxi Huan, Linwei Yue, Gui-Song Xia, Jianya Gong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | RTO-LLI: Robust Real-Time Image Orientation Method With Rapid Multilevel Matching and Third-Times Optimizations for Low-Overlap Large-Format UAV ImagesabstractUAV real-time photogrammetry is important to promote the rapid generation of photogrammetry 4D product, intelligent information extraction and rapid remote sensing mapping, and efficient large-scale 3D modeling. However, for real-time processing of low-overlap large-format image sequence, there remains two challenges: (1) Large-format images result in greater data volume and computational load, posing challenges for real-time online processing on regular-performance computing units, requiring more efficient algorithms; (2) Low-overlap images make it difficult for matching correspondences to cover the entire overlapping area at real-time, leading to significant challenges for real-time and robust relative orientation. Therefore, this paper proposes a robust Real-Time Orientation method for Low-overlap Large-format UAV Images (RTO-LLI), which can robustly handle these kind of data in real-time. Firstly, robust initialization method for real-time processing of low-overlap large-format images was designed to ensure a high-success-rate of SLAM initialization. Secondly, constant velocity hypothesis tracking enables fast orientation during constant-speed flight. Thirdly, when the second step false, using real-time pose estimation method based on multilevel matching and coarse-to-fine optimization to robustly solve the precision image pose. Fourthly, final (third-level) pose optimization method based on the IRLS algorithm with suitable search area, which can compute higher-precision image pose in real-time. Finally, real-time mapping based on parallel processing for low-overlap images can generate high-precision 3D point maps and complete feature extraction for the next frame in real-time. Experiments conducted on several different types of scenes show that: (1) the processing speed of RTO-LLI significantly surpasses traditional offline methods: PhotoScan, OpenMVG, Colmap. RTO-LLI can handle large-format UAV image sequence (single-imagery has 20-million-pixels) at a speed of 1.5 frames-per-second, meeting the demands of real-time UAV photogrammetry tasks; (2) RTO-LLI is the only method that has successfully completed real-time tasks in all 50-times repeated experiments for four different types of scenes, demonstrating robustness far superior to other classical SLAM solutions; (3) the-displacement-error of the estimated Pose by RTO-LLI is less than 1/2000 of the-trajectory-length, and the average-reprojection-error is less than 1.5 pixels, almost as well as traditional offline methods. RTO-LLI method meets the efficiency, robustness and accuracy requirements of real-time photogrammetry for low-overlap large-format UAV images. Xiongwu Xiao, Gui-Song Xia, Jianya Gong, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | ConvFormer-CD: Hybrid CNN-Transformer With Temporal Attention for Detecting Changes in Remote Sensing ImageryabstractRecently, the combination of Transformers and convolutional neural networks (CNNs) has witnessed significant advancements in change detection (CD) tasks. However, it remains unexplored how to interactively integrate long-range dependency and local information to enhance the model’s global-local context awareness for effectively mitigating pseudo-changes. In addition, accurate identification and distinction of building changes from complex backgrounds still pose challenges due to the insufficient semantic context modeling across time between bi-temporal images. To address these issues, we propose a hybrid model ConvFormer-CD with parallel convolution and multihead self-attention (MSA). This combination enables better interaction of global and local information, thereby enhancing the adaptability to complex scenarios. Moreover, we introduce a novel module called Temporal Attention to establish cross-temporal semantic relationships between image pairs, effectively highlighting change regions by learning shared and nonshared semantics. This enables our model to accurately detect changed targets even in scenarios characterized by intricate geo-spatial arrangements and distributions. To further refine the differences in bi-temporal images, we propose a difference integration module (DIM) that connects the encoder and the decoder to fuse high-level semantic features across channels. We conduct extensive experiments on four benchmark datasets, including LEVIR-CD, LEVIR-CD+, WHU-CD, and S2Looking-CD, which demonstrates that the proposed ConvFormer-CD outperforms other state-of-the-art (SOTA) methods. Our codes will be available athttps://github.com/taomi-lab/ConvFormer-CD. Feng Yang 0015, Mengtao Li, Wenqiang Shu, Anyong Qin, Tiecheng Song, Chenqiang Gao, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Minimizing Sample Redundancy for Label-Efficient Object Detection in Aerial ImagesabstractObjects in aerial images tend to be densely scattered and appear in arbitrary orientations, making the annotation process quite costly. To reduce the annotation cost, existing methods propose randomly annotating a proportion of images or objects for aerial object detection with fewer label usage. These approaches, however, can lead to redundancy in labels and inherit the biases associated with the imbalance in datasets. To minimize sample redundancy and alleviate data imbalance, we propose a novel labeling pattern that acquires heterogeneous object labels in a class-orthogonal manner, preserving a broader diversity of samples for each category with less annotation effort. To improve data utility, we design a Dynamic Multi-View Learning (DML) strategy to overcome the sample quantity-quality dilemma in current pseudo-labeling methods—a high pseudo-label threshold reduces sample quantity, while low thresholds compromise sample quality. First, DML separates model predictions into multiple hierarchies for finer screening, mitigating the suppression of unlabeled objects in binary pseudo-label strategies. With this separation, DML learns to construct a new view by injecting high-quality samples and masking low-quality regions in this view, simultaneously expanding sample quantity while ensuring sample quality. Unlike previous methods that mine pseudo labels solely from unlabelled regions, DML releases this constraint by learning to expand high-quality samples with a dynamic view. Extensive experiments on five benchmark datasets validate our method’s state-of-the-art accuracy and label efficiency. Notably, with approximately 5% DOTA-v2.0 annotations, DML achieves nearly 90% of the fully supervised performance. The codes will be available at https://github.com/ZhangRuixiang-WHU/ALOD_DML/. Ruixiang Zhang, Chang Xu 0027, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Few-Shot Oriented Object Detection in Remote Sensing Images via Memorable Contrastive LearningabstractFew-shot object detection (FSOD) has attracted significant research attention in remote sensing due to its potential to reduce reliance on large annotated datasets. However, two challenges remain in this area: (1) axis-aligned proposals, which can result in misalignment for arbitrarily oriented objects, and (2) object misclassification due to limited annotated data, which hinders generalization to unseen classes. To address these issues, we propose a novel method for few-shot oriented object detection in remote sensing images. Our approach employs oriented bounding boxes instead of horizontal ones to learn more effective feature representations for arbitrarily oriented aerial objects, enhancing detection accuracy. Additionally, we introduce a supervised contrastive learning module with a dynamically updated memory bank, enabling the model to leverage large batches of negative samples and to better learn discriminative features for unseen classes. Extensive experiments on DOTA, HRSC2016, and DIOR-R datasets demonstrate superior performance of our proposed method in few-shot oriented object detection. Code and pre-trained models will be made publicly available. Jiawei Zhou 0009, Wuzhou Li, Hongtao Cai, Tianjin Huang, Gui-Song Xia, Xiang Li 0046 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Cross-Level Multi-Instance Distillation for Self-Supervised Fine-Grained Visual CategorizationabstractHigh-quality annotation of fine-grained visual categories demands great expert knowledge, which is taxing and time consuming. Alternatively, learning fine-grained visual representation from enormous unlabeled images (e.g., species, brands) by self-supervised learning becomes a feasible solution. However, recent investigations find that existing self-supervised learning methods are less qualified to represent fine-grained categories. The bottleneck lies in that the pre-trained class-agnostic representation is built from every patch-wise embedding, while fine-grained categories are only determined by several key patches of an image. In this paper, we propose a Cross-level Multi-instance Distillation (CMD) framework to tackle this challenge. Our key idea is to consider the importance of each image patch in determining the fine-grained representation by multiple instance learning. To comprehensively learn the relation between informative patches and fine-grained semantics, the multi-instance knowledge distillation is implemented on both the region/image crop pairs from the teacher and student net, and the region-image crops inside the teacher / student net, which we term as intra-level multi-instance distillation and inter-level multi-instance distillation. Extensive experiments on several commonly used datasets, including CUB-200-2011, Stanford Cars and FGVC Aircraft, demonstrate that the proposed method outperforms the contemporary methods by up to 10.14% and existing state-of-the-art self-supervised learning approaches by up to 19.78% on both top-1 accuracy and Rank-1 retrieval metric. Source code is available at https://github.com/BiQiWHU/CMD. Qi Bi, Wei Ji 0011, Jingjun Yi, Haolan Zhan, Gui-Song Xia |
IEEE Trans. Image Process. | 5 |
| 2025 | Universal Fine-Grained Visual Categorization by Concept Guided LearningabstractExisting fine-grained visual categorization (FGVC) methods assume that the fine-grained semantics rest in the informative parts of an image. This assumption works well on favorable front-view object-centric images, but can face great challenges in many real-world scenarios, such as scene-centric images (e.g., street view) and adverse viewpoint (e.g., object reidentification, remote sensing). In such scenarios, the mis-/over-feature activation is likely to confuse the part selection and degrade the fine-grained representation. In this paper, we are motivated to design a universal FGVC framework for real-world scenarios. More precisely, we propose a concept guided learning (CGL), which models concepts of a certain fine-grained category as a combination of inherited concepts from its subordinate coarse-grained category and discriminative concepts from its own. The discriminative concepts is utilized to guide the fine-grained representation learning. Specifically, three key steps are designed, namely, concept mining, concept fusion, and concept constraint. On the other hand, to bridge the FGVC dataset gap under scene-centric and adverse viewpoint scenarios, a Fine-grained Land-cover Categorization Dataset (FGLCD) with 59,994 fine-grained samples is proposed. Extensive experiments show the proposed CGL: 1) has a competitive performance on conventional FGVC; 2) achieves state-of-the-art performance on fine-grained aerial scenes & scene-centric street scenes; 3) good generalization on object re-identification and fine-grained aerial object detection. The dataset and source code will be available at https://github.com/BiQiWHU/CGL. Qi Bi, Beichen Zhou, Wei Ji 0011, Gui-Song Xia |
IEEE Trans. Image Process. | 4 |
| 2025 | Few-Shot Learning for Annotation-Efficient Nucleus Instance SegmentationabstractNucleus instance segmentation from histopathology images suffers from the extremely laborious and expert-dependent annotation of nucleus instances. As a promising solution to this task, annotation-efficient deep learning paradigms have recently attracted much research interest, such as weakly-/semi-supervised learning, generative adversarial learning, etc. In this paper, we propose to formulate annotation-efficient nucleus instance segmentation from the perspective of few-shot learning (FSL). Our work was motivated by that, with the prosperity of computational pathology, an increasing number of fully-annotated datasets are publicly accessible, and we hope to leverage these external datasets to assist nucleus instance segmentation on the target dataset which only has very limited annotation. To achieve this goal, we adopt the meta-learning based FSL paradigm, which however has to be tailored in two substantial aspects before adapting to our task. First, since the novel classes may be inconsistent with those of the external dataset, we extend the basic definition of few-shot instance segmentation (FSIS) to generalized few-shot instance segmentation (GFSIS). Second, to cope with the intrinsic challenges of nucleus segmentation, including touching between adjacent cells, cellular heterogeneity, etc., we further introduce a structural guidance mechanism into the GFSIS network, finally leading to a unified Structurally-Guided Generalized Few-Shot Instance Segmentation (SGFSIS) framework. Extensive experiments on a couple of publicly accessible datasets demonstrate that, SGFSIS can outperform other annotation-efficient learning baselines, including semi-supervised learning, simple transfer learning, etc., with comparable performance to fully supervised learning with around 10% annotations. Zihao Wu 0004, Jie Yang 0002, Danyi Li, Yuan Gao 0015, Changxin Gao, Gui-Song Xia, Yuanqing Li 0001, Jin-Gang Yu |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Patched Line Segment Learning for Vector Road MappingabstractThis paper presents a novel approach to computing vector road maps from satellite remotely sensed images, building upon a well-defined Patched Line Segment (PaLiS) representation for road graphs that holds geometric significance. Unlike prevailing methods that derive road vector representations from satellite images using binary masks or keypoints, our method employs line segments. These segments not only convey road locations but also capture their orientations, making them a robust choice for representation. More precisely, given an input image, we divide it into non-overlapping patches and predict a suitable line segment within each patch. This strategy enables us to capture spatial and structural cues from these patch-based line segments, simplifying the process of constructing the road network graph without the necessity of additional neural networks for connectivity. In our experiments, we demonstrate how an effective representation of a road graph significantly enhances the performance of vector road mapping on established benchmarks, without requiring extensive modifications to the neural network architecture. Furthermore, our method achieves state-of-the-art performance with just 6 GPU hours of training, leading to a substantial 32-fold reduction in training costs in terms of GPU hours. Jiakun Xu, Gui-Song Xia, Nan Xue 0001 |
AAAI | 3 |
| 2024 | NEAT: Distilling 3D Wireframes from Neural Attraction FieldsabstractThis paper studies the problem of structured 3D reconstruction using wireframes that consist of line segments and junctions, focusing on the computation of structured boundary geometries of scenes. Instead of leveraging matching-based solutions from 2D wireframes (or line segments) for 3D wireframe reconstruction as done in prior arts, we present NEAT, a rendering-distilling formulation using neural fields to represent 3D line segments with 2D observations, and bipartite matching for perceiving and distilling of a sparse set of 3D global junctions. The proposed NEAT enjoys the joint optimization of the neural fields and the global junctions from scratch, using view-dependent 2D observations without precomputed cross-view feature matching. Comprehensive experiments on the DTU and BlendedMVS datasets demonstrate our NEAT's superiority over state-of-the-art alternatives for 3D wireframe reconstruction. Moreover, the distilled 3D global junctions by NEAT, are a better initialization than SfM points, for the recently-emerged 3D Gaussian Splatting for high-fidelity novel view synthesis using about 20 times fewer initial 3D points. Project page: https://xuenan.net/neat. Nan Xue 0001, Bin Tan 0002, Yuxi Xiao, Gui-Song Xia, Tianfu Wu 0001, Yujun Shen |
CVPR | 5 |
| 2024 | Anchor-based Robust Finetuning of Vision-Language ModelsabstractWe aim at finetuning a vision-language model without hurting its out-of-distribution (OOD) generalization. We address two types of OOD generalization, i.e., i) domain shift such as natural to sketch images, and ii) zero-shot capability to recognize the category that was not contained in the finetune data. Arguably, the diminished OOD generalization after finetuning stems from the excessively simplified finetuning target, which only provides the class information, such as “a photo of a [CLASS]”. This is distinct from the process in that CLIP was pretrained, where there is abundant text supervision with rich semantic information. Therefore, we propose to compensate for the finetune process using auxiliary supervision with rich semantic information, which acts as anchors to preserve the OOD generalization. Specifically, two types of anchors are elaborated in our method, including i) text-compensated anchor which uses the images from the finetune set but enriches the text supervision from a pretrained captioner, ii) image-text-pair anchor which is retrieved from the dataset similar to pretraining data of CLIP according to the downstream task, associating with the original CLIP text with rich semantics. Those anchors are utilized as auxiliary semantic information to maintain the original feature space of CLIP, thereby preserving the OOD generalization capabilities. Comprehensive experiments demonstrate that our method achieves in-distribution performance akin to conventional finetuning while attaining new state-of-the-art results on domain shift and zero-shot learning benchmarks. Jinwei Han, Zhiwen Lin, Zhongyisun Sun, Yingguo Gao, Shouhong Ding, Yuan Gao 0015, Gui-Song Xia |
CVPR | 8 |
| 2024 | Unleashing Unlabeled Data: A Paradigm for Cross-View Geo-LocalizationabstractThis paper investigates the effective utilization of unlabeled data for large-area cross-view gee-localization (CVGL), encompassing both unsupervised and semi-supervised settings. Common approaches to CVGL rely on ground-satellite image pairs and employ label-driven supervised training. However, the cost of collecting precise cross-view image pairs hinders the deployment of CVGL in real-life scenarios. Without the pairs, CVGL will be more challenging to handle the significant imaging and spatial gaps between ground and satellite images. To this end, we propose an unsupervised framework including a cross-view projection to guide the model for retrieving initial pseudo-labels and a fast re-ranking mechanism to refine the pseudo-labels by leveraging the fact that “the perfectly paired ground-satellite image is located in a unique and identical scene”. The framework exhibits competitive performance compared with supervised works on three open-source benchmarks. Our code and models will be released on https://github.com/liguopeng0923/UCVGL. Guopeng Li 0004, Ming Qian, Gui-Song Xia |
CVPR | 3 |
| 2024 | 3D Building Reconstruction from Monocular Remote Sensing Images with Multi-level Supervisionsabstract3D building reconstruction from monocular remote sensing images is an important and challenging research problem that has received increasing attention in recent years, owing to its low cost of data acquisition and availability for large-scale applications. However, existing methods rely on expensive 3D-annotated samples for fully-supervised training, restricting their application to large-scale cross-city scenarios. In this work, we propose MLS-BRN, a multi-level supervised building reconstruction network that can flexibly utilize training samples with different annotation levels to achieve better reconstruction results in an end-to-end manner. To alleviate the demand on full 3D supervision, we design two new modules, Pseudo Building Bbox Calculator and Roof-Offset guided Footprint Extractor, as well as new tasks and training strategies for different types of samples. Experimental results on several public and new datasets demonstrate that our proposed MLS-BRN achieves competitive performance using much fewer 3D-annotated samples, and significantly improves the footprint extraction and 3D reconstruction performance compared with current state-of-the-art. The code and datasets of this work will be released at https://github.com/opendatalab/MLS-BRN.git. Haote Yang, Zhenghao Hu, Juepeng Zheng, Gui-Song Xia, Conghui He |
CVPR | 5 |
| 2024 | FreePoint: Unsupervised Point Cloud Instance SegmentationabstractInstance segmentation of point clouds is a crucial task in 3D field with numerous applications that involve localizing and segmenting objects in a scene. However, achieving sat-isfactory results requires a large number of manual annotations, which is time-consuming and expensive. To alleviate dependency on annotations, we propose a novelframework, FreePoint, for underexplored unsupervised class-agnostic instance segmentation on point clouds. In detail, we represent the point features by combining coordinates, colors, and self-supervised deep features. Based on the point features, we perform a bottom-up multicut algorithm to seg-ment point clouds into coarse instance masks as pseudo labels, which are used to train a point cloud instance segmen-tation model. We propose an id-as-feature strategy at this stage to alleviate the randomness of the multicut algorithm and improve the pseudo labels' quality. During training, we propose a weakly-supervised two-step training strategy and corresponding losses to overcome the inaccuracy of coarse masks. FreePoint has achieved breakthroughs in un-supervised class-agnostic instance segmentation on point clouds and outperformed previous traditional methods by over 18.2% and a competitive concurrent work UnScene3D by 5.5% in AP. Additionally, when used as a pretext task and fine-tuned on S3DIS, FreePoint performs significantly better than existing self-supervised pre-training methods with limited annotations and surpasses CSC by 6.0% in AP with 10% annotation masks. Code will be released at https://github.com/zzk273/FreePoint. Jian Ding 0001, Li Jiang 0009, Dengxin Dai, Gui-Song Xia |
CVPR | 5 |
| 2024 | Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference CostabstractWe aim at exploiting additional auxiliary labels from an independent (auxiliary) task to boost the primary task performance which we focus on, while preserving a single task inference cost of the primary task. While most existing auxiliary learning methods are optimization-based relying on loss weights/gradients manipulation, our method is architecture-based with a flexible asymmetric structure for the primary and auxiliary tasks, which produces different networks for training and inference. Specifically, starting from two single task networks/branches (each representing a task), we propose a novel method with evolving networks where only primary-to-auxiliary links exist as the cross-task connections after convergence. These connections can be removed during the primary task inference, resulting in a single-task inference cost. We achieve this by formulating a Neural Architecture Search (NAS) problem, where we initialize bi-directional connections in the search space and guide the NAS optimization converging to an architecture with only the single-side primary-to-auxiliary connections. Moreover, our method can be incorporated with optimization-based auxiliary learning approaches. Extensive experiments with six tasks on NYU v2, CityScapes, and Taskonomy datasets using VGG, ResNet, and ViT backbones validate the promising performance. The codes are available at https://github.com/ethanygao/Aux-NAS. Yuan Gao 0015, Wenhan Luo, Lin Ma 0002, Jin-Gang Yu, Gui-Song Xia, Jiayi Ma 0001 |
ICLR | 6 |
| 2024 | DMTG: One-Shot Differentiable Multi-Task GroupingabstractWe aim to address Multi-Task Learning (MTL) with a large number of tasks by Multi-Task Grouping (MTG). Given $N$ tasks, we propose to simultaneously identify the best task groups from $2^N$ candidates and train the model weights simultaneously in one-shot, with the high-order task-affinity fully exploited. This is distinct from the pioneering methods which sequentially identify the groups and train the model weights, where the group identification often relies on heuristics. As a result, our method not only improves the training efficiency, but also mitigates the objective bias introduced by the sequential procedures that potentially leads to a suboptimal solution. Specifically, we formulate MTG as a fully differentiable pruning problem on an adaptive network architecture determined by an unknown Categorical distribution. To categorize $N$ tasks into $K$ groups (represented by $K$ encoder branches), we initially set up $KN$ task heads, where each branch connects to all $N$ task heads to exploit the high-order task-affinity. Then, we gradually prune the $KN$ heads down to $N$ by learning a relaxed differentiable Categorical distribution, ensuring that each task is exclusively and uniquely categorized into only one branch. Extensive experiments on CelebA and Taskonomy datasets with detailed ablations show the promising performance and efficiency of our method. The codes are available at https://github.com/ethanygao/DMTG. Yuan Gao 0015, Shuguo Jiang, Moran Li, Jin-Gang Yu, Gui-Song Xia |
ICML | 5 |
| 2024 | Boosting Fine-Grained Oriented Object Detection via Text Features
Beichen Zhou, Qi Bi, Jian Ding 0001, Gui-Song Xia |
ICPR (16) | 4 |
| 2024 | QuadricsNet: Learning Concise Representation for Geometric Primitives in Point CloudsabstractThis paper presents a novel framework to learn a concise geometric primitive representation for 3D point clouds. Different from representing each type of primitive individually, we focus on the challenging problem of how to achieve a concise and uniform representation robustly. We employ quadrics to represent diverse primitives with only 10 parameters and propose the first end-to-end learning-based framework, namely QuadricsNet, to parse quadrics in point clouds. The relationships between quadrics mathematical formulation and geometric attributes, including the type, scale and pose, are insightfully integrated for effective supervision of QuaidricsNet. Besides, a novel pattern-comprehensive dataset with quadrics segments and objects is collected for training and evaluation. Experiments demonstrate the effectiveness of our concise representation and the robustness of QuadricsNet. Our code is available at https://github.com/MichaelWu99-lab/QuadricsNet. Ji Wu 0012, Huai Yu, Wen Yang 0001, Gui-Song Xia |
ICRA | 4 |
| 2024 | Deeply Unsupervised Patch Re-Identification for Pre-Training Object DetectorsabstractUnsupervised pre-training aims at learning transferable features that are beneficial for downstream tasks. However, most state-of-the-art unsupervised methods concentrate on learning global representations for image-level classification tasks instead of discriminative local region representations, which limits their transferability to region-level downstream tasks, such as object detection. To improve the transferability of pre-trained features to object detection, we present Deeply Unsupervised Patch Re-ID (DUPR), a simple yet effective method for unsupervised visual representation learning. The patch Re-ID task treats individual patch as a pseudo-identity and contrastively learns its correspondence in two views, enabling us to obtain discriminative local features for object detection. Then the proposed patch Re-ID is performed in a deeply unsupervised manner, appealing to object detection, which usually requires multi-level feature maps. Extensive experiments demonstrate that DUPR outperforms state-of-the-art unsupervised pre-trainings and even the ImageNet supervised pre-training on various downstream tasks related to object detection. Jian Ding 0001, Enze Xie, Hang Xu 0004, Chenhan Jiang, Zhenguo Li, Ping Luo 0002, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Detecting Line Segments in Motion-Blurred Images With EventsabstractMaking line segment detectors more reliable under motion blurs is one of the most important challenges for practical applications, such as visual SLAM and 3D line mapping. Existing line segment detection methods face severe performance degradation for accurately detecting and locating line segments when motion blur occurs. While event data shows strong complementary characteristics to images for minimal blur and edge awareness at high-temporal resolution, potentially beneficial for reliable line segment recognition. To robustly detect line segments over motion blurs, we propose to leverage the complementary information of images and events. Specifically, we first design a general frame-event feature fusion network to extract and fuse the detailed image textures and low-latency event edges, which consists of a channel-attention-based shallow fusion module and a self-attention-based dual hourglass module. We then utilize the state-of-the-art wireframe parsing networks to detect line segments on the fused feature map. Moreover, due to the lack of line segment detection datasets with pairwise motion-blurred images and events, we contribute two datasets, i.e., synthetic FE-Wireframe and realistic FE-Blurframe, for network training and evaluation. Extensive analyses on the component configurations demonstrate the design effectiveness of our fusion network. When compared to the state-of-the-arts, the proposed approach achieves the highest detection accuracy while maintaining comparable real-time performance. In addition to being robust to motion blur, our method also exhibits superior performance for line detection under high dynamic range scenes. Huai Yu, Hao Li 0114, Wen Yang 0001, Lei Yu 0006, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | CrossZoom: Simultaneous Motion Deblurring and Event Super-ResolvingabstractEven though the collaboration between traditional and neuromorphic event cameras brings prosperity to frame-event based vision applications, the performance is still confined by the resolution gap crossing two modalities in both spatial and temporal domains. This paper is devoted to bridging the gap by increasing the temporal resolution for images, i.e., motion deblurring, and the spatial resolution for events, i.e., event super-resolving, respectively. To this end, we introduce CrossZoom, a novel unified neural Network (CZ-Net) to jointly recover sharp latent sequences within the exposure period of a blurry input and the corresponding High-Resolution (HR) events. Specifically, we present a multi-scale blur-event fusion architecture that leverages the scale-variant properties and effectively fuses cross-modal information to achieve cross-enhancement. Attention-based adaptive enhancement and cross-interaction prediction modules are devised to alleviate the distortions inherent in Low-Resolution (LR) events and enhance the final results through the prior blur-event complementary information. Furthermore, we propose a new dataset containing HR sharp-blurry images and the corresponding HR-LR event streams to facilitate future research. Extensive qualitative and quantitative experiments on synthetic and real-world datasets demonstrate the effectiveness and robustness of the proposed method. Chi Zhang 0027, Xiang Zhang 0022, Mingyuan Lin, Cheng Li 0023, Chu He, Wen Yang 0001, Gui-Song Xia, Lei Yu 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Weakly Supervised 3-D Building Reconstruction From Monocular Remote Sensing Imagesabstract3D building reconstruction from monocular remote sensing imagery is an important research problem that has been extensively studied for several decades. Although monocular remote sensing imagery is a more economic data source compared with the LiDAR data and multi-view imagery, its limited information results in great challenges and restricts the performance of existing monocular reconstruction methods. Moreover, the expensive cost and the limited quantity of 3D annotations also restrict the application scenes of existing methods, which are mostly based on fully-supervised learning. In our previous work, we have proposed MTBR-Net, a monocular building reconstruction method that consists of a fully-supervised multi-task network and a post-processing module for optimizing the reconstruction results. In this work, we further propose WS-MTBR-Net, a weakly-supervised building reconstruction network that uses fewer 3D annotations and achieves better performance in an end-to-end manner. Specifically, our WS-MTBR-Net fully leverages the relation between different components of a 3D building instance and the property of off-nadir images to improve the footprint segmentation boundary, based on six modified tasks and a new network structure with an improved feature warping module to support weakly-supervised learning. We also design a new training strategy via a hybrid loss function that enables utilizing the training samples with different annotation levels, i.e., complete 3D annotations, 2D footprint annotations, and image-level angle annotations. Results on BONAI Shanghai and Xi’an test datasets demonstrate that our method achieves competitive performance when using 50% fewer 3D-annotated samples, and improves the footprint segmentation F1-score by around 4% compared with current state-of-the-art. Zhenghao Hu, Lingxuan Meng, Jinwang Wang, Juepeng Zheng, Runmin Dong, Conghui He, Gui-Song Xia, Haohuan Fu, Dahua Lin |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Toward Generic and Controllable Attacks Against Object DetectionabstractExisting adversarial attacks against object detectors (ODs) have two inherent limitations. First, ODs have complex meta-structure designs, hence most advanced attacks for ODs concentrate on attacking specific detector-intrinsic structures [e.g., RPN and nonmaximal suppression (NMS)], which makes it hard for them to work on other new detectors. Second, most works against ODs make adversarial examples (AEs) by adding image-level perturbations into original images, which brings redundant perturbations in semantically meaningless areas (e.g., backgrounds). This article proposes a generic white-box attack on mainstream ODs with controllable perturbations. For a generic attack, LGP treats ODs as black boxes and only attacks their outputs, thereby eliminating the limitations of detector-intrinsic structures. Regarding controllability, we establish an object-wise constraint to induce the attachment of perturbations to foregrounds. Experimentally, the proposed LGP successfully attacked 16 state-of-the-art ODs on MS-COCO and DOTA datasets, with promising imperceptibility and transferability obtained. Code is publicly released inhttps://github.com/liguopeng0923/LGP.git. Guopeng Li 0004, Jian Ding 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Learning Cross-View Visual Geo-Localization Without Ground TruthabstractCross-view geo-localization (CVGL) involves determining the geographical location of a query image by matching it with a corresponding GPS-tagged reference image. Current state-of-the-art methods predominantly rely on training models with labeled paired images, incurring substantial annotation costs and training burdens. In this study, we investigate the adaptation of frozen models for CVGL without requiring ground-truth pair labels. We observe that training on unlabeled cross-view images presents significant challenges, including establishing relationships within unlabeled data and reconciling view discrepancies between uncertain queries and references. To address these challenges, we propose a self-supervised learning framework to train a learnable adapter for a frozen foundation model (FM). This adapter is designed to map feature distributions from diverse views into a uniform space using unlabeled data exclusively. To establish relationships within unlabeled data, we introduce an expectation-maximization (EM)-based pseudolabeling module, which iteratively estimates matching between cross-view features and optimizes the adapter. To maintain the robustness of the FM’s representation, we incorporate an information consistency module with a reconstruction loss, ensuring that adapted features retain strong discriminative ability across views. Experimental results demonstrate that our proposed method achieves significant improvements over vanilla FMs and competitive accuracy compared to supervised methods while necessitating fewer training parameters and relying solely on unlabeled data. Evaluation of our adaptation for task-specific models further highlights its broad applicability. Particularly, on the University-1652 dataset, our method outperforms the FM baseline by a substantial margin, achieving about 39 points improvement in Recall@1 and more than 34 points increase in average precision (AP). The project is available athttps://collebt.github.io/EM-CVGL. Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | HiCD: Change Detection in Quality-Varied Images via Hierarchical Correlation DistillationabstractAdvanced change detection techniques primarily target image pairs of equal and high quality. However, variations in imaging conditions and platforms frequently lead to image pairs with distinct qualities: one image being high-quality, while the other being low-quality. These disparities in image quality present significant challenges for understanding image pairs semantically and extracting change features, ultimately resulting in a notable decline in performance. To tackle this challenge, we introduce an innovative training strategy grounded in knowledge distillation. The core idea revolves around leveraging task knowledge acquired from high-quality image pairs to guide the model’s learning process when dealing with image pairs that exhibit differences in quality. Additionally, we develop a hierarchical correlation distillation approach (involving self-correlation, cross-correlation, and global correlation). This approach compels the student model to replicate the correlations inherent in the teacher model, rather than focusing solely on individual features. This ensures effective knowledge transfer while maintaining the student model’s training flexibility. Through extensive experimentation, we demonstrate the remarkable superiority of our methodologies in scenarios involving only resolution disparities, single-degradation, and multi-degradation quality differences. The codes will be released at https://github.com/fitzpchao/HiCD. Chao Pang 0001, Xingxing Weng, Jiang Wu 0003, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | LiST-Net: Enhanced Flood Mapping With Lightweight SAR Transformer Network and Dimension-Wise AttentionabstractDetecting flood-induced changes using synthetic aperture radar (SAR) is crucial for crisis management and damage assessment. Nevertheless, current methodologies predominantly focus on changes in buildings within optical images, struggling with the complex structures of floods. These structures are marked by widespread speckle noise and are accompanied by an increase in computational cost. These challenges hinder their success in real-world applications, necessitating a novel approach. This paper proposes LiST-Net, a lightweight SAR transformer network with dimension-wise attention to improve flood detection accuracy. LiST-Net offers three key advantages. Firstly, the graph neighbor module (GNM) is designed to enhance both detailed information of neighboring pixels and multi-date features within the encoder. Secondly, the dimension-wise interactive attention (DIA) module is proposed to effectively reduce computational complexity while enhancing feature representation. Thirdly, an attentive supervised learning module (ASLM) is incorporated to mitigate noise through a pixel mask gate, allowing change water information to pass through and improving the accuracy of water edge delineation. The effectiveness of LiST-Net is evaluated on two flood detection datasets, S1GFloods and ETCI-2021. Experimental results demonstrate that LiST-Net outperforms existing methods, showcasing a 94.7% improvement in F1 and an 88.7% enhancement in IoU on the S1GFloods datasets, with lower computational costs (11.78G) and fewer parameters (7.34M). This underscores LiST-Net as a promising strategy for precise and effective mapping of floods within SAR images in real-world applications. A public release of the demo code will be available at https://github.com/Tamer-Saleh. Tamer Saleh, Shimaa Holail, Xiongwu Xiao, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Extracting Building Footprints in SAR Images via Distilling Boundary Information From Optical ImagesabstractBuildings represent pivotal entities in remote sensing imagery for various applications like urban planning and land resource management. Predominantly, methods for building footprint extraction in the literature focus on optical imagery with visual attributes that faithfully mirror the physical world. Nevertheless, the acquisition of high-quality optical images presents formidable challenges due to the susceptibility to illumination conditions and scene visibility. In contrast, synthetic aperture radar (SAR) images can be acquired in all-weather and all-time situations, unburdened by the aforementioned constraints. However, the coherent imaging mechanism engenders intricate complexities for building footprint extraction SAR images. To address this issue, this paper introduces the Boundary Information Distillation Network (BIDNet) to improve the prediction accuracy in SAR images by distilling knowledge from optical images. The proposed approach adopts a teacher-student framework, featuring two customized components: the Explicit Distillation Module (EDM) and the Latent Distillation Module (LDM). Different from the conventional practice of directly aligning feature maps, BIDNet focuses on leveraging the more conspicuous boundary information in optical images. The EDM operates by simultaneously yielding a boundary map to emphasize the boundary area and assimilating the explicit low-level features of two modalities. The LDM represents the structural attributes within the high-level latent feature space and aligns the representations of the two modalities. Within this module, intrinsic self-correlations among features originating from boundary regions are encoded, and so are the cross-correlations established between features from boundary regions and alternative areas. The two modules also serve as the conduit for knowledge distillation from the teacher network to the student network, enabling the utilization of optical imagery for enhancing the building footprint extraction in SAR imagery. Extensive experiments demonstrate that our BIDNet achieves state-of-the-art performance on the Multi-Sensor All Weather Mapping (MSAW) dataset, outperforming the strong baseline by 4.3-7.2 points in f1-score and 4.9-8.0 points in IoU. The source code and trained models will be publicly available. Lanxin Zeng, Wen Yang 0001, Jian Kang 0005, Huai Yu, Mihai Datcu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Learning Cross-Modality High-Resolution Representation for Thermal Small-Object DetectionabstractThermal infrared (TIR) object detection plays a crucial role in diverse around-the-clock applications, such as search and rescue operations and wildlife protection. Achieving rapid and robust detection of small objects from an aerial perspective is particularly significant in these scenarios. However, the task is compounded by two interrelated challenges, rendering it even more tricky. For one, small objects only occupy a few pixels and contain limited information. For another, TIR sensors are typically low-resolution (LR) due to inherent challenges associated with the imaging mechanism of the TIR spectrum. In contrast, high-resolution (HR) RGB sensors are readily available due to their cost-effectiveness and widespread application. Recognizing the importance of HR information, especially in the context of small object detection, we propose a cross-modality high-resolution knowledge distillation framework (CMHRD), which leverages knowledge from the HR-RGB modality and provides a novel strategy for TIR small object detection. The proposed framework introduces three key components: a super-resolution generative distillation loss for cross-modal high-resolution representation learning, a cross-modality affinity distillation loss to extract scene-level cross-modality information, and a response distillation loss aimed at mimicking the HR prediction. To facilitate research on small object detection with HR-RGB and LR-TIR data, we have curated and annotated two datasets, namely NOAA-Seal and VTUAV-det-small. Experimental results on the NOAA-Seal demonstrate that CMHRD yields significant improvements, achieving a remarkable 6.39 mAP50 increase over a strong baseline without introducing additional computational cost during inference. Experiments on single-category dataset VTUAV-det-small and multi-category dataset RTDOD also show consistent improvements brought by CMHRD. The project is available at https://github.com/NNNNerd/CMHRD. Yan Zhang 0115, Xu Lei 0002, Chang Xu 0027, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Few-Shot Object Detection via Variational Feature AggregationabstractAs few-shot object detectors are often trained with abundant base samples and fine-tuned on few-shot novel examples, the learned models are usually biased to base classes and sensitive to the variance of novel examples. To address this issue, we propose a meta-learning framework with two novel feature aggregation schemes. More precisely, we first present a Class-Agnostic Aggregation (CAA) method, where the query and support features can be aggregated regardless of their categories. The interactions between different classes encourage class-agnostic representations and reduce confusion between base and novel classes. Based on the CAA, we then propose a Variational Feature Aggregation (VFA) method, which encodes support examples into class-level support features for robust feature aggregation. We use a variational autoencoder to estimate class distributions and sample variational features from distributions that are more robust to the variance of support examples. Besides, we decouple classification and regression tasks so that VFA is performed on the classification branch without affecting object localization. Extensive experiments on PASCAL VOC and COCO demonstrate that our method significantly outperforms a strong baseline (up to 16%) and previous state-of-the-art methods (4% in average). Jiaming Han, Yuqiang Ren, Jian Ding 0001, Gui-Song Xia |
AAAI | 5 |
| 2023 | HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic SegmentationabstractCurrent semantic segmentation models have achieved great success under the independent and identically distributed (i.i.d.) condition. However, in real-world applications, test data might come from a different domain than training data. Therefore, it is important to improve model robustness against domain differences. This work studies semantic segmentation under the domain generalization setting, where a model is trained only on the source domain and tested on the unseen target domain. Existing works show that Vision Transformers are more robust than CNNs and show that this is related to the visual grouping property of self-attention. In this work, we propose a novel hierarchical grouping transformer (HGFormer) to explicitly group pixels to form part-level masks and then whole-level masks. The masks at different scales aim to segment out both parts and a whole of classes. HGFormer combines mask classification results at both scales for class label prediction. We assemble multiple interesting cross-domain settings by using seven public semantic segmentation datasets. Experiments show that HGFormer yields more robust semantic segmentation results than per-pixel classification methods and flat-grouping transformers, and outperforms previous methods significantly. Code will be available at https://github.com/dingjianswl0l/HGFormer. Jian Ding 0001, Nan Xue 0001, Gui-Song Xia, Bernt Schiele, Dengxin Dai |
CVPR | 3 |
| 2023 | OmniCity: Omnipotent City Understanding with Multi-Level and Multi-View ImagesabstractThis paper presents OmniCity, a new dataset for omnipotent city understanding from multi-level and multi-view images. More precisely, OmniCity contains multi-view satellite images as well as street-level panorama and mono-view images, constituting over 100K pixel-wise annotated images that are well-aligned and collected from 25K geo-locations in New York City. To alleviate the substantial pixel-wise annotation efforts, we propose an efficient street-viewimage annotation pipeline that leverages the existing label maps of satellite view and the transformation relations between different views (satellite, panorama, and mono-view). With the new OmniCity dataset, we provide benchmarks for a variety of tasks including building footprint extraction, height estimation, and building plane/instance/fine-grained segmentation. Compared with existing multi-level and multi-view benchmarks, OmniCity contains a larger number of images with richer annotation types and more views, provides more benchmark results of state-of-the-art models, and introduces a new task for fine-grained building instance segmentation on street-level panorama images. Moreover, OmniCity provides new problem settings for existing tasks, such as cross-view image matching, synthesis, segmentation, detection, etc., and facilitates the developing of new methods for large-scale city understanding, reconstruction, and simulation. The OmniCity dataset as well as the benchmarks will be released at https://city-super.github.io/mnicity/. Yawen Lai, Linning Xu, Yuanbo Xiangli, Conghui He, Gui-Song Xia, Dahua Lin |
CVPR | 7 |
| 2023 | Level-S2fM: Structure from Motion on Neural Level Set of Implicit SurfacesabstractThis paper presents a neural incremental Structure-from-Motion (SfM) approach, Level-S2fM, which estimates the camera poses and scene geometry from a set of uncalibrated images by learning coordinate MLPs for the implicit surfaces and the radiance fields from the established key-point correspondences. Our novel formulation poses some new challenges due to inevitable two-view and few-view configurations in the incremental SfM pipeline, which complicates the optimization of coordinate MLPs for volumetric neural rendering with unknown camera poses. Nevertheless, we demonstrate that the strong inductive basis conveying in the 2D correspondences is promising to tackle those challenges by exploiting the relationship between the ray sampling schemes. Based on this, we revisit the pipeline of incremental SfM and renew the key components, including two-view geometry initialization, the camera poses registration, the 3D points triangulation, and Bundle Adjustment, with a fresh perspective based on neural implicit surfaces. By unifying the scene geometry in small MLP networks through coordinate MLPs, our Level-S2fM treats the zero-level set of the implicit surface as an informative top-down regularization to manage the reconstructed 3D points, reject the outliers in correspondences via querying SDF, and refine the estimated geometries by NBA (Neural BA). Not only does our Level-S2fM lead to promising results on camera pose estimation and scene geometry reconstruction, but it also shows a promising way for neural implicit rendering without knowing camera extrinsic beforehand. Yuxi Xiao, Nan Xue 0001, Tianfu Wu 0001, Gui-Song Xia |
CVPR | 4 |
| 2023 | Dynamic Coarse-to-Fine Learning for Oriented Tiny Object DetectionabstractDetecting arbitrarily oriented tiny objects poses intense challenges to existing detectors, especially for label assignment. Despite the exploration of adaptive label assignment in recent oriented object detectors, the extreme geometry shape and limited feature of oriented tiny objects still induce severe mismatch and imbalance issues. Specifically, the position prior, positive sample feature, and instance are mismatched, and the learning of extreme-shaped objects is biased and unbalanced due to little proper feature supervision. To tackle these issues, we propose a dynamic prior along with the coarse-to-fine assigner, dubbed DCFL. For one thing, we model the prior, label assignment, and object representation all in a dynamic manner to alleviate the mismatch issue. For another, we leverage the coarse prior matching and finer posterior constraint to dynamically assign labels, providing appropriate and relatively balanced supervision for diverse instances. Extensive experiments on six datasets show substantial improvements to the baseline. Notably, we obtain the state-of-the-art performance for one-stage detectors on the DOTA-v1.5, DOTA-v2.0, and DIOR-R datasets under single-scale training and testing. Codes are available at https://github.com/Chasel-Tsui/mmrotate-dcfl. Chang Xu 0027, Jian Ding 0001, Jinwang Wang, Wen Yang 0001, Huai Yu, Lei Yu 0006, Gui-Song Xia |
CVPR | 7 |
| 2023 | Sat2Density: Faithful Density Learning from Satellite-Ground Image PairsabstractThis paper aims to develop an accurate 3D geometry representation of satellite images using satellite-ground image pairs. Our focus is on the challenging problem of 3D-aware ground-views synthesis from a satellite image. We draw inspiration from the density field representation used in volumetric neural rendering and propose a new approach, called Sat2Density. Our method utilizes the properties of ground-view panoramas for the sky and non-sky regions to learn faithful density fields of 3D scenes in a geometric perspective. Unlike other methods that require extra depth information during training, our Sat2Density can automatically learn accurate and faithful 3D geometry via density representation without depth supervision. This advancement significantly improves the ground-view panorama synthesis task. Additionally, our study provides a new geometric perspective to understand the relationship between satellite and ground-view images in 3D space. Ming Qian, Jincheng Xiong, Gui-Song Xia, Nan Xue 0001 |
ICCV | 3 |
| 2023 | Generalizing Event-Based Motion Deblurring in Real-World ScenariosabstractEvent-based motion deblurring has shown promising results by exploiting low-latency events. However, current approaches are limited in their practical usage, as they assume the same spatial resolution of inputs and specific blurriness distributions. This work addresses these limitations and aims to generalize the performance of event-based de-blurring in real-world scenarios. We propose a scale-aware network that allows flexible input spatial scales and enables learning from different temporal scales of motion blur. A two-stage self-supervised learning scheme is then developed to fit real-world data distribution. By utilizing the relativity of blurriness, our approach efficiently ensures the restored brightness and structure of latent images and further generalizes deblurring performance to handle varying spatial and temporal scales of motion blur in a self-distillation manner. Our method is extensively evaluated, demonstrating remarkable performance, and we also introduce a real-world dataset consisting of multi-scale blurry frames and events to facilitate research in event-based deblurring. Xiang Zhang 0022, Lei Yu 0006, Wen Yang 0001, Jianzhuang Liu, Gui-Song Xia |
ICCV | 5 |
| 2023 | Depth and DOF Cues Make A Better Defocus Blur DetectorabstractDefocus blur detection (DBD) separates in-focus and out-of-focus regions in an image. Previous approaches mistakenly mistook homogeneous areas in focus for defocus blur regions, likely due to not considering the internal factors that cause defocus blur. Inspired by the law of depth, depth of field (DOF), and defocus, we propose an approach called D-DFFNet, which incorporates depth and DOF cues in an implicit manner. This allows the model to understand the defocus phenomenon in a more natural way. Our method proposes a depth feature distillation strategy to obtain depth knowledge from a pre-trained monocular depth estimation model and uses a DOF-edge loss to understand the relationship between DOF and depth. Our approach outperforms state-of-the-art methods on public benchmarks and a newly collected large benchmark dataset, EBD. Source codes and EBD dataset are available at: github.com/yuxinjin-whu/D-DFFNet. Ming Qian, Jincheng Xiong, Nan Xue 0001, Gui-Song Xia |
ICME | 5 |
| 2023 | ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud UnderstandingabstractTransformers have been recently explored for 3D point cloud understanding with impressive progress achieved. A large number of points, over 0.1 million, make the global self-attention infeasible for point cloud data. Thus, most methods propose to apply the transformer in a local region, e.g., spherical or cubic window. However, it still contains a large number of Query-Key pairs, which requires high computational costs. In addition, previous methods usually learn the query, key, and value using a linear projection without modeling the local 3D geometric structure. In this paper, we attempt to reduce the costs and model the local geometry prior by developing a new transformer block, named ConDaFormer. Technically, ConDaFormer disassembles the cubic window into three orthogonal 2D planes, leading to fewer points when modeling the attention in a similar range. The disassembling operation is beneficial to enlarging the range of attention without increasing the computational complexity, but ignores some contexts. To provide a remedy, we develop a local structure enhancement strategy that introduces a depth-wise convolution before and after the attention. This scheme can also capture the local geometric information. Taking advantage of these designs, ConDaFormer captures both long-range contextual information and local priors. The effectiveness is demonstrated by experimental results on several 3D point cloud understanding benchmarks. Our code will be available. Lunhao Duan, Shanshan Zhao 0001, Nan Xue 0001, Mingming Gong, Gui-Song Xia, Dacheng Tao |
NeurIPS | 5 |
| 2023 | Detecting building changes with off-nadir aerial images
Chao Pang 0001, Jiang Wu 0003, Jian Ding 0001, Can Song, Gui-Song Xia |
Sci. China Inf. Sci. | 5 |
| 2023 | NOPE-SAC: Neural One-Plane RANSAC for Sparse-View Planar 3D ReconstructionabstractThis article studies the challenging two-view 3D reconstruction problem in a rigorous sparse-view configuration, which is suffering from insufficient correspondences in the input image pairs for camera pose estimation. We present a novel Neural One-PlanE RANSAC framework (termed NOPE-SAC in short) that exerts excellent capability of neural networks to learn one-plane pose hypotheses from 3D plane correspondences. Building on the top of a Siamese network for plane detection, our NOPE-SAC first generates putative plane correspondences with a coarse initial pose. It then feeds the learned 3D plane correspondences into shared MLPs to estimate the one-plane camera pose hypotheses, which are subsequently reweighed in a RANSAC manner to obtain the final camera pose. Because the neural one-plane pose minimizes the number of plane correspondences for adaptive pose hypotheses generation, it enables stable pose voting and reliable pose refinement with a few of plane correspondences for the sparse-view inputs. In the experiments, we demonstrate that our NOPE-SAC significantly improves the camera pose estimation for the two-view inputs with severe viewpoint changes, setting several new state-of-the-art performances on two challenging benchmarks, i.e., MatterPort3D and ScanNet, for sparse-view 3D reconstruction. Bin Tan 0002, Nan Xue 0001, Tianfu Wu 0001, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Learning to Extract Building Footprints From Off-Nadir Aerial ImagesabstractExtracting building footprints from aerial images is essential for precise urban mapping with photogrammetric computer vision technologies. Existing approaches mainly assume that the roof and footprint of a building are well overlapped, which may not hold in off-nadir aerial images as there is often a big offset between them. In this paper, we propose an offset vector learning scheme, which turns the building footprint extraction problem in off-nadir images into an instance-level joint prediction problem of the building roof and its corresponding “roof to footprint” offset vector. Thus the footprint can be estimated by translating the predicted roof mask according to the predicted offset vector. We further propose a simple but effective feature-level offset augmentation module, which can significantly refine the offset vector prediction by introducing little extra cost. Moreover, a new dataset, Buildings in Off-Nadir Aerial Images (BONAI), is created and released in this paper. It contains 268,958 building instances across 3,300 aerial images with fully annotated instance-level roof, footprint, and corresponding offset vector for each building. Experiments on the BONAI dataset demonstrate that our method achieves the state-of-the-art, outperforming other competitors by 3.37 to 7.39 points in F1-score. The codes, datasets, and trained models are available athttps://github.com/jwwangchn/BONAI.git. Jinwang Wang, Lingxuan Meng, Wen Yang 0001, Lei Yu 0006, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Holistically-Attracted Wireframe Parsing: From Supervised to Self-Supervised LearningabstractThis article presents Holistically-Attracted Wireframe Parsing (HAWP), a method for geometric analysis of 2D images containing wireframes formed by line segments and junctions. HAWP utilizes a parsimonious Holistic Attraction (HAT) field representation that encodes line segments using a closed-form 4D geometric vector field. The proposed HAWP consists of three sequential components empowered by end-to-end and HAT-driven designs: 1) generating a dense set of line segments from HAT fields and endpoint proposals from heatmaps, 2) binding the dense line segments to sparse endpoint proposals to produce initial wireframes, and 3) filtering false positive proposals through a novel endpoint-decoupled line-of-interest aligning (EPD LOIAlign) module that captures the co-occurrence between endpoint proposals and HAT fields for better verification. Thanks to our novel designs, HAWPv2 shows strong performance in fully supervised learning, while HAWPv3 excels in self-supervised learning, achieving superior repeatability scores and efficient training (24 GPU hours on a single GPU). Furthermore, HAWPv3 exhibits a promising potential for wireframe parsing in out-of-distribution images without providing ground truth labels of wireframes. Nan Xue 0001, Tianfu Wu 0001, Song Bai 0001, Fudong Wang 0001, Gui-Song Xia, Liangpei Zhang 0001, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Learning to Super-Resolve Blurry Images With EventsabstractSuper-Resolution from a single motion Blurred image (SRB) is a severely ill-posed problem due to the joint degradation of motion blurs and low spatial resolution. In this article, we employ events to alleviate the burden of SRB and propose an Event-enhanced SRB (E-SRB) algorithm, which can generate a sequence of sharp and clear images with High Resolution (HR) from a single blurry image with Low Resolution (LR). To achieve this end, we formulate an event-enhanced degeneration model to consider the low spatial resolution, motion blurs, and event noises simultaneously. We then build an event-enhanced Sparse Learning Network (eSL-Net++) upon a dual sparse learning scheme where both events and intensity frames are modeled with sparse representations. Furthermore, we propose an event shuffle-and-merge scheme to extend the single-frame SRB to the sequence-frame SRB without any additional training process. Experimental results on synthetic and real-world datasets show that the proposed eSL-Net++ outperforms state-of-the-art methods by a large margin. Datasets, codes, and more results are available at https://github.com/ShinyWang33/eSL-Net-Plusplus. Lei Yu 0006, Bishan Wang, Xiang Zhang 0022, Wen Yang 0001, Jianzhuang Liu, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Learning to See Through With EventsabstractAlthough synthetic aperture imaging (SAI) can achieve the seeing-through effect by blurring out off-focus foreground occlusions while recovering in-focus occluded scenes from multi-view images, its performance is often deteriorated by dense occlusions and extreme lighting conditions. To address the problem, this paper presents an Event-based SAI (E-SAI) method by relying on the asynchronous events with extremely low latency and high dynamic range acquired by an event camera. Specifically, the collected events are first refocused by a Refocus-Net module to align in-focus events while scattering out off-focus ones. Following that, a hybrid network composed of spiking neural networks (SNNs) and convolutional neural networks (CNNs) is proposed to encode the spatio-temporal information from the refocused events and reconstruct a visual image of the occluded targets. Extensive experiments demonstrate that our proposed E-SAI method can achieve remarkable performance in dealing with very dense occlusions and extreme lighting conditions and produce high-quality images from pure events. Codes and datasets are available at https://dvs-whu.cn/projects/esai/. Lei Yu 0006, Xiang Zhang 0022, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | A3Track: Achieving Precise Target Tracking in Aerial Images With Receptive Field AlignmentabstractTracking arbitrary objects in aerial images presents formidable challenges to existing trackers. Among these challenges, the large scale variation and arbitrary geometry shape of visual targets are pronounced, resulting in two-fold mismatch issues between the feature receptive field and the tracking target. For one, there is a mismatch between the prior receptive field center and arbitrary-shaped targets. For another, the single receptive field mismatches the significantly scale-varied targets in the aerial imagery. To handle these challenges, we propose to Achieve precise Aerial tracking with receptive field Alignment, dubbed A3Track. The proposed A3Track is comprised of two modules: a Receptive Field Alignment (RFA) module and a Pyramid Receptive Field (PRF) module. First of all, we transform and update the receptive field center progressively, which drives the feature sampling location onto the targets’ main body, thus gradually yielding precise feature representation for arbitrary-shaped targets. We term this progressively updating process as the Receptive Field Alignment. Moreover, the PRF module constructs a set of pyramid features for the target, providing a multi-scale receptive field to handle the large scale variation of tracking objects. On four benchmarks, the new tracker A3Track achieves leading performance compared with existing methods and shows consistent improvements over baselines. The project is available at: https://chnleixu.github.io/A3Track-web/. Xu Lei 0002, Chang Xu 0027, Wensheng Cheng, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | An Empirical Study of Remote Sensing PretrainingabstractDeep learning has largely reshaped remote sensing (RS) research for aerial image understanding and made a great success. Nevertheless, most of the existing deep models are initialized with the ImageNet pretrained weights since natural images inevitably present a large domain gap relative to aerial images, probably limiting the fine-tuning performance on downstream aerial scene tasks. This issue motivates us to conduct an empirical study of RS pretraining (RSP) on aerial images. To this end, we train different networks from scratch with the help of the largest RS scene recognition dataset up to now—MillionAID—to obtain a series of RS pretrained backbones, including both convolutional neural networks (CNNs) and vision transformers, such as Swin and ViTAE, which have shown promising performance on computer vision tasks. Then, we investigate the impact of RSP on representative downstream tasks, including scene recognition, semantic segmentation, object detection, and change detection using these CNN and vision transformer backbones. Empirical study shows that RSP can help deliver distinctive performances in scene recognition tasks and in perceiving RS-related semantics, such as “Bridge” and “Airplane.” We also find that, although RSP mitigates the data discrepancies of traditional ImageNet pretraining on RS images, it may still suffer from task discrepancies, where downstream tasks require different representations from scene recognition tasks. These findings call for further research efforts on both large-scale pretraining datasets and effective pretraining methods. The codes and pretrained models will be released athttps://github.com/ViTAE-Transformer/ViTAE-Transformer-Remote-Sensing. Di Wang 0023, Jing Zhang 0037, Bo Du 0001, Gui-Song Xia, Dacheng Tao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Bayesian Collaborative Learning for Whole-Slide Image ClassificationabstractWhole-slide image (WSI) classification is fundamental to computational pathology, which is challenging in extra-high resolution, expensive manual annotation, data heterogeneity, etc. Multiple instance learning (MIL) provides a promising way towards WSI classification, which nevertheless suffers from the memory bottleneck issue inherently, due to the gigapixel high resolution. To avoid this issue, the overwhelming majority of existing approaches have to decouple the feature encoder and the MIL aggregator in MIL networks, which may largely degrade the performance. Towards this end, this paper presents a Bayesian Collaborative Learning (BCL) framework to address the memory bottleneck issue with WSI classification. Our basic idea is to introduce an auxiliary patch classifier to interact with the target MIL classifier to be learned, so that the feature encoder and the MIL aggregator in the MIL classifier can be learned collaboratively while preventing the memory bottleneck issue. Such a collaborative learning procedure is formulated under a unified Bayesian probabilistic framework and a principled Expectation-Maximization algorithm is developed to infer the optimal model parameters iteratively. As an implementation of the E-step, an effective quality-aware pseudo labeling strategy is also suggested. The proposed BCL is extensively evaluated on three publicly available WSI datasets, i.e., CAMELYON16, TCGA-NSCLC and TCGA-RCC, achieving an AUC of 95.6%, 96.0% and 97.5% respectively, which consistently outperforms all the methods compared. Comprehensive analysis and discussion will also be presented for in-depth understanding of the method. To promote future work, our source code is released at: https://github.com/Zero-We/BCL. Jin-Gang Yu, Zihao Wu 0004, Shule Deng, Qihang Wu, Zhongtang Xiong, Tianyou Yu, Gui-Song Xia, Qingping Jiang, Yuanqing Li 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Single Image Deraining With Continuous Rain Density EstimationabstractSingle image deraining (SIDR) often suffers from over/under deraining due to the nonuniformity of rain densities and the variety of raindrop scales. In this paper, we propose acontinuousdensity-guided network (CODE-Net) for SIDR. Particularly, it is composed of a rain streak extractor and a denoiser, where the convolutional sparse coding (CSC) is exploited to filter out noises from the extracted rain streaks. Inspired by the reweighted iterative soft-threshold (ISTA) for CSC, we address the problem of continuous rain density estimation by learning the weights with channel attention blocks from sparse codes. We further develop a multiscale strategy to depict rain streaks appearing at different scales. Experiments on synthetic and real-world data demonstrate the superiority of our methods over recent state-of-the-arts, in terms of both quantitative and qualitative results. Additionally, instead of quantizing rain density with several levels, our CODE-Net can provide continuous-valued estimations of rain densities, which is more desirable in real applications. Lei Yu 0006, Bishan Wang, Jingwei He, Gui-Song Xia, Wen Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | HoW-3D: Holistic 3D Wireframe Perception from a Single ImageabstractThis paper studies the problem of holistic 3D wireframe perception (HoW-3D), a new task of perceiving both the visible 3D wireframes and the invisible ones from single-view 2D images. As the non-front surfaces of an object cannot be directly observed in a single view, estimating the nonline-of-sight (NLOS) geometries in HoW-3D is a fundamentally challenging problem and remains open in computer vision. We study the problem of HoW-3D by proposing an ABC-HoW benchmark, which is created on top of CAD models sourced from the ABC-dataset with 12k single-view images and the corresponding holistic 3D wireframe models. With our large-scale ABC-HoW benchmark available, we present a novel Deep Spatial Gestalt (DSG) model to learn the visible junctions and line segments as the basis and then infer the NLOS 3D structures from the visible cues by following the Gestalt principles of human vision systems. In our experiments, we demonstrate that our DSG model performs very well in inferring the holistic 3D wireframes from single-view images. Compared with the strong baseline methods, our DSG model outperforms the previous wire-frame detectors in detecting the invisible line geometry in single-view images and is even very competitive with prior arts that take high-fidelity PointCloud as inputs on reconstructing 3D wireframes. Bin Tan 0002, Nan Xue 0001, Tianfu Wu 0001, Xianwei Zheng, Gui-Song Xia |
3DV | 6 |
| 2022 | Learning Local-Global Contextual Adaptation for Multi-Person Pose EstimationabstractThis paper studies the problem of multi-person pose estimation in a bottom-up fashion. With a new and strong observation that the localization issue of the center-offset formulation can be remedied in a local-window search scheme in an ideal situation, we propose a multi-person pose estimation approach, dubbed as LOGO-CAP, by learning the LOcal-GlObal Contextual Adaptation for human Pose. Specifically, our approach learns the keypoint attraction maps (KAMs) from the local keypoints expansion maps (KEMs) in small local windows in the first step, which are subsequently treated as dynamic convolutional kernels on the keypoints-focused global heatmaps for contextual adaptation, achieving accurate multi-person pose estimation. Our method is end-to-end trainable with near real-time inference speed in a single forward pass, obtaining state-of-the-art performance on the COCO keypoint benchmark for bottom-up human pose estimation. With the COCO trained model, our method also outperforms prior arts by a large margin on the challenging OCHuman dataset. Nan Xue 0001, Tianfu Wu 0001, Gui-Song Xia, Liangpei Zhang 0001 |
CVPR | 3 |
| 2022 | Decoupling Zero-Shot Semantic SegmentationabstractZero-shot semantic segmentation (ZS3) aims to segment the novel categories that have not been seen in the training. Existing works formulate ZS3 as a pixel-level zeroshot classification problem, and transfer semantic knowledge from seen classes to unseen ones with the help of language models pre-trained only with texts. While simple, the pixel-level ZS3 formulation shows the limited capability to integrate vision-language models that are often pre-trained with image-text pairs and currently demonstrate great potential for vision tasks. Inspired by the observation that humans often perform segment-level semantic labeling, we propose to decouple the ZS3 into two sub-tasks: 1) a classagnostic grouping task to group the pixels into segments. 2) a zero-shot classification task on segments. The former task does not involve category information and can be directly transferred to group pixels for unseen classes. The latter task performs at segment-level and provides a natural way to leverage large-scale vision-language models pre-trained with image-text pairs (e.g. CLIP) for ZS3. Based on the decoupling formulation, we propose a simple and effective zero-shot semantic segmentation model, called ZegFormer, which outperforms the previous methods on ZS3 standard benchmarks by large margins, e.g., 22 points on the PAS-CAL VOC and 3 points on the COCO-Stuff in terms of mIoU for unseen classes. Code will be released at https://github.com/dingjiansw101/ZegFormer. Jian Ding 0001, Nan Xue 0001, Gui-Song Xia, Dengxin Dai |
CVPR | 3 |
| 2022 | Expanding Low-Density Latent Regions for Open-Set Object DetectionabstractModern object detectors have achieved impressive progress under the close-set setup. However, open-set object detection (OSOD) remains challenging since objects of unknown categories are often misclassified to existing known classes. In this work, we propose to identify unknown objects by separating high/low-density regions in the latent space, based on the consensus that unknown objects are usually distributed in low-density latent regions. As traditional threshold-based methods only maintain limited low-density regions, which cannot cover all unknown objects, we present a novel Openset Detector (OpenDet) with expanded low-density regions. To this aim, we equip Open-Det with two learners, Contrastive Feature Learner (CFL) and Unknown Probability Learner (UPL). CFL performs instance-level contrastive learning to encourage compact features of known classes, leaving more low-density regions for unknown classes; UPL optimizes unknown probability based on the uncertainty of predictions, which further divides more low-density regions around the cluster of known classes. Thus, unknown objects in low-density regions can be easily identified with the learned unknown probability. Extensive experiments demonstrate that our method can significantly improve the OSOD performance, e.g., OpenDet reduces the Absolute Open-Set Errors by 25%-35% on six OSOD benchmarks. Code is available at: https://github.com/csuhan/opendet2. Jiaming Han, Yuqiang Ren, Jian Ding 0001, Xingjia Pan, Gui-Song Xia |
CVPR | 6 |
| 2022 | Revisiting Document Image Dewarping by Grid RegularizationabstractThis paper addresses the problem of document image dewarping, which aims at eliminating the geometric distortion in document images for document digitization. Instead of designing a better neural network to approximate the optical flow fields between the inputs and outputs, we pursue the best readability by taking the text lines and the document boundaries into account from a constrained optimization perspective. Specifically, our proposed method first learns the boundary points and the pixels in the text lines and then follows the most simple observation that the boundaries and text lines in both horizontal and vertical directions should be kept after dewarping to introduce a novel grid regularization scheme. To obtain the final forward mapping for dewarping, we solve an optimization problem with our proposed grid regularization. The experiments comprehensively demonstrate that our proposed approach outperforms the prior arts by large margins in terms of readability (with the metrics of Character Errors Rate and the Edit Distance) while maintaining the best image quality on the publicly-available DocUNet benchmark. Xiangwei Jiang, Rujiao Long, Nan Xue 0001, Zhibo Yang 0003, Cong Yao, Gui-Song Xia |
CVPR | 6 |
| 2022 | RFLA: Gaussian Receptive Field Based Label Assignment for Tiny Object Detection
Chang Xu 0027, Jinwang Wang, Wen Yang 0001, Huai Yu, Lei Yu 0006, Gui-Song Xia |
ECCV (9) | 6 |
| 2022 | Partial Wasserstein Adversarial Network for Non-rigid Point Set Registration
Nan Xue 0001, Gui-Song Xia |
ICLR | 4 |
| 2022 | Object Detection in Aerial Images: A Large-Scale Benchmark and ChallengesabstractIn he past decade, object detection has achieved significant progress in natural images but not in aerial images, due to the massive variations in the scale and orientation of objects caused by the bird's-eye view of aerial images. More importantly, the lack of large-scale benchmarks has become a major obstacle to the development of object detection in aerial images (ODAI). In this paper, we present a large-scale Dataset of Object deTection in Aerial images (DOTA) and comprehensive baselines for ODAI. The proposed DOTA dataset contains 1,793,658 object instances of 18 categories of oriented-bounding-box annotations collected from 11,268 aerial images. Based on this large-scale and well-annotated dataset, we build baselines covering 10 state-of-the-art algorithms with over 70 configurations, where the speed and accuracy performances of each model have been evaluated. Furthermore, we provide a code library for ODAI and build a website for evaluating different algorithms. Previous challenges run on DOTA have attracted more than 1300 teams worldwide. We believe that the expanded large-scale DOTA dataset, the extensive baselines, the code library and the challenges can facilitate the designs of robust algorithms and reproducible research on the problem of object detection in aerial images. Jian Ding 0001, Nan Xue 0001, Gui-Song Xia, Xiang Bai, Wen Yang 0001, Michael Ying Yang, Serge J. Belongie, Jiebo Luo 0001, Mihai Datcu, Marcello Pelillo, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Unmixing Convolutional Features for Crisp Edge DetectionabstractThis article presents a context-aware tracing strategy (CATS) for crisp edge detection with deep edge detectors, based on an observation that the localization ambiguity of deep edge detectors is mainly caused by the mixing phenomenon of convolutional neural networks: Feature mixing in edge classification and side mixing during fusing side predictions. The CATS consists of two modules: A novel tracing loss that performs feature unmixing by tracing boundaries for better side edge learning, and a context-aware fusion block that tackles the side mixing by aggregating the complementary merits of learned side edges. Experiments demonstrate that the proposed CATS can be integrated into modern deep edge detectors to improve localization accuracy. With the vanilla VGG16 backbone, in terms of BSDS500 dataset, our CATS improves the F-measure (ODS) of the RCF and BDCN deep edge detectors by 12 and 6 percent, respectively when evaluating without using the morphological non-maximal suppression scheme for edge detection. Linxi Huan, Nan Xue 0001, Xianwei Zheng, Wei He 0003, Jianya Gong, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | All Grains, One Scheme (AGOS): Learning Multigrain Instance Representation for Aerial Scene ClassificationabstractAerial scene classification remains challenging as: 1) the size of key objects in determining the scene scheme varies greatly; 2) many objects irrelevant to the scene scheme are often flooded in the image. Hence, how to effectively perceive the region of interests (RoIs) from a variety of sizes and build more discriminative representation from such complicated object distribution is vital to understand an aerial scene. In this paper, we propose a novelall grains, one scheme(AGOS) framework to tackle these challenges.To the best of our knowledge, it is the first work to extend the classic multiple instance learning into multi-grain formulation. Specially, it consists of a multi-grain perception module (MGP), a multi-branch multi-instance representation module (MBMIR) and a self-aligned semantic fusion (SSF) module. Firstly, our MGP preserves the differential dilated convolutional features from the backbone, which magnifies the discriminative information from multi-grains. Then, our MBMIR highlights the key instances in the multi-grain representation under the MIL formulation. Finally, our SSF allows our framework to learn the same scene scheme from multi-grain instance representations and fuses them, so that the entire framework is optimized as a whole. Notably, our AGOS is flexible and can be easily adapted to existing CNNs in a plug-and-play manner. Extensive experiments on UCM, AID and NWPU benchmarks demonstrate that our AGOS achieves a comparable performance against the state-of-the-art methods. Qi Bi, Beichen Zhou, Kun Qin, Qinghao Ye, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Align Deep Features for Oriented Object DetectionabstractThe past decade has witnessed significant progress on detecting objects in aerial images that are often distributed with large-scale variations and arbitrary orientations. However, most of existing methods rely on heuristically defined anchors with different scales, angles, and aspect ratios, and usually suffer from severe misalignment between anchor boxes (ABs) and axis-aligned convolutional features, which lead to the common inconsistency between the classification score and localization accuracy. To address this issue, we propose asingle-shot alignment network(S2A-Net) consisting of two modules: a feature alignment module (FAM) and an oriented detection module (ODM). The FAM can generate high-quality anchors with an anchor refinement network and adaptively align the convolutional features according to the ABs with a novel alignment convolution. The ODM first adopts active rotating filters to encode the orientation information and then produces orientation-sensitive and orientation-invariant features to alleviate the inconsistency between classification score and localization accuracy. Besides, we further explore the approach to detect objects in large-size images, which leads to a better trade-off between speed and accuracy. Extensive experiments demonstrate that our method can achieve the state-of-the-art performance on two commonly used aerial objects’ data sets (i.e., DOTA and HRSC2016) while keeping high efficiency. Jiaming Han, Jian Ding 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | FuTH-Net: Fusing Temporal Relations and Holistic Features for Aerial Video ClassificationabstractUnmanned aerial vehicles (UAVs) are now widely applied to data acquisition due to its low cost and fast mobility. With the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this, current research mainly focuses on extracting a holistic feature with convolutions along both spatial and temporal dimensions. However, these methods are limited by small temporal receptive fields and cannot adequately capture long-term temporal dependencies that are important for describing complicated dynamics. In this article, we propose a novel deep neural network, termed Fusing Temporal relations and Holistic features for aerial video classification (FuTH-Net), to model not only holistic features but also temporal relations for aerial video classification. Furthermore, the holistic features are refined by the multiscale temporal relations in a novel fusion module for yielding more discriminative video representations. More specially, FuTH-Net employs a two-pathway architecture: 1) a holistic representation pathway to learn a general feature of both frame appearances and short-term temporal variations and 2) a temporal relation pathway to capture multiscale temporal relations across arbitrary frames, providing long-term temporal dependencies. Afterward, a novel fusion module is proposed to spatiotemporally integrate the two features learned from the two pathways. Our model is evaluated on two aerial video classification datasets, ERA and Drone-Action, and achieves the state-of-the-art results. This demonstrates its effectiveness and good generalization capacity across different recognition tasks (event classification and human action recognition). To facilitate further research, we release the code athttps://gitlab.lrz.de/ai4eo/reasoning/futh-net. Pu Jin, Lichao Mou, Yuansheng Hua, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Anomaly Detection in Aerial Videos With TransformersabstractUnmanned aerial vehicles (UAVs) are widely applied for purposes of inspection, search, and rescue operations by the virtue of low-cost, large-coverage, real-time, and high-resolution data acquisition capacities. Massive volumes of aerial videos are produced in these processes, in which normal events often account for an overwhelming proportion. It is extremely difficult to localize and extract abnormal events containing potentially valuable information from long video streams manually. Therefore, we are dedicated to developing anomaly detection methods to solve this issue. In this paper, we create a new dataset, named Drone-Anomaly, for anomaly detection in aerial videos. This dataset provides 37 training video sequences and 22 testing video sequences from 7 different realistic scenes with various anomalous events. There are 87,488 color video frames (51,635 for training and 35,853 for testing) with the size of 640 × 640 at 30 frames per second. Based on this dataset, we evaluate existing methods and offer a benchmark for this task. Furthermore, we present a new baseline model, ANomaly Detection with Transformers (ANDT), which treats consecutive video frames as a sequence of tubelets, utilizes a Transformer encoder to learn feature representations from the sequence, and leverages a decoder to predict the next frame. Our network models normality in the training phase and identifies an event with unpredictable temporal dynamics as an anomaly in the test phase. Moreover, To comprehensively evaluate the performance of our proposed method, we use not only our Drone-Anomaly dataset but also another dataset. We will make our dataset and code publicly available. A demo video is available at https://youtu.be/ancczYryOBY. We make our dataset and code publicly available1. Pu Jin, Lichao Mou, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Locally Nonlinear Affine Verification for Multisensor Image MatchingabstractMatching local features between two overlapped images is a fundamental task in photogrammetry and remote sensing. However, images acquired by multiple sensors often differ substantially in properties, thus posing a great challenge to the robustness and flexibility of feature matching methods. In this article, we propose a locally non-linear affine verification (LAV) method for robust multisensor image matching. The main idea of the LAV is the development of a nonlinear regression formulation that practically models the nonlinear deviation of a real surface around a point from its tangent plane during affine verification. Specifically, we start by selecting a restricted set of reliable and well-distributed putative matches as the matching seeds and assign them with neighbors to construct search spaces. In each search space, the regression seeks the smoothest affine model consistent with the latent correct matches, thereby deriving a set of affine parameters to verify correspondence hypotheses for true matches. The verification can be extended to all nearest neighbor matches to discover additional inlier matches. Evaluation on multisensor image datasets with different extents of variations in viewpoint, scale, illumination, and appearance shows that the proposed LAV consistently outperforms existing methods. LAV can achieve a considerable number of high-quality matches, in cases where existing methods provide few or no correct matches. Xianwei Zheng, Mingyue Dong, Gui-Song Xia, Hanjiang Xiong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Hidden Path Selection Network for Semantic Segmentation of Remote Sensing ImagesabstractTargeting at depicting land covers with pixelwise semantic categories, semantic segmentation in remote sensing images needs to portray diverse distributions over vast geographical locations, which is difficult to be achieved by the homogeneous pixelwise forward paths in the architectures of existing deep models. Although specific algorithms have been designed to select pixelwise adaptive forward paths for natural image analysis, it still lacks theoretical supports on how to obtain optimal selections. In this article, we provide mathematical analyses in terms of the parameter optimization, which guides us to design a method called hidden path selection network (HPS-Net). With the help of hidden variables deriving from an extra mini-branch, HPS-Net is able to tackle the inherent problem about inaccessible global optimums by adjusting the direct relationships between feature maps and pixelwise path selections in existing algorithms, which we call hidden path selection. For the better training and evaluation, we further refine and expand the 5-class Gaofen image dataset (GID-5) to a new one with 15 land-cover categories, i.e., GID-15. The experimental results on both GID-5 and GID-15 demonstrate that the proposed modules can stably improve the performance of different deep structures, which validates the proposed mathematical analyses. Kunping Yang, Xin-Yi Tong 0003, Gui-Song Xia, Weiming Shen 0002, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Asymmetric Siamese Networks for Semantic Change Detection in Aerial ImagesabstractGiven two multitemporal aerial images, semantic change detection (SCD) aims to locate the land-cover variations and identify their change types with pixelwise boundaries. This problem is vital in many earth vision-related tasks, such as precise urban planning and natural resource management. Existing state-of-the-art algorithms mainly identify the changed pixels by applying homogeneous operations on each input image and comparing the extracted features. However, in changed regions, totally different land-cover distributions often require heterogeneous feature extraction procedures for images acquired at different times. In this article, we present an asymmetric Siamese network (ASN) to locate and identify semantic changes through feature pairs obtained from modules of widely different structures, which involves areas of various sizes and applies different quantities of parameters to factor in the discrepancy across land-cover distributions during different times. To better train and evaluate our model, we create a large-scale well-annotated SEmantic Change detectiON Dataset (SECOND), while an adaptive threshold learning (ATL) module and a separated kappa (SeK) coefficient are proposed to alleviate the influences of label imbalance in model training and evaluation. The experimental results demonstrate that the proposed model can stably outperform the state-of-the-art algorithms with different encoder backbones. Kunping Yang, Gui-Song Xia, Zicheng Liu 0003, Bo Du 0001, Wen Yang 0001, Marcello Pelillo, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Optical-Enhanced Oil Tank Detection in High-Resolution SAR ImagesabstractIn recent years, object detection in high-resolution SAR images has made significant progress, especially after the introduction of deep learning. However, objects like dense oil tanks, which are compactly arranged in SAR images, are still challenging to recognize due to the unique imaging mechanism of SAR. Inspired by human learning from comparison, we propose a multi-stage framework for oil tank detection in SAR images using optical image enhancement. Specifically, in the training stage, we build a teacher-student network to align the semantic information between the two modalities, where the optical features are used to guide the corresponding SAR feature learning. While in the inference stage, the learned network detects oil tanks using only SAR images as input. Besides, a pre-training stage before training is applied to further improve the network’s ability for SAR feature extraction, which is realized by the proposed paired optical-SAR self-supervised learning. To verify the effectiveness of the proposed method, we perform experiments on our newly built SpaceNet6-OTD dataset. Extensive experiments demonstrate that the proposed method can effectively improve the accuracy of detecting oil tanks in SAR images. Datasets, codes, and more results will be released at: https://EIS-VIPG.github.io/SpaceNet6-OTD/. Ruixiang Zhang, Haowen Guo, Wen Yang 0001, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2021 | Deep Graph Matching Under Quadratic ConstraintabstractRecently, deep learning based methods have demonstrated promising results on the graph matching problem, by relying on the descriptive capability of deep features extracted on graph nodes. However, one main limitation with existing deep graph matching (DGM) methods lies in their ignorance of explicit constraint of graph structures, which may lead the model to be trapped into local minimum in training. In this paper, we propose to explicitly formulate pairwise graph structures as a quadratic constraint incorporated into the DGM framework. The quadratic constraint minimizes the pairwise structural discrepancy between graphs, which can reduce the ambiguities brought by only using the extracted CNN features. Moreover, we present a differentiable implementation to the quadratic constrained-optimization such that it is compatible with the unconstrained deep learning optimizer. To give more precise and proper supervision, a well-designed false matching loss against class imbalance is proposed, which can better penalize the false negatives and false positives with less overfitting. Exhaustive experiments demonstrate that our method achieves competitive performance on real-world datasets. The code is available at: https://github.com/Zerg-Overmind/QC-DGM. Quankai Gao, Fudong Wang 0001, Nan Xue 0001, Jin-Gang Yu, Gui-Song Xia |
CVPR | 5 |
| 2021 | ReDet: A Rotation-Equivariant Detector for Aerial Object DetectionabstractRecently, object detection in aerial images has gained much attention in computer vision. Different from objects in natural images, aerial objects are often distributed with arbitrary orientation. Therefore, the detector requires more parameters to encode the orientation information, which are often highly redundant and inefficient. Moreover, as ordinary CNNs do not explicitly model the orientation variation, large amounts of rotation augmented data is needed to train an accurate object detector. In this paper, we propose a Rotation-equivariant Detector (ReDet) to address these issues, which explicitly encodes rotation equivariance and rotation invariance. More precisely, we incorporate rotation-equivariant networks into the detector to extract rotation-equivariant features, which can accurately predict the orientation and lead to a huge reduction of model size. Based on the rotation-equivariant features, we also present Rotation-invariant RoI Align (RiRoI Align), which adaptively extracts rotation-invariant features from equivariant features according to the orientation of RoI. Extensive experiments on several challenging aerial image datasets DOTA-v1.0, DOTA-v1.5 and HRSC2016, show that our method can achieve state-of-the-art performance on the task of aerial object detection. Compared with previous best results, our ReDet gains 1.2, 3.5 and 2.6 mAP on DOTA-v1.0, DOTA-v1.5 and HRSC2016 respectively while reducing the number of parameters by 60% (313 Mb vs. 121 Mb). The code is available at: https://github.com/csuhan/ReDet. Jiaming Han, Jian Ding 0001, Nan Xue 0001, Gui-Song Xia |
CVPR | 4 |
| 2021 | Event-Based Synthetic Aperture Imaging With a Hybrid NetworkabstractSynthetic aperture imaging (SAI) is able to achieve the see through effect by blurring out the off-focus foreground occlusions and reconstructing the in-focus occluded targets from multi-view images. However, very dense occlusions and extreme lighting conditions may bring significant disturbances to the SAI based on conventional frame-based cameras, leading to performance degeneration. To address these problems, we propose a novel SAI system based on the event camera which can produce asynchronous events with extremely low latency and high dynamic range. Thus, it can eliminate the interference of dense occlusions by measuring with almost continuous views, and simultaneously tackle the over/under exposure problems. To reconstruct the occluded targets, we propose a hybrid encoder-decoder network composed of spiking neural networks (SNNs) and convolutional neural networks (CNNs). In the hybrid network, the spatio-temporal information of the collected events is first encoded by SNN layers, and then transformed to the visual image of the occluded targets by a style-transfer CNN decoder. Through experiments, the proposed method shows remarkable performance in dealing with very dense occlusions and extreme lighting conditions, and high quality visual images can be reconstructed using pure event data. Xiang Zhang 0022, Lei Yu 0006, Wen Yang 0001, Gui-Song Xia |
CVPR | 5 |
| 2021 | 3D Building Reconstruction from Monocular Remote Sensing Imagesabstract3D building reconstruction from monocular remote sensing imagery is an important research problem and an economic solution to large-scale city modeling, compared with reconstruction from LiDAR data and multi-view imagery. However, several challenges such as the partial invisibility of building footprints and facades, the serious shadow effect, and the extreme variance of building height in large-scale areas, have restricted the existing monocular image based building reconstruction studies to certain application scenes, i.e., modeling simple low-rise buildings from near-nadir images. In this study, we propose a novel 3D building reconstruction method for monocular remote sensing images, which tackles the above difficulties, thus providing an appealing solution for more complicated scenarios. We design a multi-task building reconstruction network, named MTBR-Net, to learn the geometric property of oblique images, the key components of a 3D building model and their relations via four semantic-related and three offset-related tasks. The network outputs are further integrated by a prior knowledge based 3D model optimization method to produce the the final 3D building models. Results on a public 3D reconstruction dataset and a novel released dataset demonstrate that our method improves the height estimation performance by over 40% and the segmentation F1-score by 2% - 4% compared with current state-of-the-art. Lingxuan Meng, Jinwang Wang, Conghui He, Gui-Song Xia, Dahua Lin |
ICCV | 5 |
| 2021 | Parsing Table Structures in the WildabstractThis paper tackles the problem of table structure parsing (TSP) from images in the wild. In contrast to existing studies that mainly focus on parsing well-aligned tabular images with simple layouts from scanned PDF documents, we aim to establish a practical table structure parsing system for real-world scenarios where tabular input images are taken or scanned with severe deformation, bending or occlusions. For designing such a system, we propose an approach named Cycle-CenterNet on the top of CenterNet with a novel cycle-pairing module to simultaneously detect and group tabular cells into structured tables. In the cycle-pairing module, a new pairing loss function is proposed for the network training. Alongside with our Cycle-CenterNet, we also present a large-scale dataset, named Wired Table in the Wild (WTW), which includes well-annotated structure parsing of multiple style tables in several scenes like photo, scanning files, web pages, etc.. In experiments, we demonstrate that our Cycle-CenterNet consistently achieves the best accuracy of table structure parsing on the new WTW dataset by 24.6% absolute improvement evaluated by the TEDS metric. A more comprehensive experimental analysis also validates the advantages of our proposed methods for the TSP task. Rujiao Long, Nan Xue 0001, Feiyu Gao, Zhibo Yang 0003, Yongpan Wang, Gui-Song Xia |
ICCV | 7 |
| 2021 | PlaneTR: Structure-Guided Transformers for 3D Plane RecoveryabstractThis paper presents a neural network built upon Transformers, namely PlaneTR, to simultaneously detect and reconstruct planes from a single image. Different from previous methods, PlaneTR jointly leverages the context information and the geometric structures in a sequence-to-sequence way to holistically detect plane instances in one forward pass. Specifically, we represent the geometric structures as line segments and conduct the network with three main components: (i) context and line segments encoders, (ii) a structure-guided plane decoder, (iii) a pixel-wise plane embedding decoder. Given an image and its detected line segments, PlaneTR generates the context and line segment sequences via two specially designed encoders and then feeds them into a Transformers-based decoder to directly predict a sequence of plane instances by simultaneously considering the context and global structure cues. Finally, the pixel-wise embeddings are computed to assign each pixel to one predicted plane instance which is nearest to it in embedding space. Comprehensive experiments demonstrate that PlaneTR achieves state-of-the-art performance on the ScanNet and NYUv2 datasets. Bin Tan 0002, Nan Xue 0001, Song Bai 0001, Tianfu Wu 0001, Gui-Song Xia |
ICCV | 5 |
| 2021 | Motion Deblurring with Real EventsabstractIn this paper, we propose an end-to-end learning framework for event-based motion deblurring in a self-supervised manner, where real-world events are exploited to alleviate the performance degradation caused by data inconsistency. To achieve this end, optical flows are predicted from events, with which the blurry consistency and photometric consistency are exploited to enable self-supervision on the deblurring network with real-world data. Furthermore, a piecewise linear motion model is proposed to take into account motion non-linearities and thus leads to an accurate model for the physical formation of motion blurs in the real-world scenario. Extensive evaluation on both synthetic and real motion blur datasets demonstrates that the proposed algorithm bridges the gap between simulated and real-world motion blurs and shows remarkable performance for eventbased motion deblurring in real-world scenarios. Lei Yu 0006, Bishan Wang, Wen Yang 0001, Gui-Song Xia, Xu Jia 0012, Zhendong Qiao, Jianzhuang Liu |
ICCV | 5 |
| 2021 | Toward Dataset Construction for Remote Sensing Image InterpretationabstractWith the rapid advancement of remote sensing (RS) technology, RS image interpretation has made great progress and been widely used in broad applications, in which the constructed benchmark datasets for developing and testing intelligent interpretation algorithms have been playing an increasingly critical role. Motivated by the essential prerequisites of dataset in the development of RS image interpretation algorithms, this manuscript provides a discussion on dataset construction for RS image interpretation. Specifically, we first analyze the current challenges of developing algorithms for RS image interpretation and a review on the widespread RS image datasets is conducted through the bibliometric analysis. We then propose some principles and discuss the methodology on constructing benchmark datasets. An implementation on creating the RS scene classification dataset demonstrates the practicability of our proposed framework and the experimental results show that our constructed dataset can serve as a promising benchmark for RS image scene interpretation. Yang Long 0002, Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001, DeRen Li |
IGARSS | 2 |
| 2021 | Temporal Relations Matter: A Two-Pathway Network for Aerial Video RecognitionabstractWith the increasing volume of aerial videos, the demand for automatically parsing these videos is surging. To achieve this, current researches mainly focus on extracting a holistic feature with convolutions along both spatial and temporal dimensions. However, these methods are limited by small temporal receptive fields and cannot adequately capture long-term temporal dependencies which are important for describing complicated dynamics. In this paper, we propose a novel two-pathway network to model not only holistic features, but also temporal relations for aerial video classification. More specially, our model employs a two-pathway architecture: (1) a holistic representation pathway to learn a general feature of frame appearances and short-term temporal variations and (2) a temporal relation pathway to capture multi-scale temporal relations across arbitrary frames, providing long-term temporal dependencies. Our model is evaluated on event recognition dataset, ERA, and achieves the state-of-the-art results. This demonstrates its effectiveness and good generalization capacity. Pu Jin, Lichao Mou, Yuansheng Hua, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IGARSS | 4 |
| 2021 | Anomaly Detection in Aerial Videos Via Future Frame Prediction NetworksabstractBy the virtue of high flexibility, low-cost, real-time, and high-resolution data acquisition capacity, unmanned aerial vehicles (UAVs) can be exploited for a wide range of applications, especially in surveillance, inspection, and search fields. Such applications aim to detect potential suspicious events, violent human actions from an untrimmed and lengthy UAV video. Anomaly detection methods are highly in demand because it is unrealistic for human experts to manually detect all abnormal events in image scene. However, anomaly detection methods in aerial videos are rarely studied in the remote sensing community. In this paper, We propose a future frame prediction network based on convolutional variational autoencoder networks to detect anomalous events. Compared to several models, our network has a superior performance. Pu Jin, Lichao Mou, Gui-Song Xia, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2021 | Rotation adaptive correlation filter for moving object tracking in satellite videos
Shiyu Xuan, Shengyang Li, Zifei Zhao, Wanfeng Zhang, Hong Tan, Gui-Song Xia, Yanfeng Gu |
Neurocomputing | 7 |
| 2021 | Gliding Vertex on the Horizontal Bounding Box for Multi-Oriented Object DetectionabstractObject detection has recently experienced substantial progress. Yet, the widely adopted horizontal bounding box representation is not appropriate for ubiquitous oriented objects such as objects in aerial images and scene texts. In this paper, we propose a simple yet effective framework to detect multi-oriented objects. Instead of directly regressing the four vertices, we glide the vertex of the horizontal bounding box on each corresponding side to accurately describe a multi-oriented object. Specifically, We regress four length ratios characterizing the relative gliding offset on each corresponding side. This may facilitate the offset learning and avoid the confusion issue of sequential label points for oriented objects. To further remedy the confusion issue for nearly horizontal objects, we also introduce an obliquity factor based on area ratio between the object and its horizontal bounding box, guiding the selection of horizontal or oriented detection for each object. We add these five extra target variables to the regression head of faster R-CNN, which requires ignorable extra computation time. Extensive experimental results demonstrate that without bells and whistles, the proposed method achieves superior performances on multiple multi-oriented object detection benchmarks including object detection in aerial images, scene text detection, pedestrian detection in fisheye images. Yongchao Xu, Mingtao Fu, Qimeng Wang, Yukang Wang, Kai Chen 0006, Gui-Song Xia, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Learning Regional Attraction for Line Segment DetectionabstractThis paper presents regional attraction of line segment maps, and hereby poses the problem of line segment detection (LSD) as a problem of region coloring. Given a line segment map, the proposed regional attraction first establishes the relationship between line segments and regions in the image lattice. Based on this, the line segment map is equivalently transformed to an attraction field map (AFM), which can be remapped to a set of line segments without loss of information. Accordingly, we develop an end-to-end framework to learn attraction field maps for raw input images, followed by a squeeze module to detect line segments. Apart from existing works, the proposed detector properly handles the local ambiguity and does not rely on the accurate identification of edge pixels. Comprehensive experiments on the Wireframe dataset and the YorkUrban dataset demonstrate the superiority of our method. In particular, we achieve an F-measure of 0.831 on the Wireframe dataset, advancing the state-of-the-art performance by 10.3 percent. Nan Xue 0001, Song Bai 0001, Fudong Wang 0001, Gui-Song Xia, Tianfu Wu 0001, Liangpei Zhang 0001, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Siamese networks with distractor-reduction method for long-term visual object trackingabstractMany trackers which divide the tracking process into two stages have recently been proposed to solve the problem of long-term tracking. Their outstanding performance makes them become one of the mainstream algorithms of long-term tracking. To further improve the performance of two-stage tracking algorithms, some improvements are proposed in this paper. (a) A hard negative mining method is proposed. It can optimize the training process of the verification network and bridge the gap between the two sub-networks. (b) The architecture of the verification network is designed as a Siamese structure; therefore, the semantic ambiguity in classification can be alleviated. Extensive experiments performed on benchmarks demonstrate that the proposed approach significantly outperforms the state-of-the-art methods, yielding 7% relative gain in the VOT2018-LT dataset and 14.2% relative gain in the OxUvA dataset. Shiyu Xuan, Shengyang Li, Zifei Zhao, Longxuan Kou, Gui-Song Xia |
Pattern Recognit. | 6 |
| 2021 | Learning Center Probability Map for Detecting Objects in Aerial ImagesabstractOne fundamental problem in Earth Vision is to accurately find the locations and identify the categories of the interesting objects in the aerial images, for which oriented bounding boxes (OBBs) are usually employed to depict better the objects emerging with arbitrary orientations. However, the regression of the OBBs always suffers from the ambiguous problem in the definition of the regression targets, which often reduces the convergency efficiency and decreases the detection accuracy. Although there are some methods like the binary segmentation map that can handle this problem, it brings a new problem of ambiguous background pixels in the OBBs. In this article, we propose to cast the OBB regression as a center-probability-map (CenterMap)-prediction problem, thus largely eliminating the ambiguities on the target definitions and the background pixels. The predicted CenterMaps are then used to generate the OBBs. The CenterMap OBB representation is simple, yet effective. Furthermore, to distinguish better the interesting objects from the cluttered background, a weighted pseudosegmentation-guided attention network is adopted to provide the object-level features for predicting the horizontal bounding boxes and the OBBs. The experimental results on three widely used data sets, i.e., DOTA, HRSC2016, and UCAS-AOD, demonstrate the effectiveness of our proposed method. Jinwang Wang, Wen Yang 0001, Heng-Chao Li 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Local Semantic Enhanced ConvNet for Aerial Scene RecognitionabstractAerial scene recognition is challenging due to the complicated object distribution and spatial arrangement in a large-scale aerial image. Recent studies attempt to explore the local semantic representation capability of deep learning models, but how to exactly perceive the key local regions remains to be handled. In this paper, we present a local semantic enhanced ConvNet (LSE-Net) for aerial scene recognition, which mimics the human visual perception of key local regions in aerial scenes, in the hope of building a discriminative local semantic representation. Our LSE-Net consists of a context enhanced convolutional feature extractor, a local semantic perception module and a classification layer. Firstly, we design a multi-scale dilated convolution operators to fuse multi-level and multi-scale convolutional features in a trainable manner in order to fully receive the local feature responses in an aerial scene. Then, these features are fed into our two-branch local semantic perception module. In this module, we design a context-aware class peak response (CACPR) measurement to precisely depict the visual impulse of key local regions and the corresponding context information. Also, a spatial attention weight matrix is extracted to describe the importance of each key local region for the aerial scene. Finally, the refined class confidence maps are fed into the classification layer. Exhaustive experiments on three aerial scene classification benchmarks indicate that our LSE-Net achieves the state-of-the-art performance, which validates the effectiveness of our local semantic perception module and CACPR measurement. Qi Bi, Kun Qin, Han Zhang 0052, Gui-Song Xia |
IEEE Trans. Image Process. | 4 |
| 2021 | Conditional Generative ConvNets for Exemplar-Based Texture SynthesisabstractThe goal of exemplar-based texture synthesis is to generate texture images that are visually similar to a given exemplar. Recently, promising results have been reported by methods relying on convolutional neural networks (ConvNets) pretrained on large-scale image datasets. However, these methods have difficulties in synthesizing image textures with non-local structures and extending to dynamic or sound textures. In this article, we present a conditional generative ConvNet (cgCNN) model which combines deep statistics and the probabilistic framework of generative ConvNet (gCNN) model. Given a texture exemplar, cgCNN defines a conditional distribution using deep statistics of a ConvNet, and synthesizes new textures by sampling from the conditional distribution. In contrast to previous deep texture models, the proposed cgCNN does not rely on pre-trained ConvNets but learns the weights of ConvNets for each input exemplar instead. As a result, cgCNN can synthesize high quality dynamic, sound and image textures in a unified manner. We also explore the theoretical connections between our model and other texture models. Further investigations show that the cgCNN model can be easily generalized to texture expansion and inpainting. Extensive experiments demonstrate that our model can achieve better or at least comparable results than the state-of-the-art methods. Meng-Han Li, Gui-Song Xia |
IEEE Trans. Image Process. | 3 |
| 2020 | FGN: Fully Guided Network for Few-Shot Instance SegmentationabstractFew-shot instance segmentation (FSIS) conjoins the few-shot learning paradigm with general instance segmentation, which provides a possible way of tackling instance segmentation in the lack of abundant labeled data for training. This paper presents a Fully Guided Network (FGN) for few-shot instance segmentation. FGN perceives FSIS as a guided model where a so-called support set is encoded and utilized to guide the predictions of a base instance segmentation network (i.e., Mask R-CNN), critical to which is the guidance mechanism. In this view, FGN introduces different guidance mechanisms into the various key components in Mask R-CNN, including Attention-Guided RPN, Relation-Guided Detector, and Attention-Guided FCN, in order to make full use of the guidance effect from the support set and adapt better to the inter-class generalization. Experiments on public datasets demonstrate that our proposed FGN can outperform the state-of-the-art methods. Zhibo Fan, Jin-Gang Yu, Jiarong Ou, Changxin Gao, Gui-Song Xia, Yuanqing Li 0001 |
CVPR | 6 |
| 2020 | Zero-Assignment Constraint for Graph Matching With OutliersabstractGraph matching (GM), as a longstanding problem in computer vision and pattern recognition, still suffers from numerous cluttered outliers in practical applications. To address this issue, we present the zero-assignment constraint (ZAC) for approaching the graph matching problem in the presence of outliers. The underlying idea is to suppress the matchings of outliers by assigning zero-valued vectors to the potential outliers in the obtained optimal correspondence matrix. We provide elaborate theoretical analysis to the problem, i.e., GM with ZAC, and figure out that the GM problem with and without outliers are intrinsically different, which enables us to put forward a sufficient condition to construct valid and reasonable objective function. Consequently, we design an efficient outlier-robust algorithm to significantly reduce the incorrect or redundant matchings caused by numerous outliers. Extensive experiments demonstrate that our method can achieve the state-of-the-art performance in terms of accuracy and efficiency, especially in the presence of numerous outliers. Fudong Wang 0001, Nan Xue 0001, Jin-Gang Yu, Gui-Song Xia |
CVPR | 4 |
| 2020 | Holistically-Attracted Wireframe ParsingabstractThis paper presents a fast and parsimonious parsing method to accurately and robustly detect a vectorized wireframe in an input image with a single forward pass. The proposed method is end-to-end trainable, consisting of three components: (i) line segment and junction proposal generation, (ii) line segment and junction matching, and (iii) line segment and junction verification. For computing line segment proposals, a novel exact dual representation is proposed which exploits a parsimonious geometric reparameterization for line segments and forms a holistic 4-dimensional attraction field map for an input image. Junctions can be treated as the “basins” in the attraction field. The proposed method is thus called Holistically-Attracted Wireframe Parser (HAWP). In experiments, the proposed method is tested on two benchmarks, the Wireframe dataset [14] and the YorkUrban dataset [8]. On both benchmarks, it obtains state-of-the-art performance in terms of accuracy and efficiency. For example, on the Wireframe dataset, compared to the previous state-of-the-art method L-CNN [36], it improves the challenging mean structural average precision (msAP) by a large margin (2.8% absolute improvements), and achieves 29.5 FPS on a single GPU (89% relative improvement). A systematic ablation study is performed to further justify the proposed method. Nan Xue 0001, Tianfu Wu 0001, Song Bai 0001, Fudong Wang 0001, Gui-Song Xia, Liangpei Zhang 0001, Philip Torr 0001 |
CVPR | 5 |
| 2020 | Event Enhanced High-Quality Image Recovery
Bishan Wang, Jingwei He, Lei Yu 0006, Gui-Song Xia, Wen Yang 0001 |
ECCV (13) | 4 |
| 2020 | Tiny Object Detection in Aerial ImagesabstractObject detection in Earth Vision has achieved great progress in recent years. However, tiny object detection in aerial images remains a very challenging problem since the tiny objects contain a small number of pixels and are easily confused with the background. To advance tiny object detection research in aerial images, we present a new dataset for Tiny Object Detection in Aerial Images (AI-TOD). Specifically, AI-TOD comes with 700,621 object instances for eight categories across 28,036 aerial images. Compared to existing object detection datasets in aerial images, the mean size of objects in AI-TOD is about 12.8 pixels, which is much smaller than others. To build a benchmark for tiny object detection in aerial images, we evaluate the state-of-the-art object detectors on our AI-TOD dataset. Experimental results show that direct application of these approaches on AI-TOD produces suboptimal object detection results, thus new specialized detectors for tiny object detection need to be designed. Therefore, we propose a multiple center points based learning network (M-CenterNet) to improve the localization performance of tiny object detection, and experimental results show the significant performance gain over the competitors. Jinwang Wang, Wen Yang 0001, Haowen Guo, Ruixiang Zhang, Gui-Song Xia |
ICPR | 5 |
| 2020 | Look at the Big Picture: Building Area Extraction with Global Density MapabstractThe automatic extraction of building areas from high-resolution satellite imagery has become an important and challenging research issue. Many recent studies have explored different deep learning-based semantic segmentation methods for better accuracy. However, the deep network usually takes sliding window cropped satellite images as inputs, which loses the global information and causes a high false positive rate. In this paper, we propose a density map guided attention mechanism for building area extraction to make the network look at the big picture. We exploit an FCN-based building density prediction network to generate a density heatmap from large satellite images. The density factors in heatmap control the classifier's threshold of building area extraction network that optimize the FP and recall rates. Furthermore, we propose a test-time overlap augmentation mechanism to improve the segmentation results. Our method outperforms state-of-the-art approaches and increases mIoU by about 3.08% to 93.31%, and decreases FP rate to 0.91%. Haowen Guo, Wensheng Cheng, Wen Yang 0001, Gui-Song Xia |
IGARSS | 4 |
| 2020 | Instance Segmentation with Oriented Proposals for Aerial ImagesabstractInstance segmentation is a challenging issue in remote sensing. The existing state-of-the-art methods use horizontal bounding box (HBB) to infer the instance mask of object. However, the objects in the aerial images have the characteristics of being distributed in arbitrary orientation, densely packed, and so on. In this case, the HBB usually contains a lot of background information and several neighboring objects, leading to coarse and inaccurate mask prediction. To solve the aforementioned problems, we propose a new instance segmentation method, ISOP, by inferring the mask on oriented bounding box (OBB) instead of HBB. We show that the proposed method leading to more accurate mask predictions, especially for densely packed objects. We evaluate our method in the iSAID dataset, and compared to the baseline, the ISOP has achieved around 17% improvement in terms of mAP and 11% for densely packed objects. Jian Ding 0001, Jinwang Wang, Wen Yang 0001, Gui-Song Xia |
IGARSS | 5 |
| 2020 | A Functional Representation for Graph MatchingabstractGraph matching is an important and persistent problem in computer vision and pattern recognition for finding node-to-node correspondence between graphs. However, graph matching that incorporates pairwise constraints can be formulated as a quadratic assignment problem (QAP), which is NP-complete and results in intrinsic computational difficulties. This paper presents a functional representation for graph matching (FRGM) that aims to provide more geometric insights on the problem and reduce the space and time complexities. To achieve these goals, we represent each graph by a linear function space equipped with a functional such as inner product or metric, that has an explicit geometric meaning. Consequently, the correspondence matrix between graphs can be represented as a linear representation map. Furthermore, this map can be reformulated as a new parameterization for matching graphs in Euclidean space such that it is consistent with graphs under rigid or nonrigid deformations. This allows us to estimate the correspondence matrix and geometric deformations simultaneously. We use the representation of edge-attributes rather than the affinity matrix to reduce the space complexity and propose an efficient optimization strategy to reduce the time complexity. The experimental results on both synthetic and real-world datasets show that the FRGM can achieve state-of-the-art performance. Fudong Wang 0001, Nan Xue 0001, Yipeng Zhang 0001, Gui-Song Xia, Marcello Pelillo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Special Issue on Big Data From SpaceabstractThe recent multiplication of open access initiatives to Big Data from Space is giving momentum to the field by widening substantially the spectrum of scientific communities and users, as well as awareness among the public, while offering new benefits at all levels from individual citizens to the whole society. Following a detailed and rigorous review process, 14 articles have been selected out of 48 submissions for this special issue. These are briefly summarized. Mihai Datcu, Jacqueline LeMoigne-Stewart, Sveinung Loekken, Pierre Soille, Gui-Song Xia |
IEEE Trans. Big Data | 5 |
| 2020 | Mining Deep Semantic Representations for Scene Classification of High-Resolution Remote Sensing ImageryabstractScene classification is one of the most fundamental task in interpretation of high-resolution remote sensing (HRRS) images. Many recent works show that the probabilistic topic models which are capable of mining latent semantics of images can be effectively applied to HRRS scene classification. However, the existing approaches based on topic models simply utilize low-level hand-crafted features to form semantic features, which severely limit the representative capability of the semantic features derived from topic models. To alleviate this problem, this paper propose to build powerful semantic features using the probabilistic latent semantic analysis (pLSA) model, by employing the pre-trained deep convolutional neural networks (CNNs) as feature extractors rather than relying on the hand-crafted features. Specifically, we develop two methods to generate semantic features, called multi-scale deep semantic representation (MSDS) and multi-level deep semantic representation (MLDS), by extracting CNN features from different layers: (1) in MSDS, the final semantic features are learned by the pLSA with multi-scale features extracted from the convolutional layer of a pre-trained CNN; (2) in MLDS, we extract CNN features for densely sampled image patches at different size level from the fully-connected layer of a pre-trained CNN, and concatenate the sematic features learned by the pLSA at each level. We comprehensively evaluate the two methods on two public HRRS scene datasets, and achieve significant performance improvement over the state-of-the-art. The outstanding results demonstrate that the pLSA model is capable of discovering considerably discriminative semantic features from the deep CNN features. Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001 |
IEEE Trans. Big Data | 2 |
| 2020 | Exploiting Deep Features for Remote Sensing Image Retrieval: A Systematic InvestigationabstractRemote sensing (RS) image retrieval is of great significant for geological information mining. Over the past two decades, a large amount of research on this task has been carried out, which mainly focuses on the following three core issues: feature extraction, similarity metric, and relevance feedback. Due to the complexity and multiformity of ground objects in high-resolution remote sensing (HRRS) images, there is still room for improvement in the current retrieval approaches. In this article, we analyze the three core issues of RS image retrieval and provide a comprehensive review on existing methods. Furthermore, for the goal to advance the state-of-the-art in HRRS image retrieval, we focus on the feature extraction issue and delve how to use powerful deep representations to address this task. We conduct systematic investigation on evaluating correlative factors that may affect the performance of deep features. By optimizing each factor, we acquire remarkable retrieval results on publicly available HRRS datasets. Finally, we explain the experimental phenomenon in detail and draw conclusions according to our analysis. Our work can serve as a guiding role for the research of content-based RS image retrieval. Xin-Yi Tong 0003, Gui-Song Xia, Yanfei Zhong, Mihai Datcu, Liangpei Zhang 0001 |
IEEE Trans. Big Data | 2 |
| 2020 | Mental Retrieval of Remote Sensing Images via Adversarial Sketch-Image Feature LearningabstractSearching the targets of interest in large-scale remote sensing images is a fundamental problem, which becomes a very challenging issue when there is no relevant example at hand but a mental picture in mind. Hand-drawn sketch as a precise and convenient expression of the mental picture makes sketch-based remote sensing image retrieval (SBRSIR) an ideal choice to cope with this issue. However, the accuracy of SBRSIR algorithm, which is critical to effective retrieval, is still far behind classic query-by-example image retrieval. Two central limiting factors for this performance gap are: 1) the lack of effective cross-domain representations for bridging the domain gap between sketches and remote sensing images and 2) the absence of large sketch/remote sensing image data sets for developing, evaluating, and comparing the SBRSIR approaches. In this article, we first develop a novel SBRSIR model to learn a deep joint embedding space with discriminative losses, where adversarial training is used for the embedding space to learn domain-invariant representations. Then, we contribute a sketch/remote sensing image data set specifically for SBRSIR and provide a benchmark for subsequent researchers. Extensive experiments on the data set and large scene images demonstrate the effectiveness and superiority of the method for both the seen and unseen categories. Wen Yang 0001, Tianbi Jiang, Shijie Lin, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | Object Tracking in Satellite Videos by Improved Correlation Filters With Motion EstimationsabstractAs a new method of Earth observation, video satellite is capable of monitoring specific events on the Earth's surface continuously by providing high-temporal resolution remote sensing images. The video observations enable a variety of new satellite applications such as object tracking and road traffic monitoring. In this article, we address the problem of fast object tracking in satellite videos, by developing a novel tracking algorithm based on correlation filters embedded with motion estimations. Based on the kernelized correlation filter (KCF), the proposed algorithm provides the following improvements: 1) proposing a novel motion estimation (ME) algorithm by combining the Kalman filter and motion trajectory averaging and mitigating the boundary effects of KCF by using this ME algorithm and 2) solving the problem of tracking failure when a moving object is partially or completely occluded. The experimental results demonstrate that our algorithm can track the moving object in satellite videos with 95% accuracy. Shiyu Xuan, Shengyang Li, Mingfei Han 0002, Xue Wan, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | A Multiple-Instance Densely-Connected ConvNet for Aerial Scene ClassificationabstractIn contrast with nature scenes, aerial scenes are often composed of many objects crowdedly distributed on the surface in bird's view, the description of which usually demands more discriminative features as well as local semantics. However, when applied to scene classification, most of the existing convolution neural networks (ConvNets) tend to depict global semantics of images, and the loss of low- and mid-level features can hardly be avoided, especially when the model goes deeper. To tackle these challenges, in this paper, we propose a multiple-instance densely-connected ConvNet (MIDC-Net) for aerial scene classification. It regards aerial scene classification as a multiple-instance learning problem so that local semantics can be further investigated. Our classification model consists of an instance-level classifier, a multiple instance pooling and followed by a bag-level classification layer. In the instance-level classifier, we propose a simplified dense connection structure to effectively preserve features from different levels. The extracted convolution features are further converted into instance feature vectors. Then, we propose a trainable attention-based multiple instance pooling. It highlights the local semantics relevant to the scene label and outputs the bag-level probability directly. Finally, with our bag-level classification layer, this multiple instance learning framework is under the direct supervision of bag labels. Experiments on three widely-utilized aerial scene benchmarks demonstrate that our proposed method outperforms many state-of-the-art methods by a large margin with much fewer parameters. Qi Bi, Kun Qin, Zhili Li, Han Zhang 0052, Gui-Song Xia |
IEEE Trans. Image Process. | 6 |
| 2020 | Exemplar-Based Recursive Instance Segmentation With Application to Plant Image AnalysisabstractInstance segmentation is a challenging computer vision problem which lies at the intersection of object detection and semantic segmentation. Motivated by plant image analysis in the context of plant phenotyping, a recently emerging application field of computer vision, this paper presents the Exemplar-Based Recursive Instance Segmentation (ERIS) framework. A three-layer probabilistic model is firstly introduced to jointly represent hypotheses, voting elements, instance labels and their connections. Afterwards, a recursive optimization algorithm is developed to infer the maximum a posteriori (MAP) solution, which handles one instance at a time by alternating among the three steps of detection, segmentation and update. The proposed ERIS framework departs from previous works mainly in two respects. First, it is exemplar-based and model-free, which can achieve instance-level segmentation of a specific object class given only a handful of (typically less than 10) annotated exemplars. Such a merit enables its use in case that no massive manually-labeled data is available for training strong classification models, as required by most existing methods. Second, instead of attempting to infer the solution in a single shot, which suffers from extremely high computational complexity, our recursive optimization strategy allows for reasonably efficient MAP-inference in full hypothesis space. The ERIS framework is substantialized for the specific application of plant leaf segmentation in this work. Experiments are conducted on public benchmarks to demonstrate the superiority of our method in both effectiveness and efficiency in comparison with the state-of-the-art. Jin-Gang Yu, Yansheng Li 0001, Changxin Gao, Hongxia Gao, Gui-Song Xia, Zhu Liang Yu, Yuanqing Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Ultrafast Endoscopic Ultrasonography With Circular ArrayabstractRapid development of ultrafast ultrasound imaging has led to novel medical ultrasound applications, including shear wave elastography and super-resolution vascular imaging. However, these have yet to incorporate endoscopic ultrasonography (EUS) with a circular array, which provides a wider view in the alimentary canal than traditional linear and convex arrays. A coherent diverging wave compounding (CDWC) imaging method was proposed for ultrafast EUS imaging and implemented on a custom circular array. In CDWC, virtual acoustic point sources are allocated and virtually insonified diverging waves from each source are achieved by adjusting all circular array elements' emission time delays. Diverging waves emitted from different virtual sources are coherently compounded, generating synthetic transmit focusing at every location in the image plane. As the field of view of the circular array is centrally symmetric, all virtual sources are equidistantly distributed on a concentric circle of radius r . To achieve the highest frame rate possible with image quality comparable to that obtained with the traditional multi-focus imaging method, the effects of various radii r and virtual source quantities on the compounded image quality were theoretically analyzed and experimentally verified. Simulation, phantom, and ex-vivo experiments were conducted with an 8 MHz, 124-element circular array, with a 5.35 mm radius. When 16 virtual sources were used with r=1.605 mm, image quality comparable to that obtained with the multi-focus approach was achieved at a frame rate of 1000 frames/s. This demonstrates the feasibility of the proposed ultrafast EUS imaging method and promotes further development of multi-functional EUS devices. Qingyuan Tan, Congzhi Wang, Jiamei Liu, Jiqing Huang, Yongchuan Li, Yang Xiao 0012, Gui-Song Xia, Teng Ma 0004, Hairong Zheng |
IEEE Trans. Medical Imaging | 7 |
| 2019 | Learning RoI Transformer for Oriented Object Detection in Aerial ImagesabstractObject detection in aerial images is an active yet challenging task in computer vision because of the bird’s-eye view perspective, the highly complex backgrounds, and the variant appearances of objects. Especially when detecting densely packed objects in aerial images, methods relying on horizontal proposals for common object detection often introduce mismatches between the Region of Interests (RoIs) and objects. This leads to the common misalignment between the final object classification confidence and localization accuracy. In this paper, we propose a RoI Transformer to address these problems. The core idea of RoI Transformer is to apply spatial transformations on RoIs and learn the transformation parameters under the supervision of oriented bounding box (OBB) annotations. RoI Transformer is with lightweight and can be easily embedded into detectors for oriented object detection. Simply apply the RoI Transformer to light head RCNN has achieved state-of-the-art performances on two common and challenging aerial datasets, i.e., DOTA and HRSC2016, with a neglectable reduction to detection speed. Our RoI Transformer exceeds the deformable Position Sensitive RoI pooling when oriented bounding-box annotations are available. Extensive experiments have also validated the flexibility and effectiveness of our RoI Transformer. Jian Ding 0001, Nan Xue 0001, Yang Long 0002, Gui-Song Xia, Qikai Lu |
CVPR | 4 |
| 2019 | Learning Attraction Field Representation for Robust Line Segment DetectionabstractThis paper presents a region-partition based attraction field dual representation for line segment maps, and thus poses the problem of line segment detection (LSD) as the region coloring problem. The latter is then addressed by learning deep convolutional neural networks (ConvNets) for accuracy, robustness and efficiency. For a 2D line segment map, our dual representation consists of three components: (i) A region-partition map in which every pixel is assigned to one and only one line segment; (ii) An attraction field map in which every pixel in a partition region is encoded by its 2D projection vector w.r.t. the associated line segment; and (iii) A squeeze module which squashes the attraction field to a line segment map that almost perfectly recovers the input one. By leveraging the duality, we learn ConvNets to compute the attraction field maps for raw in-put images, followed by the squeeze module for LSD, in an end-to-end manner. Our method rigorously addresses several challenges in LSD such as local ambiguity and class imbalance. Our method also harnesses the best practices developed in ConvNets based semantic segmentation methods such as the encoder-decoder architecture and the a-trous convolution. In experiments, our method is tested on the WireFrame dataset and the YorkUrban dataset with state-of-the-art performance obtained. Especially, we advance the performance by 4.5 percents on the WireFramedataset. Our method is also fast with 6.6∼10.4 FPS, outperforming most of existing line segment detectors. Nan Xue 0001, Song Bai 0001, Fudong Wang 0001, Gui-Song Xia, Tianfu Wu 0001, Liangpei Zhang 0001 |
CVPR | 4 |
| 2019 | Learning to Calibrate Straight Lines for Fisheye Image RectificationabstractThis paper presents a new deep-learning based method to simultaneously calibrate the intrinsic parameters of fisheye lens and rectify the distorted images. Assuming that the distorted lines generated by fisheye projection should be straight after rectification, we propose a novel deep neural network to impose explicit geometry constraints onto processes of the fisheye lens calibration and the distorted image rectification. In addition, considering the nonlinearity of distortion distribution in fisheye images, the proposed network fully exploits multi-scale perception to equalize the rectification effects on the whole image. To train and evaluate the proposed model, we also create a new large-scale dataset labeled with corresponding distortion parameters and well-annotated distorted lines. Compared with the state-of-the-art methods, our model achieves the best published rectification quality and the most accurate estimation of distortion parameters on a large set of synthetic and real fisheye images. Zhucun Xue, Nan Xue 0001, Gui-Song Xia, Weiming Shen 0002 |
CVPR | 3 |
| 2019 | Multi-Level Fusion of the Multi-Receptive Fields Contextual Networks and Disparity Network for Pairwise Semantic StereoabstractIn this paper, we propose a multi-level fusion framework to address the pairwise semantic stereo issue. For disparity estimation, we adopt the pyramid stereo matching network. For semantic segmentation, the single segmentation network is proposed with respect to the left image, along with the disparity fusion segmentation network for the combination of semantic features and disparity features. Specifically, the multi-receptive fusion block is designed and employed to fully extract and fuse the contextual information. Finally, the refined segmentation result is obtained via yet another fusion of the multi-model results. The proposed method achieved a mean intersection over union (mIoU) of 79.05%, an average endpoint error (EPE) of 1.3966, and an mIoU-3 of 77.75%, ranking first in the Pairwise Semantic Stereo Challenge of the 2019 IEEE GRSS Data Fusion Contest [1],[2]. Hongyu Chen 0003, Manhui Lin, Hongyan Zhang 0001, Gui-Song Xia, Xianwei Zheng, Liangpei Zhang 0001 |
IGARSS | 5 |
| 2019 | Mental Retrieval of Large-Scale Satellite Images Via Learned Sketch-Image Deep FeaturesabstractSearching targets of interest in large-scale satellite images is an imperative task, which becomes a challenging issue when the targets reside only in the mind of the user as a set of subjective visual patterns. In this paper, we take the advantage of hand-drawn sketches' strong intuition of describing mental target to address the problem of no available exemplar query. We introduce a multi-level-of-detail model to learn a cross-domain representation for bridging the gap between sketches and satellite images. To train the model, we propose a novel method of generating satellite images with corresponding level of details based on generative adversarial network. Experiments on both large-scale satellite images and commonly used RS datasets demonstrate the effectiveness and superiority of our method. Ruixiang Zhang, Wen Yang 0001, Gui-Song Xia |
IGARSS | 4 |
| 2019 | GeoSay: A geometric saliency for extracting buildings in remote sensing images
Gui-Song Xia, Nan Xue 0001, Qikai Lu, Xiao Xiang Zhu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2019 | Robust visible-infrared image matching by exploiting dominant edge orientations
Nan Xue 0001, Yipeng Zhang 0001, Qikai Lu, Gui-Song Xia |
Pattern Recognit. Lett. | 5 |
| 2019 | Image Caption Generation with Part of Speech Guidance
Xinwei He 0001, Baoguang Shi, Xiang Bai, Gui-Song Xia, Zhaoxiang Zhang 0001, Weisheng Dong |
Pattern Recognit. Lett. | 4 |
| 2019 | Editorial
Junchi Yan, Minsu Cho, Francesc Serratosa, Gui-Song Xia, Yinqiang Zheng |
Pattern Recognit. Lett. | 4 |
| 2019 | Learning the Synthesizability of Dynamic Texture SamplesabstractExemplar-based dynamic texture synthesis (EDTS) is targeted to generate new samples of high quality that are perceptually similar to a given input dynamic texture exemplar. This paper addresses the issue of learning the synthesizability of dynamic texture samples. Given a dynamic texture sample, how is its possibility of being synthesized by EDTS methods estimated, and what is the most suitable EDTS algorithm to complete the task? To this end, we propose associating dynamic texture samples with synthesizability scores by learning regression models on a compiled dynamic texture dataset annotated in terms of synthesizability. More precisely, we first define the synthesizability of DT samples and characterize them by a set of spatiotemporal features. We then train regression models on the annotated dataset with feature representation to predict the synthesizability scores of the DT samples and learn classifiers to select the most suitable EDTS algorithm. We further complete the selection, partition and synthesizability prediction of the DT samples in a hierarchical scheme. The learned synthesizability is finally applied to detecting synthesizable regions in videos. Both quantitative and qualitative experiments demonstrate that our method can efficiently learn and predict the synthesizability of DT samples. Feng Yang 0015, Gui-Song Xia, Dengxin Dai, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Rotation-Sensitive Regression for Oriented Scene Text DetectionabstractText in natural images is of arbitrary orientations, requiring detection in terms of oriented bounding boxes. Normally, a multi-oriented text detector often involves two key tasks: 1) text presence detection, which is a classification problem disregarding text orientation; 2) oriented bounding box regression, which concerns about text orientation. Previous methods rely on shared features for both tasks, resulting in degraded performance due to the incompatibility of the two tasks. To address this issue, we propose to perform classification and regression on features of different characteristics, extracted by two network branches of different designs. Concretely, the regression branch extracts rotation-sensitive features by actively rotating the convolutional filters, while the classification branch extracts rotation-invariant features by pooling the rotation-sensitive features. The proposed method named Rotation-sensitive Regression Detector (RRD) achieves state-of-the-art performance on several oriented scene text benchmark datasets, including ICDAR 2015, MSRA-TD500, RCTW-17, and COCO-Text. Furthermore, RRD achieves a significant improvement on a ship collection dataset, demonstrating its generality on oriented object detection. Minghui Liao, Zhen Zhu 0006, Baoguang Shi, Gui-Song Xia, Xiang Bai |
CVPR | 4 |
| 2018 | DOTA: A Large-Scale Dataset for Object Detection in Aerial ImagesabstractObject detection is an important and challenging problem in computer vision. Although the past decade has witnessed major advances in object detection in natural scenes, such successes have been slow to aerial imagery, not only because of the huge variation in the scale, orientation and shape of the object instances on the earth's surface, but also due to the scarcity of well-annotated datasets of objects in aerial scenes. To advance object detection research in Earth Vision, also known as Earth Observation and Remote Sensing, we introduce a large-scale Dataset for Object deTection in Aerial images (DOTA). To this end, we collect 2806 aerial images from different sensors and platforms. Each image is of the size about 4000 × 4000 pixels and contains objects exhibiting a wide variety of scales, orientations, and shapes. These DOTA images are then annotated by experts in aerial image interpretation using 15 common object categories. The fully annotated DOTA images contains 188, 282 instances, each of which is labeled by an arbitrary (8 d.o.f.) quadrilateral. To build a baseline for object detection in Earth Vision, we evaluate state-of-the-art object detection algorithms on DOTA. Experiments demonstrate that DOTA well represents real Earth Vision applications and are quite challenging. Gui-Song Xia, Xiang Bai, Jian Ding 0001, Zhen Zhu 0006, Serge J. Belongie, Jiebo Luo 0001, Mihai Datcu, Marcello Pelillo, Liangpei Zhang 0001 |
CVPR | 1 |
| 2018 | Adaptively Transforming Graph Matching
Fudong Wang 0001, Nan Xue 0001, Yipeng Zhang 0001, Xiang Bai, Gui-Song Xia |
ECCV (16) | 5 |
| 2018 | ICPR2018 Contest on Object Detection in Aerial Images (ODAI-18)abstractObject detection in aerial images plays a significant role in intelligent interpretation of aerial images. Hence many effective methods, especially the new-generation data-driven methods, have been developed for this task. Here, we hold the ODAI, a new contest that focused on object detection in aerial images, based on a new large-scale aerial image dataset called DOTA [1]. This contest contains over 3000 large-size images ( 4k×4k pixels), which cover 211,581 instances divided into 15 categories. Each instance is labeled by an arbitrary (8 d.o.f.) quadrilateral. Besides, we propose two tasks for this contest, named object detection with the horizontal bounding box (OD-HBB) and object detection with the oriented bounding box (OD-OBB). The contest was opened on February 7, 2018, and ended on April 30, 2018. A website is open to the public, which provides links to download data and evaluation server. We have totally received 60 registrations. There are 8 teams that have successfully submitted results on the OD-HBB task with the top mAP as 0.719, and 9 teams that have successfully submitted results on the OD-OBB task with the top mAP as 0.705. Through the contest, we hope to draw extensive attention from a wide range of communities and call for more future research and efforts for the task of object detection in aerial images. Jian Ding 0001, Zhen Zhu 0006, Gui-Song Xia, Xiang Bai, Serge J. Belongie, Jiebo Luo 0001, Mihai Datcu, Marcello Pelillo, Liangpei Zhang 0001 |
ICPR | 3 |
| 2018 | Delving into the Synthesizability of Dynamic Texture SamplesabstractThe example-based dynamic texture synthesis (EDTS) methods have emerged in multitude, dedicated to generating new dynamic textures (DTs) of high quality from an input exemplar. The problem of EDTS has been studied for several decades, but none of the existing synthesis methods are able to tackle all kinds of dynamic textures equally well. Rather than focus on new synthesis methods, we turn to another way to help EDTS by investigating dynamic texture synthesizability - how synthesizable a specific dynamic texture sample is by EDTS. We propose to predict synthesizability score of a given dynamic texture sample, and suggest which EDTS method is best suited to synthesize it. To this end, we compiled a dynamic texture dataset and annotated each DT in terms of synthesizability. We address the problem of learning dynamic texture synthesizability by using regression model to train a predictor on the data collection. More precisely, we first characterize DT samples by a set of spatiotemporal features. Then, based on dynamic texture descriptors, we train regression models to estimate synthesizability scores and use an additional classifier to choose the optimal EDTS methods. The experiments demonstrate that our method can predict the synthesizability of DT samples effectively. Feng Yang 0015, Gui-Song Xia, Dengxin Dai, Liangpei Zhang 0001 |
ICPR | 2 |
| 2018 | Recent Advances and Opportunities in Scene Classification of Aerial Images with Deep ModelsabstractScene classification is a fundamental task in interpretation of remote sensing images, and has become an active research topic in remote sensing community due to its important role in a wide range of applications. Over the past years, tremendous efforts have been made for developing powerful approaches for scene classification of remote sensing images, evolving from the traditional bag-of-visual-words model to the new generation deep convolutional neural networks (CNNs). The deep CNN based methods have exhibited remarkable breakthrough on performance, dramatically outperforming previous methods which strongly rely on hand-crafted features. However, performance with deep CNNs has gradually plateaued on existing public scene datasets, due to the notable drawbacks of these datasets, such as the small scale and low-diversity of training samples. Therefore, to promote the development of new methods and move the scene classification task a step further, we deeply discuss the existing problems in scene classification task, and accordingly present three open directions. We believe these potential directions will be instructive for the researchers in this field. Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2018 | Accurate Building Detection in VHR Remote Sensing Images Using Geometric SaliencyabstractThis paper aims to address the problem of detecting buildings from remote sensing images with very high resolution (VHR). Inspired by the observation that buildings are always more distinguishable in geometries than in texture or spectral, we propose a new geometric building index (GBI) for accurate building detection, which relies on the geometric saliency of building structures. The geometric saliency of buildings is derived from a mid-level geometric representations based on meaningful junctions that can locally describe anisotropic geometrical structures of images. The resulting GBI is measured by integrating the derived geometric saliency of buildings. Experiments on three public datasets demonstrate that the proposed GBI achieves very promising performance, and meanwhile shows impressive generalization capability. Gui-Song Xia, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2018 | AID++: An Updated Version of AID on Scene ClassificationabstractAerial image scene classification is a fundamental problem for understanding high-resolution remote sensing images and has become an active research task in the field of remote sensing due to its important role in a wide range of applications. However, the limitations of existing datasets for scene classification, such as the small scale and low-diversity, severely hamper the potential usage of the new generation deep convolutional neural networks (CNNs). Although huge efforts have been made in building large-scale datasets very recently, e.g., the Aerial Image Dataset (AID) which contains 10,000 image samples, they are still far from sufficient to fully train a high-capacity deep CNN model. To this end, we present a larger-scale dataset in this paper, named as AID++, for aerial scene classification based on the AID dataset. The proposed AID++ consists of more than 400,000 image samples that are semi-automatically annotated by using the existing the geographical data. We evaluate several prevalent CNN models on the proposed dataset, and the results show that our dataset can be used as a promising benchmark for scene classification. Pu Jin, Gui-Song Xia, Qikai Lu, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2018 | Large-Scale Land Cover Classification in Gaofen-2 Satellite ImageryabstractMany significant applications need land cover information of remote sensing images that are acquired from different areas and times, such as change detection and disaster monitoring. However, it is difficult to find a generic land cover classification scheme for different remote sensing images due to the spectral shift caused by diverse acquisition condition. In this paper, we develop a novel land cover classification method that can deal with large-scale data captured from widely distributed areas and different times. Additionally, we establish a large-scale land cover classification dataset consisting of 150 Gaofen-2 imageries as data support for model training and performance evaluation. Our experiments achieve outstanding classification accuracy compared with traditional methods. Xin-Yi Tong 0003, Qikai Lu, Gui-Song Xia, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2018 | Image Stitching Using Smoothly Planar Homography
Tian-Zhu Xiang, Gui-Song Xia, Liangpei Zhang 0001 |
PRCV (1) | 2 |
| 2018 | Image stitching by line-guided local warping with global similarity constraint
Tian-Zhu Xiang, Gui-Song Xia, Xiang Bai, Liangpei Zhang 0001 |
Pattern Recognit. | 2 |
| 2018 | Anisotropic-Scale Junction Detection and Matching for Indoor ImagesabstractJunctions play an important role in characterizing local geometrical structures of images, and the detection of which is a longstanding but challenging task. Existing junction detectors usually focus on identifying the location and orientations of junction branches while ignoring their scales, which, however, contain rich geometries of images. This paper presents a novel approach for junction detection and characterization, which especially exploits the locally anisotropic geometries of a junction and estimates its scales by relying on an a-contrario model. The output junctions are with anisotropic scales, saying that a scale parameter is associated with each branch of a junction and are thus named as anisotropic-scale junctions (ASJs). We then apply the new detected ASJs for matching indoor images, where there are dramatic changes of viewpoints and the detected local visual features, e.g., key-points, are usually insufficient and lack distinctive ability. We propose to use the anisotropic geometries of our junctions to improve the matching precision of indoor images. The matching results on sets of indoor images demonstrate that our approach achieves the state-of-the-art performance on indoor image matching. Nan Xue 0001, Gui-Song Xia, Xiang Bai, Liangpei Zhang 0001, Weiming Shen 0002 |
IEEE Trans. Image Process. | 2 |
| 2017 | Sketch-based aerial image retrievalabstractNotwithstanding aerial image retrieval is an important and obligatory task, existing retrieval systems lose their efficiency when there is no available aerial image used as the exemplar query. In this paper, we take free-hand sketches into consideration and address the problem of sketch-based aerial image retrieval. This is an extremely challenging task due to the complex surface structures and huge variations of resolutions of aerial images, and few works have been devoted to it. For the first time to our knowledge, we propose a framework to bridge the gap between sketches and aerial images. Specifically, an aerial sketch-image dataset is first collected. Sketches and aerial images are augmented to varied levels of details and used to train a multi-scale deep hierarchical model. The fully-connected layers of the deep model are used as cross-domain features, and the similarity between aerial images and sketches is measured by the Euclidean distance. Experiments on several public aerial image datasets demonstrate the efficiency and superiority of the proposed method. Tianbi Jiang, Gui-Song Xia, Qikai Lu |
ICIP | 2 |
| 2017 | Retrieving Aerial Scene Images with Learned Deep Image-Sketch Features
Tianbi Jiang, Gui-Song Xia, Qikai Lu, Weiming Shen 0002 |
J. Comput. Sci. Technol. | 2 |
| 2017 | AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene ClassificationabstractAerial scene classification, which aims to automatically label an aerial image with a specific semantic category, is a fundamental problem for understanding high-resolution remote sensing imagery. In recent years, it has become an active task in the remote sensing area, and numerous algorithms have been proposed for this task, including many machine learning and data-driven approaches. However, the existing data sets for aerial scene classification, such as UC-Merced data set and WHU-RS19, contain relatively small sizes, and the results on them are already saturated. This largely limits the development of scene classification algorithms. This paper describes the Aerial Image data set (AID): a large-scale data set for aerial scene classification. The goal of AID is to advance the state of the arts in scene classification of remote sensing images. For creating AID, we collect and annotate more than 10000 aerial scene images. In addition, a comprehensive review of the existing aerial scene classification techniques as well as recent widely used deep learning methods is given. Finally, we provide a performance analysis of typical aerial scene classification and deep learning approaches on AID, which can be served as the baseline results on this benchmark. Gui-Song Xia, Jingwen Hu 0001, Baoguang Shi, Xiang Bai, Yanfei Zhong, Liangpei Zhang 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Unsupervised Classification of Polarimetric SAR Images via Riemannian Sparse CodingabstractUnsupervised classification plays an important role in understanding polarimetric synthetic aperture radar (PolSAR) images. One of the typical representations of PolSAR data is in the form of Hermitian positive definite (HPD) covariance matrices. Most algorithms for unsupervised classification using this representation either use statistical distribution models or adopt polarimetric target decompositions. In this paper, we propose an unsupervised classification method by introducing a sparsity-based similarity measure on HPD matrices. Specifically, we first use a novel Riemannian sparse coding scheme for representing each HPD covariance matrix as sparse linear combinations of other HPD matrices, where the sparse reconstruction loss is defined by the Riemannian geodesic distance between HPD matrices. The coefficient vectors generated by this step reflect the neighborhood structure of HPD matrices embedded in the Euclidean space and hence can be used to define a similarity measure. We apply the scheme for PolSAR data, in which we first oversegment the images into superpixels, followed by representing each superpixel by an HPD matrix. These HPD matrices are then sparse coded, and the resulting sparse coefficient vectors are then clustered by spectral clustering using the neighborhood matrix generated by our similarity measure. The experimental results on different fully PolSAR images demonstrate the superior performance of the proposed classification approach against the state-of-the-art approaches. Neng Zhong, Wen Yang 0001, Anoop Cherian, Xiangli Yang, Gui-Song Xia, Mingsheng Liao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Texture Characterization Using Shape Co-Occurrence PatternsabstractTexture characterization is a key problem in image understanding and pattern recognition. In this paper, we present a flexible shape-based texture representation using shape co-occurrence patterns. More precisely, texture images are first represented by a tree of shapes, each of which is associated with several geometrical and radiometric attributes. Then, four typical kinds of shape co-occurrence patterns based on the hierarchical relationships among the shapes in the tree are learned as codewords. Three different coding methods are investigated for learning the codewords, which can be used to encode any given texture image into a descriptive vector. In contrast with existing works, the proposed approach not only inherits the shape-based method's strong ability to capture geometrical aspects of textures and high robustness to variations in imaging conditions but also provides a flexible way to consider shape relationships and to compute high-order statistics on the tree. To the best of our knowledge, this is the first time that co-occurrence patterns of explicit shapes have been used as a tool for texture analysis. Experiments on various texture and scene data sets demonstrate the efficiency of the proposed approach. Gui-Song Xia, Gang Liu 0013, Xiang Bai, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Texture synthesis through convolutional neural networks and spectrum constraintsabstractThis paper presents a significant improvement for the synthesis of texture images using convolutional neural networks (CNNs), making use of constraints on the Fourier spectrum of the results. More precisely, the texture synthesis is regarded as a constrained optimization problem, with constraints conditioning both the Fourier spectrum and statistical features learned by CNNs. In contrast with existing methods, the presented method inherits from previous CNN approaches the ability to depict local structures and fine scale details, and at the same time yields coherent large scale structures, even in the case of quasi-periodic images. This is done at no extra computational cost. Synthesis experiments on various images show a clear improvement compared to a recent state-of-the art method relying on CNN constraints only. Gang Liu 0013, Yann Gousseau, Gui-Song Xia |
ICPR | 3 |
| 2016 | Locally warping-based image stitching by imposing line constraintsabstractWarping-based image stitching methods often suffer from perspective variations among multiple images and lead to shape and perspective distortions in stitching results. Moreover, they also quickly lose their efficiency in low-textured images, due to the lack of reliable point correspondences. To solve these problems, this paper presents a locally warping-based image stitching by imposing line constraints. First, a two-stage alignment scheme with line constraints is introduced to achieve accurate alignment. More precisely, line features are adopted as alignment constraints to jointly estimate local homographies with point correspondences, which provides strong correspondences especially in low-textured cases. Then line constraints are also imposed to the content-preserving warping framework to further reduce alignment errors and preserve image structures. Second, in order to preserve shape and perspective information, a global similarity transform is introduced to mitigate projective distortions. Experimental results demonstrate the efficiency of our method, which yields more encouraging image stitching results in contrast with state-of-the-art methods. Tian-Zhu Xiang, Gui-Song Xia, Liangpei Zhang 0001, NingNing Huang |
ICPR | 2 |
| 2016 | Mining the spatial distribution of visual words for scene classificationabstractIn the past decades, tremendous investigations have been made to classify high-spatial-resolution remote sensing (HSR-RS) images at scene level. Among them, Bag-of-Visual-Words (BoVW) model has been widely used thanks to its robustness and efficiency. However, such a representation leaves out the spatial information of the image which is very important for distinguishing various scenes. In this paper, we aim to mine the spatial distribution of the visual words in the BoVW methods so as to incorporate the spatial information of the HSR-RS image and improve the classification accuracy. More precisely, we start from a BoVW representation of each scene image, and then compute local spatial features from this representation, i.e. the encoded image with BoVW model1. The marginal distributions of these local spatial features are finally used to describe HSR-RS scene images. In particular, the local spatial features we used in this paper include the local binary pattern (LBP) and re-learned BoVW dictionaries. The method has been evaluated on a large-scale HSR-RS image dataset, i.e. WHU20, that consists of 5000 HSR-RS images with 20 semantic classes for scene classification. The experimental results show that our method can improve the classification accuracy a lot compared with the standard BoVW method. Jingwen Hu 0001, Gui-Song Xia, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2016 | Accurate object tracking by combining correlation filters and keypointsabstractObject tracking usually suffers from the geometrical deformations and occlusions of objects. This paper presents a new method for accurate object tracking by combining the multi-angle discriminative correlation filters and key-points under the framework of Discriminative Scale Space Tracker (DSST) tracker. Experimental results demonstrate that the proposed method can produce promising tracking results and outperform the-state-of-the-art methods using correlation filters. Zifeng Wang 0007, Gui-Song Xia, Liangpei Zhang 0001 |
IJCNN | 3 |
| 2016 | Multi-object tracking with inter-feedback between detection and tracking
Shu Tian, Fei Yuan 0003, Gui-Song Xia |
Neurocomputing | 3 |
| 2016 | Dynamic texture recognition by aggregating spatial and temporal features via ensemble SVMs
Feng Yang 0015, Gui-Song Xia, Gang Liu 0013, Liangpei Zhang 0001, Xin Huang 0002 |
Neurocomputing | 2 |
| 2016 | Bag-of-Visual-Words Scene Classifier With Local and Global Features for High Spatial Resolution Remote Sensing ImageryabstractScene classification has been studied to allow us to semantically interpret high spatial resolution (HSR) remote sensing imagery. The bag-of-visual-words (BOVW) model is an effective method for HSR image scene classification. However, the traditional BOVW model only captures the local patterns of images by utilizing local features. In this letter, a local-global feature bag-of-visual-words scene classifier (LGFBOVW) is proposed for HSR imagery. In LGFBOVW, the shape-based invariant texture index is designed as the global texture feature, the mean and standard deviation values are employed as the local spectral feature, and the dense scale-invariant feature transform (SIFT) feature is employed as the structural feature. The LGFBOVW can effectively combine the local and global features by an appropriate feature fusion strategy at histogram level. Experimental results on UC Merced and Google data sets of SIRI-WHU demonstrate that the proposed method outperforms the state-of-the-art scene classification methods for HSR imagery. Qiqi Zhu, Yanfei Zhong, Gui-Song Xia, Liangpei Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2016 | Globally consistent correspondence of multiple feature sets using proximal Gauss-Seidel relaxation
Jin-Gang Yu, Gui-Song Xia, Ashok Samal, Jinwen Tian |
Pattern Recognit. | 2 |
| 2016 | Iterative Time-Frequency Filtering of Sinusoidal Signals With Updated Frequency EstimationabstractIn this letter, a sinusoidal time-frequency distribution based filtering (STFD-F) algorithm is proposed and analysed for estimating mono-component stationary sinusoidal signals embedded in strong noise. An initial frequency estimation of the sinusoidal signal is required in the STFD-F algorithm. We theoretically derive the closed-form expressions of the variance and the bias of the estimated signal using the STFD-F, and show that the performance of the STFD-F is dependent on the frequency estimation accuracy, which can be gradually refined by performing an iterative STFD-F procedure. Computer simulations on synthetic sinusoidal signals are presented to corroborate the theoretical analysis. Lei Yu 0006, Gui-Song Xia |
IEEE Signal Process. Lett. | 3 |
| 2016 | Meaningful Object Segmentation From SAR Images via a Multiscale Nonlocal Active Contour ModelabstractThe segmentation of synthetic aperture radar (SAR) images is a long-standing yet challenging task, not only because of the presence of speckle but also due to the variations of surface backscattering properties in the images. Tremendous investigations have been made to suppress the speckle effects for the segmentation of SAR images, whereas few works are devoted to dealing with the variations of backscattering intensities in the images. To overcome the two difficulties, this paper presents a novel SAR image segmentation method by exploiting a multiscale active contour model based on the nonlocal processing principle. More precisely, we first formulize the SAR segmentation problem with an active contour model by integrating the nonlocal interactions between pairs of patches inside and outside the segmented regions. Second, a multiscale strategy is proposed to speed up the nonlocal active contour segmentation procedure and to avoid falling into a local minimum for achieving more accurate segmentation results. Experimental results on simulated and real SAR images demonstrate the efficiency and feasibility of the proposed method: It can not only achieve precise segmentations for images with heavy speckle and nonlocal intensity variations but also be used for SAR images from different types of sensors. Gui-Song Xia, Gang Liu 0013, Wen Yang 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Region-Based Change Detection for Polarimetric SAR Images Using Wishart Mixture ModelsabstractThe change detection of polarimetric synthetic aperture radar (PolSAR) images is a longstanding and challenging task, not only because of the speckle issue but also due to the complex texture, which generally appears highly heterogeneous. There are two widely used approaches for the change detection of PolSAR images: one is the post classification comparison algorithm, and the other is the directly unsupervised change detection algorithm. In this paper, we focus on the latter and propose a region-based change detection method for PolSAR images by means of Wishart mixture models (WMMs). The WMMs fit the distribution of PolSAR images with less errors both in the homogeneous and the extremely heterogeneous area. More precisely, two PolSAR images are first segmented into compact local regions using the customized simple-linear-iterative-clustering algorithm, while the WMMs are used to model each local region. To generate a difference map, statistical distribution differences measured by information theoretic divergence are then computed for corresponding local region pairs. The Cauchy-Schwarz divergence is adopted as its analytic expression can be derived for WMMs. Finally, the change detection results are obtained by the Kittler-Illingworth thresholding method with Markov random field-based smoothing. The proposed scheme is tested on different PolSAR data sets. Qualitative and quantitative evaluations show its superior performance comparing to the traditional pixel-level approach. Wen Yang 0001, Xiangli Yang, Tianheng Yan, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2016 | Dirichlet-Derived Multiple Topic Scene Classification Model for High Spatial Resolution Remote Sensing ImageryabstractDue to the complex arrangements of the ground objects in high spatial resolution (HSR) imagery scenes, HSR imagery scene classification is a challenging task, which is aimed at bridging the semantic gap between the low-level features and the high-level semantic concepts. A combination of multiple complementary features for HSR imagery scene classification is considered a potential way to improve the performance. However, the different types of features have different characteristics, and how to fuse the different types of features is a classic problem. In this paper, a Dirichlet-derived multiple topic model (DMTM) is proposed to fuse heterogeneous features at a topic level for HSR imagery scene classification. An efficient algorithm based on a variational expectation-maximization framework is developed to infer the DMTM and estimate the parameters of the DMTM. The proposed DMTM scene classification method is able to incorporate different types of features with different characteristics, no matter whether these features are local or global, discrete or continuous. Meanwhile, the proposed DMTM can also reduce the dimension of the features representing the HSR images. In our experiments, three types of heterogeneous features, i.e., the local spectral feature, the local structural feature, and the global textural feature, were employed. The experimental results with three different HSR imagery data sets show that the three types of features are complementary. In addition, the proposed DMTM is able to reduce the dimension of the features representing the HSR images, to fuse the different types of features efficiently, and to improve the performance of the scene classification over that of other scene classification algorithms based on spatial pyramid matching, probabilistic latent semantic analysis, and latent Dirichlet allocation. Yanfei Zhong, Gui-Song Xia, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | A Computational Model for Object-Based Visual Saliency: Spreading Attention Along Gestalt CuesabstractThe past few years have witnessed impressive progress on the research of salient object detection. Nevertheless , existing approaches still cannot perform satisfactorily in the case of complex scenes, particularly when the salient objects have non- uniform appearance or complicated shapes, and the background is complexly structured. One important reason for such limitations may be that these approaches commonly ignore the factor of perceptual grouping in saliency modeling. To address this issue, this paper presents a novel computational model for object -based visual saliency, which explicitly takes into consideration the connections between attention and perceptual grouping, and incorporates Gestalt grouping cues into saliency computation. Inspired by the sensory enhancement theory, we suggest a paradigm for object-based saliency modeling, that is, object-based saliency stems from spreading attention along Gestalt grouping cues. Computationally , three typical Gestalt cues, including proximity, similarity, and closure, are respectively extracted from the given image, which are then integrated by constructing a unified Gestalt graph. A new algorithm named personalized power iteration clustering is developed to effectively fulfill the spreading of attention information across the Gestalt graph. Intensive experiments have been carried out to demonstrate the superior performance of the proposed model in comparison to the state-of-the-art. Jin-Gang Yu, Gui-Song Xia, Changxin Gao, Ashok Samal |
IEEE Trans. Multim. | 2 |
| 2015 | A novel polarimetric-texture-structure descriptor for high-resolution PolSAR image classificationabstractA novel Polarimetric-Texture-Structure descriptor for high-resolution PolSAR image is presented in this paper. More precisely, a PolSAR image is represented by a tree of shapes, each of which is associated with several polarimetric and texture attributes. We first extract the texture properties and polarimetric characteristics from each shape, then use the shape co-occurrence patterns (SCOPs) to characterize the shape relationships, and finally use the resulting SCOPs distributions as features for PolSAR image classification. The proposed method not only has the strong ability to depict the texture and polarimetric properties, but also encodes the shape relationships on the tree. We compare the proposed method with the cluster based statistical feature (CSF) and the scattering mechanism based statistical feature (SMSF). Experimental results on high-resolution PolSAR sample dataset and a large scene for classification demonstrate the effectiveness of the proposed method. Yu Bai 0007, Wen Yang 0001, Gui-Song Xia, Mingsheng Liao |
IGARSS | 3 |
| 2015 | A benchmark for scene classification of high spatial resolution remote sensing imageryabstractScene classification for high-resolution remotely sensed imagery have been widely investigated in recent years. However, there is few public, widely accepted and large scale dataset for benchmarking different methods. This paper presents a new and large dataset consisting of 5000 high-resolution remote sensing images which is manually labeled in 20 semantic classes for scene classification. Each class includes more than 200 image samples with different appearances. Some classic classification algorithms are compared on this dataset. To our knowledge, this work is the first time to give a public benchmark dataset at this size on the problem of scene classification in high-resolution remote sensing imagery, and give comparative results and analysis of various classic classification algorithms. Jingwen Hu 0001, Tianbi Jiang, Xin-Yi Tong 0003, Gui-Song Xia, Liangpei Zhang 0001 |
IGARSS | 4 |
| 2015 | Fast binary coding for satellite image scene classificationabstractFeature extraction is at the core of satellite scene classification task. In this paper, we propose a fast binary coding (FBC) method to effectively generate the global discriminative feature representation of image scenes. Equipped with unsupervised feature learning technique, we first learn a set of optimal “filters” from large quantities of randomly sampled image patches, and then we obtain feature maps by convolving image scene with the learned filter bank. After binarizing the feature maps, a simple skillful conversion of binary-valued feature map to integer-valued feature map is performed. The final statistical histograms, which are considered as the global feature representations of scenes, are computed on the integer-valued feature map similar to the conventional BOW model. Experiments on two datasets demonstrate that the proposed FBC achieve satisfying classification performance as well as has much faster computational speed compared with traditional scene classification methods. Zifeng Wang 0007, Gui-Song Xia, Bin Luo 0005, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2015 | A comparative study of sampling analysis in scene classification of high-resolution remote sensing imageryabstractScene classification is a key problem in the interpretation of high-resolution remote sensing imagery. The state-of-the-art methods, e.g. bag-of-visual-words model and its various extensions as well as the topic models, share similar procedures: patch sampling, feature description/learning and classification. Patch sampling is the first and the key procedure which has a great influence on the results. In this paper, we focus on the effects of different sampling strategies used in the literature sa as to find a suitable sampling strategy for the scene classification of high-resolution remote sensing images. We divide the existing sampling methods into two types: random sampling and saliency-based sampling, and embed them in the bag-of-visual-words framework for comparison owing to its simplicity, robustness and efficiency. Moreover, we compare it using another framework - Fisher kernel, to validate our conclusions. The experimental results on two commonly used datasets using two different frameworks both show that random sampling can give better or comparable results than other sampling methods. Jingwen Hu 0001, Gui-Song Xia, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2015 | Finding edges of buildings via a junction process in high-resolution remotely sensed imagesabstractThis paper addresses the problem of finding edges of buildings in high-resolution remotely sensed images, which is of great help for subsequent analysis of built-up areas in urban remote sensing. More precisely, we propose a novel algorithm to extract meaningful edges associated to buildings with their saliency, by integrating an edge detection procedure with a junction process. This is inspired by the observation that meaningful junctions mainly emerge around buildings rather than on non-building objects in high-resolution remote sensing images. Thus, given a high-resolution remotely sensed image, we first use an edge detection algorithm, e.g. canny edge extractor, to compute all possible candidate edges of buildings, and then refine those candidates by meaningful junctions around buildings with their significance. The meaningful junctions are provided by an a-contrario junction detector, whose significance is related to the structural saliency in the images. For the evaluation of the proposed method, we test it on a small set of remote sensing images of half-meter resolution from WorldView-2, IKONOS and QuicBird. It demonstrates that our approach can find and locate edges of buildings with high precision and efficiency. Nan Xue 0001, Gui-Song Xia, Liangpei Zhang 0001 |
IGARSS | 3 |
| 2015 | Dissimilarity measurements for processing and analyzing PolSAR data: A surveyabstractMeasuring the pairwise similarity/dissimilarity of data is of central importance for processing and analyzing PolSAR images. In the literature, a large variety of measurements have been used, however, it is still not clear how to choose appropriate similarity measures for a given task. This paper presents a brief summary and discussion of the dissimilarity measurements used for interpreting PolSAR Images. Wen Yang 0001, Gui-Song Xia, Carlos López-Martínez |
IGARSS | 3 |
| 2015 | Learning High-level Features for Satellite Image Classification With Limited Labeled SamplesabstractThis paper presents a novel method addressing the classification task of satellite images when limited labeled data is available together with a large amount of unlabeled data. Instead of using semi-supervised classifiers, we solve the problem by learning a high-level features, called semisupervised ensemble projection (SSEP). More precisely, we propose to represent an image by projecting it onto an ensemble of weak training (WT) sets sampled from a Gaussian approximation of multiple feature spaces. Given a set of images with limited labeled ones, we first extract preliminary features, e.g., color and textures, to form a low-level image description. We then propose a new semisupervised sampling algorithm to build an ensemble of informative WT sets by exploiting these feature spaces with a Gaussian normal affinity, which ensures both the reliability and diversity of the ensemble. Discriminative functions are subsequently learned from the resulting WT sets, and each image is represented by concatenating its projected values onto such WT sets for final classification. Moreover, we consider that the potential redundant information existed in SSEP and use sparse coding to reduce it. Experiments on high-resolution remote sensing data demonstrate the efficiency of the proposed method. Wen Yang 0001, Xiaoshuang Yin, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2014 | Texture Analysis with Shape Co-occurrence PatternsabstractThis paper presents a flexible shape-based texture analysis method by investigating the co-occurrence patterns of shapes. More precisely, a texture image is represented by a tree of shapes, each of which is associated with several attributes. The modeling of texture is thus converted to characterize the tree of shapes. To this aim, we first learn a set of co-occurrence patterns of shapes from texture images, then establish a bag-of-words model on the learned shape co-occurrence patterns (SCOPs), and finally use the resulting SCOPs distributions as features for texture analysis. In contrast with existing work, the proposed method not only inherits the strong ability to depict geometrical aspects of textures and the high robustness to variations of imaging conditions from the shape-based texture analysis method, but also provides a more flexible way to model shape relationships (high-order statistics) on the tree. To our knowledge, this is the first time to use co-occurrence patterns of explicit shapes as a tool for texture analysis. Experiments of texture retrieval and classification on various databases report state-of-the-art results and demonstrate the efficiency of the proposed method. Gang Liu 0013, Gui-Song Xia, Wen Yang 0001, Liangpei Zhang 0001 |
ICPR | 2 |
| 2014 | Unsupervised feature coding on local patch manifold for satellite image scene classificationabstractThis paper presents an improved unsupervised feature learning (UFL) pipeline to discover intrinsic structures of local image patches as well as learn good feature representations automatically for image scenes. In our method, the original image patch vectors embedded in the high-dimensional pixel space are first mapped into a low-dimensional intrinsic space by linear manifold techniques, and then k-means clustering is performed on the patch manifold to learn a dictionary for feature encoding. To generate the feature representation for each local patch, triangle encoding method is applied with the learned dictionary on the same patch manifold. Finally, the holistic scene representations are obtained via the bag-of-visual-words (BOW) framework. We apply the proposed method on an aerial scene dataset. Experiments on the dataset show very promising results and demonstrate that our UFL pipeline can generate very effective local features for image scenes. Gui-Song Xia, Zifeng Wang 0007, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2014 | SAR image segmentation via non-local active contoursabstractThis paper presents a method for SAR image segmentation by relying on active contour model with the non-local processing principle [1]. The idea is to partition a SAR image via computing the patch similarity in the SAR image non-locally, and formulize the segmentation problem with an active contour model. More precisely, after computing the statistical features of SAR images, non-local comparisons between feature patches are used to calculate the active contour energy, which is defined by integrating the interactions between pairs of patches inside and outside the segmented region. A level set method is finally used to minimize the non-local energy. Compared with existing approaches for SAR image segmentation, the only requirement of this method is a local similarity between patches, and it is less sensitive to initial segmentation. The experimental results show the effectiveness and feasibility of the proposed method. Gang Liu 0013, Gui-Song Xia, Wen Yang 0001, Nan Xue 0001 |
IGARSS | 2 |
| 2014 | Spectral active clustering of remote sensing imagesabstractMining useful information from remote sensing images is a longstanding and challenging problem in earth observation, among which images clustering is used to discover meaningful scene information, by grouping similar image pixels into clusters. The main difficulty of image clustering, however, lies in the fact that imperfect similarity measure between images usually leads to bad clustering results. Supervised classification with labeled training samples can partially solve this problem, but the collection of such labeled data is usually time-consuming and sometimes impossible in many real problems. This paper presents an active remote sensing image clustering algorithm by integrating simple human queries into the clustering process. More precisely, we propose a spectral active clustering method that can actively query the oracle (such as human) to improve the image clustering performance. We first construct a k-nearest neighbor (k-NN) graph of the remote sensing images. We then iteratively select the most informative pairwise constraints and purify the k-NN graph, by removing the edges between images from different classes. The final clustering on the purified k-NN graph leads to more accurate result. The proposed method has been evaluated on three high-resolution remote sensing image datasets. It achieves the state-of-the-art performance and demonstrates high potentials in practical remote sensing applications. Zifeng Wang 0007, Gui-Song Xia, Caiming Xiong, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2014 | Anisotropic diffusion on complex tensor fields for PolSAR image filteringabstractThis paper addresses the problem of denoising for complex tensor image. In particular, we extend the anisotropic diffusion, also known as PM model (Perona and Malik, 1990) for filtering images based on PDE, from scalar or vector images to complex tensor ones and apply the new method to remove speckle noises of PolSAR images. Nan Xue 0001, Gui-Song Xia, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2014 | Semi-supervised feature learning for remote sensing image classificationabstractThis paper presents a semi-supervised method for learning informative image representations, which is a crucial but challenging step for remote sensing image classification. More precisely, we propose to represent an image by projecting it onto an ensemble of prototype sets sampled from a Gaussian approximation of multiple feature spaces. Given a set of images with a few labeled ones, we first extract preliminary features, e.g. color and textures, to form a low-level image description. We then build an ensemble of informative prototype sets by exploiting these feature spaces with a Gaussian normal affinity. Discriminative functions are subsequently learned from the resulting prototype sets, and each image is represented by concatenating their projected values onto such prototypes for final classification. Experiments on two high-resolution remote sensing image sets demonstrate the efficiency of the proposed method on remote sensing image classification with different classifiers. Xiaoshuang Yin, Wen Yang 0001, Gui-Song Xia, Lixia Dong |
IGARSS | 3 |
| 2014 | Accurate Junction Detection and Characterization in Natural Images
Gui-Song Xia, Julie Delon, Yann Gousseau |
Int. J. Comput. Vis. | 1 |
| 2014 | Synthesizing and Mixing Stationary Gaussian Texture ModelsabstractThis paper addresses the problem of modeling textures with Gaussian processes, focusing on color stationary textures that can be either static or dynamic. We detail two classes of Gaussian processes parameterized by a small number of compactly supported linear filters, the so-called textons. The first class extends the spot noise texture model to the dynamical setting, where the space-time texton is estimated to fit a translation-invariant covariance from an input exemplar. The second class is a specialization of the autoregressive dynamic texture method to the setting of space- and time-stationary textures. This enables one to parameterize the covariance with only a few spatial textons. The simplicity of these models allows us to tackle a more complex problem, texture mixing, which, in our case, amounts to interpolating between Gaussian models. We use optimal transport to derive geodesic paths and barycenters between the models learned from an input data set. This enables the user to navigate inside the set of texture models and perform texture synthesis from each new interpolated model. Numerical results on a library of exemplars show the ability of our method to generate arbitrary interpolations among unstructured natural textures. Moreover, experiments on a database of stationary textures show that the methods, despite their simplicity, provide state-of-the-art results on stationary dynamical texture synthesis and mixing. Gui-Song Xia, Sira Ferradans, Gabriel Peyré, Jean-François Aujol |
SIAM J. Imaging Sci. | 1 |
| 2013 | A Hierarchical Scheme of Multiple Feature Fusion for High-Resolution Satellite Scene Categorization
Wen Shao, Wen Yang 0001, Gui-Song Xia, Gang Liu 0013 |
ICVS | 3 |
| 2013 | A perception-inspired building index for automatic built-up area detection in high-resolution satellite imagesabstractThis paper addresses the problem of automatic extraction of built-up areas from high-resolution remote sensing images. We propose a new building presence index from the point view of perception. We argue that built-up areas usually result in significant corners and junctions in high-resolution satellite images, due to the man-made structures and occlusion, and thus can be measured by the geometrical structures they contained. More precisely, we first detect corners and junctions by relying on a perception-inspired corner detector, called an a-contrario junction detector. Each detected corner is associated with a perceptual significance, which measures the structural saliency of the corner in the image and is independent of the contrast and scale. All these detected corners together with their significance are then used to compute the building index. The proposed approach is evaluated on a high-resolution satellite image set, including 15 big images from GeoEye-1, QuickBird and IKONOS. The results demonstrated that our method achieves the state-of-the-art results and can be used in practical applications. Gang Liu 0013, Gui-Song Xia, Xin Huang 0002, Wen Yang 0001, Liangpei Zhang 0001 |
IGARSS | 2 |
| 2013 | Change detection in multi-temporal TerraSAR-X SAR images using a hierarchical Markov model on regionsabstractThis paper addresses the problem of change detection in high-resolution multi-temporal synthetic aperture radar (SAR) images (e.g. TerraSAR-X SAR images). Given two images, the proposed method first computes a difference map between them, by taking into account both the spatial and temporal correlations. Change detection is then formulated as a binary (changed/unchanged) segmentation problem of the difference map. A hierarchical Markov model (HMM) is defined on the multi-scale over-segmented regions of the difference map. The change map is finally inferred by relying on the hierarchical marginal posterior mode (HMPM) of the HMM. Experimental results on multi-temporal TerraSAR-X SAR images demonstrate the effectiveness and the reliability of the proposed approach. Wen Yang 0001, Gui-Song Xia, Mingsheng Liao |
IGARSS | 3 |
| 2012 | Compact representations of stationary dynamic texturesabstractThis paper addresses the problem of modeling stationary color dynamic textures with Gaussian processes. We detail two particular classes of such processes that are parameterized by a small number of compactly supported linear filters, so-called dynamical textons (dynTextons). The first class extends previous works on the spot noise texture model to the dynamical setting. It directly estimates the dynTexton to fit a translation-invariant covariance from the exemplar. The second class is a specialization of the auto-regressive (AR) dynamic texture method to the setting of space and time stationary textures. This allows one to parameterize the process covariance using only a few linear filters. Numerical experiments on a database of stationary textures shows that the methods, despite their extreme simplicity, provide state of the art results to synthesize space stationary dynamical texture. Gui-Song Xia, Sira Ferradans, Gabriel Peyré, Jean-François Aujol |
ICIP | 1 |
| 2012 | An accurate and contrast invariant junction detector
Gui-Song Xia, Julie Delon, Yann Gousseau |
ICPR | 1 |
| 2012 | Mid-level features and spatio-temporal context for activity recognition
Fei Yuan 0003, Gui-Song Xia, Hichem Sahbi, Véronique Prinet |
Pattern Recognit. | 2 |
| 2012 | SAR-Based Terrain Classification Using Weakly Supervised Hierarchical Markov Aspect ModelsabstractWe introduce the hierarchical Markov aspect model (HMAM), a computationally efficient graphical model for densely labeling large remote sensing images with their underlying terrain classes. HMAM resolves local ambiguities efficiently by combining the benefits of quadtree representations and aspect models-the former incorporate multiscale visual features and hierarchical smoothing to provide improved local label consistency, while the latter sharpen the labelings by focusing them on the classes that are most relevant for the broader local image context. The full HMAM model takes a grid of local hierarchical Markov quadtrees over image patches and augments it by incorporating a probabilistic latent semantic analysis aspect model over a larger local image tile at each level of the quadtree forest. Bag-of-word visual features are extracted for each level and patch, and given these, the parent-child transition probabilities from the quadtree and the label probabilities from the tile-level aspect models, an efficient forwards-backwards inference pass allows local posteriors for the class labels to be obtained for each patch. Variational expectation-maximization is then used to train the complete model from either pixel-level or tile-keyword-level labelings. Experiments on a complete TerraSAR-X synthetic aperture radar terrain map with pixel-level ground truth show that HMAM is both accurate and efficient, providing significantly better results than comparable single-scale aspect models with only a modest increase in training and test complexity. Keyword-level training greatly reduces the cost of providing training data with little loss of accuracy relative to pixel-level training. Wen Yang 0001, Dengxin Dai, Bill Triggs, Gui-Song Xia |
IEEE Trans. Image Process. | 4 |
| 2011 | Texture Segmentation by Grouping Ellipse Ensembles via Active ContoursabstractInternational audience Gui-Song Xia, Fei Yuan 0003 |
BMVC | 1 |
| 2010 | Topographic gray level multiscale analysis and its application to histogram modificationabstractThis paper describes a framework for multi-scale gray level analysis of images. It defines scales based on gray levels and organizes the basic “atoms” with a topographic map. The aim of this approach is to separate a large number of pixels concentrating in a narrow range of gray values. The main advantage of the methodology is that it allows manipulating pixels according to gray levels and spatial relations simultaneously. We apply it to histogram modification of Synthetic Aperture Radar (SAR) images. The experiments on displaying and classification prove the superiority of the approach. Chu He, Xinping Deng, Gui-Song Xia, Wen Yang 0001 |
ICIP | 3 |
| 2010 | Fast semantic scene segmentation with conditional random fieldabstractIn this paper, we present a fast approach to obtain semantic scene segmentation with high precision. We employ a two-stage classifier to label all image pixels. First, we use the regularized logistic regression to combine different appearance-based features and the improved spatial layout of labeling information. In the second stage, we incorporate the local, regional and global cues into a conditional random field model to provide a final segmentation, and a fast max-margin training method is employed to learn the parameters of the model quickly. The comparison experiments on four multi-class image segmentation databases show that our approach can achieve comparable semantic segmentation results and work faster than that of the state-of-the-art approaches. Wen Yang 0001, Dengxin Dai, Bill Triggs, Gui-Song Xia, Chu He |
ICIP | 4 |
| 2010 | Shape-based Invariant Texture Indexing
Gui-Song Xia, Julie Delon, Yann Gousseau |
Int. J. Comput. Vis. | 1 |
| 2008 | Locally invariant texture analysis from the topographic mapabstractIn this paper, we present a set of texture features that are locally invariant to similarity or affinity. The proposed indexing scheme relies on the topographic map, a shape-based representation of images. Thanks to the hierarchical organization of the topographic map, the approach gives a grip on the multi-scale structure of textures. Using simple one dimensional histograms, the method is shown to achieve state-of-the-art performances among locally invariant methods, both on the whole Brodatz and UIUC databases. Gui-Song Xia, Julie Delon, Yann Gousseau |
ICPR | 1 |
| 2007 | Compositional Boosting for Computing Hierarchical Image StructuresabstractIn this paper, we present a compositional boosting algorithm for detecting and recognizing 17 common image structures in low-middle level vision tasks. These structures, called "graphlets", are the most frequently occurring primitives, junctions and composite junctions in natural images, and are arranged in a 3-layer And-Or graph representation. In this hierarchic model, larger graphlets are decomposed (in And-nodes) into smaller graphlets in multiple alternative ways (at Or-nodes), and parts are shared and re-used between graphlets. Then we present a compositional boosting algorithm for computing the 17 graphlets categories collectively in the Bayesian framework. The algorithm runs recursively for each node A in the And-Or graph and iterates between two steps -bottom-up proposal and top-down validation. The bottom-up step includes two types of boosting methods, (i) Detecting instances of A (often in low resolutions) using Adaboosting method through a sequence of tests (weak classifiers) image feature, (ii) Proposing instances of A (often in high resolution) by binding existing children nodes of A through a sequence of compatibility tests on their attributes (e.g angles, relative size etc). The Adaboosting and binding methods generate a number of candidates for node A which are verified by a top-down process in a way similar to Data-Driven Markov Chain Monte Carlo [18]. Both the Adaboosting and binding methods are trained off-line for each graphlet category, and the compositional nature of the model means the algorithm is recursive and can be learned from a small training set. We apply this algorithm to a wide range of indoor and outdoor images with satisfactory results. Tianfu Wu 0001, Gui-Song Xia, Song-Chun Zhu |
CVPR | 2 |
| 2007 | A Rapid and Automatic MRF-Based Clustering Method for SAR ImagesabstractThis letter presents a precise and rapid clustering method for synthetic aperture radar (SAR) images by embedding a Markov random field (MRF) model in the clustering space and using graph cuts (GCs) to search the optimal clusters for the data. The proposed method is optimal in the sense of maximum a posteriori (MAP). It automatically works in a two-loop way: an outer loop and an inner loop. The outer loop determines the cluster number using a pseudolikelihood information criterion based on MRF modeling, and the inner loop is designed in a ldquohardrdquo membership expectation-maximization (EM) style: in the E step, with fixed parameters, the optimal data clusters are rapidly searched under the criterion of MAP by the GC; and in the M step, the parameters are estimated using current data clusters as ldquohardrdquo membership obtained in the E step. The two steps are iterated until the inner loop converges. Experiments on both simulated and real SAR images test the performance of the algorithm. Gui-Song Xia, Chu He |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2006 | An Adaptive and Iterative Method of Urban Area Extraction From SAR ImagesabstractThis letter presents a new method for unsupervised urban area extraction from synthetic aperture radar (SAR) images based on the ffmax algorithm proposed by C. Gouinaud specially for acquiring urban areas in SPOT imagery. According to the statistical characteristics of urban areas, an adaptive and iterative method based on the low-level extraction given by the ffmax algorithm using a large window is proposed. Experimental results on real SAR images show that the proposed automatic method works quickly and can preserve the borders of urban areas as well as avoid the disturbance of other classes and the extractions of urban areas are reliable and precise Chu He, Gui-Song Xia |
IEEE Geosci. Remote. Sens. Lett. | 2 |