VLDB 2026 Research / reviewers in the wild / expert
Zhi Gao 0005
dblp:94/7649-5
· DBLP profile ↗
70ranked-venue papers
11as first author
50since 2021 · last 2025
0000-0003-3325-1183ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 5 first-author · 31 since 2021Artificial intelligence and machine learning · 23 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 7 since 2021Systems, architecture and hardware · 10 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing the Utilization of Color Information in Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation is crucial in various applications such as autonomous driving, robotics, and virtual reality, aiming to assign labels to each point in a cloud to reflect spatial relationships and boundaries. While previous methods primarily focus on geometric features, they often overlook the auxiliary role of color information, especially in scenes where geometric structures are less distinct. In this paper, we propose the Color Point Cloud Enhancement (CPCE) method to effectively leverage color information for improved 3D scene understanding. CPCE introduces a color information enhancement module with multi-scale consistency, enriching point features throughout the encoder stages. Additionally, we develop a novel contrastive learning module that uses relative color coordinates for point cloud serialization, allowing for the capture of positive and negative samples from distant points with similar color textures. Furthermore, we design a contrastive learning module tailored for scenes with weak geometric structures, enhancing feature representation through color-augmented contrast. Our method achieved a 78.1% mIoU on the ScanNet dataset, outperforming existing models trained on a single dataset. These results highlight the effectiveness of CPCE in scenarios where traditional methods struggle, particularly in enhancing segmentation accuracy by utilizing color as a critical feature. Zhi Gao 0005, Jingshi Wang, Luliang Tang, Min Cao 0001 |
ICRA | 2 |
| 2025 | A Two-Stage Open Compound Domain Adaptation Framework for Semantic Segmentation in Remote Sensing ImageryabstractUnsupervised Domain Adaptation (UDA) has emerged as a critical research direction in remote sensing (RS) interpretation, aiming to bridge the gap between the labeled source and unlabeled target domains. However, most methods are designed for either a single target domain or multi-domain setting with clear boundaries, which necessitates retraining UDA models for each target domain, making it even harder to directly generalize to unseen domains. This paper proposes a novel two-stage open compound domain adaptation framework for semantic segmentation in RS images, which models the target domain as a composite of multiple unknown but homogeneous domains and leverages image translation and meta-learning techniques. In the first stage, we meticulously design a cross-domain image translation model (CDIT) based on contrastive learning to rapidly align the appearance of target domain images with the source domain style. In the second stage, the translated target images are first processed by a pre-trained model to generate pseudo-labels. Subsequently, a dynamic class-wise memory model (DCWM) is designed to progressively update abstract categorical features, serving as external class guidance for semantic segmentation within a meta-learning framework. Specifically, meta-training is employed to iteratively learn and update domain-agnostic categorical memory of semantic classes, while meta-testing simulates memory retrieval and gradually refines the categorical memory using pseudo-labels to adapt to new domains. Additionally, intra-class cohesion and inter-class divergence losses are incorporated to enhance the abstraction and retention of categorical features, aligning more closely with human cognitive patterns. Extensive experiments on RS benchmarks and unseen real images demonstrate the superior generalization of our method compared to state-of-the-art approaches. Zhi Gao 0005, Ziyao Li, Mengjie Xie |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Neural Radiance Fields for Sparse Satellite Images Leveraging Semantic and Geometry ConsistencyabstractHow to achieve accurate scene reconstruction models using limited views has long been a key research topic in the fields of photogrammetry and computer vision. Under sparse view conditions, the lack of sufficient semantic and geometric priors often results in suboptimal performance of reconstruction algorithms. Recently, neural radiance fields (NeRF) have gained a lot of attention in satellite scene reconstruction. However, existing NeRFs for satellite scenes require many views to render novel views and generate DSMs, which conflicts with the scarcity of satellite imagery. The key to improving performance degradation lies in how to extract rich scene semantics and geometry from a limited number of views, thereby enhancing the understanding of the scene. This paper presents a novel NeRF-based method that leverages semantic and geometry consistency to enhance the NeRF performance with sparse satellite images. Specifically, we introduce a cross-view semantic consistency loss to enhance the semantic understanding capability of the NeRF model with limited views. We utilize a vision language foundation model for remote sensing, RemoteCLIP, to extract semantic embeddings and ensure the consistency of semantic features across different views of the same scene. Additionally, we propose a geometric consistency loss to ensure the geometric consistency between surface normals and scene depth, thereby ensuring the scene geometric accuracy when NeRF model generates images from unseen views. Extensive experiments on various urban satellite scenes demonstrate that our proposed method achieves superior performance on both novel view synthesis and DSM generation tasks under sparse-view training conditions. Zhi Gao 0005, Yanzhang Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | A Semi-Supervised Domain-Adaptive Framework for Real-World Underwater Image EnhancementabstractUnderwater optical remote sensing is crucial for geoscience applications but often suffers from image degradation due to complex underwater environments. While learning-based methods have advanced underwater image enhancement (UIE), their efficacy in real-world UIE applications still faces challenges. This limitation arises from training predominantly on synthetic underwater images, resulting in a significantinter-domain gap when applied to real-world data. Additionally, diverse underwater conditions introduceintra-domain challenges, such as color casts and haze, further complicating the UIE process. To address these issues, we propose SSD-UIE, a semi-supervised domain-adaptive framework designed to mitigate bothinter- andintra-domain gaps. Our approach employs a systematic synthesis pipeline to reduce visualinter-domain discrepancies and introduces a Large Synthetic-Real Underwater Image Dataset (LSRUID) to facilitate the training of the framework. The Semantic-Blender is developed to handle semanticinter-domain differences, while the Intra-domain-aware Feature Extraction (IFE) branch and feature alignment strategy effectively addressintra-domain variability. Furthermore, the Dual-Trans Block is introduced to enhance the UIE performance while maintaining computational efficiency. Extensive experiments demonstrate that SSD-UIE outperforms state-of-the-art (SOTA) UIE methods in both qualitative and quantitative evaluations on real-world underwater images. Codes and dataset will be publicly available at https://github.com/RockWenJJ/SSD-UIE.git. Junjie Wen 0001, Guidong Yang, Benyun Zhao, Dongyue Huang, Lei Lei 0010, Bo Zhang 0019, Zhi Gao 0005, Xi Chen 0104, Ben M. Chen |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Toward End-to-End Underwater Multi-View Stereo for Real-World Dense Scene ReconstructionabstractMulti-view stereo (MVS) enables accurate and complete 3D reconstruction from multi-view imagery, serving as a core methodology in remote sensing applications across terrestrial and underwater domains. Recent advancements in learning-based MVS have demonstrated significant improvements over traditional counterparts, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi-view stereo (UwMVS) faces substantial challenges arising from the domain gap between in-air and underwater environments, leading to degraded performance when applying in-air MVS models to underwater scenarios. Furthermore, the progress of learning-based UwMVS methods has been hindered by the scarcity of underwater multi-view images with ground-truth depth maps and point clouds. In this paper, we address these challenges by introducing a physically-guided approach for synthesizing underwater multi-view images and presenting the first large-scale synthetic UwMVS dataset preserving real-world underwater degradation properties for end-to-end training and evaluation of learning-based UwMVS methods. Furthermore, we propose a novel UwMVS network that enhances geometric cue encoding to achieve more accurate and complete point cloud reconstruction. Extensive experiments on the dataset and real-world underwater scenes demonstrate that our dataset enables the trained models for underwater dense reconstruction and that our method achieves state-of-the-art performance in underwater reconstruction. Dataset, appendix, and supplementary video are available at https://yang-sober.github.io/UnderMVS/. Guidong Yang, Junjie Wen 0001, Lei Lei 0010, Benyun Zhao, Qingxiang Li, Xi Chen 0104, Zhi Gao 0005, Ben M. Chen |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Enhancing Perception of Key Changes in Remote Sensing Image Change CaptioningabstractRecently, while significant progress has been made in remote sensing image change captioning, existing methods fail to filter out areas unrelated to actual changes, making models susceptible to irrelevant features. In this article, we propose a novel multimodal model for remote sensing image change captioning, guided by Key Change Features and Instruction-tuned (KCFI). This model aims to fully leverage the intrinsic knowledge of large language models through visual instructions and enhance the effectiveness and accuracy of change features using pixel-level change detection tasks. Specifically, KCFI includes a ViTs encoder for extracting bi-temporal remote sensing image features, a key feature perceiver for identifying critical change areas, a pixel-level change detection decoder to constrain key change features, and an instruction-tuned decoder based on a large language model. Moreover, to ensure that change captioning and change detection tasks are jointly optimized, we employ a dynamic weight-averaging strategy to balance the losses between the two tasks. We also explore various feature combinations for visual fine-tuning instructions and demonstrate that using only key change features to guide the large language model is the optimal choice. To validate the effectiveness of our approach, we compare it against several state-of-the-art change captioning methods on the LEVIR-CC dataset, achieving the best performance. Our code will be available at https://github.com/yangcong356/KCFI.git. Zuchao Li, Hongzan Jiao, Zhi Gao 0005, Lefei Zhang |
IEEE Trans. Image Process. | 4 |
| 2024 | SurfOcc: Surface-Based Feature Lifting for Vision-Centric 3D Occupancy Prediction
Tonghui Ye, Zhi Gao 0005, Xinyi Liu 0002, Ronghe Jin |
ACCV (10) | 2 |
| 2024 | SGCalib: A Two-stage Camera-LiDAR Calibration Method Using Semantic Information and Geometric FeaturesabstractExtrinsic calibration is an essential prerequisite for the applications of camera-LiDAR fusion. Existing methods either suffer from the complex offline setting of man-made targets or tend to produce suboptimal and unrobust results. In this paper, we propose an online two-stage calibration method that estimates robust and accurate extrinsic parameters between camera and LiDAR. This is a novel work to use semantic information and geometric features jointly in calibration to promote accuracy and robustness. In the first stage, we detect objects in the image and point cloud and build graphs on the objects using Delaunay triangulation. Then, we design a novel graph matching algorithm to associate the objects in the two data domains and extract pairs of 2D-3D points. Using the PnP solver, we get robust initial extrinsic parameters. Then, in the second stage, we design a new optimization formulation with semantic information and geometric features to generate accurate extrinsic parameters with the initial value from the first stage. Extensive experiments on solid-state LiDAR, conventional spinning LiDAR and KITTI datasets have verified the robustness and accuracy of our method which outperforms existing works. We will share the code publicly to benefit the community (after review stages). Zhi Gao 0005, Xinyi Liu 0002, Ben M. Chen |
ICRA | 2 |
| 2024 | EnYOLO: A Real-Time Framework for Domain-Adaptive Underwater Object Detection with Image EnhancementabstractIn recent years, significant progress has been made in the field of underwater image enhancement (UIE). However, its practical utility for high-level vision tasks, such as underwater object detection (UOD) in Autonomous Underwater Vehicles (AUVs), remains relatively unexplored. It may be attributed to several factors: (1) Existing methods typically employ UIE as a pre-processing step, which inevitably introduces considerable computational overhead and latency. (2) The process of enhancing images prior to training object detectors may not necessarily yield performance improvements. (3) The complex underwater environments can induce significant domain shifts across different scenarios, seriously deteriorating the UOD performance. To address these challenges, we introduce EnYOLO, an integrated real-time framework designed for simultaneous UIE and UOD with domain-adaptation capability. Specifically, both the UIE and UOD task heads share the same network backbone and utilize a lightweight design. Furthermore, to ensure balanced training for both tasks, we present a multi-stage training strategy aimed at consistently enhancing their performance. Additionally, we propose a novel domain-adaptation strategy to align feature embeddings originating from diverse underwater environments. Comprehensive experiments demonstrate that our framework not only achieves state-of-the-art (SOTA) performance in both UIE and UOD tasks, but also shows superior adaptability when applied to different underwater scenarios. Our efficiency analysis further highlights the substantial potential of our framework for onboard deployment. Junjie Wen 0001, Jinqiang Cui, Benyun Zhao, Bingxin Han, Xuchen Liu 0001, Zhi Gao 0005, Ben M. Chen |
ICRA | 6 |
| 2024 | Adapting Vision Transformer for Few-Shot Remote Sensing Image Segmentation: Synergizing In-Domain Representations And Pretrained Model GuidanceabstractDeep neural networks, especially Transformers, struggle with generalizing to objects of unseen categories under conditions of data scarcity. In this work, we adapt a plain Vision Transformer (ViT) into an excellent few-shot learner for remote sensing image (RSI) semantic segmentation. Unlike previous few-shot segmentation methods, we introduce a dual-branch strategy featuring one branch with frozen backbones and another with tunable backbones, synergizing the strengths of both choices. In the tunable branch, we meta-learn a ViT and a refiner on RSI datasets, attaining in-domain representations while learning the few-shot capability. In the frozen branch, we leverage representations from a fixed pretrained ViT with an orientation-robust design. The synergy of the two branches is achieved explicitly through prototype fusion. Implicitly, frozen pretrained models are utilized to regularize the learning of in-domain representations through loss optimization. Extensive in-domain and cross-domain experiments validate the superiority and application potential of our method. Zhi Gao 0005 |
IGARSS | 2 |
| 2024 | CDTCL: Cross-Domain Remote Sensing Image Translation for Semantic Segmentation Leveraging Contrastive LearningabstractAlthough Deep Learning based methods for remote sensing (RS) image interpretation have reported promising results, the domain gap between RS images and the absence of sensor-specific labeled datasets result in the significant deterioration of well-trained models to adapt to new images. In practical applications, we propose an RS image translation method based on contrastive learning (CDTCL) to quickly achieve the data simulation conversion for different sensors and domains. Specifically, for unpaired images, we design a content contrastive loss for content consistency constraints and a style contrastive loss for swift alignment of appearance style. Additionally, we integrate the semantic segmentation model, a flexible model that can be retrained at the discretion of users, into the image translation framework to establish a complementary closed loop. Extensive experiments on aerial images, including visible and infrared images, verify that our method works effectively in cross-domain semantic segmentation and achieves the best performance. Ziyao Li, Zhengyi Lei, Mengjie Xie, Yanzhang Li, Zhi Gao 0005 |
IGARSS | 7 |
| 2024 | Sparse Autoencoder Based Hyperspectral Anomaly Detection with the Singular Spectrum Analysis Based Spectral DenoisingabstractAs an effective tool for monitoring surface irregularities in remote sensing, hyperspectral anomaly detection (HAD) has garnered increasing attention. However, how to improve the detection accuracy remains a formidable challenge, due mainly to the noise and variations in the spectral domain, especially when there is lack of the labelled data for training. To tackle these difficulties, a novel unsupervised HAD method is proposed. First, 1-D Singular Spectrum Analysis (SSA) is employed to eliminate outliers in the spectral domain. Second, the SSA-smoothed hypercube undergoes a sparse autoencoder for background reconstruction, where the reconstruction error is used to extract anomalous pixels. Finally, the RX algorithm is employed to segment anomalous pixels from the background. Comprehensive experiments on four publicly available datasets have validated the superior performance of our method in effectively enhancing the separability between anomaly pixels and their respective backgrounds, outperforming a few state-of-the-art methods, particularly in terms of the detection accuracy. Yinhe Li, Jinchang Ren, Zhi Gao 0005, Genyun Sun |
IGARSS | 3 |
| 2024 | Neural Radiance Fields for Multi-View Satellite Photogrammetry Leveraging Intrinsic DecompositionabstractHigh-precision geometric information derived from remote sensing scenes plays a critical role in digital surface modeling. However, acquiring such information from multi-view satellite imagery presents a challenging task due to the complexities of scene layout, illumination, and albedo. This paper presents a novel Neural Radiance Field (NeRF) that leverages intrinsic decomposition on multiview satellite image collections by inversing the photometric image formation model. Specifically, our model decomposes the satellite image into multiple intrinsic components which capture essential aspects of the scene, including albedo (reflectance property), normal (surface orientation), shadow (occlusion information), and lighting (illumination conditions). The final irradiance is derived by integrating the obtained intrinsic components via volume rendering, adhering to the physics-based photometric image formation model. Experiments on WorldView-3 images demonstrate that our model produces more photorealistic and physically plausible results for the task of novel view synthesis and digital surface modeling. Zhi Gao 0005 |
IGARSS | 7 |
| 2024 | Accurate and Efficient Loop Closure Detection With Deep Binary Image Descriptor and Augmented Point Cloud RegistrationabstractLoop Closure Detection (LCD) is an essential component of Simultaneous Localization and Mapping (SLAM), helping to correct drift errors, facilitate map merging, or both by identifying previously observed scenes. Despite its importance, traditional LCD algorithms based on single sensor such as camera or LiDAR exhibit degraded performance in challenging scenarios due to their inherent limitations. To address this issue, we propose a novel LCD method based on camera-LiDAR fusion, exploiting the rich textural information from cameras and the accurate geometric data from LiDAR to ensure robustness and speed in challenging environments. Specifically, we first employ deep hashing learning to encode deep image features into binary image descriptors for extremely fast loop candidate (LC) retrieval. Then, LiDAR points are augmented with image color for accurate geometric verification. Finally, we incorporate a spatial-temporal consistency check that mandates an LC to have consistently matched neighbors to be accepted as true. Our method is extensively verified and compared with the state-of-the-art methods on various datasets encompassing both indoor and outdoor environments. Experimental results demonstrate that our method obtains the best performance, increasing the maximum recall rate at 100% precision by a significant margin of 20% while operating in real-time at an average speed of 30 fps. Zhi Gao 0005, Jianhua Cheng, Xinyi Liu 0002, Ben M. Chen |
IROS | 2 |
| 2024 | Single image deraindrop leveraging luminance priors and context aggregation
Yi Liu 0146, Zhi Gao 0005, Tiancan Mei, Han Yi |
Neurocomputing | 2 |
| 2024 | DDformer: Dimension decomposition transformer with semi-supervised learning for underwater image enhancement
Zhi Gao 0005, Jing Yang 0041, Fengling Jiang, Xixiang Jiao, Kia Dashtipour, Mandar Gogate, Amir Hussain 0001 |
Knowl. Based Syst. | 1 |
| 2024 | Prompting-to-Distill Semantic Knowledge for Few-Shot LearningabstractRecognizing visual patterns in low-data regime necessitates deep neural networks to glean generalized representations from limited training samples. In this letter, we propose a novel few-shot classification method, namely ProDFSL, leveraging multimodal knowledge and attention mechanism. We are inspired by recent advances of large language models and the great potential they have shown across a wide range of downstream tasks and tailor it to benefit the remote sensing community. We utilize ChatGPT to produce class-specific textual inputs for enabling CLIP with rich semantic information. To promote the adaptation of CLIP in remote sensing domain, we introduce a cross-modal knowledge generation module, which dynamically generates a group of soft prompts conditioned on the few-shot visual samples and further uses a shallow Transformer to model the dependencies between language sequences. Fusing the semantic information with few-shot visual samples, we build representative class prototypes, which are conducive to both inductive and transductive inference. In extensive experiments on standard benchmarks, our ProDFSL consistently outperforms the state of the art in few-shot learning (FSL). Zhi Gao 0005, Jinchang Ren, Xing-ao Wang, Ping Ma 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | MD-NeRF: Enhancing Large-Scale Scene Rendering and Synthesis With Hybrid Point Sampling and Adaptive Scene DecompositionabstractNeural radiance fields (NeRFs) have gained great success in 3-D representation and novel-view synthesis, which attracted great efforts devoted to this area. However, when rendering large-scale scenes from a drone perspective, existing NeRF methods exhibit pronounced distortions in scene detail including absent textures and blurring of small objects. In this letter, we propose MD-NeRF to mitigate such distortions by integrating a hybrid sampling strategy and an adaptive scene decomposition method. Specifically, an anti-aliasing sampling method combining spiral sampling and sampling along rays is presented to address rendering anomalies. In addition, we decompose a large scene into multiple subscenes using a mixture of expert (MoE) modules. A shared expert is introduced to capture common features and reduce redundancy across the specialized experts. Consequently, the combination of these two methods effectively minimizes distortions when rendering large-scale scenes and enables our model to produce finer textures and more coherent details. We have conducted extensive experiments on several large-scale unbounded scene datasets, and the results demonstrate that our approach has achieved state-of-the-art performance on all datasets, most notably evidenced by a 1-dB enhancement in PSNR metrics on the Mill19 dataset. Zhi Gao 0005 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Hierarchical GNN Framework for Earth's Surface Anomaly Detection in Single Satellite ImageryabstractSudden-onset Earth’s surface anomalies, such as natural disasters and man-made incidents, pose severe threats to human life and property security, emphasizing the crucial role of accurate detection and rapid response in Humanitarian Assistance and Disaster Response (HADR). In this work, we propose a hierarchical graph neural network (GNN) based framework for Earth’s surface anomaly detection, called L2S-Net, to integrate from local to semantic (L2S) information for rapid and accurate detection of multi-class anomalies. Specifically, L2S-Net only utilizes a single very high-resolution (VHR) image as input to expedite processing speed, while employing a hierarchical graph representation for better image understanding. Meanwhile, drawing from brain-inspired research and graph theory, we design a local-to-semantic fusion network, called L2S-GNN, to explicitly learn relationships between nodes at different levels facilitating accurate detection of Earth’s surface anomaly. L2S-Net significantly reduces data requirements while capturing valuable higher-order information from images, achieving a superior balance between accuracy and efficiency. Furthermore, due to the lack of a public dataset for Earth’s surface anomaly detection, we create a novel and large-scale benchmark dataset ESADv2. Extensive experiments on the ESADv2 dataset and two real-world cases demonstrate that the proposed L2S-Net outperforms many state-of-the-art methods in both model size and performance while exhibiting exceptional generalizability and robustness. Boan Chen, Zhi Gao 0005, Ziyao Li, Aohan Hu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Query Adaptive Transformer and Multiprototype Rectification for Few-Shot Remote Sensing Image SegmentationabstractDeep learning has emerged as a powerful tool for semantic segmentation tasks. However, in some data-deficient and resource-limited scenarios, networks are constrained to learn novel concepts. To tackle this problem, few-shot segmentation (FSS) has been proposed by the machine learning community, aiming at segmenting novel objects by leveraging a handful of annotated samples. However, most advanced FSS algorithms suffer from severe performance degradation when directly applied to remote sensing image (RSI) domains. Challenges arise due to the unique characteristics of RSIs. To address the large intraclass variation brought by various imaging conditions and large-scale variations, a query adaptive transformer (QAT) is proposed. QAT incorporates query priors into the feature extraction process, adapting RSI features to novel objects from query RSIs. It is noteworthy that query priors are extracted via a query prior generation (QPG) module devoid of any query label information. To alleviate interference brought by complex object distribution, a multiprototype strategy is adopted instead of representing objects with a single prototype ambiguously. Moreover, we rectify the prototypes using a prior injection module (PIM), thereby fully leveraging the advantages offered by query priors. The superiority of our method is validated through comprehensive experiments on the public iSAID-$5^{i}$dataset and comparisons with state-of-the-art methods. Finally, we propose a novel cross-domain setting to investigate the potential and generalizability of few-shot RSI segmentation for several Earth observation applications. Zhi Gao 0005 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Semi-Supervised Few-Shot Classification With Multitask Learning and Iterative Label CorrectionabstractFew-shot learning enables rapid generalization from extremely limited training examples. While previous efforts have utilized meta-learning or data augmentation methods to mitigate the problem of data scarcity, such approaches may struggle to maintain robustness and generalize effectively due to overfitting and noise sensitivity. In this paper, we propose a novel approach, the Semi-Supervised Label Correction method for Few-Shot Learning (SSLC-FSL), which leverages the data distribution of readily available and easily obtainable unlabeled data. SSLC-FSL iteratively corrects the labels of testing samples with alternating steps of pseudo-labeling and sample selection. The objective of pseudo-labeling is to repurpose graph-based semi-supervised learning for joint prediction of the entire testing set. We then introduce a Modulation Selection Network (MSN) to rank testing samples by learning with noisy labels. The training set is expanded by selecting confident pseudo-labeled samples. In the MSN, a Modulation Aggregation Layer is designed to encode support class information into each testing sample, thereby highlighting target category features and mitigating the negative impact of incorrect labels. The iterative label correction process is repeated until all testing samples are recalled to the expanded support set. To boost the SSLC-FSL algorithm, we pre-train a feature extractor to produce general-purpose representations. Particularly, we investigate two types of auxiliary tasks and their collaborative learning to acquire transferable visual information via an end-to-end multi-task learning model. Our SSLC-FSL outperforms current state-of-the-art methods in any shot and all data settings, with up to +27.74% on standard remote sensing benchmarks and +5.70% on standard natural scene benchmarks. Zhi Gao 0005, Ziyao Li, Boan Chen, Yanzhang Li, Zhicheng Shi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Language Query-Based Transformer With Multiscale Cross-Modal Alignment for Visual Grounding on Remote Sensing ImagesabstractVisual grounding for remote sensing images (RSVG) aims to localize the referred objects in the remote sensing (RS) images according to a language expression. Existing methods tend to align visual and text features followed by concatenation and then employ a fusion Transformer to learn a token representation for final target localization. However, simple fusion Transformer structure fails to sufficiently learn the location representation of referred object from the multi-modal features. Inspired by the detection Transformer, in this paper, we propose a novel language query based Transformer framework for RSVG termed LQVG. Specifically, we adopt the extracted sentence-level text features as the queries, called language queries, to retrieve and aggregate representation information of the referred object from the multi-scale visual features in the Transformer decoder. The language queries are then converted into object embeddings for final coordinate prediction of referred object. Besides, a multi-scale cross-modal alignment module is devised before the multimodal Transformer to enhance the semantic correlation between the visual and text features, thus facilitating the cross-modal decoding process to generate more precise object representation. Moreover, a new RSVG dataset named RSVG-HR is built to evaluate the performance of the RSVG approaches on very high-resolution remote sensing images with inconspicuous objects. Experimental results on two benchmark datasets demonstrate that our proposed method significantly surpasses the comparison methods and achieves state-of-the-art performance. The dataset and code are available at https://github.com/LANMNG/LQVG. Meng Lan, Fu Rong, Hongzan Jiao, Zhi Gao 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Segmenting Remote Sensing Anomalies at Instance Level via Anomaly Map-Guided AdaptationabstractEarth anomalies can locate valuable targets in an unsupervised manner for many defense and surveillance applications. Most models assign a continuous score at the pixel-level, resulting in object-agnostic results with higher false alarms than the instance level results. However, since the anomaly objects contain a variety of categories and have a large intraclass variance, the current state-of-the-art (SOTA) query-based models designed for certain categories perform unsatisfactorily when applied to the anomaly instances. The larger intraclass variance of anomalies makes the learning of general representation more difficult. To bridge this gap, we propose general adaptations guided by the pixel-level anomaly map for any query-based model, which adapts the model from learning certain category representation to learning anomaly-aware representation in different categories. The proposed adaptation first builds a separate branch to output the pixel-level anomaly map, where anomaly information is then extracted to guide the pixel embeddings and queries to focus on a variety of anomaly categories. Especially, the anomaly rank embeddings are devised to make the pixel embeddings aware of the anomaly rank order. The queries are dynamically selected from the anomaly candidates after aligning the anomaly map and pixel embeddings for better locating different anomalies. Finally, the selected queries dot-product the anomaly-aware pixel embeddings to output the anomaly instances. The proposed adaptations are simple, general, and additive, which bring the average improvements of +4.9 box AP and +5.1 mask AP in infrared, synthetic aperture radar (SAR), and hyperspectral modalities. Yanfei Zhong, Hengwei Zhao, Zhi Gao 0005, Xinyu Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Bolstering Performance Evaluation of Image Segmentation Models With Efficacy Metrics in the Absence of a Gold StandardabstractImage segmentation using deep learning has become overwhelmingly widespread. However, routine model testing methods can encounter evaluation inconsistencies or bias, largely due to how accuracy metrics respond to variations in class share distribution. Here, we address the effects of class imbalance on model performance evaluation and demonstrate a refined approach that incorporates image classification efficacy (ICE) metrics within the context of semantic segmentation in remote sensing. This evaluation approach was applied in six segmentation experiments that involved multispectral and LiDAR data, single or multiple models tested with the same or different datasets, and binary and multiclass schemes. ICE metrics revealed unique aspects of model’s segmentation capabilities compared to precision, recall, F-score, and overall accuracy. By mitigating the class imbalance effect, per-class efficacy enables precise class-level optimization of segmentation models, while whole-class efficacy facilitates evaluating a model’s potential performance when adapted to new datasets. The suitability of the kappa coefficient, ROC-AUC, and PR-AUC for model evaluation under class imbalance was discussed in comparison with ICE metrics. This efficacy-enhanced model evaluation protocol can be implemented for deep learning model training and testing. The routine use of this evaluation approach will strengthen the dependability and applicability of segmentation tools in various fields. Lina Tang, Jinyuan Shao, Shiyan Pang, Yameng Wang, Aaron E. Maxwell, Xiangyun Hu, Zhi Gao 0005, Ting Lan 0003, Guofan Shao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | S³ANet: Spatial-Spectral Self-Attention Learning Network for Defending Against Adversarial Attacks in Hyperspectral Image ClassificationabstractDeep neural networks have demonstrated impressive capabilities in hyperspectral image (HSI) classification tasks. However, they are highly vulnerable to adversarial attacks, raising significant security concerns, especially in the remote sensing community. Even subtle adversarial perturbations that are imperceptible to human observers can mislead deep learning (DL) models and result in incorrect predictions. Therefore, ensuring the robustness of DL models has become a critical focus in addressing security-related remote sensing tasks. Considerable progress has been made in defending against adversarial attacks in HSI classification. Nevertheless, existing methods primarily concentrate on spatial relationships between pixels while overlooking the valuable spectral information present in HSI. Besides, these methods are usually limited to a specific scale and cannot accommodate the precise classification demands for ground objects with various scales. To address these limitations, we propose a spatial–spectral self-attention learning network (S3ANet) for defending against adversarial attacks in HSI classification. Our S3ANet incorporates a pyramid spatial attention learning module to effectively capture spatial dependency at multiple scales. In addition, it utilizes a global spectral transformer to establish correlations between pixels in the spectral dimension. By employing the defense method of spatial–spectral fusion, our model can effectively address adversarial attacks from a comprehensive perspective, seamlessly integrating both spatial and spectral information. Extensive experiments conducted on four benchmark HSI datasets illustrate that the proposed S3ANet achieves competitive performance compared to state-of-the-art methods when faced with adversarial attacks. The code is available online athttps://github.com/YichuXu/S3ANet. Yichu Xu, Yonghao Xu, Hongzan Jiao, Zhi Gao 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Global-Local MAV Detection Under Challenging Conditions Based on Appearance and MotionabstractVisual detection of micro aerial vehicles (MAVs) has received increasing research attention in recent years due to its importance in many applications. However, the existing approaches based on either appearance or motion features of MAVs still face challenges when the background is complex, the MAV target is small, or the computation resource is limited. In this paper, we propose a global-local MAV detector that can fuse both motion and appearance features for MAV detection under challenging conditions. This detector first searches MAV targets using a global detector and then switches to a local detector which works in an adaptive search region to enhance accuracy and efficiency. Additionally, a detector switcher is applied to coordinate the global and local detectors. A new dataset is created to train and verify the effectiveness of the proposed detector. This dataset contains more challenging scenarios that can occur in practice. Extensive experiments on three challenging datasets show that the proposed detector outperforms the state-of-the-art ones in terms of detection accuracy and computational efficiency. In particular, this detector can run with near real-time frame rate on NVIDIA Jetson NX Xavier, which demonstrates the usefulness of our approach for real-world applications. The dataset is available at https://github.com/WestlakeIntelligentRobotics/GLAD. In addition, A video summarizing this work is available at https://youtu.be/Tv473mAzHbU. Hanqing Guo, Zhi Gao 0005, Shiyu Zhao 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | TJ-FlyingFish: Design and Implementation of an Aerial-Aquatic Quadrotor with Tiltable Propulsion UnitsabstractAerial-aquatic vehicles are capable to move in the two most dominant fluids, making them more promising for a wide range of applications. We propose a prototype with special designs for propulsion and thruster configuration to cope with the vast differences in the fluid properties of water and air. For propulsion, the operating range is switched for the different mediums by the dual-speed propulsion unit, providing sufficient thrust and also ensuring output efficiency. For thruster configuration, thrust vectoring is realized by the rotation of the propulsion unit around the mount arm, thus enhancing the underwater maneuverability. This paper presents a quadrotor prototype of this concept and the design details and realization in practice. Xuchen Liu 0001, Minghao Dou, Dongyue Huang, Songqun Gao, Ruixin Yan, Biao Wang 0004, Jinqiang Cui, Qinyuan Ren, LiHua Dou, Zhi Gao 0005, Jie Chen 0003, Ben M. Chen |
ICRA | 10 |
| 2023 | SyreaNet: A Physically Guided Underwater Image Enhancement Framework Integrating Synthetic and Real ImagesabstractUnderwater image enhancement (UIE) is vital for high-level vision-related underwater tasks. Although learning-based UIE methods have made remarkable achievements in recent years, it's still challenging for them to consistently deal with various underwater conditions, which could be caused by: 1) the use of the simplified atmospheric image formation model in UIE may result in severe errors; 2) the network trained solely with synthetic images might have difficulty in generalizing well to real underwater images. In this work, we, for the first time, propose a framework SyreaNet for UIE that integrates both synthetic and real data under the guidance of the revised underwater image formation model and novel domain adaptation (DA) strategies. First, an underwater image synthesis module based on the revised model is proposed. Then, a physically guided disentangled network is designed to predict the clear images by combining both synthetic and real underwater images. The intra- and inter-domain gaps are abridged by fully exchanging the domain knowledge. Extensive experiments demonstrate the superiority of our framework over other state-of-the-art (SOTA) learning-based UIE methods qualitatively and quantitatively. The code and dataset are publicly available at https://github.com/RockWenJJ/SyreaNet.git. Junjie Wen 0001, Jinqiang Cui, Zhenjun Zhao, Ruixin Yan, Zhi Gao 0005, LiHua Dou, Ben M. Chen |
ICRA | 5 |
| 2023 | A New Multi-Level Attention Feature Fusion Method for Hyperspectral and Lidar Data Joint ClassificationabstractJoint classification of multisource data for better Earth observation becomes an interesting but challenging problem. However, existing methods usually fail to be optimal due to the limitations in the heterogeneous feature representation and complementary information fusion. In this paper, we propose a new multi-level attention-based feature fusion method for the joint classification of HSI and LiDAR data. First, a two-stream deep network is built to extract the spectral-spatial feature of HSI and the elevation feature of LiDAR, respectively. To fully use the complementary and correlated information of HSI and LiDAR data, we adopt attention-based feature extraction and fusion module to deliver a high-discrimination feature representation both for cross-source and single-source data. Then, the extracted features are fed into fully connected layers to generate class probabilities. Finally, a decision-level fusion strategy is adopted to further improve the classification results. Extensive experiments on the Houston dataset demonstrate the effectiveness of the proposed method over some state-of-the-art approaches. Zhi Gao 0005, Leyuan Fang, Yongjun Zhang 0002 |
IGARSS | 2 |
| 2023 | MFVNet: a deep adaptive fusion network with multiple field-of-views for remote sensing image semantic segmentation
Yansheng Li 0001, Wei Chen 0089, Xin Huang 0002, Zhi Gao 0005, Tao He 0002, Yongjun Zhang 0002 |
Sci. China Inf. Sci. | 4 |
| 2023 | Global-guided weakly-supervised learning for multi-label image classification
Zhi Gao 0005, Leyuan Fang |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | FG-Net: A Fast and Accurate Framework for Large-Scale LiDAR Point Cloud UnderstandingabstractThis work presents FG-Net, a general deep learning framework for large-scale point cloud understanding without voxelizations, which achieves accurate and real-time performance with a single NVIDIA GTX 1080 8G GPU and an i7 CPU. First, a novel noise and outlier filtering method is designed to facilitate the subsequent high-level understanding tasks. For effective understanding purpose, we propose a novel plug-and-play module consisting of correlated feature mining and deformable convolution-based geometric-aware modeling, in which the local feature relationships and point cloud geometric structures can be fully extracted and exploited. For the efficiency issue, we put forward a new composite inverse density sampling (IDS)-based and learning-based operation and a feature pyramid-based residual learning strategy to save the computational cost and memory consumption, respectively. Compared with current methods which are only validated on limited datasets, we have done extensive experiments on eight real-world challenging benchmarks, which demonstrates that our approaches outperform state-of-the-art (SOTA) approaches in terms of accuracy, speed, and memory efficiency. Moreover, weakly supervised transfer learning is also conducted to demonstrate the generalization capacity of our method. Kangcheng Liu, Zhi Gao 0005, Feng Lin 0003, Ben M. Chen |
IEEE Trans. Cybern. | 2 |
| 2023 | Few-Shot Aerial Image Semantic Segmentation Leveraging Pyramid Correlation FusionabstractFew-shot semantic segmentation has gained significant attention due to its ability to segment novel objects using only a limited number of labeled samples, thereby addressing the problem of overfitting caused by a lack of training data. Although this technique is widely studied in the field of computer vision, there are few methods for remote sensing images. Prevalent few-shot semantic segmentation methods can achieve remarkable results for natural images, but they are difficult to apply to remote sensing image processing because existing methods rarely take into consideration the large scale and resolution differences in remote sensing images. Consequently, it is hard for them to obtain correct semantic guidance from few annotated remote sensing images. To tackle these problems, this article proposes the pyramid correlation fusion network (PCFNet) to promote the ability to mine helpful information by calculating multi-scale pixel-wise semantic correspondence. Particularly, the dual distance correlation (DDC) module is designed to simultaneously compute the cosine similarity and Euclidean distance between query features and support features, producing adequate guidance information to determine the category of each pixel. Moreover, to improve segmentation accuracy for small objects, the scale-aware cross-entropy loss (SACELoss) is introduced to dynamically assign loss weights according to the actual sizes of objects. This enables smaller objects to be assigned larger weight values and thus receive more attention during training. Comprehensive experiments on both the iSAID-5iand DLRSD-5idatasets demonstrate that our method outperforms state-of-the-art few-shot semantic segmentation methods. Our code is available at https://github.com/TinyAway/PCFNet. Shunyi Zheng, Zhi Gao 0005 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Joint Learning of Semantic Segmentation and Height Estimation for Remote Sensing Image Leveraging Contrastive LearningabstractSemantic segmentation and height estimation are two critical tasks in remote sensing scene understanding that are highly correlated with each other. To address both tasks simultaneously, it is natural to consider designing a unified deep learning model that aims to improve performance by jointly learning complementary information among the associated tasks. In this paper, we learn the two tasks jointly under a deep multi-task learning framework and propose two novel objective functions, called cross-task contrastive loss and cross-pixel contrastive loss, respectively, to enhance multi-task learning performance through contrastive learning. Specifically, the cross-task contrastive loss is designed to maximize the mutual information of different task features and enforce the model to learn the consistency between semantic segmentation and height estimation. In addition, our method goes beyond previous approaches that only apply contrastive learning at the instance level. Instead, we design a pixel-wise contrastive loss function that pulls together pixel embeddings belonging to the same semantic class, while pushing apart pixel embeddings from different semantic classes. Furthermore, we find that this semantic-guided contrastive loss simultaneously improves the performance of the height estimation task. Our proposed approach is simple and effective and does not introduce any additional overhead to the model during the testing phase. We extensively evaluate our method on the Vaihingen and Potsdam datasets, and the experimental results demonstrate that our approach significantly outperforms the state-of-the-art methods in both height estimation and semantic segmentation. Zhi Gao 0005, Yongjun Zhang 0002, Ruifang Zhai |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Hashing-Based Deep Metric Learning for the Classification of Hyperspectral and LiDAR DataabstractMultisource remote sensing data provide abundant and complementary information for land cover classification. Existing classification methods mainly focus on designing a multi-stream deep network to extract separate features of each single-source data, then adopting a fusing strategy to combine these extracted features for final classification. However, this kind of method neglects the sample correlation of single-source and cross-source data, which may deliver an unsatisfactory classification result when dealing with high intraclass-variability and low interclass-variability samples. To this end, a novel hashing-based deep metric learning (HDML) method is proposed for hyperspectral images (HSIs) and light detection and ranging (LiDAR) data classification in this paper. First, a two-stream deep network is built to extract the spectral-spatial features of HSI and the elevation features of LiDAR, respectively. To fully use the complementary and correlated information of HSI and LiDAR data, we adopt attention-based feature fusion (AFF) modules to deliver a high-discrimination fused feature both for cross-source and single-source feature fusion. Then, the extracted features are fed into fully connected layers to generate class probabilities, respectively. Different from most existing methods that only utilize semantic information of samples, we elaborately designed a loss function to simultaneously consider the label-based semantic loss and hashing-based metric loss. Finally, a decision-level fusion strategy is adopted to further improve the classification results. Extensive experiments on three public HSI and LiDAR data sets demonstrate the effectiveness of the proposed method over some state-of-the-art approaches. Zhi Gao 0005, Leyuan Fang, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Image Deblurring With Image BlurringabstractDeep learning (DL) based methods for motion deblurring, taking advantage of large-scale datasets and sophisticated network structures, have reported promising results. However, two challenges still remain: existing methods usually perform well on synthetic datasets but cannot deal with complex real-world blur, and in addition, over- and under-estimation of the blur will result in restored images that remain blurred and even introduce unwanted distortion. We propose a motion deblurring framework that includes a Blur Space Disentangled Network (BSDNet) and a Hierarchical Scale-recurrent Deblurring Network (HSDNet) to address these issues. Specifically, we train an image blurring model to facilitate learning a better image deblurring model. Firstly, BSDNet learns how to separate the blur features from blurry images, which is adaptable for blur transferring, dataset augmentation, and ultimately directing the deblurring model. Secondly, to gradually recover sharp information in a coarse-to-fine manner, HSDNet makes full use of the blur features acquired by BSDNet as a priori and breaks down the non-uniform deblurring task into various subtasks. Moreover, the motion blur dataset created by BSDNet also bridges the gap between training images and actual blur. Extensive experiments on real-world blur datasets demonstrate that our method works effectively on complex scenarios, resulting in the best performance that significantly outperforms many state-of-the-art approaches. Ziyao Li, Zhi Gao 0005, Han Yi, Boan Chen |
IEEE Trans. Image Process. | 2 |
| 2023 | Synergizing Low Rank Representation and Deep Learning for Automatic Pavement Crack DetectionabstractDue to the critical role of pavement crack detection for road maintenance and eventually ensuring safety, remarkable efforts have been devoted to this research area, and such a trend is further intensified for the coming unmanned vehicle era. However, such crack detection task still remains unexpectedly challenging in practice since the appearance of both cracks and the background are diverse and complex in real scenarios. In this work, we propose an automatic pavement crack detection method via synergizing low rank representation (LRR) and deep learning techniques. First, leveraging LRR which facilitates anomaly detection without making any specific assumption, we can easily discriminate most of the frames with cracks from the long sequence with a consistent pavement base, followed by a straightforward algorithm to localize the cracks. In order to achieve the intelligence of detecting cracks with different pavement basis under unconstrained imaging conditions, we resort to deep learning techniques and propose a deep convolutional neural network for crack detection leveraging on multi-level features and atrous spatial pyramid pooling (ASPP). We train this network based on the training data obtained in the previous stage in an end-to-end manner. Extensive experiments on a wide range of pavements demonstrate the high performance in terms of both accuracy and automaticity. Moreover, the dataset generated by us is much more extensive and challenging than public ones. We put it online athttps://gaozhinuswhu.comto benefit the community. Zhi Gao 0005, Min Cao 0001, Ziyao Li, Kangcheng Liu, Ben M. Chen |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Weakly Supervised 3D Scene Segmentation with Region-Level Boundary Awareness and Instance Discrimination
Kangcheng Liu, Yuzhi Zhao, Qiang Nie, Zhi Gao 0005, Ben M. Chen |
ECCV (28) | 4 |
| 2022 | WeakLabel3D-Net: A Complete Framework for Real-Scene LiDAR Point Clouds Weakly Supervised Multi-Tasks UnderstandingabstractExisting state-of-the-art 3D point clouds understanding methods only perform well in a fully supervised manner. To the best of our knowledge, there exists no unified framework which simultaneously solves the downstream high-level understanding tasks, especially when labels are extremely limited. This work presents a general and simple framework to tackle point clouds understanding when labels are limited. We propose a novel unsupervised region expansion based clustering method for generating clusters. More importantly, we innovatively propose to learn to merge the over-divided clusters based on the local low-level geometric property similarities and the learned high-level feature similarities supervised by weak labels. Hence, the true weak labels guide pseudo labels merging taking both geometric and semantic feature correlations into consideration. Finally, the self-supervised data augmentation optimization module is proposed to guide the propagation of labels among semantically similar points within a scene. Experimental Results demonstrate that our framework has the best performance among the three most important weakly supervised point clouds understanding tasks including semantic segmentation, instance segmentation, and object detection even when limited points are labeled. Kangcheng Liu, Yuzhi Zhao, Zhi Gao 0005, Ben M. Chen |
ICRA | 3 |
| 2022 | Discriminative Feature Extraction and Fusion for Classification of Hyperspectral and Lidar DataabstractMultisource remote sensing data provide the abundant and complementary information for land cover classification. In this paper, we propose a deep hashing-based feature extraction and fusion framework for joint classification of hyper-spectral and LiDAR data. Firstly, HSIs and LiDAR data are fed into a two-stream network to extract deep features after data preprocessing. Then, we adopt hashing technique to constrain single-source and cross-source similarities, i.e., samples with same classes should have small feature distance and samples with different classes should have large feature distance. Furthermore, a feature-level fusion strategy is exploited to fuse the two kind of multisource information. Finally, we design an object function to consider the similarity information between sample pairs and semantic information of each sample, which can deliver the discriminative features for classification. The experiments on Houston data demonstrate the effectiveness of the proposed method over some competitive approaches. Zhi Gao 0005, Yongjun Zhang 0002 |
IGARSS | 2 |
| 2022 | Few-Shot Scene Classification Using Auxiliary Objectives and Transductive InferenceabstractFew-shot learning features the capability of generalizing from very few examples. To realize few-shot scene classification of optical remote sensing images, we propose a two-stage framework that first learns a general-purpose representation and then propagates knowledge in a transductive paradigm. Concretely, the first stage jointly learns a semantic class prediction task as well as two auxiliary objectives in a multi-task model. Therein, rotation prediction estimates the 2D transformation of an input, and contrastive prediction aims to pull together the positive pairs while pushing apart the negative pairs. The second stage aims to find an expected prototype having the minimal distance to all samples within the same class. Particularly, label propagation is applied to make joint prediction for both labeled and unlabeled data. Then the labeled set is expanded by those pseudo-labeled samples, thereby forming a rectified prototype to perform nearest-neighbor classification better. Extensive experiments on standard benchmarks including NWPU-RESISC45, AID, and WHU-RS-19 demonstrate that our method works effectively and achieves the best performance that significantly outperforms many state-of-the-art approaches. Zhi Gao 0005, Can Li 0016, Jinqiang Cui |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Rethinking Monocular Height Estimation From a Classification Task Perspective Leveraging the Vision TransformerabstractHeight estimation from a single remote sensing image has great potential in generating digital surface models (DSM) efficiently for a quick earth surface reconstruction. Recently, convolutional neural networks (CNN) have emerged as a powerful method to deal with this ill-posed problem. Most existing methods formulate height estimation as a regression problem due to the continuity of object height. However, it is difficult for the model to regress the object heights exactly to the ground-truth values with a wide range. In this letter, we reformulate the height estimation task as a classification task to improve the model performance. Specifically, we discretize the continuous ground-truth height into bins and assign each pixel to a single label according to the bin subdivision. In addition, we propose to generate a unique bin subdivision for each input image adaptively by viewing the bin generation as a set-to-set problem. Compared with the fixed bin subdivision method, a specific bin subdivision for each input image makes the model adaptively focus on the height range that is more probable to occur in the scene of the input image. In our experiments, we qualitatively and quantitatively demonstrate that the proposed method outperforms the state-of-the-art approaches on both Vaihingen and Potsdam datasets. Mingchun Lin, Ruifang Zhai, Zhi Gao 0005 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | Few-Shot Scene Classification of Optical Remote Sensing Images Leveraging Calibrated Pretext TasksabstractSmall data holds big AI potential. As one of the promising small data AI approaches, few-shot learning has the goal to learn a model efficiently that can recognize novel classes with extremely limited training samples. Therefore it is critical to accumulate useful prior knowledge obtained from large-scale base class dataset. To realize few-shot scene classification of optical remote sensing images, we start from a baseline model that trains all base classes using a standard cross-entropy loss leveraging two auxiliary objectives to capture intrinsical characteristics across the semantic classes. Specifically, rotation prediction learns to recognize the 2D rotation of an input to guide the learning of class-transferable knowledge, and contrastive learning aims to pull together the positive pairs while pushing apart the negative pairs to promote intra-class consistency and inter-class inconsistency. We jointly optimize such two pretext tasks and semantic class prediction task in an end-to-end manner. To further overcome the overfitting issue, we introduce a regularization technique, adversarial model perturbation, to calibrate the pretext tasks so as to enhance the generalization ability. Extensive experiments on public remote sensing benchmarks including NWPU-RESISC45, AID, and WHU-RS-19 demonstrate that our method works effectively and achieves best performance that significantly outperforms many state-of-the-art approaches. Zhi Gao 0005, Yongjun Zhang 0002, Can Li 0016, Tiancan Mei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Robust Camera Distortion Calibration via Unified RPC Model for Optical Remote Sensing SatellitesabstractOn-orbit geometric calibration (GC) is always performed to compensate for geometric distortion from the satellite’s camera. However, the traditional GC method is complex and difficult to apply broadly due to its reliance on a rigorous physical model (RPM), which involves not only complex processing of attitude, orbit, and time, but also the transformation among multiple coordinate systems. Additionally, the RPM is closely related to designs of satellites and cameras, which increases the complexity of the GC, and thus reduces generalizability. This paper proposes a practical and robust GC method to prevent camera distortion based on a standardized rational polynomial coefficient (RPC) model. Through a series of innovations, including stepwise optimization, a priori gross error elimination, the adjustment model with angular resolution, and the correction for the bias field-of-view (FOV) distortion, we were able to achieve robust GC for camera distortion, as well as accurate splicing and registration among segmented images. Method validation using data from the linear-array camera of the ZiYuan3-02 satellite, and the area-array camera of the GaoFen-4 satellite, produced satisfactory results, indicating that our method effectively compensates for systematic geometric distortion such that consistent GC and accuracy comparable with that of traditional RPM-based methods can be obtained. Ying-Dong Pi, Mi Wang, Zhi Gao 0005 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Asymmetric Hash Code Learning for Remote Sensing Image RetrievalabstractRemote sensing image retrieval (RSIR), aiming at searching for a set of similar items to a given query image, is a very important task in remote sensing applications. Deep hashing learning as the current mainstream method has achieved satisfactory retrieval performance. On one hand, various deep neural networks are used to extract semantic features of remote sensing images. On the other hand, the hashing techniques are subsequently adopted to map the high-dimensional deep features to the low-dimensional binary codes. This kind of method attempts to learn one hash function for both the query and database samples in a symmetric way. However, with the number of database samples increasing, it is typically time-consuming to generate the hash codes of large-scale database images. In this article, we propose a novel deep hashing method, named asymmetric hash code learning (AHCL), for RSIR. The proposed AHCL generates the hash codes of query and database images in an asymmetric way. In more detail, the hash codes of query images are obtained by binarizing the output of the network, while the hash codes of database images are directly learned by solving the designed objective function. In addition, we combine the semantic information of each image and the similarity information of pairs of images as supervised information to train a deep hashing network, which improves the representation ability of deep features and hash codes. The experimental results on three public datasets demonstrate that the proposed method outperforms symmetric methods in terms of retrieval accuracy and efficiency. The source code is available athttps://github.com/weiweisong415/Demo_AHCL_for_TGRS2022. Zhi Gao 0005, Renwei Dian, Pedram Ghamisi, Yongjun Zhang 0002, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Quadratic Pose Estimation Problems: Globally Optimal Solutions, Solvability/Observability Analysis, and Uncertainty DescriptionabstractPose estimation problems are fundamental in robotics. Most of these problems are challenging due to the nonconvex nature. This also sets up an obstacle for uncertainty description that is essential for pose integration and quality control. In this article, we show that a large class of related problems can be categorized as the quadratic pose estimation problems (QPEPs) and we propose a general quaternion-based mathematical model to unify these problems. To solve the nonconvex QPEPs, a Gröbner-basis method is investigated to derive their globally optimal and robust solutions. Furthermore, we develop the rules for characterizing the solvability and observability of these solutions. In addition, the uncertainty description, i.e., covariance matrix, as an important piece of information in robotic state estimation frameworks, is analyzed in detail. Theoretical results show that the covariance can be estimated via online optimization, in an efficient and unbiased manner. In this way, both the solution and covariance are guaranteed to be globally optimal. Through simulations and experiments, we show that the proposed QPEP-based solver is not only accurate, robust, and efficient but outperforms the representatives for covariance estimation. The designed algorithms are also assembled as a C++/MATLAB/Octave/ROS library, while these developed interfaces are built for main stream platforms and simultaneous localization and mapping schemes. Jin Wu 0002, Yu Zheng 0001, Zhi Gao 0005, Yi Jiang 0007, Xiangcheng Hu, Yilong Zhu, Jianhao Jiao, Ming Liu 0001 |
IEEE Trans. Robotics | 3 |
| 2021 | FG-Conv: Large-Scale LiDAR Point Clouds Understanding Leveraging Feature Correlation Mining and Geometric-Aware ModelingabstractThis work presents a general deep learning framework for large-scale point clouds understanding without voxelizations, called FG-Conv, which achieves an accurate and real-time understanding of point clouds. Through our novel design combining feature level correlation mining and deformable convolutions based geometric aware modeling, the local feature relationships and geometric patterns can be captured. The attention mechanism is also adopted to enhance the global long-range feature correlations. Finally, the feature pyramid residual learning network is proposed to combine patterns at different resolutions in a memory-efficient way. Extensive experiments on real-world challenging datasets demonstrated that our approaches outperform state-of-the-art methods in terms of accuracy and efficiency. Weakly supervised transfer learning demonstrates the generalization capacity of our methods. Kangcheng Liu, Zhi Gao 0005, Feng Lin 0003, Ben M. Chen |
ICRA | 2 |
| 2021 | IPMGAN: Integrating physical model and generative adversarial network for underwater image enhancement
Xiaodong Liu 0008, Zhi Gao 0005, Ben M. Chen |
Neurocomputing | 2 |
| 2021 | Single Image Deraining Integrating Physics Model and Density-Oriented Conditional GAN RefinementabstractAlthough advanced single image deraining methods have been proposed, their generalization ability to real-world images is usually limited, especially when dealing with rain patterns of different densities, shapes, and directions. In order to improve the robustness and generalization of these deraining methods, we propose a novel density-aware single image deraining method with gated multi-scale feature fusion, which consists of two stages. In the first stage, a sophisticated physics model is leveraged for initial deraining and a network branch is utilized for rain density estimation to guide the subsequent refinement. The second stage of model-independent refinement is realized using conditional Generative Adversarial Network (cGAN), attempting to eliminate artifacts and improve the restoration quality. Extensive experiments have been conducted on the representative synthetic rain datasets and real rain scenes, demonstrating the superiority of our method in terms of effectiveness and generalization ability, which outperforms the state-of-the-arts. Min Cao 0001, Zhi Gao 0005, Bharath Ramesh 0001, Tiancan Mei, Jinqiang Cui |
IEEE Signal Process. Lett. | 2 |
| 2021 | A Two-Stage Density-Aware Single Image Deraining MethodabstractAlthough advanced single image deraining methods have been proposed, one main challenge remains: the available methods usually perform well on specific rain patterns but can hardly deal with scenarios with dramatically different rain densities, especially when the impacts of rain streaks and the veiling effect caused by rain accumulation are heavily coupled. To tackle this challenge, we propose a two-stage density-aware single image deraining method with gated multi-scale feature fusion. In the first stage, a realistic physics model closer to real rain scenes is leveraged for initial deraining, and a network branch is also trained for rain density estimation to guide the subsequent refinement. The second stage of model-independent refinement is realized using conditional Generative Adversarial Network (cGAN), aiming to eliminate artifacts and improve the restoration quality. In particular, dilated convolutions are applied to extract rain features at multiple scales and gated feature fusion is exploited to better aggregate multi-level contextual information in both stages. Extensive experiments have been conducted on representative synthetic rain datasets and real rain scenes. Quantitative and qualitative results demonstrate the superiority of our method in terms of effectiveness and generalization ability, which outperforms the state-of-the-art. Min Cao 0001, Zhi Gao 0005, Bharath Ramesh 0001, Tiancan Mei, Jinqiang Cui |
IEEE Trans. Image Process. | 2 |
| 2020 | Small Object Detection Leveraging on Simultaneous Super-resolutionabstractDespite the impressive advancement achieved in object detection, the detection performance of small object is still far from satisfactory due to the lack of sufficient detailed appearance to distinguish it from similar objects. Inspired by the positive effects of super-resolution for object detection, we propose a framework that can be incorporated with detector networks to improve the performance of small object detection, in which the low-resolution image is super-resolved via generative adversarial network (GAN) in an unsupervised manner. In our method, the super-resolution network and the detection network are trained jointly. In particular, the detection loss is back-propagated into the super-resolution network during training to facilitate detection. Compared with available simultaneous super-resolution and detection methods which heavily rely on low-/high-resolution image pairs, our work breaks through such restriction via applying the CycleGAN strategy, achieving increased generality and applicability, while remaining an elegant structure. Extensive experiments on datasets from both computer vision and remote sensing communities demonstrate that our method obtains competitive performance on a wide range of complex scenarios. Zhi Gao 0005, Xiaodong Liu 0008, Yongjun Zhang 0002, Tiancan Mei |
ICPR | 2 |
| 2020 | A Target Tracking and Positioning Framework for Video Satellites Based on SLAMabstractWith the booming development in aerospace technology, the video satellite which observes the live phenomena on the ground by video shooting has gradually emerged as a new Earth observation method. And remote sensing comes into a "dynamic" era with the demand for new processing techniques, especially the near-real-time tracking and geo-positioning algorithm for ground moving targets. However, many researchers merely extract pixel-level trajectories in post-processed video products, resulting in fairly limited applications. We regard the video satellite as a robot flying in space and adopt the SLAM framework for the positioning of ground moving targets. The designed framework is based on the representative ORB-SLAM and we make improvements mainly in feature extraction, satellite pose estimation, moving target tracking and positioning. We coordinate a moving fishing boat with GPS-RTK (Real-time Kinematic) devices and a video satellite observing it simultaneously for verification and evaluation of our method. Experiments demonstrate that our framework provides reasonable geolocation of the moving target in satellite videos. Finally, some open problems and potential research directions are discussed. Zhi Gao 0005, Yongjun Zhang 0002, Ben M. Chen |
IROS | 2 |
| 2020 | Vehicle Detection in Remote Sensing Images Leveraging on Simultaneous Super-ResolutionabstractOwing to the relatively small size of vehicles in remote sensing images, lacking sufficient detailed appearance to distinguish vehicles from similar objects, the detection performance is still far from satisfactory compared with the detection results on everyday images. Inspired by the positive effects of super-resolution convolutional neural network (SRCNN) for object detection and the stunning success of deep CNN techniques, we apply generative adversarial network frameworks to realize simultaneous SRCNN and vehicle detection in an end-to-end manner, and the detection loss is backpropagated into the SRCNN during training to facilitate detection. In particular, our work is unsupervised and bypasses the requirement of low-/high-resolution image pairs during the training stage, achieving increased generality and applicability. Extensive experiments on representative data sets demonstrate that our method outperforms the state-of-the-art detectors. (The source code will be made available after the review process). Zhi Gao 0005, Tiancan Mei, Bharath Ramesh 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | MLFcGAN: Multilevel Feature Fusion-Based Conditional GAN for Underwater Image Color CorrectionabstractColor correction for underwater images has received increasing interest, due to its critical role in facilitating available mature vision algorithms for underwater scenarios. Inspired by the stunning success of deep convolutional neural network (DCNN) techniques in many vision tasks, especially the strength in extracting features in multiple scales, we propose a deep multiscale feature fusion net based on the conditional generative adversarial network (GAN) for underwater image color correction. In our network, multiscale features are extracted first, followed by augmenting local features in each scale with global features. This design was verified to facilitate more effective and faster network learning, resulting in better performance in both color correction and detail preservation. We conducted extensive experiments and compared the results with state-of-the-art approaches quantitatively and qualitatively, showing that our method achieves significant improvements. Xiaodong Liu 0008, Zhi Gao 0005, Ben M. Chen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Improved Faster R-CNN With Multiscale Feature Fusion and Homography Augmentation for Vehicle Detection in Remote Sensing ImagesabstractVehicle detection in remote sensing images has attracted remarkable attention for its important role in a variety of applications in traffic, security, and military fields. Motivated by the stunning success of region convolutional neural network (R-CNN) techniques, which have achieved the state-of-the-art performance in object detection task on benchmark data sets, we propose to improve the Faster R-CNN method with better feature extraction, multiscale feature fusion, and homography data augmentation to realize vehicle detection in remote sensing images. Extensive experiments on representative remote sensing data sets related to vehicle detection demonstrate that our method achieves better performance than the state-of-the-art approaches. The source code will be made available (after the review process). Zhi Gao 0005, Tiancan Mei |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | A novel framework for robust long-term object tracking in real-time
Xiao-Xu Zheng, Bharath Ramesh 0001, Zhi Gao 0005, Cheng Xiang 0001 |
Mach. Vis. Appl. | 3 |
| 2019 | Scalable scene understanding via saliency consensus
Bharath Ramesh 0001, Lim Zhi Jian Nicholas, Cheng Xiang 0001, Zhi Gao 0005 |
Soft Comput. | 5 |
| 2019 | Transform Learning Based Sparse Coding for LiDAR Data DenoisingabstractIn recent years, sparse coding (SC) has been exploited for light detection and ranging (LiDAR) data restoration, and promising results have been reported. However, such methods are usually time consuming, because much computational resource has been devoted to solving batches of ℓ0or ℓ1-norm optimization problems iteratively. More recently, fast SC method has been proposed to achieve nearly real-time performance at the expense of applicability. In this letter, we propose a transform learning based SC method for LiDAR data denoising. Moreover, we present a detailed evaluation for our series of SC methods with different models, and together with concluding remarks. Such remarks can be applied as a guide to apply appropriate SC model for specific application. Zhi Gao 0005 |
IEEE Signal Process. Lett. | 1 |
| 2019 | A Coarse-to-Fine Framework for Cloud Removal in Remote Sensing Image SequenceabstractClouds and accompanying shadows, which exist in optical remote sensing images with high possibility, can degrade or even completely occlude certain ground-cover information in images, limiting their applicabilities for Earth observation, change detection, or land-cover classification. In this paper, we aim to deal with cloud contamination problems with the objective of generating cloud-removed remote sensing images. Inspired by low-rank representation together with sparsity constraints, we propose a coarse-to-fine framework for cloud removal in the remote sensing image sequence. Leveraging on group-sparsity constraint, we first decompose the observed cloud image sequence of the same area into the low-rank component, group-sparse outliers, and sparse noise, corresponding to cloud-free land-covers, clouds (and accompanying shadows), and noise respectively. Subsequently, a discriminative robust principal component analysis (RPCA) algorithm is utilized to assign aggressive penalizing weights to the initially detected cloud pixels to facilitate cloud removal and scene restoration. Moreover, we incorporate geometrical transformation into a low-rank model to address the misalignment of the image sequence. Significantly superior to conventional cloud-removal methods, neither cloud-free reference image(s) nor additional operations of cloud and shadow detection are required in our method. Extensive experiments on both simulated data and real data demonstrate that our method works effectively, outperforming many state-of-the-art approaches. Yongjun Zhang 0002, Fei Wen 0004, Zhi Gao 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Long-term object tracking with a moving event camera
Bharath Ramesh 0001, Zhi Wei Lee, Zhi Gao 0005, Garrick Orchard, Cheng Xiang 0001 |
BMVC | 4 |
| 2018 | Visual adaptive tracking for monocular omnidirectional camera
Yazhe Tang, Zhi Gao 0005, Feng Lin 0003, Youfu Li 0001, Fei Wen 0004 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Two-Pass Robust Component Analysis for Cloud Removal in Satellite Image SequenceabstractDue to the inevitable existence of clouds and their shadows in optical remote sensing images, certain ground-cover information is degraded or even appears to be missing, which limits analysis and utilization. Thus, cloud removal is of great importance to facilitate downstream applications. Motivated by the sparse representation techniques which have obtained a stunning performance in a variety of applications, including target detection, anomaly detection, and so on; we propose a two-pass robust principal component analysis (RPCA) framework for cloud removal in the satellite image sequence. First, a plain RPCA is applied for initial cloud region detection, followed by a straightforward morphological operation to ensure that the cloud region is completely detected. Subsequently, a discriminative RPCA algorithm is proposed to assign aggressive penalizing weights to the detected cloud pixels to facilitate cloud removal and scene restoration. Significantly superior to currently available methods, neither a cloud-free reference image nor a specific algorithm of cloud detection is required in our method. Experiments on both simulated and real images yield visually plausible and numerically verified results, demonstrating the effectiveness of our method. Fei Wen 0004, Yongjun Zhang 0002, Zhi Gao 0005 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Synergizing Appearance and Motion With Low Rank Representation for Vehicle Counting and Traffic Flow AnalysisabstractAppearance and motion, which are complementary, account for a dominant proportion of visual information. We propose to synergize them using a low-rank representation framework for the estimation and analysis of traffic flow. Taking advantage of the downward-looking camera configuration, we do the processing only on the measure line, called virtual gantry, instead of dealing with the whole frame, resulting in much improved efficiency. Enforcing the low-rank constraint on the spatiotemporal image which is generated via stacking pixels on virtual gantry over time, we introduce the block-sparse robust principal component analysis algorithm, in which the motion cue is leveraged to highlight the foreground and realize vehicle detection with high accuracy. The motion flow is further exploited for size normalization to classify vehicles into lite, small, medium, and large categories. Benefiting from the low-rank representation, our method is parameter insensitive, robust to illumination changes, and requires no training. We perform extensive experiments on the 24/7 videos collected over the highways in China and Singapore, obtaining nearly 100% accuracy. Meanwhile, insightful observations on the obtained traffic information are given, which could be very valuable to the users, especially to the traffic management sectors. Zhi Gao 0005, Ruifang Zhai, Pengfei Wang 0011, Hailong Qin, Yazhe Tang, Bharath Ramesh 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | A brief survey of visual odometry for micro aerial vehiclesabstractRecently, visual odometry (VO) has experienced a rapid growth, which makes it viable for a range of applications. This survey paper attempts to provide a timely and comprehensive review of this field, focusing specifically on micro aerial vehicles (MAVs), with monocular, stereo or RGB-D cameras onboard. In this survey, the milestones in the development of VO will be reviewed, followed by an illustration of its general workflow, the commonly used datasets. The survey is concluded by an overall discussion. Mo Shan, Yingcai Bi, Hailong Qin, Zhi Gao 0005, Feng Lin 0003, Ben M. Chen |
IECON | 5 |
| 2016 | Adaptive and Robust Sparse Coding for Laser Range Data Denoising and InpaintingabstractSparse coding (SC) is making a significant impact in computer vision and signal processing communities, which achieves the state-of-the-art performance in a variety of applications for images, e.g., denoising, restoration, and synthesis. We propose an adaptive and robust SC algorithm exploiting the characteristics of typical laser range data and the availability of both range and reflectance data to realize range data denoising and inpainting. Specifically, our method estimates the informative level of each patch according to the variation in both range and reflectance modalities, followed by adaptive dictionary training that assigns dynamic sparsity weights to the patches with different informative levels. Furthermore, the l1-norm-based representation fidelity measure is applied to make our method robust to outliers which are common in laser range measurements. Extensive experiments on synthetic and real data demonstrate that our method works effectively, resulting in superior performance both visually and quantitatively, compared with competitive methods including the available sparse-representation-based algorithm, wavelets, partial differential equation, and non-local means. Zhi Gao 0005, Qingquan Li 0001, Ruifang Zhai, Mo Shan, Feng Lin 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Laser Range Data Denoising via Adaptive and Robust Dictionary LearningabstractSparse representation (SR) is making significant impact in the computer vision and signal processing communities due to its stunning performance in a variety of applications for images, e.g., denoising, restoration, and synthesis. We propose an adaptive and robust SR algorithm that exploits the characteristics of typical laser range data, i.e., the availability of both range and reflectance data, to realize range data denoising. Specifically, our method estimates the informative level (IL) of each patch according to the variation in both range and reflectance modalities, followed by adaptive dictionary training that assigns dynamic sparsity weights to the patches with different ILs. Furthermore, the l1-norm-based representation fidelity measure is applied to make our method robust to outliers, which are common in laser range measurements. Extensive experiments on synthesized and actual data demonstrate that our method works effectively, resulting in superior performance both visually and quantitatively. Zhi Gao 0005, Qingquan Li 0001, Ruifang Zhai, Feng Lin 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | Adaptive Sparse Coding for Painting Style Analysis
Zhi Gao 0005, Mo Shan, Loong Fah Cheong, Qingquan Li 0001 |
ACCV (2) | 1 |
| 2014 | Block-Sparse RPCA for Salient Motion DetectionabstractRecent evaluation [2], [13] of representative background subtraction techniques demonstrated that there are still considerable challenges facing these methods. Challenges in realistic environment include illumination change causing complex intensity variation, background motions (trees, waves, etc.) whose magnitude can be greater than those of the foreground, poor image quality under low light, camouflage, etc. Existing methods often handle only part of these challenges; we address all these challenges in a unified framework which makes little specific assumption of the background. We regard the observed image sequence as being made up of the sum of a low-rank background matrix and a sparse outlier matrix and solve the decomposition using the Robust Principal Component Analysis method. Our contribution lies in dynamically estimating the support of the foreground regions via a motion saliency estimation step, so as to impose spatial coherence on these regions. Unlike smoothness constraint such as MRF, our method is able to obtain crisply defined foreground regions, and in general, handles large dynamic background motion much better. Furthermore, we also introduce an image alignment step to handle camera jitter. Extensive experiments on benchmark and additional challenging data sets demonstrate that our method works effectively on a wide range of complex scenarios, resulting in best performance that significantly outperforms many state-of-the-art approaches. Zhi Gao 0005, Loong Fah Cheong, Yu-Xiang Wang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Quasi-Parallax for Nearly Parallel Frontal Eyes - A Possible Role of Binocular Overlap During Rapid Locomotion
Loong Fah Cheong, Zhi Gao 0005 |
Int. J. Comput. Vis. | 2 |
| 2012 | Block-Sparse RPCA for Consistent Foreground Detection
Zhi Gao 0005, Loong Fah Cheong, Mo Shan |
ECCV (5) | 1 |