EDBT 2026 Demo / reviewers in the wild / expert
Xin Wu 0001
dblp:13/5235-1
· DBLP profile ↗
30ranked-venue papers
11as first author
23since 2021 · last 2025
0000-0002-1733-3560ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 10 first-author · 14 since 2021Computer networks · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DMSF: A Dynamic Model Splitting Framework for Edge-Cloud Collaborative InferenceabstractEdge-cloud collaborative inference is a widely adopted approach in edge intelligence, especially for latency-sensitive tasks such as drone inspection, augmented reality, and disaster relief. To adapt to varying network bandwidth and limited computational capacity, model splitting methods divide Deep Neural Networks (DNNs) at specific split points into edge and cloud segments. However, current methods suffer from high switch latency and memory overhead, and are incompatible with Feature Pyramid Network (FPN)-based multi-branch models. To address the above issues, we propose a dynamic model splitting framework (DMSF) for edge-cloud collaborative inference. We design switchable compress–recover modules after each layer, enabling a single model to support all split points and significantly reduce memory overhead. The optimal split point is selected based on bandwidth and edge computational capacity, and the corresponding compress–recover path is activated for fast switching. For FPN-based multi-branch models, we remove the multi-scale branches from the backbone to the FPN and present hierarchical compensation modules in DMSF, making the model more suitable for splitting while maintaining perception accuracy. Experimental results show that DMSF reduces memory overhead by 75%, achieves millisecond-level switching, and outperforms the existing method with 32.0% lower latency and 12.2% higher perception accuracy. Xinyun Zhang 0002, Li Wang 0039, Xin Wu 0001, Lianming Xu, Yingyan Hou, Aiguo Fei |
GLOBECOM | 3 |
| 2025 | DULRTC-RME: A Deep Unrolled Low-rank Tensor Completion Network for Radio Map EstimationabstractRadio maps enrich radio propagation and spectrum occupancy information, which provides fundamental support for the operation and optimization of wireless communication systems. Traditional radio maps are mainly achieved by extensive manual channel measurements, which is time-consuming and inefficient. To reduce the complexity of channel measurements, radio map estimation (RME) through novel artificial intelligence techniques has emerged to attain higher resolution radio maps from sparse measurements or few observations. However, black box problems and strong dependency on training data make learning-based methods less explainable, while model-based methods offer strong theoretical grounding but perform inferior to the learning-based methods. In this paper, we develop a deep unrolled low-rank tensor completion network (DULRTC-RME) for radio map estimation, which integrates theoretical interpretability and learning ability by unrolling the tedious low-rank tensor completion optimization into a deep network. It is the first time that algorithm unrolling technology has been used in the RME field. Experimental results demonstrate that DULRTC-RME outperforms existing RME methods. Xin Wu 0001, Lianming Xu, Na Liu 0014, Li Wang 0039 |
ICASSP | 2 |
| 2025 | Real-Time Anomaly Detection of Electricity Time Series Data Based on Future-Guidance NetworkabstractElectricity data plays a pivotal role in power management systems. Smart meters, as key tools for recording this data, often encounter anomalies due to meter malfunctions, operational errors, or unauthorized electricity usage, all of which jeopardize the stability of power grids. To this end, we propose the future-guidance anomaly detection network, called FG-Net, designed for real-time analysis of electricity time series data. FG-Net is designed to memorize historical data and assimilate future data, ensuring comprehensive learning of complete data information. Specifically, we leverage the comprehensive data insights gained from a complete information network to guide the predictions of the historical information network. Subsequently, we developed a self-matching feature guidance (SFG) strategy that harnesses the strengths of the complete information network to offset the limitations of the historical information network, thus providing effective guidance. The experimental results on two power grid time series datasets with different anomaly volatility, the Low Carbon London dataset and the Ausgrid Solar Home dataset, demonstrate the proposed anomaly detection method’s accuracy and efficiency. Yilu Shi, Lianming Xu, Xin Wu 0001, Li Wang 0039, Yingyan Hou |
IEEE Signal Process. Lett. | 4 |
| 2025 | DHANet: Dual-Stream Hierarchical Interaction Networks for Multimodal Drone Object DetectionabstractDrone-based remote sensing has become pivotal for high-resolution dynamic monitoring. However, the differences between day and night modes will trigger a mismatch in multi-scale object features under extreme lighting conditions. In this paper, we propose a dual-stream hierarchical interaction network for multimodal drone object detection, called DHANet, which enhances the distinguishability between multi-scale objects and background for each modality. Specifically, DHANet is designed with a Modality-Adaptive Asymmetric Attention Module (M-AAM) that enhances object-level semantic representations through global and local attention mechanisms. The M-AAM employs global context attention and local positional attention to replace conventional multi-scale context extraction, thereby effectively integrating spatial-channel information of objects. Furthermore, the network is equipped with a Multimodal Scale-Attentive Convolution (M-SC) module that dynamically generates modality-specific feature aggregation weights. This design enables global cross-modality information fusion while reducing computational complexity. Experimental results on two multimodal remote sensing benchmark datasets (DroneVehicle and VEDAI) and two natural datasets (LLVIP and FLIR) demonstrate the robustness and generalizability of DHANet. The codes will be openly and freely available at https://github.com/Victoria-xin1009/-IEEE TGRS DHANet for the sake of reproducibility. Xin Wu 0001, Li Wang 0039, Haoyang Ji, Lianming Xu, Yingyan Hou, Aiguo Fei |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge DevicesabstractDistributed DNN inference is becoming increasingly important as the demand for intelligent services at the network edge grows. By leveraging the power of distributed computing, edge devices can perform complicated and resource-hungry inference tasks previously only possible on powerful servers, enabling new applications in areas such as autonomous vehicles, industrial automation, and smart homes. However, it is challenging to achieve accurate and efficient distributed edge inference due to the fluctuating nature of the actual resources of the devices and the processing difficulty of the input data. In this work, we propose DistrEE, a distributed DNN inference framework that can exit model inference early to meet specific quality of service requirements. In particular, the framework firstly integrates model early exit and distributed inference for multi-node collaborative inferencing scenarios. Furthermore, it designs an early exit policy to control when the model inference terminates. Extensive simulation results demonstrate that DistrEE can efficiently realize efficient collaborative inference, achieving an effective trade-off between inference latency and accuracy. Xian Peng, Xin Wu 0001, Lianming Xu, Li Wang 0039, Aiguo Fei |
GLOBECOM | 2 |
| 2024 | Design of A Collaborative Ranging Aware UE Aggregation Transmission Mechanism for Future IoT NetworksabstractWith the evolution of wireless communication ser-vices, the requirement for reliability and latency is more and more strict, especially in the field of personal consumption area and the Industrial Internet of Things (IIoT), such as autonomous vehicles, remote healthcare, smart factory, augmented reality (AR)/virtual reality (VR), and so on. Moreover, the requirement for reliability and latency is also diverse for different types of communication services. Under the requirement of the enhancement of such hyper reliability and low latency communications (HRLLC) and the fact of diver user equipment capabilities, this paper proposes a collaborative ranging aware UE aggregation transmission mechanism for future IoT networks. Under the proposed mechanism, a ranging aware scheduling (RAS) based commu-nication procedure is designed for reliability enhancement and latency reduction, wherein some UEs or particular higher-end devices play as cooperation nodes (CNs) for packet duplication through multiple paths simultaneously. Considering each CN's system signaling overhead and signal processing capability with a restriction of latency requirement, the scheduling algorithm RAS is performed at each CN. According to the presented simulation results, the research work in this paper can provide insight into the distributed cooperation system for future IoT networks. Ruijie Fang, Tao Chen 0037, Shaofu Lin, Xin Wu 0001 |
WCNC | 6 |
| 2024 | Emergency Computing: An Adaptive Collaborative Inference Method Based on Hierarchical Reinforcement LearningabstractIn achieving effective emergency response, the timely acquisition of environmental information, seamless command data transmission, and prompt decision-making are crucial. This necessitates the establishment of a resilient emergency communication dedicated network, capable of providing communication and sensing services even in the absence of basic infrastructure. In this paper, we propose an Emergency Network with Sensing, Communication, Computation, Caching, and Intelligence (E-SC3I). The framework incorporates mechanisms for emergency computing, caching, integrated communication and sensing, and intelligence empowerment. E-SC3I ensures rapid access to a large user base, reliable data transmission over unstable links, and dynamic network deployment in a changing environment. However, these advantages come at the cost of significant computation overhead. Therefore, we specifically concentrate on emergency computing and propose an adaptive collaborative inference method (ACIM) based on hierarchical reinforcement learning. Experimental results demonstrate our method's ability to achieve rapid inference of AI models with constrained computational and communication resources. Weiqi Fu, Lianming Xu, Xin Wu 0001, Li Wang 0039, Aiguo Fei |
WCNC | 3 |
| 2024 | Unlabeled Data Guided Partial Label Learning for Hyperspectral Image ClassificationabstractIncorrect labeling (i.e., noisy label learning) in HSI classification has attracted so much attention in recent years, which holds the assumption that the given pixels of an HSI may be incorrectly labeled and only one candidate label is required to provide for a typical pixel. However, instead of offering only one candidate label that may be incorrect, partial label learning often provides a candidate label set that contains the ground-truth label for each pixel in an HSI, which is also an essential problem of great practical value and has recently started to attract attention. This paper proposes a novel framework for partial label learning in HSI classification, namely unlabeled data guided partial label learning (UPLL). The proposed framework is an iterative process that can fully exploit the benefits of unlabeled data. Specifically, during each iteration, we conduct the semi-supervised label propagation; the resulting labeling confidence matrices of the original training samples and the unlabeled testing samples are further enhanced by exploiting the spatial information. Then, we select qualified original training samples and unlabeled testing samples with high confident predictions to disambiguate and expand the original training set, leading to a more robust representation of training data. Such phases are repeated until convergence. The comprehensive experiments show the superiority of the proposed UPLL method over the existing state-of-the-art methods. Especially, the classification accuracy improves more than 5% with very few training samples than the second best comparing method. Shujun Yang, Yuheng Jia, Yao Ding 0010, Xin Wu 0001, Danfeng Hong |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Multimodal Collaboration Networks for Geospatial Vehicle Detection in Dense, Occluded, and Large-Scale EventsabstractIn large-scale disaster events, the planning of optimal rescue routes depends on the object detection ability at the disaster scene, with one of the main challenges being the presence of dense and occluded objects. Existing methods, which are typically based on the RGB modality, struggle to distinguish targets with similar colors and textures in crowded environments and are unable to identify obscured objects. To this end, we first construct two multimodal dense and occlusion vehicle detection datasets for large-scale events, utilizing RGB and height map modalities. Based on these datasets, we propose a multimodal collaboration network for dense and occluded vehicle detection, MuDet for short. MuDet hierarchically enhances the completeness of discriminable information within and across modalities and differentiates between simple and complex samples. MuDet includes three main modules: Unimodal Feature Hierarchical Enhancement (Uni-Enh), Multimodal Cross Learning (Mul-Lea), and Hard-easy Discriminative (He-Dis) Pattern. Uni-Enh and Mul-Lea enhance the features within each modality and facilitate the cross-integration of features from two heterogeneous modalities. He-Dis effectively separates densely occluded vehicle targets with significant intra-class differences and minimal inter-class differences by defining and thresholding confidence values, thereby suppressing the complex background. Experimental results on two re-labeled multimodal benchmark datasets, the 4K-SAI-LCS dataset, and the ISPRS Potsdam dataset, demonstrate the robustness and generalization of the MuDet. Xin Wu 0001, Zhanchao Huang, Li Wang 0039, Jocelyn Chanussot, Jiaojiao Tian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Energy-Efficient Computation Offloading and Data Compression for UAV-Mounted MEC NetworksabstractThe advancement of mobile edge computing (MEC) is driving the utilization of mobile devices (MDs) for real-time video detection. Nonetheless, during emergency scenarios, video detection tasks encounter two primary challenges: 1) time-varying channels over time between the MDs and the unmanned aerial vehicle (UAV); 2) the need for precision while managing energy consumption on the MDs. In this paper, we collaboratively address the optimization of computation offloading, data compression, and resource allocation challenges in the context of UAV-mounted MEC networks, where the UAV offers computational support for MDs. To improve accuracy and reduce energy consumption in time-varying environments, we propose an energy-efficient Deep Deterministic Policy Gradient based computation offloading algorithm (DCOA). By pursuing the long-term goal, DCOA learns to adapt to continuously changing conditions and thus improve accuracy and energy efficiency. Additionally, we introduce a convex optimization algorithm using the Lagrange multiplier method to solve resource allocation issues for offloading tasks, further reducing energy usage. Experimental results show that DCOA achieves high accuracy with low energy consumption compared to existing algorithms. Xinyun Zhang 0002, Li Wang 0039, Xin Wu 0001, Lianming Xu, Aiguo Fei |
GLOBECOM | 3 |
| 2023 | TAttMSRecNet: Triplet-attention and multiscale reconstruction network for band selection in hyperspectral images
Utpal Nandi, Swalpa Kumar Roy, Danfeng Hong, Xin Wu 0001, Jocelyn Chanussot |
Expert Syst. Appl. | 4 |
| 2023 | UIU-Net: U-Net in U-Net for Infrared Small Object DetectionabstractLearning-based infrared small object detection methods currently rely heavily on the classification backbone network. This tends to result in tiny object loss and feature distinguishability limitations as the network depth increases. Furthermore, small objects in infrared images are frequently emerged bright and dark, posing severe demands for obtaining precise object contrast information. For this reason, we in this paper propose a simple and effective "U-Net in U-Net" framework, UIU-Net for short, and detect small objects in infrared images. As the name suggests, UIU-Net embeds a tiny U-Net into a larger U-Net backbone, enabling the multi-level and multi-scale representation learning of objects. Moreover, UIU-Net can be trained from scratch, and the learned features can enhance global and local contrast information effectively. More specifically, the UIU-Net model is divided into two modules: the resolution-maintenance deep supervision (RM-DS) module and the interactive-cross attention (IC-A) module. RM-DS integrates Residual U-blocks into a deep supervision network to generate deep multi-scale resolution-maintenance features while learning global context information. Further, IC-A encodes the local context information between the low-level details and high-level semantic features. Extensive experiments conducted on two infrared single-frame image datasets, i.e., SIRST and Synthetic datasets, show the effectiveness and superiority of the proposed UIU-Net in comparison with several state-of-the-art infrared small object detection methods. The proposed UIU-Net also produces powerful generalization performance for video sequence infrared small object datasets, e.g., ATR ground/air video sequence dataset. The codes of this work are available openly at https://github.com/danfenghong/IEEE. Xin Wu 0001, Danfeng Hong, Jocelyn Chanussot |
IEEE Trans. Image Process. | 1 |
| 2022 | Mpanet: Multi-Patch Attention for Infrared Small Target Object DetectionabstractInfrared small target detection (ISTD) has attracted widespread attention and been applied in various fields. Due to the small size of infrared targets and the noise interference from complex backgrounds, the performance of ISTD using convolutional neural networks (CNNs) is restricted. Moreover, the constriant that long-distance dependent features can not be encoded by the vanilla CNNs also impairs the robustness of capturing targets' shapes and locations in complex scenarios. To this end, a multi-patch attention network (MPANet) based on the axial-attention encoder and the multi-scale patch branch (MSPB) structure is proposed. Specially, an axial-attention-improved encoder architecture is designed to highlight the effective features of small targets and suppress background noises. Furthermore, the developed MSPB structure fuses the coarse-grained and fine-grained features from different semantic scales. Extensive experiments on the SIRST dataset show the superiority performance and effectiveness of the proposed MPANet compared to the state-of-the-art methods. Wei Li 0032, Xin Wu 0001, Zhanchao Huang, Ran Tao 0003 |
IGARSS | 3 |
| 2022 | Learning Locality-Constrained Sparse Coding for Spectral Enhancement of Multispectral ImageryabstractOwing to easy acquisition and large coverage from the space, multispectral (MS) imaging has garnered growing interest in various applications of remote sensing. However, the limited spectral information of MS data, to a great extent, leads to difficulties in classifying the materials more accurately, particularly for those classes that have very similar visual appearances. To address this issue effectively, we attempt to enhance the spectral resolution of MS imagery, enabling the identification of materials at a more precise level by the means of richer spectral information. More specifically, we propose to learn locality-constrained sparse coding (LCSC) for short, on partially overlapped hyperspectral (HS)-MS pairs (i.e., dictionary). LCSC is capable of capturing neighboring relations well by enforcing the local constraint for each pixel. Such a strategy makes it possible to better reconstruct HS products from MS images and partially overlapped HS images. Reconstruction and unmixing are explored as potential applications to assess the performance of spectral enhancement. Extensive experiments are conducted on two HS-MS data sets in comparison with several state-of-the-art baselines, which demonstrate the effectiveness of the proposed LCSC algorithm in the task of spectral enhancement. Danfeng Hong, Xin Wu 0001, Lianru Gao, Bing Zhang 0001, Jocelyn Chanussot |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Unifying Label Propagation and Graph Sparsification for Hyperspectral Image ClassificationabstractRecently, graph convolutional network (GCN) has received more and more interest in the field of hyperspectral image classification (HSIC). The existing GCN-based models for HSIC propagate and aggregate information through the GCN network based on the graph, which is constructed according to spatial location or spectral similarity. However, the constructed graph may not be ideal for the downstream classification task due to the variety of spectral characteristics. In this paper, a fully connected graph is adaptively constructed to make full use of local spatial information and global spectral information. Besides, we apply a neural sparsification technique to remove potentially task-irrelevant edges in case of misleading message propagation. Furthermore, label propagation (LP) serves as regularization to assist the graph network in learning proper edge weights that lead to improved classification performance. The resulting network is end-to-end trainable. The experimental results on three popular benchmarks, including Indian Pines, Pavia University, and Kennedy Space Center, demonstrate the superiority of our algorithm. Haojie Hu, Fang He 0012, Fenggan Zhang, Yao Ding 0010, Xin Wu 0001, Jianwei Zhao 0002, Minli Yao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Lightweight Heterogeneous Kernel Convolution for Hyperspectral Image Classification With Noisy LabelsabstractConvolutional neural networks (CNNs) have exhibited commendable performance in the hyperspectral images (HSIs) classification task with manually annotated limited available training data for supervision. The accurate classification of pixel-wise land covers using traditional CNNs is often hampered by the presence of wrong (noisy) labels in the training data and can easily be overfitted to the label noises. However, training on noisy labeled data inevitably suffers from performance degradation since CNNs tend to overfit the label noises. To overcome this problem, we propose a lightweight heterogeneous kernel convolution (HetConv3D) for HSI classification with noisy labels, whereHetConv3Duses two different types of convolutional kernels, i.e., spectral and spatial domains, and fuses them to produce the final feature maps that are less prompted to the noises and also reduces the computation time. The experiments are conducted using three well-known HSI datasets, i.e., Kennedy Space Center (KSC), Salinas Scene (SA), and University of Pavia (UP), and results are compared with traditional supervised classification methods, including support vector machine (SVM), random forest (RF), CNN3D, ContextNet, MS3DNet, lightweight dual-channel residual network (DCRN), and HetConv3DNet. The superior performance exhibited by the proposed modelHetConv3D-HSIconfirms the importance of learning a fusion of spatial and spectral kernel features. The source code will be made available publicly athttps://github.com/purbayankar/HetConv3DNet. Swalpa Kumar Roy, Danfeng Hong, Purbayan Kar, Xin Wu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Infrared Small Object Detection Using Deep Interactive U-NetabstractInfrared objects acquired from a long-distance have small sizes and are easily submerged by a complex and variable background. The existing deep network detection framework suffers greatly from the feature spatial resolution loss caused by the networks’ depth and multiple downsampling operations, which is extremely detrimental for small object detection. So, a crucial and urgent goal is, how to trade-off network depth and feature spatial resolution, while learning feature context representation and interaction to distinguish from the background. To this end, we propose a deep interactive U-Net architecture (short for DI-U-Net) with high feature learning and feature interaction ability. First, feature learning is first achieved through a multi-level and high-resolution network structure. This structure ensures feature resolution as the network depth increase, and also focus on the object’s global context information. Then, the feature interactive is further achieved by the dense feature encoder (DFI) module to learn object local context information. The proposed method yields strong object context representation and well discriminability, as well as a good fit for infrared small object detection. Extensive experiments are conducted on the SISRT dataset and Synthetic dataset, demonstrating the superiority and effectiveness of the proposed deeper U-Net compared to previous state-of-the-art detection methods. Xin Wu 0001, Danfeng Hong, Zhanchao Huang, Jocelyn Chanussot |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Hyper-Embedder: Learning a Deep Embedder for Self-Supervised Hyperspectral Dimensionality ReductionabstractHyperspectral imaging has attracted growing interest among researchers from the geoscience and remote sensing fields owing to its very rich spectral information. However, the high spectral dimensionality of hyperspectral images (HSI) tends to suffer from information redundancy. Manifold embedding is a mainstream strategy of nonlinear hyperspectral dimensionality reduction (DR). The sensitivity to the noise and the inflexibility to the out-of-sample problem (i.e., new samples) are the main drawbacks of the manifold embedding-based methods. To this end, we propose to learn a deep embedder in a self-supervised fashion for hyperspectral DR, called hyper-embedder. Hyper-embedder effectively reduces the computational complexity and storage-costing compared to conventional embedding models and improves the robustness against various noises, e.g., spectral variabilities. More significantly, hyper-embedder is capable of learning an explicit nonlinear mapping to make a one-to-one match between each original pixel (spectral signature) in the HSI and its dimension-reduced representation. These low-dimensional representations can be generated and given by existing and classic nonlinear manifold embedding methods. In this letter, we attempt to learn the correspondence or mapping by optimizing a deep regression network. The to-be-developed network cannot only capture the local topological knowledge graph of all spectral signatures of hyperspectral data but be applicable to fast prediction and inference of samples from other hyperspectral scenes. The proposed hyper-embedder outperforms existing state-of-the-art hyperspectral DR algorithms on two commonly used hyperspectral datasets, i.e., Indian pines and Augsburg scenes. Xin Wu 0001, Danfeng Hong |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | A Novel Nonlocal-Aware Pyramid and Multiscale Multitask Refinement Detector for Object Detection in Remote Sensing ImagesabstractObject detection (OD) is an important task of computer vision and has been widely used in many fields, including remote sensing (RS). However, the complex scenes, large-scale variation, and dense instances of RS bring huge challenges to OD. To meet these challenges, a novel Nonlocal-aware Pyramid and Multiscale Multitask Refinement Detector (NPMMR-Det) is proposed. Specifically, nonlocal-aware pyramid attention (NP-Attention) is designed for guiding a neural network model to focus more on efficient features and suppress background noise. Then a multiscale refinement feature pyramid network (MSR-FPN) is proposed to fuse the multiscale context features extracted by the NP-Attention guided neural network and adjust the optimal receptive field. In order to use these features more effectively, a multitask refinement head called MTR-Head, with offset sharing and a modulation mechanism, is developed to refine the feature misalignment between the localization task and the classification task. Extensive experiments performed on two public RS data sets demonstrate that the proposed NPMMR-Det achieves competitive performance compared with state-of-the-art methods. Zhanchao Huang, Wei Li 0032, Xiang-Gen Xia 0001, Xin Wu 0001, Zhaoquan Cai 0001, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Revisiting Deep Hyperspectral Feature Extraction Networks via Gradient Centralized ConvolutionabstractThe hyperspectral images are composed of a variety of textures across the different bands which increase the spectral similarity and make it difficult to predict the pixel-wise labels without inducing additional complexity at the feature level. To extract robust and discriminative features from the different regions of land cover, the hyperspectral research community is still seeking such type of convolutions which can efficiently deal with fine-grained texture information during the feature extraction phase, which often overlook this aspect by vanilla convolution. To overcome the above shortcoming, this article proposes a generalized gradient centralized 3D convolution (G2C-Conv3D) operation, which is a weighted combination between the vanilla and gradient centralized 3D convolutions (GC-Conv3D) to extract both theintensity-levelsemantic information andgradient-levelinformation. This can be easily plugged into the existing HSI feature extraction networks to boost the performance of accurate prediction for land-cover types. To validate the feasibility of the proposedG2C-Conv3D, we have considered the existing CNN3D, MS3DNet, ContextNet, and SSRN feature extraction models and as well as CAE3D, VAE3D, and SAE3D autoencoder (AE) networks, respectively. All these networks are embedded withG2C-Conv3Dconvolution to implement both generalized gradient centralized feature extraction networks (G2C-FE) and generalized gradient centralized AE networks (G2C-AE) for fine-grained spectral–spatial feature learning. In addition,G2C-Conv2Dis also considered with few networks. The extensive experiments are conducted on four most widely used hyperspectral datasets i.e., IP, KSC, UH, and UP, respectively, and compared with the nine methods. The results demonstrate that the proposedG2C-Conv3Dcan effectively enhance the feature learning ability of the existing networks and both the qualitative and quantitative results show the superiority and effectiveness of the proposedG2C-Conv3D. The source codes will be publicly available athttps://github.com/danfenghong/G2C-Conv3D-HSI. Swalpa Kumar Roy, Purbayan Kar, Danfeng Hong, Xin Wu 0001, Antonio Plaza, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Convolutional Neural Networks for Multimodal Remote Sensing Data ClassificationabstractIn recent years, enormous research has been made to improve the classification performance of single-modal remote sensing (RS) data. However, with the ever-growing availability of RS data acquired from satellite or airborne platforms, simultaneous processing and analysis of multimodal RS data pose a new challenge to researchers in the RS community. To this end, we propose a deep-learning-based new framework for multimodal RS data classification, where convolutional neural networks (CNNs) are taken as a backbone with an advanced cross-channel reconstruction module, called CCR-Net. As the name suggests, CCR-Net learns more compact fusion representations of different RS data sources by the means of the reconstruction strategy across modalities that can mutually exchange information in a more effective way. Extensive experiments conducted on two multimodal RS datasets, including hyperspectral (HS) and light detection and ranging (LiDAR) data, i.e., the Houston2013 dataset, and HS and synthetic aperture radar (SAR) data, i.e., the Berlin dataset, demonstrate the effectiveness and superiority of the proposed CCR-Net in comparison with several state-of-the-art multimodal RS data classification methods. The codes will be openly and freely available athttps://github.com/danfenghong/IEEE_TGRS_CCR-Netfor the sake of reproducibility. Xin Wu 0001, Danfeng Hong, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Semi-Active Convolutional Neural Networks for Hyperspectral Image ClassificationabstractOwing to the powerful data representation ability of deep learning (DL) techniques, tremendous progress has been recently made in hyperspectral image (HSI) classification. Convolutional neural network (CNN), as a main part of the DL family, has been proven to be considerably effective to extract spatial-spectral features for HSIs. Nevertheless, its classification performance, to a great extent, depends on the quality and quantity of samples in the network training process. To select those samples, either labeled or unlabeled, that can be used to enhance the generalization ability of CNNs and further improve the classification accuracy, we propose an iterative semi-supervised CNNs framework by means of active learning and superpixel segmentation techniques, dubbed as semi-active CNNs (SA-CNNs), for HSI classification. More specifically, we start to pre-train a CNNs-based model on a small-scale unbiased labeled set and infer unlabeled data using the trained model, i.e., generating pseudo-labels. Then, the reliable samples, which consist of two parts: high label-homogeneity and most informativeness, are actively selected from superpixel segments. These selected labeled and unlabeled samples with their labels and pseudo-labels are re-fed into the next-round network training. Moreover, three different schedules, i.e.,log-,exp-, andlinear-schedules, are progressively adopted to fully explore their potentials in sample selection, until a labeling budget is finally reached. Extensive experiments are conducted on three benchmark HSI datasets, demonstrating substantial performance improvements of the proposed SA-CNNs over other similar competitors. Jing Yao 0002, Xiangyong Cao, Danfeng Hong, Xin Wu 0001, Deyu Meng, Jocelyn Chanussot, Zongben Xu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Multimodal Convolutional Neural Networks with Cross-Channel ReconstructionabstractWith the ever-growing availability of remote sensing (RS) data from either satellite or airborne sensors, simultaneous processing and analysis of multimodal data have been paid more and more attention by researchers in various RS-related applications. In this paper, we propose a multimodal convolutional neural network with an advanced cross-channel reconstruction module, called CCR-Net. As the name suggests, CCR-Net enables a more compact fusion of different RS data sources by the means of the reconstruction strategy across modalities that can mutually exchange information in a more effective way. Experiment are conducted on a widely-used dataset, including hyperspectral and Light Detection and Ranging (LiDAR) data, i.e., Houston2013, to verify the effectiveness and superiority of the proposed CCR - N et in comparison with several state-of-the-art baseline methods. Danfeng Hong, Xin Wu 0001, Jing Yao 0002, Lianru Gao, Bing Zhang 0001, Jocelyn Chanussot |
IGARSS | 2 |
| 2020 | The FrFT convolutional face: toward robust face recognition using the fractional Fourier transform and convolutional neural networks
Xin Wu 0001, Ran Tao 0003, Danfeng Hong, Yue Wang 0001 |
Sci. China Inf. Sci. | 1 |
| 2020 | Fourier-Based Rotation-Invariant Feature Boosting: An Efficient Framework for Geospatial Object DetectionabstractGeospatial object detection (GOD) of remote sensing imagery has been attracting increasing interest in recent years, due to the rapid development in spaceborne imaging. Most of the previously proposed object detectors are very sensitive to object deformations, such as scaling and rotation. To this end, we propose a novel and efficient framework for GOD in this letter, called Fourier-based rotation-invariant feature boosting (FRIFB). A Fourier-based rotation-invariant feature is first generated in polar coordinate. Then, the extracted features can be further structurally refined using aggregate channel features. This leads to a faster feature computation and more robust feature representation, which is good fitting for the coming boosting learning. Finally, in the test phase, we achieve a fast pyramid feature extraction by estimating a scale factor instead of directly collecting all features from the image pyramid. Extensive experiments are conducted on two subsets of NWPU VHR-10 data set, demonstrating the superiority and effectiveness of the FRIFB compared to the previous state-of-the-art methods. Xin Wu 0001, Danfeng Hong, Jocelyn Chanussot, Yang Xu 0006, Ran Tao 0003, Yue Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Invariant Attribute Profiles: A Spatial-Frequency Joint Feature Extractor for Hyperspectral Image ClassificationabstractSo far, a large number of advanced techniques have been developed to enhance and extract the spatially semantic information in hyperspectral image processing and analysis. However, locally semantic change, such as scene composition, relative position between objects, spectral variability caused by illumination, atmospheric effects, and material mixture, has been less frequently investigated in modeling spatial information. Consequently, identifying the same materials from spatially different scenes or positions can be difficult. In this article, we propose a solution to address this issue by locally extracting invariant features from hyperspectral imagery (HSI) in both spatial and frequency domains, using a method called invariant attribute profiles (IAPs). IAPs extract the spatial invariant features by exploiting isotropic filter banks or convolutional kernels on HSI and spatial aggregation techniques (e.g., superpixel segmentation) in the Cartesian coordinate system. Furthermore, they model invariant behaviors (e.g., shift, rotation) by the means of a continuous histogram of oriented gradients constructed in a Fourier polar coordinate. This yields a combinatorial representation of spatial-frequency invariant features with application to HSI classification. Extensive experiments conducted on three promising hyperspectral data sets (Houston2013 and Houston2018) to demonstrate the superiority and effectiveness of the proposed IAP method in comparison with several state-of-the-art profile-related techniques. The codes will be available from the website: https://sites.google.com/view/danfeng-hong/data-code. Danfeng Hong, Xin Wu 0001, Pedram Ghamisi, Jocelyn Chanussot, Naoto Yokoya, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | LW-ODF: A Light-Weight Object Detection Framework for Optical Remote Sensing ImageryabstractIn this paper, we propose to extract the multi-scaled and rotation-insensitive deep features to address the issues of object multi-solutions and rotations in geospatial object detection. To this end, we develop a novel object detection framework where a rotation-insensitive convolution neural network is applied for extracting multi-scaled and direction-insensitive feature representation and then the learned features can be fed into the ensemble classifier learning with fast feature pyramid. Such a non-end-to-end learning strategy intuitively reduces the computational cost without the additional performance loss, yielding an effective and efficient light-weight object detection framework. Experimental results conducted on the NWPU VHR-10 dataset demonstrate that the proposed framework outperforms several state-of-the-art baselines. Xin Wu 0001, Danfeng Hong, Pedram Ghamisi, Wei Li 0032, Ran Tao 0003 |
IGARSS | 1 |
| 2019 | A Weakly-Supervised Deep Network for DSM-Aided Vehicle DetectionabstractWith the breakthrough of the spatial resolution of optical remote sensing images at the sub-meter level and the explosive development of deep learning, geospatial object detection has achieved a growing interest in remote sensing community. However, labeling large training datasets in object level is still an expensive and tedious procedure. This might lead to the poor model generalization and degraded network learning ability. To this end, a weakly-supervised deep network (WSDN) is developed for geospatial object detection by applying a digital surface model (DSM)-aided auto-labeling and a pre-trained network learned from the task-independent dataset. Experimental results conducted on the stereo aerial imagery of a large camping site are performed to demonstrate that the proposed WSDN yields better detection results, with 62.78% precision and 55.13% recall. Xin Wu 0001, Danfeng Hong, Jiaojiao Tian, Ralph Kiefl, Ran Tao 0003 |
IGARSS | 1 |
| 2019 | ORSIm Detector: A Novel Object Detection Framework in Optical Remote Sensing Imagery Using Spatial-Frequency Channel FeaturesabstractWith the rapid development of spaceborne imaging techniques, object detection in optical remote sensing imagery has drawn much attention in recent decades. While many advanced works have been developed with powerful learning algorithms, the incomplete feature representation still cannot meet the demand for effectively and efficiently handling image deformations, particularly objective scaling and rotation. To this end, we propose a novel object detection framework, called Optical Remote Sensing Imagery detector (ORSIm detector), integrating diverse channel features extraction, feature learning, fast image pyramid matching, and boosting strategy. An ORSIm detector adopts a novel spatial-frequency channel feature (SFCF) by jointly considering the rotation-invariant channel features constructed in the frequency domain and the original spatial channel features (e.g., color channel and gradient magnitude). Subsequently, we refine SFCF using learning-based strategy in order to obtain the high-level or semantically meaningful features. In the test phase, we achieve a fast and coarsely scaled channel computation by mathematically estimating a scaling factor in the image domain. Extensive experimental results conducted on the two different airborne data sets are performed to demonstrate the superiority and effectiveness in comparison with the previous state-of-the-art methods. Xin Wu 0001, Danfeng Hong, Jiaojiao Tian, Jocelyn Chanussot, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Robust palmprint recognition based on the fast variation Vese-Osher model
Danfeng Hong, Wanquan Liu, Xin Wu 0001, Zhenkuan Pan 0001, Jian Su 0001 |
Neurocomputing | 3 |