VLDB 2026 Research / reviewers in the wild / expert
Xiao Huang 0003
dblp:25/692-3
· DBLP profile ↗
37ranked-venue papers
3as first author
32since 2021 · last 2026
0000-0002-4323-382XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From forgotten to pan-sharpening
Jiaming Wang 0001, Yansong Lin, Chuanxi Chen, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140, Tao Lu 0001 |
Pattern Recognit. | 4 |
| 2025 | Multi-feature machine learning for enhanced drug-drug interaction prediction
Qiuyang Feng, Xiao Huang 0003 |
J. Biomed. Informatics | 2 |
| 2025 | Lightweight remote sensing super-resolution with multi-scale graph attention network
Yu Wang 0140, Tao Lu 0001, Xiao Huang 0003, Jiaming Wang 0001, Zhizheng Zhang 0009, Xiaolong Zuo |
Pattern Recognit. | 4 |
| 2025 | Pattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote SensingabstractRecent self-supervised learning (SSL) methods have demonstrated impressive results in learning visual representations from unlabeled remote sensing (RS) images. However, most RS images predominantly consist of scenographic scenes containing multiple ground objects without explicit foreground targets, which limits the performance of existing SSL methods that focus on foreground targets. This raises the question: Is there a method that can automatically aggregate similar objects within scenographic RS images, thereby enabling models to differentiate knowledge embedded in various geospatial patterns for improved feature representation? In this work, we present the pattern integration and enhancement vision transformer (PIEViT), a novel SSL framework designed specifically for RS imagery. PIEViT utilizes a teacher-student architecture to address both image-level and patch-level tasks. It employs a proposed, geospatial pattern cohesion (GPC) module to explore the natural clustering of patches, enhancing the differentiation of individual features. A feature integration projection (FIP) module is employed to further refine masked token reconstruction using geospatially clustered patches. We validated PIEViT across multiple downstream tasks, including object detection, semantic segmentation, and change detection. Experiments demonstrated that PIEViT enhances the representation of internal patch features, providing significant improvements over existing self-supervised baselines. It achieves excellent results in object detection, land cover classification, and change detection, underscoring its robustness, generalization, and transferability for RS image interpretation tasks. Kaixuan Lu, Ruiqian Zhang, Xiao Huang 0003, Yuxing Xie, Xiaogang Ning, Hanchao Zhang, Mengke Yuan, Pan Zhang 0001, Tao Wang 0119, Tongkui Liao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Deep Merge: Deep-Learning-Based Region Merging for Remote Sensing Image SegmentationabstractImage segmentation represents a fundamental step in analyzing very high-spatial-resolution (VHR) remote sensing imagery. Its objective is to partition an image into segments that best match with geo-objects. However, the diverse appearances of geospatial objects often lead to interobject homogeneity and intraobject heterogeneity. Existing segmentation methods often struggle to accurately segment geo-objects with varying shapes and scales. To address these challenges, we propose DeepMerge, a novel method that integrates deep learning and region adjacency graphs (RAGs) to accurately segment complete geo-objects in large VHR images. DeepMerge begins with an initial over-segmentation of the image and then iteratively merges similar regions to achieve complete geo-object segmentation. A deep learning model is employed to learn the similarity between adjacent superpixel pairs. This approach only requires labels indicating whether adjacent superpixels belong to the same geo-object eliminating the need for object-level annotations, enabling weakly supervised segmentation. A cross-scale module is incorporated to capture multiscale information, enhancing the representation of superpixels. In addition, the feature distances between neighboring super-pixels are deemed as scale parameters (thresholds) to control the merging procedure, thus yielding an interpretable, predictable, stable, and optimal scale parameter 0.5. DeepMerge can achieve high segmentation accuracy in a weakly supervised manner, which is validated on large-scale remote sensing images of 0.55-m resolution covering an area of 5660 km2. The experimental results demonstrate that DeepMerge achieves the highest F value (0.9552) and the lowest total error (TE) (0.0827), accurately segmenting geo-objects of varying sizes and outperforming all competing methods. Xianwei Lv 0002, Claudio Persello, Wangbin Li, Xiao Huang 0003, Dongping Ming, Alfred Stein |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | C2FNet: Cross-Probabilistic Weak Supervision Learning for High-Resolution Land Cover EnhancementabstractAutonomous large-scale high-resolution land cover (HRLC) mapping remains a major challenge in remote sensing due to the scarcity of reliable training data and resolution mismatches between available labels and input of massive of emerging imagery. Existing global land cover products often suffer from coarse spatial resolution and label noise, limiting their utility for fine-scale urban analysis and environmental monitoring. This article presents C2FNet, a novel Coarse-to-Fine Network designed to generate HRLC maps from noisy, coarse-resolution labels using a weak supervision strategy of the cross-probability. The C2FNet consists of three key modules: 1) edge resolution refinement backbones (ERRBs), which preserve spatial detail via multiscale feature extraction through parallel convolutional branches; 2) unsupervised dynamic shuffle and diagonal annotation (UDSDA), which enhances training reliability by identifying confident regions through spatial-consistency analysis and confidence estimation; and 3) a contrasting self-supervised loss (C2F-Loss) that integrates cross-entropy and cosine similarity terms to mitigate supervision noise and resolution gaps. Evaluations of three benchmark datasets that encompass diverse urban and rural landscapes show that C2FNet achieves state-of-the-art (SoA) performance, with 80.01% overall accuracy (OA) and a Cohen’s kappa score of 0.7567, outperforming SoA models with weak supervision. The dataset and code are available athttp://drive.google.com/file/d/1X_Fz7LQIeix3rV3K29FBfKiU1WMdROe-/view Boaz Mwubahimana, Jianguo Yan, Dingruibo Miao, Zhuohong Li, Maurice Mugabowindekwe, Swalpa Kumar Roy, Xiao Huang 0003, Elias Nyandwi, Tuyishimire Joseph, Eric Habineza, Fidele Mwizerwa, Hafashimana Athanase, Gaspard Rwanyiziri |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | Rethinking the Role of Panchromatic Images in Pan-SharpeningabstractRecent pan-sharpening methods have predominantly utilized techniques tailored for natural image scenes, often overlooking the unique features arising from non-overlapping spectral responses. In light of this, we have reevaluated the utility of panchromatic (PAN) images and introduced a theory anchored in the spectral response of satellite sensors. This posits that a PAN image is effectively a linear weighted summation of individual bands from its corresponding multi-spectral (MS) image, offset by an error map. We developed a deep unmixing network termed “DUN” that integrates an unmixing network, a fusion mechanism, and a distinctive mutual information contrastive loss function. Notably, the unmixing network is adept at decomposing a PAN image into its MS counterpart and error map. Further, the demixed image alongside the low-resolution MS image is channeled into the fusion network for pan-sharpening. Recognizing the challenges of achieving robust supervised learning directly from the unmixing phase, we have innovated a mutual information contrastive learning loss function, ensuring enhanced separation and minimizing overlap during the unmixing process. Preliminary experiments underscore both the quantitative and qualitative prowess of the proposed method. Jiaming Wang 0001, Xitong Chen, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140, Tao Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Zigzag Attention: A Structural Aware Module For Lane DetectionabstractLane detection presents a formidable challenge in the realm of autonomous driving, given the real-time processing demands, diverse acquisition conditions, and the unique elongated and angular characteristics of lane lines. While a multitude of network design strategies have been proposed to tackle this challenge, few effectively address the distinct morphology of lane lines. Approaches that leverage global relationships across all positions, e.g. attention mechanisms, are often hard to meet real-time processing requirements due to their computational complexity. To address these challenges and unique attributes of lane lines, including issues like occlusion and dashed lines, we present a specialized and plug-and-play attention module. It employs zigzag transformations to cohesively assemble spatially disparate, lane-relevant regions, thereby transforming the challenge into one of localized feature learning, which can be easily enhanced via lightweight convolutions and fully connected layers. Additionally, we harness the symmetry inherent in lane lines to bolster the learning process and enhance accuracy. Comprehensive experimentation validates the efficacy of our proposed module across a range of algorithms, demonstrating superior performance metrics, including parameters, computational complexity, and runtime, when compared to other attention approaches. Jiajun Ling, Qimin Cheng, Xiao Huang 0003 |
ICASSP | 4 |
| 2024 | A Deep Error Removal Network for Pan-SharpeningabstractThe phenomenon of nonoverlapping spectral responses is an inevitable but an overlooked problem in the deep-learning-based panchromatic (PAN) and multispectral (MS) images’ fusion task, which will introduce some error information from the PAN image. In light of this, we construct a novel prior model based on spectral response theory and develop a model-based pan-sharpening network. Specifically, we extract the initial error map from the PAN and interpolate the MS image as the initial pan-sharpened result. Then, two optimization problems regularized by the deep prior are formulated to update the error map and pan-sharpened image. By alternately optimizing the above subtasks, error information is gradually separated from PAN images and the lost texture information in MS images is gradually restored, which can effectively alleviate the negative impact of low coupling information from PAN and MS images. Plenty of experimental results on different kinds of satellite datasets demonstrate that the proposed method shows a better balance between interpretability and lightweight structure. The proposed method will be open-sourced inhttps://github.com/jiaming-wang/DERN. Jiaming Wang 0001, Tao Lu 0001, Xiao Huang 0003, Ruiqian Zhang, Dongyue Luo |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Pan-sharpening via intrinsic decomposition knowledge distillation
Jiaming Wang 0001, Xiao Huang 0003, Ruiqian Zhang, Xitong Chen, Tao Lu 0001 |
Pattern Recognit. | 3 |
| 2024 | WodNet: Weak Object Discrimination Network for Cloud DetectionabstractTo enhance the accuracy of remote sensing data analysis, cloud detection from the complex ground environment is crucial. We refer to clouds that are easily confused with similar background as weak targets clouds, including thin clouds, tiny clouds, cloud boundaries, clouds with snow’s existence or highlighted background’s existence. This paper proposes a coarse-to-fine cloud detection network for weak target recognition. The network consists of two subnetworks: the Scalable Weak Target Feature Extraction Subnetwork (SWTFES) and the Cascade Weak Target Refinement Subnetwork (CWTRS). SWTFES incorporates a Multi-scale Feature Extraction Module (MFEM) with different scale receptive field branches and an Attention-based Cross-layer Fusion Module (ACFM) to characterize cloud at various scales. The improved reverse attention operation and the Cascade Group Reverse Attention Module (CGRAM) serve as the guiding principles in CWTRS, driving the network to progressively add and refine the weak target’s details to distinguish it from the complex background surface. We evaluate our methodology on four cloud datasets with various resolutions, varying from 0.5m to 16m, and different satellites (including Gaofen-1 WFV, Sentinel-2, Gaofen-2, WorldView-2). The experimental results demonstrate that WodNet has achieved excellent results in cloud detection in a variety of complex scenarios, compared to other models, performing SOTA in four challenging datasets. Xuechao Zhou, Xinrui Xie, Xiao Huang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Optimizing pedestrian simulation based on expert trajectory guidance and deep reinforcement learning
Senlin Mu, Xiao Huang 0003, Moyang Wang, Di Zhang 0022, Dong Xu 0009, Xiang Li 0033 |
GeoInformatica | 2 |
| 2023 | AERNet: An Attention-Guided Edge Refinement Network and a Dataset for Remote Sensing Building Change DetectionabstractAdvancements in Earth observation technology enable the detection of surface changes in intricate urban environments. Building change detection (BCD) plays a crucial role in urban planning and environmental monitoring. However, existing deep learning-based BCD algorithms exhibit limited capability in feature extraction, feature relationship comprehension, sample imbalance mitigation, and accurate boundary identification for changed objects. To address these challenges, we introduce an attention-guided edge refinement network (AERNet) that employs a global context feature aggregation module (GCFAM) to aggregate information from extracted multi-layer context features. Our approach incorporates an attention decoding block (ADB) guided by enhanced coordinate attention (ECA) to capture channel and location associations between features. Furthermore, we utilize an edge refinement module (ERM) to enhance the network’s capacity to sense and refine the edges of changed areas. To tackle the issue of class imbalance and augment the algorithm’s feature learning ability, we devise a novel self-adaptive weighted binary cross-entropy (SWBCE) loss function, combined with a deep supervision (DS) strategy. Experiments are conducted on two publicly available datasets, GDSCD and LEVIR-CD, as well as our newly developed high-resolution complex urban scene BCD dataset, i.e., HRCUS-CD. The latter dataset comprises 11,388 pairs of images at 0.5-meter resolution and over 12,000 labeled change buildings. Comparative experiments indicate that AERNet surpasses advanced competitive methods, while ablation experiments demonstrate the effectiveness of AERNet’s model components and the SWBCE loss function. Efficiency comparison confirms that AERNet achieves comprehensive detection performance with superior effectiveness and robustness. Jindou Zhang, Xiao Huang 0003, Yu Wang 0140, Xuechao Zhou, DeRen Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Extraction and analysis of natural disaster-related VGI from social media: review, opportunities and challengesabstractThe idea of ‘citizen as sensors’ has gradually become a reality over the past decade. Today, Volunteered Geographic Information (VGI) from citizens is highly involved in acquiring information on natural disasters. In particular, the rapid development of deep learning techniques in computer vision and natural language processing in recent years has allowed more information related to natural disasters to be extracted from social media, such as the severity of building damage and flood water levels. Meanwhile, many recent studies have integrated information extracted from social media with that from other sources, such as remote sensing and sensor networks, to provide comprehensive and detailed information on natural disasters. Therefore, it is of great significance to review the existing work, given the rapid development of this field. In this review, we summarized eight common tasks and their solutions in social media content analysis for natural disasters. We also grouped and analyzed studies that make further use of this extracted information, either standalone or in combination with other sources. Based on the review, we identified and discussed challenges and opportunities. Yu Feng 0006, Xiao Huang 0003, Monika Sester |
Int. J. Geogr. Inf. Sci. | 2 |
| 2022 | BTS: a binary tree sampling strategy for object identification based on deep learningabstractObject-based convolutional neural networks (OCNNs) have achieved great performance in the field of land-cover and land-use classification. Studies have suggested that the generation of object convolutional positions (OCPs) largely determines the performance of OCNNs. Optimized distribution of OCPs facilitates the identification of segmented objects with irregular shapes. In this study, we propose a morphology-based binary tree sampling (BTS) method that provides a reasonable, effective, and robust strategy to generate evenly distributed OCPs. The proposed BTS algorithm consists of three major steps: 1) calculating the required number of OCPs for each object, 2) dividing a vector object into smaller sub-objects, and 3) generating OCPs based on the sub-objects. Taking the object identification in land-cover and land-use classification as a case study, we compare the proposed BTS algorithm with other competing methods. The results suggest that the BTS algorithm outperforms all other competing methods, as it yields more evenly distributed OCPs that contribute to better representation of objects, thus leading to higher object identification accuracy. Further experiments suggest that the efficiency of BTS can be improved when multi-thread technology is implemented. Xianwei Lv 0002, Xiao Huang 0003, Dongping Ming, Jiaming Wang 0001, Chengzhuo Tong |
Int. J. Geogr. Inf. Sci. | 3 |
| 2022 | Exploring the vertical dimension of street view image based on deep learning: a case study on lowest floor elevation estimationabstractStreet view imagery such as Google Street View is widely used in people’s daily lives. Many studies have been conducted to detect and map objects such as traffic signs and sidewalks for urban built-up environment analysis. While mapping objects in the horizontal dimension is common in those studies, automatic vertical measuring in large areas is underexploited. Vertical information from street view imagery can benefit a variety of studies. One notable application is estimating the lowest floor elevation, which is critical for building flood vulnerability assessment and insurance premium calculation. In this article, we explored the vertical measurement in street view imagery using the principle of tacheometric surveying. In the case study of lowest floor elevation estimation using Google Street View images, we trained a neural network (YOLO-v5) for door detection and used the fixed height of doors to measure doors’ elevation. The results suggest that the average error of estimated elevation is 0.218 m. The depthmaps of Google Street View were utilized to traverse the elevation from the roadway surface to target objects. The proposed pipeline provides a novel approach for automatic elevation estimation from street view imagery and is expected to benefit future terrain-related studies for large areas. Huan Ning, Zhenlong Li, Xinyue Ye, Shaohua Wang 0002, Xiao Huang 0003 |
Int. J. Geogr. Inf. Sci. | 6 |
| 2022 | SAR-DRDNet: A SAR image despeckling network with detail recovery
Wenfu Wu, Xiao Huang 0003, Jiahua Teng, DeRen Li |
Neurocomputing | 2 |
| 2022 | Adaptive dense pyramid network for object detection in UAV imagery
Ruiqian Zhang, Xiao Huang 0003, Jiaming Wang 0001, Yufeng Wang 0004, DeRen Li |
Neurocomputing | 3 |
| 2022 | Deep locally linear embedding network
Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Xitong Chen |
Inf. Sci. | 3 |
| 2022 | SSCAN: A Spatial-Spectral Cross Attention Network for Hyperspectral Image DenoisingabstractHyperspectral images (HSIs) have been widely used in a variety of applications thanks to the rich spectral information they are able to provide. Among all HSI processing tasks, HSI denoising is a crucial step. Recent years have seen great progress in deep learning-based image denoising methods. However, existing efforts tend to ignore the correlations between adjacent spectral bands, leading to problems such as spectral distortion and blurred edges in denoised results. In this study, we propose a novel HSI denoising network, termed spectral–spatial cross attention network (SSCAN), that combines group convolutions and attention modules. Specifically, we use a group convolution with a spatial attention module to facilitate feature extraction by directing models’ attention to bandwise important features. We also propose a spectral–spatial attention block (SSAB) to effectively exploit the spatial and spectral information in HSIs. In addition, we adopt residual learning operations with skip connections to ensure training stability. The experimental results indicate that the proposed SSCAN outperforms several state-of-the-art HSI denoising algorithms. Xiao Huang 0003, Jiaming Wang 0001, Tao Lu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | BiCSNet: A Bidirectional Cross-Scale Backbone for Recognition and LocalizationabstractRecognition and localization models can be generally decomposed into three components: encoder, decoder, and task head. In this paper, we rethink the necessity of decoder, as we observe that it brings additional computational and parametric burden. We thus propose to remove the decoder and present a bidirectional cross-scale architecture that is able to obtain rich semantic information and precise localization in a unified backbone. Extensive experiments demonstrate that, different from common encoder-decoder models and other down-sampling and up-sampling backbones, the proposed BiCSNet achieves improved performances compared to existing architectures for pixel-level tasks. In object detection, our BiCSNet brings significant performance improvement by ~ 3% AP at various scales with 13% – 23% fewer FLOPS, compared with ResNet-FPN models on COCO dataset. In Instance segmentation, the AP can be improved by 1% over SpineNet. BiCSNet is also promising for semantic segmentation tasks, as the proposed BiCSNet pre-trained on ImageNet alone significantly outperforms DeepLabv3 pre-trained on both ImageNet and COCO dataset by 1.3% in mIOU with 89% fewer FLOPs on PASCAL VOC 2012. Xiao Huang 0003, Yi Zhu 0001, Ruiqian Zhang, Junwei Zha |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Road and Car Extraction Using UAV Images via Efficient Dual Contextual Parsing NetworkabstractThe rapid development and commercialization of Unmanned Aerial Vehicle (UAV) technology has made it possible to conduct urban traffic information extraction using UAV images. However, the large variations of targets in urban environments, complex foregrounds and backgrounds in cities, and severe tree and shadow occlusions pose great challenges in car and road extraction using UAV images. In this study, we propose a lightweight, Efficient Dual Contextual Parsing Network (EDCPNet) to address the above issues. The proposed EDCP module in EDCPNet is mainly composed of spatial contextual parsing (SCP) and channel contextual parsing (CCP), which can effectively acquire rich contextual features in both spatial and channel dimensions, adaptively recalibrate the attention weights, perceive the salient features of targets in images, and suppress the importance of irrelevant elements. It thus leads to improved performance and adaptability that facilitate the practical applications of large-scale urban traffic monitoring in complex urban scenes. We conduct experiments on two benchmark datasets (UAVid and UDD) by comparing the proposed EDCPNet with six other competing methods, i.e., U-Net, PSPNet, Deelabv3+, SegNet, ESNet, and ERFNet, and validate the effectiveness of the proposed EDCP module via extensive ablation studies. The results suggest that the proposed network outperforms all competing methods in car and road extraction from UAV images with a balanced computational cost. Its great performance and low computational demand (with only 2.37M model parameters) facilitate its deployment on edge computing devices with memory constraints. Gui Cheng, Xiao Huang 0003, Zhongyuan Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Dual-Path Fusion Network for Pan-SharpeningabstractMost existing deep learning-based pan-sharpening methods own several widely recognized issues, such as spectral distortion and insufficient spatial texture enhancement. To address these challenges in pan-sharpening, we propose a novel dual-path fusion network (DPFN). The proposed DPFN includes two major components: 1) the global subnetwork (GSN) and 2) the local subnetwork (LSN). In particular, GSN aims to search similar image blocks in panchromatic (PAN) space and multispectral (MS) space and exploits HR textural information from the PAN space and spectral information from the MS space for the fine representation of pan-sharpened MS features by employing a cross nonlocal block. Meanwhile, the proposed LSN based on a high-pass modification block (HMB) is designed to learn the high-pass information, aiming to enhance bandwise spatial information from MS images. HMB forces the fused image to obtain high-frequency details from PAN images. Moreover, to facilitate the generation of visually appealing pan-sharpened images, we propose a perceptual loss function and further optimize the model based on high-level features in the near-infrared space. Experiments demonstrate the superior performance of the proposed method quantitatively and qualitatively compared to existing state-of-the-art pan-sharpening methods. The source code is available athttps://github.com/jiaming-wang/DPFN. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Pan-Sharpening via Deep Locally Linear Embedding Residual NetworkabstractThe goal of pan-sharpening tasks is to fuse panchromatic (PAN) images and low-spatial-resolution (LR) multispectral (MS) images for the purpose of aggregating texture and spectral information. Although traditional embedding-based pan-sharpening methods achieve competitive results, they are limited by the shallow network and not suitable for large-scale datasets. In this study, we design a novel multiscale locally linear embedding residual network (LLERN) that consists of two phases: the spectral preservation phase and the structural preservation phase. As the pretreatment of the structural preservation network, the spectral preservation network aims to upscale the LR MS image while retaining spectral information. The proposed locally linear embedding residual block (LLERB) in the structural preservation phase can search for similar sparse patches from the PAN image space and embed the corresponding local geometric relationship into the residual space to enhance the MS image. Extensive experiments suggest that the proposed LLERN outperforms state-of-the-art methods from visual and quantitative perspectives, and confirm the assumption that LR image patches and residual image patches in a local region share a similar manifold structure, which can be used to guide deep-learning modeling with improved interpretability. The source code is available athttps://github.com/jiaming-wang/LLERN. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Gui Cheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | From Artifact Removal to Super-ResolutionabstractDeep-learning-based super-resolution methods have been extensively studied and achieved significant performance with deep convolutional neural networks. However, the results still suffer from the ringing effect, especially in satellite image super-resolution tasks, due to the loss of image details in the satellite degradation process. In this paper, we build a novel satellite super-resolution framework by decomposing a high-resolution image into three components, i.e., low-resolution, artifact, and high-frequency information. Specifically, we propose an artifact removal network with a self-adaption difference convolution (SDC) to fully exploit the structure prior in the low-resolution image and predict the artifact map. Considering that the artifact map and the high-frequency map share a similar pattern, we introduce the supervised structure correction block (SSC) that establishes a bridge between the high-frequency generation process and the artifact removal process. Experimental results on satellite images demonstrate that the proposed method owns an improved tradeoff between the performance and the computational cost compared to existing state-of-the-art satellite and natural super-resolution methods. The source code is available at https://github.com/jiaming-wang/ARSRN. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Pan-Sharpening Via High-Pass Modification Convolutional Neural NetworkabstractMost existing deep learning-based pan-sharpening methods have several widely recognized issues, such as spectral distortion and insufficient spatial texture enhancement, we propose a novel pan-sharpening convolutional neural network based on a high-pass modification b lock. Different from existing methods, the proposed block is designed to learn the high-pass information, leading to enhance spatial information in each band of the multi-spectral-resolution images. To facilitate the generation of visually appealing pan-sharpened images, we propose a perceptual loss function and further optimize the model based on high-level features in the near-infrared space. Experiments demonstrate the superior performance of the proposed method compared to the state-of the-art pan-sharpening methods, both quantitatively and qualitatively. The proposed model is open-sourced at https://github.com/jiaming-wang/HMB. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Jiayi Ma 0001 |
ICIP | 3 |
| 2021 | Unsupervised Remoting Sensing Super-Resolution via Migration Image PriorabstractRecently, satellites with high temporal resolution have fostered wide attention in various practical applications. Due to limitations of bandwidth and hardware cost, however, the spatial resolution of such satellites is considerably low, largely limiting their potentials in scenarios that require spatially explicit information. To improve image resolution, numerous approaches based on training low-high resolution pairs have been proposed to address the super-resolution (SR) task. De-spite their success, however, low/high spatial resolution pairs are usually difficult to obtain in satellites with a high temporal resolution, making such approaches in SR impractical to use. In this paper, we proposed a new unsupervised learning framework, called "MIP", which achieves SR tasks without low/high resolution image pairs. First, random noise maps are fed into a designed generative adversarial network (GAN) for reconstruction. Then, the proposed method converts the reference image to latent space as the migration image prior. Finally, we update the input noise via an implicit method, and further transfer the texture and structured information from the reference image. Extensive experimental results on the Draper dataset show that MIP achieves significant improvements over state-of-the-art methods both quantitatively and qualitatively. The proposed MIP is open-sourced at https://github.com/jiaming-wang/MIP. Jiaming Wang 0001, Tao Lu 0001, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140 |
ICME | 4 |
| 2021 | Analysis of the performance and robustness of methods to detect base locations of individuals with geo-tagged social media dataabstractVarious methods have been proposed to detect the base locations of individuals, with their geo-tagged social media data. However, a common challenge relating to base-location detection methods (BDMs) is that, the rare availability of ground-truth data impedes the method assessment of accuracy and robustness, thus undermining research validity and reliability. To address this challenge, we collect users’ information from unstructured online content, and evaluate both the performance and robustness of BDMs. The evaluation consists of two tasks: the detection of base locations and also the differentiation between local residents and tourists. The results show BDMs can achieve high accuracies in base-location detection but tend to overestimate the number of tourists. Evaluation conducted in this study, also shows that BDMs’ accuracy is subject to the intensity of user’s activities and number of countries visited by the user but are insensitive to user’s gender. Temporally, BDMs perform better during weekends and summertime than during other periods, but the best performances appear with datasets that cover the whole time periods (whole day, week, and year). To the best of knowledge, this study is the first work to evaluate the performance and robustness of BDMs at individual level. Zhewei Liu, An-Shu Zhang, Yepeng Yao, Wenzhong Shi, Xiao Huang 0003, Xiaoqi Shen |
Int. J. Geogr. Inf. Sci. | 5 |
| 2021 | Spatial-temporal pooling for action recognition in videos
Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Xianwei Lv 0002 |
Neurocomputing | 3 |
| 2021 | Internal and external spatial-temporal constraints for person reidentification
Jiaming Wang 0001, Tao Lu 0001, Ruiqian Zhang, Xiao Huang 0003, Xianwei Lv 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Enhanced image prior for unsupervised remoting sensing super-resolution
Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Jiayi Ma 0001 |
Neural Networks | 3 |
| 2021 | An Optimal Channel Selection for EEG-Based Depression Detection via Kernel-Target AlignmentabstractDepression is a mental disorder with emotional and cognitive dysfunction. The main clinical characteristic of depression is significant and persistent low mood. As reported, depression is a leading cause of disability worldwide. Moreover, the rate of recognition and treatment for depression is low. Therefore, the detection and treatment of depression are urgent. Multichannel electroencephalogram (EEG) signals, which reflect the working status of the human brain, can be used to develop an objective and promising tool for augmenting the clinical effects in the diagnosis and detection of depression. However, when a large number of EEG channels are acquired, the information redundancy and computational complexity of the EEG signals increase; thus, effective channel selection algorithms are required not only for machine learning feasibility, but also for practicality in clinical depression detection. Consequently, we propose an optimal channel selection method for EEG-based depression detection via kernel-target alignment (KTA) to effectively resolve the abovementioned issues. In this method, we consider a modified version KTA that can measure the similarity between the kernel matrix for channel selection and the target matrix as an objective function and optimize the objective function by a proposed optimal channel selection strategy. Experimental results on two EEG datasets show that channel selection can effectively increase the classification performance and that even if we rely only on a small subset of channels, the results are still acceptable. The selected channels are in line with the expected latent cortical activity patterns in depression detection. Moreover, the experimental results demonstrate that our method outperforms the state-of-the-art channel selection approaches. Jian Shen 0004, Xiaowei Zhang 0001, Xiao Huang 0003, Manxi Wu, Zhijie Ding, Bin Hu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Spatial-temporal Joint optimization Network on Covariance Manifolds of Electroencephalography for Fatigue DetectionabstractThe World Health organization (WHO) stated that the concept of health has been widened to subjectively experienced dimensions such as fatigue and chronic fatigue syndrome (CFS). With the increasing pressure of the current life, persistent fatigue caused by sustained high-pressure work will not only be hazardous to health, but also give rise to unexpected consequences. In particularly, fatigue driving induced by long time driving has become a leading cause of accidents and death in the transportation. In this study, we investigate electroencephalography(EEG)-based fatigue detection of drivers through the spatial-temporal changes in the relations between EEG channels. EEG signals are firstly partitioned into several segments and the covariance matrices obtained from each segment are fed into a recurrent neural network to extract high-level temporal features. Then, the covariance matrices of whole signals are leveraged to extract spatial characteristics, which will be fused with temporal features to obtain comprehensive spatial-temporal information. Experimental results on a benchmark dataset showed that our method obtained an optimal classification accuracy of 91.042% and outperformed some state-of-the-art methods. These results indicate that our method is reliable and feasible for fatigue detection, which also provides a novel solution for EEG modeling. Xiaowei Zhang 0001, Jian Shen 0004, Xiao Huang 0003, Manxi Wu |
BIBM | 5 |
| 2020 | Translating Multispectral Imagery to Nighttime Imagery via Conditional Generative Adversarial NetworksabstractNighttime satellite imagery has been applied in a wide range of fields. However, our limited understanding of how observed light intensity is formed and whether it can be simulated greatly hinders its further application. This study explores the potential of conditional Generative Adversarial Networks (cGAN) in translating multispectral imagery to nighttime imagery. A popular cGAN framework, pix2pix, was adopted and modified to facilitate this translation using gridded training image pairs derived from Landsat 8 and Visible Infrared Imaging Radiometer Suite (VIIRS). The results of this study prove the possibility of multispectral-to-nighttime translation and further indicate that, with the additional social media data, the generated nighttime imagery can be very similar to the ground-truth imagery. This study fills the gap in understanding the composition of satellite observed nighttime light and provides new paradigms to solve the emerging problems in nighttime remote sensing fields, including nighttime series construction, light desaturation, and multi-sensor calibration. Xiao Huang 0003, Dong Xu 0009, Zhenlong Li, Cuizhen Wang |
IGARSS | 1 |
| 2019 | Individual Similarity Guided Transfer Modeling for EEG-based Emotion RecognitionabstractIntelligent recognition of electroencephalogram (EEG) signals has been an important means to recognize emotions. Traditional user-independent method, which treatseach individual's EEG data as independent and identically distributed (i.i.d.) samples and ignores destruction on i.i.d. condition caused by individual differences, usually has lower generalization performance. Although user-dependent method could alleviate abovementioned problem, it faces difficulty in collection of sufficient training EEG data for each individual. In order to construct user-dependent model merely based on a small amount of training EEG data, we incorporate transfer learning framework and propose a individual similarity guided transfer modeling method for EEG-based emotion recognition. We first measure the similarities between individuals using maximum mean discrepancy (MMD), then utilize pre-existing EEG data of similar individuals to assist construction of user-dependent model for the target individual using an instance-based transfer learning algorithm named TrAdaBoost. We compared this method with traditional user-independent and user-dependent methods on DEAP dataset. Experimental results showed that our method could transfer useful knowledge from other individuals for user-dependent emotion recognition, which achieved classification accuracies of 66.1% and 66.7% on arousal and valence dimentions, respectively. Xiaowei Zhang 0001, Tingzhen Ding, Jian Shen 0004, Xiao Huang 0003 |
BIBM | 6 |
| 2019 | Human Settlement Dynamics in Hurricane-Prone Zones of Conterminous U.S: A View from Nighttime Remote SensingabstractHurricane is one of the most devastating natural disasters on earth, posing great threats to people living in coastal areas. Thus, a better understanding of human settlement dynamics, especially in hurricane-prone areas, is in great need. This study examines nighttime satellite-derived human settlement in hurricane-prone areas of conterminous U.S from 1992 to 2013. To solve the lack of on-board calibration problem and saturation of luminosity problem, the original DMSP/OLS time series were intercalibrated and desaturated using NDVI products from AVHRR and MODIS. A popular index, VANUI, was derived and was later applied Mann-Kendall and Theil-Sen trend test. Four hurricane-prone zones, representing different level of hurricane proneness, were delineated from historical NAB-origin storm tracks using a wind-speed weighted track density function. The dynamic of VANUI in different hurricane-prone zones was finally quantified and analyzed. The result suggests a distinctive discrepancy in settlement intensity between northern and southern study area. It also indicates that the both the extent and the increase rate in settlement intensity positively correlate hurricane proneness in predefined hurricane-prone zones. Xiao Huang 0003, Cuizhen Wang |
IGARSS | 1 |
| 2018 | Reconstructing Flood Inundation Probability by Enhancing Near Real-Time Imagery With Real-Time Gauges and TweetsabstractFlood inundation probability is critical for situation awareness, flood mitigation, emergency response, and postevent damage assessment. Current flood inundation mapping approaches can be categorized into real-time (RT) and near-RT (NRT) processes based on the timing of data acquisition. However, the intrinsic limitations of each category largely hamper their applications for flood mapping. Taking the 2015 South Carolina flood in downtown Columbia as a case study, this paper proposes a flood inundation reconstruction model by enhancing the NRT normalized difference water index (NDWI) derived from remote sensing imagery with the RT data including stream gauge readings and social media (tweets). Splitting into three modules: water height module, global enhancement module, and local enhancement module, the proposed model first incorporates the gauge readings and the NDWI image to reconstruct a macroscale flood probability layer, which is then locally enhanced using the verified flood-related tweets. The final output of the model matches well with the U.S. Geological Survey inundation map and its surveyed high-water marks. Results suggest that by enhancing NRT imagery with RT data sources, the proposed flood inundation probability reconstruction model renders a more robust, spatially enhanced flood probability index for emergency responders to quickly identify areas in need of urgent attention. Xiao Huang 0003, Cuizhen Wang, Zhenlong Li |
IEEE Trans. Geosci. Remote. Sens. | 1 |