VLDB 2026 Research / reviewers in the wild / expert
Xiaofei Yang 0002
dblp:145/1177-2
· DBLP profile ↗
37ranked-venue papers
8as first author
32since 2021 · last 2026
0000-0003-2458-6774ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-branch attention network with multi-level spectral-spatial fusion for hyperspectral image classification
Yuanlin Dang, Lei Zhu 0004, Xiaofei Yang 0002 |
Knowl. Based Syst. | 4 |
| 2026 | RSRWKV: A Linear-Complexity 2D Attention Mechanism for Efficient Remote Sensing Vision TaskabstractDeep learning methods have got a great success in high-resolution remote sensing analysis, especially Convolution Neural Network (CNN) and Transformer. However, CNNs have a failure in modeling the long-range dependency because of their fixed receptive fields and Transformers suffer from quadratic computational complexity relative to image resolution. The RWKV model achieves breakthroughs in natural language processing (NLP) through its linear-complexity sequence modeling; however, it exhibits anisotropic limitations in vision tasks due to the constraints of its one-dimensional scanning mechanism. To address these challenges, we adapt the RWKV architecture to high-resolution remote sensing and propose the Remote Sensing RWKV (RSRWKV) model, which incorporates a Linear-Complexity 2D Attention Mechanism. Specifically, RSRWKV employs a novel 2D-WKV scanning mechanism that bridges sequential processing with two-dimensional spatial reasoning while maintaining linear computational complexity. This design facilitates the aggregation of isotropic contexts in multiple spatial directions. Then, the MVC-Shift module further optimizes multiscale receptive field coverage, whereas the Efficient Channel Attention (ECA) module improves cross-channel feature interaction and semantic saliency modeling. Experimental evaluations on the NWPU RESISC45, VHR-10 v2, SSDD and GLHWater datasets demonstrate that RSRWKV surpasses CNN and Transformer baselines in classification, detection and segmentation tasks, establishing a scalable framework for high-resolution remote sensing analysis. Code available at https://github.com/Ling-yunchi/RSRWKV. Chunshan Li, Xiaofei Yang 0002, Xishuang Han, Xiaowen Chu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | RWKVSR: Receptance Weighted Key-Value Network for Hyperspectral Image Super-ResolutionabstractDeep learning has achieved significant success in hyperspectral image super-resolution (HSISR) by leveraging advanced feature extraction techniques to reconstruct high-resolution images from low-resolution counterparts. However, existing methods predominantly utilize 2D/3D convolutions or Transformer architectures, which are often hindered by limited receptive fields, quadratic computational complexity, and inadequate fusion of spatial-spectral dependencies. To address these challenges, this paper proposes RWKVSR, a novel lightweight network that integrates a Receptance Weighted Key-Value (RWKV) architecture for efficient HSISR. The proposed RWKVSR comprises of three key components: (1) A linear-complexity RWKV module replacing quadratic self-attention, enabling efficient global spectral-spatial modeling; (2) A Spectral-Spatial Residual Module (SSRM) employing anisotropic, direction-separable 3D convolutions to hierarchically extract multi-scale features while enhancing local-global interactions; and (3) A Hyperspectral Frequency Loss (HFL) optimizing spectral consistency by prioritizing high-frequency structural alignment between reconstructed and ground-truth images in the frequency domain. Extensive experiments conducted on the CAVE and Harvard datasets demonstrate that RWKVSR outperforms the existing state-of-the-art methods, effectively balancing accuracy and efficiency, and providing a practical solution for high-quality HSI reconstruction. Our paper code is publicly available at https://github.com/backy-1/RWKVSR.git. Xiaofei Yang 0002, Sihuan Li, Weijia Cao, Yifang Ban, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | ReSeg-UNet: A Reconstruction-Guided Optimization Framework for Enhanced Medical Image Segmentation
Xiaowen Chu 0001, Xiaofei Yang 0002 |
MICCAI (3) | 4 |
| 2025 | Balancing supply and demand for ride-hailing: A preallocation hierarchical reinforcement learning approach
Jiahao Ling, Xiaohui Huang 0003, Xiaofei Yang 0002, Boxue Cheng |
Inf. Sci. | 3 |
| 2025 | Global-local prototype-based few-shot learning for cross-domain hyperspectral image classification
Haojin Tang, Yuelin Wu, Xiaofei Yang 0002, Weixin Xie |
Knowl. Based Syst. | 5 |
| 2025 | PAB-Road: A Patch-Wise Boundary for Road Network Extraction via Multitask UNetabstractRoad network extraction from remote sensing images is a fundamental task for applications like autonomous driving and urban planning. Mainstream methods, however, face a critical trade-off: segmentation-based approaches provide high geometric detail but often yield fragmented roadmaps, while graph-based approaches ensure connectivity but can sacrifice fine-grained accuracy. While hybrid models have been explored to resolve this, effectively fusing pixel-level features with structural information remains a key challenge. To address this, we propose PAB-Road, a novel framework with a unique fusion mechanism. Its core novelty is a multitask UNet that learns a patch-wise boundary representation to explicitly model local connectivity. This learned information then guides the synthesis of a geometrically accurate and structurally coherent road network. Experimental results in the real dataset show that PAB-Road achieves a compelling F1 score of 79.29%, demonstrating the effectiveness of our proposed fusion strategy. Wenhai Li, Xianhong Zhu, Xiaohui Huang 0003, Xiaofei Yang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Few-Shot Hyperspectral Image Classification With Deep Fuzzy Metric LearningabstractDeep metric learning (DML) has shown promising results in few-shot hyperspectral image (HSI) classification. The core idea of DML is to learn a generalized metric space, in which pixels from unseen classes can be effectively classified with only a few labeled samples. However, the existing DML methods mainly adopt traditional Euclidean distance to achieve the feature metric, which ignores the category uncertainty of spatial-spectral features in mixed and edge pixels. To address this issue, we fully exploit fuzzy logic theory and propose a deep fuzzy metric learning (DFML) method for few-shot HSI classification. First, we design a novel hybrid CNN-transformer spatial-spectral feature extraction network to fully capture the spatial-spectral features of HSI pixels. Then, a fuzzy set representation method based on Gaussian membership function for spatial-spectral features is proposed, which describes the inherent fuzziness of the spatial-spectral features. Finally, to perform the fuzzy similarity measure between the fuzzy sets of query samples and prototypes, we construct a spatial-spectral fuzzy metric space, in which HSI pixels with category uncertainty in their features can be better classified under the condition of small-scale labeled samples. Extensive experimental results on three public HSI datasets demonstrate that the proposed DFML method outperforms the state-of-the-art few-shot HSI classification methods. Haojin Tang, Xiaofei Yang 0002, Weixin Xie |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | ADNM-UNet: An Asymmetric Dual-Branch Noncausal Mamba U-Net With Multiscale Attention Enhancement for Cloud Mask NowcastingabstractCloud mask underpins accurate precipitation nowcasting, which in turn is vital for understanding the hydrological cycle, supporting disaster prevention, solar energy forecasting and transportation. However, cloud mask nowcasting remains challenging because meteorological data exhibit irregular temporal and spatial variations, including fine-scale structures, and often suffer from highly skewed precipitation intensity distributions. Existing methods struggle to capture complex spatiotemporal dynamics and preserve fine-scale structures due to limitations in handling sparse data from numerical weather prediction (NWP) model. To address these issues, we propose an asymmetric dual-branch non-causal mamba U-Net (ADNM-UNet) featuring three key components: (1) The Asymmetric Dual-branch Non-causal Mamba (ADNM) implements a novel asymmetric bidirectional modeling framework that resolves directional bias in conventional Mamba architectures. This design preserves precise cloud boundary delineation while capturing long-range spatiotemporal dependencies in sparse data from NWP. (2) The Multi-scale Attention Enhancement Module (MAEM) enhances discriminative feature representation and suppresses spectral redundancy through anisotropic convolution kernels and hybrid pooling. This mechanism significantly improves edge retention in precipitation systems while attenuating atmospheric noise interference. (3) Complementing these advancements, the Wavelet Decomposition and Fusion Module (WDFM) maintains cloud contour integrity across scales through multiresolution decomposition. Extensive experiments demonstrate that ADNM-UNet outperforms existing methods across all metrics, achieving 27.96% improvement in CSI and 22.89% in HSS for metrics over the best performing baseline models at high intensity scenarios. Our project is open source and available on GitHub at: https://github.com/kanyu369/ADNM-UNet. Mingzhou Li, Xiaohui Huang 0003, Xiaofei Yang 0002, Jiangtao Peng, Yifang Ban, Nan Jiang 0013 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | ACTN: Adaptive Coupling Transformer Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) and Transformer networks have shown impressive performance in hyperspectral image (HSI) classification. However, these models usually concentrate on examining either local or global representations of HSI data, frequently falling short of capturing multidimensional representations. Furthermore, these methods fail to fully leverage the strengths of CNNs and Transformers. This article presents the adaptive coupling Transformer network (ACTN), a parallel-hybrid network aiming to improve representation learning for HSI classification. ACTN can capture different types of representation and facilitate mutual learning. Specifically, we introduce a parallel-hybrid module called the adaptive coupling module (ACM), which is designed to capture multifaceted representations from the HSI cube. The ACM consists of two branches: a CNN branch that extracts local contextual representations and a Transformer branch that captures global dependency representations. Our proposal is an adaptive response fusion module (ARFM) that interacts with the hybrid module to merge local and global representations at different resolutions in an adaptive way. In addition, we utilize a cosine similarity function to restrict the loss function in mutual learning, guaranteeing the preservation of both local and global representations to the maximum extent. Extensive experiments conducted on three public HSI datasets demonstrate that ACTN outperforms state-of-the-art methods based on Transformers and CNNs. Xiaofei Yang 0002, Weijia Cao, Yicong Zhou, Yao Lu 0008 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Mamba-UNet: Dual-Branch Mamba Fusion U-Net With Multiscale Spatio-Temporal Attention for Precipitation NowcastingabstractPrecipitation nowcasting is a challenging task in the context of global climate variability. However, existing radar echo or numerical weather prediction data methods lack deep modeling between echograms at different time points and have difficulty in accurately capturing irregular variations and small-scale features of precipitable clouds. To address these challenges, we propose for the first time a U-Net short-term precipitation prediction network based on vision Mamba technology for the precipitation nowcasting mission, named Mamba-UNet. Specifically, Mamba-UNet includes two core modules: the dual-branch Mamba fusion module and the multiscale spatiotemporal attention module. Finally, we propose a loss function namely dynamic quantile weighted loss to address the problem of imbalanced precipitation intensity distribution. To validate the capacity of the proposed method, the experiments were conducted on an analysis dataset of the local analysis and prediction system model in a specific region of East China. The experimental results show that our proposed Mamba-UNet has the best overall performance. Sihao Zhao, Xiaohui Huang 0003, Xiaofei Yang 0002, Nan Jiang 0013, Jiangtao Peng, Yifang Ban |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Multi-task Domain Adaptation for Language Grounding with 3D Objects
Penglei Sun, Yaoxian Song, Xinglin Pan, Peijie Dong, Xiaofei Yang 0002, Qiang Wang 0022, Zhixu Li, Tiefeng Li, Xiaowen Chu 0001 |
ECCV (34) | 5 |
| 2024 | 3D Question Answering for City Scene Understandingabstract3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existing research has primarily focused on indoor household tasks and outdoor roadside autonomous driving tasks, there has been limited exploration of city-level scene understanding tasks. Furthermore, existing research faces challenges in understanding city scenes, due to the absence of spatial semantic information and human-environment interaction information at the city level.To address these challenges, we investigate 3D MQA from both dataset and method perspectives. From the dataset perspective, we introduce a novel 3D MQA dataset named City-3DQA for city-level scene understanding, which is the first dataset to incorporate scene semantic and human-environment interactive tasks within the city. From the method perspective, we propose a Scene graph enhanced City-level Understanding method (Sg-CityU), which utilizes the scene graph to introduce the spatial semantic. A new benchmark is reported and our proposed Sg-CityU achieves accuracy of 63.94 % and 63.76 % in different settings of City-3DQA. Compared to indoor 3D MQA methods and zero-shot using advanced large language models (LLMs), Sg-CityU demonstrates state-of-the-art (SOTA) performance in robustness and generalization. Penglei Sun, Yaoxian Song, Xiang Liu 0001, Xiaofei Yang 0002, Qiang Wang 0022, Tiefeng Li, Yang Yang 0001, Xiaowen Chu 0001 |
ACM Multimedia | 4 |
| 2024 | RDTN: Residual Densely Transformer Network for hyperspectral image classification
Yan Li 0040, Xiaofei Yang 0002 |
Expert Syst. Appl. | 2 |
| 2024 | QTU-Net: Quaternion Transformer-Based U-Net for Water Body Extraction of RGB Satellite ImageabstractDeep learning models have achieved great success in water body extraction (WBE) from remote sensing images. However, the existing deep learning-based extraction methods exhibit limitations in their ability to fully explore the intricate interconnections inherent in RGB color satellite imagery and to enhance semantic representation across diverse regions. Furthermore, these methods often struggle with challenges posed by the uneven distribution of water bodies at different scales within the image, as well as substantial color disparities between water and land areas. In this article, we tackle WBE task from quaternion domain and introduce a novel approach called quaternion transformer-based U-Net (QTU-Net) to address these challenges. Our method specifically leverages quaternion convolution operations to capture the holistic relationships among RGB channels, thereby enhancing the semantic representation of WBE. Additionally, we propose a quaternion initialization module (QIM) to determine optimal RGB weights and facilitate the generation of quaternion data. To further improve the accuracy of water body delineation, we incorporate an innovative multiscale similarity aggregation attention (MSAA) component that enhances local similarity capture across various scales. Finally, we evaluate the proposed QTU-Net based on three publicly available benchmark datasets. The experimental results demonstrate that the proposed QTU-Net outperforms state-of-the-art baseline methods. Chunshan Li, Xiaofei Yang 0002, Zhiquan Zhou 0002, Raymond Y. K. Lau |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | DCTN: Dual-Branch Convolutional Transformer Network With Efficient Interactive Self-Attention for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is an essential task in remote sensing with substantial practical significance. However, most existing convolutional neural network (CNN)-based classification methods focus only on local spatial features while neglecting global spectral dependencies. Meanwhile, Transformer-based methods exhibit robust capabilities for global spectral feature modeling but struggle to extract local spatial features effectively. To fully exploit the local spatial feature extraction capabilities of CNN-based networks and the global spectral feature extraction capabilities of Transformer-based networks, this paper proposes a dual-branch convolutional Transformer method with efficient interactive self-attention for hyperspectral image classification, namely the dual-branch convolutional Transformer network (DCTN), which can aggregate local and global spatial-spectral features fully. Specifically, DCTN includes two core modules: the spatial-spectral fusion projection module and the efficient interactive self-attention module. The former utilizes 3D convolution with adaptive pooling and 2D group convolution with residual connection to parallel extract fused and grouped spatial-spectral features, respectively. The latter performs efficient interactive self-attention across height, width and spectral dimensions, enabling deep fusion of spatial-spectral features. Extensive experiments on three real HSI datasets demonstrate that the proposed DCTN method outperforms existing classification methods, yielding state-of-the-art classification performance. The code is available at https://github.com/AllFever/DeepHyperX-DCTN for reproducibility. Xiaohui Huang 0003, Xiaofei Yang 0002, Jiangtao Peng, Yifang Ban |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MSMT-LCL: Multiscale Spatial-Spectral Masked Transformer With Local Contrastive Learning for Hyperspectral Image ClassificationabstractDeep learning plays a crucial role in hyperspectral image (HSI) classification, with the Transformer being highly favored by researchers due to its exceptional ability to model long-range dependencies. However, the Transformer necessitates a substantial amount of labeled training samples to train its numerous parameters, exacerbating the challenge of training an effective HSI classification Transformer model, particularly given the inherent scarcity of HSI data. Therefore, we propose a novel method for HSI classification, termed multiscale spatial-spectral masked Transformer with local contrastive learning (MSMT-LCL). This method consists of two stages: self-supervised pretraining and supervised fine-tuning. Initially, we utilize the multiscale augmented feature mapping module (MAFM) to project original HSI data into two mixed-scale feature maps, which are then separately fed into two masked Transformer branches for reconstruction. To facilitate the model in learning the dependency relationships between central pixel land-cover information and neighboring land cover, we introduce a novel mask strategy based on center-patch. Furthermore, in the pretraining stage, we integrate local contrastive learning (LCL) to enable the model to focus on local center information at varying scales. Upon completion of pretraining, the network undergoes fine-tuning to obtain feature maps at two different scales. Subsequently, we devise a novel adaptive multiscale feature fusion module (AMFM) to adaptively aggregate these two features and produce the final classification results. Extensive experiments on three real datasets demonstrate the superiority of our proposed MSMT-LCL method over several state-of-the-art HSI classification methods. Xiaohui Huang 0003, Xiaofei Yang 0002, Jiangtao Peng, Yifang Ban, Nan Jiang 0013 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Efficient Harmonic Neural Networks With Compound Discrete Cosine Transform Filters and Shared Reconstruction FiltersabstractThe harmonic neural network (HNN) learns a combination of discrete cosine transform (DCT) filters to obtain an integrated feature from all spectra in the frequency domain. HNN, however, faces two challenges in learning and inference processes. First, the spectrum feature learned by HNN is insufficient and limited because the number of DCT filters is much smaller than that of feature maps. In addition, the number of parameters and the computation costs of HNN are significantly high because the intermediate spectrum layers are expanded multiple times. These two challenges will severely harm the performance and efficiency of HNN. To solve these problems, we first propose the compound DCT (C-DCT) filters integrating the nearest DCT filters to retrieve rich spectrum features to improve the performance. To significantly reduce the model size and computation complexity for improving the efficiency, the shared reconstruction filter is then proposed to share and dynamically drop the meta-filters in every frequency branch. Integrating the C-DCT filters with the shared reconstruction filters, the efficient harmonic network (EH-Net) is introduced. Extensive experiments on different datasets demonstrate that the proposed EH-Nets can effectively reduce the model size and computation complexity while maintaining the model performance. The code has been released at https://github.com/zhangle408/EH-Nets. Yao Lu 0008, Le Zhang 0016, Xiaofei Yang 0002, Yicong Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Multi-view dynamic graph convolution neural network for traffic flow prediction
Xiaohui Huang 0003, Yuming Ye, Xiaofei Yang 0002, Liyan Xiong |
Expert Syst. Appl. | 3 |
| 2023 | DS-UNet: Dual-Stream U-Net for Oil Spill Detection of SAR ImageabstractThe oil spill detection of synthetic aperture radar (SAR) images has great success. Existing deep learning-based methods make predictions mainly based on the U-Net structure and Transformer, which fail to blend the local and global information generated by other different feature maps. In this letter, we proposed a Dual Stream Unet (DS-Unet) for oil spill detection of SAR images. Specially, the proposed DS-Unet consists of two modules, an edge feature extraction module for extracting the local information and an Inter-scale Alignment module for capturing the global information. Moreover, an edge extraction branch is applied for handling the speckle noise of SAR images. Extensive experiments on two real-world datasets (Palsar and Sentinel) have shown that the proposed DS-Unet outperforms many existing state-of-the-art methods. Chunshan Li, Xiaofei Yang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | QTN: Quaternion Transformer Network for Hyperspectral Image ClassificationabstractNumerous state-of-the-art transformer-based techniques with self-attention mechanisms have recently been demonstrated to be quite effective in the classification of hyperspectral images (HSIs). However, traditional transformer-based methods severely suffer from the following problems when processing HSIs with three dimensions: (1) processing the HSIs using 1D sequences misses the 3D structure information; (2) too expensive numerous parameters for hyperspectral image classification tasks; (3) only capturing spatial information while lacking the spectral information. To solve these problems, we propose a novel Quaternion Transformer Network (QTN) for recovering self-adaptive and long-range correlations in HSIs. Specially, we first develop a band adaptive selection module (BASM) for producing Quaternion data from HSIs. And then, we propose a new and novel quaternion self-attention (QSA) mechanism to capture the local and global representations. Finally, we propose a new and novel transformer method, i.e., QTN by stacking a series of QSA for hyperspectral classification. The proposed QTN could exploit computation using Quaternion algebra in hypercomplex spaces. Extensive experiments on three public datasets demonstrate that the QTN outperforms the state-of-the-art vision transformers and convolution neural networks. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multi-Agent Mix Hierarchical Deep Reinforcement Learning for Large-Scale Fleet ManagementabstractIn recent years, ride-sharing has gained popularity as a daily means of transportation. The primary challenge for large-scale online ride-sharing platforms is to design an efficient fleet management policy that reallocates vehicles to appropriate regions to receive orders, thereby improving the platform’s cumulative revenue and order response rate. Combinatorial optimization algorithms and reinforcement learning methods are commonly employed for this task, but they typically learn a unified repositioning policy for all regions. However, different regions, such as hot and cold zones, may require different repositioning policies due to varying travel patterns. In this paper, we propose a multi-agent mixed hierarchical reinforcement learning approach, called MIX-H, for efficient large-scale fleet management by formulating it as a Markov decision process. MIX-H adopts multi-level controllers, including a leader controller and follower controller, for multi-level action learning. The leader controller plans the goal to be executed by the follower controller. Additionally, to improve the algorithm’s stability, we introduce a MIX module to compute the total value of joint action. Finally, experiments on real-world datasets demonstrate that the proposed method outperforms the state-of-the-art methods. Xiaohui Huang 0003, Jiahao Ling, Xiaofei Yang 0002, Kaiming Yang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | When Convolutional Network Meets Temporal Heterogeneous Graphs: An Effective Community Detection MethodabstractCommunity detection has long been an important yet challenging task to analyze complex networks with a focus on detecting topological structures of graph data. Essentially, real-world graph data is generally heterogeneous which dynamically varies over time, and this invalidates most existing community detection approaches. To cope with these issues, this paper proposes the temporal-heterogeneous graph convolutional networks (THGCN) to detect communities using the learnt feature representations of a set of temporal heterogeneous graphs. Particularly, we first design a heterogeneous GCN component to represent features of heterogeneous graph at each time step. Then, a residual compressed aggregation component is proposed to learn temporal feature representations extracted from two consecutive heterogeneous graphs. These temporal features are considered to contain evolutionary patterns of underlying communities. To the best of our knowledge, this is the first attempt to detect communities from temporal heterogeneous graphs. To evaluate the model performance, extensive experiments are performed on two real-world datasets, i.e., DBLP and IMDB. The promising results have demonstrated that the proposed THGCN is superior to both benchmark and the state-of-the-art approaches, e.g., GCN, GAT, GNN, LGNN, HAN and STAR, with respect to a number of evaluation criteria. Yaping Zheng, Xiaofeng Zhang 0002, Shiyi Chen, Xinni Zhang, Xiaofei Yang 0002, Di Wang 0004 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | A time-dependent attention convolutional LSTM method for traffic flow prediction
Xiaohui Huang 0003, Jie Tang 0009, Xiaofei Yang 0002, Liyan Xiong |
Appl. Intell. | 3 |
| 2022 | A multi-mode traffic flow prediction method with clustering based attention convolution LSTM
Xiaohui Huang 0003, Yuming Ye, Xiaofei Yang 0002, Liyan Xiong |
Appl. Intell. | 4 |
| 2022 | Multi-mode dynamic residual graph convolution network for traffic flow prediction
Xiaohui Huang 0003, Yuming Ye, Weihua Ding, Xiaofei Yang 0002, Liyan Xiong |
Inf. Sci. | 4 |
| 2022 | GCDB-UNet: A novel robust cloud detection approach for remote sensing images
Xian Li 0007, Xiaofei Yang 0002, Xutao Li 0003, Shijian Lu, Yunming Ye, Yifang Ban |
Knowl. Based Syst. | 2 |
| 2022 | DCRS: a deep contrast reciprocal recommender system to simultaneously capture user interest and attractiveness for online dating
Linhao Luo, Xiaofeng Zhang 0002, Dan Peng, Xiaofei Yang 0002 |
Neural Comput. Appl. | 6 |
| 2022 | Neural Style Transfer With Adaptive Auto-Correlation Alignment LossabstractThe neural style transfer has achieved a significant improvement with deep learning methods. However, the existing methods are susceptible to lack the ability for handling the texture style transfer because of their less consideration of the textural structure from style images. To overcome this drawback, this letter presents a simple method to capture the textural structure by using an adaptive auto-correlation alignment loss function. Furthermore, we also introduce three metrics to quantitatively evaluate the performance. We qualitatively and quantitatively evaluate the proposed methods. The experimental results demonstrate the superiority of the proposed method and our method can synthesize the stylized images with rich texture style patterns. Yue Wu 0001, Xiaofei Yang 0002, Yicong Zhou |
IEEE Signal Process. Lett. | 3 |
| 2022 | LWCDnet: A Lightweight Network for Efficient Cloud Detection in Remote Sensing ImagesabstractCloud detection is the task of detecting cloud areas in remote sensing images, and it has attracted extensive research interest. Recently, deep learning-based methods have been proposed and achieved great performance for cloud detection. However, due to the satellite’s limitation in storage and memory, existing deep learning approaches, which suffer from extensive computation and large model size, are almost impossible to be deployed on satellites. To fill this gap, we target at studying effective and efficient cloud detection solutions that are suitable for satellites. In this paper, we develop a lightweight autoencoder-based cloud detection method, namely LWCDnet. In the encoder part, the designed novel lightweight dual-branch block (LWDBB) in the backbone extracts spatial and contextual information concurrently. Moreover, a lightweight feature pyramid module (LWFPM) is proposed to capture high-level multi-scale contextual information. In the decoder part, the lightweight feature fusion module (LWFFM) compensates for the missing spatial and detail information from the encoder to the high-level feature maps. We evaluate the proposed method on two public datasets: LandSat8 and MODIS. Extensive experiments demonstrate that the proposed LWCDnet achieves comparable accuracy as the-state-of-art cloud detection methods and lightweight semantic segmentation algorithms. Meantime LWCDnet has much less computation burden with smaller model size. Shanshan Feng 0001, Xiaofei Yang 0002, Yunming Ye, Xutao Li 0003, Baoquan Zhang, Zhihao Chen 0010, Yingling Quan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Hyperspectral Image Transformer Classification NetworksabstractHyperspectral image (HSI) classification is an important task in earth observation missions. Convolution neural networks (CNNs) with the powerful ability of feature extraction have shown prominence in HSI classification tasks. However, existing CNN-based approaches cannot sufficiently mine the sequence attributes of spectral features, hindering the further performance promotion of HSI classification. This article presents a hyperspectral image transformer (HiT) classification network by embedding convolution operations into the transformer structure to capture the subtle spectral discrepancies and convey the local spatial context information. HiT consists of two key modules, i.e., spectral-adaptive 3-D convolution projection module and convolution permutator (ConV-Permutator) to retrieve the subtle spatial–spectral discrepancies. The spectral-adaptive 3-D convolution projection module produces the local spatial–spectral information from HSIs using two spectral-adaptive 3-D convolution layers instead of the linear projection layer. In addition, the Conv-Permutator module utilizes the depthwise convolution operations to separately encode the spatial–spectral representations along the height, width, and spectral dimensions, respectively. Extensive experiments on four benchmark HSI datasets, including Indian Pines, Pavia University, Houston2013, and Xiongan (XA) datasets, show the superiority of the proposed HiT over existing transformers and the state-of-the-art CNN-based methods. Our codes of this work are available athttps://github.com/xiachangxue/DeepHyperXfor the sake of reproducibility. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Self-Supervised Learning With Prediction of Image Scale and Spectral Order for Hyperspectral Image ClassificationabstractIn recent years, Convolutional Neural Networks (CNNs) have achieved great success in hyperspectral image classification attributed to their unparalleled capacity to extract the local information. However, to successfully learn the high-level semantic image features, they always require massive amounts of manually labeled data during the training process, which is expensive, scarce, and impractical, and severely hinders the improvement of supervised deep learning methods. To alleviate these burdens, we present Self-Supervised Learning methods for hyperspectral image classification by a pre-training model using extensive unlabeled data and fine-tuning the hyperspectral image target classification. In this paper, we propose a new method for learning image characteristics by training a CNN to recognize the image scale that is applied to the hyperspectral images (HSIs). In addition, we propose a multi-pretext task method to learn stable and good feature representations combing two different pretext task methods and contrastive loss function. We evaluate the proposed methods in Self-Supervised Learning benchmarks on four benchmark HSIs datasets. The experiment results demonstrate that the proposed methods outperform the traditional supervised deep learning methods when large amounts of unlabeled HSIs data are used. Moreover, it demonstrates that the Self-Supervised Learning method is promising to alleviate dependence on manually labeled data of hyperspectral image classification. Finally, our research contributes to the creation and refinement of Self-Supervised Learning methods for pretextual tasks within the HSIs community. Xiaofei Yang 0002, Weijia Cao, Yao Lu 0008, Yicong Zhou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Road Detection via Deep Residual Dense U-NetabstractRoad extraction from aerial images is a hot research topic. With the advancement of convolutional neural network (CNN), several CNN-based road detection methods have been developed. However, most of them do not make full use of the hierarchical features from the original aerial images. In this paper, we propose a novel residual dense U-Net (RDUN), a semantic segmentation network which combines the strengths of residual learning, DenseNet, and U-Net, to overcome the drawback. Our proposed RDUN can fully exploit the hierarchical features from all the convolutional layers, which utilizes the residual dense blocks (RDB) to build up a U-Net architecture. The benefits of our model are two-fold. First, by using the RDB abundant local features can be extracted and fused effectively. Second, based the local features, hierarchical features are constructed by shortcut connections between layers in RDB. Extensive experiments are carried out on a real-world road detection dataset and the results demonstrate the proposed RDUN outperforms state-of-the-art competitors. Xiaofei Yang 0002, Xutao Li 0003, Yunming Ye, Xiaofeng Zhang 0002, Haijun Zhang 0002, Xiaohui Huang 0003, Bowen Zhang 0005 |
IJCNN | 1 |
| 2019 | Road Detection and Centerline Extraction Via Deep Recurrent Convolutional Neural Network U-NetabstractRoad information extraction based on aerial images is a critical task for many applications, and it has attracted considerable attention from researchers in the field of remote sensing. The problem is mainly composed of two subtasks, namely, road detection and centerline extraction. Most of the previous studies rely on multistage-based learning methods to solve the problem. However, these approaches may suffer from the well-known problem of propagation errors. In this paper, we propose a novel deep learning model, recurrent convolution neural network U-Net (RCNN-UNet), to tackle the aforementioned problem. Our proposed RCNN-UNet has three distinct advantages. First, the end-to-end deep learning scheme eliminates the propagation errors. Second, a carefully designed RCNN unit is leveraged to build our deep learning architecture, which can better exploit the spatial context and the rich low-level visual features. Thereby, it alleviates the detection problems caused by noises, occlusions, and complex backgrounds of roads. Third, as the tasks of road detection and centerline extraction are strongly correlated, a multitask learning scheme is designed so that two predictors can be simultaneously trained to improve both effectiveness and efficiency. Extensive experiments were carried out based on two publicly available benchmark data sets, and nine state-of-the-art baselines were used in a comparative evaluation. Our experimental results demonstrate the superiority of the proposed RCNN-UNet model for both the road detection and the centerline extraction tasks. Xiaofei Yang 0002, Xutao Li 0003, Yunming Ye, Raymond Y. K. Lau, Xiaofeng Zhang 0002, Xiaohui Huang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | A new weighting k-means type clustering framework with an l2-norm regularization
Xiaohui Huang 0003, Xiaofei Yang 0002, Junhui Zhao 0001, Liyan Xiong, Yunming Ye |
Knowl. Based Syst. | 2 |
| 2018 | Hyperspectral Image Classification With Deep Learning ModelsabstractDeep learning has achieved great successes in conventional computer vision tasks. In this paper, we exploit deep learning techniques to address the hyperspectral image classification problem. In contrast to conventional computer vision tasks that only examine the spatial context, our proposed method can exploit both spatial context and spectral correlation to enhance hyperspectral image classification. In particular, we advocate four new deep learning models, namely, 2-D convolutional neural network (2-D-CNN), 3-D-CNN, recurrent 2-D CNN (R-2-D-CNN), and recurrent 3-D-CNN (R-3-D-CNN) for hyperspectral image classification. We conducted rigorous experiments based on six publicly available data sets. Through a comparative evaluation with other state-of-the-art methods, our experimental results confirm the superiority of the proposed deep learning models, especially the R-3-D-CNN and the R-2-D-CNN deep learning models. Xiaofei Yang 0002, Yunming Ye, Xutao Li 0003, Raymond Y. K. Lau, Xiaofeng Zhang 0002, Xiaohui Huang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Clustering time-stamped data using multiple nonnegative matrices factorization
Xiaohui Huang 0003, Yunming Ye, Liyan Xiong, Shaokai Wang, Xiaofei Yang 0002 |
Knowl. Based Syst. | 5 |