EDBT 2026 Demo / reviewers in the wild / expert
Tao Gao 0001
dblp:08/17-1 · also Gao Tao 0001
· DBLP profile ↗
50ranked-venue papers
13as first author
43since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 5 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explicitly learning semantic relevance for salient object detection in remote sensing images
Tao Gao 0001, Weiguang Zhao, Mengkun Liu, Ting Chen 0003 |
Expert Syst. Appl. | 1 |
| 2026 | Unpaired iterative prompt learning for real-world image deraining
Yuanbo Wen 0002, Tao Gao 0001, Shan Liang 0002, Ting Chen 0003 |
Expert Syst. Appl. | 2 |
| 2026 | Prior-oriented specific and triple-view general prompts for multi-weather degraded image restoration
Yuanbo Wen 0002, Tao Gao 0001, Shan Liang 0002, Ting Chen 0003 |
Expert Syst. Appl. | 2 |
| 2026 | Multi-perspective prompt and assimilated self-modulation transformer for adverse weather removal
Yuanbo Wen 0002, Tao Gao 0001, Shan Liang 0002, Zixiang Liu, Ting Chen 0003 |
Expert Syst. Appl. | 2 |
| 2026 | A wavelet-guided and physics-aware network for remote sensing image dehazing
Qianxi Zhang, Ting Chen 0003, Tao Gao 0001, Yuanbo Wen 0002, Shan Liang 0002 |
Expert Syst. Appl. | 3 |
| 2026 | EN-thinking: Enhancing entity-level reasoning in large language models for knowledge graph completion
Yaoyu Chang, Tao Gao 0001 |
Knowl. Based Syst. | 2 |
| 2026 | Dual-attention cooperative and multi-view gated transformer for adverse weather removal
Shan Liang 0002, Yuanbo Wen 0002, Tao Gao 0001, Ting Chen 0003 |
Pattern Recognit. | 3 |
| 2026 | Prompt-oriented and frequency-regularized schrödinger bridge for unpaired rain streaks and raindrops removal
Yuanbo Wen 0002, Ting Chen 0003, Tao Gao 0001 |
Pattern Recognit. | 4 |
| 2026 | DynStaticNet: A biological vision-inspired dual-branch all-in-one network for video weather removal
Qianxi Zhang, Tao Gao 0001, Ting Chen 0003, Yuanbo Wen 0002, Tao Lei 0003 |
Pattern Recognit. | 2 |
| 2026 | Structure-Preserving Frequency-Regularized Text-Guided Optimal Transport for Unpaired Rain Streaks and Raindrops RemovalabstractThe removal of rain streaks and raindrops is crucial for enhancing the image visibility and mitigating the weather degradations. However, most existing approaches rely on the paired rainy and clean images, which are challenging to obtain in real-world scenarios. To this end, we propose a novel structure-preserving frequency-regularized text-guided optimal transport (SFTOT) framework, which formulates the unpaired rain streaks and raindrops removal as an optimal transport problem. Specifically, we introduce a structure-preserving transport cost, incorporating the structural similarity constraint to minimize the duality gap between the primal and dual formulations, while preserving the structural details of reconstructed images. Furthermore, by embedding the inherent frequency sparsity of rain streaks and raindrops into the transport cost, we derive a frequency-regularized optimal transport objective, ensuring consistency in frequency distributions between the generated and clean images. Additionally, we employ a pre-trained one-step stable diffusion model as the restoration network, which is fine-tuned using the low-rank adaptation (LoRA) adapters and zero convolutional layers, while integrating the domain-specific text prompts for both degraded and clean images to guide the generation process. Extensive experiments demonstrate that our method surpasses the existing well-performing unpaired learning approaches, achieving notable improvements in both the fidelity and photo-realism. Yuanbo Wen 0002, Tao Gao 0001, Qianxi Zhang, Jing Zhang 0052, Ting Chen 0003, Lidong Liu |
IEEE Trans. Multim. | 2 |
| 2025 | Multi-axis Prompt and Multi-dimension Fusion Network for All-in-one Weather-degraded Image RestorationabstractExisting approaches aiming to remove adverse weather degradations compromise the image quality and incur the long processing time. To this end, we introduce a multi-axis prompt and multi-dimension fusion network (MPMF-Net). Specifically, we develop a multi-axis prompts learning block (MPLB), which learns the prompts along three separate axis planes, requiring fewer parameters and achieving superior performance. Moreover, we present a multi-dimension feature interaction block (MFIB), which optimizes intra-scale feature fusion by segregating features along height, width and channel dimensions. This strategy enables more accurate mutual attention and adaptive weight determination. Additionally, we propose the coarse-scale degradation-free implicit neural representations (CDINR) to normalize the degradation levels of different weather conditions. Extensive experiments demonstrate the significant improvements of our model over the recent well-performing approaches in both reconstruction fidelity and inference time. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003 |
AAAI | 2 |
| 2025 | Brain-Inspired Spiking Neural Networks for Energy-Efficient Object DetectionabstractBrain-inspired spiking neural networks (SNNs) have the capability of energy-efficient processing of temporal information. However, leveraging the rich dynamic characteristics of SNNs and prior works in artificial neural networks (ANNs) to construct an effective object detection model for visual tasks remains an open question for further exploration. To develop a directly-trained , low energy consumption and high-performance multi-scale SNN model, we propose a novel interpretable object detection framework Multi-scale Spiking Detector (MSD). Initially, we propose a spiking convolutional neuron as a core component of the Optic Nerve Nucleus Block (ONNB), designed to significantly enhance the deep feature extraction capabilities of SNNs. ONNB enables direct training with improved energy efficiency, demonstrating superior performance compared to state-of-the-art ANN-to-SNN conversion and SNN techniques. In addition, we propose a Multi-scale Spiking Detection Framework to emulate the biological response and comprehension of stimuli from different objects. Wherein, spiking multi-scale fusion and the spiking detector are employed to integrate features across different depths and to detect response outcomes, respectively. Our method outperforms state-of-the-art ANN detectors, with only 7.8 M parameters and 6.43 mJ energy consumption. MSD obtains the mean average precision (mAP) of 62.0% and 66.3% on COCO and Gen1 datasets, respectively. Tao Gao 0001, Yisheng An, Ting Chen 0003, Jing Zhang 0052, Yuanbo Wen 0002, Mengkun Liu, Qianxi Zhang |
CVPR | 2 |
| 2025 | Traffic image encryption based on activation function-type chaotic map and reversible cellular automaton
Tingyu An, Tao Gao 0001, Ting Chen 0003, Donghua Jiang 0001, Lulu Xu, Yuxiu Chen |
Expert Syst. Appl. | 2 |
| 2025 | A physics prompt-based network for interpretable and effective image dehazing
Shan Liang 0002, Tao Gao 0001, Ting Chen 0003, Qianxi Zhang |
Expert Syst. Appl. | 2 |
| 2025 | MSNet: Multi-Scale Network for Object Detection in Remote Sensing Images
Tao Gao 0001, Shilin Xia, Mengkun Liu, Jing Zhang 0052, Ting Chen 0003 |
Pattern Recognit. | 1 |
| 2025 | Hyperspectral Tracker With Constrained Object Adaptive Learning and Trajectory ConstructionabstractHyperspectral imaging offers significant potential for precise object tracking, yet the scarcity of dataset volumes specifically tailored for hyperspectral tracking algorithms hinders progress, particularly for deep models with complex structures. Additionally, current deep learning-based hyperspectral trackers typically enhance model accuracy via online or adversarial learning, adversely affecting tracking speed. To address these challenges, this paper introduces the Constrained Object Adaptive Learning hyperspectral Tracker (COALT), an effective parameter-efficient fine-tuning tracker tailored for hyperspectral tracking. COALT integrates Pixel-level Object Constrained Spectral Prompt (POCSP) and Temporal Sequence Trajectory Prompt (TSTP) through Adaptive Learning with Parameter-efficient Fine-tuning (ALPEFT), enabling a transformer-based tracker to capture detailed spectral features and relationships in hyperspectral image sequences through trainable rank decomposition matrices. Specifically, POCSP is designed to retain optimal spectral information with low internal correlation and high object representativeness, enabling rapid image reconstruction. Then, the most representative spectral template and search are fused into a single stream as spectral prompts for the Encoder and Decoder layers. Concurrently, the previous coordinates within the same sequence are tokenized and utilized as temporal prompts by TSTP in the decoder layers. The model is trained with ALPEFT to optimize spectral information learning, which substantially reduces the number of training parameters, alleviating overfitting issues arising from limited data. Meanwhile, the proposed tracker not only retains the ability of pre-trained model to estimate object trajectories in an autoregressive manner but also effectively utilizes spectral information and enhances target location perception during the fine-tuning process. Extensive experiments and evaluations are conducted on two public hyperspectral tracking datasets. The results demonstrate that the proposed COALT tracker achieves satisfactory performance with leading processing speed. The code will be available at https://github.com/PING-CHUANG/COALT. Ye Wang 0020, Mingyang Ma 0004, Ge Zhang 0006, Tao Gao 0001, Shaohui Mei |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Hyperspectral Object Tracking With Context-Aware Learning and Category ConsistencyabstractHyperspectral imaging technology is of crucial importance to improve the performance of object tracking in many remote sensing surveillance areas. Previous methods primarily focused on feature fusion strategies by employing additional enhancement modules. However, these methods commonly lack contextual understanding to distinguish the target from the background and totally ignore the category information of the targets. To address these limitations, a novel hyperspectral object tracker is proposed to incorporate context-aware learning and category consistency tracker (CCTrack), which can adaptively learn context-aware representations in hyperspectral scenarios to obtain global target information with memory storage, while constructing an interframe category consistency constraint to enhance tracking process. Specifically, CCTrack integrates an adaptive context-aware learning (ACL) mechanism, which includes a feature decoupling module (FDM) to extract specific representations from decoupled features, and a Mamba layer to retain and update long-range dependencies. To align with prior knowledge of target recognition and motion patterns, an alignment transformation module (ATM) is employed with the ACL mechanism, fully leveraging spatial-spectral representations. In addition, category consistency constraint modules (C3Ms) are introduced to enforce category consistency across frames by computing the similarities between the target features and the corresponding category name, serving as the constraint to improve tracking performance. Extensive experiments over the hyperspectral object tracking (HOT) benchmark covering various remote sensing scenarios demonstrate that CCTrack outperforms state-of-the-art methods by a significant margin. Ye Wang 0020, Shaohui Mei, Mingyang Ma 0004, Tao Gao 0001, Huiyang Han |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Cross-Level Interaction and Intralevel Fusion Network for Remote Sensing Image DehazingabstractExisting approaches have significantly advanced remote sensing image dehazing. However, they often rely on conventional encoder-decoder architectures, leading to prolonged inference times. To this end, we propose a novel cross-level interaction and intra-level fusion network for remote sensing image dehazing (CINet), which shifts the focus from encoder-decoder dependencies to an innovative hierarchical architecture centered on skip connections, leading to competitive dehazing performance with decreased calculating complexity. Furthermore, we introduce a cross-level multi-view interaction module (CMIM) to facilitate effective interactions between features across hierarchical levels, mitigating the information loss commonly caused by repeated down-sampling operations. Meanwhile, we develop an intra-level dual-dimension fusion module (IDFM), which leverages height-wise and width-wise self-attention to capture rich spatial-aware information, enabling robust and efficient intra-level feature fusion. Additionally, we propose a multi-view progressive extraction block (MPEB), which decomposes features into four distinct components and applies convolutions with diverse kernel sizes, groups, and dilation factors. This design promotes progressive feature learning while significantly reducing computational overhead. Extensive experiments conducted on nine publicly available datasets validate the effectiveness and superiority of our proposed model. Yuanbo Wen 0002, Tao Gao 0001, Ting Chen 0003, Mengkun Liu, Lidong Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | FUNet: Frequency-Aware and Uncertainty-Guiding Network for Rain-Hazy Image Restoration
Mengkun Liu, Tao Gao 0001, Licheng Jiao |
IEEE Trans. Multim. | 2 |
| 2025 | All-in-One Weather-Degraded Image Restoration Via Adaptive Degradation-Aware Self-Prompting ModelabstractExisting approaches for all-in-one weather-degraded image restoration suffer from inefficiencies in leveraging degradation-aware priors, resulting in sub-optimal performance in adapting to different weather conditions. To this end, we develop an adaptive degradation-aware self-prompting model (ADSM) for all-in-one weather-degraded image restoration. Specifically, our model employs the contrastive language-image pre-training model (CLIP) to facilitate the training of our proposed latent prompt generators (LPGs), which represent three types of latent prompts to characterize the degradation type, degradation property and image caption. Moreover, we integrate the acquired degradation-aware prompts into the time embedding of diffusion model to improve degradation perception. Meanwhile, we employ the latent caption prompt to guide the reverse sampling process using the cross-attention mechanism, thereby guiding the accurate image reconstruction. Furthermore, to accelerate the reverse sampling procedure of diffusion model and address the limitations of frequency perception, we introduce a wavelet-oriented noise estimating network (WNE-Net). Extensive experiments conducted on eight publicly available datasets demonstrate the effectiveness of our proposed approach in both task-specific and all-in-one applications. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Kaihao Zhang, Ting Chen 0003 |
IEEE Trans. Multim. | 2 |
| 2024 | A Novel Scene-aware Pedestrian Detection in Dense ScenesabstractIn crowded scenes, detecting pedestrians with high density and various occlusions is always an important yet challenging task. To further improve the performance of one-stage detectors in detecting crowded pedestrians, we propose a one-stage dense pedestrian detection network called YOLO- DensePed (You Only Look Once-Dense Pedestrian Detection) to further overcome shortcoming including limited receptive field, insufficient feature fusion, and ambiguous assignment of anchor boxes for object detection. First, the proposed YOLO-DensePed utilizes a multi-head self-attention module with embedded Gaussian masks to reduce background redundant information, as well as enhance capture capability for global contextual information; Then, Deformable ConvNets v2 (DCNv2) are used instead of standard convolutions in the neck layer, which can dynamically adjust the receptive field and learn the correct feature and multi-scale position information of objects; Furthermore, SimOTA dynamic sample assignment strategy and Soft-NMS post-processing algorithm are also introduced to assist the YOLO-DensePed for better handling occlusion and dense distribution issues. Extensive experiments on the public CrowdHuman dataset demonstrate that YOLO-DensePed consistently presents the best or comparable performance, allowing for efficient and accurate detection of crowded pedestrian. Ting Chen 0003, Jinghua Chen, Tao Gao 0001, Shukang Zhu, Zongyang Guo, Zixiang Liu, Quanzhao Zhao |
CSCWD | 3 |
| 2024 | Multi-Dimension Queried and Interacting Network for Stereo Image DerainingabstractEliminating the rain degradation in stereo images poses a formidable challenge, which necessitates the efficient exploitation of mutual information present between the dual views. To this end, we devise MQINet, which employs multi-dimension queries and interactions for stereo image deraining. More specifically, our approach incorporates a context-aware dimension-wise queried block (CDQB). This module leverages dimension-wise queries that are independent of the input features and employs global context-aware attention (GCA) to capture essential features while avoiding the entanglement of redundant or irrelevant information. Meanwhile, we introduce an intra-view physics-aware attention (IPA) based on the inverse physical model of rainy images. IPA extracts shallow features that are sensitive to the physics of rain degradation, facilitating the reduction of rain-related artifacts during the early learning period. Furthermore, we integrate a cross-view multi-dimension interacting attention mechanism (CMIA) to foster comprehensive feature interaction between the two views across multiple dimensions. Extensive experimental evaluations demonstrate the superiority of our model over EPRRNet and StereoIRR, achieving respective improvements of 4.18 dB and 0.45 dB in PSNR. Code and models are available at https://github.com/chdwyb/MQINet. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003 |
ICASSP | 2 |
| 2024 | Encoder-Minimal and Decoder-Minimal Framework for Remote Sensing Image DehazingabstractHaze obscures remote sensing images, hindering valuable information extraction. To this end, we propose RSHazeNet, an encoder-minimal and decoder-minimal framework for efficient remote sensing image dehazing. Specifically, regarding the process of merging features within the same level, we develop an innovative module called intra-level transposed fusion module (ITFM). This module employs adaptive transposed self-attention to capture comprehensive context-aware information, facilitating the robust context-aware feature fusion. Meanwhile, we present a cross-level multi-view interaction module (CMIM) to enable effective interactions between features from various levels, mitigating the loss of information due to the repeated sampling operations. In addition, we propose a multi-view progressive extraction block (MPEB) that partitions the features into four distinct components and employs convolution with varying kernel sizes, groups, and dilation factors to facilitate view-progressive feature learning. Extensive experiments demonstrate the superiority of our proposed RSHazeNet. We release the source code and all pre-trained models at https://github.com/chdwyb/RSHazeNet. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003 |
ICASSP | 2 |
| 2024 | Unpaired Photo-realistic Image Deraining with Energy-informed Diffusion ModelabstractExisting unpaired image deraining approaches face challenges in accurately capture the distinguishing characteristics between the rainy and clean domains, resulting in residual degradation and color distortion within the reconstructed images. To this end, we propose an energy-informed diffusion model for unpaired photo-realistic image deraining (UPID-EDM). Initially, we delve into the intricate visual-language priors embedded within the contrastive language-image pre-training model (CLIP), and demonstrate that the CLIP priors aid in the discrimination of rainy and clean images. Furthermore, we introduce a dual-consistent energy function (DEF) that retains the rain-irrelevant characteristics while eliminating the rain-relevant features. This energy function is trained by the non-corresponding rainy and clean images. In addition, we employ the rain-relevance discarding energy function (RDEF) and the rain-irrelevance preserving energy function (RPEF) to direct the reverse sampling procedure of a pre-trained diffusion model, effectively removing the rain streaks while preserving the image contents. Extensive experiments demonstrate that our energy-informed model surpasses the existing unpaired learning approaches in terms of both supervised and no-reference metrics. Yuanbo Wen 0002, Tao Gao 0001, Ting Chen 0003 |
ACM Multimedia | 2 |
| 2024 | A Novel Algorithm for Thermal Imaging Fault Detection in Photovoltaic PanelsabstractDue to the complex background and low resolution of photovoltaic (PV) panel image, the fault detection accuracy of photovoltaic panel is low. Therefore, we devise an efficient asymptotic feature pyramid network (EAFPN) to enhance the detection accuracy. Initially, an efficient spatial and channel attention mechanism (ESCAM) module is proposed, which mitigates the impact of complex backgrounds in the feature extraction network. Meanwhile, we introduce an asymptotic feature pyramid network (AFPN) module into the feature fusion network, which effectively extracts semantic information, captures the low-level localization features, and prevents the lose and degradation of small object information. In addition, we integrate content-aware reassembly of features (CARAFE) module to the feature fusion network, which can further enhance the accuracy. Furthermore, we generate the Chang'an University Photovoltaic Panel Thermal Imaging Defect (CHD-PVTID) datasets, which containing six common faults, and conducted experiments. The results show that EAFPN algorithm has higher accuracy in identifying fault types than other algorithms. Ting Chen 0003, Yuxiu Chen, Quanzhao Zhao, Tao Gao 0001, Lulu Xu |
MSN | 4 |
| 2024 | A novel dual-stage progressive enhancement network for single image deraining
Tao Gao 0001, Yuanbo Wen 0002, Jing Zhang 0052, Ting Chen 0003 |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Neural Schrödinger bridge for unpaired real-world image deraining
Yuanbo Wen 0002, Tao Gao 0001, Ting Chen 0003 |
Inf. Sci. | 2 |
| 2024 | A Self-Supplementary and Revised Network for Remote Sensing Object DetectionabstractObject detection is an essential and crucial task in interpretation of optical remote sensing images (RSIs). However, its performance is usually limited due to the complex background and multiscale characteristics of targets. To overcome these limitations, a self-supplementary and revised anchor-free detector is proposed. First, to reduce the computational cost of detection, a partial bottleneck (PBottleneck) structure is designed to efficiently extract multiscale feature information in a lightweight manner. Second, pure spatial feature pyramid network (PSFPN) attaches importance to description of distance and suppresses environmental disturbance by a devised multidirectional distance attention (MDDA) mechanism. In addition, pure fusion strategy (PFS) is created to boost information with no occlusion between various features. Third, toward the multiscale objects issue, self-learning supplementary and revised module (SSRM) is explored to generate more abundant and balanced expression by adaptively incorporating the supplementary and corrected information from adjacent features. Finally, comprehensive experiments are conducted on several publicly available datasets, demonstrating effectiveness of our proposed detector, leading to a new benchmark. Tao Gao 0001, Zixiang Liu, Guiping Wu, Yuanbo Wen 0002, Lidong Liu, Ting Chen 0003, Jing Zhang 0052 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Restoring vision in rain-by-snow weather with simple attention-based sampling cross-hierarchy Transformer
Yuanbo Wen 0002, Tao Gao 0001, Kaihao Zhang, Peng Cheng 0002, Ting Chen 0003 |
Pattern Recognit. | 2 |
| 2024 | From heavy rain removal to detail restoration: A faster and better network
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Kaihao Zhang, Ting Chen 0003 |
Pattern Recognit. | 2 |
| 2024 | Frequency-Oriented Efficient Transformer for All-in-One Weather-Degraded Image RestorationabstractAdverse weather conditions, such as rain, raindrop, snow and haze, consistently degrade images in an unpredictable manner, thereby rendering existing task-specific and task-aligned methods inadequate in addressing this formidable problem. To this end, we investigate the application of Transformer in image restoration and introduce an efficient frequency-oriented method called AIRFormer, which is designed to restore weather-degraded images comprehensively and holistically. Specifically, we identify that the initial self-attention mechanism exhibits distinctive properties akin to a low-pass filter. Therefore, we construct a frequency-guided Transformer encoder by incorporating wavelet-based prior information to guide the extraction of image features. Additionally, considering the non-specific frequency characteristics of self-attention in the later stages, we develop a frequency-refined Transformer decoder that incorporates learnable task-specific queries across spatial dimensions, channel dimensions, and wavelet domains. To facilitate the training of our proposed method, we curate a comprehensive benchmark dataset named AIR40K that, encompasses a wide range of challenging scenarios. Extensive experimental evaluations demonstrate the superiority of our AIRFormer over both task-aligned and all-in-one methods across 15 publicly available datasets. Notably, AIRFormer achieves the best trade-off between the inference time and quality of reconstructed image, comparing with existing methods such as TransWeather and Restormer. The source code, dataset and pre-trained models will be available at https://github.com/chdwyb/AIRFormer. Tao Gao 0001, Yuanbo Wen 0002, Kaihao Zhang, Jing Zhang 0052, Ting Chen 0003, Lidong Liu, Wenhan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Attention-Free Global Multiscale Fusion Network for Remote Sensing Object DetectionabstractRemote sensing object detection (RSOD) encounters challenges in complex backgrounds and small object detection, which are interconnected and unable to address separately. To this end, we propose an attention-free global multiscale fusion network (AGMF-Net). Initially, we present a spatial bias module (SBM) to obtain long-range dependencies as a part of our proposal global information extraction module (GIEM). GIEM efficiently captures the global information, overcoming challenges posed by complex backgrounds. Moreover, we propose multitask enhanced structure (MES) and multitask feature pretreatment (MFP) to enhance the feature representation of multiscale targets, while eliminating the interference from complex backgrounds. In addition, an efficient context decoupled detector (ECDD) is presented to provide distinct features for regression and classification tasks, aiming to improve the efficiency of RSOD. Extensive experiments demonstrate that our proposed method achieves superior performance compared with the state-of-the-art detectors. Specifically, AGMF-Net obtains the mean average precision (mAP) of 73.2%, 92.03%, 95.21%, and 94.30% on detection in optical remote sensing images (DIOR), high resolution remote sensing detection (HRRSD), Northwestern Polytechnical University Very High Resolution-10 (NWPU VHR-10), and RSOD datasets, respectively. Tao Gao 0001, Yuanbo Wen 0002, Ting Chen 0003, Qianqian Niu, Zixiang Liu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Long-Term Vehicle Tracking Algorithm of Correlation Filter Optimized by Swarm Intelligence Tracking FrameworkabstractIn response to the complex situations encountered in real traffic scenarios, such as sudden acceleration or deceleration of vehicles, similar-looking vehicles, and occlusions caused by other vehicles or obstacles, which severely affect tracking accuracy, a long-term vehicle tracking algorithm based on correlation filter (CF) is proposed and optimized using the swarm intelligence (SI) tracking framework. The fast-discriminative scale space tracking (FDSST) of CF is adopted as the core tracker. Multiple features are adaptively weighted and fused. The optimal feature template is dynamically updated based on confidence. An in-depth analysis of the intrinsic relationship is conducted between SI and object tracking, leading to the design of an SI tracking framework. Within this framework, the carnivorous plant algorithm (CPA) is employed as an optimization method, further enhancing CPA functionality through phototaxis strategy and population partition mechanism. During FDSST tracking, when tracking uncertainty surpasses a set threshold, the SI tracking framework is integrated to rectify tracking outcomes, and a short-term memory module is designed to predict the object position if the object disappears for dozens of frames. The experimental results on benchmark datasets (UAV20L and LaSOT) demonstrate a success rate (SR) of 67.45%, a precision (PR) rate of 70.50%, and a speed of 22.22 ft/s. Comparative analyses with other notable tracking algorithms confirm the exceptional accuracy and robustness of the proposed approach that can effectively address diverse challenges in vehicle tracking scenarios, by achieving highly reliable tracking outcomes and enhancing long-term vehicle tracking capability. Mingbo Niu, Md. Sipon Miah, Tao Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Remote Sensing Image Dehazing Method Based on Heterogeneous PriorsabstractRemote sensing image dehazing is crucial for both military and civil applications. However, dehazed remote sensing images often suffer from pronounced artifacts and tend to overestimate the atmospheric light value. We propose a novel dehazing method based on heterogeneous priors. Specifically, superpixels are extracted from the hazy remote sensing image using a depth-based simple linear iterative clustering superpixel segmentation (DSLIC) algorithm. These superpixels serve as cells for transmission and atmospheric light estimation. To improve the robustness of atmospheric light estimation, we develop an atmospheric light value-map fusion estimation (ALFE) model that integrates the heterogeneous priors-guided haze concentration model (HP-HCM) to derive the global atmospheric light value, while utilizing the bright channel value within each superpixel as the local atmospheric light map. We also introduce a dynamic dehazing intensity parameter (DDIP) model, which refine the transmission map based on the HP-HCM. Extensive comparative experiments validate the superior performance of the proposed method. The PSNR and SSIM achieved by our method exceed those of the dark channel prior (DCP) by 22.2% and 37.5%, respectively. Shan Liang 0002, Tao Gao 0001, Ting Chen 0003, Peng Cheng 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | LRRU: Long-short Range Recurrent Updating Networks for Depth CompletionabstractExisting deep learning-based depth completion methods generally employ massive stacked layers to predict the dense depth map from sparse input data. Although such approaches greatly advance this task, their accompanied huge computational complexity hinders their practical applications. To accomplish depth completion more efficiently, we propose a novel lightweight deep network framework, the Long-short Range Recurrent Updating (LRRU) network. Without learning complex feature representations, LRRU first roughly fills the sparse input to obtain an initial dense depth map, and then iteratively updates it through learned spatially-variant kernels. Our iterative update process is content-adaptive and highly flexible, where the kernel weights are learned by jointly considering the guidance RGB images and the depth map to be updated, and large-to-small kernel scopes are dynamically adjusted to capture long-to-short range dependencies. Our initial depth map has coarse but complete scene depth information, which helps relieve the burden of directly regressing the dense depth from sparse ones, while our proposed method can effectively refine it to an accurate depth map with less learnable parameters and inference time. Experimental results demonstrate that our proposed LRRU variants achieve state-of-the-art performance across different parameter regimes. In particular, the LRRU-Base model outperforms competing approaches on the NYUv2 dataset, and ranks 1st on the KITTI depth completion benchmark at the time of submission. Project page: https://npucvr.github.io/LRRU/. Bo Li 0090, Ge Zhang 0006, Qi Liu 0054, Tao Gao 0001, Yuchao Dai |
ICCV | 5 |
| 2023 | Progressive dilation dense residual fusion network for single-image derainingabstractAbstract Rain removal is very important for many applications in computer vision, and it is a challenging problem due to its ill‐posed nature, especially for single‐image deraining. In order to remove rain streaks more thoroughly, as well as to retain more details, a progressive dilation dense residual fusion network is proposed. The entire network is designed in a cascade manner with multiple fusion blocks. The fusion block consists of a dilation dense residual block (DDRB) and a dense residual feature fusion block (DRFFB), where DDRB is created for feature extraction and DRFFB is mainly designed for feature fusion operation. Meanwhile, detail compensation memory mechanism (DCMM) is leveraged between each of two cascade modules to retain more background details. Compared with previous state‐of‐the‐art methods, extensive experiments show that the proposed method can achieve better results, in terms of rain streaks removal and background details preservation. Furthermore, the authors’ network also shows its superiority for image noise removal. Xiaolin Kong, Tao Gao 0001, Ting Chen 0003, Jing Zhang 0052 |
IET Image Process. | 2 |
| 2023 | Task Alignment Interaction and Cross-Scale Guided Enhancement for Remote Sensing Object DetectionabstractObject detection is a fundamental task in the analysis and interpretation of remote sensing images. However, compared to natural images, remote sensing images are characterized by broad diversity in object scales, fuzzy objects, and complex background, which bring great challenges to object detection. For overcoming the above problems, a task alignment interaction and cross-scale guidance enhancement network (TCNet) is proposed in this letter. Firstly, a generalized mean spatial pyramid pooling (GeMSPP) is designed and embedded in the backbone to adapt to changes of complex environment and reduce loss of features. Secondly, cross-scale guided enhancement network (CGEN) is proposed to generate high-quality non-aliasing multi-scale target features for each feature level by guiding the fusion of deep features and enhancing feature expression. Thirdly, Task alignment interactive head (TAIH) is adopted to enhance the classification and regression accuracy of the prediction box, so as to suppress background interference and highlight object features. Experiments conducted on public DIOR and RSOD datasets illustrate that the proposed modules can effectively improve the accuracy of detection and our network has superior performance compared with other state-of-the-art detectors. Guiping Wu, Lidong Liu, Zixiang Liu, Tao Gao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Multi-Scale Density-Aware Network for Single Image DehazingabstractDehazing based on deep learning has attracted a lot of attention recently. Most dehazing networks seldom consider two critical features of real outdoor-scene haze,i.e., depth and haze density, resulting in degraded performance on real hazy images compared with synthetic hazy images. Moreover, the uncertainty problem is crucial in the image restoration field, but it is often ignored. In this letter, we propose a novel multi-scale density-aware network (MSDAN) for single image dehazing, where a key dual feedback module (DFB) is proposed and embedded in the decoder part of MSDAN. Furthermore, the DFB includes a feedforward mechanism and two feedback mechanisms: feature feedback (FF) and transmission feedback (TF). Specifically, the feedforward mechanism predicts a low-scale transmission map ($t$-map), while FF and TF aim to enhance confident features to reduce model uncertainty in the training process and correct features by introducing depth and density information. In addition, two novel modules: confident feature attention module (CFA) and transmission adjustment module (TADJ) are proposed as cores for confident features estimation of FF and TF, respectively. Extensive quantitative and qualitative experiments are conducted on several public datasets, which demonstrate that the proposed algorithm outperforms the state-of-the-art algorithms. Tao Gao 0001, Peng Cheng 0002, Ting Chen 0003, Lidong Liu |
IEEE Signal Process. Lett. | 1 |
| 2023 | A Task-Balanced Multiscale Adaptive Fusion Network for Object Detection in Remote Sensing ImagesabstractObject detection is essential in the interpretation of remote sensing images. However, the blurred background and objects with vast variances are identified as the two main challenges of the task. We propose a novel detector adapted to complicated background and multi-scale objects, namely, task-balanced multi-scale adaptive fusion network (TMAFNet), targeting directly on the above two challenges. Firstly, a depth separable global context module (DSGC) is constructed to understand contextual relations among pixels from a global perspective, which is extraordinarily necessary to distinguish objects from the environment. Most importantly, DSGC reduces the computational cost by decoupling the acquisition of global information into single-channel global interaction and multi-channel single-point interaction. Secondly, in order to eliminate disturbance and enhance representation ability of objects, hidden recursive feature pyramid network (HRFPN) is explored, which encodes the information of difference before and after using the multi-scale fusion. HRFPN is proven to enhance the target features by reducing the background noise. Thirdly, a semi-coupling task-balanced head (SCTB) is presented to guarantee the consistency of detection. We have conducted comprehensive experiments on several publicly available datasets, and the results illustrate that our modules improve adaptability and robustness of the network, leading to a new state-of-the-art. Tao Gao 0001, Zixiang Liu, Jing Zhang 0052, Guiping Wu, Ting Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Global to Local: A Scale-Aware Network for Remote Sensing Object DetectionabstractWith the wide application of remote sensing images (RSIs) in military and civil fields, remote sensing object detection (RSOD) has gradually become a hot research direction. However, we observe two main challenges for remote sensing object detection, namely the complicated background and the small objects issues. Given the different appearances of generic objects and remote sensing objects, the detection algorithms designed for the former usually cannot perform well for the latter. We propose a novel global to local scale-aware detection network (GLSANet) for remote sensing object detection, aiming to solve the above mentioned two challenges. Firstly, we design a global semantic information interaction module (GSIIM) to excavate and reinforce the high-level semantic information in the deep feature map, which alleviates the obstacles of complex background on foreground objects. Secondly, we optimize the feature pyramid network to improve the performance of multiscale object detection in RSIs. Finally, a local attention pyramid (LAP) is introduced to highlight the feature representation of small objects gradually while suppressing the background and noise in the shallower feature maps. Extensive experiments on three public datasets demonstrate that the proposed method achieves superior performance compared with the state-of-the-art detectors, especially on small object detection dataset. Specifically, our algorithm reaches 94.57% mAP on NWPU VHR-10 dataset, 95.93% mAP on RSOD dataset and 77.9% mAP on DIOR dataset, respectively. Tao Gao 0001, Qianqian Niu, Jing Zhang 0052, Ting Chen 0003, Shaohui Mei, Ahmad Jubair |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Encoder-Free Multiaxis Physics-Aware Fusion Network for Remote Sensing Image DehazingabstractCurrent methods for remote sensing image dehazing confront noteworthy computational intricacies and yield suboptimal dehazed outputs, thereby circumscribing their pragmatic applicability. To this end, we propose EMPF-Net, a novel encoder-free multi-axis physics-aware fusion network that exhibits both light-weighted characteristics and computational efficiency. In our pipeline, we contend that conventional u-shaped networks allocate substantial computational resources to encode haze-degraded features, which play a subordinate role in the reconstruction process. Consequently, our encoder stages solely incorporate down-sampling operations. To improve the representation efficiency and enhance the generalization capabilities, we devise a multi-axis partial queried learning block (MPQLB) that primarily concentrates on learning dimension-wise queries, instead of relying solely on strictly-correlated content of the input features. Furthermore, we augment the reconstruction procedure by incorporating ground truth supervision into each stage via a supervised cross-scale transposed attention module (SCTAM). It calculates attention maps under the guidance of clean images, thereby suppressing less informative features to propagate to the subsequent level. In addition, to address the challenge of ineffective intral-level feature fusion, which result in insufficient elimination of haze-degraded information and negatively impact the quality of reconstructed images, we introduce a physics-aware intra-level fusion module (PIFM). This module harnesses a physical inversion model to facilitate the intra-level feature interaction and alleviate the interference of dehazing-irrelevant information. Our proposed EMPF-Net is evaluated on 12 publicly available datasets, and the experimental results substantiate our superiority in terms of both metrical scores and visual quality, despite being equipped with a modest parameter count of 300 K. Our approach is readily accessible at https://github.com/chdwyb/EMPF-Net. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Ting Chen 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Noise-insensitive image representation via multiple extended LDB and class supervised intelligent coordination feature selection
Yongxiong Liu, Ting Chen 0003, Tao Gao 0001 |
J. Supercomput. | 4 |
| 2021 | A novel face recognition method based on fusion of LBP and HOGabstractAbstract As one of the hot topics in the field of computer vision research, face recognition technology has received significant attention due to its potentiality for a wide range of applications in government as well as commercial purposes. In practical applications, although several existing face recognition methods have achieved good performances in specific scenes, they easily suffer from a sharp decline in recognition rate if affected by different conditions of light, expression, posture and occlusion. Among many factors, influences of complex illuminations on face recognition are particularly significant. To further improve the performance of the existing local binary pattern (LBP) operator, neighbourhood weighted average LBP (NWALBP) is first proposed for fully considering the strong correlations between pixel pairs in the neighbourhood, which extends the traditional LBP uni‐layer neighbourhood template window to the bi‐layer neighbourhood template window and calculates the weighted average of bi‐layer neighbourhood pixels in each direction. Then, inspired by center symmetric LBP (CS‐LBP), centre symmetric NWALBP (CS‐NWALBP) is further proposed, which can effectively reduce computation complexity by only comparing the weighted average values of the neighbourhood pixels that are symmetric about the centre pixel. Finally, by combining the merit of histogram of oriented gradient (HOG), a feature fusion algorithm named CS‐NWALBP+HOG is suggested. Several experiments have eventually demonstrated that our proposed algorithms have more robust performance under complex illumination conditions if compared with many other latest algorithms. Ting Chen 0003, Tao Gao 0001, Jinpei Cao, Dachun Yao, Y. H. Li |
IET Image Process. | 2 |
| 2020 | Description method of Illumination invariant image features
Tao Gao 0001, Shan Liang 0002, Ting Chen 0003, Mengni Liu, Yong Hui Li |
Signal Process. Image Commun. | 1 |
| 2019 | Restoration algorithm for noisy complex illuminationabstractAlthough promising results have been achieved in the restoration of complex illumination images with the Retinex algorithm, there are still some drawbacks in the processing of Retinex. Considering the noise characteristics of complex illumination images, in this study, we propose a novel restoration algorithm for noisy complex illumination, which combines guided adaptive multi‐scale Retinex (GAMSR) and improvement BayesShrink threshold filtering (IBTF) based on double‐density dual‐tree complex wavelet transform (DDDTCWT) domain. Extensive restoration experiments are conducted on three typical types images and the same image with different noises. On the basis of a series of evaluation indexes, we compare our method to those of state‐of‐the‐art algorithms. The results show that (i) SSIM of the proposed IBTF is superior to traditional Bayes threshold method by 15% as the standard variance is 100. (ii) PSNR of the proposed GAMSR enhances 15% to traditional MSR. (iii) The clarity of final results for restoration speeds up three times than that of original images, and the information entropy is improved slightly too. Therefore, the proposed method can effectively enhance the details, edges and textures of the image under complex illumination and noises. Zhanwen Liu, Tao Gao 0001, Fanjie Kong, Ziheng Jiao, Aodong Yang, Bo Liu 0006 |
IET Comput. Vis. | 2 |
| 2019 | Single sample description based on Gabor fusionabstractOwing to lack of enough face image and invalidation of many traditional face recognition algorithms, face recognition with single training sample is really a great challenge. To solve the above problem, this study proposes a novel local weighted fusion Gabor (LWFG) algorithm. First, one single sample is segmented into a series of block sub‐images, and then, each of these sub‐images is decomposed into a series of multi‐resolution Gabor wavelets with multi‐orientation and multi‐scale. Second, different orientation Gabor wavelets with the same scale are fused. Next, different scale Gabor wavelets with the same orientation are fused according to the proposed fusion criterion. Third, the fusion Gabor feature histograms are calculated in each of the divided local regions. Meanwhile, every local region's information importance is measured by the proposed local image information content model. Finally, the fusion Gabor wavelet histograms are adaptively weighed by weighting map which calculated from information content model. This study conducted simulation experiments on different face databases under the different conditions including partial occlusion, expression change and illumination variation. The results indicated that the proposed LWFG algorithm is more effective with single training sample. Ting Chen 0003, Tao Gao 0001, Xiangmo Zhao |
IET Image Process. | 2 |
| 2017 | 4G UAV communication system and hovering height optimization for public safetyabstractWhen facing with sudden terrorist attacks or natural disasters, in order to avoid the paralysis of communication networks caused by the destruction of partial ordinary 4G cellular base stations in urban area, this paper proposed a unmanned aerial vehicle (UAV) communication system based on 4G technology for public safety and investigated its optimal hovering height for maximizing its effective coverage radius. In this system, collaborative operation, among several airborne 4G cellular base stations, some unmanned aerial vehicle relays and other still worked ordinary 4G cellular base stations, could form a seamless communication coverage in the accident area, and provide the alternate communication links with QoS guarantee to people involved in the disaster relief and rescue. At meanwhile, in different urban environments, the hovering height of UAV equipped with the 4G cellular base station could be quickly optimal adjusted to maximize its effective coverage area by the configuration information sent from the emergency command management center, so as to effectively cut the cost of urban security emergency response system with the limited number of UAVs, ensure its smooth operation, and save people's life and property loss to the greatest extent. Ting Chen 0003, Xiangmo Zhao, Tao Gao 0001, Zhigang Xu 0001 |
Healthcom | 4 |
| 2017 | Image feature representation with orthogonal symmetric local weber graph structure
Tao Gao 0001, Xiangmo Zhao, Ting Chen 0003, Zhanwen Liu |
Neurocomputing | 1 |
| 2017 | Face description based on adaptive local weighted Gabor comprehensive histogram feature
Tao Gao 0001, X. M. Zhao, Ting Chen 0003, Z. W. Liu, Ce Ni |
Multim. Tools Appl. | 1 |
| 2017 | Illumination-insensitive image representation via synergistic weighted center-surround receptive field model and weber law
Tao Gao 0001, Xiangmo Zhao, Ting Chen 0003, Zhanwen Liu |
Pattern Recognit. | 1 |