Tao Ye 0002

dblp:15/9-2 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-1814-530XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Image and video processing · 100%
Artificial intelligence
2 papers
Generative modeling · 54% Segmentation and scene understanding · 46%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › image restoration
adverse weather image restoration
1.922026
AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception · IEEE Trans. Image Process. 2026
UMCFuse: A Unified Multiple Complex Scenes Infrared and Visible Image Fusion Framework · IEEE Trans. Image Process. 2025
Image and video processing › image restoration
degradation removal
1.012026
JDPNet: A Network Based on Joint Degradation Processing for Underwater Image Enhancement · IEEE Trans. Image Process. 2026
Image and video processing
image enhancement
1.012026
JDPNet: A Network Based on Joint Degradation Processing for Underwater Image Enhancement · IEEE Trans. Image Process. 2026
Image and video processing
image restoration
1.012026
JDPNet: A Network Based on Joint Degradation Processing for Underwater Image Enhancement · IEEE Trans. Image Process. 2026
Image and video processing › image fusion
multi-modal image fusion
1.012026
AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception · IEEE Trans. Image Process. 2026
Image and video processing › image enhancement
underwater image enhancement
1.012026
JDPNet: A Network Based on Joint Degradation Processing for Underwater Image Enhancement · IEEE Trans. Image Process. 2026
Image and video processing › image fusion › multi-modal image fusion
infrared and visible image fusion
0.912025
UMCFuse: A Unified Multiple Complex Scenes Infrared and Visible Image Fusion Framework · IEEE Trans. Image Process. 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312026
AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception · IEEE Trans. Image Process. 2026
Computer vision › Segmentation and scene understanding
semantic segmentation
0.312025
UMCFuse: A Unified Multiple Complex Scenes Infrared and Visible Image Fusion Framework · IEEE Trans. Image Process. 2025

Methods — techniques the papers use, named apart from their topics

text perception · 2.0ChatGPT · 2.0BLIP captioning · 2.0multi-directional energy fusion · 1.7adaptive denoising · 1.7probabilistic bootstrap distribution · 1.0loss function design · 1.0joint degradation processing · 1.0
YearPublicationVenuePosition
2026 CAWM-Mamba: A unified model for infrared-visible image fusion and compound adverse weather restoration
Huichun Liu, Xiaosong Li 0004, Zhuangfan Huang, Tao Ye 0002, Yang Liu 0335, Haishu Tan
Expert Syst. Appl.4
2026 Real-Time Underground Fire Detection on Coal Mine IoVT Systems: An Edge-Deployed Efficient YOLO-Architecture
abstract
Underground fires pose a significant threat to production safety in coal mines, and existing detection methods suffer from drawbacks such as poor adaptability to complex subterranean environments and excessive model parameters. To address the need for deploying object detection models on resource-constrained devices, this paper proposes a novel and efficient algorithm forUndergroundFireYOLOdetection, named UF-YOLO. The core innovation of this method is threefold: first, the StarNet module is introduced into the backbone to significantly reduce model parameters and computational complexity without sacrificing accuracy; second, the Cross-scale Context Fusion Module (CCFM) is integrated into the neck to enhance the model’s detection capability for fires of various scales, particularly small targets; and finally, Partial Convolution (PConv) is integrated to extract spatial features more efficiently, further reducing redundant computations and memory access. On our self-built Mine Fire Image Dataset (MFID), compared to the baseline model YOLOv11m, UF-YOLO reduces parameters by 77.1%, increases inference speed by 60.6%. Experimental results on the public COCO val 2017 dataset demonstrate that the proposed method outperforms state-of-the-art (SOTA) models such as YOLOv12. The results confirm that UF-YOLO can be efficiently deployed on the edge-side of coal mine IoVT monitoring systems to performe accurate and real-time fire detection. This work provides a new intelligent paradigm for the real-time monitoring of underground fires.
Wei Yang 0063, Jiaqi Wu 0012, Zehua Wang 0001, Qi-Chong Tian, Tao Ye 0002, Wei Chen 0036, F. Richard Yu, Victor C. M. Leung
IEEE Internet Things J.6
2026 Digital Twin-Enabled Joint Resource Optimization in THz-RIS Networks: A Graph Neural Network-Augmented PPO Approach
Fucheng Xue, Meichen Gai, Wei Chen 0036, Tao Ye 0002
IEEE Internet Things J.7
2026 FlexiSR-Diff: Flexible diffusion for multi-modal medical image fusion & super-resolution
Yushen Xu, Xiaosong Li 0004, Yang Liu 0335, Tao Ye 0002, Huafeng Li 0001
Pattern Recognit.5
2026 AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception
abstract
Multi-modality image fusion (MMIF) in adverse weather aims to address the loss of visual information caused by weather-related degradations, providing clearer scene representations. Although a few studies have attempted to incorporate textual information to improve semantic perception, they often lack effective categorization and thorough analysis of textual content. To address these limitations, we propose AWM-Fuse, a unified fusion framework that handles diverse weather degradations via global and local text perception with shared parameters. In particular, a global text perception module leverages BLIP-generated captions to extract overall scene features and identify primary degradation types, thus promoting generalization across various adverse weather conditions. Complementing this, the local module employs detailed scene descriptions produced by ChatGPT to concentrate on specific degradation effects through concrete textual cues, enabling the recovery of subtle details. Furthermore, textual descriptions are used to constrain the generation of fused images, effectively steering the network learning process toward better alignment with semantic labels, thereby promoting the learning of more meaningful visual features. To facilitate text-guided fusion under adverse weather, we construct AWMM-Text, a large-scale benchmark providing paired global and local annotations for multi-modality image pairs. Extensive experiments demonstrate that AWM-Fuse consistently outperforms state-of-the-art methods under complex weather conditions and on multiple downstream tasks. Our code is available at https://github.com/Feecuin/AWM-Fuse.
Xilai Li, Huichun Liu, Xiaosong Li 0004, Tao Ye 0002, Zhenyu Kuang, Huafeng Li 0001
IEEE Trans. Image Process.4
2026 JDPNet: A Network Based on Joint Degradation Processing for Underwater Image Enhancement
abstract
Given the complexity of underwater environments and the variability of water as a medium, underwater images are inevitably subject to various types of degradation. The degradations present nonlinear coupling rather than simple superposition, which renders the effective processing of such coupled degradations particularly challenging. Most existing methods focus on designing specific branches, modules, or strategies for specific degradations, with little attention paid to the potential information embedded in their coupling. Consequently, they struggle to effectively capture and process the nonlinear interactions of multiple degradations from a bottom-up perspective. To address this issue, we propose JDPNet, a joint degradation processing network, that mines and unifies the potential information inherent in coupled degradations within a unified framework. Specifically, we introduce a joint feature-mining module, along with a probabilistic bootstrap distribution strategy, to facilitate effective mining and unified adjustment of coupled degradation features. Furthermore, to balance color, clarity, and contrast, we design a novel AquaBalanceLoss to guide the network in learning from multiple coupled degradation losses. Experiments on six publicly available underwater datasets, as well as two new datasets constructed in this study, show that JDPNet exhibits state-of-the-art performance while offering a better tradeoff between performance, parameter size, and computational cost.
Tao Ye 0002, Hongbin Ren, Chongbing Zhang, Xiaosong Li 0004
IEEE Trans. Image Process.1
2025 UMCFuse: A Unified Multiple Complex Scenes Infrared and Visible Image Fusion Framework
abstract
Infrared and visible image fusion has emerged as a prominent research area in computer vision. However, little attention has been paid to the fusion task in complex scenes, leading to sub-optimal results under interference. To fill this gap, we propose a unified framework for infrared and visible images fusion in complex scenes, termed UMCFuse. Specifically, we classify the pixels of visible images from the degree of scattering of light transmission, allowing us to separate fine details from overall intensity. Maintaining a balance between interference removal and detail preservation is essential for the generalization capacity of the proposed method. Therefore, we propose an adaptive denoising strategy for the fusion of detail layers. Meanwhile, we fuse the energy features from different modalities by analyzing them from multiple directions. Extensive fusion experiments on real and synthetic complex scenes datasets cover adverse weather conditions, noise, blur, overexposure, fire, as well as downstream tasks including semantic segmentation, object detection, salient object detection, and depth estimation, consistently indicate the superiority of the proposed method compared with the recent representative methods. Our code is available at https://github.com/ixilai/UMCFuse.
Xilai Li, Xiaosong Li 0004, Tianshu Tan, Huafeng Li 0001, Tao Ye 0002
IEEE Trans. Image Process.5
2024 Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion
abstract
Multi-modal image fusion (MMIF) integrates valuable information from different modality images into a fused one. However, the fusion of multiple visible images with different focal regions and infrared images is a unprecedented challenge in real MMIF applications. This is because of the limited depth of the focus of visible optical lenses, which impedes the simultaneous capture of the focal information within the same scene. To address this issue, in this paper, we propose a MMIF framework for joint focused integration and modalities information extraction. Specifically, a semi-sparsity-based smoothing filter is introduced to decompose the images into structure and texture components. Subsequently, a novel multi-scale operator is proposed to fuse the texture components, capable of detecting significant information by considering the pixel focus attributes and relevant data from various modal images. Additionally, to achieve an effective capture of scene luminance and reasonable contrast maintenance, we consider the distribution of energy information in the structural components in terms of multi-directional frequency variance and information entropy. Extensive experiments on existing MMIF datasets, as well as the object detection and depth estimation tasks, consistently demonstrate that the proposed algorithm can surpass the state-of-the-art methods in visual perception and quantitative evaluation. The code is available at https://github.com/ixilai/MFIF-MMIF.
Xilai Li, Xiaosong Li 0004, Tao Ye 0002, Xiaoqi Cheng, Wuyang Liu, Haishu Tan
WACV3
2024 Parallelization Strategy of Non-Local Means Filtering Algorithm for Real-Time Denoising of Forward-Looking Multi-Beam Sonar Images
abstract
Obtaining clear sonar images is crucial for ocean exploration applications, such as marine resource detection and underwater target searches. Traditional filtering methods cannot effectively eliminate the noise generated by the complex underwater environment in sonar images and can potentially result in problems such as image blurring. Existing methods that effectively filter sonar image noise often lack real-time performance, making them impractical for ocean exploration. To address these limitations, this study proposes a real-time denoising technique for forward-looking multi-beam sonar images based on a non-local means filtering algorithm. The integral image is used to calculate the mean square error (MSE), which improves algorithm efficiency and ensures that the runtime remains unaffected by the neighbourhood window size. To further improve real-time performance, the algorithm is migrated to a graphics processing unit (GPU) and a block-wise computation method is proposed to calculate the integral image. Simultaneously, to enhance GPU thread utilisation, the three-dimensional thread structure from the compute unified device architecture (CUDA) programming model is utilised and additional threads are allocated to enhance computation. The captured images are filtered using an M1200d sonar device manufactured by Oculus. Extensive experiments demonstrate that the proposed method achieves excellent performance regarding both denoising accuracy and efficiency. Specifically, the proposed method achieves a peak signal-to-noise ratio higher than 25 dB and a structural similarity index of more than 0.85 at 50 frames per second, thus demonstrating its significant potential for real-time sonar image denoising.
Tao Ye 0002, Xiangpeng Deng, Xiao Cong, Hongkun Zhou, Xiangming Yan
IEEE Trans. Circuits Syst. Video Technol.1
2022 A Stable Lightweight and Adaptive Feature Enhanced Convolution Neural Network for Efficient Railway Transit Object Detection
abstract
Obstacles in front of a train pose a significant threat to traffic safety, and many accidents happen under shunting mode when the speed of a train is below 45 km/h. The existing track object–detection algorithms encounter difficulty in balancing the detection precision and speed in shunting mode. Additionally, their accuracy is insufficient, particularly for small objects in complex environments. To address these problems, we propose a stable lightweight feature extraction and adaptive feature fusion network for real-time detection of obstacles in railway traffic scenarios to ensure driving safety. The proposed network consists of three modules. The stable bottom feature extraction module reduces the computational load and extracts more image information stably. The lightweight feature extraction module improves feature extraction using a simple and effective network. The enhanced adaptive feature fusion module fuses the image and original features, improving the multiscale detection accuracy under complex environments, particularly in the case of small objects. With a default input size of 416$\times 416$pixels (px), the proposed method achieves a detection speed of 81 FPS and a mean average precision of 94.75% for the railway traffic dataset as well as a detection speed of 78 FPS (26 FPS faster and 0.47% higher than those of YOLOv4, respectively) and a mean average precision of 42.5% for MS COCO. This indicates its potential for real-world railway object detection and other multi-target detection tasks. Additionally, the experimental results based on PASCAL VOC2007 and VOC2012 indicate that the proposed approach is considerably better than the state-of-the-art models.
Tao Ye 0002, Zongyang Zhao, Shouan Wang, Fuqiang Zhou, Xiao Zhi Gao 0001
IEEE Trans. Intell. Transp. Syst.1
2022 Foreign Body Detection in Rail Transit Based on a Multi-Mode Feature-Enhanced Convolutional Neural Network
abstract
Detection of railway traffic objects is an important task during train driving and is implemented to ensure safe driving. Although object detection has been investigated for years, many challenges exist in precisely detecting railway objects under complex railway scenes. These challenges mainly include adverse weather states, various railway backgrounds, diverse railway objects, and low-quality images. To address these issues, we introduce a novel deep learning method, called a multi-mode feature enhanced convolutional neural network (MMFE-Net), for accurate railway object detection. The network mainly consists of three modules. 1) An improved cross-stage partial connection darknet53 (CSPDarknet53), called adaptive dilated cspdarknet53, is used as our backbone to reduce image information loss. 2) A spatial feature extraction module is used to improve the feature extraction ability of the model for blurred objects and objects in a complicated background. 3) We introduce an attention fusion enhance module to strengthen the context information between adjacent feature maps to accurately detect multiscale and small objects. The proposed method achieves 0.9439 mAP and 79 FPS with an input size of$640\times 640$pixels on the railway traffic dataset, and its performance is better than that of YOLOv4. Moreover, it is feasible to apply MMFE-Net into practical applications of railway object detection.
Tao Ye 0002, Jun Zhang 0003, Zongyang Zhao, Fuqiang Zhou
IEEE Trans. Intell. Transp. Syst.1
2021 Railway Traffic Object Detection Using Differential Feature Fusion Convolution Neural Network
abstract
Railway shunting accidents, in which trains collide with obstacles, often occur because of human error or fatigue. It is therefore necessary to detect traffic objects in front of the trains and inform the driver to take timely action. To detect these objects in railways, we proposed an object-detection method using a differential feature fusion convolutional neural network (DFF-Net). DFF-Net includes two modules: the prior object-detection module and the object-detection module. The prior module produces initial anchor boxes for the subsequent detection module. Taking the initial anchor boxes as input, the object-detection module applies a differential feature fusion sub-module to enrich the sematic information for object detection, enhancing the detection performance, particularly for small objects. In experiments conducted on a railway traffic dataset, compared with the current state-of-the-art detectors, the proposed method exhibited significant higher performance and was more effective and more efficient than the other methods for object detection in railway tracks. Additionally, evaluation results based on PASCAL VOC2007 and VOC2012 indicated that the proposed method was significantly better than the state-of-the-art methods.
Tao Ye 0002
IEEE Trans. Intell. Transp. Syst.1