Yan Wu 0011

dblp:04/3001-11 · DBLP profile ↗
← Back
39ranked-venue papers
11as first author
14since 2021 · last 2026
0000-0002-8874-8886ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Ink removal by design: Leveraging structural cues for efficient and generalizable ink removal in whole slide pathology
Huer Wen, Yan Wu 0011, De-Shuang Huang
Expert Syst. Appl.2
2026 DAPU: Distribution-aware patch upsampling for point cloud-based 3D object detection
Yan Wu 0011, Yujian Mo
Neurocomputing2
2026 LPRFusion: Asymmetric cascade RoI refinement with LiDAR-pseudo point cloud fusion for 3D object detection
Yan Wu 0011, Yujian Mo, Junqiao Zhao, Jun Yan 0009
Neurocomputing2
2025 RamIR: Reasoning and action prompting with Mamba for all-in-one image restoration
Aiqiang Tang, Yan Wu 0011
Appl. Intell.2
2025 DSFusion: a dynamic dual-scale multimodal fusion framework for robust 3D object detection
Yan Wu 0011, Yujian Mo
Multim. Syst.2
2025 NPENN: A Noise Perturbation Ensemble Neural Network for Microbiome Disease Phenotype Prediction
abstract
With advances in microbiomics, the crucial role of microbes in disease progression is increasingly recognized. However, predicting disease phenotypes using microbiome data remains challenging due to data complexity, heterogeneity, and limited model generalization. Current methods often depend on specific datasets and are vulnerable to adversarial attacks. To address these issues, this paper introduces a novel Noise Perturbation Ensemble Neural Network model (NPENN), which combines noise mechanisms with Gradient Boosting (GB) techniques for robust neural network ensemble learning. NPENN, validated on multiple microbiome datasets, shows superior accuracy and generalization compared to traditional methods, effectively handling data complexity and variability. This approach enhances model robustness and feature learning by integrating GB prior knowledge. Additionally, the study explores microbial community roles in various diseases, providing insights into disease mechanisms and potential biomarkers for personalized precision diagnosis and treatment strategies.
Yan Wu 0011, Qinhu Zhang, Siguo Wang, Zhen-Hao Guo
IEEE J. Biomed. Health Informatics2
2024 Sparse Query Dense: Enhancing 3D Object Detection with Pseudo Points
abstract
Current LiDAR-only 3D detection methods are limited by the sparsity of point clouds. The previous method used pseudo points generated by depth completion to supplement the LiDAR point cloud, but the pseudo points sampling process was complex, and the distribution of pseudo points was uneven. Meanwhile, due to the imprecision of depth completion, the pseudo points suffer from noise and local structural ambiguity, which limit the further improvement of detection accuracy. This paper presents SQDNet, a novel framework designed to address these challenges. SQDNet incorporates two key components: the SQD, which achieves sparse-to-dense matching via grid position indices, allowing for rapid sampling of large-scale pseudo points on the dense depth map directly, thus streamlining the data preprocessing pipeline. And use the density of LiDAR points within these grids to alleviate the uneven distribution and noise problems of pseudo points. Meanwhile, the sparse 3D Backbone is designed to capture long-distance dependencies, thereby improving voxel feature extraction and mitigating local structural blur in pseudo points. The experimental results validate the effectiveness of SQD and achieve considerable detection performance for difficult-to-detect instances on the KITTI test.
Yujian Mo, Yan Wu 0011, Junqiao Zhao, Zhenjie Hou, Weiquan Huang, Jun Yan 0009
ACM Multimedia2
2023 Robust Traffic Light Recognition Pipeline Based on YOLOv8 for Autonomous Driving Systems
abstract
Traffic Light Recognition (TLR) aims at detecting Traffic Lights (TLs) and then classifying the status of light signals, being an essential constituent of autonomous driving perception systems. However, it’s challenging for existing TLR methods to accurately distinguish both color and shape status of TLs due to small object sizes, illumination variations, close resemblance with other objects and varying weather conditions. Existing public datasets for TLR have three main drawbacks:(i) poor diversity, (ii) sample imbalance, and (iii) insufficient category labels, greatly hindering the development of TLR. To overcome the aforementioned problems, we propose a Robust Traffic Light Recognition Pipeline based on YOLOv8 (RTLRP-YOLO) that can recognize TLs accurately with strong robustness based on adaptively generated high-quality images. Specifically, we develop a Self-Adaptive Preprocessing Module (SAPM) which is designed to adaptively generate high-quality images under hostile conditions, followed by a Two-stage Traffic Light Recognition Model based on YOLOv8 (TTRM) to obtain both the location and status information of TLs. Moreover, We also provide our self-made Tongji Small Traffic Light Dataset (TSTLD), covering a variety of weather conditions, regions, light intensities and shooting angles. To the best of our knowledge, our proposed method is the first one to be able of simultaneously identifying three colors (i.e., red, yellow and green) and four shapes (i.e., circle, left arrow, right arrow and up arrow) of TLs, achieving 95.43% accuracy on TSTLD with the inference time of 26 ms for per image.
Yan Wu 0011, Junqiao Zhao
ICPADS2
2022 Review the state-of-the-art technologies of semantic segmentation based on deep learning
Yujian Mo, Yan Wu 0011, Xinneng Yang, Feilin Liu, Yujun Liao
Neurocomputing2
2022 Road Friction Coefficient Estimation Via Weakly Supervised Semantic Segmentation and Uncertainty Estimation
abstract
Vision-based road friction coefficient estimation received extensive attention in the field of road maintenance and autonomous driving. However, the current mainstream coarse-grained friction estimation methods are basically based on image classification tasks. This makes it difficult to deal with complex road conditions in changing weathers. Many models can correctly predict the friction coefficients of the road as a whole in consistent and simple road conditions, but perform poorly otherwise. The existing image benchmarks in this field rarely consider the above problems as well, which limits the comparable evaluations of different models. Therefore, in this paper, we first construct a challenging pixel-level friction coefficient estimation dataset WRF-P to evaluate model performances under mixed road conditions. Then, we propose a friction coefficient estimation method based on weakly supervised learning and uncertainty estimation to realize pixel-level road friction prediction with low annotation cost. The model outperforms existing weakly supervised methods and reaches 39.63% mIOU on the WRF-P dataset. The WRF-P dataset will be made publicly available at https://github.com/blackholeLFL/The-WRF-dataset soon.
Feilin Liu, Yan Wu 0011, Yujian Mo, Yujun Liao
Int. J. Pattern Recognit. Artif. Intell.2
2022 Efficient Adaptive Upsampling Module for Real-Time Semantic Segmentation
abstract
Upsampling operation is necessary for semantic segmentation and other pixel-level prediction tasks. Among the commonly used upsampling operations, some are too simple to effectively recover the spatial details lost during downsampling process, and some are too complex and have high computation complexity. In real-world applications, it is critical to achieve high accuracy and maintain real-time inference speed. Therefore, an efficient upsampling operation is essential for these tasks. In this paper, we introduce efficient adaptive upsampling module (EAUM) for real-time semantic segmentation. Inspired by dynamic filter networks, EAUM adaptively predicts the kernel weight of each point in the upsampled feature map according to the corresponding points in the input feature map. To reduce computational cost, EAUM decomposes the spatial information and channel information required for upsampling. The proposed EAUM shows impressive performance on Cityscapes and CamVid benchmarks. Specifically, DenseENet with EAUM outperforms the baseline by 1.4% [Formula: see text] and 1.6% [Formula: see text] in accuracy with a slight drop in inference speed on Cityscapes test dataset.
Xinneng Yang, Yan Wu 0011, Junqiao Zhao, Feilin Liu, Yujun Liao, Yujian Mo
Int. J. Pattern Recognit. Artif. Intell.2
2022 Identification of winter road friction coefficient based on multi-task distillation attention network
Feilin Liu, Yan Wu 0011, Xinneng Yang, Yujian Mo, Yujun Liao
Pattern Anal. Appl.2
2021 Multi Spatial Convolution Block for Lane Lines Semantic Segmentation
Yan Wu 0011, Feilin Liu, Xinneng Yang
ICIC (2)1
2021 GPU-Efficient Dense Convolutional Network for Real-time Semantic Segmentation
abstract
Real-time semantic segmentation is a challenging task as both accuracy and inference speed need to be considered simultaneously. In real-world applications, it is usually achieved by deploying a deep neural network in modern GPU device. However, most of the work focused on real-time semantic segmentation is designed by significantly reducing computation complexity and model size. There are other factors that have a significant impact on inference speed are overlooked, especially when the network is running in modern GPU device. In this paper, we focus on designing a GPU-efficient network as backbone for real-time semantic segmentation. Dense connectivity can preserve and accumulate feature maps of multiple receptive fields and is therefore ideal for semantic segmentation. Therefore, we design a GPU-efficient network (DenseENet) with dense connectivity. The proposed DenseENet shows an obvious advantage in balancing accuracy and inference speed in modern GPU device. Specifically, on Cityscapes test set, DenseENet with a simple FCN decoder achieves 75.2% mIoU with 83.6 FPS for an input of 1024 × 2048 resolution and 73.6% mIoU with 132 FPS for an input of 768 × 1536 resolution on a single GTX 1080Ti card.
Xinneng Yang, Yan Wu 0011, Junqiao Zhao, Feilin Liu
ICRA2
2020 Dense Dual-Path Network for Real-Time Semantic Segmentation
Xinneng Yang, Yan Wu 0011, Junqiao Zhao, Feilin Liu
ACCV (1)2
2020 A Survey of Vision-Based Road Parameter Estimating Methods
Yan Wu 0011, Feilin Liu, Linting Guan, Xinneng Yang
ICIC (3)1
2019 DFNet: Semantic Segmentation on Panoramic Images with Dynamic Loss Weights and Residual Fusion Block
abstract
For the domain of self-driving and automatic parking, perception is a basic and critical technique, moreover, the detection of lane markings and parking slots is an important part of visual perception. Compared with front sight images, panoramic images(PI) can capture more comprehensive pavement information. However, the imbalance of different classes in PI is even more serious. Additionally, the judgment of boundary information between areas is a hard problem in deep models. Therefore, we propose a new model named DFNet to solve these problems. The proposed model has two main contributions, one is dynamic loss weights, and the other is residual fusion block(RFB). DFNet use dynamic loss weights to overcome the negative effect of imbalance dataset, which are calculated according to the pixel number of each class in a batch. RFB is composed of several convolutional layers, a pooling layer, and a fusion layer to combine the feature maps by pixel multiplication, which can reduce boundary information loss. We evaluate our method on PSV dataset, and the achieved advanced results demonstrate the effectiveness of the proposed model.
Yan Wu 0011, Linting Guan, Junqiao Zhao
ICRA2
2019 Efficient Large Margin-Based Feature Extraction
Yan Wu 0011
Neural Process. Lett.2
2018 Learn to Detect Objects Incrementally
abstract
Intelligent vehicles need to detect new classes of traffic objects while keeping the performance of old ones. Deep convolution neural network (DCNN) based detector has shown superior performance, however, DCNN is ill-equipped for incremental learning, i.e., a DCNN based vehicle detector trained on traffic sign dataset will catastrophic forget how to detect vehicles. In this paper, we propose a novel method to alleviate this problem, our key insight is that the original class of objects also appears in new task data, by utilizing these objects, our method effectively keeps the detection accuracy of original models while incremental learning to detect new classes of objects. Detailed experiments on PASCAL VOC dataset and TSD-max database verified the effectiveness of our method.
Linting Guan, Yan Wu 0011, Junqiao Zhao, Chen Ye 0002
Intelligent Vehicles Symposium2
2018 VH-HFCN based Parking Slot and Lane Markings Segmentation on Panoramic Surround View
abstract
The automatic parking is being massively developed by car manufacturers and providers. Until now, there are two problems with the automatic parking. First, there is no openly-available segmentation labels of parking slot on panoramic surround view (PSV) dataset. Second, how to detect parking slot and road structure robustly. Therefore, in this paper, we build up a public PSV dataset. At the same time, we proposed a highly fused convolutional network (HFCN) based segmentation method for parking slot and lane markings based on the PSV dataset. A surround-view image is made of four calibrated images captured from four fisheye cameras. We collect and label more than 4,200 surround view images for this task, which contain various illuminated scenes of different types of parking slots. A VH-HFCN network is proposed, which adopts an HFCN as the base, with an extra efficient VH-stage for better segmenting various markings. The VH-stage consists of two independent linear convolution paths with vertical and horizontal convolution kernels respectively. This modification enables the network to robustly and precisely extract linear features. We evaluated our model on the PSV dataset and the results showed outstanding performance in ground markings segmentation. Based on the segmented markings, parking slots and lanes are acquired by skeletonization, hough line transform and line arrangement.
Yan Wu 0011, Tao Yang 0044, Junqiao Zhao, Linting Guan
Intelligent Vehicles Symposium1
2017 Speeding Up Dilated Convolution Based Pedestrian Detection with Tensor Decomposition
Yan Wu 0011, Jiqian Li, Tao Yang 0044
ICIC (3)1
2017 Fully Combined Convolutional Network with Soft Cost Function for Traffic Scene Parsing
Yan Wu 0011, Tao Yang 0044, Junqiao Zhao, Linting Guan, Jiqian Li
ICIC (1)1
2017 Pedestrian detection with dilated convolution, region proposal network and boosted decision trees
abstract
With the rapid development of driverless cars, pedestrian detection has been a canonical instance of object detection. Although recent deep learning detectors such as RPN+BF and MS-CNN have shown excellent performance for pedestrian detection, they have limited success for detecting pedestrian, and the importance of final feature receptive field has been awared by previous leading deep learning pedestrian detectors. Applying the dilated convolution to the feature learning of pedestrian detection, we constructed a pedestrian detection framework along with the region proposal network and boosted decision trees. Pipeline of our proposed framework can be briefly generalized as follows: firstly, the fine-tuned RPN with specified aspect ratio is used to get boxes and scores. Secondly, the designed dilated convolution feature extraction model is used to get features. As different dilation factors provide different receptive field scales, we concat the features of different layers with the dilated convolutional features to get the final features. Finally, the candidate boxes are sent to the boosted decision trees to be classified using the scores and features. We evaluated our method on the Caltech Pedestrian Detection Benchmark. Comparing with other state-of-the-art detection methods, the proposed framework with dilated convolution has better performance.
Jiqian Li, Yan Wu 0011, Junqiao Zhao, Linting Guan, Chen Ye 0002, Tao Yang 0044
IJCNN2
2017 Multiple Classifiers-Based Feature Fusion for RGB-D Object Recognition
abstract
RGB-D-based object recognition has been enthusiastically investigated in the past few years. RGB and depth images provide useful and complementary information. Fusing RGB and depth features can significantly increase the accuracy of object recognition. However, previous works just simply take the depth image as the fourth channel of the RGB image and concatenate the RGB and depth features, ignoring the different power of RGB and depth information for different objects. In this paper, a new method which contains three different classifiers is proposed to fuse features extracted from RGB image and depth image for RGB-D-based object recognition. Firstly, a RGB classifier and a depth classifier are trained by cross-validation to get the accuracy difference between RGB and depth features for each object. Then a variant RGB-D classifier is trained with different initialization parameters for each class according to the accuracy difference. The variant RGB-D-classifier can result in a more robust classification performance. The proposed method is evaluated on two benchmark RGB-D datasets. Compared with previous methods, ours achieves comparable performance with the state-of-the-art method.
Yan Wu 0011, Jiqian Li
Int. J. Pattern Recognit. Artif. Intell.1
2015 Subset based deep learning for RGB-D object recognition
Yan Wu 0011, Fuqiang Chen
Neurocomputing2
2015 Effective feature selection using feature vector graph for classification
Yan Wu 0011, Fuqiang Chen
Neurocomputing2
2014 SAE-RNN Deep Learning for RGB-D Based Object Recognition
Yan Wu 0011
ICIC (1)2
2014 Contractive De-noising Auto-Encoder
Fuqiang Chen, Yan Wu 0011
ICIC (1)2
2014 Convolutional deep belief networks for feature extraction of EEG signal
abstract
In recent years, deep learning approaches have been successfully used to learn hierarchical representations of image data, audio data etc. However, to our knowledge, these deep learning approaches have not been extensively studied for electroencephalographic (EEG) data. Considering the properties of EEG data, high-dimensional and multichannel, we applied convolutional deep belief networks to the feature learning of EEG data and evaluated it on the datasets from previous BCI competitions. Compared with other state-of-the-art feature extraction methods, the learned features using convolutional deep belief network have better performance.
Yuanfang Ren, Yan Wu 0011
IJCNN2
2014 A co-training algorithm for EEG classification with biomimetic pattern recognition and sparse representation
Yuanfang Ren, Yan Wu 0011, Yanbin Ge
Neurocomputing2
2013 A novel method for motor imagery EEG adaptive classification based biomimetic pattern recognition
Yan Wu 0011, Yanbin Ge
Neurocomputing1
2013 An efficient algorithm for high-dimensional function optimization
Yuanfang Ren, Yan Wu 0011
Soft Comput.2
2012 A New Hybrid Method with Biomimetic Pattern Recognition and Sparse Representation for EEG Classification
Yanbin Ge, Yan Wu 0011
ICIC (3)2
2011 Towards Adaptive Classification of Motor Imagery EEG Using Biomimetic Pattern Recognition
Yanbin Ge, Yan Wu 0011
ICIC (2)2
2006 A New Subspace Analysis Approach Based on Laplacianfaces
Yan Wu 0011, Ren-Min Gu
ICONIP (2)1
2006 A New Method for Feature Selection
Yan Wu 0011
ISNN (1)1
2006 Face Detection Method Based on Kernel Independent Component Analysis and Boosting Chain Algorithm
Yan Wu 0011, Yin-Fang Zhuang
ISNN (2)1
2005 Speech Recognition of Finite Words Based on Multi-weight Neural Network
Yan Wu 0011, Mingxi Jin, Shoujue Wang
ISNN (2)1
2004 Local Face Recognition Based on the Combination of ICA and NFL
Yisong Ye, Yan Wu 0011, Mingliang Sun, Mingxi Jin
ISNN (1)2