EDBT 2026 Demo / reviewers in the wild / expert
Xuefeng Yan 0001
dblp:13/4745-1
· DBLP profile ↗
35ranked-venue papers
0as first author
35since 2021 · last 2026
0000-0001-6030-0855ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A node embedded representation based gravity model for evaluating the importance of nodes in complex networks
Haoming Guo, Xuefeng Yan 0001, Yusong Liu, Juping Zhang |
Neurocomputing | 2 |
| 2026 | CrossTracker: Robust Multi-Modal 3D Multi-Object Tracking via Cross CorrectionabstractInaccurate detections remain a critical bottleneck in 3D multi-object tracking (MOT). Recent detection fusion-based methods incorporate camera detections as supplementary to reduce false detections and compensate for missing ones in LiDAR. However, their unidirectional camera-LiDAR correction lacks a feedback mechanism, precluding iterative mutual refinement between modalities for more robust LiDAR-based tracking. Inspired by the coarse-to-fine strategy in two-stage object detection, we introduceCrossTracker, a novel two-stage framework for online multi-modal 3D MOT. CrossTracker first constructs coarse camera and LiDAR trajectories independently, then performs trajectory fusion using both current and historical frames, without requiring future data. This ensures more robust mutual refinement between modalities. Specifically, CrossTracker comprises three core modules: i) the multi-modal modeling (M3) module, which fuses data from images, point clouds, and even planar geometry derived from images to establish a robust tracking constraint; ii) the coarse trajectory generation (C-TG) module, which independently generates coarse trajectories for both modalities using the M3constraint; and iii) the trajectory fusion (TF) module, which applies mutual refinement between coarse LiDAR and camera trajectories through cross correction to ensure robust LiDAR trajectories. Extensive experiments show that CrossTracker outperforms 19 state-of-the-art methods, highlighting its effectiveness in leveraging the synergistic strengths of camera and LiDAR sensors for robust multi-modal 3D MOT. The code is available at https://github.com/lipeng-gu/CrossTracker. Lipeng Gu, Xuefeng Yan 0001, Weiming Wang 0002, Honghua Chen, Dingkun Zhu, Liangliang Nan, Mingqiang Wei |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Rethinking mixture of rain removal via depth-guided adversarial learning
Yongzhen Wang 0001, Xuefeng Yan 0001, Yanbiao Niu, Lina Gong, Yanwen Guo 0001, Mingqiang Wei |
Neural Networks | 2 |
| 2025 | RSHazeDiff: A Unified Fourier-Aware Diffusion Model for Remote Sensing Image DehazingabstractHaze severely degrades the visual quality of remote sensing images and hampers the performance of road extraction, vehicle detection, and traffic flow monitoring. The emerging denoising diffusion probabilistic model (DDPM) exhibits the significant potential for dense haze removal with its strong generation ability. Since remote sensing images contain extensive small-scale texture structures, it is important to effectively restore image details from hazy images. However, current wisdom of DDPM fails to preserve image details and color fidelity well, limiting its dehazing capacity for remote sensing images. In this paper, we propose a novel unified Fourier-aware diffusion model for remote sensing image dehazing, termed RSHazeDiff. From a new perspective, RSHazeDiff explores the conditional DDPM to improve image quality in dense hazy scenarios, and it makes three key contributions. First, RSHazeDiff refines the training phase of diffusion process by performing noise estimation and reconstruction constraints in a coarse-to-fine fashion. Thus, it remedies the unpleasing results caused by the simple noise estimation constraint in DDPM. Second, by taking the frequency information as important prior knowledge during iterative sampling steps, RSHazeDiff can preserve more texture details and color fidelity in dehazed images. Third, we design a global compensated learning module to utilize the Fourier transform to capture the global dependency features of input images, which can effectively mitigate the effects of boundary artifacts when processing fixed-size patches. Experiments on both synthetic and real-world benchmarks validate the favorable performance of RSHazeDiff over state-of-the-art methods. Source code will be released athttps://github.com/jm-xiong/RSHazeDiff Jiamei Xiong, Xuefeng Yan 0001, Yongzhen Wang 0001, Wei Zhao 0039, Xiao-Ping Zhang 0002, Mingqiang Wei |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Revisiting Tradition and Beyond: A Customized Bilateral Filtering Framework for Point Cloud DenoisingabstractDeep learning-based methods have become the dominant solution for point cloud denoising, offering strong generalization capabilities through data-driven training. However, traditional methods, despite their drawbacks of heavy parameter tuning and weak generalization, retain unique advantages in interpretability and theoretical robustness. This complementarity motivates us to explore a hybrid solution that leverages data-driven paradigms to overcome the performance constraints of traditional methods. In this paper, we revisit the classic bilateral filter (BF) as a case study and identify three key limitations hindering its performance: excessive parameter tuning, suboptimal neighborhood quality, and fixed parameters across the entire model. To address them, we propose CustomBF, a novel framework for customizing BF components at a per-point level. CustomBF employs multigraph encoders and a mutual guidance strategy to analyze local patches, enabling the customization of BF components including center point normal, neighborhood point coordinates, Gaussian function parameters, and neighborhood radius for each point. Experimental results demonstrate that this component-customized bilateral filter outperforms state-of-the-art methods and achieves robust denoising even in complex scenarios. It highlights the potential of hybrid methods to extend the applicability and effectiveness of traditional techniques. Peng Li 0064, Zeyong Wei, Honghua Chen, Xuefeng Yan 0001, Mingqiang Wei |
ACM Trans. Graph. | 4 |
| 2025 | PointCG: Self-Supervised Point Cloud Learning via Joint Completion and GenerationabstractThe core of self-supervised point cloud learning lies in setting up appropriate pretext tasks, to construct a pre-training framework that enables the encoder to perceive 3D objects effectively. In this article, we integrate two prevalent methods, masked point modeling (MPM) and 3D-to-2D generation, as pretext tasks within a pre-training framework. We leverage the spatial awareness and precise supervision offered by these two methods to address their respective limitations: ambiguous supervision signals and insensitivity to geometric information. Specifically, the proposed framework, abbreviated as PointCG, consists of a Hidden Point Completion (HPC) module and an Arbitrary-view Image Generation (AIG) module. We first capture visible points from arbitrary views as inputs by removing hidden points. Then, HPC extracts representations of the inputs with an encoder and completes the entire shape with a decoder, while AIG is used to generate rendered images based on the visible points' representations. Extensive experiments demonstrate the superiority of the proposed method over the baselines in various downstream tasks. Our code will be made available upon acceptance. Yun Liu 0002, Peng Li 0064, Xuefeng Yan 0001, Liangliang Nan, Bing Wang 0013, Honghua Chen, Lina Gong, Wei Zhao 0039, Mingqiang Wei |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Semi-UFormer: Semi-supervised Uncertainty-aware Transformer for Image DehazingabstractImage dehazing is fundamental yet not well-solved in computer vision. Most cutting-edge models are trained in synthetic data, leading to the poor performance on real-world hazy scenarios. Besides, they commonly give deterministic dehazed images while neglecting to mine their uncertainty. To bridge the domain gap and enhance the dehazing performance, we propose a novel semi-supervised uncertainty-aware transformer network, called Semi-UFormer. Semi-UFormer can well leverage both the real-world hazy images and their uncertainty guidance information. Specifically, Semi-UFormer builds itself on the knowledge distillation framework. Such teacher-student networks effectively absorb real-world haze information for quality dehazing. Furthermore, an uncertainty estimation block is introduced into the model to estimate the pixel uncertainty representations, which is then used as a guidance signal to help the student network produce haze-free images more accurately. Extensive experiments demonstrate that Semi-UFormer generalizes well from synthetic to real-world images. Ming Tong, Xuefeng Yan 0001, Yongzhen Wang 0001, Mingqiang Wei |
IJCNN | 2 |
| 2024 | PointeNet: A lightweight framework for effective and efficient point cloud analysis
Lipeng Gu, Xuefeng Yan 0001, Liangliang Nan, Dingkun Zhu, Honghua Chen, Weiming Wang 0002, Mingqiang Wei |
Comput. Aided Geom. Des. | 2 |
| 2024 | An improved sand cat swarm optimization for moving target search by UAV
Yanbiao Niu, Xuefeng Yan 0001, Yongzhen Wang 0001, Yanzhao Niu |
Expert Syst. Appl. | 2 |
| 2024 | GeoDC: Geometry-Constrained Depth Completion With Depth Distribution ModelingabstractDepth completion is a fundamental, yet not well-solved problem in 3-D vision. Current wisdom attempts to employ implicit geometric spatial cues from point clouds to assist in depth completion. However, these methods encounter challenges in extracting rich geometric features due to the absence of explicit constraints. In this article, we propose GeoDC, a geometry-constrained depth completion network with depth distribution modeling. GeoDC employs point cloud upsampling as an auxiliary task to guide the network in learning more robust and effective geometric features. Simultaneously, a novel image and point cloud fusion module, denoted as IP-Interaction, is implemented to holistically integrate features from images and point clouds. Besides, recognizing the presence of uncertainty and ambiguity in the ground-truth (GT) data, we construct a prior network and a posterior network to model depth feature distributions and leverage the distributions to guide depth map inference. GeoDC can solve both the problems of geometric constraint inadequacies in feature extraction and data uncertainty within depth maps well. Extensive experiments underscore the efficacy of our method, demonstrating comparable or superior performance when compared to existing state-of-the-art methods. Peng Li 0064, Xuefeng Yan 0001, Honghua Chen, Mingqiang Wei |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | UCL-Dehaze: Toward Real-World Image Dehazing via Unsupervised Contrastive LearningabstractWhile the wisdom of training an image dehazing model on synthetic hazy data can alleviate the difficulty of collecting real-world hazy/clean image pairs, it brings the well-known domain shift problem. From a different yet new perspective, this paper explores contrastive learning with an adversarial training effort to leverage unpaired real-world hazy and clean images, thus alleviating the domain shift problem and enhancing the network's generalization ability in real-world scenarios. We propose an effective unsupervised contrastive learning paradigm for image dehazing, dubbed UCL-Dehaze. Unpaired real-world clean and hazy images are easily captured, and will serve as the important positive and negative samples respectively when training our UCL-Dehaze network. To train the network more effectively, we formulate a new self-contrastive perceptual loss function, which encourages the restored images to approach the positive samples and keep away from the negative samples in the embedding space. Besides the overall network architecture of UCL-Dehaze, adversarial training is utilized to align the distributions between the positive samples and the dehazed images. Compared with recent image dehazing works, UCL-Dehaze does not require paired data during training and utilizes unpaired positive/negative data to better enhance the dehazing performance. We conduct comprehensive experiments to evaluate our UCL-Dehaze and demonstrate its superiority over the state-of-the-arts, even only 1,800 unpaired real-world images are used to train our network. Source code is publicly available at https://github.com/yz-wang/UCL-Dehaze. Yongzhen Wang 0001, Xuefeng Yan 0001, Fu Lee Wang, Haoran Xie 0001, Wenhan Yang, Xiao-Ping Zhang 0002, Harry Qin, Mingqiang Wei |
IEEE Trans. Image Process. | 2 |
| 2024 | PointSee: Image Enhances Point CloudabstractThere is a prevailing trend towards fusing multi-modal information for 3D object detection (3OD). However, challenges related to computational efficiency, plug-and-play capabilities, and accurate feature alignment have not been adequately addressed in the design of multi-modal fusion networks. In this paper, we present PointSee, a lightweight, flexible, and effective multi-modal fusion solution to facilitate various 3OD networks by semantic feature enhancement of point clouds (e.g., LiDAR or RGB-D data) assembled with scene images. Beyond the existing wisdom of 3OD, PointSee consists of a hidden module (HM) and a seen module (SM): HM decorates point clouds using 2D image information in an offline fusion manner, leading to minimal or even no adaptations of existing 3OD networks; SM further enriches the point clouds by acquiring point-wise representative semantic features, leading to enhanced performance of existing 3OD networks. Besides the new architecture of PointSee, we propose a simple yet efficient training strategy, to ease the potential inaccurate regressions of 2D object detection networks. Extensive experiments on the popular outdoor/indoor benchmarks show quantitative and qualitative improvements of our PointSee over thirty-five state-of-the-art methods. Lipeng Gu, Xuefeng Yan 0001, Peng Cui 0013, Lina Gong, Haoran Xie 0001, Fu Lee Wang, Harry Qin, Mingqiang Wei |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | GeoSegNet: point cloud semantic segmentation via geometric encoder-decoder modeling
Chen Chen 0161, Yisen Wang 0003, Honghua Chen, Xuefeng Yan 0001, Dayong Ren, Yanwen Guo 0001, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei |
Vis. Comput. | 4 |
| 2023 | Adaptive Dehazing YOLO for Object Detection
Kaiwen Zhang 0011, Xuefeng Yan 0001, Yongzhen Wang 0001, Junchen Qi |
ICANN (7) | 2 |
| 2023 | ifUNet++: Iterative Feedback UNet++ for Infrared Small Target DetectionabstractSmall targets are often submerged in the cluttered backgrounds of infrared images. In this paper, we propose an iterative feedback UNet++ for infrared small target detection, dubbed ifUNet++. Unlike most of existing methods, ifU-Net++ enables to concentrate on small targets while weakening the interference of clutter backgrounds. ifUNet++ contains two parts: a simplified UNet++ and an iterative feedback strategy. We reduce the unnecessary nodes of UNet++ and have the simplified UNet++ as our backbone network, avoiding the loss of infrared small targets. Based on the simplified network, we search the infrared small targets in an iterative feedback manner, avoiding the interference of cluttered backgrounds. Besides, to optimize the iterative results, we propose Contextual Multiple Attention (CMA) to enhance the features in each iteration. Experimental results exhibit the clear promotion of ifUNet++ over eight state-of-the-art methods, in terms of noise-robustness and detection accuracy. Zhangying Weng, Peng Li 0064, Xin Zhuang, Xuefeng Yan 0001, Lina Gong, Haoran Xie 0001, Mingqiang Wei |
ICASSP | 4 |
| 2023 | Three-dimensional UCAV path planning using a novel modified artificial ecosystem optimizer
Yanbiao Niu, Xuefeng Yan 0001, Yongzhen Wang 0001, Yanzhao Niu |
Expert Syst. Appl. | 2 |
| 2023 | Three-dimensional collaborative path planning for multiple UCAVs based on improved artificial ecosystem optimizer and reinforcement learning
Yanbiao Niu, Xuefeng Yan 0001, Yongzhen Wang 0001, Yanzhao Niu |
Knowl. Based Syst. | 2 |
| 2023 | AGConv: Adaptive Graph Convolution on 3D Point CloudsabstractConvolution on 3D point clouds is widely researched yet far from perfect in geometric deep learning. The traditional wisdom of convolution characterises feature correspondences indistinguishably among 3D points, arising an intrinsic limitation of poor distinctive feature learning. In this article, we propose Adaptive Graph Convolution (AGConv) for wide applications of point cloud analysis. AGConv generates adaptive kernels for points according to their dynamically learned features. Compared with the solution of using fixed/isotropic kernels, AGConv improves the flexibility of point cloud convolutions, effectively and precisely capturing the diverse relations between points from different semantic parts. Unlike the popular attentional weight schemes, AGConv implements the adaptiveness inside the convolution operation instead of simply assigning different weights to the neighboring points. Extensive evaluations clearly show that our method outperforms state-of-the-arts of point cloud classification and segmentation on various benchmark datasets. Meanwhile, AGConv can flexibly serve more point cloud analysis approaches to boost their performance. To validate its flexibility and effectiveness, we explore AGConv-based paradigms of completion, denoising, upsampling, registration and circle extraction, which are comparable or even superior to their competitors. Mingqiang Wei, Zeyong Wei, Huajian Si, Zhilei Chen, Zhe Zhu, Jingbo Qiu, Xuefeng Yan 0001, Yanwen Guo 0001, Jun Wang 0039, Harry Qin |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2023 | PointGame: Geometrically and Adaptively Masked Autoencoder on Point CloudsabstractSelf-supervised learning is attracting large attention in point cloud understanding. However, exploring discriminative and transferable features still remains challenging due to their nature of irregularity. We propose a geometrically and adaptively masked auto-encoder on point clouds for self-supervised learning, termedPointGame. PointGame contains two core components: GATE and EAT. GATE stands for the geometrical and adaptive token embedding module; it not only absorbs the conventional wisdom of geometric descriptors that captures the surface shape effectively, but also exploits adaptive saliency to focus on the salient part of a point cloud. EAT stands for the external attention-based Transformer encoder with linear computational complexity, which increases the efficiency of the whole pipeline. Unlike cutting-edge unsupervised learning models, PointGame leverages geometric descriptors to perceive surface shapes and adaptively mines discriminative features from training data. PointGame showcases clear advantages over its competitors on various downstream tasks under both global and local fine-tuning strategies. The code and pre-trained models will be publicly available. Yun Liu 0002, Xuefeng Yan 0001, Zhiqi Li 0002, Zhilei Chen, Zeyong Wei, Mingqiang Wei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CF-YOLO: Cross Fusion YOLO for Object Detection in Adverse Weather With a High-Quality Real Snow DatasetabstractSnow is one of the toughest adverse weather conditions for object detection (OD). Currently, not only there is a lack of snowy OD datasets to train cutting-edge detectors, but also these detectors have difficulties of learning latent information beneficial for detection in snow. To alleviate the two above problems, we first establish a real-world snowy OD dataset, named RSOD. Besides, we develop an unsupervised training strategy with a distinctive activation function, called$Peak Act$, to quantitatively evaluate the effect of snow on each object. Peak Act helps grade the images in RSOD into four-difficulty levels. To our knowledge, RSOD is the first quantitatively evaluated and graded real-world snowy OD dataset. Then, we propose a novel Cross Fusion (CF) block to construct a lightweight OD network based on YOLOv5s (called CF-YOLO). CF is a plug-and-play feature aggregation module, which integrates the advantages of Feature Pyramid Network and Path Aggregation Network in a simpler yet more flexible form. Both RSOD and CF lead our CF-YOLO to possess an optimization ability for OD in real-world snow. That is, CF-YOLO can handle unfavorable detection problems of vagueness, distortion and covering of snow. Experiments show that our CF-YOLO achieves better detection results on RSOD, compared to SOTAs. The code and dataset are available athttps://github.com/qqding77/CF-YOLO-and-RSOD. Qiqi Ding, Peng Li 0064, Xuefeng Yan 0001, Ding Shi, Luming Liang, Weiming Wang 0002, Haoran Xie 0001, Jonathan Li 0001, Mingqiang Wei |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | USCFormer: Unified Transformer With Semantically Contrastive Learning for Image DehazingabstractHaze severely degrades the visibility of scene objects and deteriorates the performance of autonomous driving, traffic monitoring, and other vision-based intelligent transportation systems. As a potential remedy, we propose a novel unified Transformer with semantically contrastive learning for image dehazing, dubbed USCFormer. USCFormer has three key contributions. First, USCFormer absorbs the respective strengths of CNN and Transformer by incorporating them into a unified Transformer format. Thus, it allows the simultaneous capture of global-local dependency features for better image dehazing. Second, by casting clean/hazy images as the positive/negative samples, the contrastive constraint encourages the restored image to be closer to the ground-truth images (positives) and away from the hazy ones (negatives). Third, we regard the semantic information as important prior knowledge to help USCFormer mitigate the effects of haze on the scene and preserve image details and colors by leveraging intra-object semantic correlation. Experiments on synthetic datasets and real-world hazy photos fully validate the superiority of USCFormer in both perceptual quality assessment and subjective evaluation. Code is available athttps://github.com/yz-wang/USCFormer. Yongzhen Wang 0001, Jiamei Xiong, Xuefeng Yan 0001, Mingqiang Wei |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | PV-RCNN++: semantical point-voxel feature interaction for 3D object detection
Lipeng Gu, Xuefeng Yan 0001, Haoran Xie 0001, Fu Lee Wang, Gary Cheng 0001, Mingqiang Wei |
Vis. Comput. | 3 |
| 2022 | I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object DetectionabstractCan you find me? By simulating how humans to discover the so-called 'perfectly'-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to highlight the separator (or say the camouflaged object's boundary) between an image's background and foreground: the reverse attention stream helps erase the camouflaged object's interior to focus on the background, while the normal attention stream recovers the interior and thus pay more attention to the foreground; and both streams are followed by a boundary guider module and combined to strengthen the understanding of boundary. The core design of such separated attention is motivated by the COD procedure of humans: find the subtle difference between the foreground and background to delineate the boundary of a camouflaged object, then the boundary can help further enhance the COD accuracy. We validate on three benchmark datasets that the proposed BSA-Net is very beneficial to detect camouflaged objects with the blurred boundaries and similar colors/patterns with their backgrounds. Extensive results exhibit very clear COD improvements on our BSA-Net over sixteen SOTAs. Peng Li 0064, Haoran Xie 0001, Xuefeng Yan 0001, Dong Liang 0008, Dapeng Chen, Mingqiang Wei, Harry Qin |
AAAI | 4 |
| 2022 | UTOPIC: Uncertainty-aware Overlap Prediction Network for Partial Point Cloud RegistrationabstractAbstract High‐confidence overlap prediction and accurate correspondences are critical for cutting‐edge models to align paired point clouds in a partial‐to‐partial manner. However, there inherently exists uncertainty between the overlapping and non‐overlapping regions, which has always been neglected and significantly affects the registration performance. Beyond the current wisdom, we propose a novel uncertainty‐aware overlap prediction network, dubbed UTOPIC, to tackle the ambiguous overlap prediction problem; to our knowledge, this is the first to explicitly introduce overlap uncertainty to point cloud registration. Moreover, we induce the feature extractor to implicitly perceive the shape knowledge through a completion decoder, and present a geometric relation embedding for Transformer to obtain transformation‐invariant geometry‐aware feature representations. With the merits of more reliable overlap scores and more precise dense correspondences, UTOPIC can achieve stable and accurate registration results, even for the inputs with limited overlapping areas. Extensive quantitative and qualitative experiments on synthetic and real benchmarks demonstrate the superiority of our approach over state‐of‐the‐art methods. Zhilei Chen, Honghua Chen, Lina Gong, Xuefeng Yan 0001, Jun Wang 0039, Yanwen Guo 0001, Harry Qin, Mingqiang Wei |
Comput. Graph. Forum | 4 |
| 2022 | SO(3)-Pose: SO(3)-Equivariance Learning for 6D Object Pose EstimationabstractAbstract 6D pose estimation of rigid objects from RGB‐D images is crucial for object grasping and manipulation in robotics. Although RGB channels and the depth (D) channel are often complementary, providing respectively the appearance and geometry information, it is still non‐trivial on how to fully benefit from the two cross‐modal data. From the simple yet new observation, when an object rotates, its semantic label is invariant to the pose while its keypoint offset direction is variant to the pose. To this end, we present SO(3)‐Pose, a new representation learning network to explore SO(3)‐equivariant and SO(3)‐invariant features from the depth channel for pose estimation. The SO(3)‐invariant features facilitate to learn more distinctive representations for segmenting objects with similar appearance from RGB channels. The SO(3)‐equivariant features communicate with RGB features to deduce the (missed) geometry for detecting keypoints of an object with the reflective surface from the depth channel. Unlike most of existing pose estimation methods, our SO(3)‐Pose not only implements the information communication between the RGB and depth channels, but also naturally absorbs the SO(3)‐equivariance geometry knowledge from depth images, leading to better appearance and geometry representation learning. Comprehensive experiments show that our method achieves the state‐of‐the‐art performance on three benchmarks. Code is available at https://github.com/phaoran9999/SO3-Pose . Haoran Pan, Jun Zhou 0007, Xuequan Lu, Weiming Wang 0002, Xuefeng Yan 0001, Mingqiang Wei |
Comput. Graph. Forum | 6 |
| 2022 | Contrastive Semantic-Guided Image Smoothing NetworkabstractAbstract Image smoothing is a fundamental low‐level vision task that aims to preserve salient structures of an image while removing insignificant details. Deep learning has been explored in image smoothing to deal with the complex entanglement of semantic structures and trivial details. However, current methods neglect two important facts in smoothing: 1) naive pixel‐level regression supervised by the limited number of high‐quality smoothing ground‐truth could lead to domain shift and cause generalization problems towards real‐world images; 2) texture appearance is closely related to object semantics, so that image smoothing requires awareness of semantic difference to apply adaptive smoothing strengths. To address these issues, we propose a novel Contrastive Semantic‐Guided Image Smoothing Network (CSGIS‐Net) that combines both contrastive prior and semantic prior to facilitate robust image smoothing. The supervision signal is augmented by leveraging undesired smoothing effects as negative teachers, and by incorporating segmentation tasks to encourage semantic distinctiveness. To realize the proposed network, we also enrich the original VOC dataset with texture enhancement and smoothing labels, namely VOC‐smooth, which first bridges image smoothing and semantic segmentation. Extensive experiments demonstrate that the proposed CSGIS‐Net outperforms state‐of‐the‐art algorithms by a large margin. Code and dataset are available at https://github.com/wangjie6866/CSGIS-Net . Jie Wang 0069, Yongzhen Wang 0001, Yidan Feng, Lina Gong, Xuefeng Yan 0001, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei |
Comput. Graph. Forum | 5 |
| 2022 | TogetherNet: Bridging Image Restoration and Object Detection Together via Dynamic Enhancement LearningabstractAbstract Adverse weather conditions such as haze, rain, and snow often impair the quality of captured images, causing detection networks trained on normal images to generalize poorly in these scenarios. In this paper, we raise an intriguing question – if the combination of image restoration and object detection, can boost the performance of cutting‐edge detectors in adverse weather conditions. To answer it, we propose an effective yet unified detection paradigm that bridges these two subtasks together via dynamic enhancement learning to discern objects in adverse weather conditions, called TogetherNet. Different from existing efforts that intuitively apply image dehazing/deraining as a pre‐processing step, TogetherNet considers a multi‐task joint learning problem. Following the joint learning scheme, clean features produced by the restoration network can be shared to learn better object detection in the detection network, thus helping TogetherNet enhance the detection capacity in adverse weather conditions. Besides the joint learning architecture, we design a new Dynamic Transformer Feature Enhancement module to improve the feature extraction and representation capabilities of TogetherNet. Extensive experiments on both synthetic and real‐world datasets demonstrate that our TogetherNet outperforms the state‐of‐the‐art detection approaches by a large margin both quantitatively and qualitatively. Source code is available at https://github.com/yz-wang/TogetherNet . Yongzhen Wang 0001, Xuefeng Yan 0001, Kaiwen Zhang 0011, Lina Gong, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei |
Comput. Graph. Forum | 2 |
| 2022 | GlassNet: Label Decoupling-based Three-stream Neural Network for Robust Image Glass DetectionabstractAbstract Most of the existing object detection methods generate poor glass detection results, due to the fact that the transparent glass shares the same appearance with arbitrary objects behind it in an image. Different from traditional deep learning‐based wisdoms that simply use the object boundary as an auxiliary supervision, we exploit label decoupling to decompose the original labelled ground‐truth (GT) map into an interior‐diffusion map and a boundary‐diffusion map. The GT map in collaboration with the two newly generated maps breaks the imbalanced distribution of the object boundary, leading to improved glass detection quality. We have three key contributions to solve the transparent glass detection problem: (1) We propose a three‐stream neural network (call GlassNet for short) to fully absorb beneficial features in the three maps. (2) We design a multi‐scale interactive dilation module to explore a wider range of contextual information. (3) We develop an attention‐based boundary‐aware feature Mosaic module to integrate multi‐modal information. Extensive experiments on the benchmark dataset exhibit clear improvements of our method over SOTAs, in terms of both the overall glass detection accuracy and boundary clearness. Ding Shi, Xuefeng Yan 0001, Dong Liang 0008, Mingqiang Wei, Xin Yang 0011, Yanwen Guo 0001, Haoran Xie 0001 |
Comput. Graph. Forum | 3 |
| 2022 | An adaptive neighborhood-based search enhanced artificial ecosystem optimizer for UCAV path planning
Yanbiao Niu, Xuefeng Yan 0001, Yongzhen Wang 0001, Yanzhao Niu |
Expert Syst. Appl. | 2 |
| 2022 | Detecting Occluded and Dense Trees in Urban Terrestrial Views With a High-Quality Tree Detection DatasetabstractUrban trees are often densely planted along the two sides of a street. When observing these trees from a fixed view, they are inevitably occluded with each other and the passing vehicles. The high density and occlusion of urban tree scenes significantly degrade the performance of object detectors. This paper raises an intriguing learning-related question – if a module is developed to enable the network to adaptively cope with occluded and un-occluded regions while enhancing its feature extraction capabilities, can the performance of a cutting-edge detection model be improved? To answer it, a lightweight yet effective object detection network is proposed for discerning occluded and dense urban trees, called OD-UTDNet. The main contribution is a newly-designed Dilated Attention Cross Stage Partial (DACSP) module. DACSP can expand the fields-of-view of OD-UTDNet for paying more attention to the un-occluded region, while enhancing the network’s feature extraction ability in the occluded region. This work further explores both the self-calibrated convolution module and GFocal loss, which enhance the OD-UTDNet’s ability to resolve the challenging problem of high densities and occlusions. Finally, to facilitate the detection task of urban trees, a high-quality urban tree detection dataset is established, named UTD; to our knowledge, this is the first time. Extensive experiments show clear improvements of the proposed OD-UTDNet over twelve representative object detectors on UTD. The code and dataset are available at https://github.com/yzwang/OD-UTDNet. Yongzhen Wang 0001, Xuefeng Yan 0001, Hexiang Bao, Yiping Chen 0002, Lina Gong, Mingqiang Wei, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | FindNet: Can You Find Me? Boundary-and-Texture Enhancement Network for Camouflaged Object DetectionabstractCamouflaged objects share very similar colors but have different semantics with the surroundings. Cognitive scientists observe that both the global contour (i.e., boundary) and the local pattern (i.e., texture) of camouflaged objects are key cues to help humans find them successfully. Inspired by the cognitive scientist's observation, we propose a novel boundary-and-texture enhancement network (FindNet) for camouflaged object detection (COD) from single images. Different from most of existing COD methods, FindNet embeds both the boundary-and-texture information into the camouflaged object features. The boundary enhancement (BE) module is leveraged to focus on the global contour of the camouflaged object, and the texture enhancement (TE) module is utilized to focus on the local pattern. The enhanced features from BE and TE, which complement each other, are combined to obtain the final prediction. FindNet performs competently on various conditions of COD, including slightly clear boundaries but very similar textures, fuzzy boundaries but slightly differentiated textures, and simultaneous fuzzy boundaries and textures. Experimental results exhibit clear improvements of FindNet over fifteen state-of-the-art methods on four benchmark datasets, in terms of detection accuracy and boundary clearness. The code will be publicly released. Peng Li 0064, Xuefeng Yan 0001, Mingqiang Wei, Xiao-Ping Zhang 0002, Harry Qin |
IEEE Trans. Image Process. | 2 |
| 2022 | Cycle-SNSPGAN: Towards Real-World Image Dehazing via Cycle Spectral Normalized Soft Likelihood Estimation Patch GANabstractImage dehazing is a common operation in autonomous driving, traffic monitoring and surveillance. Learning-based image dehazing has achieved excellent performance recently. However, it is nearly impossible to capture pairs of hazy/clean images from the real world to train an image dehazing network. Most of existing dehazing models that are learnt from synthetically generated hazy images generalize poorly on real-world hazy scenarios due to the obvious domain shift. To deal with this unpaired problem arisen by real-world hazy images, we present Cycle Spectral Normalized Soft likelihood estimation Patch Generative Adversarial Network (Cycle-SNSPGAN) for image dehazing. Cycle-SNSPGAN is an unsupervised dehazing framework to boost the generalization ability on real-world hazy images. To leverage unpaired samples of real-world hazy images without relying on their clean counterparts, we design an SN-Soft-Patch GAN and exploit a new cyclic self-perceptual loss which avoids using the ground-truth image to compute the perceptual similarity. Moreover, a significant color loss is adopted to brighten the dehazed images as human expects. Both visual and numerical results show clear improvements of the proposed Cycle-SNSPGAN over state-of-the-arts in terms of hazy-robustness and image detail recovery, with even only a small dataset training our Cycle-SNSPGAN. Code has been available athttps://github.com/yz-wang/Cycle-SNSPGAN. Yongzhen Wang 0001, Xuefeng Yan 0001, Donghai Guan, Mingqiang Wei, Yiping Chen 0002, Xiao-Ping Zhang 0002, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Dynamic opposite learning enhanced artificial ecosystem optimizer for IIR system identification
Yanbiao Niu, Xuefeng Yan 0001, Yongzhen Wang 0001, Yanzhao Niu |
J. Supercomput. | 2 |
| 2022 | Easy2Hard: Learning to Solve the Intractables From a Synthetic Dataset for Structure-Preserving Image SmoothingabstractImage smoothing is a prerequisite for many computer vision and graphics applications. In this article, we raise an intriguing question whether a dataset that semantically describes meaningful structures and unimportant details can facilitate a deep learning model to smooth complex natural images. To answer it, we generate ground-truth labels from easy samples by candidate generation and a screening test and synthesize hard samples in structure-preserving smoothing by blending intricate and multifarious details with the labels. To take full advantage of this dataset, we present a joint edge detection and structure-preserving image smoothing neural network (JESS-Net). Moreover, we propose the distinctive total variation loss as prior knowledge to narrow the gap between synthetic and real data. Experiments on different datasets and real images show clear improvements of our method over the state of the arts in terms of both the image cleanness and structure-preserving ability. Code and dataset are available at https://github.com/YidFeng/Easy2Hard. Yidan Feng, Xuefeng Yan 0001, Xin Yang 0011, Mingqiang Wei, Ligang Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Research on GPU parallel algorithm for direct numerical solution of two-dimensional compressible flows
Yongzhen Wang 0001, Xuefeng Yan 0001, Jun-an Zhang |
J. Supercomput. | 2 |