Peng Li 0064

dblp:83/6353-64 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoreEditor: Correspondence-Constrained Diffusion for Consistent 3D Editing
abstract
Text-driven 3D editing is an emerging task that focuses on modifying scenes based on text prompts. Current methods often adapt pre-trained 2D image editors to multi-view observations, using specific strategies to combine information across views. However, these approaches still struggle with ensuring consistency across views, as they lack precise control over the sharing of information, resulting in edits with insufficient visual changes and blurry details. In this paper, we propose CoreEditor, a novel framework for consistent text-to-3D editing. At the core of our approach is a novel correspondence-constrained attention mechanism, which enforces structured interactions between corresponding pixels that are expected to remain visually consistent during the diffusion denoising process. Unlike conventional wisdom that relies solely on scene geometry, we enhance the correspondence by incorporating semantic similarity derived from the diffusion denoising process. This combined support from both geometry and semantics ensures a robust multi-view editing process. Additionally, we introduce a selective editing pipeline that enables users to choose their preferred edits from multiple candidates, creating a more flexible and user-centered 3D editing process. Extensive experiments demonstrate the effectiveness of CoreEditor, showing its ability to generate high-quality 3D edits, significantly outperforming existing methods.
Zhe Zhu, Honghua Chen, Peng Li 0064, Mingqiang Wei
IEEE Trans. Vis. Comput. Graph.3
2025 FPMamba: A Frequency-Prior Mamba for Brain MRI Synthesis with Randomly Missing Modalities
abstract
With the diversified development of medical data modalities, traditional methods have exposed significant limitations when dealing with scenarios involving arbitrary missing combinations of multimodal data. They require independent model training for each specific modality combination and struggle to adapt to dynamically changing data missing patterns in clinical practice. To solve these problems, we propose FPMamba, an innovative multimodal MRI synthesis model aiming to perform many-to-many synthesis for arbitrary modal combinations. This specific capability stems from our integration of modality masks during model training. Meanwhile, due to the weak contextual awareness of traditional scanning trajectories for nonrectangular directions such as diagonals and rings, we design a scanning path optimization algorithm based on frequency domain energy analysis, which quantifies the spatial distribution of image frequency features to enhance the state space model's ability to model key structures. Comparative experiments against various mainstream methods highlight FPMamba's superior performance in high-frequency detail synthesis, modality adaptability, and complex combination processing. Our method achieves a 0.64 increase in peak signal-to-noise ratio, demonstrating its efficacy and significant clinical application potential.
Peng Li 0064, Mingqiang Wei
CW2
2025 TaP-SAM: Text-and-Point Guided SAM for Nuclei Instance Segmentation
abstract
Nuclei instance segmentation plays a crucial role in pathological image analysis. Recent Segment Anything Model (SAM) and Vision-Language Pre-trained Models (VLPMs) are being actively investigated in medical imaging for their excellent generalization and semantic alignment capabilities, yet their application in nuclei segmentation remains insufficiently explored. We propose a text-point prompt fusion framework for nuclei segmentation, termed TaP-SAM. The framework introduces text prompts that describe nuclei morphology and color attributes to enhance SAM's semantic understanding and improve its segmentation performance. These prompts are generated automatically by a dedicated module that incorporates VLPMs to produce high-quality semantic descriptions. In parallel, we utilize anchor optimization and graph matching to automatically generate point prompts, enabling collaborative guidance via both spatial and semantic cues. Experimental results on CPM17, MoNuSeg, and PanNuke demonstrate that our method achieves excellent segmentation performance and strong generalization ability.
Ketian Li, Peng Li 0064, Mingqiang Wei
CW2
2025 Revisiting Tradition and Beyond: A Customized Bilateral Filtering Framework for Point Cloud Denoising
abstract
Deep learning-based methods have become the dominant solution for point cloud denoising, offering strong generalization capabilities through data-driven training. However, traditional methods, despite their drawbacks of heavy parameter tuning and weak generalization, retain unique advantages in interpretability and theoretical robustness. This complementarity motivates us to explore a hybrid solution that leverages data-driven paradigms to overcome the performance constraints of traditional methods. In this paper, we revisit the classic bilateral filter (BF) as a case study and identify three key limitations hindering its performance: excessive parameter tuning, suboptimal neighborhood quality, and fixed parameters across the entire model. To address them, we propose CustomBF, a novel framework for customizing BF components at a per-point level. CustomBF employs multigraph encoders and a mutual guidance strategy to analyze local patches, enabling the customization of BF components including center point normal, neighborhood point coordinates, Gaussian function parameters, and neighborhood radius for each point. Experimental results demonstrate that this component-customized bilateral filter outperforms state-of-the-art methods and achieves robust denoising even in complex scenarios. It highlights the potential of hybrid methods to extend the applicability and effectiveness of traditional techniques.
Peng Li 0064, Zeyong Wei, Honghua Chen, Xuefeng Yan 0001, Mingqiang Wei
ACM Trans. Graph.1
2025 PointCG: Self-Supervised Point Cloud Learning via Joint Completion and Generation
abstract
The core of self-supervised point cloud learning lies in setting up appropriate pretext tasks, to construct a pre-training framework that enables the encoder to perceive 3D objects effectively. In this article, we integrate two prevalent methods, masked point modeling (MPM) and 3D-to-2D generation, as pretext tasks within a pre-training framework. We leverage the spatial awareness and precise supervision offered by these two methods to address their respective limitations: ambiguous supervision signals and insensitivity to geometric information. Specifically, the proposed framework, abbreviated as PointCG, consists of a Hidden Point Completion (HPC) module and an Arbitrary-view Image Generation (AIG) module. We first capture visible points from arbitrary views as inputs by removing hidden points. Then, HPC extracts representations of the inputs with an encoder and completes the entire shape with a decoder, while AIG is used to generate rendered images based on the visible points' representations. Extensive experiments demonstrate the superiority of the proposed method over the baselines in various downstream tasks. Our code will be made available upon acceptance.
Yun Liu 0002, Peng Li 0064, Xuefeng Yan 0001, Liangliang Nan, Bing Wang 0013, Honghua Chen, Lina Gong, Wei Zhao 0039, Mingqiang Wei
IEEE Trans. Vis. Comput. Graph.2
2024 GeoDC: Geometry-Constrained Depth Completion With Depth Distribution Modeling
abstract
Depth completion is a fundamental, yet not well-solved problem in 3-D vision. Current wisdom attempts to employ implicit geometric spatial cues from point clouds to assist in depth completion. However, these methods encounter challenges in extracting rich geometric features due to the absence of explicit constraints. In this article, we propose GeoDC, a geometry-constrained depth completion network with depth distribution modeling. GeoDC employs point cloud upsampling as an auxiliary task to guide the network in learning more robust and effective geometric features. Simultaneously, a novel image and point cloud fusion module, denoted as IP-Interaction, is implemented to holistically integrate features from images and point clouds. Besides, recognizing the presence of uncertainty and ambiguity in the ground-truth (GT) data, we construct a prior network and a posterior network to model depth feature distributions and leverage the distributions to guide depth map inference. GeoDC can solve both the problems of geometric constraint inadequacies in feature extraction and data uncertainty within depth maps well. Extensive experiments underscore the efficacy of our method, demonstrating comparable or superior performance when compared to existing state-of-the-art methods.
Peng Li 0064, Xuefeng Yan 0001, Honghua Chen, Mingqiang Wei
IEEE Trans. Geosci. Remote. Sens.1
2024 eViTBins: Edge-Enhanced Vision-Transformer Bins for Monocular Depth Estimation on Edge Devices
abstract
Monocular depth estimation (MDE) remains a fundamental yet not well-solved problem in computer vision. Current wisdom of MDE often achieves blurred or even indistinct depth boundaries, degenerating the quality of vision-based intelligent transportation systems. This paper presents an edge-enhanced vision transformer bins network for monocular depth estimation, termed eViTBins. eViTBins has three core modules to predict monocular depth maps with exceptional smoothness, accuracy, and fidelity to scene structures and object edges. First, a multi-scale feature fusion module is proposed to circumvent the loss of depth information at various levels during depth regression. Second, an image-guided edge-enhancement module is proposed to accurately infer depth values around image boundaries. Third, a vision transformer-based depth discretization module is introduced to comprehend the global depth distribution. Meanwhile, unlike most MDE models that rely on high-performance GPUs, eViTBins is optimized for seamless deployment on edge devices, such as NVIDIA Jetson Nano and Google Coral SBC, making it ideal for real-time intelligent transportation systems applications. Extensive experimental evaluations corroborate the superiority of eViTBins over competing methods, notably in terms of preserving depth edges and global depth representations.
Yutong She, Peng Li 0064, Mingqiang Wei, Dong Liang 0008, Yiping Chen 0002, Haoran Xie 0001, Fu Lee Wang
IEEE Trans. Intell. Transp. Syst.2
2023 Geogcn: Geometric Dual-Domain Graph Convolution Network For Point Cloud Denoising
abstract
We propose GeoGCN, a novel geometric dual-domain graph convolution network for point cloud denoising (PCD). Beyond the traditional wisdom of PCD, to fully exploit the geometric information of point clouds, we define two kinds of surface normals, one is called Real Normal (RN), and the other is Virtual Normal (VN). RN preserves the local details of noisy point clouds while VN avoids the global shape shrinkage during denoising. GeoGCN is a new PCD paradigm that, 1) first regresses point positions by spatial-based GCN with the help of VNs, 2) then estimates initial RNs by performing Principal Component Analysis on the regressed points, and 3) finally regresses fine RNs by normal-based GCN. Unlike existing PCD methods, GeoGCN not only exploits two kinds of geometry expertise (i.e., RN and VN) but also benefits from training data. Experiments validate that GeoGCN outperforms SOTAs in terms of both noise-robustness and local-and-global feature preservation.
Zhaowei Chen, Peng Li 0064, Zeyong Wei, Honghua Chen, Haoran Xie 0001, Mingqiang Wei, Fu Lee Wang
ICASSP2
2023 ISmallNet: Densely Nested Network with Label Decoupling for Infrared Small Target Detection
abstract
Small targets are often submerged in cluttered backgrounds of infrared images. Conventional detectors tend to generate false alarms, while CNN-based detectors lose small targets in deep layers. To this end, we propose iSmallNet, a multi-stream densely nested network with label decoupling for infrared small object detection. On the one hand, to fully exploit the shape information of small targets, we decouple the original labeled ground-truth (GT) map into an interior map and a boundary one. The GT map, in collaboration with the two additional maps, tackles the unbalanced distribution of small object boundaries. On the other hand, two key modules are delicately designed and incorporated into the proposed network to boost the overall performance. First, to maintain small targets in deep layers, we develop a multi-scale nested interaction module to explore a wide range of context information. Second, we develop an interior-boundary fusion module to integrate multi-granularity information. Experiments on NUAA-SIRST and NUDT-SIRST clearly show the superiority of iSmallNet over 11 state-of-the-art detectors.
Zhiheng Hu, Yongzhen Wang 0001, Peng Li 0064, Haoran Xie 0001, Mingqiang Wei
ICASSP3
2023 ifUNet++: Iterative Feedback UNet++ for Infrared Small Target Detection
abstract
Small targets are often submerged in the cluttered backgrounds of infrared images. In this paper, we propose an iterative feedback UNet++ for infrared small target detection, dubbed ifUNet++. Unlike most of existing methods, ifU-Net++ enables to concentrate on small targets while weakening the interference of clutter backgrounds. ifUNet++ contains two parts: a simplified UNet++ and an iterative feedback strategy. We reduce the unnecessary nodes of UNet++ and have the simplified UNet++ as our backbone network, avoiding the loss of infrared small targets. Based on the simplified network, we search the infrared small targets in an iterative feedback manner, avoiding the interference of cluttered backgrounds. Besides, to optimize the iterative results, we propose Contextual Multiple Attention (CMA) to enhance the features in each iteration. Experimental results exhibit the clear promotion of ifUNet++ over eight state-of-the-art methods, in terms of noise-robustness and detection accuracy.
Zhangying Weng, Peng Li 0064, Xin Zhuang, Xuefeng Yan 0001, Lina Gong, Haoran Xie 0001, Mingqiang Wei
ICASSP2
2023 CF-YOLO: Cross Fusion YOLO for Object Detection in Adverse Weather With a High-Quality Real Snow Dataset
abstract
Snow is one of the toughest adverse weather conditions for object detection (OD). Currently, not only there is a lack of snowy OD datasets to train cutting-edge detectors, but also these detectors have difficulties of learning latent information beneficial for detection in snow. To alleviate the two above problems, we first establish a real-world snowy OD dataset, named RSOD. Besides, we develop an unsupervised training strategy with a distinctive activation function, called$Peak Act$, to quantitatively evaluate the effect of snow on each object. Peak Act helps grade the images in RSOD into four-difficulty levels. To our knowledge, RSOD is the first quantitatively evaluated and graded real-world snowy OD dataset. Then, we propose a novel Cross Fusion (CF) block to construct a lightweight OD network based on YOLOv5s (called CF-YOLO). CF is a plug-and-play feature aggregation module, which integrates the advantages of Feature Pyramid Network and Path Aggregation Network in a simpler yet more flexible form. Both RSOD and CF lead our CF-YOLO to possess an optimization ability for OD in real-world snow. That is, CF-YOLO can handle unfavorable detection problems of vagueness, distortion and covering of snow. Experiments show that our CF-YOLO achieves better detection results on RSOD, compared to SOTAs. The code and dataset are available athttps://github.com/qqding77/CF-YOLO-and-RSOD.
Qiqi Ding, Peng Li 0064, Xuefeng Yan 0001, Ding Shi, Luming Liang, Weiming Wang 0002, Haoran Xie 0001, Jonathan Li 0001, Mingqiang Wei
IEEE Trans. Intell. Transp. Syst.2
2022 I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection
abstract
Can you find me? By simulating how humans to discover the so-called 'perfectly'-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to highlight the separator (or say the camouflaged object's boundary) between an image's background and foreground: the reverse attention stream helps erase the camouflaged object's interior to focus on the background, while the normal attention stream recovers the interior and thus pay more attention to the foreground; and both streams are followed by a boundary guider module and combined to strengthen the understanding of boundary. The core design of such separated attention is motivated by the COD procedure of humans: find the subtle difference between the foreground and background to delineate the boundary of a camouflaged object, then the boundary can help further enhance the COD accuracy. We validate on three benchmark datasets that the proposed BSA-Net is very beneficial to detect camouflaged objects with the blurred boundaries and similar colors/patterns with their backgrounds. Extensive results exhibit very clear COD improvements on our BSA-Net over sixteen SOTAs.
Peng Li 0064, Haoran Xie 0001, Xuefeng Yan 0001, Dong Liang 0008, Dapeng Chen, Mingqiang Wei, Harry Qin
AAAI2
2022 FindNet: Can You Find Me? Boundary-and-Texture Enhancement Network for Camouflaged Object Detection
abstract
Camouflaged objects share very similar colors but have different semantics with the surroundings. Cognitive scientists observe that both the global contour (i.e., boundary) and the local pattern (i.e., texture) of camouflaged objects are key cues to help humans find them successfully. Inspired by the cognitive scientist's observation, we propose a novel boundary-and-texture enhancement network (FindNet) for camouflaged object detection (COD) from single images. Different from most of existing COD methods, FindNet embeds both the boundary-and-texture information into the camouflaged object features. The boundary enhancement (BE) module is leveraged to focus on the global contour of the camouflaged object, and the texture enhancement (TE) module is utilized to focus on the local pattern. The enhanced features from BE and TE, which complement each other, are combined to obtain the final prediction. FindNet performs competently on various conditions of COD, including slightly clear boundaries but very similar textures, fuzzy boundaries but slightly differentiated textures, and simultaneous fuzzy boundaries and textures. Experimental results exhibit clear improvements of FindNet over fifteen state-of-the-art methods on four benchmark datasets, in terms of detection accuracy and boundary clearness. The code will be publicly released.
Peng Li 0064, Xuefeng Yan 0001, Mingqiang Wei, Xiao-Ping Zhang 0002, Harry Qin
IEEE Trans. Image Process.1
2015 Road Boundaries Detection Based on Local Normal Saliency From Mobile Laser Scanning Data
abstract
The accurate extraction of roads is a prerequisite for the automatic extraction of other road features. This letter describes a method for detecting road boundaries from mobile laser scanning (MLS) point clouds in an urban environment. The key idea of our method is directly constructing a saliency map on 3-D unorganized point clouds to extract road boundaries. The method consists of four major steps, i.e., road partition with the assistance of the vehicle trajectory, salient map construction and salient points extraction, curb detection and curb lowest points extraction, and road boundaries fitting. The performance of the proposed method is evaluated on the point clouds of an urban scene collected by a RIEGL VMX-450 MLS system. The completeness, correctness, and quality of the extracted road boundaries are 95.41%, 99.35%, and 94.81%, respectively. Experimental results demonstrate that our method is feasible for detecting road boundaries in MLS point clouds.
Hanyun Wang, Huan Luo 0001, Chenglu Wen, Jun Cheng 0002, Peng Li 0064, Yiping Chen 0002, Cheng Wang 0003, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2014 Object Detection in Terrestrial Laser Scanning Point Clouds Based on Hough Forest
abstract
This letter presents a novel rotation-invariant method for object detection from terrestrial 3-D laser scanning point clouds acquired in complex urban environments. We utilize the Implicit Shape Model to describe object categories, and extend the Hough Forest framework for object detection in 3-D point clouds. A 3-D local patch is described by structure and reflectance features and then mapped to the probabilistic vote about the possible location of the object center. Objects are detected at the peak points in the 3-D Hough voting space. To deal with the arbitrary azimuths of objects in real world, circular voting strategy is introduced by rotating the offset vector. To deal with the interference of adjacent objects, distance weighted voting is proposed. Large-scale real-world point cloud data collected by terrestrial mobile laser scanning systems are used to evaluate the performance. Experimental results demonstrate that the proposed method outperforms the state-of-the-art 3-D object detection methods.
Hanyun Wang, Cheng Wang 0003, Huan Luo 0001, Peng Li 0064, Ming Cheng 0002, Chenglu Wen, Jonathan Li 0001
IEEE Geosci. Remote. Sens. Lett.4