Lunhao Duan

dblp:261/9477 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 43% Generative modeling · 16% Vision and language · 11%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
point cloud analysis
1.422024
Local-consistent Transformation Learning for Rotation-invariant Point Cloud Analysis · CVPR 2024
ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud Understanding · NeurIPS 2023
Machine learning › Generative modeling
diffusion model
1.122025
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation · CVPR 2025
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning · ICCV 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding · ACL (1) 2026
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning
0.912025
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning · ICCV 2025
Computer vision › Vision and language
multimodal fusion
0.912025
High-Quality Pseudo-Labeling for Point Cloud Segmentation With Scene-Level Annotation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud representation learning
0.912025
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning · ICCV 2025
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
0.912025
High-Quality Pseudo-Labeling for Point Cloud Segmentation With Scene-Level Annotation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label generation
0.912025
High-Quality Pseudo-Labeling for Point Cloud Segmentation With Scene-Level Annotation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Visual content generation and editing › image generation
controllable image generation
0.912025
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation · CVPR 2025
Computer vision › 3D vision › invariant feature extraction
rotation invariance
0.812024
Local-consistent Transformation Learning for Rotation-invariant Point Cloud Analysis · CVPR 2024
Computer vision › 3D vision › point cloud analysis › point cloud learning
transformer-based point cloud learning
0.712023
ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud Understanding · NeurIPS 2023
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
0.312026
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding · ACL (1) 2026
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.312025
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning · ICCV 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312025
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation · CVPR 2025
Computer vision › Segmentation and scene understanding
part segmentation
0.212024
Local-consistent Transformation Learning for Rotation-invariant Point Cloud Analysis · CVPR 2024
Computer vision › Image recognition and object detection
shape recognition
0.212024
Local-consistent Transformation Learning for Rotation-invariant Point Cloud Analysis · CVPR 2024

Methods — techniques the papers use, named apart from their topics

rotary position embedding · 1.7cross-attention · 1.7multimodal retrieval · 1.0region-voting · 0.9pseudo-labeling · 0.9multimodal diffusion transformer · 0.9multi-modal diffusion transformer · 0.9feature alignment · 0.9diffusion model · 0.9cross-modal feature guidance · 0.9local reference frame · 0.8
YearPublicationVenuePosition
2026 Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
abstract
Sensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan, Yong Xien Chng, Qing-Guo Chen, Weihua Luo, Kaifu Zhang, Jia-Wang Bian, Mingming Gong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sensen Gao, Lunhao Duan, Yong Xien Chng, Weihua Luo, Kaifu Zhang, Jiawang Bian, Mingming Gong
ACL (1)4
2025 UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
abstract
Recently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descriptions. However, these models often struggle with achieving precise control over pixel-level layouts, object appearances, and global styles when using text prompts alone. To mitigate this issue, previous works introduce conditional images as auxiliary inputs for image generation, enhancing control but typically necessitating specialized models tailored to different types of reference inputs. In this paper, we explore a new approach to unify controllable generation within a single framework. Specifically, we propose the unified image-instruction adapter (UNIC-Adapter) built on the Multi-Modal-Diffusion Transformer architecture, to enable flexible and controllable generation across diverse conditions without the need for multiple specialized models. Our UNIC-Adapter effectively extracts multi-modal instruction information by incorporating both conditional images and task instructions, injecting this information into the image generation process through a cross-attention mechanism enhanced by Rotary Position Embedding. Experimental results across a variety of tasks, including pixel-level spatial control, subject-driven image generation, and style-image-based image synthesis, demonstrate the effectiveness of our UNIC-Adapter in unified controllable image generation.
Lunhao Duan, Shanshan Zhao 0001, Yinglun Li, Weihua Luo, Kaifu Zhang, Mingming Gong, Gui-Song Xia
CVPR1
2025 Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
abstract
Diffusion-based models, widely used in text-to-image generation, have proven effective in 2D representation learning. Recently, this framework has been extended to 3D self-supervised learning by constructing a conditional point generator for enhancing 3D representations. However, its performance remains constrained by the 3D diffusion model, which is trained on the available 3D datasets with limited size. We hypothesize that the robust capabilities of text-to-image diffusion models, particularly Stable Diffusion (SD), which is trained on large-scale datasets, can help overcome these limitations. To investigate this hypothesis, we propose PointSD, a framework that leverages the SD model for 3D self-supervised learning. By replacing the SD model's text encoder with a 3D encoder, we train a point-to-image diffusion model that allows point clouds to guide the denoising of rendered noisy images. With the trained point-to-image diffusion model, we use noise-free images as the input and point clouds as the condition to extract SD features. Next, we train a 3D backbone by aligning its features with these SD features, thereby facilitating direct semantic learning. Comprehensive experiments on downstream point cloud tasks and ablation studies demonstrate that the SD model can enhance point cloud self-supervised learning. Code is publicly available at https://github.com/wdttt/PointSD.
Yiyang Chen 0002, Shanshan Zhao 0001, Lunhao Duan, Changxing Ding, Dacheng Tao
ICCV3
2025 High-Quality Pseudo-Labeling for Point Cloud Segmentation With Scene-Level Annotation
abstract
This paper investigates indoor point cloud semantic segmentation under scene-level annotation, which is less explored compared to methods relying on sparse point-level labels. In the absence of precise point-level labels, current methods first generate point-level pseudo-labels, which are then used to train segmentation models. However, generating accurate pseudo-labels for each point solely based on scene-level annotations poses a considerable challenge, substantially affecting segmentation performance. Consequently, to enhance accuracy, this paper proposes a high-quality pseudo-label generation framework by exploring contemporary multi-modal information and region-point semantic consistency. Specifically, with a cross-modal feature guidance module, our method utilizes 2D-3D correspondences to align point cloud features with corresponding 2D image pixels, thereby assisting point cloud feature learning. To further alleviate the challenge presented by the scene-level annotation, we introduce a region-point semantic consistency module. It produces regional semantics through a region-voting strategy derived from point-level semantics, which are subsequently employed to guide the point-level semantic predictions. Leveraging the aforementioned modules, our method can rectify inaccurate point-level semantic predictions during training and obtain high-quality pseudo-labels. Significant improvements over previous works on ScanNet v2 and S3DIS datasets under scene-level annotation can demonstrate the effectiveness. Additionally, comprehensive ablation studies validate the contributions of our approach's individual components.
Lunhao Duan, Shanshan Zhao 0001, Xingxing Weng, Jing Zhang 0037, Gui-Song Xia
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Localization of Ground-Based Periodic Pulse Interferers Using Time Difference of Arrival Estimation in SAR Satellite Systems
Shengqi Zhou, Xingyu Lu 0003, Jianchao Yang, Huizhang Yang, Junpeng Du, Lunhao Duan, Wenchao Yu, Ke Tan 0007, Shaojia Ge, Hong Gu 0002
IEEE Trans. Geosci. Remote. Sens.6
2024 Local-consistent Transformation Learning for Rotation-invariant Point Cloud Analysis
abstract
Rotation invariance is an important requirement for point shape analysis. To achieve this, current state-of-the-art methods attempt to construct the local rotation-invariant representation through learning or defining the local reference frame (LRF). Although efficient, these LRF-based methods suffer from perturbation of local geometric relations, resulting in suboptimal local rotation invariance. To alleviate this issue, we propose a Local-consistent Transformation (LocoTrans) learning strategy. Specifically, we first construct the local-consistent reference frame (LCRF) by considering the symmetry of the two axes in LRF. In comparison with previous LRFs, our LCRF is able to preserve local geometric relationships better through performing local-consistent transformation. However, as the consistency only exists in local regions, the relative pose information is still lost in the intermediate layers of the network. We mitigate such a relative pose issue by developing a relative pose recovery (RPR) module. RPR aims to restore the relative pose between adjacent transformed patches. Equipped with LCRF and RPR, our LocoTrans is capable of learning local-consistent transformation and preserving local geometry, which benefits rotation invariance learning. Competitive performance under arbitrary rotations on both shape classification and part segmentation tasks and ablations can demonstrate the effectiveness of our method. Code will be available publicly at https://github.com/wdttt/LocoTrans.
Yiyang Chen 0002, Lunhao Duan, Shanshan Zhao 0001, Changxing Ding, Dacheng Tao
CVPR2
2024 A New Method of Noise Frequency Modulated Interference Suppression for SAR
abstract
Synthetic aperture radar (SAR) is vulnerable to interference, including intentional and unintentional-ones. Noise frequency modulated (FM) interference is a kind of intentional interference, which has the characteristics of broadband and randomness, which makes the noise FM signal become a kind of most commonly used interference signal. Noise FM interference will have a serious impact on the SAR image, but the current algorithms for interference suppression are not sufficiently studied. This paper extends a time-domain cancellation algorithm for suppressing the noise FM interference of SAR. This algorithm can reconstruct the noise FM interference signal from the contaminated SAR echo, and then suppress the interference component in the echo by time-domain cancellation. Finally, this paper validates the superior performance of the algorithm by point target simulation and Radarsat-1 data. The proposed method is valid even when the signal-to-interference ratio is lower than -40dB.
Lunhao Duan, Xingyu Lu 0003, Shengqi Zhou, Jianchao Yang, Ke Tan 0007, Zheng Dai, Wenchao Yu, Hong Gu 0002
IGARSS1
2024 A multi-view references image super-resolution framework for generating the large-FOV and high-resolution image
Jiaqin Jiang, Li Li 0047, Lunhao Duan, Jian Yao 0002
J. Vis. Commun. Image Represent.4
2023 ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud Understanding
abstract
Transformers have been recently explored for 3D point cloud understanding with impressive progress achieved. A large number of points, over 0.1 million, make the global self-attention infeasible for point cloud data. Thus, most methods propose to apply the transformer in a local region, e.g., spherical or cubic window. However, it still contains a large number of Query-Key pairs, which requires high computational costs. In addition, previous methods usually learn the query, key, and value using a linear projection without modeling the local 3D geometric structure. In this paper, we attempt to reduce the costs and model the local geometry prior by developing a new transformer block, named ConDaFormer. Technically, ConDaFormer disassembles the cubic window into three orthogonal 2D planes, leading to fewer points when modeling the attention in a similar range. The disassembling operation is beneficial to enlarging the range of attention without increasing the computational complexity, but ignores some contexts. To provide a remedy, we develop a local structure enhancement strategy that introduces a depth-wise convolution before and after the attention. This scheme can also capture the local geometric information. Taking advantage of these designs, ConDaFormer captures both long-range contextual information and local priors. The effectiveness is demonstrated by experimental results on several 3D point cloud understanding benchmarks. Our code will be available.
Lunhao Duan, Shanshan Zhao 0001, Nan Xue 0001, Mingming Gong, Gui-Song Xia, Dacheng Tao
NeurIPS1
2023 Robust Extraction of Vectorized Buildings via Bidirectional Tracing of Keypoints From Remotely Sensed Imagery
abstract
Automatic extraction of vector polygons of buildings from remotely sensed images is an important but difficult task. Recent existing methods based on deep learning usually adopt a multi-stage solution of semantic segmentation, contour detection, and polygon simplification. Such a long processing chain may lead to unreliable results as the boundary regularization and optimization processes are ultimately completed by utilizing low-level features, which ignores the potential of deep features in polygon generation. In this paper, we present an algorithm for directly extracting simplified polygons of buildings in remotely sensed images. The key of this task is the encoding of the polygon structure. PolyMapper [1] utilizes a recurrent neural network (RNN) to produce vertices of a polygon sequentially. Due to the limitation of RNN, this approach is unstable and difficult to deal with objects with complex shapes. In this work, we encode the polygon into a tensor representation and utilize a non-recurrent manner to recover the polygon structure. In our algorithm, two types of points are utilized, i.e., the corner point and the connecting point. Corner points are utilized to delineate the building outlines and form the vertices of the final polygon. Meanwhile, connecting points are sampled from the edges of the buildings for the assistance of the connection of the corner points. Furthermore, we predict the forward and backward directions of each keypoint in a polygon and propose a bidirectional tracing strategy for the polygon structure recovery. Our approach is simple, effective and robust. Experiments on public datasets demonstrate the superiority of the proposed algorithm. The code is made publicly available at https://github.com/sz94/bldvec.
Zhen Shu, Xiangyun Hu, Hengming Dai, Lunhao Duan, Litong Zhang
IEEE Trans. Geosci. Remote. Sens.4
2020 Multiscale Refinement Network for Water-Body Segmentation in High-Resolution Satellite Imagery
abstract
Water-body segmentation in high-resolution satellite imagery is challenging because of the significant variations in the appearance, size, and shape of water bodies. In this letter, a novel multiscale refinement network (MSR-Net) is proposed for water-body segmentation. Similar to most learning-based methods, the MSR-Net resorts to the multiscale information for segmentation, but it improves existing networks in two ways: First, it uses the multiscale information in a new perspective. Instead of the traditional one-off manner that concatenates features and conducts segmentation on one uniform scale, the MSR-Net adopts a new multiscale refinement scheme that makes full use of the multiscale features for more accurate water-body segmentation. In addition, a novel erasing-attention module is designed for an effective feature embedding during the refinement scheme. Experiments on the Gaofen Image Data Set and the DeepGlobe Data Set demonstrate the superiority of MSR-Net when compared with the other state-of-the-art semantic segmentation methods, including U-Net, SegNet, DeepLabv3+, and ExFuse.
Lunhao Duan, Xiangyun Hu
IEEE Geosci. Remote. Sens. Lett.1