Liqun Kuang

dblp:120/6755 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0003-3276-5748ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Agglomerative Model Meets Multi-Scale Adaptive Fusion for Cross-Modal Unsupervised Domain Adaptation
abstract
Although cross-modal unsupervised domain adaptation (UDA) has emerged as an effective means to reduce annotation costs for semantic segmentation, existing approaches often struggle to learn semantically consistent, domain-invariant representations under severe domain shifts. To narrow this representational gap, we develop a new strategy that couples an agglomerative visual foundation model–based encoding paradigm with a Multi-scale Adaptive Fusion Module (MAFM). Specifically, we adopt the visual foundation model C-RADIOv3 as the 2D encoder and tailor its representations for dense prediction while keeping the encoder frozen, which yields robust, high-quality features under varying image resolutions. We further introduce MAFM that builds parallel multi-resolution branches on top of the enhanced 2D features and leverages dynamic large and small-kernel convolutions to broaden the receptive field. After cross-scale alignment, MAFM performs per-location adaptive reweighting to fuse complementary cues, thereby strengthening the joint representation of global semantics and fine-grained boundaries. Experiments across multiple scenarios, including USA → Singapore and A2D2 → SemanticKITTI, demonstrate consistent improvements over strong baselines on the 2D and fusion branches; qualitative visualizations further reveal notably better recognition of small-object classes and sharper semantic boundaries, corroborating the effectiveness of the proposed method.
Zhixun Wang, Liqun Kuang, Shichao Jiao, Fengguang Xiong
ICMR2
2026 ACT-Agent: Affinity-Cross Transformer for Point Cloud Registration via Reinforcement Learning
abstract
ABSTRACT Point cloud registration, a core task in 3D computer vision for aligning two point clouds via rotation and translation, underpins critical applications like robotic navigation and 3D reconstruction. Classical methods (e.g., Iterative Closest Point) easily converge to local minima under poor initial alignment. Deep learning–based approaches, while efficient, suffer from high annotation costs for large‐scale data. Existing reinforcement learning (RL)‐based methods rely on simple PointNet feature extractors, which are insensitive to local geometric details and thus yield suboptimal registration precision. To address these challenges, we propose ACT‐Agent: Affinity‐Cross Transformer for point cloud registration via reinforcement learning, a novel method that formulates point cloud registration as an RL Markov decision process for iterative optimisation. We leverage Pointnet and Affinity‐Cross Transformer to extract and enhance expressive salient features and assign adaptive weights to channels based on their relative importance. We use RL to autonomously learn from feedback in the environment, freeing ourselves from dependence on data annotation. Experimental results on ModelNet40 (synthetic data) and ScanObjectNN (real‐world data) demonstrate that our proposed ACT‐Agent achieves higher accuracy, efficiency, and generalisation ability than the state‐of‐the‐art methods of point cloud registration.
Fengguang Xiong, Haixin Gong, Qiao Ma, Yingbo Jia, Ruize Guo, Ligang He, Liqun Kuang, Xie Han 0001
IET Image Process.8
2026 DeepPAT: Deep position-aware transformer with mixed dataset for robust point cloud registration
Fengguang Xiong, Ligang He, Liqun Kuang, Xie Han 0001
Neurocomputing4
2025 The Source Image Is the Best Attention for Infrared and Visible Image Fusion
Liqun Kuang, Boying Wang, Zherui Qiao, Bingyu Zhang, Zhixun Wang
ICCV3
2025 Multi-modal semantic embedding network for 3D shape recognition and retrieval
Shichao Jiao, Liye Long, Liqun Kuang, Fengguang Xiong, Xie Han 0001
J. Vis. Commun. Image Represent.3
2025 A novel iterative deformable joint attention network for remote sensing image change detection
Qiao Ma, Yingbo Jia, Haixin Gong, Ruize Guo, Zhengtang Li, Xie Han 0001, Liqun Kuang, Fengguang Xiong
Multim. Syst.8
2024 Physical-Simulation-Based Dynamic Template Matching Method for Remote Sensing Small Object Detection
abstract
Object detection is a fundamental challenge encountered in understanding and analyzing remote sensing images. Many current detection methods struggle to identify objects in remote sensing images due to poor feature saliency, diverse directions, and unique viewing angles of small remote sensing targets. As most remote sensing images are captured through overhead photography, the structural contour features of objects in these images remain relatively stable, such as the cross-shaped structure of the aircraft and the rectangular structure of the vehicle. Therefore, we utilize the physical simulation images as prior knowledge to supplement the invariant structural features of the target, extract the target’s key geometric features through the dynamic matching method, and fuse it with the features extracted by the neural network to detect the remote sensing small target more effectively in this article. We propose a template matching method based on physical simulation images (TMSI) and add it to the modified Darknet-53 (named TMSI-Net) for small target detection. Based on this, we perform specific transformations on existing templates in terms of target scale and direction to achieve better adaptation to the corresponding goals and propose DTMSI-Net with a dynamic template library (Dynamic TMSI-Net). Experiments on several datasets demonstrate that the DTMSI-Net exhibits higher detection accuracy and more stable performance compared with the state-of-the-art methods, 4% and 2.7% higher than the second-ranked model on the VEDAI and NWPU datasets, respectively.
Yaming Cao, Lei Guo 0019, Fengguang Xiong, Liqun Kuang, Xie Han 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Exploring the Point Feature Relation on Point Cloud for Multi-View Stereo
abstract
Learning-based multi-view stereo (MVS) is gaining prominence as a method for 3D reconstruction. However, existing methods in the process of feature learning fail to focus on the structural information implied in the scene. This oversight prevents the network from perceiving the geometric properties of the scene and weakens the generalizability of the network. Therefore, we propose a novel framework named Point Feature Relation Network for Multi-view Stereo (PFR-MVSNet), which is composed of a Dynamic Structure Perception (DSP) module, an Adaptive Feature Exploration (AFE) module, and a Point Transformer Block (PTB) module, to solve the problems caused by the oversight. The DSP module first augments the feature of the 3D point cloud from multi-view 2D features, then establishes spatial structure relations within local regions on the point cloud and guides the feature learning of points through the aggregated structure information. After the network has fully learned the intra-region structure features, the AFE module repartitions perception regions with similar features. The point features within the regions are further learned by the PTB module. We evaluate our method on three benchmark datasets: DTU, Tanks & Temples, and ETH3D. The experimental results show that our method achieves superior accuracy of 0.289 mm on the DTU dataset and exhibits more robust generalization on the Tanks & Temples and ETH3D datasets compared with other learning-based MVS methods.
Xie Han 0001, Xindong Guo, Liqun Kuang, Xiaowen Yang, Fusheng Sun
IEEE Trans. Circuits Syst. Video Technol.4
2022 SWPT: Spherical Window-Based Point Cloud Transformer
Xindong Guo, Liqun Kuang, Xie Han 0001
ACCV (1)4
2022 SHREC'22 track: Open-Set 3D Object Retrieval
Yifan Feng 0001, Yue Gao 0002, Xibin Zhao, Yandong Guo, Nihar Bagewadi, Nhat-Tan Bui, Hieu Dao, Shankar Gangisetty, Ripeng Guan, Xie Han 0001, Cong Hua, Chidambar Hunakunti, Yu Jiang 0006, Shichao Jiao, Yuqi Ke, Liqun Kuang, Anan Liu, Dinh-Huan Nguyen, Hai-Dang Nguyen, Weizhi Nie, Bang-Dang Pham, Karthik Raikar, Qingmei Tang, Minh-Triet Tran, Jialong Wan, Chenggang Yan 0001, Haoxuan You, Difei Zhu
Comput. Graph.16
2022 Deep cross-modal discriminant adversarial learning for zero-shot sketch-based image retrieval
Shichao Jiao, Xie Han 0001, Fengguang Xiong, Xiaowen Yang, Huiyan Han, Ligang He, Liqun Kuang
Neural Comput. Appl.7