Xianghui Pan

dblp:381/2743 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Rotation-Equivariant Robot Vision: A Perspective via Correspondence-Matching and Pre-training
abstract
Correspondence matching is a fundamental and crucial task in robot vision. In recent years, deep learning-based keypoint matching techniques have shown outstanding performance in downstream tasks. Conventional learning-based correspondence matching methods rely on large datasets and a specific training procedure. Correspondence techniques based on pre-trained features have been preliminarily explored by researchers. Unfortunately, traditional convolutional neural networks only possess translation invariance but lack rotational invariance, hence, their performance suffers significantly under heavy rotations. Therefore, we propose a correspondence matching method based on pre-trained group-equivariant neural networks and compare the performance of various rotation-equivariant to rotation-invariant transformers. We conducted experiments on the Rotated-Hpatches and Rotated-MegaDepth datasets, and the results indicate that our proposed method is concise and effective, achieving state-of-the-art performance without the need for retraining in downstream tasks.
Shuai Su, Xianghui Pan, Jiayuan Du
IROS2
2025 LGPR: Local Feature Learning Brings More Generalizable Visual Place Recognition
abstract
We propose a Visual Place Recognition (VPR) framework by sharing lightweight keypoint extraction modules for local features. Current research on the joint learning of local keypoint matching and VPR is relatively scarce, and the application deployment of real-time spatial computing on edge devices has a high learning cost. There is also a significant spatial structural difference between existing VPR methods and the scenarios in practical applications. To address these issues, we design a joint learning framework for local keypoint extraction and VPR, which shares local features and fuses irregularly distributed key features in space through self-attention and cross-attention mechanisms. Our framework achieves excellent results on several VPR datasets. In particular, we introduce a new VPR dataset, called TJPark, which has a significant spatial information difference from common street view data. Our method demonstrates that local features with strong generalization capabilities effectively help enhance the generalization of VPR. Our open source code and dataset are available at: https://github.com/ShuaiAlger/LGPR.
Shuai Su, Jingwei Yang 0002, Jiayuan Du, Xianghui Pan
IROS4
2025 A Multi-Step ADC With Lightweight Input Buffer Distortion, Sub-Stage Coarse-Fine Gain, and Sampling Skew Background Calibrations
abstract
This paper presents a lightweight background calibration for the distortion of the analog-to-digital converter (ADC)’s input buffer. The buffer’s nonlinearity is calibrated on-chip by the Harmonic-Compensated Diode Load (HC-DL), whose bias voltage is determined by the calibration algorithm facilitated by dither injection at the input of the buffer. A two-step ADC is exploited to demonstrate the calibration, where the coarse-fine ADC gain mismatch and sampling skew at the 1ststep are calibrated by monitoring the occupation of the correction range (OCR) at the 2ndstep. Verified in a 12b 1 GS/s pipe-SAR ADC in 28 nm CMOS, the SNDR and SFDR at Nyquist input are 60.96 dB/76.8 dB, respectively and the SFDR keeps >75 dB over PVT.
Xianghui Pan, Buhui Rui, Yuefeng Cao, Rui Paulo Martins, Yan Zhu 0001, Chi-Hang Chan
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 DVT: Decoupled Dual-Branch View Transformation for Monocular Bird's Eye View Semantic Segmentation
abstract
Monocular Bird’s Eye View (BEV) semantic segmentation is critical for autonomous driving for its inherent advantages in spatial representation and downstream tasks. However, it is challenging to simultaneously learn view transformation and pixel-wise classification. Previous works suffer from non-flat region distortion, distant depth ambiguity, and visual occlusion. To address these aforementioned concerns, we propose dual-branch view transformation (DVT), a novel framework for monocular BEV semantic segmentation. Our method consists of: (i) A dual-branch view transformation to decouple features into flat region and non-flat region and process them independently. (ii) A depth-aware weighting method to make the model pay more attention to the distant depth. (iii) An auxiliary task to introduce more inductive biases to alleviate the inaccuracy caused by visual occlusion. Furthermore, we design a class-aware weighting method to address the class and size imbalance of datasets. Experimental results on nuScenes and KITTI-360 datasets demonstrate that DVT outperforms previous state-of-the-art (SOTA). Our codes are available at https://github.com/MrPicklesGG/DVT.
Jiayuan Du, Xianghui Pan, Mengjiao Shen, Shuai Su, Jingwei Yang 0002
IROS2
2024 GenerOcc: Self-supervised Framework of Real-time 3D Occupancy Prediction for Monocular Generic Cameras
abstract
In the context of 3D scene perception tasks, the significance of 3D occupancy prediction has been progressively growing, aiming to forecast the occupancy state of voxels in a discrete 3D space. However, existing methods typically exhibit several limitations, such as restricted adaptability to non-pinhole cameras due to fixed camera parameters, heavy reliance on 3D annotations because of the inability to project 3D output back to the camera plane, and inferior real-time inference performance resulting from the conversion process from 2D to 3D features. To address these constrains, we introduce GenerOcc, a self-supervised framework of real-time 3D occupancy prediction for monocular generic cameras. We have collected the fisheye Dominant dataset to confirm the compatibility of our ray-based camera model with non-pinhole cameras. By transforming the occupancy prediction task into a depth estimation task in a self-supervised manner, we eliminate dependency on 3D annotations. Furthermore, we propose a parametric voxel probability distribution module that leverages 2D features to quickly predict 3D occupancy without 3D representations of the scene. Additionally, our GenerOcc has been extensively evaluated on public pinhole Occ3D-nuScenes dataset and our proprietary fisheye Dominant dataset, both yielding impressive performance.
Xianghui Pan, Jiayuan Du, Shuai Su, Wenhao Zong
IROS1
2024 A wireframe-detection based method for correcting credential photos
Xianghui Pan, Hangting Lv
Multim. Tools Appl.1