EDBT 2026 Demo / reviewers in the wild / expert
Xian-Feng Han
dblp:206/7959 · also Xianfeng Han
· DBLP profile ↗
24ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0002-4869-4537ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Context Modeling for Unsupervised Non-Rigid Point Cloud CorrespondenceabstractWe address the “long-range ambiguity” problem for unsupervised non-rigid point cloud correspondence, where corresponding points own inconsistent features while different local regions are spatially or geometrically similar. Previous methods struggle with this problem, since local reference frames (LRF) or coordinate-based methods struggle to exclude locally similar or spatially near mismatches, and widely used independent geometric relations might be inconsistent under non-rigid deformation, introducing extra ambiguity. To this end, we propose a novel robust context modeling module (RCM) to alleviatelong-range ambiguityin two aspects: 1) RCM tackles the ambiguity problem by introducing inter-relation attention (IRA), which mines robust cues from the interplay between relative geometric relations. 2) RCM enhances features with accessiblelong-rangeinformation from IRAs, following a local-to-global manner. Our method shows significant improvements in multiple benchmarks, with accurate correspondence over rotation and large deformation perturbation. Specifically, our method achieves a new state-of-the-art performance with correspondence accuracy of 33.9% and mean error of 4.2 on the SURREAL benchmark. Rui Li 0013, Jiaming Guo, Ya'nan He, Zhengbao Wang, Xian-Feng Han, Kun Sun 0002, Jiaqi Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | OW3DPCA: Toward Open-World 3-D Point Cloud AnalysisabstractThe inherent complexity of real-world environments necessitates robust open-world understanding across both 2-D and 3-D domains. To this end, vision-language models (VLMs) have emerged as a promising solution, demonstrating remarkable performance in 2-D visual tasks. However, the prohibitive cost of collecting and annotating 3-D point clouds, as well as the scarcity of large-scale 3-D point cloud-language paired datasets, present significant challenges in generalizing these successes to the 3-D domain. To address these challenges, we propose OW3DPCA, a novel cross-modal contrastive pretraining framework for open-world 3-D point cloud analysis, designed to unleash the power of VLMs for 3-D understanding. Specifically, we synthesize single front-view RGB images that are more compatible with contrastive language-image pretraining’s visual encoder, facilitating open-world knowledge transfer. Furthermore, we leverage a pretrained image-to-caption VLM to generate semantically rich textual prompts. Subsequently, a cross-modal interaction transformer module is designed to enhance feature representations by integrating cross-domain knowledge, thereby mitigating domain gaps between 3-D features and the VLM’s embedding space. Through contrastive learning, the 3-D backbone network is pretrained using point-cloud–image–language triplets to achieve multimodal feature alignment. Extensive experiments on ModelNet10, ModelNet40, ScanObjectNN, and ShapeNet benchmarks demonstrate OW3DPCA’s competitive performance across zero-shot, few-shot, and fully supervised 3-D understanding tasks compared with state-of-the-art methods. Jing-Wei Zhou, Xian-Feng Han |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | HM3D: A Lightweight Hierarchical Mamba Model for Efficient 3D Point Cloud AnalysisabstractTransformer models have revolutionized 3D point cloud analysis tasks, such as classification and segmentation, due to their permutation invariance and exceptional ability to capture global context. However, the self-attention operation in Transformers faces computational complexity that grows quadratically with the number of input points, leading to substantial computational overhead. To mitigate this challenge, we introduce the Mamba mechanism into 3D point cloud learning, and propose an efficient and lightweight hierarchical Mamba architecture, termed HM3D. Specifically, we develop the FPS-Causal and BallQuery-Sort algorithms to handle the unidirectionality problem inherent in State Space Model. In addition, we design a BallMamba layer, which integrates a Local-BallMamba block for capturing fine-grained local details and a Global-BallMamba block to learn global geometric features. Extensive experiments conducted on ModelNet, ShapeNetPart and S3DIS benchmarks demonstrate that our HM3D achieves competitive performance compared to state-of-the-art Transformer-based and Mamba-based models, while significantly reducing the number of parameters and computational FLOPs. The source codes are available at https://github.com/Rtandi8comingWing/HM3D. Xian-Feng Han |
ICMR | 2 |
| 2025 | Open-World 3D Scene Understanding with Cross-Modal Dual Consistency LearningabstractLarge Vision-Language Pre-training models have achieved remarkable advancements in 2D zero-shot or few-shot visual tasks in the 2D domain. However, promoting their potential to benefit the 3D counterparts is full of challenging due to the notable hindrance of limited 3D-text pairs, which results in the open-world 3D scene understanding remaining an unexplored problem. In this paper, we pretrain a cross-modal dual consistency learning based 3D visual-language model to learn semantically-rich 3D point cloud representation. Specifically, we first introduce a Visual Feature Distribution Consistency strategy to bridge the gap between the point clouds and images. Then, a Visual-Semantic Enhancement Feature Distribution Consistency approach is developed to narrow the distance between enhanced visual and language information. Finally, under the supervision of profound knowledge from 2D Vsion-Language Models (VLMs), the learned 3D features achieve powerful generalization capability, facilitating open-world 3D scene understanding. Quantitative and qualitative evaluations on ScanNet and Matterport3D benchmarks demonstrate the effectiveness of our pre-trained method in open-world 3D semantic segmentation. Xian-Feng Han, Yuhang Wang 0036, Mingjie Wang 0002 |
ICMR | 1 |
| 2025 | Draw What You Hear: High-Fidelity Image Generation and Manipulation via SoundAdapterabstractCurrently, the text-to-image (T2I) generation has established itself as a cornerstone within the realm of AI-generated content (AIGC), due its remarkable success to the availability of extensive datasets comprising paired text-vision samples. Nevertheless, the absence of audio-visual pairs hinders the growth of audio-to-image (A2I). Although prior approaches have pioneered the A2I task, the tight entanglement between initial audio and image encoders imposes the challenge of gathering audio-visual samples, resulting in degraded performance and limited sound flexibility. Therefore, this article proposes a novel SoundAdapter to draw what you hear. Specifically, the SoundAdapter's structure is meticulously designed around transformer blocks, which are critical for capturing overarching patterns and dependencies within the data. In addition, it integrates a sophisticated multigranularity approach coupled with a hybrid supervisory signal, ensuring both fine-grained semantic alignment and seamless optimization across various levels of representation. Extensive tests demonstrate that the SoundAdapter excels in training, setting new benchmarks in zero-shot audio classification, as well as in creating and modifying images across a variety of datasets. The implementation code and several demos supporting this study are openly accessible at https://github.com/CV-MM-Lab/SoundAdapter, facilitating reproducibility and further research. Mingjie Wang 0002, Song Yuan, Xian-Feng Han, Zili Yi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Language-Driven Open-Vocabulary 3D Semantic Segmentation with Knowledge Distillationabstract3D open-vocabulary semantic segmentation is a challenge in the task of 3D scene understanding, as most current models trained on closed-set datasets struggle to effectively identify categories that were not seen during training. To address this, we introduce a framework called LSWKD. It distills knowledge from a pre-trained 3D open-world model, thereby enhancing the alignment between visual and semantic features. Furthermore, we employ Point-discriminative Contrastive Learning to compute caption loss in the teacher model instead of CLIP-style Contrastive Loss in order to let each point be supervised with its all related language captions, which improves the teacher model’s performance. We conducted experiments on ScanNet and S3DIS datasets. The results demonstrate that our approach achieves better hIoU compared with state-of-the-art models. Code will be released at https://github.com/wu39848/LSWKD. Xian-Feng Han, Guoqiang Xiao 0001 |
ICASSP | 2 |
| 2024 | PCB-RandNet: Rethinking Random Sampling for LiDAR Semantic Segmentation in Autonomous Driving SceneabstractFast and efficient semantic segmentation of large-scale LiDAR point clouds is a fundamental problem in autonomous driving. To achieve this goal, the existing point-based methods mainly choose to adopt Random Sampling strategy to process large-scale point clouds. However, our quantative and qualitative studies have found that Random Sampling may be less suitable for the autonomous driving scenario, since the LiDAR points follow an uneven or even long-tailed distribution across the space, which prevents the model from capturing sufficient information from points in different distance ranges and reduces the model’s learning capability. To alleviate this problem, we propose a new Polar Cylinder Balanced Random Sampling method that enables the downsampled point clouds to maintain a more balanced distribution and improve the segmentation performance under different spatial distributions. In addition, a sampling consistency loss is introduced to further improve the segmentation performance and reduce the model’s variance under different sampling methods. Extensive experiments confirm that our approach produces excellent performance on both SemanticKITTI and SemanticPOSS benchmarks, achieving a 2.8% and 4.0% improvement, respectively. The source code is available at PCB-RandNet. Xian-Feng Han, Hui-Xian Cheng, Dehong He, Guoqiang Xiao 0001 |
ICRA | 1 |
| 2024 | Yolo-3DMM for Simultaneous Multiple Object Detection and Tracking in Traffic ScenariosabstractVideo-based multiple object tracking (MOT) is a fundamental task in intelligent transportation with applications ranging from automated traffic surveillance to autonomous driving. MOT methods commonly follow a tracking-by-detection paradigm, tracking objects by associating their detections across video frames. However, insofar, these methods have not used the entire vehicle trajectory motion characteristics to perform tracking, which converts the vehicle localization problem into a motion parameter estimation problem. Moreover, MOT methods mainly rely on off-the-shelf detectors. An independently trained detector is sub-optimal for the tracking-by-detection paradigm and adversely affects the overall system performance. In this article, we address these issues by proposing a novel MOT method for moving vehicles in traffic scenarios. Our tracker treats the vehicle tracks as unified 3D spatio-temporal trajectory instances and leverages the power of deep learning to extract vehicle motion from the 3D instances. We propose a new simultaneous detection and tracking network, called YOLO-3D Motion Model Network (Yolo-3DMM) that employs spatio-temporal features of traffic videos for simultaneous vehicle detection and tracking in an end-to-end manner. We adopt a variety of different vehicle tracking datasets to evaluate our method. Moreover, we also propose a tunnel MOT dataset from real highway tunnel surveillance in Guangdong, China to expand the experimental scenarios. To establish the efficacy of our method, we evaluate it on 100 different roadside traffic scenarios. Our method shows excellent performance on UA-DETRAC and Omni-MOT datasets. It achieves a PR-MOTA score of 29.40% on UA-DETRAC and gets a 69.7% MOTA score on the Omni-MOT dataset. Lichen Liu, Huansheng Song, Shijie Sun 0001, Xian-Feng Han, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | CA-SSD: Channel-Independent Rare Class Attention for 3D Object DetectionabstractThe current 3D object detection datasets suffer from a serve problem of uneven data distribution. In autonomous driving scenarios, rare classes are also critical. The existing methods perform poorly in these classes. Generic methods of solving long-tailed distributions such as re-sampling and re-weighting would over-fit or destroy the original distribution. The popular solution is to design detectors for different classes, which improves detection accuracy but reduces generalizability. Motived by this, we propose a simple and effective single-stage detector named CA-SSD, which can detect multiple categories simultaneously. The key to our approach is the channel transformation of the features, which corrects for the bias in task distribution from the training phase to the testing phase. In addition, to reduce the information loss due to multiple downsampling, a self-attention module is introduced to retain richer contextual information. Experimental results on the KITTI benchmark demonstrate the superiority of CA-SSD, especially on small samples. Besides, we designed ablation experiments to verify the effectiveness of both modules. Hui-Xian Cheng, Guoqiang Xiao 0001, Xian-Feng Han |
IJCNN | 4 |
| 2023 | PointCMT: An MLP-Transformer Network for Contrastive Learning of Point RepresentationabstractThe irregular and unordered structure makes deep learning based point cloud analysis become a challenging task. Although the recent success of Transformer frameworks in Natural Language Processing (NLP) leads to a series of key advances in dealing with point cloud, most approaches in this category mainly focus on global information representation. In this paper, we propose a Point based MLP-Transformer Contrastive Learning architecture, termed as PointCMT. Specifically, by the introduction of the key ideas and advantages behind the MLP and Transformer, a locally Multi-scale Contrastive Feature Representation Learning module is developed as the core component of our network, which involves a local neighborhood embedding layer, and a contrastive learning model with our well-designed MLP-Transformer feature extractor (MFormer) as fundamental blocks. Experimental evidences on public benchmark datasets indicate that, compared with most state-of-the-arts, our proposed framework can obtain competitive or even better performance in terms of 3D point cloud analysis tasks. The code and trained models will be available at Github. Xian-Feng Han, Guoqiang Xiao 0001 |
IJCNN | 2 |
| 2023 | IFA-Net: Isomerous Feature-aware Network for Single-view 3D ReconstructionabstractSingle-view 3D reconstruction has long been an intractable and fundamental problem in computer vision. Objects with complex topological structures are difficult to be accurately reconstructed, which makes the existing methods suffer from blurred shape boundaries between multiple components in the object. Recently, convolutional neural network and vision transformer have begun to appear in the field of 3D reconstruction and have been widely used with excellent performance. However, the existing transformer-based methods mainly focus on the global long-term context dependency, and ignore the local details of the part space features, resulting in poor reconstruction of the detail part. In this paper, we propose a novel dual-branch network architecture, called IFA-Net, to capture local spatial perception information and retain global structural features for single-view 3D reconstruction. In addition, we propose an isomerous feature-aware module, which enables the dynamic fusion of different resolution features under the two branches. Thus, high-fidelity and detail-rich 3D object reconstruction can be achieved. Extensive experimental results demonstrate that our method is able to produce high-quality voxels, particularly with diverse topologies, as compared with the state-of-the-art methods. Zecheng Zhang, Xian-Feng Han, Guoqiang Xiao 0001 |
IJCNN | 2 |
| 2023 | Joint Geometric-Semantic Driven Character Line Drawing GenerationabstractCharacter line drawing synthesis can be formulated as a special case of image-to-image translation problem that automatically manipulates the photo-to-line drawing style transformation. In this paper, we present the first generative adversarial network based end-to-end trainable translation architecture, dubbed P2LDGAN, for automatic generation of high-quality character drawings from input photos/images. The core component of our approach is the joint geometric-semantic driven generator, which uses our well-designed cross-scale dense skip connections framework to embed learned geometric and semantic information for generating delicate line drawings. In order to support the evaluation of our model, we release a new dataset including 1,532 well-matched pairs of freehand character line drawings as well as corresponding character images/photos, where these line drawings with diverse styles are manually drawn by skilled artists. Extensive experiments on our introduced dataset demonstrate the superior performance of our proposed models against the state-of-the-art approaches in terms of quantitative, qualitative and human evaluations. Our code, models and dataset will be available at Github. Chengyu Fang 0001, Xian-Feng Han |
ICMR | 2 |
| 2023 | Guided normal filter for 3D point clouds
Zhi-Ao Feng, Xian-Feng Han |
Multim. Tools Appl. | 2 |
| 2023 | TransRVNet: LiDAR Semantic Segmentation With TransformerabstractEffective and efficient 3D semantic segmentation from large-scale LiDAR point cloud is a fundamental problem in the field of autonomous driving. In this paper, we present Transformer-Range-View Network (TransRVNet), a novel and powerful projection-based CNN-Transformer architecture to infer point-wise semantics. First, a Multi Residual Channel Interaction Attention Module (MRCIAM) is introduced to capture channel-level multi-scale feature and model intra-channel, inter-channel correlations based on attention mechanism. Then, in the encoder stage, we use a well-designed Residual Context Aggregation Module (RCAM), including a residual dilated convolution structure and a context aggregation module, to fuse information from different receptive fields while reducing the impact of missing points. Finally, a Balanced Non-square-Transformer Module (BNTM) is employed as fundamental component of decoder to achieve locally feature dependencies for more discriminative feature learning by introducing the non-square shifted window strategy. Extensive qualitative and quantitative experiments conducted on challenging SemanticKITTI and SemanticPOSS benchmarks have verified the effectiveness of our proposed technique. Our TransRVNet presents superior performance over most existing state-of-the-art approaches. The source code and trained model are available athttps://github.com/huixiancheng/TransRVNet. Hui-Xian Cheng, Xian-Feng Han, Guoqiang Xiao 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Dual Transformer for Point Cloud AnalysisabstractFeature representation learning is a key component in 3D point cloud analysis. However, the powerful convolutional neural networks (CNNs) cannot be applied due to the irregular structure of point clouds. Therefore, following the tremendous success of transformer in natural language processing and image understanding tasks, in this paper, we present a novel point cloud representation learning architecture, named Dual Transformer Network (DTNet), which mainly consists of Dual Point Cloud Transformer (DPCT) module. Specifically, by aggregating the well-designed point-wise and channel-wise self-attention models simultaneously, DPCT module can capture much richer contextual dependencies semantically from the perspective of position and channel. With the DPCT model as a fundamental component, we construct the DTNet for performing 3D point cloud analysis in an end-to-end manner. Extensive quantitative and qualitative experiments on publicly available benchmarks demonstrate the effectiveness of our transformer framework for the tasks of 3D point cloud classification, segmentation and visual object affordance understanding, achieving highly competitive performance in comparison with the state-of-the-art approaches. Xian-Feng Han, Yi-Fei Jin, Hui-Xian Cheng, Guoqiang Xiao 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | STNet: Scale Tree Network With Multi-Level Auxiliator for Crowd CountingabstractState-of-the-art approaches for crowd counting resort to deepneural networks to predict density maps. However, counting people in congested scenes remains a challenging task because the presence of drastic scale variation, density inconsistency, and complex background can seriously degrade their counting accuracy. To battle the ingrained issue of accuracy degradation, in this paper, we propose a novel and powerful network called Scale Tree Network (STNet) for accurate crowd counting. STNet consists of two key components: a Scale-Tree Diversity Enhancer and a Multi-level Auxiliator. Specifically, the Diversity Enhancer is designed to enrich scale diversity, which alleviates limitations of existing methods caused by insufficient level of scales. A novel tree structure is adopted to hierarchically parse coarse-to-fine crowd regions. Furthermore, a simple yet effective Multi-level Auxiliator is presented to aid in exploiting generalisable shared characteristics at multiple levels, allowing more accurate pixel-wise background cognition. The overall STNet is trained in an end-to-end manner, without the needs for manually tuning loss weights between the main and the auxiliary tasks. Extensive experiments on five challenging crowd datasets demonstrate the superiority of the proposed method. Mingjie Wang 0002, Hao Cai 0004, Xian-Feng Han, Jun Zhou 0023, Minglun Gong |
IEEE Trans. Multim. | 3 |
| 2022 | Cenet: Toward Concise and Efficient Lidar Semantic Segmentation for Autonomous DrivingabstractAccurate and fast scene understanding is one of the chal-lenging task for autonomous driving, which requires to take full advantage of LiDAR point clouds for semantic segmen-tation. In this paper, we present a concise and efficient image-based semantic segmentation network, named CENet. In order to improve the descriptive power of learned features and reduce the computational as well as time complex-ity, our CEN et integrates the convolution with larger ker-nel size instead of MLP, carefully-selected activation functions, and multiple auxiliary segmentation heads with cor-responding loss functions into architecture. Quantitative and qualitative experiments conducted on publicly available benchmarks, SemanticKITTI and SemanticPOSS, demon-strate that our pipeline achieves much better mIoU and in-ference performance compared with state-of-the-art models. The code will be available at https://github.com/huixiancheng/CENet. Hui-Xian Cheng, Xian-Feng Han, Guoqiang Xiao 0001 |
ICME | 2 |
| 2021 | Cascade Attention Blend Residual Network For Single Image Super-ResolutionabstractNowadays, deep convolutional neural networks are playing an increasingly important role in single-image super-resolution vision applications. Yet, most of the existing deep convolution-based methodologies are insufficiently intelligent to capture targeted information when the distribution of spatial and channel information is uneven for low-resolution images. To address this research issue, we propose a cascade attention blend residual network, with the non-local channel and multi-scale attention being considered for channel-wise dependencies and multi-scale receptive fields, respectively. Cascading both attentions in a potent blend residual block aims to learn more spatial and channel correlations between low- and super-resolution images. Experimental results demonstrate that the proposed method achieves promising performance for super-resolution image reconstruction, as well as gains an average reduction of 50.9% network parameters, compared to some state-of-the-art methods. Guoqiang Xiao 0001, Xiaoqin Tang, Xian-Feng Han, Wenzhuo Ma, Xinye Gou |
ICIP | 4 |
| 2021 | PGMANet: Pose-Guided Mixed Attention Network for Occluded Person Re-IdentificationabstractRecently, the person re-identification task becomes increasingly crucial in crowded scenarios, e.g., airports and schools. Many methods with high performance have been proposed to solve this problem. However, the existence of occlusion still challenges the development of person re-identification. In this work, we present the novel Pose - Guided Mixed Attention Network (PGMANet), an end-to-end framework to deal with pedestrian reidentification under occluded situations by fusing posture and second-order information. Specially, we employ two models. The first is Human Part - level Attention Model. We use key point information of a pedestrian to generate a heat map to enhance the pedestrian body part's feature. Simultaneously, we design Second - order Information Attention Model to investigate the correlation among features of different parts. Experimental results show that our method achieves state-of-the-art person re-identification performance on two challenging occlusion datasets Occluded-DukeMTMC and Occluded-Reid. You Zhai, Xian-Feng Han, Wenzhuo Ma, Xinye Gou, Guoqiang Xiao 0001 |
IJCNN | 2 |
| 2021 | Novel methods for noisy 3D point cloud based object recognition
Xian-Feng Han, Xin-Yu Yan, Shijie Sun 0001 |
Multim. Tools Appl. | 1 |
| 2021 | Image-Based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Eraabstract3D reconstruction is a longstanding ill-posed problem, which has been explored for decades by the computer vision, computer graphics, and machine learning communities. Since 2015, image-based 3D reconstruction using convolutional neural networks (CNN) has attracted increasing interest and demonstrated an impressive performance. Given this new era of rapid evolution, this article provides a comprehensive survey of the recent developments in this field. We focus on the works which use deep learning techniques to estimate the 3D shape of generic objects either from a single or multiple RGB images. We organize the literature based on the shape representations, the network architectures, and the training mechanisms they use. While this survey is intended for methods which reconstruct generic objects, we also review some of the recent works which focus on specific object classes such as human body shapes and faces. We provide an analysis and comparison of the performance of some key papers, summarize some of the open problems in this field, and discuss promising directions for future research. Xian-Feng Han, Hamid Laga, Mohammed Bennamoun |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Iterative guidance normal filter for point cloud
Xian-Feng Han, Jesse S. Jin, Ming-Jie Wang, Wei Jiang 0038 |
Multim. Tools Appl. | 1 |
| 2018 | Guided 3D point cloud filtering
Xian-Feng Han, Jesse S. Jin, Ming-Jie Wang, Wei Jiang 0038 |
Multim. Tools Appl. | 1 |
| 2017 | A review of algorithms for filtering the 3D point cloud
Xian-Feng Han, Jesse S. Jin, Ming-Jie Wang, Wei Jiang 0038, Liping Xiao |
Signal Process. Image Commun. | 1 |