VLDB 2026 Research / reviewers in the wild / expert
Xiaoshui Huang
dblp:167/9599
· DBLP profile ↗
58ranked-venue papers
9as first author
50since 2021 · last 2026
0000-0002-3579-538XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 7 first-author · 36 since 2021Artificial intelligence and machine learning · 34 · 6 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-CancerabstractSingle-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a comprehensive evaluation benchmark. In this paper, we introduce \algoname, a lightweight hybrid neural network that combines the strengths of Transformers and state-space models to achieve a balance between performance and efficiency. \algoname consists of a front-end local-context encoder with shared self-attention layers to capture complex, order-independent gene interactions; and a back-end global sequential feature decoder that efficiently integrates global context using a linear-time state-space model. This modular design preserves the expressive power of Transformers while leveraging the scalability of Mamba to enable transcriptome modeling, effectively capturing both local and global regulatory signals. To enable robust evaluation, we also construct a large-scale pan-cancer single-cell benchmark, \algoname Bench, containing over 3.5 million high-quality cells across 33 cancer subtypes, curated through a rigorous preprocessing pipeline. Experimental results show that \algoname outperforms state-of-the-art models on our pan-cancer benchmark (+4.0\%) and across multiple public tasks, including cell type annotation (+7.4\%), batch integration (+4.0\%) and multi-omics integration (+3.1\%). Xiaoshui Huang, Tianlin Zhu, Yifan Zuo 0001, Xue Xia 0005, Zonghan Wu, Jiebin Yan, Dingli Hua, Zongyi Xu, Yuming Fang 0001, Jian Zhang 0002 |
AAAI | 1 |
| 2026 | Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent DiffusionabstractDenoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To address this, we propose a Robust single-stage fully Sparse 3D object Detection Network with a Detachable Latent Framework (DLF) of DDPMs, named RSDNet. Specifically, RSDNet learns the denoising process in latent feature spaces through lightweight denoising networks like multi-level denoising autoencoders (DAEs). This enables RSDNet to effectively understand scene distributions under multi-level perturbations, achieving robust and reliable detection. Meanwhile, we reformulate the noising and denoising mechanisms of DDPMs, enabling DLF to construct multi-type and multi-level noise samples and targets, enhancing RSDNet robustness to multiple perturbations. Furthermore, a semantic-geometric conditional guidance is introduced to perceive the object boundaries and shapes, alleviating the center feature missing problem in sparse representations, enabling RSDNet to perform in a fully sparse detection pipeline. Moreover, the detachable denoising network design of DLF enables RSDNet to perform single-step detection in inference, further enhancing detection efficiency. Extensive experiments on public benchmarks show that RSDNet can outperform existing methods, achieving state-of-the-art detection. Wentao Qu, Guofeng Mei, Jing Wang 0201, Yujiao Wu, Xiaoshui Huang, Liang Xiao 0001 |
AAAI | 5 |
| 2026 | Masked Clustering Prediction for Unsupervised Point Cloud Pre-trainingabstractVision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We propose MaskClu, a novel unsupervised pre-training method for ViTs on 3D point clouds that integrates masked point modeling with clustering-based learning. MaskClu is designed to reconstruct both cluster assignments and cluster centers from masked point clouds, thus encouraging the model to capture dense semantic information. Additionally, we introduce a global contrastive learning mechanism that enhances instance-level feature learning by contrasting different masked views of the same point cloud. By jointly optimizing these complementary objectives, i.e., dense semantic reconstruction, and instance-level contrastive learning. MaskClu enables ViTs to learn richer and more semantically meaningful representations from 3D point clouds. We validate the effectiveness of MaskClu via multiple 3D tasks, including part segmentation, semantic segmentation, object detection, and classification, setting new competitive results. Bin Ren 0005, Xiaoshui Huang, Mengyuan Liu 0001, Hong Liu 0008, Fabio Poiesi, Nicu Sebe, Guofeng Mei |
AAAI | 2 |
| 2026 | DiffPano++: Scalable and Consistent Multi-View Panorama Generation with Spherical Epipolar-Aware Diffusion
Chenhao Ji, Weicai Ye, Zheng Chen 0016, Junyao Gao 0002, Xiaoshui Huang, Xuekuan Wang, Guofeng Zhang 0001, Song-Hai Zhang, Tong He 0001, Wanli Ouyang, Cairong Zhao |
Int. J. Comput. Vis. | 5 |
| 2026 | Non-target divergence hypothesis: Toward understanding modality differences in cross-modal knowledge distillation
Zongyi Xu, Xiaoshui Huang, Shanshan Zhao 0001, Xinbo Gao 0001 |
Neural Networks | 3 |
| 2025 | PSReg: Prior-guided Sparse Mixture of Experts for Point Cloud RegistrationabstractThe discriminative feature is crucial for point cloud registration. Recent methods improve the feature discriminative by distinguishing between non-overlapping and overlapping region points. However, they still face challenges in distinguishing the ambiguous structures in the overlapping regions. Therefore, the ambiguous features they extracted resulted in a significant number of outlier matches from overlapping regions. To solve this problem, we propose a prior-guided SMoE-based registration method to improve the feature distinctiveness by dispatching the potential correspondences to the same experts. Specifically, we propose a prior-guided SMoE module by fusing prior overlap and potential correspondence embeddings for routing, assigning tokens to the most suitable experts for processing. In addition, we propose a registration framework by a specific combination of Transformer layer and prior-guided SMoE module. The proposed method not only pays attention to the importance of locating the overlapping areas of point clouds, but also commits to finding more accurate correspondences in overlapping areas. Our extensive experiments demonstrate the effectiveness of our method, achieving state-of-the-art registration recall (95.7%/79.3%) on the 3DMatch/3DLoMatch benchmark. Moreover, we also test the performance on ModelNet40 and demonstrate excellent performance. Xiaoshui Huang, Zhou Huang 0006, Yifan Zuo 0001, Yongshun Gong, Chengdong Zhang, Deyang Liu, Yuming Fang 0001 |
AAAI | 1 |
| 2025 | LPCG: A Self-conditional Architecture for Labeled Point Cloud GenerationabstractRecently, there has been considerable exploration of methods for generating 3D point clouds, which is crucial for numerous 3D vision applications. Though conditional generation methods show promising performance, it depends on the additional paired label. On the other hand, unconditional generation methods usually fail to annotate the generated 3D point cloud. In this paper, we introduce a novel self-conditional architecture that trains on unlabeled data and then generates high-quality labeled 3D point clouds. Specifically, we design a module to extract geometry and view features, and then use a feature fusion module to integrate them as a substitute for label embedding in conditional point cloud generation. Then the point cloud generator is trained using the fused features. LPCG also harnesses CLIP to handle the view features of point clouds for generating label information. Besides, we train two feature diffusion modules to capture the essence of multimodal features and obtain diverse fused features for use as conditions in generating point clouds. Experiments on the ShapeNet dataset demonstrate that LPCG achieves state-of-the-art performance for single class generation. Our experimental results show that the accuracy of our generated label annotations reaches around 97.44% for a two-class generation task. Dongshuo Huang, Xiaoshui Huang, Chengdong Zhang, Yilei Shi |
AAAI | 2 |
| 2025 | BrepGiff: Lightweight Generation of Complex B-rep with 3D GAT DiffusionabstractDespite advancements in Computer-Aided-Design (CAD) generation, direct generation of complex Boundary Representation (B-rep) CAD models remains challenging. The difficulty arises from the parametric nature of B-rep data, complicating the encoding and generation of its geometric and topological information. In this paper, we introduce BrepGiff, a lightweight generation approach for high-quality and complex B-rep based on 3D Graph Diffusion. First, we transfer B-rep models into 3D graphs representation. Specifically, BrepGiff extracts and integrates topological and geometric features to construct a 3D graph where nodes correspond to face centroids in 3D space, preserving adjacency and degree information. Geometric features are derived by sampling points in the UV domain and extracting face and edge features. BrepGiff then applies Graph Attention Network (GAT) to enforce topological constraints from local to global during the degree-guided diffusion process. With the 3D graph representation and diffusion process, BrepGiff significantly reduces the computational cost and improves the quality, thus achieving lightweight generation of complex models. Experiments show that BrepGiff can generate complex B-rep models (>100 faces) using only 2 RTX4090 GPUs, achieving state-of-the-art performance in B-rep generation. Xiaoshui Huang, Jiacheng Hao, Yunpeng Bai, Hongping Gan, Yilei Shi |
CVPR | 2 |
| 2025 | An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion ModelsabstractExisting conditional Denoising Diffusion Probabilistic Models (DDPMs) with a Noise-Conditional Framework (NCF) remain challenging for 3D scene understanding tasks, as the complex geometric details in scenes increase the difficulty of fitting the gradients of the data distribution (the scores) from semantic labels. This also results in longer training and inference time for DDPMs compared to non-DDPMs. From a different perspective, we delve deeply into the model paradigm dominated by the Conditional Network. In this paper, we propose an end-to-end robust semantic Segmentation Network based on a Conditional-Noise Framework (CNF) of DDPMs, named CDSegNet. Specifically, CDSegNet models the Noise Network (NN) as a learnable noise-feature generator. This enables the Conditional Network (CN) to understand 3D scene semantics under multi-level feature perturbations, enhancing the generalization in unseen scenes. Meanwhile, benefiting from the noise system of DDPMs, CDSegNet exhibits strong robustness for data noise and sparsity in experiments. Moreover, thanks to CNF, CDSegNet can generate the semantic labels in a single-step inference like non-DDPMs, due to avoiding directly fitting the scores from semantic labels in the dominant network of CDSegNet. On public indoor and outdoor benchmarks, CDSegNet significantly outperforms existing methods, achieving state-of-the-art performance. Wentao Qu, Jing Wang 0201, Yongshun Gong, Xiaoshui Huang, Liang Xiao 0001 |
CVPR | 4 |
| 2025 | Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationabstractTalking head video generation aims to generate a realistic talking head video that preserves the person’s identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a challenging and critical problem to generate videos with accurate poses and fine-grained facial details simultaneously. Essentially, facial motion is often highly complex to model precisely, and the one-shot source face image cannot provide sufficient appearance guidance during generation due to dynamic pose changes. To tackle the problem, we propose to jointly learn motion and appearance codebooks and perform multi-scale codebook compensation to effectively refine both the facial motion conditions and appearance features for talking face image decoding. Specifically, the designed multi-scale motion and appearance codebooks are learned simultaneously in a unified framework to store representative global facial motion flow and appearance patterns. Then, we present a novel multi-scale motion and appearance compensation module, which utilizes a transformer-based codebook retrieval strategy to query complementary information from the two codebooks for joint motion and appearance compensation. The entire process produces motion flows of greater flexibility and appearance features with fewer distortions across different scales, resulting in a high-quality talking head video generation framework. Extensive experiments on various benchmarks validate the effectiveness of our approach and demonstrate superior generation results from both qualitative and quantitative perspectives when compared to state-of-the-art competitors. The project page is available at https://shaelynz.github.io/synergize-motion-appearance/. Shuling Zhao, Fa-Ting Hong, Xiaoshui Huang, Dan Xu 0002 |
CVPR | 3 |
| 2025 | MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence
Liyuan Deng, Yunpeng Bai, Yongkang Dai, Xiaoshui Huang, Hongping Gan, Dongshuo Huang, Jiacheng Hao, Yilei Shi |
ICCV | 4 |
| 2025 | Dynamic 3D Gaussian Reconstruction with Specular Reflectionabstract3D Gaussian Splatting (3DGS) has shown remarkable potential in novel view synthesis. However, it still encounters significant challenges in reconstructing dynamic scenes, particularly when dealing with reflective surfaces. To address this issue, we propose a novel 3DGS-based method for dynamic scene reconstruction with explicit reflection modeling. Our approach integrates deferred shading with a dual-environment map that combines static and dynamic components, enabling effective modeling of specular reflections. This allows our method to capture both steady and temporally varying lighting, resulting in more realistic renderings. We evaluate the proposed method on the benchmark dynamic reflection dataset, NERF-DS, and compare it with state-of-the-art approaches. Experimental results show that our method achieves superior or comparable performance in terms of PSNR, SSIM, and LPIPS metrics compared to competing approaches. Mingyang Zhao 0001, Yuanzhi Xu, Yifan Zuo 0001, Xiaoshui Huang, Yuming Fang 0001 |
ICIP | 4 |
| 2025 | THOR: Text to Human-Object Interaction Diffusion via Relation InterventionabstractThis paper addresses the challenging task of generating dynamic Human-Object Interactions from textual descriptions, named Text2HOI. While most existing works assume interactions with limited body parts or static objects, our task involves addressing the variation in human motion, the diversity of object shapes, and the semantic vagueness of object motion simultaneously. To tackle this, we propose a novel Text-guided Human-Object Interaction diffusion model with Relation Intervention (THOR). THOR is a cohesive diffusion model equipped with a relation intervention mechanism. In each diffusion step, we initiate text-guided human and object motion and then leverage human-object relations to intervene in object motion. This intervention enhances the spatial-temporal relations between humans and objects, with human-centric motion providing additional guidance for synthesizing consistent motion from text. To achieve more reasonable and realistic results, relation intervention loss is introduced at different levels of motion granularity. Qianyang Wu, Ye Shi 0001, Xiaoshui Huang, Lan Xu 0003, Jingyi Yu 0001, Jingya Wang 0001 |
ICME | 3 |
| 2025 | BRepFormer: Transformer-Based B-rep Geometric Feature Recognition
Yongkang Dai, Xiaoshui Huang, Yunpeng Bai, Hongping Gan, Yilei Shi |
ICMR | 2 |
| 2025 | LLaMA-Berry: Pairwise Optimization for Olympiad-level Mathematical Reasoning via O1-like Monte Carlo Tree SearchabstractDi Zhang, Jianbo Wu, Jingdi Lei, Tong Che, Jiatong Li, Tong Xie, Xiaoshui Huang, Shufei Zhang, Marco Pavone, Yuqiang Li, Wanli Ouyang, Dongzhan Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Di Zhang 0026, Jingdi Lei, Tong Che, Jiatong Li 0003, Tong Xie, Xiaoshui Huang, Shufei Zhang, Marco Pavone 0001, Wanli Ouyang, Dongzhan Zhou |
NAACL (Long Papers) | 7 |
| 2025 | Controllable text-to-3D multi-object generation via integrating layout and multiview patterns
Shaorong Sun, Shuchao Pang, Yazhou Yao, Xiaoshui Huang |
Comput. Graph. | 4 |
| 2025 | Fine-grained visual tracking via distribution-aware mask modeling and temporal propagation
Junjie Zhang 0002, Hongwen Yu, Fangyu Wu 0001, Xiaoshui Huang, Jian Zhang 0002 |
Knowl. Based Syst. | 5 |
| 2025 | 3DBench: A scalable benchmark for object and scene-level instruction-tuning of 3D large language models
Tianci Hu, Junjie Zhang 0002, Yutao Rao, Dan Zeng 0001, Hongwen Yu, Xiaoshui Huang |
Neural Networks | 6 |
| 2025 | Diverse Teacher-Students for deep safe semi-supervised learning under class mismatch
Qikai Wang, Rundong He, Yongshun Gong, Chunxiao Ren, Haoliang Sun, Xiaoshui Huang, Yilong Yin |
Neural Networks | 6 |
| 2025 | NeRF-Det++: Incorporating Semantic Cues and Perspective-Aware Depth Supervision for Indoor Multi-View 3D DetectionabstractNeRF-Det has achieved impressive performance in indoor multi-view 3D detection by innovatively utilizing NeRF to enhance representation learning. Despite its notable performance, we uncover three decisive shortcomings in its current design, including semantic ambiguity, inappropriate sampling, and insufficient utilization of depth supervision. To combat the aforementioned problems, we present three corresponding solutions: 1) Semantic Enhancement. We project the freely available 3D segmentation annotations onto the 2D plane and leverage the corresponding 2D semantic maps as the supervision signal, significantly enhancing the semantic awareness of multi-view detectors. 2) Perspective-Aware Sampling. Instead of employing the uniform sampling strategy, we put forward the perspective-aware sampling policy that samples densely near the camera while sparsely in the distance, more effectively collecting the valuable geometric clues. 3) Ordinal Residual Depth Supervision. As opposed to directly regressing the depth values that are difficult to optimize, we divide the depth range of each scene into a fixed number of ordinal bins and reformulate the depth prediction as the combination of the classification of depth bins as well as the regression of the residual depth values, thereby benefiting the depth learning process. The resulting algorithm, NeRF-Det++, has exhibited appealing performance in the ScanNetV2 and ARKITScenes datasets. Notably, in ScanNetV2, NeRF-Det++ outperforms the competitive NeRF-Det by +1.9% in mAP $\text{@}0.25$ and +3.5% in mAP $\text{@}0.50$ . The code will be publicly available at https://github.com/mrsempress/NeRF-Detplusplus. Chenxi Huang 0004, Yuenan Hou, Weicai Ye, Xiaoshui Huang, Binbin Lin 0001, Deng Cai 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Weakly Supervised LiDAR Semantic Segmentation via Scatter Image AnnotationabstractWeakly supervised LiDAR semantic segmentation has made significant strides with limited labeled data. However, most existing methods focus on the network training under weak supervision, while efficient annotation strategies remain largely unexplored. To tackle this gap, we implement LiDAR semantic segmentation using scatter image annotation, effectively integrating an efficient annotation strategy with network training. Specifically, we propose employing scatter images to annotate LiDAR point clouds, combining a pre-trained optical flow estimation network with a foundational image segmentation model to rapidly propagate manual annotations into dense labels for both images and point clouds. Moreover, we propose ScatterNet, a network that includes three pivotal strategies to reduce the performance gap caused by such annotations. First, it utilizes dense semantic labels as supervision for the image branch, alleviating the modality imbalance between point clouds and images. Second, an intermediate fusion branch is proposed to obtain multimodal texture and structural features. Finally, a perception consistency loss is introduced to determine which information needs to be fused and which needs to be discarded during the fusion process. Extensive experiments on the nuScenes and SemanticKITTI datasets demonstrate that our method requires less than 0.02% of the labeled points to achieve over 95% of the performance of fully-supervised methods. Notably, our labeled points are only 5% of those used in the most advanced weakly supervised methods. Zongyi Xu, Xiaoshui Huang, Shanshan Zhao 0001, Xinqi Jiang, Xinbo Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Frozen CLIP Transformer Is an Efficient Point Cloud EncoderabstractThe pretrain-finetune paradigm has achieved great success in NLP and 2D image fields because of the high-quality representation ability and transferability of their pretrained models. However, pretraining such a strong model is difficult in the 3D point cloud field due to the limited amount of point cloud sequences. This paper introduces Efficient Point Cloud Learning (EPCL), an effective and efficient point cloud learner for directly training high-quality point cloud models with a frozen CLIP transformer. Our EPCL connects the 2D and 3D modalities by semantically aligning the image features and point cloud features without paired 2D-3D data. Specifically, the input point cloud is divided into a series of local patches, which are converted to token embeddings by the designed point cloud tokenizer. These token embeddings are concatenated with a task token and fed into the frozen CLIP transformer to learn point cloud representation. The intuition is that the proposed point cloud tokenizer projects the input point cloud into a unified token space that is similar to the 2D images. Comprehensive experiments on 3D detection, semantic segmentation, classification and few-shot learning demonstrate that the CLIP transformer can serve as an efficient point cloud encoder and our method achieves promising performance on both indoor and outdoor benchmarks. In particular, performance gains brought by our EPCL are 19.7 AP50 on ScanNet V2 detection, 4.4 mIoU on S3DIS segmentation and 1.2 mIoU on SemanticKITTI segmentation compared to contemporary pretrained models. Code is available at \url{https://github.com/XiaoshuiHuang/EPCL}. Xiaoshui Huang, Zhou Huang 0006, Sheng Li 0020, Wentao Qu, Tong He 0001, Yuenan Hou, Yifan Zuo 0001, Wanli Ouyang |
AAAI | 1 |
| 2024 | Semi-supervised 3D Object Detection with PatchTeacher and PillarMixabstractSemi-supervised learning aims to leverage numerous unlabeled data to improve the model performance. Current semi-supervised 3D object detection methods typically use a teacher to generate pseudo labels for a student, and the quality of the pseudo labels is essential for the final performance. In this paper, we propose PatchTeacher, which focuses on partial scene 3D object detection to provide high-quality pseudo labels for the student. Specifically, we divide a complete scene into a series of patches and feed them to our PatchTeacher sequentially. PatchTeacher leverages the low memory consumption advantage of partial scene detection to process point clouds with a high-resolution voxelization, which can minimize the information loss of quantization and extract more fine-grained features. However, it is non-trivial to train a detector on fractions of the scene. Therefore, we introduce three key techniques, i.e., Patch Normalizer, Quadrant Align, and Fovea Selection, to improve the performance of PatchTeacher. Moreover, we devise PillarMix, a strong data augmentation strategy that mixes truncated pillars from different LiDAR scans to generate diverse training samples and thus help the model learn more general representation. Extensive experiments conducted on Waymo and ONCE datasets verify the effectiveness and superiority of our method and we achieve new state-of-the-art results, surpassing existing methods by a large margin. Codes are available at https://github.com/LittlePey/PTPM. Xiaopei Wu, Liang Xie 0003, Yuenan Hou, Binbin Lin 0001, Xiaoshui Huang, Haifeng Liu 0001, Deng Cai 0001, Wanli Ouyang |
AAAI | 6 |
| 2024 | A Conditional Denoising Diffusion Probabilistic Model for Point Cloud UpsamplingabstractPoint cloud upsampling (PCU) enriches the representation of raw point clouds, significantly improving the performance in downstream tasks such as classification and reconstruction. Most of the existing point cloud upsampling methods focus on sparse point cloud feature extraction and upsampling module design. In a different way, we dive deeper into directly modelling the gradient of data distribution from dense point clouds. In this paper, we proposed a conditional denoising diffusion probabilistic model (DDPM) for point cloud upsampling, called PUDM. Specifically, PUDM treats the sparse point cloud as a condition, and iteratively learns the transformation relationship between the dense point cloud and the noise. Simultaneously, PUDM aligns with a dual mapping paradigm to further improve the discernment of point features. In this context, PUDM enables learning complex geometry details in the ground truth through the dominant features, while avoiding an additional upsampling module design. Furthermore, to generate high-quality arbitrary-scale point clouds during inference, PUDM exploits the prior knowledge of the scale between sparse point clouds and dense point clouds during training by parameterizing a rate factor. Moreover, PUDM exhibits strong noise robustness in experimental results. In the quantitative and qualitative evaluations on PU1K and PUGAN, PUDM significantly outperformed existing methods in terms of Chamfer Distance (CD) and Hausdorff Distance (HD), achieving state of the art (SOTA) performance. Wentao Qu, Yuantian Shao, Lingwu Meng, Xiaoshui Huang, Liang Xiao 0001 |
CVPR | 4 |
| 2024 | TASeg: Temporal Aggregation Network for LiDAR Semantic SegmentationabstractTraining deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the spar-sity problem as it makes the input signal denser. However, previous multi-frame fusion algorithms fall short in utilizing sufficient temporal information due to the memory constraint, and they also ignore the informative temporal images. To fully exploit rich information hidden in long-term temporal point clouds and images, we present the Temporal Aggre-gation Network, termed TASeg. Specifically, we propose a Temporal LiDAR Aggregation and Distillation (TLAD) algorithm, which leverages historical priors to assign dif-ferent aggregation steps for different classes. It can largely reduce memory and time overhead while achieving higher accuracy. Besides, TLAD trains a teacher injected with gt priors to distill the model, further boosting the performance. To make full use of temporal images, we design a Temporal Image Aggregation and Fusion (TIAF) module, which can greatly expand the camera FOVand enhance the present features. Temporal LiDAR points in the camera FOV are used as mediums to transform temporal image features to the present coordinate for temporal multi-modal fusion. Moreover, we develop a Static-Moving Switch Augmentation (SMSA) algorithm, which utilizes sufficient temporal information to enable objects to switch their motion states freely, thus greatly increasing static and moving training samples. Our TASeg ranks 1st††the date of CVPR deadline, i.e., 2023-11-18 07:59 AM UTC. on three challenging tracks, i.e., SemanticKITTI single-scan track, multi-scan track and nuScenes LiDAR segmentation track, strongly demonstrating the superiority of our method. Codes are available at https://github.com/LittlePey/TASeg. Xiaopei Wu, Yuenan Hou, Xiaoshui Huang, Binbin Lin 0001, Tong He 0001, Xinge Zhu, Yuexin Ma, Boxi Wu 0001, Haifeng Liu 0001, Deng Cai 0001, Wanli Ouyang |
CVPR | 3 |
| 2024 | Taming Stable Diffusion for Text to 360° Panorama Image GenerationabstractGenerative models, e.g., Stable Diffusion, have enabled the creation of photorealistic images from text prompts. Yet, the generation of 360-degree panorama images from text remains a challenge, particularly due to the dearth of paired text-panorama data and the domain gap between panorama and perspective images. In this paper, we introduce a novel dual-branch diffusion model named PanFusion to generate a 360-degree image from a text prompt. We leverage the stable diffusion model as one branch to provide prior knowledge in natural image generation and register it to another panorama branch for holistic image generation. We propose a unique cross-attention mechanism with projection awareness to minimize distortion during the collaborative denoising process. Our experiments validate that PanFusion surpasses existing methods and, thanks to its dual-branch structure, can integrate additional constraints like room layout for customized panorama outputs. Qianyi Wu, Camilo Cruz Gambardella, Xiaoshui Huang, Dinh Q. Phung, Wanli Ouyang, Jianfei Cai 0001 |
CVPR | 4 |
| 2024 | Point Cloud Pre-Training with Diffusion ModelsabstractPre-training a model and then fine-tuning it on down-stream tasks has demonstrated significant success in the 2D image and NLP domains. However, due to the unordered and non-uniform density characteristics of point clouds, it is non-trivial to explore the prior knowledge of point clouds and pre-train a point cloud backbone. In this paper, we propose a novel pre-training method called Point cloud Diffusion pre-training (PointDif). We consider the point cloud pre-training task as a conditional point-to-point gen-eration problem and introduce a conditional point genera-tor. This generator aggregates the features extracted by the backbone and employs them as the condition to guide the point-to-point recovery from the noisy point cloud, thereby assisting the backbone in capturing both local and global geometric priors as well as the global point density distri-bution of the object. We also present a recurrent uniform sampling optimization strategy, which enables the model to uniformly recover from various noise levels and learn from balanced supervision. Our PointDif achieves substan-tial improvement across various real-world datasets for di-verse downstream tasks such as classification, segmentation and detection. Specifically, PointDif attains 70.0% mIoU on S3DIS Area 5 for the segmentation task and achieves an average improvement of 2.4% on ScanObjectNN for the classification task compared to TAP. Furthermore, our pre-training framework can be flexibly applied to diverse point cloud backbones and bring considerable gains. Code is available at https://github.com/zhengxiaozx/PointDif. Xiaoshui Huang, Guofeng Mei, Yuenan Hou, Zhaoyang Lyu, Bo Dai 0002, Wanli Ouyang, Yongshun Gong |
CVPR | 2 |
| 2024 | GVGEN: Text-to-3D Generation with Volumetric Representation
Xianglong He, Sida Peng, Yangguang Li 0001, Xiaoshui Huang, Chun Yuan 0003, Wanli Ouyang, Tong He 0001 |
ECCV (8) | 6 |
| 2024 | UniDream: Unifying Diffusion Priors for Relightable Text-to-3D Generation
Zexiang Liu, Yangguang Li 0001, Youtian Lin, Xin Yu 0004, Sida Peng, Yan-Pei Cao 0001, Xiaojuan Qi 0001, Xiaoshui Huang, Ding Liang, Wanli Ouyang |
ECCV (5) | 8 |
| 2024 | 3D Point Cloud Pre-Training with Knowledge Distilled from 2D ImagesabstractThe success of pre-trained 2D vision models can largely be attributed to their ability to learn from large-scale datasets. However, compared with 2D image datasets, current pre-training data for 3D point clouds are limited. In this paper, we propose a knowledge distillation method for pre-training 3D point cloud models by directly acquiring knowledge from a 2D representation learning model, specifically the image encoder of CLIP. To close the significant domain gap between 2D images and 3D point clouds, we propose to align the features from the two domains at the concept level. Our method utilizes a cross-attention mechanism to extract concept features from 3D point clouds and compares them with the corresponding information from 2D images. This approach bridges the two domains and allows point cloud models to learn directly from the rich information contained in 2D teacher models. Extensive experiments show that our proposed knowledge distillation scheme achieves higher accuracy than the state-of-the-art 3D pre-training methods for synthetic and real-world datasets on various downstream tasks, including object classification, object detection, semantic segmentation and part segmentation. Yuanhan Zhang, Zhenfei Yin, Jiebo Luo 0001, Wanli Ouyang, Xiaoshui Huang |
ICME | 6 |
| 2024 | 3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset
Junjie Zhang 0002, Tianci Hu, Xiaoshui Huang, Yongshun Gong, Dan Zeng 0001 |
IJCAI | 3 |
| 2024 | DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionabstractDiffusion-based methods have achieved remarkable achievements in 2D image or 3D object generation, however, the generation of 3D scenes and even $360^{\circ}$ images remains constrained, due to the limited number of scene datasets, the complexity of 3D scenes themselves, and the difficulty of generating consistent multi-view images. To address these issues, we first establish a large-scale panoramic video-text dataset containing millions of consecutive panoramic keyframes with corresponding panoramic depths, camera poses, and text descriptions. Then, we propose a novel text-driven panoramic generation framework, termed DiffPano, to achieve scalable, consistent, and diverse panoramic scene generation. Specifically, benefiting from the powerful generative capabilities of stable diffusion, we fine-tune a single-view text-to-panorama diffusion model with LoRA on the established panoramic video-text dataset. We further design a spherical epipolar-aware multi-view diffusion model to ensure the multi-view consistency of the generated panoramic images. Extensive experiments demonstrate that DiffPano can generate scalable, consistent, and diverse panoramic images with given unseen text descriptions and camera poses. Weicai Ye, Chenhao Ji, Zheng Chen 0016, Junyao Gao 0002, Xiaoshui Huang, Song-Hai Zhang, Wanli Ouyang, Tong He 0001, Cairong Zhao, Guofeng Zhang 0001 |
NeurIPS | 5 |
| 2024 | Retrieval-and-alignment based large-scale indoor point cloud semantic segmentation
Zongyi Xu, Xiaoshui Huang, Yangfu Wang, Qianni Zhang, Weisheng Li 0001, Xinbo Gao 0001 |
Sci. China Inf. Sci. | 2 |
| 2024 | Learning content-aware feature fusion for guided depth map super-resolution
Yifan Zuo 0001, Xiaoshui Huang, Xue Xia 0005, Yuming Fang 0001 |
Signal Process. Image Commun. | 5 |
| 2024 | A2 GSTran: Depth Map Super-Resolution via Asymmetric Attention With Guidance SelectionabstractCurrently, Convolutional Neural Network (CNN) has dominated guided depth map super-resolution (SR). However, the inefficient receptive field growing and input-independent convolution limit the generalization of CNN. Motivated by vision transformer, this paper proposes an efficient transformer-based backbone A2GSTran for guided depth map SR, which resolves the above intrinsic defect of CNN. In addition, state-of-the-art (SOTA) models only refine depth features with the guidance which is implicitly selected without supervision. So, there is no explicit guarantee to mitigate the artifacts of texture copying and edge blurring. Accordingly, the proposed A2GSTran simultaneously solves two sub-problems,i.e., guided monocular depth estimation and guided depth SR, in separate branches. Specifically, the explicit supervision upon monocular depth estimation lifts the efficiency of guidance selection. The feature fusion between branches is designed via bi-directional cross attention. Moreover, since guidance domain is defined in high resolution (HR), we propose asymmetric cross attention to maintain the guidance information via pixel unshuffle instead of pooling which has unequal channel number to depth features. Based on the supervisions to depth reconstruction and guidance selection, the final depth features are refined by fusing the output features of the corresponding branches via channel attention to generate the HR depth map. Sufficient experimental results on synthetic and real datasets for multiple scales validate our contributions compared with SOTA models. The code and models are public via https://github.com/alex-cate/Depth_Map_Super-resolution_via_Asymmetric_Attention_with_Guidance_Selection Yifan Zuo 0001, Yifeng Zeng, Yuming Fang 0001, Xiaoshui Huang, Jiebin Yan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Frequency-Aware Multi-Modal Fine-Tuning for Few-Shot Open-Set Remote Sensing Scene ClassificationabstractFew-shot open-set recognition, as a new paradigm, leveraging a limited amount of supervised data to identify specific Remote Sensing (RS) scene categories and generalize to novel ones. However, the data bias induced by the small sample size not only causes severe overfitting within base classes, but also impairs the capacity for inference to identify RS scenes in hitherto unobserved categories. Furthermore, owing to environmental influences, RS images frequently manifest notable intra-class disparities and comparatively low inter-class distinctions, intensifying the challenge in obtaining suitable classifiers. To address above issues, we investigate the utilization of a Multi-modal Foundational Model (MFM) infused with essential domain knowledge to mitigate the generalization limitations encountered in few-shot scenarios. Recognizing that existing MFMs with a visual-text dual-branch structure are primarily tailored for natural scenes, we propose a custom Frequency Distribution-based Multi-modal Fine-Tuning strategy (FreqDiMFT) in a parameter-efficient manner. More specifically, within the vision branch, we address the high inter-class similarity and intra-class diversity in RS images by embedding the local-global frequency distribution information to facilitate the recognition of RS scenes. To further amplify the model's generalization ability post transfer, we introduce an adaptive feature refinement module designed for Transformers, proficient in filtering redundant features resulting from domain disparities. To mitigate the domain drift on the textual branch, we adopt an input format that combines basic templates with domain expertise from RS end to generate more discriminative class prototypes. To fully verify the effectiveness of our FreqDiMFT in a more practical setting, we collect a Large-Scale hybrid dataset (LSRS). Extensive experiments demonstrate that, even with a scant number of training samples, our strategy yields advanced performances compared to state-of-the-art models. Junjie Zhang 0002, Yutao Rao, Xiaoshui Huang, Guanyi Li, Dan Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Accurate Registration of Cross-Modality Geometry via Consistent ClusteringabstractThe registration of unitary-modality geometric data has been successfully explored over past decades. However, existing approaches typically struggle to handle cross-modality data due to the intrinsic difference between different models. To address this problem, in this article, we formulate the cross-modality registration problem as a consistent clustering process. First, we study the structure similarity between different modalities based on an adaptive fuzzy shape clustering, from which a coarse alignment is successfully operated. Then, we optimize the result using fuzzy clustering consistently, in which the source and target models are formulated as clustering memberships and centroids, respectively. This optimization casts new insight into point set registration, and substantially improves the robustness against outliers. Additionally, we investigate the effect of fuzzier in fuzzy clustering on the cross-modality registration problem, from which we theoretically prove that the classical Iterative Closest Point (ICP) algorithm is a special case of our newly defined objective function. Comprehensive experiments and analysis are conducted on both synthetic and real-world cross-modality datasets. Qualitative and quantitative results demonstrate that our method outperforms state-of-the-art approaches with higher accuracy and robustness. Our code is publicly available at https://github.com/zikai1/CrossModReg. Mingyang Zhao 0001, Xiaoshui Huang, Jingen Jiang 0001, Luntian Mou, Dong-Ming Yan 0001, Lei Ma 0008 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Unsupervised Deep Probabilistic Approach for Partial Point Cloud RegistrationabstractDeep point cloud registration methods face challenges to partial overlaps and rely on labeled data. To address these issues, we propose UDPReg, an unsupervised deep probabilistic registration framework for point clouds with partial overlaps. Specifically, we first adopt a network to learn posterior probability distributions of Gaussian mixture models (GMMs) from point clouds. To handle partial point cloud registration, we apply the Sinkhorn algorithm to predict the distribution-level correspondences under the constraint of the mixing weights of GMMs. To enable unsupervised learning, we design three distribution consistency-based losses: self-consistency, cross-consistency, and local contrastive. The self-consistency loss is formulated by encouraging GMMs in Euclidean and feature spaces to share identical posterior distributions. The cross-consistency loss derives from the fact that the points of two partially overlapping point clouds belonging to the same clusters share the cluster centroids. The cross-consistency loss allows the network to flexibly learn a transformation-invariant posterior distribution of two aligned point clouds. The local contrastive loss facilitates the network to extract discriminative local features. Our UDPReg achieves competitive performance on the 3DMatch/3DLoMatch and ModelNet/ModelLoNet benchmarks. Guofeng Mei, Hao Tang 0005, Xiaoshui Huang, Weijie Wang 0002, Juan Liu 0006, Jian Zhang 0002, Luc Van Gool, Qiang Wu 0001 |
CVPR | 3 |
| 2023 | CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-TrainingabstractPre-training across 3D vision and language remains under development because of limited training data. Recent works attempt to transfer vision-language (V-L) pre-training methods to 3D vision. However, the domain gap between 3D and images is unsolved, so that V-L pre-trained models are restricted in 3D downstream tasks. To address this issue, we propose CLIP2Point, an image-depth pre-training method by contrastive learning to transfer CLIP to the 3D domain, and adapt it to point cloud classification. We introduce a new depth rendering setting that forms a better visual effect, and then render 52,460 pairs of images and depth maps from ShapeNet for pre-training. The pre-training scheme of CLIP2Point combines cross-modality learning to enforce the depth features for capturing expressive visual and textual features and intra-modality learning to enhance the invariance of depth aggregation. Additionally, we propose a novel Gated Dual-Path Adapter (GDPA), i.e., a dual-path structure with global-view aggregators and gated fusion for downstream representative learning. It allows the ensemble of CLIP and CLIP2Point, tuning pre-training knowledge to downstream tasks in an efficient adaptation. Experimental results show that CLIP2Point is effective in transferring CLIP knowledge to 3D vision. CLIP2Point outperforms other 3D transfer learning and pre-training networks, achieving state-of-the-art results on zero-shot, few-shot, and fully-supervised classification. Codes are available at: https://github.com/tyhuang0428/CLIP2Point. Bowen Dong 0001, Yunhan Yang, Xiaoshui Huang, Rynson W. H. Lau, Wanli Ouyang, Wangmeng Zuo |
ICCV | 4 |
| 2023 | Boosting 3D Point Cloud Registration by Transferring Multi-modality KnowledgeabstractThe recent multi-modality models have achieved great performance in many vision tasks because the extracted features contain the multi-modality knowledge. However, most of the current registration descriptors have only concentrated on local geometric structures. This paper proposes a method to boost point cloud registration accuracy by transferring the multi-modality knowledge of pre-trained multi-modality model to a new descriptor neural network. Different to the previous multi-modality methods that requires both modalities, the proposed method only requires point clouds during inference. Specifically, we propose an ensemble descriptor neural network combining pre-trained sparse convolution branch and a new point-based convolution branch. By fine-tuning on a single modality data, the proposed method achieves new state-of-the-art results on 3DMatch and competitive accuracy on 3DLoMatch and KITTI. The code and the trained model will be released at https://github.com/phdymz/DBENet.git. Mingzhi Yuan, Xiaoshui Huang, Kexue Fu 0001, Manning Wang |
ICRA | 2 |
| 2023 | LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and BenchmarkabstractLarge language models have emerged as a promising approach towards achieving general-purpose AI agents. The thriving open-source LLM community has greatly accelerated the development of agents that support human-machine dialogue interaction through natural language processing. However, human interaction with the world extends beyond only text as a modality, and other modalities such as vision are also crucial. Recent works on multi-modal large language models, such as GPT-4V and Bard, have demonstrated their effectiveness in handling visual modalities. However, the transparency of these works is limited and insufficient to support academic research. To the best of our knowledge, we present one of the very first open-source endeavors in the field, LAMM, encompassing a Language-Assisted Multi-Modal instruction tuning dataset, framework, and benchmark. Our aim is to establish LAMM as a growing ecosystem for training and evaluating MLLMs, with a specific focus on facilitating AI agents capable of bridging the gap between ideas and execution, thereby enabling seamless human-AI interaction. Our main contribution is three-fold: 1) We present a comprehensive dataset and benchmark, which cover a wide range of vision tasks for 2D and 3D vision. Extensive experiments validate the effectiveness of our dataset and benchmark. 2) We outline the detailed methodology of constructing multi-modal instruction tuning datasets and benchmarks for MLLMs, enabling rapid scaling and extension of MLLM research to diverse domains, tasks, and modalities. 3) We provide a primary but potential MLLM training framework optimized for modality extension. We also provide baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained within 24 A100 GPU hours, framework supports training with V100 and RTX3090 is available thanks to the open-source society. Codes and data are now available at https://openlamm.github.io. Zhenfei Yin, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Xiaoshui Huang, Zhiyong Wang 0001, Lu Sheng, Lei Bai 0001, Wanli Ouyang |
NeurIPS | 7 |
| 2023 | Cross-source point cloud registration: Challenges, progress and prospects
Xiaoshui Huang, Guofeng Mei, Jian Zhang 0002 |
Neurocomputing | 1 |
| 2023 | Beyond CNNs: Exploiting Further Inherent Symmetries in Medical Image SegmentationabstractAutomatic tumor or lesion segmentation is a crucial step in medical image analysis for computer-aided diagnosis. Although the existing methods based on convolutional neural networks (CNNs) have achieved the state-of-the-art performance, many challenges still remain in medical tumor segmentation. This is because, although the human visual system can detect symmetries in 2-D images effectively, regular CNNs can only exploit translation invariance, overlooking further inherent symmetries existing in medical images, such as rotations and reflections. To solve this problem, we propose a novel group equivariant segmentation framework by encoding those inherent symmetries for learning more precise representations. First, kernel-based equivariant operations are devised on each orientation, which allows it to effectively address the gaps of learning symmetries in existing approaches. Then, to keep segmentation networks globally equivariant, we design distinctive group layers with layer-wise symmetry constraints. Finally, based on our novel framework, extensive experiments conducted on real-world clinical data demonstrate that a group equivariant Res-UNet (called GER-UNet) outperforms its regular CNN-based counterpart and the state-of-the-art segmentation methods in the tasks of hepatic tumor segmentation, COVID-19 lung infection segmentation, and retinal vessel detection. More importantly, the newly built GER-UNet also shows potential in reducing the sample complexity and the redundancy of filters, upgrading current segmentation CNNs, and delineating organs on other medical imaging modalities. Shuchao Pang, Anan Du, Mehmet A. Orgun, Yan Wang 0002, Quan Z. Sheng, Shoujin Wang, Xiaoshui Huang, Zhenmei Yu |
IEEE Trans. Cybern. | 7 |
| 2022 | Unsupervised Point Cloud Pre-Training Via Contrasting and ClusteringabstractThe annotation for large-scale point clouds is still time-consuming and unavailable for many complex real-world tasks. Point cloud pre-training is a promising direction to auto-extract features without labeled data. Therefore, this paper proposes a general unsupervised approach, named ConClu for point cloud pre-training by jointly performing contrasting and clustering. Specifically, the contrasting is formulated by maximizing the similarity feature vectors produced by encoders fed with two augmentations of the same point cloud. The clustering simultaneously clusters the data while enforcing consistency between cluster assignments produced different augmentations. Experimental evaluations on downstream applications outperform state-of-the-art techniques, which demonstrates the effectiveness of our framework. Guofeng Mei, Xiaoshui Huang, Juan Liu 0006, Jian Zhang 0002, Qiang Wu 0001 |
ICIP | 2 |
| 2022 | Partial Point Cloud Registration Via Soft SegmentationabstractMost existing correspondence-free registration methods suffer from performance degradation in partial overlapped point clouds. To solve the partial overlapped point cloud registration, this paper proposes, SegReg, a soft Segmentation-based correspondence-free Registration approach. Specifically, we first softly segment both source and target point clouds into a discrete number of geometric partitions, respectively. Then registration is achieved through iteratively using the IC-LK algorithm to minimize the distance between the feature descriptors of the corresponded partitions. Extensive experiments on synthetic synthetic dataset ModelNet40 and real dataset 7Scene show that the proposed method achieves state-of-the-art performance. Guofeng Mei, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001 |
ICIP | 2 |
| 2022 | Overlap-Guided Coarse-to-Fine Correspondence Prediction for Point Cloud RegistrationabstractEstablishing reliable correspondences between a pair of point clouds is essential for registration with partial overlaps. However, existing correspondence estimation works usually struggle to distinguish the points in overlap and non-overlap regions. This paper thus proposes an Overlap-guided Coarse-to-Fine Network, named OCFNet, which first establishes correspondences at a coarse level and then refines them at a point level. Specifically, at the coarse level, our model first aggregates two point clouds into smaller sets of super-points with associated features and overlap scores, followed by establishing coarse-level correspondences between the two sets of super-points under the guidance of overlap scores. On the fine stage, a decoder recovers the raw points while jointly learning the associated features and overlap scores. Coarse-level proposals are then expanded to patches, and point-level correspondences are sequentially refined from the corresponding patches. We conducted comprehensive experiments on 3DMatch, 3DLoMatch, and KITTI benchmarks to show the effectiveness of the proposed method. [code] Guofeng Mei, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001 |
ICME | 2 |
| 2022 | Unsupervised Pre-training for 3D Object Detection with Transformer
Maosheng Sun, Xiaoshui Huang, Zeren Sun, Qiong Wang 0003, Yazhou Yao |
PRCV (3) | 2 |
| 2022 | Robust real-world point cloud registration by inlier detection
Xiaoshui Huang, Yangfu Wang, Sheng Li 0020, Guofeng Mei, Zongyi Xu, Yucheng Wang 0003, Jian Zhang 0002, Mohammed Bennamoun |
Comput. Vis. Image Underst. | 1 |
| 2022 | MIG-Net: Multi-Scale Network Alternatively Guided by Intensity and Gradient Features for Depth Map Super-ResolutionabstractThe studies of previous decades have shown that the quality of depth maps can be significantly lifted by introducing the guidance from intensity images describing the same scenes. With the rising of deep convolutional neural network, the performance of guided depth map super-resolution is further improved. The variants always consider deep structure, optimized gradient flow and feature reusing. Nevertheless, it is difficult to obtain sufficient and appropriate guidance from intensity features without any prior. In fact, features in the gradient domain, e.g., edges, present strong correlations between the intensity image and the corresponding depth map. Therefore, the guidance in the gradient domain can be more efficiently explored. In this paper, the depth features are iteratively upsampled by 2×. In each upsampling stage, the low-quality depth features and the corresponding gradient features are iteratively refined by the guidance from the intensity features via two parallel streams. Then, to make full use of depth features in the image and gradient domains, the depth features and gradient features are alternatively complemented with each other. Compared with state-of-the-art counterparts, the sufficient experimental results show improvements according to the objective and subjective assessments. The code is available athttps://github.com/Yifan-Zuo/MIG-net-gradient_guided_depth_enhancement. Yifan Zuo 0001, Yuming Fang 0001, Xiaoshui Huang, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | DeepMMSA: A Novel Multimodal Deep Learning Method for Non-small Cell Lung Cancer Survival AnalysisabstractLung cancer is the leading cause of cancer death worldwide. The critical reason for the deaths is delayed diagnosis and poor prognosis. With the accelerated development of deep learning techniques, it has been successfully applied extensively in many real-world applications, including health sectors such as medical image interpretation and disease diagnosis. By combining more modalities that being engaged in the processing of information, multimodal learning can extract better features and improve the predictive ability. The conventional methods for lung cancer survival analysis normally utilize clinical data and only provide a statistical probability. To improve the survival prediction accuracy and help prognostic decision-making in clinical practice for medical experts, we for the first time propose a multimodal deep learning framework for non-small cell lung cancer (NSCLC) survival analysis, named DeepMMSA. This framework leverages CT images in combination with clinical data, enabling the abundant information held within medical images to be associate with lung cancer survival information. We validate our model on the data of 422 NSCLC patients from The Cancer Imaging Archive (TCIA). Experimental results support our hypothesis that there is an underlying relationship between prognostic information and radiomic images. Besides, quantitative results show that our method could surpass the state-of-the-art methods by 4% on concordance. Yujiao Wu, Xiaoshui Huang, Sai-Ho Ling, Steven W. Su |
SMC | 3 |
| 2020 | Feature-Metric Registration: A Fast Semi-Supervised Approach for Robust Point Cloud Registration Without CorrespondencesabstractWe present a fast feature-metric point cloud registration framework, which enforces the optimisation of registration by minimising a feature-metric projection error without correspondences. The advantage of the feature-metric projection error is robust to noise, outliers and density difference in contrast to the geometric projection error. Besides, minimising the feature-metric projection error does not need to search the correspondences so that the optimisation speed is fast. The principle behind the proposed method is that the feature difference is smallest if point clouds are aligned very well. We train the proposed method in a semi-supervised or unsupervised approach, which requires limited or no registration label data. Experiments demonstrate our method obtains higher accuracy and robustness than the state-of-the-art methods. Besides, experimental results show that the proposed method can handle significant noise and density difference, and solve both same-source and cross-source point cloud registration. Xiaoshui Huang, Guofeng Mei, Jian Zhang 0002 |
CVPR | 1 |
| 2020 | Exploring Long-Short-Term Context For Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation attracts numerous attention following the success of the point-based convolution neural network. Due to the ambiguity of the point-based feature, many methods study on integrating contextual information to solve the ambiguous problem. However, the extracted context is severely limited to the small input blocks. Few prior works exploit contextual information beyond the blocks to capture long-range dependencies. To address this limitation, we propose a novel long-short-term context framework, which adopts a long-short-term feature bank to exploit both the local context within each block and the long-range context beyond the current task block. The proposed framework is flexible and easy to be combined with existing models, thereby enables existing models to capture the larger range context. Extensive experiments demonstrate that the proposed model achieves improved segmentation performance, and augmenting existing models with a long-short-term feature bank consistently increases the performance. Anan Du, Shuchao Pang, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001 |
ICIP | 3 |
| 2020 | Classification Constrained Discriminator For Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation for semantic segmentation aims to transfer knowledge from label-rich synthetic datasets to real-world images without any annotation. The traditional adversarial learning methods for domain adaptation learn to extract domain-invariant feature representations by aligning the feature distributions of both domains. However, these methods suffer from an imbalance in adversarial training and feature distortion. In this work, we propose a classification constrained discriminator to alleviate these problems. Specifically, we first propose to balance the adversarial training by eliminating any pooling layers or strided convolutions in the discriminator. Then, we propose to constrain the discriminator with an auxiliary classification loss to help the feature generator extract the domain-invariant features that are useful for segmentation rather than just ambiguous features to fool the domain discriminator. Extensive experiments demonstrate the superiority of our proposed approach. The source code and models have been made available at https://github.com/NUSTMachine-Intelligence-Laboratory/ccd. Tao Chen 0012, Jian Zhang 0002, Guosen Xie, Yazhou Yao, Xiaoshui Huang, Zhenmin Tang |
ICME | 5 |
| 2019 | Kpsnet: Keypoint Detection and Feature Extraction for Point Cloud RegistrationabstractThis paper presents the KPSNet, a KeyPoint Siamese Network to simultaneously learn task-desirable keypoint detector and feature extractor. The keypoint detector is optimized to predict a score vector, which signifies the probability of each candidate being a keypoint. The feature extractor is optimized to learn robust features of keypoints by exploiting the correspondence between the keypoints generated from two inputs, respectively. For training, the KPSNet does not require to manually annotate keypoints and local patches pairwise. Instead, we design an alignment module to establish the correspondence between the two inputs and generate positive and negative samples on-the-fly. Therefore, our method can be easily extended to new scenes. We test the proposed method on the open-source benchmark and experiments show the validity of our method. Anan Du, Xiaoshui Huang, Jian Zhang 0002, Lingxiang Yao, Qiang Wu 0001 |
ICIP | 2 |
| 2019 | Fast Registration for Cross-Source Point Clouds by using Weak Regional Affinity and Pixel-Wise RefinementabstractMany types of 3D acquisition sensors have emerged in recent years and point cloud has been widely used in many areas. Accurate and fast registration of cross-source 3D point clouds from different sensors is an emerged research problem in computer vision. This problem is extremely challenging because cross-source point clouds contain a mixture of various variances, such as density, partial overlap, large noise and outliers, viewpoint changing. In this paper, an algorithm is proposed to align cross-source point clouds with both high accuracy and high efficiency. There are two main contributions: firstly, two components, the weak region affinity and pixel-wise refinement, are proposed to maintain the global and local information of 3D point clouds. Then, these two components are integrated into an iterative tensor-based registration algorithm to solve the cross-source point cloud registration problem. We conduct experiments on a synthetic cross-source benchmark dataset and real cross-source datasets. Comparison with six state-of-the-art methods, the proposed method obtains both higher efficiency and accuracy. Xiaoshui Huang, Lixin Fan, Qiang Wu 0001, Jian Zhang 0002, Chun Yuan 0003 |
ICME | 1 |
| 2018 | Attention-Based Transactional Context Embedding for Next-Item RecommendationabstractTo recommend the next item to a user in a transactional context is practical yet challenging in applications such as marketing campaigns. Transactional context refers to the items that are observable in a transaction. Most existing transaction based recommender systems (TBRSs) make recommendations by mainly considering recently occurring items instead of all the ones observed in the current context. Moreover, they often assume a rigid order between items within a transaction, which is not always practical. More importantly, a long transaction often contains many items irreverent to the next choice, which tends to overwhelm the influence of a few truly relevant ones. Therefore, we posit that a good TBRS should not only consider all the observed items in the current transaction but also weight them with different relevance to build an attentive context that outputs the proper next item with a high probability. To this end, we design an effective attention based transaction embedding model (ATEM) for context embedding to weight each observed item in a transaction without assuming order. The empirical study on real-world transaction datasets proves that ATEM significantly outperforms the state-of-the-art methods in terms of both accuracy and novelty. Shoujin Wang, Liang Hu 0004, Longbing Cao, Xiaoshui Huang, Defu Lian, Wei Liu 0007 |
AAAI | 4 |
| 2018 | A Coarse-to-Fine Algorithm for Matching and Registration in 3D Cross-Source Point CloudsabstractWe propose an efficient method to deal with the matching and registration problem found in cross-source point clouds captured by different types of sensors. This task is especially challenging due to the presence of density variation, scale difference, a large proportion of noise and outliers, missing data, and viewpoint variation. The proposed method has two stages: in the coarse matching stage, we use the ensemble of shape functions descriptor to select potential K regions from the candidate point clouds for the target. In the fine stage, we propose a scale embedded generative Gaussian mixture models registration method to refine the results from the coarse matching stage. Following the fine stage, both the best region and accurate camera pose relationships between the candidates and target are found. We conduct experiments in which we apply the method to two applications: one is 3D object detection and localization in street-view outdoor (LiDAR/VSFM) cross-source point clouds and the other is 3D scene matching and registration in indoor (KinectFusion/VSFM) cross-source point clouds. The experiment results show that the proposed method performs well when compared with the existing methods. It also shows that the proposed method is robust under various sensing techniques, such as LiDAR, Kinect, and RGB camera. Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001, Lixin Fan, Chun Yuan 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | A Systematic Approach for Cross-Source Point Cloud Registration by Preserving Macro and Micro StructuresabstractWe propose a systematic approach for registering cross-source point clouds that come from different kinds of sensors. This task is especially challenging due to the presence of significant missing data, large variations in point density, scale difference, large proportion of noise, and outliers. The robustness of the method is attributed to the extraction of macro and micro structures. Macro structure is the overall structure that maintains similar geometric layout in cross-source point clouds. Micro structure is the element (e.g., local segment) being used to build the macro structure. We use graph to organize these structures and convert the registration into graph matching. With a novel proposed descriptor, we conduct the graph matching in a discriminative feature space. The graph matching problem is solved by an improved graph matching solution, which considers global geometrical constraints. Robust cross source registration results are obtained by incorporating graph matching outcome with RANSAC and ICP refinements. Compared with eight state-of-the-art registration algorithms, the proposed method invariably outperforms on Pisa Cathedral and other challenging cases. In order to compare quantitatively, we propose two challenging cross-source data sets and conduct comparative experiments on more than 27 cases, and the results show we obtain much better performance than other methods. The proposed method also shows high accuracy in same-source data sets. Xiaoshui Huang, Jian Zhang 0002, Lixin Fan, Qiang Wu 0001, Chun Yuan 0003 |
IEEE Trans. Image Process. | 1 |