Le Hui

dblp:211/6859 · DBLP profile ↗
← Back
51ranked-venue papers
10as first author
45since 2021 · last 2026
0000-0003-0851-6805ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 9 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 8 first-author · 30 since 2021Systems, architecture and hardware · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Diffusion-Based Contextual Reconstruction for Point Cloud Segmentation with Limited Annotations
abstract
Point cloud semantic segmentation is fundamental to 3D scene understanding, but dense annotation requirements limit scalability. Although recent label propagation and contrastive learning methods enhance local consistency, the incomplete object coverage caused by sparse annotations hinders global context modeling, ultimately limiting overall performance. To this end, we propose a diffusion-based contextual reconstruction framework for point cloud semantic segmentation with limited annotations. At its core, our framework guides denoising with semantic predictions, using better context reconstruction to enhance the conditional model for better segmentation. Specifically, our contributions include: (1) Diffusion-based segmentation framework: reconstructs contextual semantics from noise under conditional guidance, sharing the decoder with the segmentation module for robust contextual semantic learning. (2) Dynamically aggregates local context from segmentation features and guides denoising with global spatial structure, significantly enhancing denoising quality and contextual awareness. Notably, we pioneer diffusion models for 3D semantic segmentation with limited annotations, enabling efficient single-step inference. Experiments show robustness across varying annotation ratios and state-of-the-art performance on benchmarks.
Jiawei Lian, Zhengxue Wang, Wentao Qu, Haobo Jiang, Le Hui, Jian Yang 0003
AAAI5
2026 Discriminative region learning for point cloud-based place recognition
Le Hui, Yun Zhu 0011, Jianjun Qian, Yigong Zhang, Jin Xie 0001
Neural Networks2
2026 Adaptive geometry-semantic fusion for few-shot 3D point cloud classification
Le Hui, Shigang Liu
Pattern Recognit.3
2026 Learning 3D Representation From Auto-Labeled 2D Object Boxes
abstract
Recent advances in LiDAR representation learning with limited annotations show strong promise. Existing well-performed methods mainly focus on distilling the 2D representation into the 3D representation via superpixels. Superpixels are used to construct the cross-modal contrastive learning, leading to semantic ambiguity of 3D features belonging to the same object and impairing the performance. To this end, we aim to leverage unlabeled LiDAR-camera pairs to design a novel pre-training pipeline, which learns from category space directly and pulls the 3D features belonging to the same object close. Specifically, we obtain autolabeled 2D object boxes with a fixed 2D open-vocabulary object detector and transform the labeled 2D object boxes into high-quality pixel-wise label maps with a box-to-label-maps generation algorithm. Based on the pseudo labels, we present a dual-space pre-training 3D network that recognizes accurate categories from the semantic priors of paired 3D points and segments complete objects. Furthermore, we propose a module named AdaptPro to improve performance further when fine-tuning the 3D network under limited annotations, aiming to explore the unpaired 3D features that lack 2D correspondences via category prototypes. The experimental results show that our method achieves state-of-the-art performances on both the nuScenes and SemanticKITTI benchmark datasets. Code is avialable at https://github.com/dengq7/Box4Scene.
Le Hui, Jian Yang 0003, Jin Xie 0001
IEEE Trans. Image Process.2
2026 Multi-Granularity Superpoint Graph Learning for Weakly Supervised 3D Semantic Segmentation
abstract
Weakly supervised 3D semantic segmentation has proven effective in alleviating the heavy dependence on dense annotations by generating high-quality pseudo-labels. However, due to the scene complexity and disorder of the point cloud, merely applying the model semantic prediction or hand-crafted feature similarity for pseudo labeling is inefficient and biased. This limitation inevitably results in incorrect pseudo labels. To tackle this challenge, we propose a new method called Multi-granularity Superpoint Graph Learning (MSGL) that leverages the multi-scale local features of point clouds to improve the quality of pseudo labels. We first design a multi-granularity local representation learning module on the superpoint graph to capture the neighboring structure information of each superpoint within complex scenes. Subsequently, the generated structural embedding is utilized to enhance the affinity matrix of label propagation, thereby yielding high-quality pseudo labels. To further enforce the generalization of the structural representation module under scenario changes or data fluctuations, we present a multi-granularity consistency loss in MSGL. This loss is applied across different views of the superpoint graph within each scene to ensure a robust and consistent learning process. Our experiments conducted on three benchmarks show that the proposed method outperforms existing weakly supervised methods under several sparse label settings, and improves the baseline by an average of 7.7% with only 1% extra computation cost. Moreover, our approach even compares favorably to some fully supervised methods with only one point labeled for each thing.
Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Bin Xiao 0002, Qinghua Hu
IEEE Trans. Multim.4
2025 Geometry-Aware 3D Salient Object Detection Network
abstract
Point cloud salient object detection has attracted the attention of researchers in recent years. Since existing works do not fully utilize the geometry context of 3D objects, blurry boundaries are generated when segmenting objects with complex backgrounds. In this paper, we propose a geometry-aware 3D salient object detection network that explicitly clusters points into superpoints to enhance the geometric boundaries of objects, thereby segmenting complete objects with clear boundaries. Specifically, we first propose a simple yet effective superpoint partition module to cluster points into superpoints. In order to improve the quality of superpoints, we present a point cloud class-agnostic loss to learn discriminative point features for clustering superpoints from the object. After obtaining superpoints, we then propose a geometry enhancement module that utilizes superpoint-point attention to aggregate geometric information into point features for predicting the salient map of the object with clear boundaries. Extensive experiments show that our method achieves new state-of-the-art performance on the PCSOD dataset.
Chen Wang 0049, Le Hui, Qi Liu 0054, Yuchao Dai
AAAI3
2025 Sketchy Bounding-box Supervision for 3D Instance Segmentation
abstract
Bounding box supervision has gained considerable attention in weakly supervised 3D instance segmentation. While this approach alleviates the need for extensive point-level annotations, obtaining accurate bounding boxes in practical applications remains challenging. To this end, we explore the inaccurate bounding box, named sketchy bounding box, which is imitated through perturbing ground truth bounding box by adding scaling, translation, and rotation. In this paper, we propose Sketchy-3DIS, a novel weakly 3D instance segmentation framework, which jointly learns pseudo labeler and segmentator to improve the performance under the sketchy bounding-box supervisions. Specifically, we first propose an adaptive box-to-point pseudo labeler that adaptively learns to assign points located in the overlapped parts between two sketchy bounding boxes to the correct instance, resulting in compact and pure pseudo instance labels. Then, we present a coarse-to-fine instance segmentator that first predicts coarse instances from the entire point cloud and then learns fine instances based on the region of coarse instances. Finally, by using the pseudo instance labels to supervise the instance segmentator, we can gradually generate high-quality instances through joint training. Extensive experiments show that our method achieves state-of-the-art performance on both the ScanNetV2 and S3DIS benchmarks, and even outperforms several fully supervised methods using sketchy bounding boxes. Code is available at https://github.com/dengq7/Sketchy-3DIS.
Le Hui, Jin Xie 0001, Jian Yang 0003
CVPR2
2025 Learning Class Prototypes for Unified Sparse-Supervised 3D Object Detection
abstract
Both indoor and outdoor scene perceptions are essential for embodied intelligence. However, current sparse supervised 3D object detection methods focus solely on outdoor scenes without considering indoor settings. To this end, we propose a unified sparse supervised 3D object detection method for both indoor and outdoor scenes through learning class prototypes to effectively utilize unlabeled objects. Specifically, we first propose a prototype-based object mining module that converts the unlabeled object mining into a matching problem between class prototypes and unlabeled features. By using optimal transport matching results, we assign prototype labels to high-confidence features, thereby achieving the mining of unlabeled objects. We then present a multi-label cooperative refinement module to effectively recover missed detections through pseudo label quality control and prototype label cooperation. Experiments show that our method achieves state-of-the-art performance under the one object per scene sparse supervised setting across indoor and outdoor datasets. With only one labeled object per scene, our method achieves about 78%, 90%, and 96% performance compared to the fully supervised detector on ScanNet V2, SUN RGB-D, and KITTI, respectively, highlighting the scalability of our method. Code is available at https://github.com/zyrant/CPDet3D.
Yun Zhu 0011, Le Hui, Jianjun Qian, Jin Xie 0001, Jian Yang 0003
CVPR2
2025 GSRecon: Efficient Generalizable Gaussian Splatting for Surface Reconstruction from Sparse Views
Le Hui, Jianjun Qian, Jin Xie 0001, Jian Yang 0003
ICCV2
2025 Deep Height Decoupling for Precise Vision-Based 3D Occupancy Prediction
abstract
The task of vision-based 3D occupancy prediction aims to reconstruct 3D geometry and estimate its semantic classes from 2D color images, where the 2D-to-3D view transformation is an indispensable step. Most previous methods conduct forward projection, such as BEVPooling and VoxelPooling, both of which map the 2D image features into 3D grids. However, the current grid representing features within a certain height range usually introduces many confusing features that belong to other height ranges. To address this challenge, we present Deep Height Decoupling (DHD), a novel framework that incorporates explicit height prior to filter out the confusing features. Specifically, DHD first predicts height maps via explicit supervision. Based on the height distribution statistics, DHD designs Mask Guided Height Sampling (MGHS) to adaptively decouple the height map into multiple binary masks. MGHS projects the 2D image features into multiple subspaces, where each grid contains features within reasonable height ranges. Finally, a Synergistic Feature Aggregation (SFA) module is deployed to enhance the feature representation through channel and spatial affinities, enabling further occupancy refinement. On the popular Occ3D-nuScenes benchmark, our method achieves state-of-the-art performance even with minimal input frames. Source code is released at https://github.com/yanzq95/DHD.
Zhiqiang Yan 0001, Zhengxue Wang, Xiang Li 0041, Le Hui, Jian Yang 0003
ICRA5
2025 Cross-View Geometric Collaboration for Generalizable Sparse View Neural Surface Reconstruction
abstract
Generalizable neural implicit surface reconstruction aims to recover accurate surfaces with sparse views from unseen scenes. Most existing methods suffer from severe incompleteness and inaccuracies in the case of reconstruction with large viewpoint variations, as significant perspective distortions across views lead to unreliable feature correspondence and geometry representations. In this paper, we propose a cross-view geometric collaboration framework for generalizable neural surface reconstruction, which exploits cross-view complementary geometric information to improve the accuracy and robustness of reconstruction from sparse views. Specifically, we propose a cross-view geometry complement module that utilizes the reliable geometric information of different views to refine geometric representations. In addition, we construct a distortion-robust patch-based consistency volume to provide supplementary geometric cues for uncertain regions. For the rendering process, we develop a cross-view geometry transformer to adaptively aggregate reliable cross-view point features by considering geometric context along the ray. Finally, we render per-view depth maps and fuse them to reconstruct the final surface. Extensive experimental results on the DTU, BlendedMVS, and Tanks and Temples datasets demonstrate the superior reconstruction quality and view-combination generalizability of our solution.
Le Hui, Jianjun Qian, Jian Yang 0003, Yigong Zhang, Jin Xie 0001
ACM Multimedia2
2025 Instance-Level Moving Object Segmentation from a Single Image with Events
Zhexiong Wan, Bin Fan 0002, Le Hui, Yuchao Dai, Gim Hee Lee
Int. J. Comput. Vis.3
2025 RigNet++: Semantic Assisted Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan 0001, Xiang Li 0041, Le Hui, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
Int. J. Comput. Vis.3
2025 Uncertainty-Aware Superpoint Graph Transformer for Weakly Supervised 3-D Semantic Segmentation
abstract
Weakly supervised 3-D semantic segmentation has successfully mitigated the labor-intensive and time-consuming task of annotating 3-D point clouds. However, reliably utilizing the minimal point-wise annotations for unlabeled data in complex and large-scale scenes is still challenging, such as only 20 points labeled in 2 million points. To tackle this challenge, we propose a new Uncertainty-aware Superpoint Graph Transformer (UaSGT) framework that utilizes minimal annotations for unlabeled data learning through reliable long-range supervision propagation from labeled superpoints to unlabeled superpoints. First, we propose a superpoint graph transformer to achieve long-range supervision propagation along the attention-based fuzzy subsets defined on superpoints. The attention-based fuzzy subset measures the membership of unlabeled superpoints to clusters centered on labeled superpoints. Second, we employ an uncertainty-aware membership rectification technique on the fuzzy subset to ensure reliable propagation among superpoints within the same category. This technique integrates an uncertainty prediction module to mask the influence of unreliable membership and a spatial prior refinement module to reduce uncertainty in intraclass membership degrees. Finally, experimental results on two large-scale benchmarks S3DIS and ScanNet-V2 demonstrate the superiority of our approach compared to the state-of-the-art with at least 90% annotation reduction, and our method also achieves comparable performance to fully supervised methods with less than 0.1% labeled points.
Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Qinghua Hu
IEEE Trans. Fuzzy Syst.4
2025 Cross-Modal Driven Object Restoration for 3D Point Cloud Backdoor Defense
abstract
3D point cloud recognition plays a critical role in autonomous driving, robotics, and medical diagnostics. However, its vulnerability to backdoor attacks remains underexplored, posing significant security risks in real-world applications. Current defense mechanisms against 3D point cloud backdoor attacks are still in their infancy and lacking effective solutions. To address this, we propose a cross-modal driven object restoration framework that leverages 3D reconstruction to mitigate backdoor attacks. Specifically, we introduce a cross-modal semantic encoding module that projects 3D point clouds into multi-view depth maps and utilizes CLIP to extract aligned text-image features, providing semantic guidance for 3D reconstruction. Furthermore, we leverage cross-modal information as conditional guidance to drive dynamic diffusion-based 3D reconstruction and adaptively fuse semantic and geometric features through a gated self-conditioned modulator. This module dynamically selects features for fusion, effectively mitigating noise interference and distribution shifts during latent diffusion, significantly enhancing robustness to noise, and thereby achieving precise restoration of clean point clouds. Extensive experiments on ModelNet40, and ShapeNetPart datasets demonstrate that our method robustly defends against adaptive attacks under varying noise levels and significantly restores classification performance degraded by backdoor triggers.
Jiawei Lian, Xia Du, Jianghua Liu 0001, Le Hui, Jian Yang 0003
IEEE Trans. Inf. Forensics Secur.4
2025 ZS-VAT: Learning Unbiased Attribute Knowledge for Zero-Shot Recognition Through Visual Attribute Transformer
abstract
In zero-shot learning (ZSL), attribute knowledge plays a vital role in transferring knowledge from seen classes to unseen classes. However, most existing ZSL methods learn biased attribute knowledge, which usually results in biased attribute prediction and a decline in zero-shot recognition performance. To solve this problem and learn unbiased attribute knowledge, we propose a visual attribute Transformer for zero-shot recognition (ZS-VAT), which is an effective and interpretable Transformer designed specifically for ZSL. In ZS-VAT, we design an attribute-head self-attention (AHSA) that is capable of learning unbiased attribute knowledge. Specifically, each attribute head in AHSA first transforms the local features into attribute-reinforced features and then accumulates the attribute knowledge from all corresponding reinforced features, reducing the mutual influence between attributes and avoiding information loss. AHSA finally preserves unbiased attribute knowledge through attribute embeddings. We also propose an attribute fusion model (AFM) that learns to recover the correct category knowledge from the attribute knowledge. In particular, AFM takes all features from AHSA as input and generates global embeddings. We carried out experiments to demonstrate that the attribute knowledge from AHSA and the category knowledge from AFM are able to assist each other. During the final semantic prediction, we combine the attribute embedding prediction (AEP) and global embedding prediction (GEP). We evaluated the proposed scheme on three benchmark datasets. ZS-VAT outperformed the state-of-the-art generalized ZSL (GZSL) methods on two datasets and achieved competitive results on the other dataset.
Zongyan Han, Zhenyong Fu, Shuo Chen 0003, Le Hui, Jian Yang 0003, Chang Wen Chen
IEEE Trans. Neural Networks Learn. Syst.4
2025 Point Cloud Registration-Driven Robust Feature Matching for 3-D Siamese Object Tracking
abstract
Learning robust feature matching between the template and search area is crucial for 3-D Siamese tracking. The core of Siamese feature matching is how to assign high feature similarity to the corresponding points between the template and the search area for precise object localization. In this article, we propose a novel point cloud registration-driven Siamese tracking framework, with the intuition that spatially aligned corresponding points (via 3-D registration) tend to achieve consistent feature representations. Specifically, our method consists of two modules, including a tracking-specific nonlocal registration (TSNR) module and a registration-aided Sinkhorn template-feature aggregation module. The registration module targets the precise spatial alignment between the template and the search area. The tracking-specific spatial distance constraint is proposed to refine the cross-attention weights in the nonlocal module for discriminative feature learning. Then, we use the weighted singular value decomposition (SVD) to compute the rigid transformation between the template and the search area and align them to achieve the desired spatially aligned corresponding points. For the feature aggregation model, we formulate the feature matching between the transformed template and the search area as an optimal transport problem and utilize the Sinkhorn optimization to search for the outlier-robust matching solution. Also, a registration-aided spatial distance map is built to improve the matching robustness in indistinguishable regions (e.g., smooth surfaces). Finally, guided by the obtained feature matching map, we aggregate the target information from the template into the search area to construct the target-specific feature, which is then fed into a CenterPoint-like detection head for object localization. Extensive experiments on KITTI, NuScenes, and Waymo datasets verify the effectiveness of our proposed method.
Haobo Jiang, Kaihao Lan, Le Hui, Jin Xie 0001, Shangbing Gao, Jian Yang 0003
IEEE Trans. Neural Networks Learn. Syst.3
2025 Weakly Supervised Object Localization With Progressive Activation Diffusion
abstract
Weakly supervised object localization (WSOL) aims to locate objects with only image-level labels. Previous works mainly follow the framework of class activation map (CAM), which discovers the objects by estimating the contribution of each pixel position to the category prediction. However, most of them overlook the pixel-level spatial and semantic contextual correlation, resulting in: 1) limited activation ranges that only highlight the most discriminative parts rather than the entire object and 2) low activation values for some foreground parts, especially regions near the boundary between foreground and background. To alleviate this issue, we propose an activation diffusion network (ADNet) to progressively refine both the range and value of activations on the localization map. Specifically, a context propagation module is first developed to learn the top-down spatial dependency between adjacent feature maps, which helps back-propagate the activation from the discriminative part to its surroundings for more complete objects. Then, a diffusion probability distillation module (DPDM) is proposed, which transfers the pixel-level semantic correlation emerging in the image generation process to the localization map generation in a teacher-student learning manner. This helps boost the value of the activated foreground region and stimulates the value of neighboring inactivated foreground positions to sharpen the object boundary. Experiments on various datasets and backbones demonstrate the superiority of our ADNet over state-of-the-art (SOTA) methods in object localization and segmentation, yielding 82.2% and 62.2% Top-1 Loc on Caltech-UCSD Birds-200-2011 (CUB) and ImageNet Large-ScaleVisual Recognition Challenge (ILSVRC) datasets and 76.6% pixel average precision (PxAP) on OpenImages dataset. Qualitative results also show that we can achieve a more complete and consistent activation covering the whole object.
Can Xu 0006, Le Hui, Jin Xie 0001, Jian Yang 0003
IEEE Trans. Neural Networks Learn. Syst.2
2024 SPGroup3D: Superpoint Grouping Network for Indoor 3D Object Detection
abstract
Current 3D object detection methods for indoor scenes mainly follow the voting-and-grouping strategy to generate proposals. However, most methods utilize instance-agnostic groupings, such as ball query, leading to inconsistent semantic information and inaccurate regression of the proposals. To this end, we propose a novel superpoint grouping network for indoor anchor-free one-stage 3D object detection. Specifically, we first adopt an unsupervised manner to partition raw point clouds into superpoints, areas with semantic consistency and spatial similarity. Then, we design a geometry-aware voting module that adapts to the centerness in anchor-free detection by constraining the spatial relationship between superpoints and object centers. Next, we present a superpoint-based grouping module to explore the consistent representation within proposals. This module includes a superpoint attention layer to learn feature interaction between neighboring superpoints, and a superpoint-voxel fusion layer to propagate the superpoint-level information to the voxel level. Finally, we employ effective multiple matching to capitalize on the dynamic receptive fields of proposals based on superpoints during the training. Experimental results demonstrate our method achieves state-of-the-art performance on ScanNet V2, SUN RGB-D, and S3DIS datasets in the indoor one-stage 3D object detection. Source code is available at https://github.com/zyrant/SPGroup3D.
Yun Zhu 0011, Le Hui, Yaqi Shen, Jin Xie 0001
AAAI2
2024 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis
abstract
In this paper, we propose a 3D geometry-aware deformable Gaussian Splatting method for dynamic view synthesis. Existing neural radiance fields (NeRF) based solutions learn the deformation in an implicit manner, which cannot incorporate 3D scene geometry. Therefore, the learned deformation is not necessarily geometrically coherent, which results in unsatisfactory dynamic view synthesis and 3D dynamic reconstruction. Recently, 3D Gaussian Splatting provides a new representation of the 3D scene, building upon which the 3D geometry could be exploited in learning the complex 3D deformation. Specifically, the scenes are represented as a collection of 3D Gaussian, where each 3D Gaussian is optimized to move and rotate over time to model the deformation. To enforce the 3D scene geometry constraint during deformation, we explicitly extract 3D geometry features and integrate them in learning the 3D deformation. In this way, our solution achieves 3D geometry-aware deformation modeling, which enables improved dynamic view synthesis and 3D dynamic reconstruction. Extensive experimental results on both synthetic and real datasets prove the superiority of our solution, which achieves new state-of-the-art performance. The project is available at https://npucvr.github.io/GaGS/.
Zhicheng Lu, Le Hui, Yuchao Dai
CVPR3
2024 Improving Depth Completion via Depth Feature Upsampling
abstract
The encoder-decoder network (ED-Net) is a commonly employed choice for existing depth completion methods, but its working mechanism is ambiguous. In this paper, we vi-sualize the internal feature maps to analyze how the net-work densifies the input sparse depth. We find that the en-coder feature of ED-Net focus on the areas with input depth points around. To obtain a dense feature and thus esti-mate complete depth, the decoder feature tends to comple-ment and enhance the encoder feature by skip-connection to make the fused encoder-decoder feature dense, resulting in the decoder feature also exhibits sparse. However, ED-Net obtains the sparse decoder feature from the dense fused feature at the previous stage, where the “dense-i-sparse‘’ process destroys the completeness of features and loses in-formation. To address this issue, we present a depth feature upsampling network (DFU) that explicitly utilizes these dense features to guide the upsampling of a low-resolution (LR) depth feature to a high-resolution (HR) one. The completeness of features is maintained throughout the up-sampling process, thus avoiding information loss. Fur-thermore, we propose a confidence-aware guidance module (CGM), which is confidence-aware and performs guidance with adaptive receptive fields (GARF), to fully exploit the potential of these dense features as guidance. Experimental results show that our DFU, a plug-and-play module, can significantly improve the performance of existing ED-Net based methods with limited computational overheads, and new SOTA results are achieved. Besides, the generalization capability on sparser depth is also enhanced. Project page: https://npucvr.github.iolDFU.
Ge Zhang 0006, Shaoqian Wang, Bo Li 0090, Qi Liu 0054, Le Hui, Yuchao Dai
CVPR6
2024 Multi-Attribute Interactions Matter for 3D Visual Grounding
abstract
3D visual grounding aims to localize 3D objects described by free-form language sentences. Following the detection-then-matching paradigm, existing methods mainly focus on embedding object attributes in unimodal feature extraction and multimodal feature fusion, to enhance the discriminability of the proposal feature for accurate grounding. However, most of them ignore the explicit interaction of multiple attributes, causing a bias in unimodal representation and misalignment in multimodal fusion. In this paper, we propose a multi-attribute aware Transformer for 3D visual grounding, learning the multi-attribute interactions to refine the intra-modal and inter-modal grounding cues. Specifically, we first develop an attribute causal analysis module to quantify the causal effect of different attributes for the final prediction, which provides powerful supervision to correct the misleading attributes and adaptively capture other discriminative features. Then, we design an exchanging-based multimodal fusion module, which dynamically replaces tokens with low attribute attention between modalities before directly integrating low-dimensional global features. This ensures an attribute-level multimodal information fusion and helps align the language and vision details more efficiently for fine-grained multimodal features. Extensive experiments show that our method can achieve state-of-the-art performance on ScanRefer and Sr3D/Nr3D datasets. The code is publicly available at https://github.com/volcanoXC/MA2TransVG.
Can Xu 0006, Yuehui Han, Rui Xu 0021, Le Hui, Jin Xie 0001, Jian Yang 0003
CVPR4
2024 3D Focusing-and-Matching Network for Multi-Instance Point Cloud Registration
abstract
Multi-instance point cloud registration aims to estimate the pose of all instances of a model point cloud in the whole scene. Existing methods all adopt the strategy of first obtaining the global correspondence and then clustering to obtain the pose of each instance. However, due to the cluttered and occluded objects in the scene, it is difficult to obtain an accurate correspondence between the model point cloud and all instances in the scene. To this end, we propose a simple yet powerful 3D focusing-and-matching network for multi-instance point cloud registration by learning the multiple pair-wise point cloud registration. Specifically, we first present a 3D multi-object focusing module to locate the center of each object and generate object proposals. By using self-attention and cross-attention to associate the model point cloud with structurally similar objects, we can locate potential matching instances by regressing object centers. Then, we propose a 3D dual-masking instance matching module to estimate the pose between the model point cloud and each object proposal. It performs instance mask and overlap mask masks to accurately predict the pair-wise correspondence. Extensive experiments on two public benchmarks, Scan2CAD and ROBI, show that our method achieves a new state-of-the-art performance on the multi-instance point cloud registration task.
Le Hui, Qi Liu 0054, Bo Li 0090, Yuchao Dai
NeurIPS2
2024 Learning Local Semantic Region Activations for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) aims to train instance-level locators by exploiting accessible image-level labels. By multiplying channel-wise features with classification weights and then adding them together, most prior works follow the pipeline of the Class Activation Map (CAM) to collect the semantic responses, thereby highlighting regions that contribute to class prediction to achieve WSOL. However, CAM-based methods treat the class contributions of all pixel positions in a channel equally and assign dominant weights for the discriminative channels biasedly. This fails to express the fine-grained pixel-level semantic response of each channel and model the complex contextual relations between channels, resulting in the mixup of the activation value between non-discriminative foreground regions and the background. To alleviate these issues, we present a Local Semantic activation enhancement and Global Spatial correlation mining network (LSGS-Net) for accurate WSOL. Specifically, we first propose a local activation generation module to explicitly learn the semantic response of each pixel position from channels. Then, we design a regularization loss to supervise the consistency between similar local activations, which utilizes the cross-image information to improve the accuracy of local activations. We further propose a K-nearest Neighbors graph module to capture the spatial correlation between different local activations, which can adaptively assign more proper weights when fusing all local activation. In the inference stage, the bounding box will be determined with a foreground threshold. Extensive experiments show that LSGS-Net achieves significant and consistent improvement with various backbones on the CUB, ILSVRC, and OpenImages benchmarks, with a 97.5% and 75.3% GT-Known LOC on CUB and ILSVRC, respectively. For segmentation quality on OpenImages, LSGS-Net already exceeds the SOTA method by 1.2% pIoU and 1.9% PxAP.
Can Xu 0006, Le Hui, Yuehui Han, Haobo Jiang, Jiaxin Chen 0001, Jin Xie 0001, Jian Yang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2023 Self-Supervised 3D Scene Flow Estimation Guided by Superpoints
abstract
3D scene flow estimation aims to estimate point-wise motions between two consecutive frames of point clouds. Superpoints, i.e., points with similar geometric features, are usually employed to capture similar motions of local regions in 3D scenes for scene flow estimation. However, in existing methods, superpoints are generated with the offline clustering methods, which cannot characterize local regions with similar motions for complex 3D scenes well, leading to inaccurate scene flow estimation. To this end, we propose an iterative end-to-end super point based scene flow estimation framework, where the superpoints can be dynamically updated to guide the point-level flow prediction. Specifically, our framework consists of a flow guided superpoint generation module and a superpoint guided flow refinement module. In our superpoint generation module, we utilize the bidirectional flow information at the previous iteration to obtain the matching points of points and superpoint centers for soft point-to-superpoint association construction, in which the superpoints are generated for pairwise point clouds. With the generated superpoints, we first reconstruct the flow for each point by adaptively aggregating the superpoint-level flow, and then encode the consistency between the reconstructed flow of pairwise point clouds. Finally, we feed the consistency encoding along with the reconstructed flow into GRU to refine point-level flow. Extensive experiments on several different datasets show that our method can achieve promising performance. Code is available at https://github.com/supersyq/SPFlowNet.
Yaqi Shen, Le Hui, Jin Xie 0001, Jian Yang 0003
CVPR2
2023 Efficient LiDAR Point Cloud Oversegmentation Network
abstract
Point cloud oversegmentation is a challenging task since it needs to produce perceptually meaningful partitions (i.e., superpoints) of a point cloud. Most existing oversegmentation methods cannot efficiently generate superpoints from large-scale LiDAR point clouds due to complex and inefficient procedures. In this paper, we propose a simple yet efficient end-to-end LiDAR oversegmentation network, which segments superpoints from the LiDAR point cloud by grouping points based on low-level point embeddings. Specifically, we first learn the similarity of points from the constructed local neighborhoods to obtain low-level point embeddings through the local discriminative loss. Then, to generate homogeneous superpoints from the sparse LiDAR point cloud, we propose a LiDAR point grouping algorithm that simultaneously considers the similarity of point embeddings and the Euclidean distance of points in 3D space. Finally, we design a superpoint refinement module for accurately assigning the hard boundary points to the corresponding superpoints. Extensive results on two large-scale outdoor datasets, SemanticKITTI and nuScenes, show that our method achieves a new state-of-the-art in LiDAR oversegmentation. Notably, the inference time of our method is 100× faster than that of other methods. Furthermore, we apply the learned superpoints to the LiDAR semantic segmentation task and the results show that using superpoints can significantly improve the LiDAR semantic segmentation of the baseline network. Code is available at https://github.com/fpthink/SuperLiDAR.
Le Hui, Linghua Tang, Yuchao Dai, Jin Xie 0001, Jian Yang 0003
ICCV1
2023 Transformer-based Point Cloud Generation Network
abstract
Point cloud generation is an important research topic in 3D computer vision, which can provide high-quality datasets for various downstream tasks. However, efficiently capturing the geometry of point clouds remains a challenging problem due to their irregularities. In this paper, we propose a novel transformer-based 3D point cloud generation network to generate realistic point clouds. Specifically, we first develop a transformer-based interpolation module that utilizes k-nearest neighbors at different scales to learn global and local information about point clouds in the feature space. Based on geometric information, we interpolate new point features to upsample the point cloud features. Then, the upsampled features are used to generate a coarse point cloud with spatial coordinate information. We construct a transformer-based refinement module to enhance the upsampled features in feature space with geometric information in coordinate space. Finally, we use a multi-layer perceptron on the upsampled features to generate the final point cloud. Extensive experiments on ShapeNet and ModelNet demonstrate the effectiveness of our proposed method.
Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001
ACM Multimedia2
2023 Scene Graph Masked Variational Autoencoders for 3D Scene Generation
abstract
Generating realistic 3D indoor scenes requires a deep understanding of objects and their spatial relationships. However, existing methods often fail to generate realistic 3D scenes due to the limited understanding of object relationships. To tackle this problem, we propose a Scene Graph Masked Variational Auto-Encoder (SG-MVAE) framework that fully captures the relationships between objects to generate more realistic 3D scenes. Specifically, we first introduce a relationship completion module that adaptively learns the missing relationships between objects in the scene graph. To accurately predict the missing relationships, we employ multi-group attention to capture the correlations between the objects with missing relationships and other objects in the scene. After obtaining the complete scene relationships, we mask the relationships between objects and use a decoder to reconstruct the scene. The reconstruction process enhances the model's understanding of relationships, generating more realistic scenes. Extensive experiments on benchmark datasets show that our model outperforms state-of-the-art methods.
Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001
ACM Multimedia2
2023 Multi-Scale Superpoint Network for 3D Point Cloud Semantic Segmentation
abstract
3D point cloud semantic segmentation is a fundamental task for 3D scene understanding. However, most existing pipelines usually use k-NN or ball query operation to form hard neighborhoods, which may cross different semantic objects, resulting low-quality local features. To address this issue, we propose a multi-scale superpoint network that gradually generates multi-scale soft neighborhoods to extract geometric local features, thereby boosting the 3D semantic segmentation performance. Specifically, we present a simple yet efficient superpoint merging module that merge small-scale superpoints to obtain large-scale superpoint by considering the feature similarity of superpoints, so that we can obtain multi-scale geometric features of point clouds. We also develop a superpoint upsampling module that adopt inverse mapping function to propagate multi-scale features from low-resolution point cloud to high-resolution point cloud. By integrating our multi-scale superpoint network into a simple point based semantic segmentation network, our method can obtain SOTA results on S3DIS Area 5 and 6-fold, and competitive results on ScanNet v2.
Ft Zheng, Le Hui, Jin Xie 0001, Haofeng Zhang 0001
MMAsia2
2023 FlowFormer: 3D scene flow estimation for point clouds with transformers
Yaqi Shen, Le Hui
Knowl. Based Syst.2
2022 Reliable Inlier Evaluation for Unsupervised Point Cloud Registration
abstract
Unsupervised point cloud registration algorithm usually suffers from the unsatisfied registration precision in the partially overlapping problem due to the lack of effective inlier evaluation. In this paper, we propose a neighborhood consensus based reliable inlier evaluation method for robust unsupervised point cloud registration. It is expected to capture the discriminative geometric difference between the source neighborhood and the corresponding pseudo target neighborhood for effective inlier distinction. Specifically, our model consists of a matching map refinement module and an inlier evaluation module. In our matching map refinement module, we improve the point-wise matching map estimation by integrating the matching scores of neighbors into it. The aggregated neighborhood information potentially facilitates the discriminative map construction so that high-quality correspondences can be provided for generating the pseudo target point cloud. Based on the observation that the outlier has the significant structure-wise difference between its source neighborhood and corresponding pseudo target neighborhood while this difference for inlier is small, the inlier evaluation module exploits this difference to score the inlier confidence for each estimated correspondence. In particular, we construct an effective graph representation for capturing this geometric difference between the neighborhoods. Finally, with the learned correspondences and the corresponding inlier confidence, we use the weighted SVD algorithm for transformation estimation.Under the unsupervised setting, we exploit the Huber function based global alignment loss, the local neighborhood consensus loss and spatial consistency loss for model optimization. The experimental results on extensive datasets demonstrate that our unsupervised point cloud registration method can yield comparable performance.
Yaqi Shen, Le Hui, Haobo Jiang, Jin Xie 0001, Jian Yang 0003
AAAI2
2022 Domain Disentangled Generative Adversarial Network for Zero-Shot Sketch-Based 3D Shape Retrieval
abstract
Sketch-based 3D shape retrieval is a challenging task due to the large domain discrepancy between sketches and 3D shapes. Since existing methods are trained and evaluated on the same categories, they cannot effectively recognize the categories that have not been used during training. In this paper, we propose a novel domain disentangled generative adversarial network (DD-GAN) for zero-shot sketch-based 3D retrieval, which can retrieve the unseen categories that are not accessed during training. Specifically, we first generate domain-invariant features and domain-specific features by disentangling the learned features of sketches and 3D shapes, where the domain-invariant features are used to align with the corresponding word embeddings. Then, we develop a generative adversarial network that combines the domain-specific features of the seen categories with the aligned domain-invariant features to synthesize samples, where the synthesized samples of the unseen categories are generated by using the corresponding word embeddings. Finally, we use the synthesized samples of the unseen categories combined with the real samples of the seen categories to train the network for retrieval, so that the unseen categories can be recognized. In order to reduce the domain shift problem, we utilize unlabeled unseen samples to enhance the discrimination ability of the discriminator. With the discriminator distinguishing the generated samples from the unlabeled unseen samples, the generator can generate more realistic unseen samples. Extensive experiments on the SHREC'13 and SHREC'14 datasets show that our method significantly improves the retrieval performance of the unseen categories.
Rui Xu 0021, Zongyan Han, Le Hui, Jianjun Qian, Jin Xie 0001
AAAI3
2022 Learning Inter-superpoint Affinity for Weakly Supervised 3D Instance Segmentation
Linghua Tang, Le Hui, Jin Xie 0001
ACCV (1)2
2022 Generative Subgraph Contrast for Self-Supervised Graph Representation Learning
Yuehui Han, Le Hui, Haobo Jiang, Jianjun Qian, Jin Xie 0001
ECCV (30)2
2022 RA-Depth: Resolution Adaptive Self-supervised Monocular Depth Estimation
Le Hui, Yikai Bian, Jin Xie 0001, Jian Yang 0003
ECCV (27)2
2022 3D Siamese Transformer Network for Single Object Tracking on Point Clouds
Le Hui, Lingpeng Wang, Linghua Tang, Kaihao Lan, Jin Xie 0001, Jian Yang 0003
ECCV (2)1
2022 Unsupervised Domain Adaptation for Point Cloud Semantic Segmentation via Graph Matching
abstract
Unsupervised domain adaptation for point cloud semantic segmentation has attracted great attention due to its effectiveness in learning with unlabeled data. Most of existing methods use global-level feature alignment to transfer the knowledge from the source domain to the target domain, which may cause the semantic ambiguity of the feature space. In this paper, we propose a graph-based framework to explore the local-level feature alignment between the two domains, which can reserve semantic discrimination during adaptation. Specifically, in order to extract local-level features, we first dynamically construct local feature graphs on both domains and build a memory bank with the graphs from the source domain. In particular, we use optimal transport to generate the graph matching pairs. Then, based on the assignment matrix, we can align the feature distributions between the two domains with the graph-based local feature loss. Furthermore, we consider the correlation between the features of different categories and formulate a category-guided contrastive loss to guide the segmentation model to learn discriminative features on the target domain. Extensive experiments on different synthetic-to-real and real-to-real domain adaptation scenarios demonstrate that our method can achieve state-of-the-art performance. Our code is available at https://github.com/BianYikai/PointUDA.
Yikai Bian, Le Hui, Jianjun Qian, Jin Xie 0001
IROS2
2022 Learning Superpoint Graph Cut for 3D Instance Segmentation
abstract
3D instance segmentation is a challenging task due to the complex local geometric structures of objects in point clouds. In this paper, we propose a learning-based superpoint graph cut method that explicitly learns the local geometric structures of the point cloud for 3D instance segmentation. Specifically, we first oversegment the raw point clouds into superpoints and construct the superpoint graph. Then, we propose an edge score prediction network to predict the edge scores of the superpoint graph, where the similarity vectors of two adjacent nodes learned through cross-graph attention in the coordinate and feature spaces are used for regressing edge scores. By forcing two adjacent nodes of the same instance to be close to the instance center in the coordinate and feature spaces, we formulate a geometry-aware edge loss to train the edge score prediction network. Finally, we develop a superpoint graph cut network that employs the learned edge scores and the predicted semantic classes of nodes to generate instances, where bilateral graph attention is proposed to extract discriminative features on both the coordinate and feature spaces for predicting semantic labels and scores of instances. Extensive experiments on two challenging datasets, ScanNet v2 and S3DIS, show that our method achieves new state-of-the-art performance on 3D instance segmentation.
Le Hui, Linghua Tang, Yaqi Shen, Jin Xie 0001, Jian Yang 0003
NeurIPS1
2022 Efficient 3D Point Cloud Feature Learning for Large-Scale Place Recognition
abstract
Point cloud based retrieval for place recognition is still a challenging problem since the drastic appearance changes of scenes due to seasonal or artificial changes in the environments. Existing deep learning based global descriptors for the retrieval task usually consume a large amount of computational resources ( e.g ., memory), which may not be suitable for the cases of limited hardware resources. In this paper, we develop an efficient point cloud learning network (EPC-Net) to generate global descriptors of point clouds for place recognition. While obtaining good performance, it can greatly reduce computational memory and inference time. First, we propose a lightweight but effective neural network module, called ProxyConv, to aggregate the local geometric features of point clouds. We leverage the adjacency matrix and proxy points to simplify the original edge convolution for lower memory consumption. Then, we design a lightweight grouped VLAD network to form global descriptors for retrieval. Compared with the original VLAD network, we propose a grouped fully connected layer to decompose the high-dimensional vectors into a group of low-dimensional vectors, which can reduce the number of parameters of the network and maintain the discrimination of the feature vector. Finally, we further develop a simple version of EPC-Net, called EPC-Net-L, which consists of two ProxyConv modules and one max pooling layer to aggregate global descriptors. By distilling the knowledge from EPC-Net, EPC-Net-L can obtain discriminative global descriptors for retrieval. Extensive experiments on the Oxford dataset and three in-house datasets demonstrate that our method achieves good results with lower parameters, FLOPs, GPU memory, and shorter inference time. Our code is available at https://github.com/fpthink/EPC-Net.
Le Hui, Mingmei Cheng, Jin Xie 0001, Jian Yang 0003, Ming-Ming Cheng
IEEE Trans. Image Process.1
2021 SSPC-Net: Semi-supervised Semantic 3D Point Cloud Segmentation Network
abstract
Point cloud semantic segmentation is a crucial task in 3D scene understanding. Existing methods mainly focus on employing a large number of annotated labels for supervised semantic segmentation. Nonetheless, manually labeling such large point clouds for the supervised segmentation task is time-consuming. In order to reduce the number of annotated labels, we propose a semi-supervised semantic point cloud segmentation network, named SSPC-Net, where we train the semantic segmentation network by inferring the labels of unlabeled points from the few annotated 3D points. In our method, we first partition the whole point cloud into superpoints and build superpoint graphs to mine the long-range dependencies in point clouds. Based on the constructed superpoint graph, we then develop a dynamic label propagation method to generate the pseudo labels for the unsupervised superpoints. Particularly, we adopt a superpoint dropout strategy to dynamically select the generated pseudo labels. In order to fully exploit the generated pseudo labels of the unsupervised superpoints, we furthermore propose a coupled attention mechanism for superpoint feature embedding. Finally, we employ the cross-entropy loss to train the semantic segmentation network with the labels of the supervised superpoints and the pseudo labels of the unsupervised superpoints. Experiments on various datasets demonstrate that our semisupervised segmentation method can achieve better performance than the current semi-supervised segmentation method with fewer annotated 3D points.
Mingmei Cheng, Le Hui, Jin Xie 0001, Jian Yang 0003
AAAI2
2021 Pyramid Point Cloud Transformer for Large-Scale Place Recognition
abstract
Recently, deep learning based point cloud descriptors have achieved impressive results in the place recognition task. Nonetheless, due to the sparsity of point clouds, how to extract discriminative local features of point clouds to efficiently form a global descriptor is still a challenging problem. In this paper, we propose a pyramid point cloud transformer network (PPT-Net) to learn the discriminative global descriptors from point clouds for efficient retrieval. Specifically, we first develop a pyramid point transformer module that adaptively learns the spatial relationship of the different k-NN neighboring points of point clouds, where the grouped self-attention is proposed to extract discriminative local features of the point clouds. The grouped self-attention not only enhances long-term dependencies of the point clouds, but also reduces the computational cost. In order to obtain discriminative global descriptors, we construct a pyramid VLAD module to aggregate the multi-scale feature maps of point clouds into the global descriptors. By applying VLAD pooling on multi-scale feature maps, we utilize the context gating mechanism on the multiple global descriptors to adaptively weight the multi-scale global context information into the final global descriptor. Experimental results on the Oxford dataset and three in-house datasets show that our method achieves the state-of-the-art on the point cloud based place recognition task. Code is available at https://github.com/fpthink/PPT-Net.
Le Hui, Mingmei Cheng, Jin Xie 0001, Jian Yang 0003
ICCV1
2021 Superpoint Network for Point Cloud Oversegmentation
abstract
Superpoints are formed by grouping similar points with local geometric structures, which can effectively reduce the number of primitives of point clouds for subsequent point cloud processing. Existing superpoint methods mainly focus on employing clustering or graph partition to generate superpoints with handcrafted or learned features. Nonetheless, these methods cannot learn superpoints of point clouds with an end-to-end network. In this paper, we develop a new deep iterative clustering network to directly generate superpoints from irregular 3D point clouds in an end-to-end manner. Specifically, in our clustering network, we first jointly learn a soft point-superpoint association map from the coordinate and feature spaces of point clouds, where each point is assigned to the superpoint with a learned weight. Furthermore, we then iteratively update the association map and superpoint centers so that we can more accurately group the points into the corresponding superpoints with locally similar geometric structures. Finally, by predicting the pseudo labels of the superpoint centers, we formulate a label consistency loss on the points and superpoint centers to train the network. Extensive experiments on various datasets indicate that our method not only achieves the state-of-the-art on superpoint generation but also improves the performance of point cloud semantic segmentation. Code is available at https://github.com/fpthink/SPNet.
Le Hui, Jia Yuan, Mingmei Cheng, Jin Xie 0001, Jian Yang 0003
ICCV1
2021 SSPU-Net: Self-Supervised Point Cloud Upsampling via Differentiable Rendering
abstract
Point clouds obtained from 3D sensors are usually sparse. Existing methods mainly focus on upsampling sparse point clouds in a supervised manner by using dense ground truth point clouds. In this paper, we propose a self-supervised point cloud upsampling network (SSPU-Net) to generate dense point clouds without using ground truth. To achieve this, we exploit the consistency between the input sparse point cloud and generated dense point cloud for the shapes and rendered images. Specifically, we first propose a neighbor expansion unit (NEU) to upsample the sparse point clouds, where the local geometric structures of the sparse point clouds are exploited to learn weights for point interpolation. Then, we develop a differentiable point cloud rendering unit (DRU) as an end-to-end module in our network to render the point cloud into multi-view images. Finally, we formulate a shape-consistent loss and an image-consistent loss to train the network so that the shapes of the sparse and dense point clouds are as consistent as possible. Extensive results on the CAD and scanned datasets demonstrate that our method can achieve impressive results in a self-supervised manner.
Le Hui, Jin Xie 0001
ACM Multimedia2
2021 3D Siamese Voxel-to-BEV Tracker for Sparse Point Clouds
abstract
3D object tracking in point clouds is still a challenging problem due to the sparsity of LiDAR points in dynamic environments. In this work, we propose a Siamese voxel-to-BEV tracker, which can significantly improve the tracking performance in sparse 3D point clouds. Specifically, it consists of a Siamese shape-aware feature learning network and a voxel-to-BEV target localization network. The Siamese shape-aware feature learning network can capture 3D shape information of the object to learn the discriminative features of the object so that the potential target from the background in sparse point clouds can be identified. To this end, we first perform template feature embedding to embed the template's feature into the potential target and then generate a dense 3D shape to characterize the shape information of the potential target. For localizing the tracked target, the voxel-to-BEV target localization network regresses the target's 2D center and the z-axis center from the dense bird's eye view (BEV) feature map in an anchor-free manner. Concretely, we compress the voxelized point cloud along z-axis through max pooling to obtain a dense BEV feature map, where the regression of the 2D center and the z-axis center can be performed more effectively. Extensive evaluation on the KITTI tracking dataset shows that our method significantly outperforms the current state-of-the-art methods by a large margin. Code is available at https://github.com/fpthink/V2B.
Le Hui, Lingpeng Wang, Mingmei Cheng, Jin Xie 0001, Jian Yang 0003
NeurIPS1
2021 Facilitating 3D Object Tracking in Point Clouds with Image Semantics and Geometry
Lingpeng Wang, Le Hui, Jin Xie 0001
PRCV (1)2
2020 Progressive Point Cloud Deconvolution Generation Network
Le Hui, Rui Xu 0021, Jin Xie 0001, Jianjun Qian, Jian Yang 0003
ECCV (15)1
2020 Cascaded Non-local Neural Network for Point Cloud Semantic Segmentation
abstract
In this paper, we propose a cascaded non-local neural network for point cloud segmentation. The proposed network aims to build the long-range dependencies of point clouds for the accurate segmentation. Specifically, we develop a novel cascaded non-local module, which consists of the neighborhood-level, superpoint-level and global-level non-local blocks. First, in the neighborhood-level block, we extract the local features of the centroid points of point clouds by assigning different weights to the neighboring points. The extracted local features of the centroid points are then used to encode the superpoint-level block with the non-local operation. Finally, the global-level block aggregates the non-local features of the superpoints for semantic segmentation in an encoder-decoder framework. Benefiting from the cascaded structure, geometric structure information of different neighborhoods with the same label can be propagated. In addition, the cascaded structure can largely reduce the computational cost of the original non-local operation on point clouds. Experiments on different indoor and outdoor datasets show that our method achieves state-of-the-art performance and effectively reduces the time consumption and memory occupation.
Mingmei Cheng, Le Hui, Jin Xie 0001, Jian Yang 0003, Hui Kong 0001
IROS2
2019 Data-Adaptive Metric Learning with Scale Alignment
abstract
The central problem for most existing metric learning methods is to find a suitable projection matrix on the differences of all pairs of data points. However, a single unified projection matrix can hardly characterize all data similarities accurately as the practical data are usually very complicated, and simply adopting one global projection matrix might ignore important local patterns hidden in the dataset. To address this issue, this paper proposes a novel method dubbed “Data-Adaptive Metric Learning” (DAML), which constructs a data-adaptive projection matrix for each data pair by selectively combining a set of learned candidate matrices. As a result, every data pair can obtain a specific projection matrix, enabling the proposed DAML to flexibly fit the training data and produce discriminative projection results. The model of DAML is formulated as an optimization problem which jointly learns candidate projection matrices and their sparse combination for every data pair. Nevertheless, the over-fitting problem may occur due to the large amount of parameters to be learned. To tackle this issue, we adopt the Total Variation (TV) regularizer to align the scales of data embedding produced by all candidate projection matrices, and thus the generated metrics of these learned candidates are generally comparable. Furthermore, we extend the basic linear DAML model to the kernerlized version (denoted “KDAML”) to handle the non-linear cases, and the Iterative Shrinkage-Thresholding Algorithm (ISTA) is employed to solve the optimization model. Intensive experimental results on various applications including retrieval, classification, and verification clearly demonstrate the superiority of our algorithm to other state-of-the-art metric learning methodologies.
Shuo Chen 0003, Chen Gong 0002, Jian Yang 0003, Ying Tai, Le Hui, Jun Li 0027
AAAI5
2019 Inter-Class Angular Loss for Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have shown great power in various classification tasks and have achieved remarkable results in practical applications. However, the distinct learning difficulties in discriminating different pairs of classes are largely ignored by the existing networks. For instance, in CIFAR-10 dataset, distinguishing cats from dogs is usually harder than distinguishing horses from ships. By carefully studying the behavior of CNN models in the training process, we observe that the confusion level of two classes is strongly correlated with their angular separability in the feature space. That is, the larger the inter-class angle is, the lower the confusion will be. Based on this observation, we propose a novel loss function dubbed “Inter-Class Angular Loss” (ICAL), which explicitly models the class correlation and can be directly applied to many existing deep networks. By minimizing the proposed ICAL, the networks can effectively discriminate the examples in similar classes by enlarging the angle between their corresponding class vectors. Thorough experimental results on a series of vision and nonvision datasets confirm that ICAL critically improves the discriminative ability of various representative deep neural networks and generates superior performance to the original networks with conventional softmax loss.
Le Hui, Xiang Li 0041, Chen Gong 0002, Joey Tianyi Zhou, Jian Yang 0003
AAAI1
2018 Unsupervised Multi-Domain Image Translation with Domain-Specific Encoders/Decoders
abstract
Unsupervised Image-to-Image Translation achieves spectacularly advanced developments nowadays. However, recent approaches mainly focus on one model with two domains, which may face heavy burdens with the large cost of training time and the huge model parameters, under such a requirement that domains are freely transferred to each other in a general setting. To address this problem, we propose a novel and unified framework named Domain-Bank, which consists of a globally shared auto-encoder and n domain-specific encoders/decoders, assuming that there is a universal shared-latent space can be projected. Thus, we not only reduce the parameters of the model but also have a huge reduction of the time budgets. Besides the high efficiency, we show the comparable (or even better) image translation results over state-of-the-arts on various challenging unsupervised image translation tasks, including face image translation and painting style translation. We also apply the proposed framework to the domain adaptation task and achieve state-of-the-art performance on digit benchmark datasets.
Le Hui, Xiang Li 0041, Jiaxin Chen 0001, Jian Yang 0003
ICPR1
2018 Attention-based Neural Network for Traffic Sign Detection
abstract
Existing object detection pipelines can show superior performance for large objects with high resolution but fail to detect very small objects such as traffic signs. So, detecting traffic signs is a proverbially challenging problem. In this paper, we propose a novel end-to-end architecture that improves small object detection by combining Faster R-CNN with the attention mechanism. Specifically, we focus on channel-wise features and utilize the attention mechanism to enhance the feature responses by explicitly modeling the interdependencies between channel-wise features. Finally, the regression of bounding boxes and the classification of traffic signs are generated after selecting the discriminative features by the attention mechanism. Extensive evaluations of the largest traffic sign dataset demonstrate that the attention mechanism improves the performance of detecting objects, especially the small targets. For traffic sign detection task, our method achieves better performance compared with many state-of-the-art approaches on the largest traffic sign detection dataset, Tsinghua-Tencent 100K.
Le Hui, Jianfeng Lu 0003, Yuhua Zhu
ICPR2