VLDB 2026 Research / reviewers in the wild / expert
Jin Xie 0001
dblp:80/1949-1
· DBLP profile ↗
106ranked-venue papers
9as first author
75since 2021 · last 2026
0000-0002-4933-5229ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 6 first-author · 53 since 2021Artificial intelligence and machine learning · 72 · 6 first-author · 51 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Systems, architecture and hardware · 5 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discriminative region learning for point cloud-based place recognition
Le Hui, Yun Zhu 0011, Jianjun Qian, Yigong Zhang, Jin Xie 0001 |
Neural Networks | 6 |
| 2026 | Learning 3D Representation From Auto-Labeled 2D Object BoxesabstractRecent advances in LiDAR representation learning with limited annotations show strong promise. Existing well-performed methods mainly focus on distilling the 2D representation into the 3D representation via superpixels. Superpixels are used to construct the cross-modal contrastive learning, leading to semantic ambiguity of 3D features belonging to the same object and impairing the performance. To this end, we aim to leverage unlabeled LiDAR-camera pairs to design a novel pre-training pipeline, which learns from category space directly and pulls the 3D features belonging to the same object close. Specifically, we obtain autolabeled 2D object boxes with a fixed 2D open-vocabulary object detector and transform the labeled 2D object boxes into high-quality pixel-wise label maps with a box-to-label-maps generation algorithm. Based on the pseudo labels, we present a dual-space pre-training 3D network that recognizes accurate categories from the semantic priors of paired 3D points and segments complete objects. Furthermore, we propose a module named AdaptPro to improve performance further when fine-tuning the 3D network under limited annotations, aiming to explore the unpaired 3D features that lack 2D correspondences via category prototypes. The experimental results show that our method achieves state-of-the-art performances on both the nuScenes and SemanticKITTI benchmark datasets. Code is avialable at https://github.com/dengq7/Box4Scene. Le Hui, Jian Yang 0003, Jin Xie 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Multi-Granularity Superpoint Graph Learning for Weakly Supervised 3D Semantic SegmentationabstractWeakly supervised 3D semantic segmentation has proven effective in alleviating the heavy dependence on dense annotations by generating high-quality pseudo-labels. However, due to the scene complexity and disorder of the point cloud, merely applying the model semantic prediction or hand-crafted feature similarity for pseudo labeling is inefficient and biased. This limitation inevitably results in incorrect pseudo labels. To tackle this challenge, we propose a new method called Multi-granularity Superpoint Graph Learning (MSGL) that leverages the multi-scale local features of point clouds to improve the quality of pseudo labels. We first design a multi-granularity local representation learning module on the superpoint graph to capture the neighboring structure information of each superpoint within complex scenes. Subsequently, the generated structural embedding is utilized to enhance the affinity matrix of label propagation, thereby yielding high-quality pseudo labels. To further enforce the generalization of the structural representation module under scenario changes or data fluctuations, we present a multi-granularity consistency loss in MSGL. This loss is applied across different views of the superpoint graph within each scene to ensure a robust and consistent learning process. Our experiments conducted on three benchmarks show that the proposed method outperforms existing weakly supervised methods under several sparse label settings, and improves the baseline by an average of 7.7% with only 1% extra computation cost. Moreover, our approach even compares favorably to some fully supervised methods with only one point labeled for each thing. Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Bin Xiao 0002, Qinghua Hu |
IEEE Trans. Multim. | 5 |
| 2026 | DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset GenerationabstractThe labor- and experience-intensive creation of 3D assets with physically based rendering (PBR) materials demands an autonomous 3D asset creation pipeline. However, most existing 3D generation methods focus on geometry modeling, either baking textures into simple vertex colors or leaving texture synthesis to post-processing with image diffusion models. To achieve end-to-end PBR-ready 3D asset generation, we present Lightweight Gaussian Asset Adapter (LGAA), a novel framework that unifies the modeling of geometry and PBR materials by exploiting multi-view (MV) diffusion priors from a novel perspective. The LGAA features a modular design with three components. Specifically, the LGAA Wrapper reuses and adapts network layers from MV diffusion models, which encapsulate knowledge acquired from billions of images, enabling better convergence in a data-efficient manner. To incorporate multiple diffusion priors for geometry and PBR synthesis, the LGAA Switcher aligns multiple LGAA Wrapper layers encapsulating different knowledge. Then, a tamed variational autoencoder (VAE), termed LGAA Decoder, is designed to predict 2D Gaussian Splatting (2DGS) with PBR channels. Finally, we introduce a dedicated post-processing procedure to effectively extract high-quality, relightable mesh assets from the resulting 2DGS. Extensive quantitative and qualitative experiments demonstrate the superior performance of LGAA with both text- and image-conditioned MV diffusion models. Additionally, the modular design enables flexible incorporation of multiple diffusion priors, and the knowledge-preserving scheme effectively preseves the 2D priors learned on massive image dataset, which leads to data efficient finetuning to lift the MV diffuison models for 3D generation with merely 69k multi-view instances. Our code, pre-trained weights, and the dataset used will be publicly available via our project page: https://zx-yin.github.io/dreamlifting/. Ze-Xin Yin, Jiaxiong Qiu, Wei Sui, Zhizhong Su, Jian Yang 0003, Jin Xie 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object NavigationabstractLearning discriminative state representations of agents, encompassing the spatial layout and temporal pose trajectory, is essential for effective navigation decisions. However, existing approaches often rely on simplistic plain networks for navigation information fusion, overlooking the complex long-range dependencies across spatio-temporal cues, which leads to suboptimal state perception and potential decision failures. In this paper, we introduce NaviFormer, an effective encoder-decoder navigation transformer, to aggregate discriminative spatio-temporal context information for object navigation. Our navigation encoder not only encodes spatial layouts and temporal agent poses but also innovatively constructs and encodes a passable frontier map, enriching the original state encoding with cues of potential exploration regions. Furthermore, our navigation decoder employs spatio-temporal self-attention and cross-attention mechanisms to model the dependencies among spatial layout encoding, temporal pose encoding, and passable frontier encoding, thereby facilitating comprehensive contextual state feature aggregation. Finally, we leverage these learned spatio-temporal contextual state representations for PPO-based navigation decisions. Extensive experiments on the Gibson, Habitat-Matterport3D (HM3D) and Matterport3D (MP3D) datasets demonstrate the superiority of our approach. Wei Xie 0019, Haobo Jiang, Yun Zhu 0011, Jianjun Qian, Jin Xie 0001 |
AAAI | 5 |
| 2025 | Sketchy Bounding-box Supervision for 3D Instance SegmentationabstractBounding box supervision has gained considerable attention in weakly supervised 3D instance segmentation. While this approach alleviates the need for extensive point-level annotations, obtaining accurate bounding boxes in practical applications remains challenging. To this end, we explore the inaccurate bounding box, named sketchy bounding box, which is imitated through perturbing ground truth bounding box by adding scaling, translation, and rotation. In this paper, we propose Sketchy-3DIS, a novel weakly 3D instance segmentation framework, which jointly learns pseudo labeler and segmentator to improve the performance under the sketchy bounding-box supervisions. Specifically, we first propose an adaptive box-to-point pseudo labeler that adaptively learns to assign points located in the overlapped parts between two sketchy bounding boxes to the correct instance, resulting in compact and pure pseudo instance labels. Then, we present a coarse-to-fine instance segmentator that first predicts coarse instances from the entire point cloud and then learns fine instances based on the region of coarse instances. Finally, by using the pseudo instance labels to supervise the instance segmentator, we can gradually generate high-quality instances through joint training. Extensive experiments show that our method achieves state-of-the-art performance on both the ScanNetV2 and S3DIS benchmarks, and even outperforms several fully supervised methods using sketchy bounding boxes. Code is available at https://github.com/dengq7/Sketchy-3DIS. Le Hui, Jin Xie 0001, Jian Yang 0003 |
CVPR | 3 |
| 2025 | Zero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision ModelabstractThis paper introduces ZeroMatch, a novel zero-shot RGB-D point cloud registration framework, aimed at achieving robust 3D matching on unseen data without any task-specific training. Our core idea is to utilize the powerful zero-shot image representation of Stable Diffusion, achieved through extensive pre-training on large-scale data, to enhance point-cloud geometric descriptors for robust matching. Specifically, we combine the handcrafted geometric descriptor FPFH with Stable-Diffusion features to create point descriptors that are both locally and contextually aware, enabling reliable RGB-D registration with zero-shot capability. This approach is based on our observation that Stable-Diffusion features effectively encode discriminative global contextual cues, naturally alleviating the feature ambiguity that FPFH often encounters in scenes with repetitive patterns or low overlap. To further enhance cross-view consistency of Stable-Diffusion features for improved matching, we propose a coupled-image input mode that concatenates the source and target images into a single input, replacing the original single-image mode. This design achieves both inter-image and prompt-to-image consistency attentions, facilitating robust cross-view feature interaction and alignment. Finally, we leverage feature nearest neighbors to construct putative correspondences for hypothesize-and-verify transformation estimation. Extensive experiments on 3DMatch, ScanNet, and ScanLoNet verify the excellent zero-shot matching ability of our method. [Code] Haobo Jiang, Jin Xie 0001, Jian Yang 0003, Liang Yu 0005, Jianmin Zheng |
CVPR | 2 |
| 2025 | SVG-IR: Spatially-Varying Gaussian Splatting for Inverse RenderingabstractReconstructing 3D assets from images, known as inverse rendering (IR), remains a challenging task due to its ill-posed nature. 3D Gaussian Splatting (3DGS) has demonstrated impressive capabilities for novel view synthesis (NVS) tasks. Methods apply it to relighting by separating radiance into BRDF parameters and lighting, yet produce inferior relighting quality with artifacts and unnatural indirect illumination due to the limited capability of each Gaussian, which has constant material parameters and normal, alongside the absence of physical constraints for indirect lighting. In this paper, we present a novel framework called Spatially-vayring Gaussian Inverse Rendering (SVG-IR), aimed at enhancing both NVS and relighting quality. To this end, we propose a new representation—Spatially-varying Gaussian (SVG)—that allows per-Gaussian spatially varying parameters. This enhanced representation is complemented by a SVG splatting scheme akin to vertex/fragment shading in traditional graphics pipelines. Furthermore, we integrate a physically-based indirect lighting model, enabling more realistic relighting. The proposed SVG-IR framework significantly improves rendering quality, outperforming state-of-the-art NeRF-based methods by 2.5 dB in peak signal-to-noise ratio (PSNR) and surpassing existing Gaussian-based techniques by 3.5 dB in relighting tasks, all while maintaining a real-time rendering speed. The source code is available at https://github.com/learner-shx/SVG-IR. Hanxiao Sun, Yupeng Gao, Jin Xie 0001, Jian Yang 0003, Beibei Wang 0002 |
CVPR | 3 |
| 2025 | WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusionabstract3D scene perception demands a large amount of adverse-weather LiDAR data, yet the cost of LiDAR data collection presents a significant scaling-up challenge. To this end, a series of LiDAR simulators have been proposed. Yet, they can only simulate a single adverse weather with a single physical model, and the fidelity of the generated data is quite limited. This paper presents WeatherGen, the first unified diverse-weather LiDAR data diffusion generation framework, significantly improving fidelity. Specifically, we first design a map-based data producer, which can provide a vast amount of high-quality diverse-weather data for training purposes. Then, we utilize the diffusion-denoising paradigm to construct a diffusion model. Among them, we propose a spider mamba generator to restore the disturbed diverse weather data gradually. The spider mamba models the feature interactions by scanning the Li-Dar beam circle or central ray, excellently maintaining the physical structure of the LiDAR data. Subsequently, following the generator to transfer real-world knowledge, we design a latent feature aligner. Afterward, we devise a contrastive learning-based controller, which equips weather control signals with compact semantic knowledge through language supervision, guiding the diffusion model to generate more discriminative data. Extensive evaluations demonstrate the high generation quality of WeatherGen. Through WeatherGen, we construct the mini-weather dataset, promoting the performance of the downstream task under adverse weather conditions. Code is available: https://github.com/wuyang98/weathergen Yun Zhu 0011, Kaihua Zhang 0001, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
CVPR | 5 |
| 2025 | Learning Class Prototypes for Unified Sparse-Supervised 3D Object DetectionabstractBoth indoor and outdoor scene perceptions are essential for embodied intelligence. However, current sparse supervised 3D object detection methods focus solely on outdoor scenes without considering indoor settings. To this end, we propose a unified sparse supervised 3D object detection method for both indoor and outdoor scenes through learning class prototypes to effectively utilize unlabeled objects. Specifically, we first propose a prototype-based object mining module that converts the unlabeled object mining into a matching problem between class prototypes and unlabeled features. By using optimal transport matching results, we assign prototype labels to high-confidence features, thereby achieving the mining of unlabeled objects. We then present a multi-label cooperative refinement module to effectively recover missed detections through pseudo label quality control and prototype label cooperation. Experiments show that our method achieves state-of-the-art performance under the one object per scene sparse supervised setting across indoor and outdoor datasets. With only one labeled object per scene, our method achieves about 78%, 90%, and 96% performance compared to the fully supervised detector on ScanNet V2, SUN RGB-D, and KITTI, respectively, highlighting the scalability of our method. Code is available at https://github.com/zyrant/CPDet3D. Yun Zhu 0011, Le Hui, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
CVPR | 5 |
| 2025 | VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow PredictionabstractRecent advancements in camera-based occupancy prediction have focused on the simultaneous prediction of 3D semantics and scene flow, a task that presents significant challenges due to specific difficulties, e.g., occlusions and unbalanced dynamic environments. In this paper, we analyze these challenges and their underlying causes. To address them, we propose a novel regularization framework called VoxelSplat. This framework leverages recent developments in 3D Gaussian Splatting to enhance model performance in two key ways: (i) Enhanced Semantics Supervision through 2D Projection: During training, our method decodes sparse semantic 3D Gaussians from 3D representations and projects them onto the 2D camera view. This provides additional supervision signals in the camera-visible space, allowing 2D labels to improve the learning of 3D semantics. (ii) Scene Flow Learning: Our framework uses the predicted scene flow to model the motion of Gaussians, and is thus able to learn the scene flow of moving objects in a self-supervised manner using the labels of adjacent frames. Our method can be seamlessly integrated into various existing occupancy models, enhancing performance without increasing inference time. Extensive experiments on benchmark datasets demonstrate the effectiveness of Voxel-Splat in improving the accuracy of both semantic occupancy and scene flow estimation. The project page and codes are available at https://zzy816.github.io/VoxelSplat-Demo/. Ziyue Zhu, Shenlong Wang, Jin Xie 0001, Jiang-jiang Liu, Jingdong Wang 0001, Jian Yang 0003 |
CVPR | 3 |
| 2025 | AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous DrivingabstractModeling and rendering dynamic urban driving scenes is crucial for self-driving simulation. Current high-quality methods typically rely on costly manual object tracklet annotations, while self-supervised approaches fail to capture dynamic object motions accurately and decompose scenes properly, resulting in rendering artifacts. We introduce AD-GS, a novel self-supervised framework for high-quality free-viewpoint rendering of driving scenes from a single log. At its core is a novel learnable motion model that integrates locality-aware B-spline curves with global-aware trigonometric functions, enabling flexible yet precise dynamic object modeling. Rather than requiring comprehensive semantic labeling, AD-GS automatically segments scenes into objects and background with the simplified pseudo 2D segmentation, representing objects using dynamic Gaussians and bidirectional temporal visibility masks. Further, our model incorporates visibility reasoning and physically rigid regularization to enhance robustness. Extensive evaluations demonstrate that our annotation-free model significantly outperforms current state-of-the-art annotation-free methods and is competitive with annotation-dependent approaches. Zexin Fan, Shenlong Wang, Jin Xie 0001, Jian Yang 0003 |
ICCV | 5 |
| 2025 | GSRecon: Efficient Generalizable Gaussian Splatting for Surface Reconstruction from Sparse Views
Le Hui, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
ICCV | 4 |
| 2025 | DiffPCI: Large Motion Point Cloud Frame Interpolation with Diffusion Model
Haobo Jiang, Jian Yang 0003, Jin Xie 0001 |
ICCV | 4 |
| 2025 | Generative Point Cloud RegistrationabstractIn this paper, we propose a novel 3D registration paradigm, Generative Point Cloud Registration, which bridges advanced 2D generative models with 3D matching tasks to enhance registration performance. Our key idea is to generate cross-view consistent image pairs that are well-aligned with the source and target point clouds, enabling geometric-color feature fusion to facilitate robust matching. To ensure high-quality matching, the generated image pair should feature both 2D-3D geometric consistency and cross-view texture consistency. To achieve this, we introduce Match-ControlNet, a matching-specific, controllable 2D generative model. Specifically, it leverages the depth-conditioned generation capability of ControlNet to produce images that are geometrically aligned with depth maps derived from point clouds, ensuring 2D-3D geometric consistency. Additionally, by incorporating a coupled conditional denoising scheme and coupled prompt guidance, Match-ControlNet further promotes cross-view feature interaction, guiding texture consistency generation. Our generative 3D registration paradigm is general and could be seamlessly integrated into various registration methods to enhance their performance. Extensive experiments on 3DMatch and ScanNet datasets verify the effectiveness of our approach. Haobo Jiang, Jin Xie 0001, Jian Yang 0003, Liang Yu 0005, Jianmin Zheng |
ICML | 2 |
| 2025 | Cross-View Geometric Collaboration for Generalizable Sparse View Neural Surface ReconstructionabstractGeneralizable neural implicit surface reconstruction aims to recover accurate surfaces with sparse views from unseen scenes. Most existing methods suffer from severe incompleteness and inaccuracies in the case of reconstruction with large viewpoint variations, as significant perspective distortions across views lead to unreliable feature correspondence and geometry representations. In this paper, we propose a cross-view geometric collaboration framework for generalizable neural surface reconstruction, which exploits cross-view complementary geometric information to improve the accuracy and robustness of reconstruction from sparse views. Specifically, we propose a cross-view geometry complement module that utilizes the reliable geometric information of different views to refine geometric representations. In addition, we construct a distortion-robust patch-based consistency volume to provide supplementary geometric cues for uncertain regions. For the rendering process, we develop a cross-view geometry transformer to adaptively aggregate reliable cross-view point features by considering geometric context along the ray. Finally, we render per-view depth maps and fuse them to reconstruct the final surface. Extensive experimental results on the DTU, BlendedMVS, and Tanks and Temples datasets demonstrate the superior reconstruction quality and view-combination generalizability of our solution. Le Hui, Jianjun Qian, Jian Yang 0003, Yigong Zhang, Jin Xie 0001 |
ACM Multimedia | 6 |
| 2025 | GigaSLAM: Large-Scale Monocular SLAM with Hierarchical Gaussian SplatsabstractTracking and mapping in large-scale, unbounded outdoor environments using only monocular RGB input presents substantial challenges for existing SLAM systems. Traditional Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) SLAM methods are typically limited to small, bounded indoor settings. To overcome these challenges, we introduce GigaSLAM, the first NeRF/3DGS-based SLAM framework for kilometer-scale outdoor environments, as mainly demonstrated on the 4 kilometer-scale datasets. Our approach employs a hierarchical sparse voxel map representation, where Gaussians are decoded by neural networks at multiple levels of detail. This design enables efficient, scalable mapping and high-fidelity viewpoint rendering across expansive, unbounded scenes. For front-end tracking, GigaSLAM utilizes a metric depth model combined with epipolar geometry and PnP algorithms to accurately estimate poses, while incorporating a Bag-of-Words-based loop closure mechanism to maintain robust alignment over long trajectories. Consequently, GigaSLAM delivers high-precision tracking and visually faithful rendering on urban outdoor benchmarks, establishing a robust SLAM solution for large-scale, long-term scenarios, and significantly extending the applicability of Gaussian Splatting SLAM systems to unbounded outdoor environments. Yigong Zhang, Jian Yang 0003, Jin Xie 0001 |
SIGGRAPH Asia | 4 |
| 2025 | Uncertainty-Aware Superpoint Graph Transformer for Weakly Supervised 3-D Semantic SegmentationabstractWeakly supervised 3-D semantic segmentation has successfully mitigated the labor-intensive and time-consuming task of annotating 3-D point clouds. However, reliably utilizing the minimal point-wise annotations for unlabeled data in complex and large-scale scenes is still challenging, such as only 20 points labeled in 2 million points. To tackle this challenge, we propose a new Uncertainty-aware Superpoint Graph Transformer (UaSGT) framework that utilizes minimal annotations for unlabeled data learning through reliable long-range supervision propagation from labeled superpoints to unlabeled superpoints. First, we propose a superpoint graph transformer to achieve long-range supervision propagation along the attention-based fuzzy subsets defined on superpoints. The attention-based fuzzy subset measures the membership of unlabeled superpoints to clusters centered on labeled superpoints. Second, we employ an uncertainty-aware membership rectification technique on the fuzzy subset to ensure reliable propagation among superpoints within the same category. This technique integrates an uncertainty prediction module to mask the influence of unreliable membership and a spatial prior refinement module to reduce uncertainty in intraclass membership degrees. Finally, experimental results on two large-scale benchmarks S3DIS and ScanNet-V2 demonstrate the superiority of our approach compared to the state-of-the-art with at least 90% annotation reduction, and our method also achieves comparable performance to fully supervised methods with less than 0.1% labeled points. Yan Fan 0002, Yu Wang 0106, Pengfei Zhu 0001, Le Hui, Jin Xie 0001, Qinghua Hu |
IEEE Trans. Fuzzy Syst. | 5 |
| 2025 | Point Cloud Registration-Driven Robust Feature Matching for 3-D Siamese Object TrackingabstractLearning robust feature matching between the template and search area is crucial for 3-D Siamese tracking. The core of Siamese feature matching is how to assign high feature similarity to the corresponding points between the template and the search area for precise object localization. In this article, we propose a novel point cloud registration-driven Siamese tracking framework, with the intuition that spatially aligned corresponding points (via 3-D registration) tend to achieve consistent feature representations. Specifically, our method consists of two modules, including a tracking-specific nonlocal registration (TSNR) module and a registration-aided Sinkhorn template-feature aggregation module. The registration module targets the precise spatial alignment between the template and the search area. The tracking-specific spatial distance constraint is proposed to refine the cross-attention weights in the nonlocal module for discriminative feature learning. Then, we use the weighted singular value decomposition (SVD) to compute the rigid transformation between the template and the search area and align them to achieve the desired spatially aligned corresponding points. For the feature aggregation model, we formulate the feature matching between the transformed template and the search area as an optimal transport problem and utilize the Sinkhorn optimization to search for the outlier-robust matching solution. Also, a registration-aided spatial distance map is built to improve the matching robustness in indistinguishable regions (e.g., smooth surfaces). Finally, guided by the obtained feature matching map, we aggregate the target information from the template into the search area to construct the target-specific feature, which is then fed into a CenterPoint-like detection head for object localization. Extensive experiments on KITTI, NuScenes, and Waymo datasets verify the effectiveness of our proposed method. Haobo Jiang, Kaihao Lan, Le Hui, Jin Xie 0001, Shangbing Gao, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Weakly Supervised Object Localization With Progressive Activation DiffusionabstractWeakly supervised object localization (WSOL) aims to locate objects with only image-level labels. Previous works mainly follow the framework of class activation map (CAM), which discovers the objects by estimating the contribution of each pixel position to the category prediction. However, most of them overlook the pixel-level spatial and semantic contextual correlation, resulting in: 1) limited activation ranges that only highlight the most discriminative parts rather than the entire object and 2) low activation values for some foreground parts, especially regions near the boundary between foreground and background. To alleviate this issue, we propose an activation diffusion network (ADNet) to progressively refine both the range and value of activations on the localization map. Specifically, a context propagation module is first developed to learn the top-down spatial dependency between adjacent feature maps, which helps back-propagate the activation from the discriminative part to its surroundings for more complete objects. Then, a diffusion probability distillation module (DPDM) is proposed, which transfers the pixel-level semantic correlation emerging in the image generation process to the localization map generation in a teacher-student learning manner. This helps boost the value of the activated foreground region and stimulates the value of neighboring inactivated foreground positions to sharpen the object boundary. Experiments on various datasets and backbones demonstrate the superiority of our ADNet over state-of-the-art (SOTA) methods in object localization and segmentation, yielding 82.2% and 62.2% Top-1 Loc on Caltech-UCSD Birds-200-2011 (CUB) and ImageNet Large-ScaleVisual Recognition Challenge (ILSVRC) datasets and 76.6% pixel average precision (PxAP) on OpenImages dataset. Qualitative results also show that we can achieve a more complete and consistent activation covering the whole object. Can Xu 0006, Le Hui, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | SPGroup3D: Superpoint Grouping Network for Indoor 3D Object DetectionabstractCurrent 3D object detection methods for indoor scenes mainly follow the voting-and-grouping strategy to generate proposals. However, most methods utilize instance-agnostic groupings, such as ball query, leading to inconsistent semantic information and inaccurate regression of the proposals. To this end, we propose a novel superpoint grouping network for indoor anchor-free one-stage 3D object detection. Specifically, we first adopt an unsupervised manner to partition raw point clouds into superpoints, areas with semantic consistency and spatial similarity. Then, we design a geometry-aware voting module that adapts to the centerness in anchor-free detection by constraining the spatial relationship between superpoints and object centers. Next, we present a superpoint-based grouping module to explore the consistent representation within proposals. This module includes a superpoint attention layer to learn feature interaction between neighboring superpoints, and a superpoint-voxel fusion layer to propagate the superpoint-level information to the voxel level. Finally, we employ effective multiple matching to capitalize on the dynamic receptive fields of proposals based on superpoints during the training. Experimental results demonstrate our method achieves state-of-the-art performance on ScanNet V2, SUN RGB-D, and S3DIS datasets in the indoor one-stage 3D object detection. Source code is available at https://github.com/zyrant/SPGroup3D. Yun Zhu 0011, Le Hui, Yaqi Shen, Jin Xie 0001 |
AAAI | 4 |
| 2024 | Multi-Attribute Interactions Matter for 3D Visual Groundingabstract3D visual grounding aims to localize 3D objects described by free-form language sentences. Following the detection-then-matching paradigm, existing methods mainly focus on embedding object attributes in unimodal feature extraction and multimodal feature fusion, to enhance the discriminability of the proposal feature for accurate grounding. However, most of them ignore the explicit interaction of multiple attributes, causing a bias in unimodal representation and misalignment in multimodal fusion. In this paper, we propose a multi-attribute aware Transformer for 3D visual grounding, learning the multi-attribute interactions to refine the intra-modal and inter-modal grounding cues. Specifically, we first develop an attribute causal analysis module to quantify the causal effect of different attributes for the final prediction, which provides powerful supervision to correct the misleading attributes and adaptively capture other discriminative features. Then, we design an exchanging-based multimodal fusion module, which dynamically replaces tokens with low attribute attention between modalities before directly integrating low-dimensional global features. This ensures an attribute-level multimodal information fusion and helps align the language and vision details more efficiently for fine-grained multimodal features. Extensive experiments show that our method can achieve state-of-the-art performance on ScanRefer and Sr3D/Nr3D datasets. The code is publicly available at https://github.com/volcanoXC/MA2TransVG. Can Xu 0006, Yuehui Han, Rui Xu 0021, Le Hui, Jin Xie 0001, Jian Yang 0003 |
CVPR | 5 |
| 2024 | Masked Motion Prediction with Semantic Contrast for Point Cloud Sequence Learning
Yuehui Han, Can Xu 0006, Rui Xu 0021, Jianjun Qian, Jin Xie 0001 |
ECCV (76) | 5 |
| 2024 | Diff-Reg: Diffusion Model in Doubly Stochastic Matrix Space for Registration Problem
Qianliang Wu, Haobo Jiang, Lei Luo 0001, Jun Li 0027, Yaqing Ding 0001, Jin Xie 0001, Jian Yang 0003 |
ECCV (65) | 6 |
| 2024 | Text2LiDAR: Text-Guided LiDAR Point Cloud Generation via Equirectangular Transformer
Kaihua Zhang 0001, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
ECCV (56) | 4 |
| 2024 | FastPCI: Motion-Structure Guided Fast Point Cloud Frame Interpolation
Guocheng Qian, Jin Xie 0001, Jian Yang 0003 |
ECCV (74) | 3 |
| 2024 | Mask6D: Masked Pose Priors for 6D Object Pose EstimationabstractRobust 6D object pose estimation in cluttered or occluded conditions using monocular RGB images remains a challenging task. One reason is that current pose estimation networks struggle to extract discriminative, pose-aware features using 2D feature backbones, especially when the available RGB information is limited due to target occlusion in cluttered scenes. To mitigate this, we propose a novel pose estimation-specific pre-training strategy named Mask6D. Our approach incorporates pose-aware 2D-3D correspondence maps and visible mask maps as additional modal information, which is combined with RGB images for the reconstruction-based model pre-training. Essentially, this 2D-3D correspondence maps a transformed 3D object model to 2D pixels, reflecting the pose information of the target in camera coordinate system. Meanwhile, the integrated visible mask map can effectively guide our model to disregard cluttered background information. In addition, an object-focused pre-training loss function is designed to further facilitate our network to remove the background interference. Finally, we fine-tune our pre-trained pose prior-aware network via conventional pose training strategy to realize the reliable pose prediction. Extensive experiments verify that our method outperforms previous end-to-end pose estimation methods. Yuechen Xie, Haobo Jiang, Jin Xie 0001 |
ICASSP | 3 |
| 2024 | Dense Voxel Representation Network for Implicit Scene CompletionabstractImplicit scene completion aims to learn an implicit representation of dense point clouds from incomplete ones. Since point clouds are disordered and irregular, some implicit scene completion methods learn representations from voxelized point clouds with sparse convolution. Despite achieving promising results, they lack deep exploration of feature learning on empty voxels, which is beneficial for implicit scene completion task. To address this, we propose a dense voxel representation network for implicit scene completion. First, we design a Bird’s-Eye View (BEV) assisted enhancement module to enhance non-empty voxel features by incorporating the information contained in the learned dense BEV features into them through deformable cross-attention. Second, we construct a feature adaptive completion module to adaptively complete voxel features using deformable self-attention, realizing the transfer of the information from non-empty voxels to empty voxels. Extensive experiments on SemanticKITTI and SemanticPOSS datasets demonstrate our method achieves state-of-the-art performance. Fan Dai, Yun Zhu 0011, Yaqi Shen, Jin Xie 0001, Jianjun Qian |
ICME | 4 |
| 2024 | SGNet: Salient Geometric Network for Point Cloud RegistrationabstractPoint Cloud Registration (PCR) is a critical and challenging task in computer vision and robotics. One of the primary difficulties in PCR is identifying salient and meaningful points that exhibit consistent semantic and geometric properties across different scans. Previous methods have encountered challenges with ambiguous matching due to the similarity among patch blocks throughout the entire point cloud and the lack of consideration for efficient global geometric consistency. To address these issues, we propose a new framework that includes several novel techniques. Firstly, we introduce a semantic-aware geometric encoder that combines object-level and patch-level semantic information. This encoder significantly improves registration recall by reducing ambiguity in patch-level superpoint matching. Additionally, we incorporate a prior knowledge approach that utilizes an intrinsic shape signature to identify salient points. This enables us to extract the most salient super points and meaningful dense points in the scene. Secondly, we introduce an innovative transformer that encodes High-Order (HO) geometric features. These features are crucial for identifying salient points within initial overlap regions while considering global high-order geometric consistency. We introduce an anchor node selection strategy to optimize this high-order transformer further. By encoding inter-frame triangle or polyhedron consistency features based on these anchor nodes, we can effectively learn high-order geometric features of salient super points. These high-order features are then propagated to dense points and utilized by a Sinkhorn matching module to identify critical correspondences for successful registration. The experiments conducted on the 3DMatch/3DLoMatch and KITTI datasets demonstrate the effectiveness of our method. Qianliang Wu, Yaqing Ding 0001, Lei Luo 0001, Haobo Jiang, Shuo Gu, Chuanwei Zhou, Jin Xie 0001, Jian Yang 0003 |
IROS | 7 |
| 2024 | Grid4D: 4D Decomposed Hash Encoding for High-Fidelity Dynamic Gaussian SplattingabstractRecently, Gaussian splatting has received more and more attention in the field of static scene rendering. Due to the low computational overhead and inherent flexibility of explicit representations, plane-based explicit methods are popular ways to predict deformations for Gaussian-based dynamic scene rendering models. However, plane-based methods rely on the inappropriate low-rank assumption and excessively decompose the space-time 4D encoding, resulting in overmuch feature overlap and unsatisfactory rendering quality. To tackle these problems, we propose Grid4D, a dynamic scene rendering model based on Gaussian splatting and employing a novel explicit encoding method for the 4D input through the hash encoding. Different from plane-based explicit representations, we decompose the 4D encoding into one spatial and three temporal 3D hash encodings without the low-rank assumption. Additionally, we design a novel attention module that generates the attention scores in a directional range to aggregate the spatial and temporal features. The directional attention enables Grid4D to more accurately fit the diverse deformations across distinct scene components based on the spatial encoded features. Moreover, to mitigate the inherent lack of smoothness in explicit representation methods, we introduce a smooth regularization term that keeps our model from the chaos of deformation prediction. Our experiments demonstrate that Grid4D significantly outperforms the state-of-the-art models in visual quality and rendering speed. Zexin Fan, Jian Yang 0003, Jin Xie 0001 |
NeurIPS | 4 |
| 2024 | Learning Robust Facial Representation From the View of Diversity and Closeness
Chaoyu Zhao, Jianjun Qian, Shumin Zhu, Jin Xie 0001, Jian Yang 0003 |
Int. J. Comput. Vis. | 4 |
| 2024 | Learning Local Semantic Region Activations for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to train instance-level locators by exploiting accessible image-level labels. By multiplying channel-wise features with classification weights and then adding them together, most prior works follow the pipeline of the Class Activation Map (CAM) to collect the semantic responses, thereby highlighting regions that contribute to class prediction to achieve WSOL. However, CAM-based methods treat the class contributions of all pixel positions in a channel equally and assign dominant weights for the discriminative channels biasedly. This fails to express the fine-grained pixel-level semantic response of each channel and model the complex contextual relations between channels, resulting in the mixup of the activation value between non-discriminative foreground regions and the background. To alleviate these issues, we present a Local Semantic activation enhancement and Global Spatial correlation mining network (LSGS-Net) for accurate WSOL. Specifically, we first propose a local activation generation module to explicitly learn the semantic response of each pixel position from channels. Then, we design a regularization loss to supervise the consistency between similar local activations, which utilizes the cross-image information to improve the accuracy of local activations. We further propose a K-nearest Neighbors graph module to capture the spatial correlation between different local activations, which can adaptively assign more proper weights when fusing all local activation. In the inference stage, the bounding box will be determined with a foreground threshold. Extensive experiments show that LSGS-Net achieves significant and consistent improvement with various backbones on the CUB, ILSVRC, and OpenImages benchmarks, with a 97.5% and 75.3% GT-Known LOC on CUB and ILSVRC, respectively. For segmentation quality on OpenImages, LSGS-Net already exceeds the SOTA method by 1.2% pIoU and 1.9% PxAP. Can Xu 0006, Le Hui, Yuehui Han, Haobo Jiang, Jiaxin Chen 0001, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Action Candidate Driven Clipped Double Q-Learning for Discrete and Continuous Action TasksabstractDouble Q-learning is a popular reinforcement learning algorithm in Markov decision process (MDP) problems. Clipped double Q-learning, as an effective variant of double Q-learning, employs the clipped double estimator to approximate the maximum expected action value. Due to the underestimation bias of the clipped double estimator, the performance of clipped Double Q-learning may be degraded in some stochastic environments. In this article, in order to reduce the underestimation bias, we propose an action candidate-based clipped double estimator (AC-CDE) for Double Q-learning. Specifically, we first select a set of elite action candidates with high action values from one set of estimators. Then, among these candidates, we choose the highest valued action from the other set of estimators. Finally, we use the maximum value in the second set of estimators to clip the action value of the chosen action in the first set of estimators and the clipped value is used for approximating the maximum expected action value. Theoretically, the underestimation bias in our clipped Double Q-learning decays monotonically as the number of action candidates decreases. Moreover, the number of action candidates controls the tradeoff between the overestimation and underestimation biases. In addition, we also extend our clipped Double Q-learning to continuous action tasks via approximating the elite continuous action candidates. We empirically verify that our algorithm can more accurately estimate the maximum expected action value on some toy environments and yield good performance on several benchmark problems. Code is available at https://github.com/Jiang-HB/ac_CDQ. Haobo Jiang, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Robust Outlier Rejection for 3D Registration with Variational BayesabstractLearning-based outlier (mismatched correspondence) rejection for robust 3D registration generally formulates the outlier removal as an inlier/outlier classification problem. The core for this to be successful is to learn the discriminative inlier/outlier feature representations. In this paper, we develop a novel variational non-local network-based outlier rejection framework for robust alignment. By reformulating the non-local feature learning with variational Bayesian inference, the Bayesian-driven long-range dependencies can be modeled to aggregate discriminative geometric context information for inlier/outlier distinction. Specifically, to achieve such Bayesian-driven contextual dependencies, each query/key/value component in our nonlocal network predicts a prior feature distribution and a posterior one. Embedded with the inlier/outlier label, the posterior feature distribution is label-dependent and discriminative. Thus, pushing the prior to be close to the discriminative posterior in the training step enables the features sampled from this prior at test time to model highquality long-range dependencies. Notably, to achieve effective posterior feature guidance, a specific probabilistic graphical model is designed over our non-local model, which lets us derive a variational low bound as our optimization objective for model training. Finally, we propose a voting-based inlier searching strategy to cluster the high-quality hypothetical inliers for transformation estimation. Extensive experiments on 3DMatch, 3DLoMatch, and KITTI datasets verify the effectiveness of our method. Code is available at https://github.com/Jiang-HB/VBReg. Haobo Jiang, Zheng Dang, Zhen Wei 0001, Jin Xie 0001, Jian Yang 0003, Mathieu Salzmann |
CVPR | 4 |
| 2023 | Self-Supervised 3D Scene Flow Estimation Guided by Superpointsabstract3D scene flow estimation aims to estimate point-wise motions between two consecutive frames of point clouds. Superpoints, i.e., points with similar geometric features, are usually employed to capture similar motions of local regions in 3D scenes for scene flow estimation. However, in existing methods, superpoints are generated with the offline clustering methods, which cannot characterize local regions with similar motions for complex 3D scenes well, leading to inaccurate scene flow estimation. To this end, we propose an iterative end-to-end super point based scene flow estimation framework, where the superpoints can be dynamically updated to guide the point-level flow prediction. Specifically, our framework consists of a flow guided superpoint generation module and a superpoint guided flow refinement module. In our superpoint generation module, we utilize the bidirectional flow information at the previous iteration to obtain the matching points of points and superpoint centers for soft point-to-superpoint association construction, in which the superpoints are generated for pairwise point clouds. With the generated superpoints, we first reconstruct the flow for each point by adaptively aggregating the superpoint-level flow, and then encode the consistency between the reconstructed flow of pairwise point clouds. Finally, we feed the consistency encoding along with the reconstructed flow into GRU to refine point-level flow. Extensive experiments on several different datasets show that our method can achieve promising performance. Code is available at https://github.com/supersyq/SPFlowNet. Yaqi Shen, Le Hui, Jin Xie 0001, Jian Yang 0003 |
CVPR | 3 |
| 2023 | Efficient LiDAR Point Cloud Oversegmentation NetworkabstractPoint cloud oversegmentation is a challenging task since it needs to produce perceptually meaningful partitions (i.e., superpoints) of a point cloud. Most existing oversegmentation methods cannot efficiently generate superpoints from large-scale LiDAR point clouds due to complex and inefficient procedures. In this paper, we propose a simple yet efficient end-to-end LiDAR oversegmentation network, which segments superpoints from the LiDAR point cloud by grouping points based on low-level point embeddings. Specifically, we first learn the similarity of points from the constructed local neighborhoods to obtain low-level point embeddings through the local discriminative loss. Then, to generate homogeneous superpoints from the sparse LiDAR point cloud, we propose a LiDAR point grouping algorithm that simultaneously considers the similarity of point embeddings and the Euclidean distance of points in 3D space. Finally, we design a superpoint refinement module for accurately assigning the hard boundary points to the corresponding superpoints. Extensive results on two large-scale outdoor datasets, SemanticKITTI and nuScenes, show that our method achieves a new state-of-the-art in LiDAR oversegmentation. Notably, the inference time of our method is 100× faster than that of other methods. Furthermore, we apply the learned superpoints to the LiDAR semantic segmentation task and the results show that using superpoints can significantly improve the LiDAR semantic segmentation of the baseline network. Code is available at https://github.com/fpthink/SuperLiDAR. Le Hui, Linghua Tang, Yuchao Dai, Jin Xie 0001, Jian Yang 0003 |
ICCV | 4 |
| 2023 | Center-Based Decoupled Point Cloud Registration for 6D Object Pose EstimationabstractIn this paper, we propose a novel center-based decoupled point cloud registration framework for robust 6D object pose estimation in real-world scenarios. Our method decouples the translation from the entire transformation by predicting the object center and estimating the rotation in a center-aware manner. This center offset-based translation estimation is correspondence-free, freeing us from the difficulty of constructing correspondences in challenging scenarios, thus improving robustness. To obtain reliable center predictions, we use a multi-view (bird’s eye view and front view) object shape description of the source-point features, with both views jointly voting for the object center. Additionally, we propose an effective shape embedding module to augment the source features, largely completing the missing shape information due to partial scanning, thus facilitating the center prediction. With the center-aligned source and model point clouds, the rotation predictor utilizes feature similarity to establish putative correspondences for SVD-based rotation estimation. In particular, we introduce a center-aware hybrid feature descriptor with a normal correction technique to extract discriminative, part-aware features for high-quality correspondence construction. Our experiments show that our method outperforms the state-of-the-art methods by a large margin on real-world datasets such as TUD-L, LINEMOD, and Occluded-LINEMOD. Code is available at https://github.com/JiangHB/CenterReg. Haobo Jiang, Zheng Dang, Shuo Gu, Jin Xie 0001, Mathieu Salzmann, Jian Yang 0003 |
ICCV | 4 |
| 2023 | Graph Matching Optimization Network for Point Cloud RegistrationabstractPoint Cloud Registration is a fundamental and challenging problem in 3D computer vision. Recent works often utilize geometric structure features in downsampled points (patches) to seek correspondences, then propagate these sparse patch correspondences to the dense level in the corresponding patches' neighborhood. However, they neglect the explicit global scale rigid constraint at the dense level point matching. We claim that the explicit isometry-preserving constraint in the dense level on a global scale is also important for improving feature representation in the training stage. To this end, we propose a Graph Matching Optimization based Network (GMONet for short), which utilizes the graph-matching optimizer to explicitly exert the isometry preserving constraints in the point feature training to improve the point feature representation. Specifically, we exploit a partial graph-matching optimizer to enhance the super point (i.e., down-sampled key points) features and a full graph-matching optimizer to improve the dense level point features in the overlap region. Meanwhile, we leverage the inexact proximal point method and the mini-batch sampling technique to accelerate these two graph-matching optimizers. Given high discriminative point features in the evaluation stage, we utilize the RANSAC approach to estimate the transformation between the scanned pairs. The proposed method has been evaluated on the 3DMatch/3DLoMatch and the KITTI datasets. The experimental results show that our method performs competitively compared to state-of-the-art baselines. Qianliang Wu, Yaqi Shen, Haobo Jiang, Guofeng Mei, Yaqing Ding 0001, Lei Luo 0001, Jin Xie 0001, Jian Yang 0003 |
IROS | 7 |
| 2023 | Graph Spectral Perturbation for 3D Point Cloud Contrastive Learningabstract3D point cloud contrastive learning has attracted increasing attention due to its efficient learning ability. By distinguishing the similarity relationship between positive and negative samples in the feature space, it can learn effective point cloud feature representations without manual annotation. However, most point cloud contrastive learning methods construct contrastive samples by perturbing point clouds in data space or introducing multi-modality/format data, which may be difficult to control the intensity of the perturbation or introduce interference from different modalities/formats. To this end, in this paper, we propose a novel graph spectral perturbation based contrastive learning framework (GSPCon) for efficient and robust self-supervised 3D point cloud representation learning. It aims to perform perturbations in the graph spectral domain to construct contrastive samples of the point cloud. Specifically, we first naturally represent the point cloud as a k-nearest neighbors (KNN) graph, and adaptively transform the coordinates of the points into the graph spectral domain based on the graph Fourier transform (GFT). Then we implement data augmentation in the graph spectral domain by perturbing the spectral representations. Finally, the contrastive samples are generated by employing the inverse graph Fourier transform (IGFT) to transform the augmented spectral representations back to the point clouds. Experimental results show that our method achieves the state-of-the-art performance on various downstream tasks. Source code is available at https://github.com/yh-han/GSPCon.git. Yuehui Han, Jiaxin Chen 0001, Jianjun Qian, Jin Xie 0001 |
ACM Multimedia | 4 |
| 2023 | Implicit Obstacle Map-driven Indoor Navigation Model for Robust Obstacle AvoidanceabstractRobust obstacle avoidance is one of the critical steps for successful goal-driven indoor navigation tasks. Due to the obstacle missing in the visual image and the possible missed detection issue, visual image-based obstacle avoidance techniques still suffer from unsatisfactory robustness. To mitigate it, in this paper, we propose a novel implicit obstacle map-driven indoor navigation framework for robust obstacle avoidance, where an implicit obstacle map is learned based on the historical trial-and-error experience rather than the visual image. In order to further improve the navigation efficiency, a non-local target memory aggregation module is designed to leverage a non-local network to model the intrinsic relationship between the target semantic and the target orientation clues during the navigation process so as to mine the most target-correlated object clues for the navigation decision. Extensive experimental results on AI2-Thor and RoboTHOR benchmarks verify the excellent obstacle avoidance and navigation efficiency of our proposed method.The core source code is available at https://github.com/xwaiyy123/object-navigation. Wei Xie 0019, Haobo Jiang, Shuo Gu, Jin Xie 0001 |
ACM Multimedia | 4 |
| 2023 | Transformer-based Point Cloud Generation NetworkabstractPoint cloud generation is an important research topic in 3D computer vision, which can provide high-quality datasets for various downstream tasks. However, efficiently capturing the geometry of point clouds remains a challenging problem due to their irregularities. In this paper, we propose a novel transformer-based 3D point cloud generation network to generate realistic point clouds. Specifically, we first develop a transformer-based interpolation module that utilizes k-nearest neighbors at different scales to learn global and local information about point clouds in the feature space. Based on geometric information, we interpolate new point features to upsample the point cloud features. Then, the upsampled features are used to generate a coarse point cloud with spatial coordinate information. We construct a transformer-based refinement module to enhance the upsampled features in feature space with geometric information in coordinate space. Finally, we use a multi-layer perceptron on the upsampled features to generate the final point cloud. Extensive experiments on ShapeNet and ModelNet demonstrate the effectiveness of our proposed method. Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001 |
ACM Multimedia | 5 |
| 2023 | Scene Graph Masked Variational Autoencoders for 3D Scene GenerationabstractGenerating realistic 3D indoor scenes requires a deep understanding of objects and their spatial relationships. However, existing methods often fail to generate realistic 3D scenes due to the limited understanding of object relationships. To tackle this problem, we propose a Scene Graph Masked Variational Auto-Encoder (SG-MVAE) framework that fully captures the relationships between objects to generate more realistic 3D scenes. Specifically, we first introduce a relationship completion module that adaptively learns the missing relationships between objects in the scene graph. To accurately predict the missing relationships, we employ multi-group attention to capture the correlations between the objects with missing relationships and other objects in the scene. After obtaining the complete scene relationships, we mask the relationships between objects and use a decoder to reconstruct the scene. The reconstruction process enhances the model's understanding of relationships, generating more realistic scenes. Extensive experiments on benchmark datasets show that our model outperforms state-of-the-art methods. Rui Xu 0021, Le Hui, Yuehui Han, Jianjun Qian, Jin Xie 0001 |
ACM Multimedia | 5 |
| 2023 | Multi-Scale Superpoint Network for 3D Point Cloud Semantic Segmentationabstract3D point cloud semantic segmentation is a fundamental task for 3D scene understanding. However, most existing pipelines usually use k-NN or ball query operation to form hard neighborhoods, which may cross different semantic objects, resulting low-quality local features. To address this issue, we propose a multi-scale superpoint network that gradually generates multi-scale soft neighborhoods to extract geometric local features, thereby boosting the 3D semantic segmentation performance. Specifically, we present a simple yet efficient superpoint merging module that merge small-scale superpoints to obtain large-scale superpoint by considering the feature similarity of superpoints, so that we can obtain multi-scale geometric features of point clouds. We also develop a superpoint upsampling module that adopt inverse mapping function to propagate multi-scale features from low-resolution point cloud to high-resolution point cloud. By integrating our multi-scale superpoint network into a simple point based semantic segmentation network, our method can obtain SOTA results on S3DIS Area 5 and 6-fold, and competitive results on ScanNet v2. Ft Zheng, Le Hui, Jin Xie 0001, Haofeng Zhang 0001 |
MMAsia | 3 |
| 2023 | SE(3) Diffusion Model-based Point Cloud Registration for Robust 6D Object Pose EstimationabstractIn this paper, we introduce an SE(3) diffusion model-based point cloud registration framework for 6D object pose estimation in real-world scenarios. Our approach formulates the 3D registration task as a denoising diffusion process, which progressively refines the pose of the source point cloud to obtain a precise alignment with the model point cloud. Training our framework involves two operations: An SE(3) diffusion process and an SE(3) reverse process. The SE(3) diffusion process gradually perturbs the optimal rigid transformation of a pair of point clouds by continuously injecting noise (perturbation transformation). By contrast, the SE(3) reverse process focuses on learning a denoising network that refines the noisy transformation step-by-step, bringing it closer to the optimal transformation for accurate pose estimation. Unlike standard diffusion models used in linear Euclidean spaces, our diffusion model operates on the SE(3) manifold. This requires exploiting the linear Lie algebra $\mathfrak{se}(3)$ associated with SE(3) to constrain the transformation transitions during the diffusion and reverse processes. Additionally, to effectively train our denoising network, we derive a registration-specific variational lower bound as the optimization objective for model learning. Furthermore, we show that our denoising network can be constructed with a surrogate registration model, making our approach applicable to different deep registration networks. Extensive experiments demonstrate that our diffusion registration framework presents outstanding pose estimation performance on the real-world TUD-L, LINEMOD, and Occluded-LINEMOD datasets. Haobo Jiang, Mathieu Salzmann, Zheng Dang, Jin Xie 0001, Jian Yang 0003 |
NeurIPS | 4 |
| 2023 | Unsupervised Cross-Spectrum Depth Estimation by Visible-Light and Thermal CamerasabstractCross-spectrum depth estimation aims to provide a reliable depth map under variant-illumination conditions with a pair of dual-spectrum images. It is valuable for autonomous driving applications when vehicles are equipped with two cameras of different modalities. However, images captured by different-modality cameras can be photometrically quite different, which makes cross-spectrum depth estimation a very challenging problem. Moreover, the shortage of large-scale open-source datasets also retards further research in this field. In this paper, we propose an unsupervised visible light(VIS)-image-guided cross-spectrum (i.e., thermal and visible-light, TIR-VIS in short) depth-estimation framework. The input of the framework consists of a cross-spectrum stereo pair (one VIS image and one thermal image). First, we train a depth-estimation base network using VIS-image stereo pairs. To adapt the trained depth-estimation network to the cross-spectrum images, we propose a multi-scale feature-transfer network to transfer features from the TIR domain to the VIS domain at the feature level. Furthermore, we introduce a mechanism of cross-spectrum depth cycle-consistency to improve the depth estimation result of dual-spectrum image pairs. Meanwhile, we release to society a large cross-spectrum dataset with visible-light and thermal stereo images captured in different scenes. The experiment result shows that our method achieves better depth-estimation results than the compared existing methods. Our code and dataset are available onhttps://github.com/whitecrow1027/CrossSP_Depth. Yubin Guo, Xinlei Qi, Jin Xie 0001, Cheng-Zhong Xu 0001, Hui Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Geometry-Aware Network for Unsupervised Learning of Monocular Camera's Ego-MotionabstractDeep neural networks have been shown to be effective for unsupervised monocular visual odometry that can predict the camera’s ego-motion based on an input of monocular video sequence. However, most existing unsupervised monocular methods haven’t fully exploited the extracted information from both local geometric structure and visual appearance of the scenes, resulting in degraded performance. In this paper, a novel geometry-aware network is proposed to predict the camera’s ego-motion by learning representations in both 2D and 3D space. First, to extract geometry-aware features, we design an RGB-PointCloud feature fusion module to capture information from both geometric structure and the visual appearance of the scenes by fusing local geometric features from depth-map-derived point clouds and visual features from RGB images. Furthermore, the fusion module can adaptively allocate different weights to the two types of features to emphasize important regions. Then, we devise a relevant feature filtering module to build consistency between the two views and preserve informative features with high relevance. It can capture the correlation of frame pairs in the feature-embedding space by attention mechanisms. Finally, the obtained features are fed into the pose estimator to recover the 6-DoF poses of the camera. Extensive experiments show that our method achieves promising results among the unsupervised monocular deep learning methods on the KITTI odometry and TUM-RGBD datasets. Beibei Zhou, Jin Xie 0001, Zhong Jin, Hui Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Tube-Embedded Transformer for Pixel PredictionabstractMulti-task pixel-level learning, which aims to exploit the inter-task interactions to improve the learning of each task, is an important but challenging issue in visual perception and multimedia applications. Measuring the inter-task correlation and intra-task specificity, we propose a tube-embedded transformer (TET) framework for robust multi-task pixel prediction. To facilitate inter-task interactions, we aggregate and project all tasks into a shared tube pool to generate the latent multi-task representation during the coarse-to-fine decoding stages. The resulting task-tube interactions replace the two-by-two task-task interactions to reduce the model complexity significantly. In addition, we introduce the transformer mechanism to adaptively transfer tube features to the target task. Concretely, on the one hand, multi-task features aggregate in the tube to generate the shared feature representation bases; on the other hand, based on the task-tube association and complementarity, the tube outputs the query entry and the weighting coefficients of the target task. Experimentally, on the joint learning of semantic segmentation, depth estimation, and surface normal estimation, the comparison experiments show the superiority of the TET multi-task learning method over other state-of-the-art approaches, and the ablation experiments verify the effectiveness of the TET mechanism. Zhen Cui 0001, Zechao Li, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Multim. | 5 |
| 2022 | Reliable Inlier Evaluation for Unsupervised Point Cloud RegistrationabstractUnsupervised point cloud registration algorithm usually suffers from the unsatisfied registration precision in the partially overlapping problem due to the lack of effective inlier evaluation. In this paper, we propose a neighborhood consensus based reliable inlier evaluation method for robust unsupervised point cloud registration. It is expected to capture the discriminative geometric difference between the source neighborhood and the corresponding pseudo target neighborhood for effective inlier distinction. Specifically, our model consists of a matching map refinement module and an inlier evaluation module. In our matching map refinement module, we improve the point-wise matching map estimation by integrating the matching scores of neighbors into it. The aggregated neighborhood information potentially facilitates the discriminative map construction so that high-quality correspondences can be provided for generating the pseudo target point cloud. Based on the observation that the outlier has the significant structure-wise difference between its source neighborhood and corresponding pseudo target neighborhood while this difference for inlier is small, the inlier evaluation module exploits this difference to score the inlier confidence for each estimated correspondence. In particular, we construct an effective graph representation for capturing this geometric difference between the neighborhoods. Finally, with the learned correspondences and the corresponding inlier confidence, we use the weighted SVD algorithm for transformation estimation.Under the unsupervised setting, we exploit the Huber function based global alignment loss, the local neighborhood consensus loss and spatial consistency loss for model optimization. The experimental results on extensive datasets demonstrate that our unsupervised point cloud registration method can yield comparable performance. Yaqi Shen, Le Hui, Haobo Jiang, Jin Xie 0001, Jian Yang 0003 |
AAAI | 4 |
| 2022 | Domain Disentangled Generative Adversarial Network for Zero-Shot Sketch-Based 3D Shape RetrievalabstractSketch-based 3D shape retrieval is a challenging task due to the large domain discrepancy between sketches and 3D shapes. Since existing methods are trained and evaluated on the same categories, they cannot effectively recognize the categories that have not been used during training. In this paper, we propose a novel domain disentangled generative adversarial network (DD-GAN) for zero-shot sketch-based 3D retrieval, which can retrieve the unseen categories that are not accessed during training. Specifically, we first generate domain-invariant features and domain-specific features by disentangling the learned features of sketches and 3D shapes, where the domain-invariant features are used to align with the corresponding word embeddings. Then, we develop a generative adversarial network that combines the domain-specific features of the seen categories with the aligned domain-invariant features to synthesize samples, where the synthesized samples of the unseen categories are generated by using the corresponding word embeddings. Finally, we use the synthesized samples of the unseen categories combined with the real samples of the seen categories to train the network for retrieval, so that the unseen categories can be recognized. In order to reduce the domain shift problem, we utilize unlabeled unseen samples to enhance the discrimination ability of the discriminator. With the discriminator distinguishing the generated samples from the unlabeled unseen samples, the generator can generate more realistic unseen samples. Extensive experiments on the SHREC'13 and SHREC'14 datasets show that our method significantly improves the retrieval performance of the unseen categories. Rui Xu 0021, Zongyan Han, Le Hui, Jianjun Qian, Jin Xie 0001 |
AAAI | 5 |
| 2022 | Temporal-Aware Siamese Tracker: Integrate Temporal Context for 3D Object Tracking
Kaihao Lan, Haobo Jiang, Jin Xie 0001 |
ACCV (1) | 3 |
| 2022 | Learning Inter-superpoint Affinity for Weakly Supervised 3D Instance Segmentation
Linghua Tang, Le Hui, Jin Xie 0001 |
ACCV (1) | 3 |
| 2022 | Emphasizing Closeness and Diversity Simultaneously for Deep Face Representation
Chaoyu Zhao, Jianjun Qian, Shumin Zhu, Jin Xie 0001, Jian Yang 0003 |
ACCV (4) | 4 |
| 2022 | Generative Subgraph Contrast for Self-Supervised Graph Representation Learning
Yuehui Han, Le Hui, Haobo Jiang, Jianjun Qian, Jin Xie 0001 |
ECCV (30) | 5 |
| 2022 | RA-Depth: Resolution Adaptive Self-supervised Monocular Depth Estimation
Le Hui, Yikai Bian, Jin Xie 0001, Jian Yang 0003 |
ECCV (27) | 5 |
| 2022 | 3D Siamese Transformer Network for Single Object Tracking on Point Clouds
Le Hui, Lingpeng Wang, Linghua Tang, Kaihao Lan, Jin Xie 0001, Jian Yang 0003 |
ECCV (2) | 5 |
| 2022 | Globally Optimal Relative Pose Estimation for Multi-Camera Systems with Known Gravity DirectionabstractMultiple-camera systems have been widely used in self-driving cars, robots, and smartphones. In addition, they are typically also equipped with IMUs (inertial measurement units). Using the gravity direction extracted from the IMU data, the y-axis of the body frame of the multi-camera system can be aligned with this common direction, reducing the original three degree-of-freedom(DOF) relative rotation to a single DOF one. This paper presents a novel globally optimal solver to compute the relative pose of a generalized camera. Existing optimal solvers based on LM (Levenberg-Marquardt) method or SDP (semidefinite program) are either iterative or have high computational complexity. Our proposed optimal solver is based on minimizing the algebraic residual objective function. According to our derivation, using the least-squares algorithm, the original optimization problem can be converted into a system of two polynomials with only two variables. The proposed solvers have been tested on synthetic data and the KITTI benchmark. The experimental results show that the proposed methods have competitive robustness and accuracy compared with the existing state-of-the-art solvers. Qianliang Wu, Yaqing Ding 0001, Xinlei Qi, Jin Xie 0001, Jian Yang 0003 |
ICRA | 4 |
| 2022 | Unsupervised Domain Adaptation for Point Cloud Semantic Segmentation via Graph MatchingabstractUnsupervised domain adaptation for point cloud semantic segmentation has attracted great attention due to its effectiveness in learning with unlabeled data. Most of existing methods use global-level feature alignment to transfer the knowledge from the source domain to the target domain, which may cause the semantic ambiguity of the feature space. In this paper, we propose a graph-based framework to explore the local-level feature alignment between the two domains, which can reserve semantic discrimination during adaptation. Specifically, in order to extract local-level features, we first dynamically construct local feature graphs on both domains and build a memory bank with the graphs from the source domain. In particular, we use optimal transport to generate the graph matching pairs. Then, based on the assignment matrix, we can align the feature distributions between the two domains with the graph-based local feature loss. Furthermore, we consider the correlation between the features of different categories and formulate a category-guided contrastive loss to guide the segmentation model to learn discriminative features on the target domain. Extensive experiments on different synthetic-to-real and real-to-real domain adaptation scenarios demonstrate that our method can achieve state-of-the-art performance. Our code is available at https://github.com/BianYikai/PointUDA. Yikai Bian, Le Hui, Jianjun Qian, Jin Xie 0001 |
IROS | 4 |
| 2022 | Learning Superpoint Graph Cut for 3D Instance Segmentationabstract3D instance segmentation is a challenging task due to the complex local geometric structures of objects in point clouds. In this paper, we propose a learning-based superpoint graph cut method that explicitly learns the local geometric structures of the point cloud for 3D instance segmentation. Specifically, we first oversegment the raw point clouds into superpoints and construct the superpoint graph. Then, we propose an edge score prediction network to predict the edge scores of the superpoint graph, where the similarity vectors of two adjacent nodes learned through cross-graph attention in the coordinate and feature spaces are used for regressing edge scores. By forcing two adjacent nodes of the same instance to be close to the instance center in the coordinate and feature spaces, we formulate a geometry-aware edge loss to train the edge score prediction network. Finally, we develop a superpoint graph cut network that employs the learned edge scores and the predicted semantic classes of nodes to generate instances, where bilateral graph attention is proposed to extract discriminative features on both the coordinate and feature spaces for predicting semantic labels and scores of instances. Extensive experiments on two challenging datasets, ScanNet v2 and S3DIS, show that our method achieves new state-of-the-art performance on 3D instance segmentation. Le Hui, Linghua Tang, Yaqi Shen, Jin Xie 0001, Jian Yang 0003 |
NeurIPS | 4 |
| 2022 | Joint Optimal Transport With Convex Regularization for Robust Image ClassificationabstractThe critical step of learning the robust regression model from high-dimensional visual data is how to characterize the error term. The existing methods mainly employ the nuclear norm to describe the error term, which are robust against structure noises (e.g., illumination changes and occlusions). Although the nuclear norm can describe the structure property of the error term, global distribution information is ignored in most of these methods. It is known that optimal transport (OT) is a robust distribution metric scheme due to that it can handle correspondences between different elements in the two distributions. Leveraging this property, this article presents a novel robust regression scheme by integrating OT with convex regularization. The OT-based regression with$L_{2} $norm regularization (OTR) is first proposed to perform image classification. The alternating direction method of multipliers is developed to handle the model. To further address the occlusion problem in image classification, the extended OTR (EOTR) model is then presented by integrating the nuclear norm error term with an OTR model. In addition, we apply the alternating direction method of multipliers with Gaussian back substitution to solve EOTR and also provide the complexity and convergence analysis of our algorithms. Experiments were conducted on five benchmark datasets, including illumination changes and various occlusions. The experimental results demonstrate the performance of our robust regression model on biometric image classification against several state-of-the-art regression-based classification methods. Jianjun Qian, Wai Keung Wong, Hengmin Zhang, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Cybern. | 4 |
| 2022 | Efficient 3D Point Cloud Feature Learning for Large-Scale Place RecognitionabstractPoint cloud based retrieval for place recognition is still a challenging problem since the drastic appearance changes of scenes due to seasonal or artificial changes in the environments. Existing deep learning based global descriptors for the retrieval task usually consume a large amount of computational resources ( e.g ., memory), which may not be suitable for the cases of limited hardware resources. In this paper, we develop an efficient point cloud learning network (EPC-Net) to generate global descriptors of point clouds for place recognition. While obtaining good performance, it can greatly reduce computational memory and inference time. First, we propose a lightweight but effective neural network module, called ProxyConv, to aggregate the local geometric features of point clouds. We leverage the adjacency matrix and proxy points to simplify the original edge convolution for lower memory consumption. Then, we design a lightweight grouped VLAD network to form global descriptors for retrieval. Compared with the original VLAD network, we propose a grouped fully connected layer to decompose the high-dimensional vectors into a group of low-dimensional vectors, which can reduce the number of parameters of the network and maintain the discrimination of the feature vector. Finally, we further develop a simple version of EPC-Net, called EPC-Net-L, which consists of two ProxyConv modules and one max pooling layer to aggregate global descriptors. By distilling the knowledge from EPC-Net, EPC-Net-L can obtain discriminative global descriptors for retrieval. Extensive experiments on the Oxford dataset and three in-house datasets demonstrate that our method achieves good results with lower parameters, FLOPs, GPU memory, and shorter inference time. Our code is available at https://github.com/fpthink/EPC-Net. Le Hui, Mingmei Cheng, Jin Xie 0001, Jian Yang 0003, Ming-Ming Cheng |
IEEE Trans. Image Process. | 3 |
| 2022 | CBi-GNN: Cross-Scale Bilateral Graph Neural Network for 3D Object Detectionabstract3D object detection from LiDAR point clouds is a challenging task, since the point clouds are irregular and sparse. Existing one-stage methods mainly predict the 3D bounding box of 3D objects by extracting deep down-scaled features of point clouds from low-level (high-resolution, HR) feature maps to high-level (low-resolution, LR). Nonetheless, most of these methods ignore geometric context information of the down-scaled feature maps across scales, especially only using the LR feature will result in incomplete structure and less location accuracy of 3D objects. In this paper, we propose a novel cross-scale graph network-based one-stage 3D object detector to fully exploit the geometric contexts of the voxels between the down-scaled feature maps. Specifically, we first employ a 3D sparse convolution neural network to form different resolutions of feature maps of voxels. We then dynamically construct a cross-scale bilateral graph to search the neighbor non-empty voxels in the HR feature map with a fixed radius for each non-empty voxel in the LR feature map. In the constructed graph, we present a bilateral attention mechanism (i.e., self-attention and spatial attention) in the HR feature map and encode each non-empty voxel in the LR feature map by aggregating the HR features to obtain the attention features. In addition, we design a non-local part pooling operation to improve the score of the detected bounding box of 3D objects. Finally, we formulate a multi-task loss to train our network for regression of the 3D bounding box of the 3D objects. Experiments on the challenging KITTI’s 3D/BEV benchmark show that our proposed detector outperforms all one-stage 3D object detectors and is comparable to two-stage 3D object detectors. Our code is available athttps://github.com/csjxchen/CBi-GNN. Jiaxin Chen 0001, Xiang Li 0041, Jin Xie 0001, Jun Li 0027, Jianjun Qian, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | MSCFNet: A Lightweight Network With Multi-Scale Context Fusion for Real-Time Semantic SegmentationabstractIn recent years, how to strike a good trade-off between accuracy, inference speed, and model size has become the core issue for real-time semantic segmentation applications, which plays a vital role in real-world scenarios such as autonomous driving systems and drones. In this study, we devise a novel lightweight network using a multi-scale context fusion (MSCFNet) scheme, which explores an asymmetric encoder-decoder architecture to alleviate these problems. More specifically, the encoder adopts some developed efficient asymmetric residual (EAR) modules, which are composed of factorization depth-wise convolution and dilation convolution. Meanwhile, instead of complicated computation, simple deconvolution is applied in the decoder to further reduce the amount of parameters while still maintaining the high segmentation accuracy. Also, MSCFNet has branches with efficient attention modules from different stages of the network to well capture multi-scale contextual information. Then we combine them before the final classification to enhance the expression of the features and improve the segmentation efficiency. Comprehensive experiments on challenging datasets have demonstrated that the proposed MSCFNet, which contains only 1.15M parameters, achieves 71.9% Mean IoU on the Cityscapes testing dataset and can run at over 50 FPS on a single Titan XP GPU configuration. Guangwei Gao, Guoan Xu, Yi Yu 0001, Jin Xie 0001, Jian Yang 0003, Dong Yue 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Capitalizing on RGB-FIR Hybrid Imaging for Road DetectionabstractTraditionally, road detection approaches mostly capitalize on RGB images, 3D LiDAR point cloud or their fusion. However, RGB camera is sensitive to light conditions, while LiDAR point cloud is sparse compared with dense image pixels. In this work, a new hybrid image dataset is provided for the task of road detection based on cameras. In this dataset, the hybrid images are acquired by an optically aligned hybrid imaging device, consisting of a far-infrared (FIR) imager and an RGB camera to output pixel-wise registration of thermal and RGB frames. Then we investigate on three methods based on fully convolutional neural network (F-CNN) to demonstrate the advantages by fusing RGB-FIR images in road detection. First, a middle-fusion based model is built, where the output feature maps of encoder branches from RGB and FIR images are directly concatenated into a single-fusion branch as the decoder. Next, the originally discarded layers after fusion operation for both RGB and FIR branches are recovered as the mimic branches to imitate the distributions of the fusion outputs, which constitutes an extended cross model (ECM). Moreover, the outputs of mimic branches at different scales are also used to imitate the corresponding outputs in the fusion branch, called a hierarchical cross model (HCM). The experimental results demonstrate the effectiveness and efficiency of our fusion strategies. Yigong Zhang, Jin Xie 0001, José M. Álvarez 0004, Cheng-Zhong Xu 0001, Jian Yang 0003, Hui Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | SSPC-Net: Semi-supervised Semantic 3D Point Cloud Segmentation NetworkabstractPoint cloud semantic segmentation is a crucial task in 3D scene understanding. Existing methods mainly focus on employing a large number of annotated labels for supervised semantic segmentation. Nonetheless, manually labeling such large point clouds for the supervised segmentation task is time-consuming. In order to reduce the number of annotated labels, we propose a semi-supervised semantic point cloud segmentation network, named SSPC-Net, where we train the semantic segmentation network by inferring the labels of unlabeled points from the few annotated 3D points. In our method, we first partition the whole point cloud into superpoints and build superpoint graphs to mine the long-range dependencies in point clouds. Based on the constructed superpoint graph, we then develop a dynamic label propagation method to generate the pseudo labels for the unsupervised superpoints. Particularly, we adopt a superpoint dropout strategy to dynamically select the generated pseudo labels. In order to fully exploit the generated pseudo labels of the unsupervised superpoints, we furthermore propose a coupled attention mechanism for superpoint feature embedding. Finally, we employ the cross-entropy loss to train the semantic segmentation network with the labels of the supervised superpoints and the pseudo labels of the unsupervised superpoints. Experiments on various datasets demonstrate that our semisupervised segmentation method can achieve better performance than the current semi-supervised segmentation method with fewer annotated 3D points. Mingmei Cheng, Le Hui, Jin Xie 0001, Jian Yang 0003 |
AAAI | 3 |
| 2021 | Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action TasksabstractDouble Q-learning is a popular reinforcement learning algorithm in Markov decision process (MDP) problems. Clipped Double Q-learning, as an effective variant of Double Q-learning, employs the clipped double estimator to approximate the maximum expected action value. Due to the underestimation bias of the clipped double estimator, performance of clipped Double Q-learning may be degraded in some stochastic environments. In this paper, in order to reduce the underestimation bias, we propose an action candidate based clipped double estimator for Double Q-learning. Specifically, we first select a set of elite action candidates with the high action values from one set of estimators. Then, among these candidates, we choose the highest valued action from the other set of estimators. Finally, we use the maximum value in the second set of estimators to clip the action value of the chosen action in the first set of estimators and the clipped value is used for approximating the maximum expected action value. Theoretically, the underestimation bias in our clipped Double Q-learning decays monotonically as the number of the action candidates decreases. Moreover, the number of action candidates controls the trade-off between the overestimation and underestimation biases. In addition, we also extend our clipped Double Q-learning to continuous action tasks via approximating the elite continuous action candidates. We empirically verify that our algorithm can more accurately estimate the maximum expected action value on some toy environments and yield good performance on several benchmark problems. Haobo Jiang, Jin Xie 0001, Jian Yang 0003 |
AAAI | 2 |
| 2021 | Pyramid Point Cloud Transformer for Large-Scale Place RecognitionabstractRecently, deep learning based point cloud descriptors have achieved impressive results in the place recognition task. Nonetheless, due to the sparsity of point clouds, how to extract discriminative local features of point clouds to efficiently form a global descriptor is still a challenging problem. In this paper, we propose a pyramid point cloud transformer network (PPT-Net) to learn the discriminative global descriptors from point clouds for efficient retrieval. Specifically, we first develop a pyramid point transformer module that adaptively learns the spatial relationship of the different k-NN neighboring points of point clouds, where the grouped self-attention is proposed to extract discriminative local features of the point clouds. The grouped self-attention not only enhances long-term dependencies of the point clouds, but also reduces the computational cost. In order to obtain discriminative global descriptors, we construct a pyramid VLAD module to aggregate the multi-scale feature maps of point clouds into the global descriptors. By applying VLAD pooling on multi-scale feature maps, we utilize the context gating mechanism on the multiple global descriptors to adaptively weight the multi-scale global context information into the final global descriptor. Experimental results on the Oxford dataset and three in-house datasets show that our method achieves the state-of-the-art on the point cloud based place recognition task. Code is available at https://github.com/fpthink/PPT-Net. Le Hui, Mingmei Cheng, Jin Xie 0001, Jian Yang 0003 |
ICCV | 4 |
| 2021 | Superpoint Network for Point Cloud OversegmentationabstractSuperpoints are formed by grouping similar points with local geometric structures, which can effectively reduce the number of primitives of point clouds for subsequent point cloud processing. Existing superpoint methods mainly focus on employing clustering or graph partition to generate superpoints with handcrafted or learned features. Nonetheless, these methods cannot learn superpoints of point clouds with an end-to-end network. In this paper, we develop a new deep iterative clustering network to directly generate superpoints from irregular 3D point clouds in an end-to-end manner. Specifically, in our clustering network, we first jointly learn a soft point-superpoint association map from the coordinate and feature spaces of point clouds, where each point is assigned to the superpoint with a learned weight. Furthermore, we then iteratively update the association map and superpoint centers so that we can more accurately group the points into the corresponding superpoints with locally similar geometric structures. Finally, by predicting the pseudo labels of the superpoint centers, we formulate a label consistency loss on the points and superpoint centers to train the network. Extensive experiments on various datasets indicate that our method not only achieves the state-of-the-art on superpoint generation but also improves the performance of point cloud semantic segmentation. Code is available at https://github.com/fpthink/SPNet. Le Hui, Jia Yuan, Mingmei Cheng, Jin Xie 0001, Jian Yang 0003 |
ICCV | 4 |
| 2021 | Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud RegistrationabstractIn this paper, by modeling the point cloud registration task as a Markov decision process, we propose an end-to-end deep model embedded with the cross-entropy method (CEM) for unsupervised 3D registration. Our model consists of a sampling network module and a differentiable CEM module. In our sampling network module, given a pair of point clouds, the sampling network learns a prior sampling distribution over the transformation space. The learned sampling distribution can be used as a "good" initialization of the differentiable CEM module. In our differentiable CEM module, we first propose a maximum consensus criterion based alignment metric as the reward function for the point cloud registration task. Based on the reward function, for each state, we then construct a fused score function to evaluate the sampled transformations, where we weight the current and future rewards of the transformations. Particularly, the future rewards of the sampled transforms are obtained by performing the iterative closest point (ICP) algorithm on the transformed state. By selecting the top-k transformations with the highest scores, we iteratively update the sampling distribution. Furthermore, in order to make the CEM differentiable, we use the sparse-max function to replace the hard top-k selection. Finally, we formulate a Geman-McClure estimator based loss to train our end-to-end registration model. Extensive experimental results demonstrate the good registration performance of our method on benchmark datasets. Code is available at https://github.com/Jiang-HB/CEMNet. Haobo Jiang, Yaqi Shen, Jin Xie 0001, Jun Li 0027, Jianjun Qian, Jian Yang 0003 |
ICCV | 3 |
| 2021 | Planning with Learned Dynamic Model for Unsupervised Point Cloud RegistrationabstractPoint cloud registration is a fundamental problem in 3D computer vision. In this paper, we cast point cloud registration into a planning problem in reinforcement learning, which can seek the transformation between the source and target point clouds through trial and error. By modeling the point cloud registration process as a Markov decision process (MDP), we develop a latent dynamic model of point clouds, consisting of a transformation network and evaluation network. The transformation network aims to predict the new transformed feature of the point cloud after performing a rigid transformation (i.e., action) on it while the evaluation network aims to predict the alignment precision between the transformed source point cloud and target point cloud as the reward signal. Once the dynamic model of the point cloud is trained, we employ the cross-entropy method (CEM) to iteratively update the planning policy by maximizing the rewards in the point cloud registration process. Thus, the optimal policy, i.e., the transformation between the source and target point clouds, can be obtained via gradually narrowing the search space of the transformation. Experimental results on ModelNet40 and 7Scene benchmark datasets demonstrate that our method can yield good registration performance in an unsupervised manner. Haobo Jiang, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
IJCAI | 3 |
| 2021 | Transfer Vision Patterns for Multi-Task Pixel LearningabstractMulti-task pixel perception is one of the most important topics in the field of machine intelligence. Inspired by the observation of cross-task interdependencies of visual patterns, we propose a multi-task vision pattern transformation (VPT) method to adaptively correlate and transfer cross-task visual patterns by leveraging the powerful transformer mechanism. To better transfer visual patterns, specifically, we build two types of pattern transformation based on the statistic prior that the affinity relations across tasks are correlated. One aims to transfer feature patterns for the integration of different task features; the other aims to exchange structure patterns for mining and leveraging the latent interaction cues. These two types of transformations are encapsulated into two VPT units, which provide universal matching interfaces for multi-task learning, complement each other to guide the transmission of feature/structure patterns, and finally realize an adaptive selection of important patterns across tasks. Extensive experiments on the joint learning of semantic segmentation, depth prediction and surface normal estimation demonstrate that our proposed method is more effective than those baselines and achieve the state-of-that-art performance in three pixel-level visual tasks. Yong Li 0032, Zhen Cui 0001, Jin Xie 0001, Jian Yang 0003 |
ACM Multimedia | 5 |
| 2021 | SSPU-Net: Self-Supervised Point Cloud Upsampling via Differentiable RenderingabstractPoint clouds obtained from 3D sensors are usually sparse. Existing methods mainly focus on upsampling sparse point clouds in a supervised manner by using dense ground truth point clouds. In this paper, we propose a self-supervised point cloud upsampling network (SSPU-Net) to generate dense point clouds without using ground truth. To achieve this, we exploit the consistency between the input sparse point cloud and generated dense point cloud for the shapes and rendered images. Specifically, we first propose a neighbor expansion unit (NEU) to upsample the sparse point clouds, where the local geometric structures of the sparse point clouds are exploited to learn weights for point interpolation. Then, we develop a differentiable point cloud rendering unit (DRU) as an end-to-end module in our network to render the point cloud into multi-view images. Finally, we formulate a shape-consistent loss and an image-consistent loss to train the network so that the shapes of the sparse and dense point clouds are as consistent as possible. Extensive results on the CAD and scanned datasets demonstrate that our method can achieve impressive results in a self-supervised manner. Le Hui, Jin Xie 0001 |
ACM Multimedia | 3 |
| 2021 | 3D Siamese Voxel-to-BEV Tracker for Sparse Point Cloudsabstract3D object tracking in point clouds is still a challenging problem due to the sparsity of LiDAR points in dynamic environments. In this work, we propose a Siamese voxel-to-BEV tracker, which can significantly improve the tracking performance in sparse 3D point clouds. Specifically, it consists of a Siamese shape-aware feature learning network and a voxel-to-BEV target localization network. The Siamese shape-aware feature learning network can capture 3D shape information of the object to learn the discriminative features of the object so that the potential target from the background in sparse point clouds can be identified. To this end, we first perform template feature embedding to embed the template's feature into the potential target and then generate a dense 3D shape to characterize the shape information of the potential target. For localizing the tracked target, the voxel-to-BEV target localization network regresses the target's 2D center and the z-axis center from the dense bird's eye view (BEV) feature map in an anchor-free manner. Concretely, we compress the voxelized point cloud along z-axis through max pooling to obtain a dense BEV feature map, where the regression of the 2D center and the z-axis center can be performed more effectively. Extensive evaluation on the KITTI tracking dataset shows that our method significantly outperforms the current state-of-the-art methods by a large margin. Code is available at https://github.com/fpthink/V2B. Le Hui, Lingpeng Wang, Mingmei Cheng, Jin Xie 0001, Jian Yang 0003 |
NeurIPS | 4 |
| 2021 | Improving Unsupervised Learning of Monocular Depth and Ego-Motion via Stereo Network
Jin Xie 0001, Jian Yang 0003 |
PRCV (2) | 2 |
| 2021 | Facilitating 3D Object Tracking in Point Clouds with Image Semantics and Geometry
Lingpeng Wang, Le Hui, Jin Xie 0001 |
PRCV (1) | 3 |
| 2021 | Constructing multilayer locality-constrained matrix regression framework for noise robust face super-resolution
Guangwei Gao, Yi Yu 0001, Jin Xie 0001, Jian Yang 0003, Meng Yang 0001, Jian Zhang 0002 |
Pattern Recognit. | 3 |
| 2020 | Sequential 3D Human Pose and Shape Estimation From Point CloudsabstractThis work addresses the problem of 3D human pose and shape estimation from a sequence of point clouds. Existing sequential 3D human shape estimation methods mainly focus on the template model fitting from a sequence of depth images or the parametric model regression from a sequence of RGB images. In this paper, we propose a novel sequential 3D human pose and shape estimation framework from a sequence of point clouds. Specifically, the proposed framework can regress 3D coordinates of mesh vertices at different resolutions from the latent features of point clouds. Based on the estimated 3D coordinates and features at the low resolution, we develop a spatial-temporal mesh attention convolution (MAC) to predict the 3D coordinates of mesh vertices at the high resolution. By assigning specific attentional weights to different neighboring points in the spatial and temporal domains, our spatial-temporal MAC can capture structured spatial and temporal features of point clouds. We further generalize our framework to the real data of human bodies with a weakly supervised fine-tuning method. The experimental results on SURREAL, Human3.6M, DFAUST and the real detailed data demonstrate that the proposed approach can accurately recover the 3D body model sequence from a sequence of point clouds. Kangkan Wang, Jin Xie 0001, Guofeng Zhang 0001, Jian Yang 0003 |
CVPR | 2 |
| 2020 | Progressive Point Cloud Deconvolution Generation Network
Le Hui, Rui Xu 0021, Jin Xie 0001, Jianjun Qian, Jian Yang 0003 |
ECCV (15) | 3 |
| 2020 | Cascaded Non-local Neural Network for Point Cloud Semantic SegmentationabstractIn this paper, we propose a cascaded non-local neural network for point cloud segmentation. The proposed network aims to build the long-range dependencies of point clouds for the accurate segmentation. Specifically, we develop a novel cascaded non-local module, which consists of the neighborhood-level, superpoint-level and global-level non-local blocks. First, in the neighborhood-level block, we extract the local features of the centroid points of point clouds by assigning different weights to the neighboring points. The extracted local features of the centroid points are then used to encode the superpoint-level block with the non-local operation. Finally, the global-level block aggregates the non-local features of the superpoints for semantic segmentation in an encoder-decoder framework. Benefiting from the cascaded structure, geometric structure information of different neighborhoods with the same label can be propagated. In addition, the cascaded structure can largely reduce the computational cost of the original non-local operation on point clouds. Experiments on different indoor and outdoor datasets show that our method achieves state-of-the-art performance and effectively reduces the time consumption and memory occupation. Mingmei Cheng, Le Hui, Jin Xie 0001, Jian Yang 0003, Hui Kong 0001 |
IROS | 3 |
| 2020 | PUI-Net: A Point Cloud Upsampling and Inpainting Network
Jin Xie 0001, Jianjun Qian, Jian Yang 0003 |
PRCV (1) | 2 |
| 2020 | Human-in-the-loop image segmentation and annotation
Lianjie Wang, Jin Xie 0001, Pengfei Zhu 0001 |
Sci. China Inf. Sci. | 3 |
| 2020 | Image decomposition based matrix regression with applications to robust face recognition
Jianjun Qian, Jian Yang 0003, Yong Xu 0001, Jin Xie 0001, Zhihui Lai 0001, Bob Zhang 0001 |
Pattern Recognit. | 4 |
| 2019 | Deep Sketch-Shape Hashing With Segmented 3D Stochastic ViewingabstractSketch-based 3D shape retrieval has been extensively studied in recent works, most of which focus on improving the retrieval accuracy, whilst neglecting the efficiency. In this paper, we propose a novel framework for efficient sketch-based 3D shape retrieval, i.e., Deep Sketch-Shape Hashing (DSSH), which tackles the challenging problem from two perspectives. Firstly, we propose an intuitive 3D shape representation method to deal with unaligned shapes with arbitrary poses. Specifically, the proposed Segmented Stochastic-viewing Shape Network models discriminative 3D representations by a set of 2D images rendered from multiple views, which are stochastically selected from non-overlapping spatial segments of a 3D sphere. Secondly, Batch-Hard Binary Coding (BHBC) is developed to learn semantics-preserving compact binary codes by mining the hardest samples. The overall framework is jointly learned by developing an alternating iteration algorithm. Extensive experimental results on three benchmarks show that DSSH improves both the retrieval efficiency and accuracy remarkably, compared to the state-of-the-art methods. Jiaxin Chen 0002, Jie Qin 0004, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Jin Xie 0001, Ling Shao 0001 |
CVPR | 6 |
| 2019 | Assignment Problem Based Deep Embedding
Ruishen Zheng, Jin Xie 0001, Jianjun Qian, Jian Yang 0003 |
PRCV (2) | 2 |
| 2018 | Learning Adversarial 3D Model Generation With 2D Image EnhancerabstractRecent advancements in generative adversarial nets (GANs) and volumetric convolutional neural networks (CNNs) enable generating 3D models from a probabilistic space. In this paper, we have developed a novel GAN-based deep neural network to obtain a better latent space for the generation of 3D models. In the proposed method, an enhancer neural network is introduced to extract information from other corresponding domains (e.g. image) to improve the performance of the 3D model generator, and the discriminative power of the unsupervised shape features learned from the 3D model discriminator. Specifically, we train the 3D generative adversarial networks on 3D volumetric models, and at the same time, the enhancer network learns image features from rendered images. Different from the traditional GAN architecture that uses uninformative random vectors as inputs, we feed the high-level image features learned from the enhancer into the 3D model generator for better training. The evaluations on two large-scale 3D model datasets, ShapeNet and ModelNet, demonstrate that our proposed method can not only generate high-quality 3D models, but also successfully learn discriminative shape representation for classification and retrieval without supervision. Jing Zhu 0002, Jin Xie 0001, Yi Fang 0006 |
AAAI | 2 |
| 2018 | Structure-Aware 3D Shape Synthesis from Single-View Images
Xuyang Hu, Fan Zhu 0001, Li Liu 0004, Jin Xie 0001, Jun Tang 0007, Nian Wang 0002, Fumin Shen, Ling Shao 0001 |
BMVC | 4 |
| 2018 | Siamese CNN-BiLSTM Architecture for 3D Shape Representation LearningabstractLearning a 3D shape representation from a collection of its rendered 2D images has been extensively studied. However, existing view-based techniques have not yet fully exploited the information among all the views of projections. In this paper, by employing recurrent neural network to efficiently capture features across different views, we propose a siamese CNN-BiLSTM network for 3D shape representation learning. The proposed method minimizes a discriminative loss function to learn a deep nonlinear transformation, mapping 3D shapes from the original space into a nonlinear feature space. In the transformed space, the distance of 3D shapes with the same label is minimized, otherwise the distance is maximized to a large margin. Specifically, the 3D shapes are first projected into a group of 2D images from different views. Then convolutional neural network (CNN) is adopted to extract features from different view images, followed by a bidirectional long short-term memory (LSTM) to aggregate information across different views. Finally, we construct the whole CNN-BiLSTM network into a siamese structure with contrastive loss function. Our proposed method is evaluated on two benchmarks, ModelNet40 and SHREC 2014, demonstrating superiority over the state-of-the-art methods. Guoxian Dai, Jin Xie 0001, Yi Fang 0006 |
IJCAI | 2 |
| 2018 | Learning to Synthesize 3D Indoor Scenes from Monocular ImagesabstractDepth images have always been playing critical roles for indoor scene understanding problems, and are particularly important for tasks in which 3D inferences are involved. However, since depth images are not universally available, abandoning them from the testing stage can significantly improve the generality of a method. In this work, we consider the scenarios where depth images are not available in the testing data, and propose to learn a convolutional long short-term memory (Conv LSTM) network and a regression convolutional neural network (regression ConvNet) using only monocular RGB images. The proposed networks benefit from 2D segmentations, object-level spatial context, object-scene dependencies and objects' geometric information, where optimization is governed by the semantic label loss, which measures the label consistencies of both objects and scenes, and the 3D geometrical loss, which measures the correctness of objects' 6Dof estimation. Conv LSTM and regression ConvNet are applied to scene/object classification, object detection and 6Dof estimation tasks respectively, where we utilize the joint inference from both networks and further provide the perspective of synthesizing fully rigged 3D scenes according to objects' arrangements in monocular images. Both quantitative and qualitative experimental results are provided on the NYU-v2 dataset, and we demonstrate that the proposed Conv LSTM can achieve state-of-the-art performance without requiring the depth information. Fan Zhu 0001, Li Liu 0004, Jin Xie 0001, Fumin Shen, Ling Shao 0001, Yi Fang 0006 |
ACM Multimedia | 3 |
| 2018 | Episode-Experience Replay Based Tree-Backup Method for Off-Policy Actor-Critic Algorithm
Haobo Jiang, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
PRCV (1) | 3 |
| 2018 | Deep Nonlinear Metric Learning for 3-D Shape RetrievalabstractEffective 3-D shape retrieval is an important problem in 3-D shape analysis. Recently, feature learning-based shape retrieval methods have been widely studied, where the distance metrics between 3-D shape descriptors are usually hand-crafted. In this paper, motivated by the fact that deep neural network has the good ability to model nonlinearity, we propose to learn an effective nonlinear distance metric between 3-D shape descriptors for retrieval. First, the locality-constrained linear coding method is employed to encode each vertex on the shape and the encoding coefficient histogram is formed as the global 3-D shape descriptor to represent the shape. Then, a novel deep metric network is proposed to learn a nonlinear transformation to map the 3-D shape descriptors to a nonlinear feature space. The proposed deep metric network minimizes a discriminative loss function that can enforce the similarity between a pair of samples from the same class to be small and the similarity between a pair of samples from different classes to be large. Finally, the distance between the outputs of the metric network is used as the similarity for shape retrieval. The proposed method is evaluated on the McGill, SHREC'10 ShapeGoogle, and SHREC'14 Human shape datasets. Experimental results on the three datasets validate the effectiveness of the proposed method. Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Ling Shao 0001, Yi Fang 0006 |
IEEE Trans. Cybern. | 1 |
| 2018 | Deep Correlated Holistic Metric Learning for Sketch-Based 3D Shape RetrievalabstractHow to effectively retrieve desired 3D models with simple queries is a long-standing problem in computer vision community. The model-based approach is quite straightforward but nontrivial, since people could not always have the desired 3D query model available by side. Recently, large amounts of wide-screen electronic devices are prevail in our daily lives, which makes the sketch-based 3D shape retrieval a promising candidate due to its simpleness and efficiency. The main challenge of sketch-based approach is the huge modality gap between sketch and 3D shape. In this paper, we proposed a novel deep correlated holistic metric learning (DCHML) method to mitigate the discrepancy between sketch and 3D shape domains. The proposed DCHML trains two distinct deep neural networks (one for each domain) jointly, which learns two deep nonlinear transformations to map features from both domains into a new feature space. The proposed loss, including discriminative loss and correlation loss, aims to increase the discrimination of features within each domain as well as the correlation between different domains. In the new feature space, the discriminative loss minimizes the intra-class distance of the deep transformed features and maximizes the inter-class distance of the deep transformed features to a large margin within each domain, while the correlation loss focused on mitigating the distribution discrepancy across different domains. Different from existing deep metric learning methods only with loss at the output layer, our proposed DCHML is trained with loss at both hidden layer and output layer to further improve the performance by encouraging features in the hidden layer also with desired properties. Our proposed method is evaluated on three benchmarks, including 3D Shape Retrieval Contest 2013, 2014, and 2016 benchmarks, and the experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods. Guoxian Dai, Jin Xie 0001, Yi Fang 0006 |
IEEE Trans. Image Process. | 2 |
| 2017 | Deep Correlated Metric Learning for Sketch-based 3D Shape RetrievalabstractThe explosive growth of 3D models has led to the pressing demand for an efficient searching system. Traditional model-based search is usually not convenient, since people don't always have 3D model available by side. The sketch-based 3D shape retrieval is a promising candidate due to its simpleness and efficiency. The main challenge for sketch-based 3D shape retrieval is the discrepancy across different domains. In the paper, we propose a novel deep correlated metric learning (DCML) method to mitigate the discrepancy between sketch and 3D shape domains. The proposed DCML trains two distinct deep neural networks (one for each domain) jointly with one loss, which learns two deep nonlinear transformations to map features from both domains into a nonlinear feature space. The proposed loss, including discriminative loss and correlation loss, aims to increase the discrimination of features within each domain as well as the correlation between different domains. In the transfered space, the discriminative loss minimizes the intra-class distance of the deep transformed features and maximizes the inter-class distance of the deep transformed features at least a predefined margin within each domain, while the correlation loss focuses on minimizing the distribution discrepancy across different domains. Our proposed method is evaluated on SHREC 2013 and 2014 benchmarks, and the experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods. Guoxian Dai, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006 |
AAAI | 2 |
| 2017 | Learning Barycentric Representations of 3D Shapes for Sketch-Based 3D Shape RetrievalabstractRetrieving 3D shapes with sketches is a challenging problem since 2D sketches and 3D shapes are from two heterogeneous domains, which results in large discrepancy between them. In this paper, we propose to learn barycenters of 2D projections of 3D shapes for sketch-based 3D shape retrieval. Specifically, we first use two deep convolutional neural networks (CNNs) to extract deep features of sketches and 2D projections of 3D shapes. For 3D shapes, we then compute the Wasserstein barycenters of deep features of multiple projections to form a barycentric representation. Finally, by constructing a metric network, a discriminative loss is formulated on the Wasserstein barycenters of 3D shapes and sketches in the deep feature space to learn discriminative and compact 3D shape and sketch features for retrieval. The proposed method is evaluated on the SHREC13 and SHREC14 sketch track benchmark datasets. Compared to the state-of-the-art methods, our proposed method can significantly improve the retrieval performance. Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Yi Fang 0006 |
CVPR | 1 |
| 2017 | Metric-based Generative Adversarial NetworkabstractExisting methods of generative adversarial network (GAN) use different criteria to distinguish between real and fake samples, such as probability [9],energy [44] energy or other losses [30]. In this paper, by employing the merits of deep metric learning, we propose a novel metric-based generative adversarial network (MBGAN), which uses the distance-criteria to distinguish between real and fake samples. Specifically, the discriminator of MBGAN adopts a triplet structure and learns a deep nonlinear transformation, which maps input samples into a new feature space. In the transformed space, the distance between real samples is minimized, while the distance between real sample and fake sample is maximized. Similar to the adversarial procedure of existing GANs, a generator is trained to produce synthesized examples, which are close to real examples, while a discriminator is trained to maximize the distance between real and fake samples to a large margin. Meanwhile, instead of using a fixed margin, we adopt a data-dependent margin [30], so that the generator could focus on improving the synthesized samples with poor quality, instead of wasting energy on well-produce samples. Our proposed method is verified on various benchmarks, such as CIFAR-10, SVHN and CelebA, and generates high-quality samples. Guoxian Dai, Jin Xie 0001, Yi Fang 0006 |
ACM Multimedia | 2 |
| 2017 | DeepShape: Deep-Learned Shape Descriptor for 3D Shape RetrievalabstractComplex geometric variations of 3D models usually pose great challenges in 3D shape matching and retrieval. In this paper, we propose a novel 3D shape feature learning method to extract high-level shape features that are insensitive to geometric deformations of shapes. Our method uses a discriminative deep auto-encoder to learn deformation-invariant shape features. First, a multiscale shape distribution is computed and used as input to the auto-encoder. We then impose the Fisher discrimination criterion on the neurons in the hidden layer to develop a deep discriminative auto-encoder. Finally, the outputs from the hidden layers of the discriminative auto-encoders at different scales are concatenated to form the shape descriptor. The proposed method is evaluated on four benchmark datasets that contain 3D models with large geometric variations: McGill, SHREC'10 ShapeGoogle, SHREC'14 Human and SHREC'14 Large Scale Comprehensive Retrieval Track Benchmark datasets. Experimental results on the benchmark datasets demonstrate the effectiveness of the proposed method for 3D shape retrieval. Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Edward K. Wong, Yi Fang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Progressive Shape-Distribution-Encoder for Learning 3D Shape RepresentationabstractSince there are complex geometric variations with 3D shapes, extracting efficient 3D shape features is one of the most challenging tasks in shape matching and retrieval. In this paper, we propose a deep shape descriptor by learning shape distributions at different diffusion time via a progressive shape-distribution-encoder (PSDE). First, we develop a shape distribution representation with the kernel density estimator to characterize the intrinsic geometry structures of 3D shapes. Then, we propose to learn a deep shape feature through an unsupervised PSDE. Specially, the unsupervised PSDE aims at modeling the complex non-linear transform of the estimated shape distributions between consecutive diffusion time. In order to characterize the intrinsic structures of 3D shapes more efficiently, we stack multiple PSDEs to form a network structure. Finally, we concatenate all neurons in the middle hidden layers of the unsupervised PSDE network to form an unsupervised shape descriptor for retrieval. Furthermore, by imposing an additional constraint on the outputs of all hidden layers, we propose a supervised PSDE to form a supervised shape descriptor. For each hidden layer, the similarity between a pair of outputs from the same class is as large as possible and the similarity between a pair of outputs from different classes is as small as possible. The proposed method is evaluated on three benchmark 3D shape data sets with large geometric variations, i.e., McGill, SHREC'10 ShapeGoogle, and SHREC'14 Human data sets, and the experimental results demonstrate the superiority of the proposed method to the existing approaches. Jin Xie 0001, Fan Zhu 0001, Guoxian Dai, Ling Shao 0001, Yi Fang 0006 |
IEEE Trans. Image Process. | 1 |
| 2017 | Deep Multimetric Learning for Shape-Based 3D Model RetrievalabstractRecently, feature-learning-based 3D shape retrieval methods have been receiving more and more attention in the 3D shape analysis community. In these methods, the hand-crafted metrics or the learned linear metrics are usually used to compute the distances between shape features. Since there are complex geometric structural variations with 3D shapes, the single hand-crafted metric or learned linear metric cannot characterize the manifold, where 3D shapes lie well. In this paper, by exploring the nonlinearity of the deep neural network and the complementarity among multiple shape features, we propose a novel deep multimetric network for 3D shape retrieval. The developed multimetric network minimizes a discriminative loss function that, for each type of shape feature, the outputs of the network from the same class are encouraged to be as similar as possible and the outputs from different classes are encouraged to be as dissimilar as possible. Meanwhile, the Hilbert-Schmidt independence criterion is employed to enforce the outputs of different types of shape features to be as complementary as possible. Furthermore, the weights of the learned multiple distance metrics can be adaptively determined in our developed deep metric network. The weighted distance metric is then used as the similarity for shape retrieval. We conduct experiments with the proposed method on the four benchmark shape datasets. Experimental results demonstrate that the proposed method can obtain better performance than the learned deep single metric and outperform the state-of-the-art 3D shape retrieval methods. Jin Xie 0001, Guoxian Dai, Yi Fang 0006 |
IEEE Trans. Multim. | 1 |
| 2016 | Learning Cross-Domain Neural Networks for Sketch-Based 3D Shape RetrievalabstractSketch-based 3D shape retrieval, which returns a set of relevant 3D shapes based on users' input sketch queries, has been receiving increasing attentions in both graphics community and vision community. In this work, we address the sketch-based 3D shape retrieval problem with a novel Cross-Domain Neural Networks (CDNN) approach, which is further extended to Pyramid Cross-Domain Neural Networks (PCDNN) by cooperating with a hierarchical structure. In order to alleviate the discrepancies between sketch features and 3D shape features, a neural network pair that forces identical representations at the target layer for instances of the same class is trained for sketches and 3D shapes respectively. By constructing cross-domain neural networks at multiple pyramid levels, a many-to-one relationship is established between a 3D shape feature and sketch features extracted from different scales. We evaluate the effectiveness of both CDNN and PCDNN approach on the extended large-scale SHREC 2014 benchmark and compare with some other well established methods. Experimental results suggest that both CDNN and PCDNN can outperform state-of-the-art performance, where PCDNN can further improve CDNN when employing a hierarchical structure. Fan Zhu 0001, Jin Xie 0001, Yi Fang 0006 |
AAAI | 2 |
| 2016 | Learned Binary Spectral Shape Descriptor for 3D Shape CorrespondenceabstractDense 3D shape correspondence is an important problem in computer vision and computer graphics. Recently, the local shape descriptor based 3D shape correspondence approaches have been widely studied, where the local shape descriptor is a real-valued vector to characterize the geometrical structure of the shape. Different from these realvalued local shape descriptors, in this paper, we propose to learn a novel binary spectral shape descriptor with the deep neural network for 3D shape correspondence. The binary spectral shape descriptor can require less storage space and enable fast matching. First, based on the eigenvectors of the Laplace-Beltrami operator, we construct a neural network to form a nonlinear spectral representation to characterize the shape. Then, for the defined positive and negative points on the shapes, we train the constructed neural network by minimizing the errors between the outputs and their corresponding binary descriptors, minimizing the variations of the outputs of the positive points and maximizing the variations of the outputs of the negative points, simultaneously. Finally, we binarize the output of the neural network to form the binary spectral shape descriptor for shape correspondence. The proposed binary spectral shape descriptor is evaluated on the SCAPE and TOSCA 3D shape datasets for shape correspondence. The experimental results demonstrate the effectiveness of the proposed binary shape descriptor for the shape correspondence task. Jin Xie 0001, Meng Wang 0001, Yi Fang 0006 |
CVPR | 1 |
| 2016 | Heat Diffusion Long-Short Term Memory Learning for 3D Shape Analysis
Fan Zhu 0001, Jin Xie 0001, Yi Fang 0006 |
ECCV (7) | 2 |
| 2016 | Dynamic texture recognition with video set based collaborative representation
Jin Xie 0001, Yi Fang 0006 |
Image Vis. Comput. | 1 |
| 2016 | From handcrafted to learned representations for human action recognition: A survey
Fan Zhu 0001, Ling Shao 0001, Jin Xie 0001, Yi Fang 0006 |
Image Vis. Comput. | 3 |
| 2016 | Learning a discriminative deformation-invariant 3D shape descriptor via many-to-one encoder
Guoxian Dai, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006 |
Pattern Recognit. Lett. | 2 |
| 2016 | Linear discrimination dictionary learning for shape descriptors
Meng Wang 0001, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006 |
Pattern Recognit. Lett. | 2 |
| 2015 | 3D deep shape descriptorabstractShape descriptor is a concise yet informative representation that provides a 3D object with an identification as a member of some category. We have developed a concise deep shape descriptor to address challenging issues from ever-growing 3D datasets in areas as diverse as engineering, medicine, and biology. Specifically, in this paper, we developed novel techniques to extract concise but geometrically informative shape descriptor and new methods of defining Eigen-shape descriptor and Fisher-shape descriptor to guide the training of a deep neural network. Our deep shape descriptor tends to maximize the inter-class margin while minimize the intra-class variance. Our new shape descriptor addresses the challenges posed by the high complexity of 3D model and data representation, and the structural variations and noise present in 3D models. Experimental results on 3D shape retrieval demonstrate the superior performance of deep shape descriptor over other state-of-the-art techniques in handling noise, incompleteness and structural variations. Yi Fang 0006, Jin Xie 0001, Guoxian Dai, Meng Wang 0001, Fan Zhu 0001, Edward K. Wong |
CVPR | 2 |
| 2015 | Deepshape: Deep learned shape descriptor for 3D shape matching and retrievalabstractComplex geometric structural variations of 3D model usually pose great challenges in 3D shape matching and retrieval. In this paper, we propose a high-level shape feature learning scheme to extract features that are insensitive to deformations via a novel discriminative deep auto-encoder. First, a multiscale shape distribution is developed for use as input to the auto-encoder. Then, by imposing the Fisher discrimination criterion on the neurons in the hidden layer, we developed a novel discriminative deep auto-encoder for shape feature learning. Finally, the neurons in the hidden layers from multiple discriminative auto-encoders are concatenated to form a shape descriptor for 3D shape matching and retrieval. The proposed method is evaluated on the representative datasets that contain 3D models with large geometric variations, i.e., Mcgill and SHREC'10 ShapeGoogle datasets. Experimental results on the benchmark datasets demonstrate the effectiveness of the proposed method for 3D shape matching and retrieval. Jin Xie 0001, Yi Fang 0006, Fan Zhu 0001, Edward K. Wong |
CVPR | 1 |
| 2015 | Progressive Shape-Distribution-Encoder for 3D Shape RetrievalabstractIn this paper, we propose a deep shape descriptor by learning the shape distributions at different diffusion time via a progressive deep shape-distribution-encoder. First, we develop a shape distribution representation with the kernel density estimator to characterize the intrinsic geometrical structure of the shape. Then, we propose to learn discriminative shape features through a progressive shape-distribution-encoder. Specially, the progressive shape-distribution-encoder aims at modeling the complex non-linear transform of the estimated shape distributions between consecutive diffusion time. Furthermore, in order to characterize the intrinsic structure of the shape more efficiently, we stack multiple proposed progressive shape-distribution-encoders to form a neural network structure. Finally, we concatenated all neurons in the hidden layers of the progressive shape-distribution-encoder network to form a discriminative shape descriptor for retrieval. The proposed method is evaluated on three benchmark 3D shape datasets %with large geometric variations, i.e., McGill, SHREC'10 ShapeGoogle and SHREC'14 Human datasets, and the experimental results demonstrate the superiority of our method to the existing approaches. Jin Xie 0001, Fan Zhu 0001, Guoxian Dai, Yi Fang 0006 |
ACM Multimedia | 1 |