EDBT 2026 Demo / reviewers in the wild / expert
Xiang Li 0046
dblp:40/1491-46
· DBLP profile ↗
30ranked-venue papers
9as first author
23since 2021 · last 2025
0000-0002-9946-7000ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3DCoMPaT++: An Improved Large-Scale 3D Vision Dataset for Compositional RecognitionabstractIn this work, we present 3DCOMPAT++, a multimodal 2D/3D dataset with 160 million rendered views of more than 10 million stylized 3D shapes carefully annotated at the partinstance level, alongside matching RGB point clouds, 3D textured meshes, depth maps, and segmentation masks. 3DCOMPAT ++ covers 42 shape categories, 275 fine-grained part categories, and 293 fine-grained material classes that can be compositionally applied to parts of 3D objects. We render a subset of one million stylized shapes from four equally spaced views as well as four randomized views, leading to a total of 160 million renderings. Parts are segmented at the instance level, with coarse-grained and fine-grained semantic levels. We introduce a new task, called Grounded CoMPaT Recognition (GCR), to collectively recognize and ground compositions of materials on parts of 3D objects. Additionally, we report the outcomes of a data challenge organized at the CVPR conference, showcasing the winning method's utilization of a modified PointNet++model trained on 6D inputs, and exploring alternative techniques for GCR enhancement. We hope our work will help ease future research on compositional 3D Vision. The dataset and code have been made publicly available at https://3dcompat-dataset.org/v2/.3D vision, dataset, 3D modeling, multimodal learning, compositional learning. Habib Slim, Xiang Li 0046, Yuchen Li 0010, Mohamed Ayman, Ujjwal Upadhyay, Ahmed Abdelreheem 0002, Arpit Prajapati, Suhail Pothigara, Peter Wonka, Mohamed Elhoseiny 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Few-Shot Oriented Object Detection in Remote Sensing Images via Memorable Contrastive LearningabstractFew-shot object detection (FSOD) has attracted significant research attention in remote sensing due to its potential to reduce reliance on large annotated datasets. However, two challenges remain in this area: (1) axis-aligned proposals, which can result in misalignment for arbitrarily oriented objects, and (2) object misclassification due to limited annotated data, which hinders generalization to unseen classes. To address these issues, we propose a novel method for few-shot oriented object detection in remote sensing images. Our approach employs oriented bounding boxes instead of horizontal ones to learn more effective feature representations for arbitrarily oriented aerial objects, enhancing detection accuracy. Additionally, we introduce a supervised contrastive learning module with a dynamically updated memory bank, enabling the model to leverage large batches of negative samples and to better learn discriminative features for unseen classes. Extensive experiments on DOTA, HRSC2016, and DIOR-R datasets demonstrate superior performance of our proposed method in few-shot oriented object detection. Code and pre-trained models will be made publicly available. Jiawei Zhou 0009, Wuzhou Li, Hongtao Cai, Tianjin Huang, Gui-Song Xia, Xiang Li 0046 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | 3D Shape Contrastive Representation Learning With Adversarial ExamplesabstractCurrent supervised methods for 3D shape representation learning have achieved satisfying performance, yet require extensive human-labeled datasets. Unsupervised learning-based methods provide a viable solution by learning shape representations without using ground truth labels. In this study, we develop a contrastive learning framework for unsupervised representation learning of 3D shapes. Specifically, in order to encourage models to pay more attention to useful information during representation learning, we first introduce a new paradigm for critical points search based on the adversarial mechanism. We extract critical points with a larger impact on the global feature by attacking a pre-trained auto-encoder model, and apply data augmentations on these points to generate adversarial examples. Taking a pair of adversarial examples as inputs, we obtain their intermediate embeddings and global representations of corresponding inputs, which are then transformed into latent spaces by two predictor heads. Finally, we train the proposed model by maximizing the agreements on these latent spaces via Normalized Temperature-scaled Cross Entropy (NT-Xent) loss and a newly designed Cross-layer Normalized Temperature-scaled Cross Entropy (Cross-NT-Xent) loss, where the latter is proposed in this paper to enforce cross-layer feature similarities. The effectiveness, robustness, and transferability of learned representations are validated on three downstream tasks, including object classification, few-shot classification, and shape retrieval. Experiments on three benchmark datasets show that our learned representations achieve better or competitive performance than current state-of-the-art methods in these downstream tasks. Moreover, our model can easily be extended to 3D part segmentation and scene segmentation tasks. Congcong Wen, Xiang Li 0046, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006 |
IEEE Trans. Multim. | 2 |
| 2024 | Uni3DL: A Unified Model for 3D Vision-Language Understanding
Xiang Li 0046, Jian Ding 0001, Mohamed Elhoseiny 0001 |
ECCV (23) | 1 |
| 2024 | MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsabstractThe recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in previous vision-language models. However, the technical details behind GPT-4 continue to remain undisclosed.
We believe that the enhanced multi-modal generation capabilities of GPT-4 stem from the utilization of sophisticated large language models (LLM).
To examine this phenomenon, we present MiniGPT-4, which aligns a frozen visual encoder with a frozen advanced LLM, Vicuna, using one projection layer.
Our work, for the first time, uncovers that properly aligning the visual features with an advanced large language model can possess numerous advanced multi-modal abilities demonstrated by GPT-4,
such as detailed image description generation and website creation from hand-drawn drafts.
Furthermore, we also observe other emerging capabilities in MiniGPT-4, including writing stories and poems inspired by given images, teaching users how to cook based on food photos, and so on.
In our experiment, we found that the model trained on short image caption pairs could produce unnatural language outputs (e.g., repetition and fragmentation). To address this problem, we curate a detailed image description dataset in the second stage to finetune the model, which consequently improves the model's generation reliability and overall usability. Deyao Zhu, Jun Chen 0021, Xiaoqian Shen, Xiang Li 0046, Mohamed Elhoseiny 0001 |
ICLR | 4 |
| 2024 | VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image UnderstandingabstractWe introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been proposed to pursue this goal, existing datasets are typically tailored to single tasks, lack detailed object information, or suffer from inadequate quality control. Exploring these improvement opportunities, we present a Versatile vision-language Benchmark for Remote Sensing image understanding, termed VRSBench. This benchmark comprises 29,614 images, with 29,614 human-verified detailed captions, 52,472 object references, and 123,221 question-answer pairs. It facilitates the training and evaluation of vision-language models across a broad spectrum of remote sensing image understanding tasks. We further evaluated state-of-the-art models on this benchmark for three vision-language tasks: image captioning, visual grounding, and visual question answering. Our work aims to significantly contribute to the development of advanced vision-language models in the field of remote sensing. The data and code can be accessed at https://vrsbench.github.io. Xiang Li 0046, Jian Ding 0001, Mohamed Elhoseiny 0001 |
NeurIPS | 1 |
| 2024 | Learning to learn point signature for 3D shape geometry
Hao Huang 0003, Lingjing Wang, Xiang Li 0046, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
Pattern Recognit. Lett. | 3 |
| 2024 | InfRS: Incremental Few-Shot Object Detection in Remote Sensing ImagesabstractFew-shot detection in remote sensing images has witnessed significant advancements recently. Despite these progresses, the capacity for continuous conceptual learning still poses a significant challenge to existing methodologies. In this article, we explore the intricate task of incremental few-shot object detection (iFSOD) in remote sensing images. We present a pioneering transfer-learning-based technique, termed InfRS, designed to enable the incremental learning of novel classes using a restricted set of examples, while simultaneously preserving the knowledge learned from previously seen classes without the need to revisit old data. Specifically, we pretrain the detector using sufficient data from base datasets and then generate a set of classwise prototypes that represent the intrinsic characteristics of the data. In the incremental learning session, we design a hybrid prototypical contrastive (HPC) encoding module for learning discriminative representations. Furthermore, we develop a prototypical calibration strategy based on the Wasserstein distance to overcome the catastrophic forgetting problem. Comprehensive evaluations conducted with two aerial imagery datasets show that our InfRS effectively addresses the iFSOD issue in remote sensing imagery. Code is available athttps://github.com/lyanna4869/InfRS.git. Wuzhou Li, Jiawei Zhou 0009, Xiang Li 0046, Guang Jin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | MoStGAN-V: Video Generation with Temporal Motion StylesabstractVideo generation remains a challenging task due to spatiotemporal complexity and the requirement of synthesizing diverse motions with temporal consistency. Previous works attempt to generate videos in arbitrary lengths either in an autoregressive manner or regarding time as a continuous signal. However, they struggle to synthesize detailed and diverse motions with temporal coherence and tend to generate repetitive scenes after a few time steps. In this work, we argue that a single time-agnostic latent vector of style-based generator is insufficient to model various and temporally-consistent motions. Hence, we introduce additional time-dependent motion styles to model diverse motion patterns. In addition, a Motion Style Attention modulation mechanism, dubbed as MoStAtt, is proposed to augment frames with vivid dynamics for each specific scale (i.e., layer), which assigns attention score for each motion style w.r.t deconvolution filter weights in the target synthesis layer and softly attends different motion styles for weight modulation. Experimental results show our model achieves state-of-the-art performance on four unconditional 2562video synthesis benchmarks trained with only 3 frames per clip and produces better qualitative results with respect to dynamic motions. Code and videos have been made available at https:/github.com/xiaoqian-shen/MoStGAN-V Xiaoqian Shen, Xiang Li 0046, Mohamed Elhoseiny 0001 |
CVPR | 2 |
| 2023 | FishNet: A Large-scale Dataset and Benchmark for Fish Recognition, Detection, and Functional Trait PredictionabstractAquatic species are essential components of the world’s ecosystem, and the preservation of aquatic biodiversity is crucial for maintaining proper ecosystem functioning. Unfortunately, increasing anthropogenic pressures such as overfishing, climate change, and coastal development pose significant threats to aquatic biodiversity. To address this challenge, it is necessary to design an automatic aquatic species monitoring systems that can help researchers and policymakers better understand changes in aquatic ecosystems and take appropriate actions to preserve biodiversity. However, the development of such systems is impeded by a lack of large-scale diverse aquatic species datasets. Existing aquatic species recognition datasets generally have a limited number of species, nor do they provide functional trait data, and so have only narrow potential for application. To address the need for generalized systems that can recognize, locate, and predict a wide array of species and their functional traits, we present FishNet, a large-scale diverse dataset containing 94,532 meticulously organized images from 17,357 aquatic species, organized according to aquatic biological taxonomy (order, family, genus, and species). We further build three benchmarks, i.e., fish classification, fish detection, and functional trait prediction, inspired by ecological research needs, to facilitate the development of aquatic species recognition systems, and promote further research in the field of aquatic ecology. Our FishNet dataset has the potential to encourage the development of more accurate and effective tools for the monitoring and protection of aquatic ecosystems, and hence take effective action toward the conservation of our planet’s aquatic biodiversity. Our dataset and code will be released at https://fishnet-2023.github.io/. Faizan Farooq Khan, Xiang Li 0046, Andrew J. Temple, Mohamed Elhoseiny 0001 |
ICCV | 2 |
| 2023 | An Explainable Multi-view Semantic Fusion Model for Multimodal Fake News DetectionabstractThe existing models have been achieved great success in capturing and fusing miltimodal semantics of news. However, they paid more attention to the global information, ignoring the interactions of global and local semantics and the inconsistency between different modalities. Therefore, we propose an explainable multi-view semantic fusion model (EMSFM), where we aggregate the important inconsistent semantics from local and global views to compensate the global information. Inspired by various forms of artificial fake news and real news, we summarize four views of multimodal correlation: consistency and inconsistency in the local and global views. Integrating these four views, our EMSFM can interpretatively establish global and local fusion between consistent and inconsistent semantics in multimodal relations for fake news detection. The extensive experimental results show that the EMSFM can improve the performance of multimodal fake news detection and provide a novel paradigm for explainable multi-view semantic fusion. Zhi Zeng 0001, Mingmin Wu, Xiang Li 0046, Zhongqiang Huang, Ying Sha |
ICME | 4 |
| 2023 | Correcting the Bias: Mitigating Multimodal Inconsistency Contrastive Learning for Multimodal Fake News DetectionabstractMultimodal fake news detection has become a topical research of fake news detection. Existing models have made great efforts in capturing and fusing multimodal semantics of news for classification. However, they overlooked mitigating inconsistency between different modalities, which may result in learning biased statistical information. Therefore, we propose a mitigating multimodal inconsistency contrastive learning framework (MMICF), which mitigates inconsistency in multi-modal relations for fake news detection. Inspired by various forms of artificial fake news, we summarize two patterns of multimodal inconsistency: local and global inconsistency. To mitigate local inconsistency in multimodal relations, we use a causal-relation reasoning module by causally removing the direct effects of the textual and visual entities. Considering the influence of global inconsistency in multimodal semantics, our contrastive learning framework mitigates the semantic deviation of contrastive text-image objectives, which are constrained into a unified semantic space by a modal unified module. Thus, our MMICF can jointly mitigate local and global inconsistency for further maximally exploiting multimodal consistent semantics for fake news detection. The extensive experimental results show that the MMICF can improve the performance of multimodal fake news detection and provide a novel paradigm for mitigating multimodal inconsistency contrastive learning. Zhi Zeng 0001, Mingmin Wu, Xiang Li 0046, Zhongqiang Huang, Ying Sha |
ICME | 4 |
| 2023 | Unsupervised Category-Specific Partial Point Set Registration via Joint Shape Completion and RegistrationabstractWe propose a self-supervised method for partial point set registration. Although recently proposed learning-based methods demonstrate impressive registration performance on full shape observations, these methods often suffer from performance degradation when dealing with partial shapes. To bridge the performance gap between partial and full point set registration, we propose to incorporate a shape completion network to benefit the registration process. To achieve this, we introduce a learnable latent code for each pair of shapes, which can be regarded as the geometric encoding of the target shape. By doing so, our model does not require an explicit feature embedding network to learn the feature encodings. More importantly, both our shape completion and point set registration networks take the shared latent codes as input, which are optimized simultaneously with the parameters of two decoder networks in the training process. Therefore, the point set registration process can benefit from the joint optimization process of latent codes, which are enforced to represent the information of full shapes instead of partial ones. In the inference stage, we fix the network parameters and optimize the latent codes to obtain the optimal shape completion and registration results. Our proposed method is purely unsupervised and does not require ground truth supervision. Experiments on the ModelNet40 dataset demonstrate the effectiveness of our model for partial point set registration. Xiang Li 0046, Lingjing Wang, Yi Fang 0006 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | Unsupervised 3D Shape Representation Learning Using Normalizing Flow
Xiang Li 0046, Congcong Wen, Hao Huang 0003 |
ACCV (1) | 1 |
| 2022 | Meta-Det3D: Learn to Learn Few-Shot 3D Object Detection
Shuaihang Yuan, Xiang Li 0046, Hao Huang 0003, Yi Fang 0006 |
ACCV (1) | 2 |
| 2022 | Road Extraction From Remote Sensing Images in Wildland-Urban Interface AreasabstractIn this letter, we address the problem of road extraction in Wildland–urban interface (WUI) areas. In recent years, with the great success of convolutional neural networks (CNNs) in various vision-related tasks, researchers have developed many CNN-based methods for road extraction on remote sensing images. Nevertheless, these methods mostly treat road extraction as a binary classification problem on semantic labeling. In WUI areas, the road is narrower and tends to be occluded by trees, which may result in the serious discontinuous problem of inferred road maps. To address this issue, we propose transforming the input representation of the binary classification map into a continuous signed distance map. In this way, our model is forced to predict the continuous distance representations and, thus, improve the spatial continuities of inferred roads. In addition, a real-value regression task is designed to train along with the original binary classification task to generate spatially continuous and semantically accurate road maps. Then, we conduct experiments on the public Massachusetts road data set and a homemade data set collected from Yajishan Mountain, Beijing, China. Finally, our proposed method achieves intersection-over-unions (IoUs) of 64.11% and 65.92% for the Massachusetts and WUI-Yajishan data sets, respectively, without any postprocessing. In addition, the ablation analysis shows that introducing the regression task on the proposed signed distance representation can effectively alleviate the problem of discontinuous road prediction. Furthermore, comparing with the state-of-the-art methods demonstrates the superiority of our method for road extraction in WUI areas. Xiang Li 0046, Yuan Hu 0004, Congcong Wen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Few-Shot Segmentation of Remote Sensing Images Using Deep Metric LearningabstractCurrent convolutional neural network (CNN)-based methods for remote sensing image segmentation require a large number of densely annotated images for model training and have limited generalization abilities for unseen object categories. In this letter, we propose a novel few-shot learning-based method for the semantic segmentation of remote sensing images. Our method can perform semantic labeling for unseen object categories with only a few annotated samples. More specifically, our model starts by using a deep CNN to extract high-level semantic features. The prototype representation of each class is then generated by using a masked average pooling on the feature embeddings of the support images with ground truth masks. Finally, our model performs semantic labeling over the query images by matching the feature embedding of each pixel to its nearest prototypes in the embedding space. Our model is optimized with a nonparametric metric learning-based loss function to maximize the intra-class similarity of learned prototypes while minimizing the inter-class similarity. Experiments on International Society for Photogrammetry and Remote Sensing (ISPRS) 2-D semantic labeling dataset demonstrate satisfying in-domain and cross-domain transferring abilities of our model. Xufeng Jiang, Xiang Li 0046 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Height Estimation From Single Aerial Images Using a Deep Ordinal Regression NetworkabstractUnderstanding the 3-D geometric structure of the Earth’s surface has been an active research topic in photogrammetry and remote sensing community for decades, serving as an essential building block for various applications such as 3-D digital city modeling, change detection, and city management. Previous research studies have extensively studied the problem of height estimation from aerial images based on stereo or multiview image matching. These methods require two or more images from different perspectives to reconstruct 3-D coordinates with camera information provided. In this letter, we deal with the ambiguous and unsolved problem of height estimation from a single aerial image. Driven by the great success of deep learning, especially deep convolutional neural networks (CNNs), some research studies have proposed to estimate height information from a single aerial image by training a deep CNN model with large-scale annotated data sets. These methods treat height estimation as a regression problem and directly use an encoder–decoder network to regress the height values. In this letter, we propose to divide height values into spacing-increasing intervals and transform the regression problem into an ordinal regression problem, using an ordinal loss for network training. To enable multiscale feature extraction, we further incorporate an Atrous Spatial Pyramid Pooling (ASPP) module to extract features from multiple dilated convolution layers. After that, a postprocessing technique is designed to transform the predicted height map of each patch into a seamless height map. Finally, we conduct extensive experiments on International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam data sets. Experimental results demonstrate significantly better performance of our method compared to state-of-the-art methods. Xiang Li 0046, Yi Fang 0006 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Geometry-Aware Segmentation of Remote Sensing Images via Joint Height EstimationabstractRecent studies have shown the benefits of using additional elevation data [e.g., digital surface model (DSM) or normalized DSM (nDSM)] for enhancing the performance of the semantic labeling of aerial images. However, previous methods mostly adopt 3-D elevation information as additional inputs, while, in many real-world applications, one does not have the corresponding DSM images at hand, and the spatial resolution of acquired DSM images usually does not match the aerial images. To alleviate this data constraint and also take advantage of 3-D elevation information, in this letter, a geometry-aware segmentation model is introduced to achieve accurate semantic labeling of aerial images via joint height estimation. Instead of using a single-stream encoder–decoder network for semantic labeling, we design a separate decoder branch to predict the height map and use the DSM images as side supervision to train this newly designed decoder branch. With the newly designed decoder branch, our model can distill the 3-D geometric features from 2-D appearance features under the supervision of ground-truth DSM images. Moreover, we develop a new geometry-aware convolution module that fuses the 3-D geometric features from the height decoder branch and the 2-D contextual features from the semantic segmentation branch. The fused feature embeddings can produce geometry-aware segmentation maps with enhanced performance. Our model is trained with DSM images as side supervision, while, in the inference stage, it does not require DSM data and directly predicts the semantic labels. Experiments on International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam data sets demonstrate the effectiveness of the proposed method for the semantic segmentation of aerial images. Xiang Li 0046, Congcong Wen, Lingjing Wang, Yi Fang 0006 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Few-Shot Object Detection on Remote Sensing ImagesabstractIn this article, we deal with the problem of object detection on remote sensing images. Previous researchers have developed numerous deep convolutional neural network (CNN)-based methods for object detection on remote sensing images, and they have reported remarkable achievements in detection performance and efficiency. However, current CNN-based methods often require a large number of annotated samples to train deep neural networks and tend to have limited generalization abilities for unseen object categories. In this article, we introduce a metalearning-based method for few-shot object detection on remote sensing images where only a few annotated samples are needed for the unseen object categories. More specifically, our model contains three main components: a metafeature extractor that learns to extract metafeature maps from input images, a feature reweighting module that learns class-specific reweighting vectors from the support images and use them to recalibrate the metafeature maps, and a bounding box prediction module that carries out object detection on the reweighted feature maps. We build our few-shot object detection model upon the YOLOv3 architecture and develop a multiscale object detection framework. Experiments on two benchmark data sets demonstrate that with only a few annotated samples, our model can still achieve a satisfying detection performance on remote sensing images, and the performance of our model is significantly better than the well-established baseline models. Xiang Li 0046, Jingyu Deng, Yi Fang 0006 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | GP-Aligner: Unsupervised Groupwise Nonrigid Point Set Registration Based on Optimizable Group Latent DescriptorabstractIn this paper, we propose a novel unsupervised method named GP-Aligner to address the problem of groupwise non-rigid point set registration. Compared to previous non-learning-based approaches, the proposed method gains competitive advantages by leveraging deep neural networks to effectively and efficiently align a large number of highly deformed 3D shapes with superior performance. Unlike most learning-based methods that use an explicit feature encoding network to extract per-shape features and their correlations, our model leverages a model-free learnable latent descriptor to characterize shape correlations among groups. More specifically, for a given group we first define an optimizable Group Latent Descriptor (GLD) to characterize the relationship among a group of point sets. Each GLD is randomly initialized from a Gaussian distribution and then concatenated with the coordinates of each point of the associated point sets in the group. A neural network-based decoder network is further constructed to predict the coherent flow fields to optimally deform the input groups of shapes to the aligned ones. During the optimization process, GP-Aligner jointly updates all GLDs and weight parameters of the decoder network towards the minimization of an unsupervised groupwise alignment loss. After optimization, for each group, our model coherently drives each point set towards a mean position (shape) without specifying one as the target. GP-Aligner does not require large-scale training data for network training, and it can directly align groups of point sets in a one-stage optimization process. GP-Aligner shows both accuracy and computational efficiency improvement in comparison with the state-of-the-art methods for groupwise point set registration. Moreover, GP-Aligner exhibits high efficiency in aligning a large number of groups of real-world 3D shapes. Lingjing Wang, Hao Huang 0003, Jifei Wang, Xiang Li 0046, Yi Fang 0006 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | 3D-MetaConNet: Meta-learning for 3D Shape Classification and SegmentationabstractSupervised learning on 3D shapes are extensively studied by prior literature, among which PointNet [29] and its variants PointNet++ [31] are representatives. However, these methods tackle 3D shape learning problems by training from scratch using a fixed learning algorithm over large amounts of labeled data, potentially challenged by data and computation bottlenecks. In the paper, we design a novel model, under the framework of meta-learning, to learn 3D shape representation. By training over multiple 3D tasks, each of which is defined as a supervised learning problem, our method can fast adapt to unseen tasks containing limited labeled data. Specifically, our model consists of a 3Dmeta-learner and a task-oriented 3D-learner, where the 3D-meta-learner produces parameter initialization for the 3D-learner after being trained over different tasks. With adaptively initialized parameters, the 3D-learner can be tuned rapidly in a few steps to achieve good performance on novel tasks with a small amount of training data. To further facilitate discriminative shape feature learning, we introduce a novel task-aware feature adaptation module under a contrastive learning scheme, in which all shapes in each task are considered as a whole and task-oriented compact features are learned. Therefore, we dub our model as 3DMetaConNet. Experiments on three public 3D datasets for few-shot shape classification and segmentation demonstrate that our method can learn compact and discriminative 3D shape features efficiently and robustly in a fast adaptation manner. Our method particularly outperforms the methods without a meta-learning framework and is also superior to existing meta-learning approaches. Hao Huang 0003, Xiang Li 0046, Lingjing Wang, Yi Fang 0006 |
3DV | 2 |
| 2021 | Topology Constrained Shape CorrespondenceabstractTo better address the deformation and structural variation challenges inherently present in 3D shapes, researchers have shifted their focus from designing handcrafted point descriptors to learning point descriptors and their correspondences in a data-driven manner. Recent studies have developed deep neural networks for robust point descriptor and shape correspondence learning in consideration of local structural information. In this article, we developed a novel shape correspondence learning network, called TC-NET, which further enhances performance by encouraging the topological consistency between the embedding feature space and the input shape space. Specifically, in this article, we first calculate the topology-associated edge weights to represent the topological structure of each point. Then, in order to preserve this topological structure in high-dimensional feature space, a structural regularization term is defined to minimize the topology-consistent feature reconstruction loss (Topo-Loss) during the correspondence learning process. Our proposed method achieved state-of-the-art performance on three shape correspondence benchmark datasets. In addition, the proposed topology preservation concept can be easily generalized to other learning-based shape analysis tasks to regularize the topological structure of high-dimensional feature spaces. Xiang Li 0046, Congcong Wen, Lingjing Wang, Yi Fang 0006 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Robust Image Matching By Dynamic Feature Selection
Hao Huang 0003, Jianchun Chen, Xiang Li 0046, Lingjing Wang, Yi Fang 0006 |
BMVC | 3 |
| 2020 | Few-Shot Learning of Part-Specific Probability Space for 3D Shape SegmentationabstractRecently, deep neural networks are introduced as supervised discriminative models for the learning of 3D point cloud segmentation. Most previous supervised methods require a large number of training data with human annotation part labels to guide the training process to ensure the model's generalization abilities on test data. In comparison, we propose a novel 3D shape segmentation method that requires few labeled data for training. Given an input 3D shape, the training of our model starts with identifying a similar 3D shape with part annotations from a mini-pool of shape templates (e.g. 10 shapes). With the selected template shape, a novel Coherent Point Transformer is proposed to fully leverage the power of a deep neural network to smoothly morph the template shape towards the input shape. Then, based on the transformed template shapes with part labels, a newly proposed Part-specific Density Estimator is developed to learn a continuous part-specific probability distribution function on the entire 3D space with a batch consistency regularization term. With the learned part-specific probability distribution, our model is able to predict the part labels of a new input 3D shape in an end-to-end manner. We demonstrate that our proposed method can achieve remarkable segmentation results on the ShapeNet dataset with few shots, compared to previous supervised learning approaches. Lingjing Wang, Xiang Li 0046, Yi Fang 0006 |
CVPR | 2 |
| 2020 | 3DMotion-Net: Learning Continuous Flow Function for 3D Motion PredictionabstractThis paper deals with predicting future 3D motions of 3D object scans from the previous two consecutive frames. Previous methods mostly focus on sparse motion prediction in the form of skeletons. While in this paper, we focus on predicting dense 3D motions in the form of 3D point clouds. To approach this problem, we propose a self-supervised approach that leverages the power of the deep neural network to learn a continuous flow function of 3D point clouds that can predict temporally consistent future motions and naturally bring out the correspondences among consecutive point clouds at the same time. More specifically, in our approach, to eliminate the unsolved and challenging process of defining a discrete point convolution on 3D point cloud sequences to encode spatial and temporal information, we introduce a learnable latent code to represent the temporal-aware shape descriptor, which is optimized during the model training. Moreover, a temporally consistent motion Morpher is proposed to learn a continuous flow field which deforms a 3D scan from the current frame to the next frame. We perform extensive experiments on D-FAUST, SCAPE, and TOSCA benchmark data sets. The results demonstrate that our approach is capable of handling temporally inconsistent input and produces consistent future 3D motion while requiring no ground truth supervision. Shuaihang Yuan, Xiang Li 0046, Anthony Tzes, Yi Fang 0006 |
IROS | 2 |
| 2019 | PC-Net: Unsupervised Point Correspondence Learning with Neural NetworksabstractPoint sets correspondence concerns with the establishment of point-wise correspondence for a group of 2D or 3D point sets with similar shape description. Existing methods often iteratively search for the optimal point-wise correspondence assignment for two sets of points, driven by maximizing the similarity between two sets of explicitly designed point features or by determining the parametric transformation for the best alignment between two point sets. In contrast, without depending on the explicit definitions of point features or transformation, our paper introduces a novel point correspondence neural networks (PC-Net) that is able to learn and predict the point correspondence among the populations of a specific object (e.g. fish, human, chair, etc) in an unsupervised manner. Specifically, in this paper, we first develop an encoder to learn the shape descriptor from a point set that captures essential global and deformation-insensitive geometric properties. Then followed with a novel motion-driven process, our PC-Net drives a template shape, that consists of a set of landmark points, morph and conform around a target shape object which is reconstructed through decoding the previously characterized shape descriptor. As a result, the motion-driven process progressively and coherently drifts all landmark points from the template shape to corresponding positions on the target object shape. The experimental results demonstrate that PC-Net can establish robust unsupervised point correspondence over a group of deformable object shapes in the presence of geometric noise and missing points. More importantly, with great generalization capability, PC-Net is capable of instantly predicting group point corresponding for unseen point sets. Xiang Li 0046, Lingjing Wang, Yi Fang 0006 |
3DV | 1 |
| 2019 | Dynamic Feature Fusion for Semantic Edge DetectionabstractFeatures from multiple scales can greatly benefit the semantic edge detection task if they are well fused. However, the prevalent semantic edge detection methods apply a fixed weight fusion strategy where images with different semantics are forced to share the same weights, resulting in universal fusion weights for all images and locations regardless of their different semantics or local context. In this work, we propose a novel dynamic feature fusion strategy that assigns different fusion weights for different input images and locations adaptively. This is achieved by a proposed weight learner to infer proper fusion weights over multi-level features for each location of the feature map, conditioned on the specific input. In this way, the heterogeneity in contributions made by different locations of feature maps and input images can be better considered and thus help produce more accurate and sharper edge predictions. We show that our model with the novel dynamic feature fusion is superior to fixed weight fusion and also the na\"ive location-invariant weight fusion methods, via comprehensive experiments on benchmarks Cityscapes and SBD. In particular, our method outperforms all existing well established methods and achieves new state-of-the-art. Yuan Hu 0004, Yunpeng Chen, Xiang Li 0046, Jiashi Feng |
IJCAI | 3 |
| 2019 | Arbicon-Net: Arbitrary Continuous Geometric Transformation Networks for Image RegistrationabstractThis paper concerns the undetermined problem of estimating geometric transformation between image pairs. Recent methods introduce deep neural networks to predict the controlling parameters of hand-crafted geometric transformation models (e.g. thin-plate spline) for image registration and matching. However, the low-dimension parametric models are incapable of estimating a highly complex geometric transform with limited flexibility to model the actual geometric deformation from image pairs. To address this issue, we present an end-to-end trainable deep neural networks, named Arbitrary Continuous Geometric Transformation Networks (Arbicon-Net), to directly predict the dense displacement field for pairwise image alignment. Arbicon-Net is generalized from training data to predict the desired arbitrary continuous geometric transformation in a data-driven manner for unseen new pair of images. Particularly, without imposing penalization terms, the predicted displacement vector function is proven to be spatially continuous and smooth. To verify the performance of Arbicon-Net, we conducted semantic alignment tests over both synthetic and real image dataset with various experimental settings. The results demonstrate that Arbicon-Net outperforms the previous image alignment techniques in identifying the image correspondences. Jianchun Chen, Lingjing Wang, Xiang Li 0046, Yi Fang 0006 |
NeurIPS | 3 |
| 2019 | A Sample Update-Based Convolutional Neural Network Framework for Object Detection in Large-Area Remote Sensing ImagesabstractThis letter addresses the issue of accurate object detection in large-area remote sensing images. Although many convolutional neural network (CNN)-based object detection models can achieve high accuracy in small image patches, the models perform poorly in large-area images due to the large quantity of false and missing detections that arise from complex backgrounds and diverse groundcover types. To address this challenge, this letter proposes a sample update-based CNN (SUCNN) framework for object detection in large-area remote sensing images. The proposed framework contains two stages. In the first stage, a base model—single-shot multibox detector—is trained with the training data set. In the second stage, artificial composite samples are generated to update the training set. The parameters of the first-stage model are fine-tuned with the updated data set to obtain the second-stage model. The first- and second-stage models are evaluated using the large-area remote sensing image test set. Comparison experiments show the effectiveness and superiority of the proposed SUCNN framework for object detection in large-area remote sensing images. Yuan Hu 0004, Xiang Li 0046, Sha Xiao |
IEEE Geosci. Remote. Sens. Lett. | 2 |