Yi Fang 0006

dblp:96/361-6 · DBLP profile ↗
← Back
109ranked-venue papers
6as first author
56since 2021 · last 2026
0000-0001-9427-3883ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 69 · 3 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 3 first-author · 29 since 2021Systems, architecture and hardware · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Phys-Liquid: A Physics-Informed Dataset for Estimating 3D Geometry and Volume of Transparent Deformable Liquids
abstract
Estimating the geometric and volumetric properties of transparent deformable liquids is challenging due to optical complexities and dynamic surface deformations induced by container movements. Autonomous robots performing precise liquid manipulation tasks—such as dispensing, aspiration, and mixing—must handle containers in ways that inevitably induce these deformations, complicating accurate liquid state assessment. Current datasets lack comprehensive physics-informed simulation data representing realistic liquid behaviors under diverse dynamic scenarios. To bridge this gap, we introduce Phys-Liquid, a physics-informed dataset comprising 97,200 simulation images and corresponding 3D meshes, capturing liquid dynamics across multiple laboratory scenes, lighting conditions, liquid colors, and container rotations. To validate the realism and effectiveness of Phys-Liquid, we propose a four-stage reconstruction and estimation pipeline involving liquid segmentation, multi-view mask generation, 3D mesh reconstruction, and real-world scaling. Experimental results demonstrate improved accuracy and consistency in reconstructing liquid geometry and volume, outperforming existing benchmarks. The dataset and associated validation methods facilitate future advancements in transparent liquid perception tasks.
Ke Ma 0012, Yizhou Fang, Jean-Baptiste Weibel, Xinggang Wang, Yang Xiao 0007, Yi Fang 0006
AAAI7
2026 Deep learning regression for photovoltaic soiling quantification using multi-source drone-ground imaging
Muhammad Faizan Tahir, Samyam Lamichhane, Anthony Tzes, Yi Fang 0006, Dongliang Xiao
Expert Syst. Appl.4
2026 On Demographic Group Fairness Guarantees in Deep Learning
abstract
We present a theoretical framework analyzing the relationship between data distributions and fairness guarantees in deep learning. Our work establishes novel bounds that explicitly account for data distribution heterogeneity across demographic groups, while introducing a formal analysis framework that minimizes expected loss differences across these groups. Moreover, we derive bounds for fairness errors and convergence rates, characterizing how distributional differences between groups affect the fundamental trade-off between fairness and accuracy. Through extensive experiments on diverse datasets across various modalities (image, tabular data, and text), including FairVision (eye disease detection), CheXpert (pleural effusion detection), HAM10000 (skin lesion classification), FairFace (facial attribute recognition), ACS Income (income prediction), CivilComments-WILDS (toxic comment detection), we validate our theoretical findings and demonstrate that differences in feature distributions across demographic groups significantly impact model fairness, with performance disparities particularly pronounced in racial categories. The theoretical bounds we derive corroborate these empirical observations, providing insights into the fundamental limits of achieving fairness in deep learning models when faced with heterogeneous data distributions. This work advances our understanding of fairness in AI and provides a theoretical foundation for developing more equitable algorithms. Motivated by these theoretical insights, particularly the link between feature distribution shifts and fairness gaps, we propose Fairness-Aware Regularization (FAR), a practical training objective that directly minimizes inter-group discrepancies in feature centroids and covariances to improve equitable performance. We validate the effectiveness of FAR across all datasets considered in this study, consistently observing improvements in overall AUC, ES-AUC, and subgroup performance.
Yan Luo 0002, Congcong Wen, Min Shi 0001, Hao Huang 0003, Yi Fang 0006, Mengyu Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 FairFedMed: Benchmarking Group Fairness in Federated Medical Imaging With FairLoRA
abstract
Fairness remains a critical concern in healthcare, where unequal access to services and treatment outcomes can adversely affect patient health. While Federated Learning (FL) presents a collaborative and privacy-preserving approach to model training, ensuring fairness is challenging due to heterogeneous data across institutions, and current research primarily addresses non-medical applications. To fill this gap, we establish the first experimental benchmark for fairness in medical FL, evaluating six representative FL methods across diverse demographic attributes and imaging modalities. We introduce FairFedMed, the first medical FL dataset specifically designed to study group fairness (i.e., consistent performance across demographic groups). It comprises two parts: FairFedMed-Oph, featuring 2D fundus and 3D OCT ophthalmology samples with six demographic attributes; and FairFedMed-Chest, which simulates real cross-institutional FL using subsets of CheXpert and MIMIC-CXR. Together, they support both simulated and real-world FL across diverse medical modalities and demographic groups. Existing FL models often underperform on medical images and overlook fairness across demographic groups. To address this, we propose FairLoRA, a fairness-aware FL framework based on SVD-based low-rank approximation. It customizes singular value matrices per demographic group while sharing singular vectors, ensuring both fairness and efficiency. Experimental results on the FairFedMed dataset demonstrate that FairLoRA not only achieves state-of-the-art performance in medical image classification but also significantly improves fairness across diverse populations. Our code and dataset can be accessible via GitHub link: https://github.com/Harvard-AI-and-Robotics-Lab/FairFedMed.
Minghan Li 0001, Congcong Wen, Yu Tian 0001, Min Shi 0001, Yan Luo 0002, Hao Huang 0003, Yi Fang 0006, Mengyu Wang 0001
IEEE Trans. Medical Imaging7
2026 PFF-Net: Patch Feature Fitting for Point Cloud Normal Estimation
abstract
Estimating the normal of a point requires constructing a local patch to provide center-surrounding context, but determining the appropriate neighborhood size is difficult when dealing with different data or geometries. Existing methods commonly employ various parameter-heavy strategies to extract a full feature description from the input patch. However, they still have difficulties in accurately and efficiently predicting normals for various point clouds. In this work, we present a new idea of feature extraction for robust normal estimation of point clouds. We use the fusion of multi-scale features from different neighborhood sizes to address the issue of selecting reasonable patch sizes for various data or geometries. We seek to model a patch feature fitting (PFF) based on multi-scale features to approximate the optimal geometric description for normal estimation and implement the approximation process via multi-scale feature aggregation and cross-scale feature compensation. The feature aggregation module progressively aggregates the patch features of different scales to the center of the patch and shrinks the patch size by removing points far from the center. It not only enables the network to precisely capture the structure characteristic in a wide range, but also describes highly detailed geometries. The feature compensation module ensures the reusability of features from earlier layers of large scales and reveals associated information in different patch sizes. Our approximation strategy based on aggregating the features of multiple scales enables the model to achieve scale adaptation of varying local patches and deliver the optimal feature description. Extensive experiments demonstrate that our method achieves state-of-the-art performance on both synthetic and real-world datasets with fewer network parameters and running time.
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yue Gao 0002, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
IEEE Trans. Vis. Comput. Graph.5
2025 Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon Tasks
abstract
Policy learning focuses on devising strategies for agents in embodied artificial intelligence systems to perform optimal actions based on their perceived states. One of the key challenges in policy learning involves handling complex, long-horizon tasks that require managing extensive sequences of actions and observations with multiple modes. Wavelet analysis offers significant advantages in signal processing, notably in decomposing signals at multiple scales to capture both global trends and fine-grained details. In this work, we introduce a novel wavelet policy learning framework that utilizes wavelet transformations to enhance policy learning. Our approach leverages learnable multi-scale wavelet decomposition to facilitate detailed observation analysis and robust action planning over extended sequences. We detail the design and implementation of our wavelet policy, which incorporates lifting schemes for effective multi-resolution analysis and action generation. This framework is evaluated across multiple complex scenarios, including robotic manipulation, self-driving, and multi-robot collaboration, demonstrating the effectiveness of our method in improving the precision and reliability of the learned policy.
Hao Huang 0003, Shuaihang Yuan, Geeta Chandra Raju Bethala, Congcong Wen, Anthony Tzes, Yi Fang 0006
ICCV6
2025 MultiTalk: Introspective and Extrospective Dialogue for Human-Environment-LLM Alignment
abstract
LLMs have shown promising results in task planning due to their strong natural language understanding and reasoning capabilities. However, issues such as hallucinations, ambiguities in human instructions, environmental constraints, and limitations in the executing agent's capabilities often lead to flawed or incomplete plans. This paper proposes MultiTalk, an LLM-based task planning methodology that addresses these issues through a framework of introspective and extrospective dialogue loops. This approach helps ground generated plans in the context of the environment and the agent's capabilities, while also resolving uncertainties and ambiguities in the given task. These loops are enabled by specialized systems designed to extract and predict task-specific states, and flag mismatches or misalignments among the human user, the LLM agent, and the environment. Effective feedback pathways between these systems and the LLM planner foster meaningful dialogue. The efficacy of this methodology is demonstrated through its application to robotic manipulation tasks. Experiments and ablations highlight the robustness and reliability of our method, and comparisons with baselines further illustrate the superiority of MultiTalk in task planning for embodied agents. Project Website: https://llm-multitalk.github.io/
Venkata Naren Devarakonda, Ali Umut Kaypak, Shuaihang Yuan, Prashanth Krishnamurthy, Yi Fang 0006, Farshad Khorrami
ICRA5
2025 Tracking Everything in Robotic-Assisted Surgery
abstract
Accurate tracking of tissues and instruments in videos is crucial for Robotic-Assisted Minimally Invasive Surgery (RAMIS), as it enables the robot to comprehend the surgical scene with precise locations and interactions of tissues and tools. Traditional keypoint-based sparse tracking is limited by featured points, while flow-based dense two-view matching suffers from long-term drifts. Recently, the Tracking Any Point (TAP) algorithm was proposed to overcome these limitations and achieve dense accurate long-term tracking. However, its efficacy in surgical scenarios remains untested, largely due to the lack of a comprehensive surgical tracking dataset for evaluation. To address this gap, we introduce a new annotated surgical tracking dataset for benchmarking tracking methods for surgical scenarios, comprising real-world surgical videos with complex tissue and instrument motions. We extensively evaluate state-of-the-art (SOTA) TAP-based algorithms on this dataset and reveal their limitations in challenging surgical scenarios, including fast instrument motion, severe occlusions, and motion blur, etc. Furthermore, we propose a new tracking method, namely SurgMotion, to solve the challenges and further improve the tracking performance. Our proposed method outperforms most TAP-based algorithms in surgical instruments tracking, and especially demonstrates significant improvements over baselines in challenging medical videos. Our code and dataset are available at https://github.com/zhanbh1019/SurgicalMotion.
Bohan Zhan, Yi Fang 0006, Francisco Vasconcelos 0001, Danail Stoyanov, Daniel S. Elson, Baoru Huang
ICRA3
2025 Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language Models
abstract
Robotic navigation plays a pivotal role in a wide range of real-world applications. While traditional navigation systems focus on efficiency and obstacle avoidance, their inability to model complex human behaviors in shared spaces has underscored the growing need for socially aware navigation. In this work, we explore a novel paradigm of socially aware robot navigation empowered by large language models (LLMs), and propose HSAC-LLM, a hybrid framework that seamlessly integrates deep reinforcement learning with the reasoning and communication capabilities of LLMs. Unlike prior approaches that passively predict pedestrian trajectories or issue pre-scripted alerts, HSAC-LLM enables bidirectional natural language interaction, allowing robots to proactively engage in dialogue with pedestrians to resolve potential conflicts and negotiate path decisions. Extensive evaluations across 2D simulations, Gazebo environments, and real-world deployments demonstrate that HSAC-LLM consistently outperforms state-of-the-art DRL baselines under our proposed socially aware navigation metric, which covers safety, efficiency, and human comfort. By bridging linguistic reasoning and interactive motion planning, our results highlight the potential of LLM-augmented agents for robust, adaptive, and human-aligned navigation in real-world settings. Project page: https://hsacllm.github.io/.
Congcong Wen, Geeta Chandra Raju Bethala, Shuaihang Yuan, Hao Huang 0003, Mengyu Wang 0001, Yu-Shen Liu, Anthony Tzes, Yi Fang 0006
IROS10
2025 MaterialRefGS: Reflective Gaussian Splatting with Multi-view Consistent Material Inference
abstract
Modeling reflections from 2D images is essential for photorealistic rendering and novel view synthesis. Recent approaches enhance Gaussian primitives with reflection-related material attributes to enable physically based rendering (PBR) with Gaussian Splatting. However, the material inference often lacks sufficient constraints, especially under limited environment modeling, resulting in illumination aliasing and reduced generalization. In this work, we revisit the problem from a multi-view perspective and show that multi-view consistent material inference with more physically-based environment modeling is key to learning accurate reflections with Gaussian Splatting. To this end, we enforce 2D Gaussians to produce multi-view consistent material maps during deferred shading. We also track photometric variations across views to identify highly reflective regions, which serve as strong priors for reflection strength terms. To handle indirect illumination caused by inter-object occlusions, we further introduce an environment modeling strategy through ray tracing with 2DGS, enabling photorealistic rendering of indirect radiance. Experiments on widely used benchmarks show that our method faithfully recovers both illumination and geometry, achieving state-of-the-art rendering quality in novel views synthesis. Project Page: https://wen-yuan-zhang.github.io/MaterialRefGS.
Jimin Tang, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
NeurIPS4
2025 U-CAN: Unsupervised Point Cloud Denoising with Consistency-Aware Noise2Noise Matching
abstract
Point clouds captured by scanning sensors are often perturbed by noise, which have a highly negative impact on downstream tasks (e.g. surface reconstruction and shape understanding). Previous works mostly focus on training neural networks with noisy-clean point cloud pairs for learning denoising priors, which requires extensively manual efforts. In this work, we introduce U-CAN, an Unsupervised framework for point cloud denoising with Consistency-Aware Noise2Noise matching. Specifically, we leverage a neural network to infer a multi-step denoising path for each point of a shape or scene with a noise to noise matching schema. We achieve this by a novel loss which enables statistical reasoning on multiple noisy point cloud observations. We further introduce a novel constraint on the denoised geometry consistency for learning consistency-aware denoising patterns. We justify that the proposed constraint is a general term which is not limited to 3D domain and can also contribute to the area of 2D image denoising. Our evaluations under the widely used benchmarks in point cloud denoising, upsampling and image denoising show significant improvement over the state-of-the-art unsupervised methods, where U-CAN also produces comparable results with the supervised methods. Project page: https://gloriasze.github.io/U-CAN/.
Junsheng Zhou, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
NeurIPS4
2025 3D Shape Contrastive Representation Learning With Adversarial Examples
abstract
Current supervised methods for 3D shape representation learning have achieved satisfying performance, yet require extensive human-labeled datasets. Unsupervised learning-based methods provide a viable solution by learning shape representations without using ground truth labels. In this study, we develop a contrastive learning framework for unsupervised representation learning of 3D shapes. Specifically, in order to encourage models to pay more attention to useful information during representation learning, we first introduce a new paradigm for critical points search based on the adversarial mechanism. We extract critical points with a larger impact on the global feature by attacking a pre-trained auto-encoder model, and apply data augmentations on these points to generate adversarial examples. Taking a pair of adversarial examples as inputs, we obtain their intermediate embeddings and global representations of corresponding inputs, which are then transformed into latent spaces by two predictor heads. Finally, we train the proposed model by maximizing the agreements on these latent spaces via Normalized Temperature-scaled Cross Entropy (NT-Xent) loss and a newly designed Cross-layer Normalized Temperature-scaled Cross Entropy (Cross-NT-Xent) loss, where the latter is proposed in this paper to enforce cross-layer feature similarities. The effectiveness, robustness, and transferability of learned representations are validated on three downstream tasks, including object classification, few-shot classification, and shape retrieval. Experiments on three benchmark datasets show that our learned representations achieve better or competitive performance than current state-of-the-art methods in these downstream tasks. Moreover, our model can easily be extended to 3D part segmentation and scene segmentation tasks.
Congcong Wen, Xiang Li 0046, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006
IEEE Trans. Multim.5
2025 Retrieval-Specific View Learning for Sketch-to-Shape Retrieval
abstract
Sketch-based 3D shape retrieval (SBSR) can be approached by learning domain-invariant descriptors or ranking metrics from sketches and 2D view images of 3D shapes rendered through numerous viewpoints. However, determining the most appropriate viewpoints that convey discriminative geometric features to benefit the task of SBSR became an essential yet not fully explored area. Existing works extract 3D features from multi-view images observed through pre-defined viewpoints to match 2D sketches. Those methods, however, fail to dynamically select viewpoints by considering the SBSR task. In this work, we introduce a fully differentiable viewpoint learning paradigm driven by the downstream SBSR task, which supports the task-aware and sketch-dependent dynamic viewpoint determination process. We naturally integrate this task-specific and sketch-dependent viewpoint learning process into a meta-learning framework to develop a novel Dynamic Viewer (DV) module for SBSR. DV module comprises a Meta View Learner (MVL) block and a View Generator (VG) block. Specifically, as the first part of the DV module, the MVL block learns to initiate the necessary network parameters of the VG block. Then, the VG block that serves as the second part learns the best viewpoints to render 2D images. To learn the optimal viewpoints for SBSR, we further introduce a view mining loss that aims to maximize the similarity of feature-level information among rendered 2D views and the query sketch. Further, we adopt a variational autoencoder (VAE) to retrieve 3D shapes by setting the newly rendered images and query sketch as inputs. As evidenced by the comprehensive experimental results conducted on popular SBSR datasets, the proposed framework has been demonstrated to outperform recent methods in both category-level sketch-based and fine-grained SBSR.
Shuaihang Yuan, Congcong Wen, Yu-Shen Liu, Yi Fang 0006
IEEE Trans. Multim.4
2024 FairCLIP: Harnessing Fairness in Vision-Language Learning
abstract
Fairness is a critical concern in deep learning, especially in healthcare, where these models influence diagnoses and treatment decisions. Although fairness has been investigated in the vision-only domain, the fairness of medical vision-language (VL) models remains unexplored due to the scarcity of medical VL datasets for studying fairness. To bridge this research gap, we introduce the first fair vision-language medical dataset (Harvard-FairVLMed) that provides detailed demographic attributes, ground-truth labels, and clinical notes to facilitate an in-depth examination of fairness within VL foundation models. Using Harvard-FairVLMed, we conduct a comprehensive fairness analysis of two widely-used VL models (CLIP and BLIP2), pre-trained on both natural and medical domains, across four different protected attributes. Our results highlight significant biases in all VL models, with Asian, Male, Non-Hispanic, and Spanish being the preferred subgroups across the protected attributes of race, gender, ethnicity, and language, respectively. In order to alleviate these biases, we propose FairCLIP an optimal-transport-based approach that achieves a favorable trade-off between performance and fairness by reducing the Sinkhorn distance between the overall sample distribution and the distributions corresponding to each demographic group. As the first VL dataset of its kind, Harvard-FairVLMed holds the potential to catalyze advancements in the development of machine learning models that are both ethically aware and clinically effective. Our dataset and code are available at https://ophai.hms.harvard.edu/datasets/harvard-fairvlmed10k.
Yan Luo 0002, Min Shi 0001, Muhammad Osama Khan, Muhammad Muneeb Afzal, Hao Huang 0003, Shuaihang Yuan, Yu Tian 0001, Luo Song, Ava Kouhana, Tobias Elze, Yi Fang 0006, Mengyu Wang 0001
CVPR11
2024 FairDomain: Achieving Fairness in Cross-Domain Medical Image Segmentation and Classification
Yu Tian 0001, Congcong Wen, Min Shi 0001, Muhammad Muneeb Afzal, Hao Huang 0003, Muhammad Osama Khan, Yan Luo 0002, Yi Fang 0006, Mengyu Wang 0001
ECCV (76)8
2024 Towards Dual Transparent Liquid Level Estimation in Biomedical Lab: Dataset, Methods and Practices
Xiayu Wang, Ke Ma 0012, Ruiyun Zhong, Xinggang Wang, Yi Fang 0006, Yang Xiao 0007
ECCV (65)5
2024 Optimizing Personalized Robot Actions with Ranking of Trajectories
Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006
ICPR (29)6
2024 Reliable Semantic Understanding for Real World Zero-Shot Object Goal Navigation
Halil Utku Unlu, Shuaihang Yuan, Congcong Wen, Hao Huang 0003, Anthony Tzes, Yi Fang 0006
ICPR (30)6
2024 Zero-Shot Object Navigation with Vision-Language Models Reasoning
Congcong Wen, Yisiyuan Huang, Hao Huang 0003, Yanjia Huang, Shuaihang Yuan, Yu-Shen Liu, Yi Fang 0006
ICPR (18)9
2024 Goal-Driven Transformer for Robot Behavior Learning from Play Data
Congcong Wen, Jiazhao Liang, Shuaihang Yuan, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006
ICPR (30)8
2024 Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
Shuaihang Yuan, Halil Utku Unlu, Hao Huang 0003, Congcong Wen, Anthony Tzes, Yi Fang 0006
ICPR (30)6
2024 Noisy Few-shot 3D Point Cloud Scene Segmentation
abstract
3D scene semantic segmentation plays a crucial role in robotics by enabling robots to understand and interpret their environment in a detailed and context-aware manner, facilitating tasks such as navigation, object manipulation, and interaction within complex spaces. A preponderance of methodology predominantly adopts a fully supervised framework for 3D point cloud scene semantic segmentation. Such paradigms exhibit an intrinsic dependency on extensive labeled datasets, presenting challenges in acquisition and exhibiting incapacity to segment novel classes, especially when the training data are contaminated by noisy samples. To address these limitations, this study introduces a novel few-shot segmentation approach to robustly segment 3D point cloud scenes with noisy labels using a meta-learning scheme. Specifically, we first build a multi-prototype graph and then suppress samples with noisy labels based on the graph structure. A subgraph bagging scheme is then proposed to conduct semi-supervised transductive learning to propagate labels. To optimize the graph structure to learn discriminative prototype features, we design a triplet contrastive loss to increase the compactness of these subgraphs. We evaluated our method on two widely used 3D point cloud scene segmentation benchmarks within few-shot (i.e., 2/3-way 5-shot) segmentation settings with noisy samples. Experimental results demonstrate the improvement of our method over the compared baselines, illustrating the robustness of our method in few-shot 3D scene segmentation against noisy samples. The code is available at: https://github.com/hhuang-code/Noisy_Fewshot_Segmentation.
Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006
ICRA5
2024 3D-OAE: Occlusion Auto-Encoders for Self-Supervised Learning on Point Clouds
abstract
The manual annotation for large-scale point clouds is still tedious and unavailable for many harsh real-world tasks. Self-supervised learning, which is used on raw and unlabeled data to pre-train deep neural networks, is a promising approach to address this issue. Existing works usually take the common aid from auto-encoders to establish the self-supervision by the self-reconstruction schema. However, the previous auto-encoders merely focus on the global shapes and do not distinguish the local and global geometric features apart. To address this problem, we present a novel and efficient self-supervised point cloud representation learning framework, named 3D Occlusion Auto-Encoder (3D-OAE), to facilitate the detailed supervision inherited in local regions and global shapes. We propose to randomly occlude some local patches of point clouds and establish the supervision via inpainting the occluded patches using the remaining ones. Specifically, we design an asymmetrical encoder-decoder architecture based on standard Transformer, where the encoder operates only on the visible subset of patches to learn local patterns, and a lightweight decoder is designed to leverage these visible patterns to infer the missing geometries via self-attention. We find that occluding a very high proportion of the input point cloud (e.g. 75%) will still yield a nontrivial self-supervisory performance, which enables us to achieve 3-4 times faster during training but also improve accuracy. Experimental results show that our approach outperforms the state-of-the-art on a diverse range of down-stream discriminative and generative tasks. Code is available at https://github.com/junshengzhou/3D-OAE.
Junsheng Zhou, Xin Wen 0003, Baorui Ma, Yu-Shen Liu, Yue Gao 0002, Yi Fang 0006, Zhizhong Han
ICRA6
2024 Weakly Scene Segmentation Using Efficient Transformer
abstract
Current methods for large-scale point cloud scene semantic segmentation rely on manually annotated dense point-wise labels, which are costly, labor-intensive, and prone to errors. Consequently, gathering point cloud scenes with billions of labeled points is impractical in real-world scenarios. In this paper, we introduce a novel weak supervision approach to semantically segment large-scale indoor scenes, requiring only 1‰ of the points to be labeled. Specifically, we develop an efficient point neighbor Transformer to capture the geometry of local point cloud patches. To address the quadratic complexity of self-attention computation in Transformers, particularly for large-scale point clouds, we propose approximating the self-attention matrix using low-rank and sparse decomposition. Building on the point neighbor Transformer as foundational blocks, we design a Low-rank Sparse Transformer Network (LST-Net) for weakly supervised large-scale point cloud scene semantic segmentation. Experimental results on two commonly used indoor point cloud scene segmentation benchmarks demonstrate that our model achieves performance comparable to those of both weakly supervised and fully supervised methods. Our code can be found in https://github.com/hhuang-code/LST-Net.
Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006
IROS5
2024 ChatMap: A Wearable Platform Based on the Multi-modal Foundation Model to Augment Spatial Cognition for People with Blindness and Low Vision
abstract
Spatial cognition refers to the ability to gain knowledge about their surroundings and utilize this information to identify their location, acquire resources, and navigate their way back to familiar places. People with blindness and low vision (pBLV) face significant challenges with spatial cognition due to the reliance on visual input. Without the full range of visual cues, pBLV individuals often find it difficult to grasp a comprehensive understanding of their environment, leading to obstacles in scene recognition and precise object localization, especially in unfamiliar environments. This limitation extends to their ability to independently detect and avoid potential tripping hazards, making navigation and interaction with their environment more challenging. In this paper, we present a pioneering wearable platform tailored to enhance the spatial cognition of pBLV through the integration of multi-modal foundation model. The proposed platform integrates a wearable camera with audio module and leverages the advanced capabilities of vision language foundation model (i.e., GPT-4 and GPT-4V), for the nuanced processing of visual and textual data. Specifically, we employ vision language models to bridge the gap between visual information and the proprioception of visually impaired users, offering more intelligible guidance by aligning visual data with the natural perception of space and movement. Then we apply prompt engineering to guide the large language model to act as an assistant tailored specifically for pBLV users to produce accurate answers. Another innovation in our model is the incorporation of a chain of thought reasoning process, which enhances the accuracy and interpretability of the model, facilitating the generation of more precise responses to complex user inquiries across diverse environmental contexts. To assess the practical impact of our proposed wearable platform, we carried out a series of real-world experiments across three tasks that are commonly challenging for people with blindness and low vision: risk assessment, object localization, and scene recognition. Additionally, through an ablation study conducted on the VizWiz dataset, we rigorously assess the contribution of each individual module, substantiating the integral role in the model’s overall performance.
Alexey Magay, Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006
IROS6
2024 GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
abstract
Zero-Shot Object Goal Navigation (ZS-OGN) enables robots to navigate toward objects of unseen categories without prior training. Traditional approaches often leverage categorical semantic information for navigation guidance, which struggles when only partial objects are observed or detailed and functional representations of the environment are lacking. To resolve the above two issues, we propose \textit{Geometric-part and Affordance Maps} (GAMap), a novel method that integrates object parts and affordance attributes for navigation guidance. Our method includes a multi-scale scoring approach to capture geometric-part and affordance attributes of objects at different scales. Comprehensive experiments conducted on the HM3D and Gibson benchmark datasets demonstrate improvements in Success Rates and Success weighted by Path Length, underscoring the efficacy of our geometric-part and affordance-guided navigation approach in enhancing robot autonomy and versatility, without any additional task-specific training or fine-tuning with the semantics of unseen objects and/or the locomotions of the robot.
Shuaihang Yuan, Hao Huang 0003, Congcong Wen, Anthony Tzes, Yi Fang 0006
NeurIPS6
2024 A single 3D shape wavelet-based generative model
Hao Huang 0003, Shuaihang Yuan, Zheng Peng 0013, Congcong Wen, Yi Fang 0006
Comput. Graph.6
2024 Learning Signed Hyper Surfaces for Oriented Point Cloud Normal Estimation
abstract
We propose a novel method called SHS-Net for point cloud normal estimation by learning signed hyper surfaces, which can accurately predict normals with global consistent orientation from various point clouds. Almost all existing methods estimate oriented normals through a two-stage pipeline, i.e., unoriented normal estimation and normal orientation, and each step is implemented by a separate algorithm. However, previous methods are sensitive to parameter settings, resulting in poor results from point clouds with noise, density variations and complex geometries. In this work, we introduce signed hyper surfaces (SHS), which are parameterized by multi-layer perceptron (MLP) layers, to learn to estimate oriented normals from point clouds in an end-to-end manner. The signed hyper surfaces are implicitly learned in a high-dimensional feature space where the local and global information is aggregated. Specifically, we introduce a patch encoding module and a shape encoding module to encode a 3D point cloud into a local latent code and a global latent code, respectively. Then, an attention-weighted normal prediction module is proposed as a decoder, which takes the local and global latent codes as input to predict oriented normals. Experimental results show that our algorithm outperforms the state-of-the-art methods in both unoriented and oriented normal estimation.
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yue Gao 0002, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 CAP-UDF: Learning Unsigned Distance Functions Progressively From Raw Point Clouds With Consistency-Aware Field Optimization
abstract
Surface reconstruction for point clouds is an important task in 3D computer vision. Most of the latest methods resolve this problem by learning signed distance functions from point clouds, which are limited to reconstructing closed surfaces. Some other methods tried to represent open surfaces using unsigned distance functions (UDF) which are learned from ground truth distances. However, the learned UDF is hard to provide smooth distance fields due to the discontinuous character of point clouds. In this paper, we propose CAP-UDF, a novel method to learn consistency-aware UDF from raw point clouds. We achieve this by learning to move queries onto the surface with a field consistency constraint, where we also enable to progressively estimate a more accurate surface. Specifically, we train a neural network to gradually infer the relationship between queries and the approximated surface by searching for the moving target of queries in a dynamic way. Meanwhile, we introduce a polygonization algorithm to extract surfaces using the gradients of the learned UDF. We conduct comprehensive experiments in surface reconstruction for point clouds, real scans or depth maps, and further explore our performance in unsupervised point normal estimation, which demonstrate non-trivial improvements of CAP-UDF over the state-of-the-art methods.
Junsheng Zhou, Baorui Ma, Shujuan Li, Yu-Shen Liu, Yi Fang 0006, Zhizhong Han
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Learning to learn point signature for 3D shape geometry
Hao Huang 0003, Lingjing Wang, Xiang Li 0046, Shuaihang Yuan, Congcong Wen, Yi Fang 0006
Pattern Recognit. Lett.7
2023 SHS-Net: Learning Signed Hyper Surfaces for Oriented Normal Estimation of Point Clouds
abstract
We propose a novel method called SHS-Net for oriented normal estimation of point clouds by learning signed hyper surfaces, which can accurately predict normals with global consistent orientation from various point clouds. Almost all existing methods estimate oriented normals through a two-stage pipeline, i.e., unoriented normal estimation and normal orientation, and each step is implemented by a separate algorithm. However, previous methods are sensitive to parameter settings, resulting in poor results from point clouds with noise, density variations and complex geometries. In this work, we introduce signed hyper surfaces (SHS), which are parameterized by multi-layer perceptron (MLP) layers, to learn to estimate oriented normals from point clouds in an end-to-end manner. The signed hyper surfaces are implicitly learned in a high-dimensional feature space where the local and global information is aggregated. Specifically, we introduce a patch encoding module and a shape encoding module to encode a 3D point cloud into a local latent code and a global latent code, respectively. Then, an attention-weighted normal prediction module is proposed as a decoder, which takes the local and global latent codes as input to predict oriented normals. Experimental results show that our SHS-Net outperforms the state-of-the-art methods in both unoriented and oriented normal estimation on the widely used benchmarks. The code, data and pretrained models are available at https://github.com/LeoQLi/SHS-Net.
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yue Gao 0002, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
CVPR5
2023 LP-DIF: Learning Local Pattern-Specific Deep Implicit Function for 3D Objects and Scenes
abstract
Deep Implicit Function (DIF) has gained much popularity as an efficient 3D shape representation. To capture geometry details, current mainstream methods divide 3D shapes into local regions and then learn each one with a local latent code via a decoder. Such local methods can capture more local details due to less diversity among local regions than global shapes. Although the diversity of local regions has been decreased compared to global approaches, the diversity in different local regions still poses a challenge in learning an implicit function when treating all regions equally using only a single decoder. What is worse, these local regions often exhibit imbalanced distributions, where certain regions have significantly fewer observations. This leads that fine geometry details could not be preserved well. To solve this problem, we propose a novel Local Pattern-specific Implicit Function, named LP-DIF, to represent a shape with clusters of local regions and multiple decoders, where each decoder only focuses on one cluster of local regions which share a certain pattern. Specifically, we first extract local codes for all regions, and then cluster them into multiple groups in the latent space, where similar regions sharing a common pattern fall into one group. After that, we train multiple decoders for mining local patterns of different groups, which simplifies the learning of fine geometric details by reducing the diversity of local regions seen by each decoder. To further alleviate the data-imbalance problem, we introduce a region re-weighting module to each pattern-specific decoder using a kernel density estimator, which dynamically re-weights the regions during learning. Our LP-DIF can restore more geometry details, and thus improve the quality of 3D reconstruction. Experiments demonstrate that our method can achieve the state-of-the-art performance over previous methods. Code is available at https://github.com/gtyxyz/lpdif.
Meng Wang 0001, Yu-Shen Liu, Yue Gao 0002, Kanle Shi, Yi Fang 0006, Zhizhong Han
CVPR5
2023 Retro-FPN: Retrospective Feature Pyramid Network for Point Cloud Semantic Segmentation
abstract
Learning per-point semantic features from the hierarchical feature pyramid is essential for point cloud semantic segmentation. However, most previous methods suffered from ambiguous region features or failed to refine per-point features effectively, which leads to information loss and ambiguous semantic identification. To resolve this, we propose Retro-FPN to model the per-point feature prediction as an explicit and retrospective refining process, which goes through all the pyramid layers to extract semantic features explicitly for each point. Its key novelty is a retro-transformer for summarizing semantic contexts from the previous layer and accordingly refining the features in the current stage. In this way, the categorization of each point is conditioned on its local semantic pattern. Specifically, the retro-transformer consists of a local cross-attention block and a semantic gate unit. The cross-attention serves to summarize the semantic pattern retrospectively from the previous layer. And the gate unit carefully incorporates the summarized contexts and refines the current semantic features. Retro-FPN is a pluggable neural network that applies to hierarchical decoders. By integrating Retro-FPN with three representative backbones, including both point-based and voxel-based methods, we show that Retro-FPN can significantly improve performance over state-of-the-art backbones. Comprehensive experiments on widely used benchmarks can justify the effectiveness of our design. The source is available at https://github.com/AllenXiangX/Retro-FPN.
Peng Xiang 0002, Xin Wen 0003, Yu-Shen Liu, Hui Zhang 0013, Yi Fang 0006, Zhizhong Han
ICCV5
2023 3D Unsupervised Region-Aware Registration Transformer
abstract
This paper concerns the research problem of point cloud registration to find the rigid transformation to optimally align the source point set with the target one. Learning robust point cloud registration models with deep neural networks has emerged as a powerful paradigm, offering promising performance in predicting the global geometric transformation for a pair of point sets. Existing methods first leverage an encoder to regress the global shape descriptor, which is then decoded into a shape-conditioned transformation via concatenation-based conditioning. However, different regions of a 3D shape vary in their geometric structures which makes it more sense that we have a region-conditioned transformation instead of the shape-conditioned one. In this paper, we define our 3D registration function through the introduction of a new design of 3D region partition module that is able to divide the input shape to different regions with a self-supervised 3D shape reconstruction loss without the need for ground truth labels. We further propose the 3D shape transformer module to efficiently and effectively capture short-and long-range geometric dependencies for regions on the 3D shape Consequently, the region-aware decoder module is proposed to predict the transformations for different regions respectively. The global geometric transformation from the source point set to the target one is then formed by the weighted fusion of region-aware transformation. Compared to the state-of-the-art approaches, our experiments show that our 3D-URRT achieves superior registration performance over various benchmark datasets (e.g. ModelNet40).
Yi Fang 0006
ICIP2
2023 Pyramid Learnable Tokens for 3D LiDAR Place Recognition
abstract
3D LiDAR place recognition plays a vital role in various robot applications' including robotic navigation, autonomous driving, and simultaneous localization and mapping. However, most previous studies evaluated their models on accumulated 2D scans instead of real-world 3D LiDAR scans with a larger number of points, which limits the application in real scenarios. To address this limitation, we propose a point transformer network with pyramid learnable tokens (PTNet-PLT) to learn global descriptors for an actual scanned 3D LiDAR place recognition. Specifically, we first present a novel shifted cube attention module that consists of a self-attention module for local feature extraction and a cross-attention module for regional feature aggregation. The self-attention module constrains attention computation on a locally partitioned cube and builds connections across cubes based on the shifted cube scheme. In addition, the cross-attention module introduces several learnable tokens to separately aggregate features of points with similar features but spatially distant into an arbitrarily shaped region, which enables the model to capture long-term dependencies of the points. Next, we build a pyramid architecture network to learn multi-scale features and involve a decreasing number of tokens at each layer to aggregate features over a larger region. Finally, we obtain the global descriptor by concatenating learned region tokens of all layers. Experiments on three datasets, including USyd Campus, Oxford Robot-Car, and KITTI, demonstrate the effectiveness and generalization of the proposed model for large-scale 3D LiDAR place recognition.
Congcong Wen, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006
ICRA4
2023 Understanding the Impact of Image Quality and Distance of Objects to Object Detection Performance
abstract
Object detection is a fundamental task for autonomous driving, which aim to identify and localize objects within an image. Deep learning has made great strides for object detection, with popular models including Faster R-CNN, YOLO, and SSD. The detection accuracy and computational cost of object detection depend on the spatial resolution of an image, which may be constrained by both the camera and storage considerations. Furthermore, original images are often compressed and uploaded to a remote server for object detection. Compression is often achieved by reducing either spatial or amplitude resolution or, at times, both, both of which have well-known effects on performance. Detection accuracy also depends on the distance of the object of interest from the camera. Our work examines the impact of spatial and amplitude resolution, as well as object distance, on object detection accuracy and computational cost. As existing models are optimized for uncompressed (or lightly compressed) images over a narrow range of spatial resolution, we develop a resolution-adaptive variant of YOLOv5 (RA-YOLO), which varies the number of scales in the feature pyramid and detection head based on the spatial resolution of the input image. To train and evaluate this new method, we created a dataset of images with diverse spatial and amplitude resolutions by combining images from the TJU and Eurocity datasets and generating different resolutions by applying spatial resizing and compression. We first show that RA-YOLO achieves a good trade-off between detection accuracy and inference time over a large range of spatial resolutions. We then evaluate the impact of spatial and amplitude resolutions on object detection accuracy using the proposed RA-YOLO model. We demonstrate that the optimal spatial resolution that leads to the highest detection accuracy depends on the ‘tolerated’ image size (constrained by the available bandwidth or storage). We further assess the impact of the distance of an object to the camera on the detection accuracy and show that higher spatial resolution enables a greater detection range. These results provide important guidelines for choosing the image spatial resolution and compression settings predicated on available bandwidth, storage, desired inference time, and/or desired detection range, in practical applications.
Haoyang Pei, Yixuan Lyu, Zhongzheng Yuan, John-Ross Rizzo, Yao Wang 0001, Yi Fang 0006
IROS7
2023 NeuralGF: Unsupervised Point Normal Estimation by Learning Neural Gradient Function
abstract
Normal estimation for 3D point clouds is a fundamental task in 3D geometry processing. The state-of-the-art methods rely on priors of fitting local surfaces learned from normal supervision. However, normal supervision in benchmarks comes from synthetic shapes and is usually not available from real scans, thereby limiting the learned priors of these methods. In addition, normal orientation consistency across shapes remains difficult to achieve without a separate post-processing procedure. To resolve these issues, we propose a novel method for estimating oriented normals directly from point clouds without using ground truth normals as supervision. We achieve this by introducing a new paradigm for learning neural gradient functions, which encourages the neural network to fit the input point clouds and yield unit-norm gradients at the points. Specifically, we introduce loss functions to facilitate query points to iteratively reach the moving targets and aggregate onto the approximated surface, thereby learning a global surface representation of the data. Meanwhile, we incorporate gradients into the surface approximation to measure the minimum signed deviation of queries, resulting in a consistent gradient field associated with the surface. These techniques lead to our deep unsupervised oriented normal estimator that is robust to noise, outliers and density variations. Our excellent results on widely used benchmarks demonstrate that our method can learn more accurate normals for both unoriented and oriented normal estimation tasks than the latest methods. The source code and pre-trained model are publicly available.
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yue Gao 0002, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
NeurIPS5
2023 Differentiable Registration of Images and LiDAR Point Clouds with VoxelPoint-to-Pixel Matching
abstract
Cross-modality registration between 2D images captured by cameras and 3D point clouds from LiDARs is a crucial task in computer vision and robotic. Previous methods estimate 2D-3D correspondences by matching point and pixel patterns learned by neural networks, and use Perspective-n-Points (PnP) to estimate rigid transformation during post-processing. However, these methods struggle to map points and pixels to a shared latent space robustly since points and pixels have very different characteristics with patterns learned in different manners (MLP and CNN), and they also fail to construct supervision directly on the transformation since the PnP is non-differentiable, which leads to unstable registration results. To address these problems, we propose to learn a structured cross-modality latent space to represent pixel features and 3D features via a differentiable probabilistic PnP solver. Specifically, we design a triplet network to learn VoxelPoint-to-Pixel matching, where we represent 3D elements using both voxels and points to learn the cross-modality latent space with pixels. We design both the voxel and pixel branch based on CNNs to operate convolutions on voxels/pixels represented in grids, and integrate an additional point branch to regain the information lost during voxelization. We train our framework end-to-end by imposing supervisions directly on the predicted pose distribution with a probabilistic PnP solver. To explore distinctive patterns of cross-modality features, we design a novel loss with adaptive-weighted optimization for cross-modality feature description. The experimental results on KITTI and nuScenes datasets show significant improvements over the state-of-the-art methods.
Junsheng Zhou, Baorui Ma, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
NeurIPS4
2023 Neural Gradient Learning and Optimization for Oriented Point Normal Estimation
abstract
We propose Neural Gradient Learning (NGL), a deep learning approach to learn gradient vectors with consistent orientation from 3D point clouds for normal estimation. It has excellent gradient approximation properties for the underlying geometry of the data. We utilize a simple neural network to parameterize the objective function to produce gradients at points using a global implicit representation. However, the derived gradients usually drift away from the ground-truth oriented normals due to the lack of local detail descriptions. Therefore, we introduce Gradient Vector Optimization (GVO) to learn an angular distance field based on local plane geometry to refine the coarse gradient vectors. Finally, we formulate our method with a two-phase pipeline of coarse estimation followed by refinement. Moreover, we integrate two weighting functions, i.e., anisotropic kernel and inlier score, into the optimization to improve the robust and detail-preserving performance. Our method efficiently conducts global gradient approximation while achieving better accuracy and generalization ability of local feature description. This leads to a state-of-the-art normal estimator that is robust to noise, outliers and point density variations. Extensive evaluations show that our method outperforms previous works in both unoriented and oriented normal estimation on widely used benchmarks. The source code and pre-trained models are available at https://github.com/LeoQLi/NGLO .
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
SIGGRAPH Asia4
2023 Unsupervised Category-Specific Partial Point Set Registration via Joint Shape Completion and Registration
abstract
We propose a self-supervised method for partial point set registration. Although recently proposed learning-based methods demonstrate impressive registration performance on full shape observations, these methods often suffer from performance degradation when dealing with partial shapes. To bridge the performance gap between partial and full point set registration, we propose to incorporate a shape completion network to benefit the registration process. To achieve this, we introduce a learnable latent code for each pair of shapes, which can be regarded as the geometric encoding of the target shape. By doing so, our model does not require an explicit feature embedding network to learn the feature encodings. More importantly, both our shape completion and point set registration networks take the shared latent codes as input, which are optimized simultaneously with the parameters of two decoder networks in the training process. Therefore, the point set registration process can benefit from the joint optimization process of latent codes, which are enforced to represent the information of full shapes instead of partial ones. In the inference stage, we fix the network parameters and optimize the latent codes to obtain the optimal shape completion and registration results. Our proposed method is purely unsupervised and does not require ground truth supervision. Experiments on the ModelNet40 dataset demonstrate the effectiveness of our model for partial point set registration.
Xiang Li 0046, Lingjing Wang, Yi Fang 0006
IEEE Trans. Vis. Comput. Graph.3
2022 Meta-Det3D: Learn to Learn Few-Shot 3D Object Detection
Shuaihang Yuan, Xiang Li 0046, Hao Huang 0003, Yi Fang 0006
ACCV (1)4
2022 Manifold Adversarial Learning for Cross-domain 3D Shape Representation
Hao Huang 0003, Cheng Chen 0078, Yi Fang 0006
ECCV (26)3
2022 Non-Rigid Multiple Point Set Registration Using Latent Gaussian Mixture
abstract
Point set registration is a fundamental task in 3D computer vision. Existing registration approaches mainly focus on either pair-wise or rigid registration. In this paper, we propose a robust group-wise registration method from a probabilistic view and adopt non-rigid transformations to register multiple point sets without bias toward any given set. The proposed method lessens the need of point correspondences by representing each point set as Gaussian Mixture Model and the registration is equivalent to multiple distributions alignment. Closed-form of Jensen-Rényi divergence and L2distance are used as cost functions. We further design a neural network to extract correspondences between raw point sets and Gaussian Mixture Model (GMM) parameters, and recover the optimal diffeomorphic non-rigid transformations from the matched GMM parameters. The proposed method is compared against two well-known probabilistic methods for group-wise point-set registration on several public 2D and 3D datasets. The results demonstrate that our method improves registration accuracy.
Hao Huang 0003, Cheng Chen 0078, Yi Fang 0006
ICIP3
2022 Adaptive Wavelet Transformer Network for 3D Shape Representation Learning
Hao Huang 0003, Yi Fang 0006
ICLR2
2022 3D Point Cloud Completion with Geometric-Aware Adversarial Augmentation
abstract
With the popularity of 3D sensors in self-driving and other robotics applications, extensive research has focused on designing novel neural network architectures for accurate 3D point cloud completion. However, unlike point cloud classification and reconstruction, the role of adversarial samples in 3D point cloud completion has seldom been explored. In this work, we demonstrate that adversarial samples can benefit neural networks on 3D point cloud completion tasks. We propose a novel approach to craft adversarial samples that improve the performance of models on both clean and adversarial inputs. In contrast to the Projected Gradient Descent (PGD) attack, our method generates adversarial samples that keep the geometric features in clean samples and contain few outliers. In particular, we use minimum absolute curvature directions to constrain the adversarial perturbations for each input point. The gradient components in the minimum absolute curvature directions are taken as adversarial perturbations. In addition, we adopt attack strength accumulation and auxiliary Batch Normalization layers to speed up the training process and alleviate the distribution mismatch between clean and adversarial samples. Experimental results demonstrate that training with the adversarial samples crafted by our method under the geometric-aware constraint effectively enhances the performance of the Point Completion Network (PCN) on the ShapeNet dataset.
Mengxi Wu, Hao Huang 0003, Yi Fang 0006
ICPR3
2022 Implicit Neural Representations for Medical Imaging Segmentation
Muhammad Osama Khan, Yi Fang 0006
MICCAI (5)2
2022 HSurf-Net: Normal Estimation for 3D Point Clouds by Learning Hyper Surfaces
abstract
We propose a novel normal estimation method called HSurf-Net, which can accurately predict normals from point clouds with noise and density variations. Previous methods focus on learning point weights to fit neighborhoods into a geometric surface approximated by a polynomial function with a predefined order, based on which normals are estimated. However, fitting surfaces explicitly from raw point clouds suffers from overfitting or underfitting issues caused by inappropriate polynomial orders and outliers, which significantly limits the performance of existing methods. To address these issues, we introduce hyper surface fitting to implicitly learn hyper surfaces, which are represented by multi-layer perceptron (MLP) layers that take point features as input and output surface patterns in a high dimensional feature space. We introduce a novel space transformation module, which consists of a sequence of local aggregation layers and global shift layers, to learn an optimal feature space, and a relative position encoding module to effectively convert point clouds into the learned feature space. Our model learns hyper surfaces from the noise-less features and directly predicts normal vectors. We jointly optimize the MLP weights and module parameters in a data-driven manner to make the model adaptively find the most suitable surface pattern for various points. Experimental results show that our HSurf-Net achieves the state-of-the-art performance on the synthetic shape dataset, the real-world indoor and outdoor scene datasets. The code, data and pretrained models are publicly available.
Qing Li 0032, Yu-Shen Liu, Jin-San Cheng, Cheng Wang 0003, Yi Fang 0006, Zhizhong Han
NeurIPS5
2022 Learning Consistency-Aware Unsigned Distance Functions Progressively from Raw Point Clouds
abstract
Surface reconstruction for point clouds is an important task in 3D computer vision. Most of the latest methods resolve this problem by learning signed distance functions (SDF) from point clouds, which are limited to reconstructing shapes or scenes with closed surfaces. Some other methods tried to represent shapes or scenes with open surfaces using unsigned distance functions (UDF) which are learned from large scale ground truth unsigned distances. However, the learned UDF is hard to provide smooth distance fields near the surface due to the noncontinuous character of point clouds. In this paper, we propose a novel method to learn consistency-aware unsigned distance functions directly from raw point clouds. We achieve this by learning to move 3D queries to reach the surface with a field consistency constraint, where we also enable to progressively estimate a more accurate surface. Specifically, we train a neural network to gradually infer the relationship between 3D queries and the approximated surface by searching for the moving target of queries in a dynamic way, which results in a consistent field around the surface. Meanwhile, we introduce a polygonization algorithm to extract surfaces directly from the gradient field of the learned UDF. The experimental results in surface reconstruction for synthetic and real scan data show significant improvements over the state-of-the-art under the widely used benchmarks.
Junsheng Zhou, Baorui Ma, Yu-Shen Liu, Yi Fang 0006, Zhizhong Han
NeurIPS4
2022 SHREC 2022: Protein-ligand binding site recognition
Luca Gagliardi, Andrea Raffo, Ulderico Fugacci, Silvia Biasotti, Walter Rocchia, Hao Huang 0003, Boulbaba Ben Amor, Yi Fang 0006, Charles Christoffer, Daisuke Kihara, Apostolos Axenopoulos, Stelios K. Mylonas, Petros Daras
Comput. Graph.8
2022 SHREC'22 track: Sketch-based 3D shape retrieval in the wild
Jie Qin 0004, Shuaihang Yuan, Jiaxin Chen 0002, Boulbaba Ben Amor, Yi Fang 0006, Nhat Hoang-Xuan, Chi-Bien Chu, Khoi-Nguyen Nguyen-Ngoc, Thien-Tri Cao, Nhat-Khang Ngô, Tuan-Luc Huynh, Hai-Dang Nguyen, Minh-Triet Tran, Haoyang Luo, Jianning Wang, Zheng Zhang 0006, Zihao Xin, Yang Wang 0023, Haiqin Chen, Qunying Zhou
Comput. Graph.5
2022 Height Estimation From Single Aerial Images Using a Deep Ordinal Regression Network
abstract
Understanding the 3-D geometric structure of the Earth’s surface has been an active research topic in photogrammetry and remote sensing community for decades, serving as an essential building block for various applications such as 3-D digital city modeling, change detection, and city management. Previous research studies have extensively studied the problem of height estimation from aerial images based on stereo or multiview image matching. These methods require two or more images from different perspectives to reconstruct 3-D coordinates with camera information provided. In this letter, we deal with the ambiguous and unsolved problem of height estimation from a single aerial image. Driven by the great success of deep learning, especially deep convolutional neural networks (CNNs), some research studies have proposed to estimate height information from a single aerial image by training a deep CNN model with large-scale annotated data sets. These methods treat height estimation as a regression problem and directly use an encoder–decoder network to regress the height values. In this letter, we propose to divide height values into spacing-increasing intervals and transform the regression problem into an ordinal regression problem, using an ordinal loss for network training. To enable multiscale feature extraction, we further incorporate an Atrous Spatial Pyramid Pooling (ASPP) module to extract features from multiple dilated convolution layers. After that, a postprocessing technique is designed to transform the predicted height map of each patch into a seamless height map. Finally, we conduct extensive experiments on International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam data sets. Experimental results demonstrate significantly better performance of our method compared to state-of-the-art methods.
Xiang Li 0046, Yi Fang 0006
IEEE Geosci. Remote. Sens. Lett.3
2022 Geometry-Aware Segmentation of Remote Sensing Images via Joint Height Estimation
abstract
Recent studies have shown the benefits of using additional elevation data [e.g., digital surface model (DSM) or normalized DSM (nDSM)] for enhancing the performance of the semantic labeling of aerial images. However, previous methods mostly adopt 3-D elevation information as additional inputs, while, in many real-world applications, one does not have the corresponding DSM images at hand, and the spatial resolution of acquired DSM images usually does not match the aerial images. To alleviate this data constraint and also take advantage of 3-D elevation information, in this letter, a geometry-aware segmentation model is introduced to achieve accurate semantic labeling of aerial images via joint height estimation. Instead of using a single-stream encoder–decoder network for semantic labeling, we design a separate decoder branch to predict the height map and use the DSM images as side supervision to train this newly designed decoder branch. With the newly designed decoder branch, our model can distill the 3-D geometric features from 2-D appearance features under the supervision of ground-truth DSM images. Moreover, we develop a new geometry-aware convolution module that fuses the 3-D geometric features from the height decoder branch and the 2-D contextual features from the semantic segmentation branch. The fused feature embeddings can produce geometry-aware segmentation maps with enhanced performance. Our model is trained with DSM images as side supervision, while, in the inference stage, it does not require DSM data and directly predicts the semantic labels. Experiments on International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam data sets demonstrate the effectiveness of the proposed method for the semantic segmentation of aerial images.
Xiang Li 0046, Congcong Wen, Lingjing Wang, Yi Fang 0006
IEEE Geosci. Remote. Sens. Lett.4
2022 Few-Shot Object Detection on Remote Sensing Images
abstract
In this article, we deal with the problem of object detection on remote sensing images. Previous researchers have developed numerous deep convolutional neural network (CNN)-based methods for object detection on remote sensing images, and they have reported remarkable achievements in detection performance and efficiency. However, current CNN-based methods often require a large number of annotated samples to train deep neural networks and tend to have limited generalization abilities for unseen object categories. In this article, we introduce a metalearning-based method for few-shot object detection on remote sensing images where only a few annotated samples are needed for the unseen object categories. More specifically, our model contains three main components: a metafeature extractor that learns to extract metafeature maps from input images, a feature reweighting module that learns class-specific reweighting vectors from the support images and use them to recalibrate the metafeature maps, and a bounding box prediction module that carries out object detection on the reweighted feature maps. We build our few-shot object detection model upon the YOLOv3 architecture and develop a multiscale object detection framework. Experiments on two benchmark data sets demonstrate that with only a few annotated samples, our model can still achieve a satisfying detection performance on remote sensing images, and the performance of our model is significantly better than the well-established baseline models.
Xiang Li 0046, Jingyu Deng, Yi Fang 0006
IEEE Trans. Geosci. Remote. Sens.3
2022 GP-Aligner: Unsupervised Groupwise Nonrigid Point Set Registration Based on Optimizable Group Latent Descriptor
abstract
In this paper, we propose a novel unsupervised method named GP-Aligner to address the problem of groupwise non-rigid point set registration. Compared to previous non-learning-based approaches, the proposed method gains competitive advantages by leveraging deep neural networks to effectively and efficiently align a large number of highly deformed 3D shapes with superior performance. Unlike most learning-based methods that use an explicit feature encoding network to extract per-shape features and their correlations, our model leverages a model-free learnable latent descriptor to characterize shape correlations among groups. More specifically, for a given group we first define an optimizable Group Latent Descriptor (GLD) to characterize the relationship among a group of point sets. Each GLD is randomly initialized from a Gaussian distribution and then concatenated with the coordinates of each point of the associated point sets in the group. A neural network-based decoder network is further constructed to predict the coherent flow fields to optimally deform the input groups of shapes to the aligned ones. During the optimization process, GP-Aligner jointly updates all GLDs and weight parameters of the decoder network towards the minimization of an unsupervised groupwise alignment loss. After optimization, for each group, our model coherently drives each point set towards a mean position (shape) without specifying one as the target. GP-Aligner does not require large-scale training data for network training, and it can directly align groups of point sets in a one-stage optimization process. GP-Aligner shows both accuracy and computational efficiency improvement in comparison with the state-of-the-art methods for groupwise point set registration. Moreover, GP-Aligner exhibits high efficiency in aligning a large number of groups of real-world 3D shapes.
Lingjing Wang, Hao Huang 0003, Jifei Wang, Xiang Li 0046, Yi Fang 0006
IEEE Trans. Geosci. Remote. Sens.6
2021 3D-MetaConNet: Meta-learning for 3D Shape Classification and Segmentation
abstract
Supervised learning on 3D shapes are extensively studied by prior literature, among which PointNet [29] and its variants PointNet++ [31] are representatives. However, these methods tackle 3D shape learning problems by training from scratch using a fixed learning algorithm over large amounts of labeled data, potentially challenged by data and computation bottlenecks. In the paper, we design a novel model, under the framework of meta-learning, to learn 3D shape representation. By training over multiple 3D tasks, each of which is defined as a supervised learning problem, our method can fast adapt to unseen tasks containing limited labeled data. Specifically, our model consists of a 3Dmeta-learner and a task-oriented 3D-learner, where the 3D-meta-learner produces parameter initialization for the 3D-learner after being trained over different tasks. With adaptively initialized parameters, the 3D-learner can be tuned rapidly in a few steps to achieve good performance on novel tasks with a small amount of training data. To further facilitate discriminative shape feature learning, we introduce a novel task-aware feature adaptation module under a contrastive learning scheme, in which all shapes in each task are considered as a whole and task-oriented compact features are learned. Therefore, we dub our model as 3DMetaConNet. Experiments on three public 3D datasets for few-shot shape classification and segmentation demonstrate that our method can learn compact and discriminative 3D shape features efficiently and robustly in a fast adaptation manner. Our method particularly outperforms the methods without a meta-learning framework and is also superior to existing meta-learning approaches.
Hao Huang 0003, Xiang Li 0046, Lingjing Wang, Yi Fang 0006
3DV4
2021 Topology Constrained Shape Correspondence
abstract
To better address the deformation and structural variation challenges inherently present in 3D shapes, researchers have shifted their focus from designing handcrafted point descriptors to learning point descriptors and their correspondences in a data-driven manner. Recent studies have developed deep neural networks for robust point descriptor and shape correspondence learning in consideration of local structural information. In this article, we developed a novel shape correspondence learning network, called TC-NET, which further enhances performance by encouraging the topological consistency between the embedding feature space and the input shape space. Specifically, in this article, we first calculate the topology-associated edge weights to represent the topological structure of each point. Then, in order to preserve this topological structure in high-dimensional feature space, a structural regularization term is defined to minimize the topology-consistent feature reconstruction loss (Topo-Loss) during the correspondence learning process. Our proposed method achieved state-of-the-art performance on three shape correspondence benchmark datasets. In addition, the proposed topology preservation concept can be easily generalized to other learning-based shape analysis tasks to regularize the topological structure of high-dimensional feature spaces.
Xiang Li 0046, Congcong Wen, Lingjing Wang, Yi Fang 0006
IEEE Trans. Vis. Comput. Graph.4
2020 Robust Image Matching By Dynamic Feature Selection
Hao Huang 0003, Jianchun Chen, Xiang Li 0046, Lingjing Wang, Yi Fang 0006
BMVC5
2020 MDA-Net: Memorable Domain Adaptation Network for Monocular Depth Estimation
Jing Zhu 0002, Yunxiao Shi, Mengwei Ren, Yi Fang 0006
BMVC4
2020 Unsupervised Deep Shape Descriptor With Point Distribution Learning
abstract
Deep learning models have achieved great success in supervised shape descriptor learning for 3D shape retrieval, classification, and correspondence. However, the unsupervised shape descriptor calculated via deep learning is less studied than that of supervised ones due to the design challenges of unsupervised neural network architecture. This paper proposes a novel probabilistic framework for the learning of unsupervised deep shape descriptors with point distribution learning. In our approach, we firstly associate each point with a Gaussian, and the point clouds are modeled as the distribution of the points. We then use deep neural networks (DNNs) to model a maximum likelihood estimation process that is traditionally solved with an iterative Expectation-Maximization (EM) process. Our key novelty is that ``training'' these DNNs with unsupervised self-correspondence L2 distance loss will elegantly reveal the statically significant deep shape descriptor representation for the distribution of the point clouds. We have conducted experiments over various 3D datasets. Qualitative and quantitative comparisons demonstrate that our proposed method achieves superior classification performance over existing unsupervised 3D shape descriptors. In addition, we verified the following attractive properties of our shape descriptor through experiments: multi-scale shape representation, robustness to shape rotation, and robustness to noise.
Yi Shi 0008, Mengchen Xu, Shuaihang Yuan, Yi Fang 0006
CVPR4
2020 Few-Shot Learning of Part-Specific Probability Space for 3D Shape Segmentation
abstract
Recently, deep neural networks are introduced as supervised discriminative models for the learning of 3D point cloud segmentation. Most previous supervised methods require a large number of training data with human annotation part labels to guide the training process to ensure the model's generalization abilities on test data. In comparison, we propose a novel 3D shape segmentation method that requires few labeled data for training. Given an input 3D shape, the training of our model starts with identifying a similar 3D shape with part annotations from a mini-pool of shape templates (e.g. 10 shapes). With the selected template shape, a novel Coherent Point Transformer is proposed to fully leverage the power of a deep neural network to smoothly morph the template shape towards the input shape. Then, based on the transformed template shapes with part labels, a newly proposed Part-specific Density Estimator is developed to learn a continuous part-specific probability distribution function on the entire 3D space with a batch consistency regularization term. With the learned part-specific probability distribution, our model is able to predict the part labels of a new input 3D shape in an end-to-end manner. We demonstrate that our proposed method can achieve remarkable segmentation results on the ShapeNet dataset with few shots, compared to previous supervised learning approaches.
Lingjing Wang, Xiang Li 0046, Yi Fang 0006
CVPR3
2020 3DMotion-Net: Learning Continuous Flow Function for 3D Motion Prediction
abstract
This paper deals with predicting future 3D motions of 3D object scans from the previous two consecutive frames. Previous methods mostly focus on sparse motion prediction in the form of skeletons. While in this paper, we focus on predicting dense 3D motions in the form of 3D point clouds. To approach this problem, we propose a self-supervised approach that leverages the power of the deep neural network to learn a continuous flow function of 3D point clouds that can predict temporally consistent future motions and naturally bring out the correspondences among consecutive point clouds at the same time. More specifically, in our approach, to eliminate the unsolved and challenging process of defining a discrete point convolution on 3D point cloud sequences to encode spatial and temporal information, we introduce a learnable latent code to represent the temporal-aware shape descriptor, which is optimized during the model training. Moreover, a temporally consistent motion Morpher is proposed to learn a continuous flow field which deforms a 3D scan from the current frame to the next frame. We perform extensive experiments on D-FAUST, SCAPE, and TOSCA benchmark data sets. The results demonstrate that our approach is capable of handling temporally inconsistent input and produces consistent future 3D motion while requiring no ground truth supervision.
Shuaihang Yuan, Xiang Li 0046, Anthony Tzes, Yi Fang 0006
IROS4
2020 ROSS: Robust Learning of One-Shot 3D Shape Segmentation
abstract
3D shape segmentation is a fundamental computer vision task that partitions the object into labeled semantic parts. Recent approaches to 3D shape segmentation learning heavily rely on high-quality labeled training datasets. This limits their use in applications to handle the large scale unannotated datasets. In this paper, we proposed a novel semi-supervised approach, named Robust Learning of OneShot 3D Shape Segmentation (ROSS), which only requires one single exemplar labeled shape for training. The proposed ROSS can generalize its ability from a one-shot training process to predict the segmentation for previously unseen 3D shape models. The proposed ROSS is composed of three major modules for 3D shape segmentation as follows. The global shape descriptor generator is the first module that utilizes the proposed reference weighted convolution to learn a 3D shape descriptor. The second module is a part-aware shape descriptor constructor that can generate weighted descriptors from a learned 3D shape descriptor according to semantic parts without supervision. The shape morphing with label transferring works as the last module. It morphs the exemplar shape and then transfers labels from the transformed exemplar shape to the target shape. The extensive experimental results on 3D mesh datasets demonstrate the ROSS is robust to noise and incomplete shapes and it can be applied to unannotated datasets. The experiment shows the proposed ROSS can achieve comparable performance with the supervised method.
Shuaihang Yuan, Yi Fang 0006
WACV2
2020 Reference Grid-assisted Network for 3D Point Signature Learning from Point Clouds
abstract
Learning a robust 3D point signature from point clouds is an interesting but challenging task in the computer vision field due to the irregular and unordered structure characteristics of the point cloud data. In this paper, we propose to learn a 3D point signature by exploring the implicit relation between keypoints and their neighbors (grouped as patches) among the given scene point clouds. We design a uniform reference grid to represent the raw relation between each keypoint and its neighbors from the raw point clouds. In order to learn a 3D point signature gradually from expanding perceptive region, we create a novel siamese framework with a multi-layer perceptron (MLP)-based unit feature network and a 3D convolutional neural network (CNN)-based grid feature network. Specifically, the unit feature network aims to dig the connections among points fallen into the same unit of the reference grid, while the grid feature network is used to discover the grid-wise relations across the whole reference grid with concatenation of the learned unit- wise features. Moreover, we introduce an attention network upon the unit feature network to enhance the discriminative ability of our learned 3D point signature. Our proposed 3D point signature achieves superior performance over other state-of-the-art methods on keypoint matching and geometric registration on the real-world scenes datasets, e.g. SUN3D, 7-scenes and the synthetic scan augmented scenes in ICL-NUIM dataset. More importantly, our learned 3D point signature successfully handles the point cloud fragment alignment challenges by producing correct transformations with RANSAC algorithm.
Jing Zhu 0002, Yi Fang 0006
WACV2
2019 PC-Net: Unsupervised Point Correspondence Learning with Neural Networks
abstract
Point sets correspondence concerns with the establishment of point-wise correspondence for a group of 2D or 3D point sets with similar shape description. Existing methods often iteratively search for the optimal point-wise correspondence assignment for two sets of points, driven by maximizing the similarity between two sets of explicitly designed point features or by determining the parametric transformation for the best alignment between two point sets. In contrast, without depending on the explicit definitions of point features or transformation, our paper introduces a novel point correspondence neural networks (PC-Net) that is able to learn and predict the point correspondence among the populations of a specific object (e.g. fish, human, chair, etc) in an unsupervised manner. Specifically, in this paper, we first develop an encoder to learn the shape descriptor from a point set that captures essential global and deformation-insensitive geometric properties. Then followed with a novel motion-driven process, our PC-Net drives a template shape, that consists of a set of landmark points, morph and conform around a target shape object which is reconstructed through decoding the previously characterized shape descriptor. As a result, the motion-driven process progressively and coherently drifts all landmark points from the template shape to corresponding positions on the target object shape. The experimental results demonstrate that PC-Net can establish robust unsupervised point correspondence over a group of deformable object shapes in the presence of geometric noise and missing points. More importantly, with great generalization capability, PC-Net is capable of instantly predicting group point corresponding for unseen point sets.
Xiang Li 0046, Lingjing Wang, Yi Fang 0006
3DV3
2019 Pairwise Attention Encoding for Point Cloud Feature Learning
abstract
Compared to hand-crafted ones, learning a 3D point signature has attracted increasing attention in the research community to better address challenging issues such as deformation and structural variation in 3D objects. PointNet is a pioneering work in introducing learning 3D point signature directly by consuming raw point cloud as input and applying convolution on each one of these points. Ground-breaking as it is, PointNet has limited capability in capturing local structure when learning visual features from each individual point. Recent variants of PointNet improved the quality of 3D point signature learning by taking neighbourhood information into account, but typically do so through hard-coded mechanisms (e.g. manually setting 'k' for k-Nearest Neighbour search, radius 'r' for Ball Query, etc). In this paper, we developed a novel point signature learning approach by considering pairwise interaction between every two individual points that moves beyond hard-coded neighbourhood exploitation, which further improves the quality of 3D point signature learning by encouraging the model to be aware of both neighbourhood information and global context. Specifically, we first introduce a novel pairwise reference tensor (PRT) in the original input point space to represent the influence of every two individual points that have on each other. Then, by passing the pairwise reference tensor through a multi-layer perceptron (MLP), we obtain a high-dimensional attention tensor that encodes pairwise relationships in high dimensional space that acts as an attention mechanism. Next we further fuse learned point features with the attention weights to obtain global visual features. Our proposed method has demonstrated superior performance on various 3D visual recognition tasks (e.g. object classification, part segmentation and scene semantic segmentation).
Yunxiao Shi, Haoyu Fang, Jing Zhu 0002, Yi Fang 0006
3DV4
2019 Learning Object-Specific Distance From a Monocular Image
abstract
Environment perception, including object detection and distance estimation, is one of the most crucial tasks for autonomous driving. Many attentions have been paid on the object detection task, but distance estimation only arouse few interests in the computer vision community. Observing that the traditional inverse perspective mapping algorithm performs poorly for objects far away from the camera or on the curved road, in this paper, we address the challenging distance estimation problem by developing the first end-to-end learning-based model to directly predict distances for given objects in the images. Besides the introduction of a learning-based base model, we further design an enhanced model with a keypoint regressor, where a projection loss is defined to enforce a better distance estimation, especially for objects close to the camera. To facilitate the research on this task, we construct the extented KITTI and nuScenes (mini) object detection datasets with a distance for each object. Our experiments demonstrate that our proposed methods outperform alternative approaches (e.g., the traditional IPM, SVR) on object-specific distance estimation, particularly for the challenging cases that objects are on a curved road. Moreover, the performance margin implies the effectiveness of our enhanced method.
Jing Zhu 0002, Yi Fang 0006
ICCV2
2019 An assistive low-vision platform that augments spatial cognition through proprioceptive guidance: Point-to-Tell-and-Touch
abstract
Spatial cognition, as gained through the sense of vision, is one of the most important capabilities of human beings. However, for the visually impaired (VI), lack of this perceptual capability poses great challenges in their life. Therefore, we have designed Point-to-Tell-and-Touch, a wearable system with an ergonomic human-machine interface, for assisting the VI with active environmental exploration, with a particular focus on spatial intelligence and navigation to objects of interest in an alien environment. Our key idea is to link visual signals, as decoded synthetically, to the VI's proprioception for more intelligible guidance, in addition to vision-to-audio assistance, i.e., finger pose, as indicated by pointing, is used as “proprioceptive laser pointer” to target an object in that line of sight. The whole system consists of two features, Point-to-Tell and Point-to-Touch, both of which can work independently or cooperatively. The Point-to-Tell feature contains a camera with a novel one-stage neural network tailored for blind-centered object detection and recognition, and a headphone telling the VI the semantic label and distance from the pointed object. the Point-to-Touch, the second feature, leverages a vibrating wrist band to create a haptic feedback tool that supplements the initial vectorial guidance provided by the first stage (hand pose being direction and the distance being the extent, offered through audio cues). Both platform features utilize proprioception or joint position sense. Through hand pose, the VI end user knows where he or she is pointing relative to their egocentric coordinate system and we are able to use this foundation to build spatial intelligence. Our successful indoor experiments demonstrate the proposed system to be effective and reliable in helping the VI gain spatial cognition and explore the world in a more intuitive way.
Wenjun Gui, Shuaihang Yuan, John-Ross Rizzo, Lakshay Sharma, Chen Feng 0002, Anthony Tzes, Yi Fang 0006
IROS8
2019 Arbicon-Net: Arbitrary Continuous Geometric Transformation Networks for Image Registration
abstract
This paper concerns the undetermined problem of estimating geometric transformation between image pairs. Recent methods introduce deep neural networks to predict the controlling parameters of hand-crafted geometric transformation models (e.g. thin-plate spline) for image registration and matching. However, the low-dimension parametric models are incapable of estimating a highly complex geometric transform with limited flexibility to model the actual geometric deformation from image pairs. To address this issue, we present an end-to-end trainable deep neural networks, named Arbitrary Continuous Geometric Transformation Networks (Arbicon-Net), to directly predict the dense displacement field for pairwise image alignment. Arbicon-Net is generalized from training data to predict the desired arbitrary continuous geometric transformation in a data-driven manner for unseen new pair of images. Particularly, without imposing penalization terms, the predicted displacement vector function is proven to be spatially continuous and smooth. To verify the performance of Arbicon-Net, we conducted semantic alignment tests over both synthetic and real image dataset with various experimental settings. The results demonstrate that Arbicon-Net outperforms the previous image alignment techniques in identifying the image correspondences.
Jianchun Chen, Lingjing Wang, Xiang Li 0046, Yi Fang 0006
NeurIPS4
2019 Deep Learning for Pattern Recognition
Zhaoxiang Zhang 0001, Shiguang Shan, Yi Fang 0006, Ling Shao 0001
Pattern Recognit. Lett.3
2019 Learning domain-invariant feature for robust depth-image-based 3D shape retrieval
Jing Zhu 0002, John-Ross Rizzo, Yi Fang 0006
Pattern Recognit. Lett.3
2018 Learning Adversarial 3D Model Generation With 2D Image Enhancer
abstract
Recent advancements in generative adversarial nets (GANs) and volumetric convolutional neural networks (CNNs) enable generating 3D models from a probabilistic space. In this paper, we have developed a novel GAN-based deep neural network to obtain a better latent space for the generation of 3D models. In the proposed method, an enhancer neural network is introduced to extract information from other corresponding domains (e.g. image) to improve the performance of the 3D model generator, and the discriminative power of the unsupervised shape features learned from the 3D model discriminator. Specifically, we train the 3D generative adversarial networks on 3D volumetric models, and at the same time, the enhancer network learns image features from rendered images. Different from the traditional GAN architecture that uses uninformative random vectors as inputs, we feed the high-level image features learned from the enhancer into the 3D model generator for better training. The evaluations on two large-scale 3D model datasets, ShapeNet and ModelNet, demonstrate that our proposed method can not only generate high-quality 3D models, but also successfully learn discriminative shape representation for classification and retrieval without supervision.
Jing Zhu 0002, Jin Xie 0001, Yi Fang 0006
AAAI3
2018 Deep Cross-Modality Adaptation via Semantics Preserving Adversarial Learning for Sketch-Based 3D Shape Retrieval
Jiaxin Chen 0002, Yi Fang 0006
ECCV (13)2
2018 Siamese CNN-BiLSTM Architecture for 3D Shape Representation Learning
abstract
Learning a 3D shape representation from a collection of its rendered 2D images has been extensively studied. However, existing view-based techniques have not yet fully exploited the information among all the views of projections. In this paper, by employing recurrent neural network to efficiently capture features across different views, we propose a siamese CNN-BiLSTM network for 3D shape representation learning. The proposed method minimizes a discriminative loss function to learn a deep nonlinear transformation, mapping 3D shapes from the original space into a nonlinear feature space. In the transformed space, the distance of 3D shapes with the same label is minimized, otherwise the distance is maximized to a large margin. Specifically, the 3D shapes are first projected into a group of 2D images from different views. Then convolutional neural network (CNN) is adopted to extract features from different view images, followed by a bidirectional long short-term memory (LSTM) to aggregate information across different views. Finally, we construct the whole CNN-BiLSTM network into a siamese structure with contrastive loss function. Our proposed method is evaluated on two benchmarks, ModelNet40 and SHREC 2014, demonstrating superiority over the state-of-the-art methods.
Guoxian Dai, Jin Xie 0001, Yi Fang 0006
IJCAI3
2018 Learning to Synthesize 3D Indoor Scenes from Monocular Images
abstract
Depth images have always been playing critical roles for indoor scene understanding problems, and are particularly important for tasks in which 3D inferences are involved. However, since depth images are not universally available, abandoning them from the testing stage can significantly improve the generality of a method. In this work, we consider the scenarios where depth images are not available in the testing data, and propose to learn a convolutional long short-term memory (Conv LSTM) network and a regression convolutional neural network (regression ConvNet) using only monocular RGB images. The proposed networks benefit from 2D segmentations, object-level spatial context, object-scene dependencies and objects' geometric information, where optimization is governed by the semantic label loss, which measures the label consistencies of both objects and scenes, and the 3D geometrical loss, which measures the correctness of objects' 6Dof estimation. Conv LSTM and regression ConvNet are applied to scene/object classification, object detection and 6Dof estimation tasks respectively, where we utilize the joint inference from both networks and further provide the perspective of synthesizing fully rigged 3D scenes according to objects' arrangements in monocular images. Both quantitative and qualitative experimental results are provided on the NYU-v2 dataset, and we demonstrate that the proposed Conv LSTM can achieve state-of-the-art performance without requiring the depth information.
Fan Zhu 0001, Li Liu 0004, Jin Xie 0001, Fumin Shen, Ling Shao 0001, Yi Fang 0006
ACM Multimedia6
2018 Unsupervised Learning of 3D Model Reconstruction from Hand-Drawn Sketches
abstract
3D objects modeling has gained considerable attention in the visual computing community. We propose a low-cost unsupervised learning model for 3D objects reconstruction from hand-drawn sketches. Recent advancements in deep learning opened new opportunities to learn high-quality 3D objects from 2D sketches via supervised networks. However, the limited availability of labeled 2D hand-drawn sketches data (i.e. sketches and its corresponding 3D ground truth models) hinders the training process of supervised methods. In this paper, driven by a novel design of combination of retrieval and reconstruction process, we developed a learning paradigm to reconstruct 3D objects from hand-drawn sketches, without the use of well-labeled hand-drawn sketch data during the entire training process. Specifically, the paradigm begins with the training of an adaption network via autoencoder with adversarial loss, embedding the unpaired 2D rendered image domain with the hand-drawn sketch domain to a shared latent vector space. Then from the embedding latent space, for each testing sketch image, we retrieve a few (e.g. five) nearest neighbors from the training 3D data set as prior knowledge for a 3D Generative Adversarial Network. Our experiments verify our network's robust and superior performance in handling 3D volumetric object generation from single hand-drawn sketch without requiring any 3D ground truth labels.
Lingjing Wang, Jifei Wang, Yi Fang 0006
ACM Multimedia4
2018 Learning Local Descriptors with Adversarial Enhancer from Volumetric Geometry Patches
abstract
Local matching problems (e.g. key point matching, geometry registration) are significant but challenging tasks in computer vision field. In this paper, we propose to learn a robust local 3D descriptor from volumetric point patches to tackle the local matching tasks. Intuitively, given two inputs, it would be easy for a network to map the inputs to a space with similar characteristics (e.g. similar outputs for similar inputs, far different outputs for far different inputs), but the difficult case for a network would be to map the inputs into a space with opposite characteristics (e.g. far different outputs for very similar inputs but very similar outputs for far different inputs). Inspired by this intuition, in our proposed method, we design a siamese-network-based local descriptor generator to learn a local descriptor with small distances between match pairs and large distances between non-match pairs. Specifically, an adversarial enhancer is introduced to map the outputs of the local descriptor generator into an opposite space that match pairs have the maximum differences and non-match pairs have the minimum differences. The local descriptor generator and the adversarial enhancer are trained in an adversarial manner. By competing with the adversarial enhancer, the local descriptor generator learns to generate a much stronger descriptor for given volumetric point patches. The experiments conducted on real-world scan datasets, including 7-scenes and SUN3D, and the synthetic scan augmented ICL-NUIM dataset show that our method can achieve superior performance over other state-of-the-art approaches on both keypoint matching and geometry registration, such as fragment alignment and scene reconstruction.
Jing Zhu 0002, Yi Fang 0006
ACM Multimedia2
2018 Deep Nonlinear Metric Learning for 3-D Shape Retrieval
abstract
Effective 3-D shape retrieval is an important problem in 3-D shape analysis. Recently, feature learning-based shape retrieval methods have been widely studied, where the distance metrics between 3-D shape descriptors are usually hand-crafted. In this paper, motivated by the fact that deep neural network has the good ability to model nonlinearity, we propose to learn an effective nonlinear distance metric between 3-D shape descriptors for retrieval. First, the locality-constrained linear coding method is employed to encode each vertex on the shape and the encoding coefficient histogram is formed as the global 3-D shape descriptor to represent the shape. Then, a novel deep metric network is proposed to learn a nonlinear transformation to map the 3-D shape descriptors to a nonlinear feature space. The proposed deep metric network minimizes a discriminative loss function that can enforce the similarity between a pair of samples from the same class to be small and the similarity between a pair of samples from different classes to be large. Finally, the distance between the outputs of the metric network is used as the similarity for shape retrieval. The proposed method is evaluated on the McGill, SHREC'10 ShapeGoogle, and SHREC'14 Human shape datasets. Experimental results on the three datasets validate the effectiveness of the proposed method.
Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Ling Shao 0001, Yi Fang 0006
IEEE Trans. Cybern.5
2018 Greedy Criterion in Orthogonal Greedy Learning
abstract
Orthogonal greedy learning (OGL) is a stepwise learning scheme that starts with selecting a new atom from a specified dictionary via the steepest gradient descent (SGD) and then builds the estimator through orthogonal projection. In this paper, we found that SGD is not the unique greedy criterion and introduced a new greedy criterion, called as " -greedy threshold" for learning. Based on this new greedy criterion, we derived a straightforward termination rule for OGL. Our theoretical study shows that the new learning scheme can achieve the existing (almost) optimal learning rate of OGL. Numerical experiments are also provided to support that this new scheme can achieve almost optimal generalization performance while requiring less computation than OGL.
Lin Xu 0001, Shaobo Lin, Jinshan Zeng, Yi Fang 0006, Zongben Xu
IEEE Trans. Cybern.5
2018 Deep Correlated Holistic Metric Learning for Sketch-Based 3D Shape Retrieval
abstract
How to effectively retrieve desired 3D models with simple queries is a long-standing problem in computer vision community. The model-based approach is quite straightforward but nontrivial, since people could not always have the desired 3D query model available by side. Recently, large amounts of wide-screen electronic devices are prevail in our daily lives, which makes the sketch-based 3D shape retrieval a promising candidate due to its simpleness and efficiency. The main challenge of sketch-based approach is the huge modality gap between sketch and 3D shape. In this paper, we proposed a novel deep correlated holistic metric learning (DCHML) method to mitigate the discrepancy between sketch and 3D shape domains. The proposed DCHML trains two distinct deep neural networks (one for each domain) jointly, which learns two deep nonlinear transformations to map features from both domains into a new feature space. The proposed loss, including discriminative loss and correlation loss, aims to increase the discrimination of features within each domain as well as the correlation between different domains. In the new feature space, the discriminative loss minimizes the intra-class distance of the deep transformed features and maximizes the inter-class distance of the deep transformed features to a large margin within each domain, while the correlation loss focused on mitigating the distribution discrepancy across different domains. Different from existing deep metric learning methods only with loss at the output layer, our proposed DCHML is trained with loss at both hidden layer and output layer to further improve the performance by encouraging features in the hidden layer also with desired properties. Our proposed method is evaluated on three benchmarks, including 3D Shape Retrieval Contest 2013, 2014, and 2016 benchmarks, and the experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods.
Guoxian Dai, Jin Xie 0001, Yi Fang 0006
IEEE Trans. Image Process.3
2017 Deep Correlated Metric Learning for Sketch-based 3D Shape Retrieval
abstract
The explosive growth of 3D models has led to the pressing demand for an efficient searching system. Traditional model-based search is usually not convenient, since people don't always have 3D model available by side. The sketch-based 3D shape retrieval is a promising candidate due to its simpleness and efficiency. The main challenge for sketch-based 3D shape retrieval is the discrepancy across different domains. In the paper, we propose a novel deep correlated metric learning (DCML) method to mitigate the discrepancy between sketch and 3D shape domains. The proposed DCML trains two distinct deep neural networks (one for each domain) jointly with one loss, which learns two deep nonlinear transformations to map features from both domains into a nonlinear feature space. The proposed loss, including discriminative loss and correlation loss, aims to increase the discrimination of features within each domain as well as the correlation between different domains. In the transfered space, the discriminative loss minimizes the intra-class distance of the deep transformed features and maximizes the inter-class distance of the deep transformed features at least a predefined margin within each domain, while the correlation loss focuses on minimizing the distribution discrepancy across different domains. Our proposed method is evaluated on SHREC 2013 and 2014 benchmarks, and the experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods.
Guoxian Dai, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006
AAAI4
2017 Learning Barycentric Representations of 3D Shapes for Sketch-Based 3D Shape Retrieval
abstract
Retrieving 3D shapes with sketches is a challenging problem since 2D sketches and 3D shapes are from two heterogeneous domains, which results in large discrepancy between them. In this paper, we propose to learn barycenters of 2D projections of 3D shapes for sketch-based 3D shape retrieval. Specifically, we first use two deep convolutional neural networks (CNNs) to extract deep features of sketches and 2D projections of 3D shapes. For 3D shapes, we then compute the Wasserstein barycenters of deep features of multiple projections to form a barycentric representation. Finally, by constructing a metric network, a discriminative loss is formulated on the Wasserstein barycenters of 3D shapes and sketches in the deep feature space to learn discriminative and compact 3D shape and sketch features for retrieval. The proposed method is evaluated on the SHREC13 and SHREC14 sketch track benchmark datasets. Compared to the state-of-the-art methods, our proposed method can significantly improve the retrieval performance.
Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Yi Fang 0006
CVPR4
2017 Metric-based Generative Adversarial Network
abstract
Existing methods of generative adversarial network (GAN) use different criteria to distinguish between real and fake samples, such as probability [9],energy [44] energy or other losses [30]. In this paper, by employing the merits of deep metric learning, we propose a novel metric-based generative adversarial network (MBGAN), which uses the distance-criteria to distinguish between real and fake samples. Specifically, the discriminator of MBGAN adopts a triplet structure and learns a deep nonlinear transformation, which maps input samples into a new feature space. In the transformed space, the distance between real samples is minimized, while the distance between real sample and fake sample is maximized. Similar to the adversarial procedure of existing GANs, a generator is trained to produce synthesized examples, which are close to real examples, while a discriminator is trained to maximize the distance between real and fake samples to a large margin. Meanwhile, instead of using a fixed margin, we adopt a data-dependent margin [30], so that the generator could focus on improving the synthesized samples with poor quality, instead of wasting energy on well-produce samples. Our proposed method is verified on various benchmarks, such as CIFAR-10, SVHN and CelebA, and generates high-quality samples.
Guoxian Dai, Jin Xie 0001, Yi Fang 0006
ACM Multimedia3
2017 3DensiNet: A Robust Neural Network Architecture towards 3D Volumetric Object Prediction from 2D Image
abstract
3D volumetric object generation/prediction from single 2D image is a quite challenging but meaningful task in 3D visual computing. In this paper, we propose a novel neural network architecture, named "3DensiNet", which uses density heat-map as an intermediate supervision tool for 2D-to-3D transformation. Specifically, we firstly present a 2D density heat-map to 3D volumetric object encoding-decoding network, which outperforms classical 3D autoencoder. Then we show that using 2D image to predict its density heat-map via a 2D to 2D encoding-decoding network is feasible. In addition, we leverage adversarial loss to fine tune our network, which improves the generated/predicted 3D voxel objects to be more similar to the ground truth voxel object. Experimental results on 3D volumetric prediction from 2D images demonstrates superior performance of 3DensiNet over other state-of-the-art techniques in handling 3D volumetric object generation/prediction from single 2D image.
Meng Wang 0001, Lingjing Wang, Yi Fang 0006
ACM Multimedia3
2017 DeepShape: Deep-Learned Shape Descriptor for 3D Shape Retrieval
abstract
Complex geometric variations of 3D models usually pose great challenges in 3D shape matching and retrieval. In this paper, we propose a novel 3D shape feature learning method to extract high-level shape features that are insensitive to geometric deformations of shapes. Our method uses a discriminative deep auto-encoder to learn deformation-invariant shape features. First, a multiscale shape distribution is computed and used as input to the auto-encoder. We then impose the Fisher discrimination criterion on the neurons in the hidden layer to develop a deep discriminative auto-encoder. Finally, the outputs from the hidden layers of the discriminative auto-encoders at different scales are concatenated to form the shape descriptor. The proposed method is evaluated on four benchmark datasets that contain 3D models with large geometric variations: McGill, SHREC'10 ShapeGoogle, SHREC'14 Human and SHREC'14 Large Scale Comprehensive Retrieval Track Benchmark datasets. Experimental results on the benchmark datasets demonstrate the effectiveness of the proposed method for 3D shape retrieval.
Jin Xie 0001, Guoxian Dai, Fan Zhu 0001, Edward K. Wong, Yi Fang 0006
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 Progressive Shape-Distribution-Encoder for Learning 3D Shape Representation
abstract
Since there are complex geometric variations with 3D shapes, extracting efficient 3D shape features is one of the most challenging tasks in shape matching and retrieval. In this paper, we propose a deep shape descriptor by learning shape distributions at different diffusion time via a progressive shape-distribution-encoder (PSDE). First, we develop a shape distribution representation with the kernel density estimator to characterize the intrinsic geometry structures of 3D shapes. Then, we propose to learn a deep shape feature through an unsupervised PSDE. Specially, the unsupervised PSDE aims at modeling the complex non-linear transform of the estimated shape distributions between consecutive diffusion time. In order to characterize the intrinsic structures of 3D shapes more efficiently, we stack multiple PSDEs to form a network structure. Finally, we concatenate all neurons in the middle hidden layers of the unsupervised PSDE network to form an unsupervised shape descriptor for retrieval. Furthermore, by imposing an additional constraint on the outputs of all hidden layers, we propose a supervised PSDE to form a supervised shape descriptor. For each hidden layer, the similarity between a pair of outputs from the same class is as large as possible and the similarity between a pair of outputs from different classes is as small as possible. The proposed method is evaluated on three benchmark 3D shape data sets with large geometric variations, i.e., McGill, SHREC'10 ShapeGoogle, and SHREC'14 Human data sets, and the experimental results demonstrate the superiority of the proposed method to the existing approaches.
Jin Xie 0001, Fan Zhu 0001, Guoxian Dai, Ling Shao 0001, Yi Fang 0006
IEEE Trans. Image Process.5
2017 Deep Multimetric Learning for Shape-Based 3D Model Retrieval
abstract
Recently, feature-learning-based 3D shape retrieval methods have been receiving more and more attention in the 3D shape analysis community. In these methods, the hand-crafted metrics or the learned linear metrics are usually used to compute the distances between shape features. Since there are complex geometric structural variations with 3D shapes, the single hand-crafted metric or learned linear metric cannot characterize the manifold, where 3D shapes lie well. In this paper, by exploring the nonlinearity of the deep neural network and the complementarity among multiple shape features, we propose a novel deep multimetric network for 3D shape retrieval. The developed multimetric network minimizes a discriminative loss function that, for each type of shape feature, the outputs of the network from the same class are encouraged to be as similar as possible and the outputs from different classes are encouraged to be as dissimilar as possible. Meanwhile, the Hilbert-Schmidt independence criterion is employed to enforce the outputs of different types of shape features to be as complementary as possible. Furthermore, the weights of the learned multiple distance metrics can be adaptively determined in our developed deep metric network. The weighted distance metric is then used as the similarity for shape retrieval. We conduct experiments with the proposed method on the four benchmark shape datasets. Experimental results demonstrate that the proposed method can obtain better performance than the learned deep single metric and outperform the state-of-the-art 3D shape retrieval methods.
Jin Xie 0001, Guoxian Dai, Yi Fang 0006
IEEE Trans. Multim.3
2016 Learning Cross-Domain Neural Networks for Sketch-Based 3D Shape Retrieval
abstract
Sketch-based 3D shape retrieval, which returns a set of relevant 3D shapes based on users' input sketch queries, has been receiving increasing attentions in both graphics community and vision community. In this work, we address the sketch-based 3D shape retrieval problem with a novel Cross-Domain Neural Networks (CDNN) approach, which is further extended to Pyramid Cross-Domain Neural Networks (PCDNN) by cooperating with a hierarchical structure. In order to alleviate the discrepancies between sketch features and 3D shape features, a neural network pair that forces identical representations at the target layer for instances of the same class is trained for sketches and 3D shapes respectively. By constructing cross-domain neural networks at multiple pyramid levels, a many-to-one relationship is established between a 3D shape feature and sketch features extracted from different scales. We evaluate the effectiveness of both CDNN and PCDNN approach on the extended large-scale SHREC 2014 benchmark and compare with some other well established methods. Experimental results suggest that both CDNN and PCDNN can outperform state-of-the-art performance, where PCDNN can further improve CDNN when employing a hierarchical structure.
Fan Zhu 0001, Jin Xie 0001, Yi Fang 0006
AAAI3
2016 Learned Binary Spectral Shape Descriptor for 3D Shape Correspondence
abstract
Dense 3D shape correspondence is an important problem in computer vision and computer graphics. Recently, the local shape descriptor based 3D shape correspondence approaches have been widely studied, where the local shape descriptor is a real-valued vector to characterize the geometrical structure of the shape. Different from these realvalued local shape descriptors, in this paper, we propose to learn a novel binary spectral shape descriptor with the deep neural network for 3D shape correspondence. The binary spectral shape descriptor can require less storage space and enable fast matching. First, based on the eigenvectors of the Laplace-Beltrami operator, we construct a neural network to form a nonlinear spectral representation to characterize the shape. Then, for the defined positive and negative points on the shapes, we train the constructed neural network by minimizing the errors between the outputs and their corresponding binary descriptors, minimizing the variations of the outputs of the positive points and maximizing the variations of the outputs of the negative points, simultaneously. Finally, we binarize the output of the neural network to form the binary spectral shape descriptor for shape correspondence. The proposed binary spectral shape descriptor is evaluated on the SCAPE and TOSCA 3D shape datasets for shape correspondence. The experimental results demonstrate the effectiveness of the proposed binary shape descriptor for the shape correspondence task.
Jin Xie 0001, Meng Wang 0001, Yi Fang 0006
CVPR3
2016 Heat Diffusion Long-Short Term Memory Learning for 3D Shape Analysis
Fan Zhu 0001, Jin Xie 0001, Yi Fang 0006
ECCV (7)3
2016 Local Diffusion Map Signature for Symmetry-aware Non-rigid Shape Correspondence
abstract
Identifying accurate correspondences information among different shapes is of great importance in shape analysis such as shape registration, segmentation and retrieval. This paper aims to develop a paradigm to address the challenging issues posed by shape structural variation and symmetry ambiguity. Specifically, the proposed research developed a novel shape signature based on local diffusion map on 3D surface, which is used to identify the shape correspondence through graph matching process. The developed shape signature, named local diffusion map signature (LDMS), is obtained by projecting heat diffusion distribution on 3D surface into 2D images along the surface normal direction with orientation determined by gradients of heat diffusion field. The local diffusion map signature is able to capture the concise geometric essence that is deformation-insensitive and symmetry-aware. Experimental results on 3D shape correspondence demonstrate the superior performance of our proposed method over other state-of-the-art techniques in identifying correspondences for non-rigid shapes with symmetry ambiguity.
Meng Wang 0001, Yi Fang 0006
ACM Multimedia2
2016 Global Consistent Shape Correspondence for Efficient and Effective Active Shape Models
abstract
Finding the accurate corresponded landmarks from a collection of shape instances plays critical role in constructing active shape models (ASMs). We have developed a global consistent shape correspondence paradigm for efficient and effective active shape models to address challenging issues in statistical shape modelling. Specifically, in this paper, we developed techniques to perform a fast multiple shape matching to identify global consistent shape correspondence from a set of training shape instances via efficient low-rank recovery optimization. High quality ASMs can then be constructed based on the identified corresponded points. The entire process is unsupervised without manual annotation as well as free of selection of anatomically significant point. Experimental results on mobile hand image data demonstrate the superior performance of our proposed method over other state-of-the-art techniques like MDL in constructing active shape models.
Meng Wang 0001, Yi Fang 0006
ACM Multimedia2
2016 Dynamic texture recognition with video set based collaborative representation
Jin Xie 0001, Yi Fang 0006
Image Vis. Comput.2
2016 Dual many-to-one-encoder-based transfer learning for cross-dataset human action recognition
Fan Zhu 0001, Edward K. Wong, Yi Fang 0006
Image Vis. Comput.4
2016 From handcrafted to learned representations for human action recognition: A survey
Fan Zhu 0001, Ling Shao 0001, Jin Xie 0001, Yi Fang 0006
Image Vis. Comput.4
2016 Learning a discriminative deformation-invariant 3D shape descriptor via many-to-one encoder
Guoxian Dai, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006
Pattern Recognit. Lett.4
2016 Linear discrimination dictionary learning for shape descriptors
Meng Wang 0001, Jin Xie 0001, Fan Zhu 0001, Yi Fang 0006
Pattern Recognit. Lett.4
2016 The Intrinsic Geometric Structure of Protein-Protein Interaction Networks for Protein Interaction Prediction
abstract
Recent developments in high-throughput technologies for measuring protein-protein interaction (PPI) have profoundly advanced our ability to systematically infer protein function and regulation. However, inherently high false positive and false negative rates in measurement have posed great challenges in computational approaches for the prediction of PPI. A good PPI predictor should be 1) resistant to high rate of missing and spurious PPIs, and 2) robust against incompleteness of observed PPI networks. To predict PPI in a network, we developed an intrinsic geometry structure (IGS) for network, which exploits the intrinsic and hidden relationship among proteins in network through a heat diffusion process. In this process, all explicit PPIs participate simultaneously to glue local infinitesimal and noisy experimental interaction data to generate a global macroscopic descriptions about relationships among proteins. The revealed implicit relationship can be interpreted as the probability of two proteins interacting with each other. The revealed relationship is intrinsic and robust against individual, local and explicit protein interactions in the original network. We apply our approach to publicly available PPI network data for the evaluation of the performance of PPI prediction. Experimental results indicate that, under different levels of the missing and spurious PPIs, IGS is able to robustly exploit the intrinsic and hidden relationship for PPI prediction with a higher sensitivity and specificity compared to that of recently proposed methods.
Yi Fang 0006, Mengtian Sun, Guoxian Dai, Karthik Ramani
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 3D deep shape descriptor
abstract
Shape descriptor is a concise yet informative representation that provides a 3D object with an identification as a member of some category. We have developed a concise deep shape descriptor to address challenging issues from ever-growing 3D datasets in areas as diverse as engineering, medicine, and biology. Specifically, in this paper, we developed novel techniques to extract concise but geometrically informative shape descriptor and new methods of defining Eigen-shape descriptor and Fisher-shape descriptor to guide the training of a deep neural network. Our deep shape descriptor tends to maximize the inter-class margin while minimize the intra-class variance. Our new shape descriptor addresses the challenges posed by the high complexity of 3D model and data representation, and the structural variations and noise present in 3D models. Experimental results on 3D shape retrieval demonstrate the superior performance of deep shape descriptor over other state-of-the-art techniques in handling noise, incompleteness and structural variations.
Yi Fang 0006, Jin Xie 0001, Guoxian Dai, Meng Wang 0001, Fan Zhu 0001, Edward K. Wong
CVPR1
2015 Deepshape: Deep learned shape descriptor for 3D shape matching and retrieval
abstract
Complex geometric structural variations of 3D model usually pose great challenges in 3D shape matching and retrieval. In this paper, we propose a high-level shape feature learning scheme to extract features that are insensitive to deformations via a novel discriminative deep auto-encoder. First, a multiscale shape distribution is developed for use as input to the auto-encoder. Then, by imposing the Fisher discrimination criterion on the neurons in the hidden layer, we developed a novel discriminative deep auto-encoder for shape feature learning. Finally, the neurons in the hidden layers from multiple discriminative auto-encoders are concatenated to form a shape descriptor for 3D shape matching and retrieval. The proposed method is evaluated on the representative datasets that contain 3D models with large geometric variations, i.e., Mcgill and SHREC'10 ShapeGoogle datasets. Experimental results on the benchmark datasets demonstrate the effectiveness of the proposed method for 3D shape matching and retrieval.
Jin Xie 0001, Yi Fang 0006, Fan Zhu 0001, Edward K. Wong
CVPR2
2015 Progressive Shape-Distribution-Encoder for 3D Shape Retrieval
abstract
In this paper, we propose a deep shape descriptor by learning the shape distributions at different diffusion time via a progressive deep shape-distribution-encoder. First, we develop a shape distribution representation with the kernel density estimator to characterize the intrinsic geometrical structure of the shape. Then, we propose to learn discriminative shape features through a progressive shape-distribution-encoder. Specially, the progressive shape-distribution-encoder aims at modeling the complex non-linear transform of the estimated shape distributions between consecutive diffusion time. Furthermore, in order to characterize the intrinsic structure of the shape more efficiently, we stack multiple proposed progressive shape-distribution-encoders to form a neural network structure. Finally, we concatenated all neurons in the hidden layers of the progressive shape-distribution-encoder network to form a discriminative shape descriptor for retrieval. The proposed method is evaluated on three benchmark 3D shape datasets %with large geometric variations, i.e., McGill, SHREC'10 ShapeGoogle and SHREC'14 Human datasets, and the experimental results demonstrate the superiority of our method to the existing approaches.
Jin Xie 0001, Fan Zhu 0001, Guoxian Dai, Yi Fang 0006
ACM Multimedia4
2015 Learning Pairwise Neural Network Encoder for Depth Image-based 3D Model Retrieval
abstract
With the emergence of RGB-D cameras (e.g., Kinect), the sensing capability of artificial intelligence systems has been dramatically increased, and as a consequence, a wide range of depth image-based human-machine interaction applications are proposed. In design industry, a 3D model always contains abundant information, which are required for manufacture. Since depth images can be conveniently acquired, a retrieval system that can return 3D models based on depth image inputs can assist or improve the traditional product design process. In this work, we address the depth image-based 3D model retrieval problem. By extending the neural network to a neural network pair with identical output layers for objects of the same category, unified domain-invariant representations can be learned based on the low-level mismatched depth image features and 3D model features. A unique advantage of the framework is that the correspondence information between depth images and 3D models are not required, so that it can easily be generalized to large-scale databases. In order to evaluate the effectiveness of our approach, depth images (with Kinect-type noise) in the NYU Depth V2 dataset are used as queries to retrieve 3D models of the same categories in the SHREC 2014 dataset. Experimental results suggest that our approach can outperform the state-of-the-arts methods, and the paradigm that directly uses the original representations of depth images and 3D models for retrieval.
Jing Zhu 0002, Fan Zhu 0001, Edward K. Wong, Yi Fang 0006
ACM Multimedia4
2014 Global Voting Model for Protein Function Prediction from Protein-Protein Interaction Networks
Yi Fang 0006, Mengtian Sun, Guoxian Dai, Karthik Ramani
ICIC (3)1
2014 The Intrinsic Geometric Structure of Protein-Protein Interaction Networks for Protein Interaction Prediction
Yi Fang 0006, Mengtian Sun, Guoxian Dai, Karthik Ramani
ICIC (3)1
2012 Center-Shift: An approach towards automatic robust mesh segmentation (ARMS)
abstract
In the area of 3D shape analysis, research in mesh segmentation has always been an important topic, as it is a fundamental low-level task which can be utilized in many applications including computer-aided design, computer animation, biomedical applications and many other fields. We define the automatic robust mesh segmentation (ARMS) method in this paper, which 1) is invariant to isometric transformation, 2) is insensitive to noise and deformation, 3) performs closely to human perception, 4) is efficient in computation, and 5) is minimally dependent on prior knowledge. In this work, we develop a new framework, namely the Center-Shift, which discovers meaningful segments of a 3D object by exploring the intrinsic geometric structure encoded in the biharmonic kernel. Our Center-Shift framework has three main steps: First, we construct a feature space where every vertex on the mesh surface is associated with the corresponding biharmonic kernel density function value. Second, we apply the Center-Shift algorithm for initial segmentation. Third, the initial segmentation result is refined through an efficient iterative process which leads to visually salient segmentation of the shape. The performance of this segmentation method is demonstrated through extensive experiments on various sets of 3D shapes and different types of noise and deformation. The experimental results of 3D shape segmentation have shown better performance of Center-Shift, compared to state-of-the-art segmentation methods.
Mengtian Sun, Yi Fang 0006, Karthik Ramani
CVPR2
2012 Towards locally and globally shape-aware reverse 3D modeling
Manish Goyal 0001, Sundar Murugappan, Cecil Piya, William Benjamin, Yi Fang 0006, Min Liu 0018, Karthik Ramani
Comput. Aided Des.5
2011 Heat-mapping: A robust approach toward perceptually consistent mesh segmentation
abstract
3D mesh segmentation is a fundamental low-level task with applications in areas as diverse as computer vision, computer-aided design, bio-informatics, and 3D medical imaging. A perceptually consistent mesh segmentation (PCMS), as defined in this paper is one that satisfies 1) in-variance to isometric transformation of the underlying surface, 2) robust to the perturbations of the surface, 3) robustness to numerical noise on the surface, and 4) close conformation to human perception. We exploit the intelligence of the heat as a global structure-aware message on a meshed surface and develop a robust PCMS scheme, called Heat-Mapping based on the heat kernel. There are three main steps in Heat-Mapping. First, the number of the segments is estimated based on the analysis of the behavior of the Laplacian spectrum. Second, the heat center, which is defined as the most representative vertex on each segment, is discovered by a proposed heat center hunting algorithm. Third, a heat center driven segmentation scheme reveals the PCMS with a high consistency towards human perception. Extensive experimental results on various types of models verify the performance of Heat-Mapping with respect to the consistent segmentation of articulated bodies, the topological changes, and various levels of numerical noise.
Yi Fang 0006, Mengtian Sun, Minhyong Kim, Karthik Ramani
CVPR1
2011 sLLE: Spherical locally linear embedding with applications to tomography
abstract
The tomographic reconstruction of a planar object from its projections taken at random unknown view angles is a problem that occurs often in medical imaging. Therefore, there is a need to robustly estimate the view angles given random observations of the projections. The widely used locally linear embedding (LLE) technique provides nonlinear embedding of points on a flat manifold. In our case, the projections belong to a sphere. Therefore, we extend LLE and develop a spherical locally linear embedding (sLLE) algorithm, which is capable of embedding data points on a non-flat spherically constrained manifold. Our algorithm, sLLE, transforms the problem of the angle estimation to a spherically constrained embedding problem. It considers each projection as a high dimensional vector with dimensionality equal to the number of sampling points on the projection. The projections are then embedded onto a sphere, which parametrizes the projections with respect to view angles in a globally consistent manner. The image is reconstructed from parametrized projections through the inverse Radon transform. A number of experiments demonstrate that sLLE is particularly effective for the tomography application we consider. We evaluate its performance in terms of the computational efficiency and noise tolerance, and show that sLLE can be used to shed light on the other constrained applications of LLE.
Yi Fang 0006, Mengtian Sun, S. V. N. Vishwanathan, Karthik Ramani
CVPR1
2009 Using least median of squares for structural superposition of flexible proteins
abstract
BACKGROUND: The conventional superposition methods use an ordinary least squares (LS) fit for structural comparison of two different conformations of the same protein. The main problem of the LS fit that it is sensitive to outliers, i.e. large displacements of the original structures superimposed. RESULTS: To overcome this problem, we present a new algorithm to overlap two protein conformations by their atomic coordinates using a robust statistics technique: least median of squares (LMS). In order to effectively approximate the LMS optimization, the forward search technique is utilized. Our algorithm can automatically detect and superimpose the rigid core regions of two conformations with small or large displacements. In contrast, most existing superposition techniques strongly depend on the initial LS estimating for the entire atom sets of proteins. They may fail on structural superposition of two conformations with large displacements. The presented LMS fit can be considered as an alternative and complementary tool for structural superposition. CONCLUSION: The proposed algorithm is robust and does not require any prior knowledge of the flexible regions. Furthermore, we show that the LMS fit can be extended to multiple level superposition between two conformations with several rigid domains. Our fit tool has produced successful superpositions when applied to proteins for which two conformations are known. The binary executable program for Windows platform, tested examples, and database are available from https://engineering.purdue.edu/PRECISE/LMSfit.
Yu-Shen Liu, Yi Fang 0006, Karthik Ramani
BMC Bioinform.2
2009 IDSS: deformation invariant signatures for molecular shape comparison
abstract
BACKGROUND: Many molecules of interest are flexible and undergo significant shape deformation as part of their function, but most existing methods of molecular shape comparison (MSC) treat them as rigid bodies, which may lead to incorrect measure of the shape similarity of flexible molecules. RESULTS: To address the issue we introduce a new shape descriptor, called Inner Distance Shape Signature (IDSS), for describing the 3D shapes of flexible molecules. The inner distance is defined as the length of the shortest path between landmark points within the molecular shape, and it reflects well the molecular structure and deformation without explicit decomposition. Our IDSS is stored as a histogram which is a probability distribution of inner distances between all sample point pairs on the molecular surface. We show that IDSS is insensitive to shape deformation of flexible molecules and more effective at capturing molecular structures than traditional shape descriptors. Our approach reduces the 3D shape comparison problem of flexible molecules to the comparison of IDSS histograms. CONCLUSION: The proposed algorithm is robust and does not require any prior knowledge of the flexible regions. We demonstrate the effectiveness of IDSS within a molecular search engine application for a benchmark containing abundant conformational changes of molecules. Such comparisons in several thousands per second can be carried out. The presented IDSS method can be considered as an alternative and complementary tool for the existing methods for rigid MSC. The binary executable program for Windows platform and database are available from https://engineering.purdue.edu/PRECISE/IDSS.
Yu-Shen Liu, Yi Fang 0006, Karthik Ramani
BMC Bioinform.2