EDBT 2026 Demo / reviewers in the wild / expert
Hao Huang 0003
dblp:04/5616-3
· DBLP profile ↗
32ranked-venue papers
12as first author
30since 2021 · last 2026
0000-0002-9131-5854ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 19 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vision-Only Gaussian Splatting for Collaborative Semantic Occupancy PredictionabstractCollaborative perception enables connected vehicles to share information, overcoming occlusions and extending the limited sensing range inherent in single-agent (non-collaborative) systems. Existing vision-only methods for 3D semantic occupancy prediction commonly rely on dense 3D voxels, which incur high communication costs, or 2D planar features, which require accurate depth estimation or additional supervision, limiting their applicability to collaborative scenarios. To address these challenges, we propose the first approach leveraging sparse 3D semantic Gaussian splatting for collaborative 3D semantic occupancy prediction. By sharing and fusing intermediate Gaussian primitives, our method provides three benefits: a neighborhood-based cross-agent fusion that removes duplicates and suppresses noisy or inconsistent Gaussians; a joint encoding of geometry and semantics in each primitive, which reduces reliance on depth supervision and allows simple rigid alignment; and sparse, object-centric messages that preserve structural information while reducing communication volume. Extensive experiments demonstrate that our approach outperforms single-agent perception and baseline collaborative methods by +8.42 and +3.28 points in mIoU, and +5.11 and +22.41 points in IoU, respectively. When further reducing the number of transmitted Gaussians, our method still achieves a +1.9 improvement in mIoU, using only 34.6% communication volume, highlighting robust performance under limited communication budgets. Cheng Chen 0078, Hao Huang 0003, Saurabh Bagchi |
AAAI | 2 |
| 2026 | On Demographic Group Fairness Guarantees in Deep LearningabstractWe present a theoretical framework analyzing the relationship between data distributions and fairness guarantees in deep learning. Our work establishes novel bounds that explicitly account for data distribution heterogeneity across demographic groups, while introducing a formal analysis framework that minimizes expected loss differences across these groups. Moreover, we derive bounds for fairness errors and convergence rates, characterizing how distributional differences between groups affect the fundamental trade-off between fairness and accuracy. Through extensive experiments on diverse datasets across various modalities (image, tabular data, and text), including FairVision (eye disease detection), CheXpert (pleural effusion detection), HAM10000 (skin lesion classification), FairFace (facial attribute recognition), ACS Income (income prediction), CivilComments-WILDS (toxic comment detection), we validate our theoretical findings and demonstrate that differences in feature distributions across demographic groups significantly impact model fairness, with performance disparities particularly pronounced in racial categories. The theoretical bounds we derive corroborate these empirical observations, providing insights into the fundamental limits of achieving fairness in deep learning models when faced with heterogeneous data distributions. This work advances our understanding of fairness in AI and provides a theoretical foundation for developing more equitable algorithms. Motivated by these theoretical insights, particularly the link between feature distribution shifts and fairness gaps, we propose Fairness-Aware Regularization (FAR), a practical training objective that directly minimizes inter-group discrepancies in feature centroids and covariances to improve equitable performance. We validate the effectiveness of FAR across all datasets considered in this study, consistently observing improvements in overall AUC, ES-AUC, and subgroup performance. Yan Luo 0002, Congcong Wen, Min Shi 0001, Hao Huang 0003, Yi Fang 0006, Mengyu Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | FairFedMed: Benchmarking Group Fairness in Federated Medical Imaging With FairLoRAabstractFairness remains a critical concern in healthcare, where unequal access to services and treatment outcomes can adversely affect patient health. While Federated Learning (FL) presents a collaborative and privacy-preserving approach to model training, ensuring fairness is challenging due to heterogeneous data across institutions, and current research primarily addresses non-medical applications. To fill this gap, we establish the first experimental benchmark for fairness in medical FL, evaluating six representative FL methods across diverse demographic attributes and imaging modalities. We introduce FairFedMed, the first medical FL dataset specifically designed to study group fairness (i.e., consistent performance across demographic groups). It comprises two parts: FairFedMed-Oph, featuring 2D fundus and 3D OCT ophthalmology samples with six demographic attributes; and FairFedMed-Chest, which simulates real cross-institutional FL using subsets of CheXpert and MIMIC-CXR. Together, they support both simulated and real-world FL across diverse medical modalities and demographic groups. Existing FL models often underperform on medical images and overlook fairness across demographic groups. To address this, we propose FairLoRA, a fairness-aware FL framework based on SVD-based low-rank approximation. It customizes singular value matrices per demographic group while sharing singular vectors, ensuring both fairness and efficiency. Experimental results on the FairFedMed dataset demonstrate that FairLoRA not only achieves state-of-the-art performance in medical image classification but also significantly improves fairness across diverse populations. Our code and dataset can be accessible via GitHub link: https://github.com/Harvard-AI-and-Robotics-Lab/FairFedMed. Minghan Li 0001, Congcong Wen, Yu Tian 0001, Min Shi 0001, Yan Luo 0002, Hao Huang 0003, Yi Fang 0006, Mengyu Wang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon TasksabstractPolicy learning focuses on devising strategies for agents in embodied artificial intelligence systems to perform optimal actions based on their perceived states. One of the key challenges in policy learning involves handling complex, long-horizon tasks that require managing extensive sequences of actions and observations with multiple modes. Wavelet analysis offers significant advantages in signal processing, notably in decomposing signals at multiple scales to capture both global trends and fine-grained details. In this work, we introduce a novel wavelet policy learning framework that utilizes wavelet transformations to enhance policy learning. Our approach leverages learnable multi-scale wavelet decomposition to facilitate detailed observation analysis and robust action planning over extended sequences. We detail the design and implementation of our wavelet policy, which incorporates lifting schemes for effective multi-resolution analysis and action generation. This framework is evaluated across multiple complex scenarios, including robotic manipulation, self-driving, and multi-robot collaboration, demonstrating the effectiveness of our method in improving the precision and reliability of the learned policy. Hao Huang 0003, Shuaihang Yuan, Geeta Chandra Raju Bethala, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
ICCV | 1 |
| 2025 | Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language ModelsabstractRobotic navigation plays a pivotal role in a wide range of real-world applications. While traditional navigation systems focus on efficiency and obstacle avoidance, their inability to model complex human behaviors in shared spaces has underscored the growing need for socially aware navigation. In this work, we explore a novel paradigm of socially aware robot navigation empowered by large language models (LLMs), and propose HSAC-LLM, a hybrid framework that seamlessly integrates deep reinforcement learning with the reasoning and communication capabilities of LLMs. Unlike prior approaches that passively predict pedestrian trajectories or issue pre-scripted alerts, HSAC-LLM enables bidirectional natural language interaction, allowing robots to proactively engage in dialogue with pedestrians to resolve potential conflicts and negotiate path decisions. Extensive evaluations across 2D simulations, Gazebo environments, and real-world deployments demonstrate that HSAC-LLM consistently outperforms state-of-the-art DRL baselines under our proposed socially aware navigation metric, which covers safety, efficiency, and human comfort. By bridging linguistic reasoning and interactive motion planning, our results highlight the potential of LLM-augmented agents for robust, adaptive, and human-aligned navigation in real-world settings. Project page: https://hsacllm.github.io/. Congcong Wen, Geeta Chandra Raju Bethala, Shuaihang Yuan, Hao Huang 0003, Mengyu Wang 0001, Yu-Shen Liu, Anthony Tzes, Yi Fang 0006 |
IROS | 5 |
| 2025 | SHREC 2025: Protein surface shape retrieval including electrostatic potentialabstractThis SHREC 2025 track dedicated to protein surface shape retrieval involved 9 participating teams. We evaluated the performance in retrieval of 15 proposed methods on a large dataset of 11,555 protein surfaces with calculated electrostatic potential (a key molecular surface descriptor). The performance in retrieval of the proposed methods was evaluated through different metrics (Accuracy, Balanced accuracy, F1 score, Precision and Recall). The best retrieval performance was achieved by the proposed methods that used the electrostatic potential complementary to molecular surface shape. This observation was also valid for classes with limited data which highlights the importance of taking into account additional molecular surface descriptors. Taher Yacoub, Camille Depenveiller, Atsushi Tatsuma, Tin Barisin, Eugen Rusakov, Udo Göbel, Yuxu Peng, Shiqiang Deng, Yuki Kagaya, Joon Hong Park, Daisuke Kihara, Marco Guerra, Giorgio Palmieri, Andrea Ranieri, Ulderico Fugacci, Silvia Biasotti, He Ruiwen, Halim Benhabiles, Adnane Cabani, Karim Hammoudi, Hao Huang 0003, Chunyan Li 0002, Alireza Tehrani, Fanwang Meng, Farnaz Heidar-Zadeh, Tuan-Anh Yang, Matthieu Montès |
Comput. Graph. | 22 |
| 2025 | 3D Shape Contrastive Representation Learning With Adversarial ExamplesabstractCurrent supervised methods for 3D shape representation learning have achieved satisfying performance, yet require extensive human-labeled datasets. Unsupervised learning-based methods provide a viable solution by learning shape representations without using ground truth labels. In this study, we develop a contrastive learning framework for unsupervised representation learning of 3D shapes. Specifically, in order to encourage models to pay more attention to useful information during representation learning, we first introduce a new paradigm for critical points search based on the adversarial mechanism. We extract critical points with a larger impact on the global feature by attacking a pre-trained auto-encoder model, and apply data augmentations on these points to generate adversarial examples. Taking a pair of adversarial examples as inputs, we obtain their intermediate embeddings and global representations of corresponding inputs, which are then transformed into latent spaces by two predictor heads. Finally, we train the proposed model by maximizing the agreements on these latent spaces via Normalized Temperature-scaled Cross Entropy (NT-Xent) loss and a newly designed Cross-layer Normalized Temperature-scaled Cross Entropy (Cross-NT-Xent) loss, where the latter is proposed in this paper to enforce cross-layer feature similarities. The effectiveness, robustness, and transferability of learned representations are validated on three downstream tasks, including object classification, few-shot classification, and shape retrieval. Experiments on three benchmark datasets show that our learned representations achieve better or competitive performance than current state-of-the-art methods in these downstream tasks. Moreover, our model can easily be extended to 3D part segmentation and scene segmentation tasks. Congcong Wen, Xiang Li 0046, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006 |
IEEE Trans. Multim. | 3 |
| 2024 | FairCLIP: Harnessing Fairness in Vision-Language LearningabstractFairness is a critical concern in deep learning, especially in healthcare, where these models influence diagnoses and treatment decisions. Although fairness has been investigated in the vision-only domain, the fairness of medical vision-language (VL) models remains unexplored due to the scarcity of medical VL datasets for studying fairness. To bridge this research gap, we introduce the first fair vision-language medical dataset (Harvard-FairVLMed) that provides detailed demographic attributes, ground-truth labels, and clinical notes to facilitate an in-depth examination of fairness within VL foundation models. Using Harvard-FairVLMed, we conduct a comprehensive fairness analysis of two widely-used VL models (CLIP and BLIP2), pre-trained on both natural and medical domains, across four different protected attributes. Our results highlight significant biases in all VL models, with Asian, Male, Non-Hispanic, and Spanish being the preferred subgroups across the protected attributes of race, gender, ethnicity, and language, respectively. In order to alleviate these biases, we propose FairCLIP an optimal-transport-based approach that achieves a favorable trade-off between performance and fairness by reducing the Sinkhorn distance between the overall sample distribution and the distributions corresponding to each demographic group. As the first VL dataset of its kind, Harvard-FairVLMed holds the potential to catalyze advancements in the development of machine learning models that are both ethically aware and clinically effective. Our dataset and code are available at https://ophai.hms.harvard.edu/datasets/harvard-fairvlmed10k. Yan Luo 0002, Min Shi 0001, Muhammad Osama Khan, Muhammad Muneeb Afzal, Hao Huang 0003, Shuaihang Yuan, Yu Tian 0001, Luo Song, Ava Kouhana, Tobias Elze, Yi Fang 0006, Mengyu Wang 0001 |
CVPR | 5 |
| 2024 | FairDomain: Achieving Fairness in Cross-Domain Medical Image Segmentation and Classification
Yu Tian 0001, Congcong Wen, Min Shi 0001, Muhammad Muneeb Afzal, Hao Huang 0003, Muhammad Osama Khan, Yan Luo 0002, Yi Fang 0006, Mengyu Wang 0001 |
ECCV (76) | 5 |
| 2024 | Optimizing Personalized Robot Actions with Ranking of Trajectories
Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
ICPR (29) | 1 |
| 2024 | Reliable Semantic Understanding for Real World Zero-Shot Object Goal Navigation
Halil Utku Unlu, Shuaihang Yuan, Congcong Wen, Hao Huang 0003, Anthony Tzes, Yi Fang 0006 |
ICPR (30) | 4 |
| 2024 | Zero-Shot Object Navigation with Vision-Language Models Reasoning
Congcong Wen, Yisiyuan Huang, Hao Huang 0003, Yanjia Huang, Shuaihang Yuan, Yu-Shen Liu, Yi Fang 0006 |
ICPR (18) | 3 |
| 2024 | Goal-Driven Transformer for Robot Behavior Learning from Play Data
Congcong Wen, Jiazhao Liang, Shuaihang Yuan, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006 |
ICPR (30) | 4 |
| 2024 | Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
Shuaihang Yuan, Halil Utku Unlu, Hao Huang 0003, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
ICPR (30) | 3 |
| 2024 | Noisy Few-shot 3D Point Cloud Scene Segmentationabstract3D scene semantic segmentation plays a crucial role in robotics by enabling robots to understand and interpret their environment in a detailed and context-aware manner, facilitating tasks such as navigation, object manipulation, and interaction within complex spaces. A preponderance of methodology predominantly adopts a fully supervised framework for 3D point cloud scene semantic segmentation. Such paradigms exhibit an intrinsic dependency on extensive labeled datasets, presenting challenges in acquisition and exhibiting incapacity to segment novel classes, especially when the training data are contaminated by noisy samples. To address these limitations, this study introduces a novel few-shot segmentation approach to robustly segment 3D point cloud scenes with noisy labels using a meta-learning scheme. Specifically, we first build a multi-prototype graph and then suppress samples with noisy labels based on the graph structure. A subgraph bagging scheme is then proposed to conduct semi-supervised transductive learning to propagate labels. To optimize the graph structure to learn discriminative prototype features, we design a triplet contrastive loss to increase the compactness of these subgraphs. We evaluated our method on two widely used 3D point cloud scene segmentation benchmarks within few-shot (i.e., 2/3-way 5-shot) segmentation settings with noisy samples. Experimental results demonstrate the improvement of our method over the compared baselines, illustrating the robustness of our method in few-shot 3D scene segmentation against noisy samples. The code is available at: https://github.com/hhuang-code/Noisy_Fewshot_Segmentation. Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
ICRA | 1 |
| 2024 | Weakly Scene Segmentation Using Efficient TransformerabstractCurrent methods for large-scale point cloud scene semantic segmentation rely on manually annotated dense point-wise labels, which are costly, labor-intensive, and prone to errors. Consequently, gathering point cloud scenes with billions of labeled points is impractical in real-world scenarios. In this paper, we introduce a novel weak supervision approach to semantically segment large-scale indoor scenes, requiring only 1‰ of the points to be labeled. Specifically, we develop an efficient point neighbor Transformer to capture the geometry of local point cloud patches. To address the quadratic complexity of self-attention computation in Transformers, particularly for large-scale point clouds, we propose approximating the self-attention matrix using low-rank and sparse decomposition. Building on the point neighbor Transformer as foundational blocks, we design a Low-rank Sparse Transformer Network (LST-Net) for weakly supervised large-scale point cloud scene semantic segmentation. Experimental results on two commonly used indoor point cloud scene segmentation benchmarks demonstrate that our model achieves performance comparable to those of both weakly supervised and fully supervised methods. Our code can be found in https://github.com/hhuang-code/LST-Net. Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
IROS | 1 |
| 2024 | ChatMap: A Wearable Platform Based on the Multi-modal Foundation Model to Augment Spatial Cognition for People with Blindness and Low VisionabstractSpatial cognition refers to the ability to gain knowledge about their surroundings and utilize this information to identify their location, acquire resources, and navigate their way back to familiar places. People with blindness and low vision (pBLV) face significant challenges with spatial cognition due to the reliance on visual input. Without the full range of visual cues, pBLV individuals often find it difficult to grasp a comprehensive understanding of their environment, leading to obstacles in scene recognition and precise object localization, especially in unfamiliar environments. This limitation extends to their ability to independently detect and avoid potential tripping hazards, making navigation and interaction with their environment more challenging. In this paper, we present a pioneering wearable platform tailored to enhance the spatial cognition of pBLV through the integration of multi-modal foundation model. The proposed platform integrates a wearable camera with audio module and leverages the advanced capabilities of vision language foundation model (i.e., GPT-4 and GPT-4V), for the nuanced processing of visual and textual data. Specifically, we employ vision language models to bridge the gap between visual information and the proprioception of visually impaired users, offering more intelligible guidance by aligning visual data with the natural perception of space and movement. Then we apply prompt engineering to guide the large language model to act as an assistant tailored specifically for pBLV users to produce accurate answers. Another innovation in our model is the incorporation of a chain of thought reasoning process, which enhances the accuracy and interpretability of the model, facilitating the generation of more precise responses to complex user inquiries across diverse environmental contexts. To assess the practical impact of our proposed wearable platform, we carried out a series of real-world experiments across three tasks that are commonly challenging for people with blindness and low vision: risk assessment, object localization, and scene recognition. Additionally, through an ablation study conducted on the VizWiz dataset, we rigorously assess the contribution of each individual module, substantiating the integral role in the model’s overall performance. Alexey Magay, Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
IROS | 3 |
| 2024 | GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance GuidanceabstractZero-Shot Object Goal Navigation (ZS-OGN) enables robots to navigate toward objects of unseen categories without prior training. Traditional approaches often leverage categorical semantic information for navigation guidance, which struggles when only partial objects are observed or detailed and functional representations of the environment are lacking. To resolve the above two issues, we propose \textit{Geometric-part and Affordance Maps} (GAMap), a novel method that integrates object parts and affordance attributes for navigation guidance. Our method includes a multi-scale scoring approach to capture geometric-part and affordance attributes of objects at different scales. Comprehensive experiments conducted on the HM3D and Gibson benchmark datasets demonstrate improvements in Success Rates and Success weighted by Path Length, underscoring the efficacy of our geometric-part and affordance-guided navigation approach in enhancing robot autonomy and versatility, without any additional task-specific training or fine-tuning with the semantics of unseen objects and/or the locomotions of the robot. Shuaihang Yuan, Hao Huang 0003, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
NeurIPS | 2 |
| 2024 | A single 3D shape wavelet-based generative model
Hao Huang 0003, Shuaihang Yuan, Zheng Peng 0013, Congcong Wen, Yi Fang 0006 |
Comput. Graph. | 1 |
| 2024 | Learning to learn point signature for 3D shape geometry
Hao Huang 0003, Lingjing Wang, Xiang Li 0046, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
Pattern Recognit. Lett. | 1 |
| 2023 | Pyramid Learnable Tokens for 3D LiDAR Place Recognitionabstract3D LiDAR place recognition plays a vital role in various robot applications' including robotic navigation, autonomous driving, and simultaneous localization and mapping. However, most previous studies evaluated their models on accumulated 2D scans instead of real-world 3D LiDAR scans with a larger number of points, which limits the application in real scenarios. To address this limitation, we propose a point transformer network with pyramid learnable tokens (PTNet-PLT) to learn global descriptors for an actual scanned 3D LiDAR place recognition. Specifically, we first present a novel shifted cube attention module that consists of a self-attention module for local feature extraction and a cross-attention module for regional feature aggregation. The self-attention module constrains attention computation on a locally partitioned cube and builds connections across cubes based on the shifted cube scheme. In addition, the cross-attention module introduces several learnable tokens to separately aggregate features of points with similar features but spatially distant into an arbitrarily shaped region, which enables the model to capture long-term dependencies of the points. Next, we build a pyramid architecture network to learn multi-scale features and involve a decreasing number of tokens at each layer to aggregate features over a larger region. Finally, we obtain the global descriptor by concatenating learned region tokens of all layers. Experiments on three datasets, including USyd Campus, Oxford Robot-Car, and KITTI, demonstrate the effectiveness and generalization of the proposed model for large-scale 3D LiDAR place recognition. Congcong Wen, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006 |
ICRA | 2 |
| 2022 | Unsupervised 3D Shape Representation Learning Using Normalizing Flow
Xiang Li 0046, Congcong Wen, Hao Huang 0003 |
ACCV (1) | 3 |
| 2022 | Meta-Det3D: Learn to Learn Few-Shot 3D Object Detection
Shuaihang Yuan, Xiang Li 0046, Hao Huang 0003, Yi Fang 0006 |
ACCV (1) | 3 |
| 2022 | Manifold Adversarial Learning for Cross-domain 3D Shape Representation
Hao Huang 0003, Cheng Chen 0078, Yi Fang 0006 |
ECCV (26) | 1 |
| 2022 | Non-Rigid Multiple Point Set Registration Using Latent Gaussian MixtureabstractPoint set registration is a fundamental task in 3D computer vision. Existing registration approaches mainly focus on either pair-wise or rigid registration. In this paper, we propose a robust group-wise registration method from a probabilistic view and adopt non-rigid transformations to register multiple point sets without bias toward any given set. The proposed method lessens the need of point correspondences by representing each point set as Gaussian Mixture Model and the registration is equivalent to multiple distributions alignment. Closed-form of Jensen-Rényi divergence and L2distance are used as cost functions. We further design a neural network to extract correspondences between raw point sets and Gaussian Mixture Model (GMM) parameters, and recover the optimal diffeomorphic non-rigid transformations from the matched GMM parameters. The proposed method is compared against two well-known probabilistic methods for group-wise point-set registration on several public 2D and 3D datasets. The results demonstrate that our method improves registration accuracy. Hao Huang 0003, Cheng Chen 0078, Yi Fang 0006 |
ICIP | 1 |
| 2022 | Adaptive Wavelet Transformer Network for 3D Shape Representation Learning
Hao Huang 0003, Yi Fang 0006 |
ICLR | 1 |
| 2022 | 3D Point Cloud Completion with Geometric-Aware Adversarial AugmentationabstractWith the popularity of 3D sensors in self-driving and other robotics applications, extensive research has focused on designing novel neural network architectures for accurate 3D point cloud completion. However, unlike point cloud classification and reconstruction, the role of adversarial samples in 3D point cloud completion has seldom been explored. In this work, we demonstrate that adversarial samples can benefit neural networks on 3D point cloud completion tasks. We propose a novel approach to craft adversarial samples that improve the performance of models on both clean and adversarial inputs. In contrast to the Projected Gradient Descent (PGD) attack, our method generates adversarial samples that keep the geometric features in clean samples and contain few outliers. In particular, we use minimum absolute curvature directions to constrain the adversarial perturbations for each input point. The gradient components in the minimum absolute curvature directions are taken as adversarial perturbations. In addition, we adopt attack strength accumulation and auxiliary Batch Normalization layers to speed up the training process and alleviate the distribution mismatch between clean and adversarial samples. Experimental results demonstrate that training with the adversarial samples crafted by our method under the geometric-aware constraint effectively enhances the performance of the Point Completion Network (PCN) on the ShapeNet dataset. Mengxi Wu, Hao Huang 0003, Yi Fang 0006 |
ICPR | 2 |
| 2022 | SHREC 2022: Protein-ligand binding site recognition
Luca Gagliardi, Andrea Raffo, Ulderico Fugacci, Silvia Biasotti, Walter Rocchia, Hao Huang 0003, Boulbaba Ben Amor, Yi Fang 0006, Charles Christoffer, Daisuke Kihara, Apostolos Axenopoulos, Stelios K. Mylonas, Petros Daras |
Comput. Graph. | 6 |
| 2022 | GP-Aligner: Unsupervised Groupwise Nonrigid Point Set Registration Based on Optimizable Group Latent DescriptorabstractIn this paper, we propose a novel unsupervised method named GP-Aligner to address the problem of groupwise non-rigid point set registration. Compared to previous non-learning-based approaches, the proposed method gains competitive advantages by leveraging deep neural networks to effectively and efficiently align a large number of highly deformed 3D shapes with superior performance. Unlike most learning-based methods that use an explicit feature encoding network to extract per-shape features and their correlations, our model leverages a model-free learnable latent descriptor to characterize shape correlations among groups. More specifically, for a given group we first define an optimizable Group Latent Descriptor (GLD) to characterize the relationship among a group of point sets. Each GLD is randomly initialized from a Gaussian distribution and then concatenated with the coordinates of each point of the associated point sets in the group. A neural network-based decoder network is further constructed to predict the coherent flow fields to optimally deform the input groups of shapes to the aligned ones. During the optimization process, GP-Aligner jointly updates all GLDs and weight parameters of the decoder network towards the minimization of an unsupervised groupwise alignment loss. After optimization, for each group, our model coherently drives each point set towards a mean position (shape) without specifying one as the target. GP-Aligner does not require large-scale training data for network training, and it can directly align groups of point sets in a one-stage optimization process. GP-Aligner shows both accuracy and computational efficiency improvement in comparison with the state-of-the-art methods for groupwise point set registration. Moreover, GP-Aligner exhibits high efficiency in aligning a large number of groups of real-world 3D shapes. Lingjing Wang, Hao Huang 0003, Jifei Wang, Xiang Li 0046, Yi Fang 0006 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | 3D-MetaConNet: Meta-learning for 3D Shape Classification and SegmentationabstractSupervised learning on 3D shapes are extensively studied by prior literature, among which PointNet [29] and its variants PointNet++ [31] are representatives. However, these methods tackle 3D shape learning problems by training from scratch using a fixed learning algorithm over large amounts of labeled data, potentially challenged by data and computation bottlenecks. In the paper, we design a novel model, under the framework of meta-learning, to learn 3D shape representation. By training over multiple 3D tasks, each of which is defined as a supervised learning problem, our method can fast adapt to unseen tasks containing limited labeled data. Specifically, our model consists of a 3Dmeta-learner and a task-oriented 3D-learner, where the 3D-meta-learner produces parameter initialization for the 3D-learner after being trained over different tasks. With adaptively initialized parameters, the 3D-learner can be tuned rapidly in a few steps to achieve good performance on novel tasks with a small amount of training data. To further facilitate discriminative shape feature learning, we introduce a novel task-aware feature adaptation module under a contrastive learning scheme, in which all shapes in each task are considered as a whole and task-oriented compact features are learned. Therefore, we dub our model as 3DMetaConNet. Experiments on three public 3D datasets for few-shot shape classification and segmentation demonstrate that our method can learn compact and discriminative 3D shape features efficiently and robustly in a fast adaptation manner. Our method particularly outperforms the methods without a meta-learning framework and is also superior to existing meta-learning approaches. Hao Huang 0003, Xiang Li 0046, Lingjing Wang, Yi Fang 0006 |
3DV | 1 |
| 2020 | Robust Image Matching By Dynamic Feature Selection
Hao Huang 0003, Jianchun Chen, Xiang Li 0046, Lingjing Wang, Yi Fang 0006 |
BMVC | 1 |
| 2019 | Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition
Hao Huang 0003, Luowei Zhou, Jason J. Corso, Chenliang Xu |
BMVC | 1 |