EDBT 2026 Demo / reviewers in the wild / expert
Congcong Wen
dblp:218/4638
· DBLP profile ↗
34ranked-venue papers
5as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 4 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 17 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On Demographic Group Fairness Guarantees in Deep LearningabstractWe present a theoretical framework analyzing the relationship between data distributions and fairness guarantees in deep learning. Our work establishes novel bounds that explicitly account for data distribution heterogeneity across demographic groups, while introducing a formal analysis framework that minimizes expected loss differences across these groups. Moreover, we derive bounds for fairness errors and convergence rates, characterizing how distributional differences between groups affect the fundamental trade-off between fairness and accuracy. Through extensive experiments on diverse datasets across various modalities (image, tabular data, and text), including FairVision (eye disease detection), CheXpert (pleural effusion detection), HAM10000 (skin lesion classification), FairFace (facial attribute recognition), ACS Income (income prediction), CivilComments-WILDS (toxic comment detection), we validate our theoretical findings and demonstrate that differences in feature distributions across demographic groups significantly impact model fairness, with performance disparities particularly pronounced in racial categories. The theoretical bounds we derive corroborate these empirical observations, providing insights into the fundamental limits of achieving fairness in deep learning models when faced with heterogeneous data distributions. This work advances our understanding of fairness in AI and provides a theoretical foundation for developing more equitable algorithms. Motivated by these theoretical insights, particularly the link between feature distribution shifts and fairness gaps, we propose Fairness-Aware Regularization (FAR), a practical training objective that directly minimizes inter-group discrepancies in feature centroids and covariances to improve equitable performance. We validate the effectiveness of FAR across all datasets considered in this study, consistently observing improvements in overall AUC, ES-AUC, and subgroup performance. Yan Luo 0002, Congcong Wen, Min Shi 0001, Hao Huang 0003, Yi Fang 0006, Mengyu Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | FairFedMed: Benchmarking Group Fairness in Federated Medical Imaging With FairLoRAabstractFairness remains a critical concern in healthcare, where unequal access to services and treatment outcomes can adversely affect patient health. While Federated Learning (FL) presents a collaborative and privacy-preserving approach to model training, ensuring fairness is challenging due to heterogeneous data across institutions, and current research primarily addresses non-medical applications. To fill this gap, we establish the first experimental benchmark for fairness in medical FL, evaluating six representative FL methods across diverse demographic attributes and imaging modalities. We introduce FairFedMed, the first medical FL dataset specifically designed to study group fairness (i.e., consistent performance across demographic groups). It comprises two parts: FairFedMed-Oph, featuring 2D fundus and 3D OCT ophthalmology samples with six demographic attributes; and FairFedMed-Chest, which simulates real cross-institutional FL using subsets of CheXpert and MIMIC-CXR. Together, they support both simulated and real-world FL across diverse medical modalities and demographic groups. Existing FL models often underperform on medical images and overlook fairness across demographic groups. To address this, we propose FairLoRA, a fairness-aware FL framework based on SVD-based low-rank approximation. It customizes singular value matrices per demographic group while sharing singular vectors, ensuring both fairness and efficiency. Experimental results on the FairFedMed dataset demonstrate that FairLoRA not only achieves state-of-the-art performance in medical image classification but also significantly improves fairness across diverse populations. Our code and dataset can be accessible via GitHub link: https://github.com/Harvard-AI-and-Robotics-Lab/FairFedMed. Minghan Li 0001, Congcong Wen, Yu Tian 0001, Min Shi 0001, Yan Luo 0002, Hao Huang 0003, Yi Fang 0006, Mengyu Wang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Towards Robust Visual Question Answering via Prompt-Driven Geometric HarmonizationabstractVisual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations posed by language bias and imbalanced distributions. To address these challenges, this paper proposes a novel Prompt-Driven Geometric Harmonization (PDGH) paradigm, which integrates both geometric structure and information entropy principles to enhance the ability of VQA models to generalize effectively across diverse scenarios. Specifically, our PDGH approach is meticulously designed to generate image-generated prompts that are guided by specific question cues, facilitating a more accurate and context-aware understanding of the visual content. Moreover, we project the prompt-visual-question and visual-question joint representations into a unified hypersphere space, applying feature weight self-orthogonality and prompt-information entropy correction constraints to optimize the margin, further alleviating minority class collapse and correcting language bias. To maintain the geometric integrity of the representation space, we introduce multi-space geometric contrast constraints to minimize the impact of spurious priors introduced during training. Finally, a semantic matrix is constructed for the coordinated joint representation to ensure that the learned instances are semantically consistent and improve reasoning ability. Extensive experiments on various general and medical VQA datasets demonstrate the consistent superiority of our PDGH approach over existing state-of-the-art baselines. Yishu Liu 0001, Congcong Wen, Guangming Lu 0002, Bingzhi Chen |
AAAI | 3 |
| 2025 | A2GP-SF: Enhancing Few-shot Class Incremental Learning via Attribute Generative Prompting and Adaptive Sharpness FlatteningabstractFew-shot Class Incremental Learning (FSCIL) aims to incrementally learn new classes with limited examples while retaining knowledge of previously learned classes. Recent advancements in prompt tuning for large pre-trained models have shown promise in FSCIL. However, current FSCIL methods still suffer from challenges like insufficient plasticity and limited generalization. To tackle these challenges, we propose a novel prompt tuning-based framework named A2GP-SF, which integrates attribute generative prompting (AGP) and adaptive sharpness flattening (ASF). The proposed AGP paradigm dynamically generates attribute-aware prompts for each instance, facilitating better semantics learning and enhancing plasticity. Additionally, the ASF mechanism aims to mitigate overfitting by applying adaptive perturbations to flatten sharpness, with these perturbations adjusted based on gradient norm changes, thereby enhancing the model’s robustness and generalization. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our proposed A2GP-SF framework. Desen Wang, Sisi Fu, Congcong Wen, Bingzhi Chen |
ICASSP | 4 |
| 2025 | Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon TasksabstractPolicy learning focuses on devising strategies for agents in embodied artificial intelligence systems to perform optimal actions based on their perceived states. One of the key challenges in policy learning involves handling complex, long-horizon tasks that require managing extensive sequences of actions and observations with multiple modes. Wavelet analysis offers significant advantages in signal processing, notably in decomposing signals at multiple scales to capture both global trends and fine-grained details. In this work, we introduce a novel wavelet policy learning framework that utilizes wavelet transformations to enhance policy learning. Our approach leverages learnable multi-scale wavelet decomposition to facilitate detailed observation analysis and robust action planning over extended sequences. We detail the design and implementation of our wavelet policy, which incorporates lifting schemes for effective multi-resolution analysis and action generation. This framework is evaluated across multiple complex scenarios, including robotic manipulation, self-driving, and multi-robot collaboration, demonstrating the effectiveness of our method in improving the precision and reliability of the learned policy. Hao Huang 0003, Shuaihang Yuan, Geeta Chandra Raju Bethala, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
ICCV | 4 |
| 2025 | DiffGAD: A Diffusion-based Unsupervised Graph Anomaly DetectorabstractGraph Anomaly Detection (GAD) is crucial for identifying abnormal entities within networks, garnering significant attention across various fields. Traditional unsupervised methods, which decode encoded latent representations of unlabeled data with a reconstruction focus, often fail to capture critical discriminative content, leading to suboptimal anomaly detection.
To address these challenges, we present a Diffusion-based Graph Anomaly Detector (DiffGAD). At the heart of DiffGAD is a novel latent space learning paradigm, meticulously designed to enhance the model's proficiency by guiding it with discriminative content. This innovative approach leverages diffusion sampling to infuse the latent space with discriminative content and introduces a content-preservation mechanism that retains valuable information across different scales, significantly improving the model’s adeptness at identifying anomalies with limited time and space complexity.
Our comprehensive evaluation of DiffGAD, conducted on six real-world and large-scale datasets with various metrics, demonstrated its exceptional performance. Our code is available at https://github.com/fortunato-all/DiffGAD Jinghan Li, Jinda Lu, Junfeng Fang, Congcong Wen, Xiang Wang 0010 |
ICLR | 5 |
| 2025 | ResLPR: A LiDAR Data Restoration Network and Benchmark for Robust Place Recognition Against Weather CorruptionsabstractLiDAR-based place recognition (LPR) is a key component for autonomous driving, and its resilience to environmental corruption is critical for safety in high-stakes applications. While state-of-the-art (SOTA) LPR methods perform well in clean weather, they still struggle with weather-induced corruption commonly encountered in driving scenarios. To tackle this, we propose ResLPRNet, a novel LiDAR data restoration network that largely enhances LPR performance under adverse weather by restoring corrupted LiDAR scans using a wavelet transform-based network. ResLPRNet is efficient, lightweight and can be integrated plug-and-play with pretrained LPR models without substantial additional computational cost. Given the lack of LPR datasets under adverse weather, we introduce ResLPR, a novel benchmark that examines SOTA LPR methods under a wide range of LiDAR distortions induced by severe snow, fog, and rain conditions. Experiments on our proposed WeatherKITTI and WeatherNCLT datasets demonstrate the resilience and notable gains achieved by using our restoration method with multiple LPR approaches in challenging weather scenarios. Our code and benchmark are publicly available here: https://github.com/nubot-nudt/ResLPR. Wenqing Kuang, Xiongwei Zhao, Yehui Shen, Congcong Wen, Huimin Lu 0002, Zongtan Zhou, Xieyuanli Chen |
IROS | 4 |
| 2025 | Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language ModelsabstractRobotic navigation plays a pivotal role in a wide range of real-world applications. While traditional navigation systems focus on efficiency and obstacle avoidance, their inability to model complex human behaviors in shared spaces has underscored the growing need for socially aware navigation. In this work, we explore a novel paradigm of socially aware robot navigation empowered by large language models (LLMs), and propose HSAC-LLM, a hybrid framework that seamlessly integrates deep reinforcement learning with the reasoning and communication capabilities of LLMs. Unlike prior approaches that passively predict pedestrian trajectories or issue pre-scripted alerts, HSAC-LLM enables bidirectional natural language interaction, allowing robots to proactively engage in dialogue with pedestrians to resolve potential conflicts and negotiate path decisions. Extensive evaluations across 2D simulations, Gazebo environments, and real-world deployments demonstrate that HSAC-LLM consistently outperforms state-of-the-art DRL baselines under our proposed socially aware navigation metric, which covers safety, efficiency, and human comfort. By bridging linguistic reasoning and interactive motion planning, our results highlight the potential of LLM-augmented agents for robust, adaptive, and human-aligned navigation in real-world settings. Project page: https://hsacllm.github.io/. Congcong Wen, Geeta Chandra Raju Bethala, Shuaihang Yuan, Hao Huang 0003, Mengyu Wang 0001, Yu-Shen Liu, Anthony Tzes, Yi Fang 0006 |
IROS | 1 |
| 2025 | SplitFlow: Flow Decomposition for Inversion-Free Text-to-Image EditingabstractRectified flow models have become a $\textit{de facto}$ standard in image generation due to their stable sampling trajectories and high-fidelity outputs. Despite their strong generative capabilities, they face critical limitations in image editing tasks: inaccurate inversion processes for mapping real images back into the latent space, and gradient entanglement issues during editing often result in outputs that do not faithfully reflect the target prompt. Recent efforts have attempted to directly map source and target distributions via ODE-based approaches without inversion; however, these methods still yield suboptimal editing quality. In this work, we propose a flow decomposition-and-aggregation framework built upon an inversion-free formulation to address these limitations. Specifically, we semantically decompose the target prompt into multiple sub-prompts, compute an independent flow for each, and aggregate them to form a unified editing trajectory. While we empirically observe that decomposing the original flow enhances diversity in the target space, generating semantically aligned outputs still requires consistent guidance toward the full target prompt. To this end, we design a projection and soft-aggregation mechanism for flow, inspired by gradient conflict resolution in multi-task learning. This approach adaptively weights the sub-target velocity fields, suppressing semantic redundancy while emphasizing distinct directions, thereby preserving both diversity and consistency in the final edited output. Experimental results demonstrate that our method outperforms existing zero-shot editing approaches in terms of semantic fidelity and attribute disentanglement. The code is available at https://github.com/Harvard-AI-and-Robotics-Lab/SplitFlow. Sung-Hoon Yoon 0001, Gaspard Beaudouin, Congcong Wen, Muhammad Rafay Azhar |
NeurIPS | 4 |
| 2025 | Earthquake Spatio-Temporal Patterns Around Precursory Thermal Infrared (TIR) AnomaliesabstractIn this work, we analyze outgoing longwave radiation (OLR) data from the National Oceanic and Atmospheric Administration (NOAA) to extract thermal infrared (TIR) anomalies over mainland China from 2007 to 2019. These anomalies are statistically significantly correlated with earthquakes of magnitude$\geq 4.5$, 5.0 and 5.5. We then calculate the distance ($d/^{\circ } $) and time lag$(t/\text {days})$between each TIR anomaly and the corresponding earthquake. By examining the distribution of earthquakes in$t \times d$space, our results show that the normalized frequency of earthquakes as a function of distance and time follows a power-law distribution:$f(d,t,M)=\phi (M) *t^{-d_{t}(M)}*d^{-d_{d}(M)}$, where$f(d,t,M)$represents the frequency of earthquakes with magnitude${\geq }M$occurring within the time-space unit$(t, d)$around TIR anomalies. The coefficient$\phi (M)$is negatively correlated with magnitude$[\phi (4.5)=0.31,\phi (5.0)=0.30,\phi (5.5)=0.28]$, while$d_{d}(M)$and$d_{t}(M)$are two positive coefficients$[d_{t}(4.5)=1.59,d_{t}(5.0)=1.58,d_{t}(5.5)=1.50;d_{d}(4.5)=2.76,d_{d}(5.0)=2.66,d_{d}(5.5)=2.59]$. A larger$\phi $suggests that earthquakes are more spatially and temporally concentrated around TIR anomalies, while$d_{d}$and$d_{t}$represent the decay rates of frequency with distance and time, respectively. These findings imply that earthquakes are more likely to occur near TIR anomalies, and the model$f(d,t,M)$could potentially be used to construct a probabilistic earthquake forecasting system based on TIR anomalies. Jizhe Zhou 0002, Congcong Wen |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Generalization-Enhanced Few-Shot Object Detection in Remote SensingabstractObject detection is a fundamental task in computer vision that involves accurately locating and classifying objects within images or video frames. In remote sensing, this task is particularly challenging due to the high resolution, multi-scale features, and diverse ground object characteristics inherent in satellite and UAV imagery. These challenges necessitate more advanced approaches for effective object detection in such environments. While deep learning methods have achieved remarkable success in remote sensing object detection, they typically rely on large amounts of labeled data. Acquiring sufficient labeled data, particularly for novel or rare objects, is both challenging and time-consuming in remote sensing scenarios, limiting the generalization capabilities of existing models. To address these challenges, few-shot learning (FSL) has emerged as a promising approach, aiming to enable models to learn new classes from limited labeled examples. Building on this concept, few-shot object detection (FSOD) specifically targets object detection challenges in data-limited conditions. However, the generalization capability of FSOD models, particularly in remote sensing, is often constrained by the complex and diverse characteristics of the objects present in such environments. In this paper, we propose the Generalization-Enhanced Few-Shot Object Detection (GE-FSOD) model to improve the generalization capability in remote sensing FSOD tasks. Our model introduces three key innovations: the Cross-Level Fusion Pyramid Attention Network (CFPAN) for enhanced multi-scale feature representation, the Multi-Stage Refinement Region Proposal Network (MRRPN) for more accurate region proposals, and the Generalized Classification Loss (GCL) for improved classification performance in few-shot scenarios. GE-FSOD demonstrates superior robustness and accuracy in remote sensing FSOD tasks through these enhancements. Extensive experiments on the DIOR and NWPU VHR-10 datasets show that our model achieves state-of-the-art performance, significantly advancing the field of few-shot object detection in remote sensing. The source code is available at (https://github.com/leenamx/GE-FSOD). Nan Li 0052, Pengjuan Yao, Kexin Dong, Danfeng Hong, Congcong Wen |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | RS-MoE: A Vision-Language Model With Mixture of Experts for Remote Sensing Image Captioning and Visual Question AnsweringabstractRemote sensing image captioning (RSIC) presents unique challenges and plays a critical role in applications such as environmental monitoring, urban planning, and disaster management. Traditional RSIC methods often struggle to produce rich and diverse descriptions. Recently, with significant advancements in vision-language models (VLMs), efforts have emerged to integrate these models into the remote sensing domain and to introduce richly descriptive datasets specifically designed to enhance VLM training. However, most current RSIC models generally apply only fine-tuning to these datasets without developing models tailored to the unique characteristics of remote sensing imagery. This article proposes RS-MoE, the first mixture of expert (MoE)-based VLM specifically customized for remote sensing domain. Unlike traditional MoE models, the core of RS-MoE is the MoE block, which incorporates a novel instruction router and multiple lightweight large language models (LLMs) as expert models. The instruction router is designed to generate specific prompts tailored for each corresponding LLM, guiding them to focus on distinct aspects of the RSIC task. This design not only allows each expert LLM to concentrate on a specific subset of the task, thereby enhancing the specificity and accuracy of the generated captions, but also improves the scalability of the model by facilitating parallel processing of subtasks. In addition, we present a two-stage training strategy for tuning our RS-MoE model to prevent performance degradation due to sparsity. We fine-tuned our model on the RSICap dataset using our proposed training strategy. Experimental results on the RSICap dataset, along with evaluations on other traditional datasets where no additional fine-tuning was applied, demonstrate that our model achieves state-of-the-art performance in generating precise and contextually relevant captions. Notably, our RS-MoE-1B variant achieves performance comparable to 13B VLMs, demonstrating the efficiency of our model design. Moreover, our model demonstrates promising generalization capabilities by consistently achieving state-of-the-art performance on the remote sensing visual question answering (RSVQA) task. Danfeng Hong, Shuhang Ge, Chuyao Luo, Congcong Wen |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | TripleMixer: A Triple-Domain Mixing Model for Point Cloud Denoising Under Adverse WeatherabstractAdverse weather conditions such as snow, fog, and rain pose significant challenges to LiDAR-based perception models by introducing noise and corrupting point cloud measurements. To address this issue, we propose TripleMixer, a robust and efficient point cloud denoising network that integrates spatial, frequency, and channel-wise processing through three specialized mixer modules. TripleMixer effectively suppresses high-frequency noise while preserving essential geometric structures and can be seamlessly deployed as a plug-and-play module within existing LiDAR perception pipelines. To support the development and evaluation of denoising methods, we construct two large-scale simulated datasets, Weather-KITTI and Weather-NuScenes, covering diverse weather scenarios with dense point-wise semantic and noise annotations. Based on these datasets, we establish four benchmarks: Denoising, Semantic Segmentation (SS), Place Recognition (PR), and Object Detection (OD). These benchmarks enable systematic evaluation of denoising generalization, transferability, and downstream impact under both simulated and real-world adverse weather conditions. Extensive experiments demonstrate that TripleMixer achieves state-of-the-art denoising performance and yields substantial improvements across all downstream tasks without requiring retraining. Our results highlight the potential of denoising as a task-agnostic preprocessing strategy to enhance LiDAR robustness in real-world autonomous driving applications. Xiongwei Zhao, Congcong Wen, Xu Zhu 0001, Yang Wang 0029, Haojie Bai 0001, Wenhao Dou |
IEEE Trans. Image Process. | 2 |
| 2025 | 3D Shape Contrastive Representation Learning With Adversarial ExamplesabstractCurrent supervised methods for 3D shape representation learning have achieved satisfying performance, yet require extensive human-labeled datasets. Unsupervised learning-based methods provide a viable solution by learning shape representations without using ground truth labels. In this study, we develop a contrastive learning framework for unsupervised representation learning of 3D shapes. Specifically, in order to encourage models to pay more attention to useful information during representation learning, we first introduce a new paradigm for critical points search based on the adversarial mechanism. We extract critical points with a larger impact on the global feature by attacking a pre-trained auto-encoder model, and apply data augmentations on these points to generate adversarial examples. Taking a pair of adversarial examples as inputs, we obtain their intermediate embeddings and global representations of corresponding inputs, which are then transformed into latent spaces by two predictor heads. Finally, we train the proposed model by maximizing the agreements on these latent spaces via Normalized Temperature-scaled Cross Entropy (NT-Xent) loss and a newly designed Cross-layer Normalized Temperature-scaled Cross Entropy (Cross-NT-Xent) loss, where the latter is proposed in this paper to enforce cross-layer feature similarities. The effectiveness, robustness, and transferability of learned representations are validated on three downstream tasks, including object classification, few-shot classification, and shape retrieval. Experiments on three benchmark datasets show that our learned representations achieve better or competitive performance than current state-of-the-art methods in these downstream tasks. Moreover, our model can easily be extended to 3D part segmentation and scene segmentation tasks. Congcong Wen, Xiang Li 0046, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006 |
IEEE Trans. Multim. | 1 |
| 2025 | Retrieval-Specific View Learning for Sketch-to-Shape RetrievalabstractSketch-based 3D shape retrieval (SBSR) can be approached by learning domain-invariant descriptors or ranking metrics from sketches and 2D view images of 3D shapes rendered through numerous viewpoints. However, determining the most appropriate viewpoints that convey discriminative geometric features to benefit the task of SBSR became an essential yet not fully explored area. Existing works extract 3D features from multi-view images observed through pre-defined viewpoints to match 2D sketches. Those methods, however, fail to dynamically select viewpoints by considering the SBSR task. In this work, we introduce a fully differentiable viewpoint learning paradigm driven by the downstream SBSR task, which supports the task-aware and sketch-dependent dynamic viewpoint determination process. We naturally integrate this task-specific and sketch-dependent viewpoint learning process into a meta-learning framework to develop a novel Dynamic Viewer (DV) module for SBSR. DV module comprises a Meta View Learner (MVL) block and a View Generator (VG) block. Specifically, as the first part of the DV module, the MVL block learns to initiate the necessary network parameters of the VG block. Then, the VG block that serves as the second part learns the best viewpoints to render 2D images. To learn the optimal viewpoints for SBSR, we further introduce a view mining loss that aims to maximize the similarity of feature-level information among rendered 2D views and the query sketch. Further, we adopt a variational autoencoder (VAE) to retrieve 3D shapes by setting the newly rendered images and query sketch as inputs. As evidenced by the comprehensive experimental results conducted on popular SBSR datasets, the proposed framework has been demonstrated to outperform recent methods in both category-level sketch-based and fine-grained SBSR. Shuaihang Yuan, Congcong Wen, Yu-Shen Liu, Yi Fang 0006 |
IEEE Trans. Multim. | 2 |
| 2024 | FairDomain: Achieving Fairness in Cross-Domain Medical Image Segmentation and Classification
Yu Tian 0001, Congcong Wen, Min Shi 0001, Muhammad Muneeb Afzal, Hao Huang 0003, Muhammad Osama Khan, Yan Luo 0002, Yi Fang 0006, Mengyu Wang 0001 |
ECCV (76) | 2 |
| 2024 | Optimizing Personalized Robot Actions with Ranking of Trajectories
Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
ICPR (29) | 4 |
| 2024 | Reliable Semantic Understanding for Real World Zero-Shot Object Goal Navigation
Halil Utku Unlu, Shuaihang Yuan, Congcong Wen, Hao Huang 0003, Anthony Tzes, Yi Fang 0006 |
ICPR (30) | 3 |
| 2024 | Zero-Shot Object Navigation with Vision-Language Models Reasoning
Congcong Wen, Yisiyuan Huang, Hao Huang 0003, Yanjia Huang, Shuaihang Yuan, Yu-Shen Liu, Yi Fang 0006 |
ICPR (18) | 1 |
| 2024 | Goal-Driven Transformer for Robot Behavior Learning from Play Data
Congcong Wen, Jiazhao Liang, Shuaihang Yuan, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006 |
ICPR (30) | 1 |
| 2024 | Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
Shuaihang Yuan, Halil Utku Unlu, Hao Huang 0003, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
ICPR (30) | 4 |
| 2024 | Noisy Few-shot 3D Point Cloud Scene Segmentationabstract3D scene semantic segmentation plays a crucial role in robotics by enabling robots to understand and interpret their environment in a detailed and context-aware manner, facilitating tasks such as navigation, object manipulation, and interaction within complex spaces. A preponderance of methodology predominantly adopts a fully supervised framework for 3D point cloud scene semantic segmentation. Such paradigms exhibit an intrinsic dependency on extensive labeled datasets, presenting challenges in acquisition and exhibiting incapacity to segment novel classes, especially when the training data are contaminated by noisy samples. To address these limitations, this study introduces a novel few-shot segmentation approach to robustly segment 3D point cloud scenes with noisy labels using a meta-learning scheme. Specifically, we first build a multi-prototype graph and then suppress samples with noisy labels based on the graph structure. A subgraph bagging scheme is then proposed to conduct semi-supervised transductive learning to propagate labels. To optimize the graph structure to learn discriminative prototype features, we design a triplet contrastive loss to increase the compactness of these subgraphs. We evaluated our method on two widely used 3D point cloud scene segmentation benchmarks within few-shot (i.e., 2/3-way 5-shot) segmentation settings with noisy samples. Experimental results demonstrate the improvement of our method over the compared baselines, illustrating the robustness of our method in few-shot 3D scene segmentation against noisy samples. The code is available at: https://github.com/hhuang-code/Noisy_Fewshot_Segmentation. Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
ICRA | 3 |
| 2024 | Weakly Scene Segmentation Using Efficient TransformerabstractCurrent methods for large-scale point cloud scene semantic segmentation rely on manually annotated dense point-wise labels, which are costly, labor-intensive, and prone to errors. Consequently, gathering point cloud scenes with billions of labeled points is impractical in real-world scenarios. In this paper, we introduce a novel weak supervision approach to semantically segment large-scale indoor scenes, requiring only 1‰ of the points to be labeled. Specifically, we develop an efficient point neighbor Transformer to capture the geometry of local point cloud patches. To address the quadratic complexity of self-attention computation in Transformers, particularly for large-scale point clouds, we propose approximating the self-attention matrix using low-rank and sparse decomposition. Building on the point neighbor Transformer as foundational blocks, we design a Low-rank Sparse Transformer Network (LST-Net) for weakly supervised large-scale point cloud scene semantic segmentation. Experimental results on two commonly used indoor point cloud scene segmentation benchmarks demonstrate that our model achieves performance comparable to those of both weakly supervised and fully supervised methods. Our code can be found in https://github.com/hhuang-code/LST-Net. Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
IROS | 3 |
| 2024 | ChatMap: A Wearable Platform Based on the Multi-modal Foundation Model to Augment Spatial Cognition for People with Blindness and Low VisionabstractSpatial cognition refers to the ability to gain knowledge about their surroundings and utilize this information to identify their location, acquire resources, and navigate their way back to familiar places. People with blindness and low vision (pBLV) face significant challenges with spatial cognition due to the reliance on visual input. Without the full range of visual cues, pBLV individuals often find it difficult to grasp a comprehensive understanding of their environment, leading to obstacles in scene recognition and precise object localization, especially in unfamiliar environments. This limitation extends to their ability to independently detect and avoid potential tripping hazards, making navigation and interaction with their environment more challenging. In this paper, we present a pioneering wearable platform tailored to enhance the spatial cognition of pBLV through the integration of multi-modal foundation model. The proposed platform integrates a wearable camera with audio module and leverages the advanced capabilities of vision language foundation model (i.e., GPT-4 and GPT-4V), for the nuanced processing of visual and textual data. Specifically, we employ vision language models to bridge the gap between visual information and the proprioception of visually impaired users, offering more intelligible guidance by aligning visual data with the natural perception of space and movement. Then we apply prompt engineering to guide the large language model to act as an assistant tailored specifically for pBLV users to produce accurate answers. Another innovation in our model is the incorporation of a chain of thought reasoning process, which enhances the accuracy and interpretability of the model, facilitating the generation of more precise responses to complex user inquiries across diverse environmental contexts. To assess the practical impact of our proposed wearable platform, we carried out a series of real-world experiments across three tasks that are commonly challenging for people with blindness and low vision: risk assessment, object localization, and scene recognition. Additionally, through an ablation study conducted on the VizWiz dataset, we rigorously assess the contribution of each individual module, substantiating the integral role in the model’s overall performance. Alexey Magay, Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
IROS | 5 |
| 2024 | GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance GuidanceabstractZero-Shot Object Goal Navigation (ZS-OGN) enables robots to navigate toward objects of unseen categories without prior training. Traditional approaches often leverage categorical semantic information for navigation guidance, which struggles when only partial objects are observed or detailed and functional representations of the environment are lacking. To resolve the above two issues, we propose \textit{Geometric-part and Affordance Maps} (GAMap), a novel method that integrates object parts and affordance attributes for navigation guidance. Our method includes a multi-scale scoring approach to capture geometric-part and affordance attributes of objects at different scales. Comprehensive experiments conducted on the HM3D and Gibson benchmark datasets demonstrate improvements in Success Rates and Success weighted by Path Length, underscoring the efficacy of our geometric-part and affordance-guided navigation approach in enhancing robot autonomy and versatility, without any additional task-specific training or fine-tuning with the semantics of unseen objects and/or the locomotions of the robot. Shuaihang Yuan, Hao Huang 0003, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
NeurIPS | 4 |
| 2024 | A single 3D shape wavelet-based generative model
Hao Huang 0003, Shuaihang Yuan, Zheng Peng 0013, Congcong Wen, Yi Fang 0006 |
Comput. Graph. | 5 |
| 2024 | Learning to learn point signature for 3D shape geometry
Hao Huang 0003, Lingjing Wang, Xiang Li 0046, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
Pattern Recognit. Lett. | 5 |
| 2023 | Pyramid Learnable Tokens for 3D LiDAR Place Recognitionabstract3D LiDAR place recognition plays a vital role in various robot applications' including robotic navigation, autonomous driving, and simultaneous localization and mapping. However, most previous studies evaluated their models on accumulated 2D scans instead of real-world 3D LiDAR scans with a larger number of points, which limits the application in real scenarios. To address this limitation, we propose a point transformer network with pyramid learnable tokens (PTNet-PLT) to learn global descriptors for an actual scanned 3D LiDAR place recognition. Specifically, we first present a novel shifted cube attention module that consists of a self-attention module for local feature extraction and a cross-attention module for regional feature aggregation. The self-attention module constrains attention computation on a locally partitioned cube and builds connections across cubes based on the shifted cube scheme. In addition, the cross-attention module introduces several learnable tokens to separately aggregate features of points with similar features but spatially distant into an arbitrarily shaped region, which enables the model to capture long-term dependencies of the points. Next, we build a pyramid architecture network to learn multi-scale features and involve a decreasing number of tokens at each layer to aggregate features over a larger region. Finally, we obtain the global descriptor by concatenating learned region tokens of all layers. Experiments on three datasets, including USyd Campus, Oxford Robot-Car, and KITTI, demonstrate the effectiveness and generalization of the proposed model for large-scale 3D LiDAR place recognition. Congcong Wen, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006 |
ICRA | 1 |
| 2022 | Unsupervised 3D Shape Representation Learning Using Normalizing Flow
Xiang Li 0046, Congcong Wen, Hao Huang 0003 |
ACCV (1) | 2 |
| 2022 | Road Extraction From Remote Sensing Images in Wildland-Urban Interface AreasabstractIn this letter, we address the problem of road extraction in Wildland–urban interface (WUI) areas. In recent years, with the great success of convolutional neural networks (CNNs) in various vision-related tasks, researchers have developed many CNN-based methods for road extraction on remote sensing images. Nevertheless, these methods mostly treat road extraction as a binary classification problem on semantic labeling. In WUI areas, the road is narrower and tends to be occluded by trees, which may result in the serious discontinuous problem of inferred road maps. To address this issue, we propose transforming the input representation of the binary classification map into a continuous signed distance map. In this way, our model is forced to predict the continuous distance representations and, thus, improve the spatial continuities of inferred roads. In addition, a real-value regression task is designed to train along with the original binary classification task to generate spatially continuous and semantically accurate road maps. Then, we conduct experiments on the public Massachusetts road data set and a homemade data set collected from Yajishan Mountain, Beijing, China. Finally, our proposed method achieves intersection-over-unions (IoUs) of 64.11% and 65.92% for the Massachusetts and WUI-Yajishan data sets, respectively, without any postprocessing. In addition, the ablation analysis shows that introducing the regression task on the proposed signed distance representation can effectively alleviate the problem of discontinuous road prediction. Furthermore, comparing with the state-of-the-art methods demonstrates the superiority of our method for road extraction in WUI areas. Xiang Li 0046, Yuan Hu 0004, Congcong Wen |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Geometry-Aware Segmentation of Remote Sensing Images via Joint Height EstimationabstractRecent studies have shown the benefits of using additional elevation data [e.g., digital surface model (DSM) or normalized DSM (nDSM)] for enhancing the performance of the semantic labeling of aerial images. However, previous methods mostly adopt 3-D elevation information as additional inputs, while, in many real-world applications, one does not have the corresponding DSM images at hand, and the spatial resolution of acquired DSM images usually does not match the aerial images. To alleviate this data constraint and also take advantage of 3-D elevation information, in this letter, a geometry-aware segmentation model is introduced to achieve accurate semantic labeling of aerial images via joint height estimation. Instead of using a single-stream encoder–decoder network for semantic labeling, we design a separate decoder branch to predict the height map and use the DSM images as side supervision to train this newly designed decoder branch. With the newly designed decoder branch, our model can distill the 3-D geometric features from 2-D appearance features under the supervision of ground-truth DSM images. Moreover, we develop a new geometry-aware convolution module that fuses the 3-D geometric features from the height decoder branch and the 2-D contextual features from the semantic segmentation branch. The fused feature embeddings can produce geometry-aware segmentation maps with enhanced performance. Our model is trained with DSM images as side supervision, while, in the inference stage, it does not require DSM data and directly predicts the semantic labels. Experiments on International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam data sets demonstrate the effectiveness of the proposed method for the semantic segmentation of aerial images. Xiang Li 0046, Congcong Wen, Lingjing Wang, Yi Fang 0006 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Fooling LiDAR Perception via Adversarial Trajectory PerturbationabstractLiDAR point clouds collected from a moving vehicle are functions of its trajectories, because the sensor motion needs to be compensated to avoid distortions. When autonomous vehicles are sending LiDAR point clouds to deep networks for perception and planning, could the motion compensation consequently become a wide-open backdoor in those networks, due to both the adversarial vulnerability of deep learning and GPS-based vehicle trajectory estimation that is susceptible to wireless spoofing? We demonstrate such possibilities for the first time: instead of directly attacking point cloud coordinates which requires tampering with the raw LiDAR readings, only adversarial spoofing of a self-driving car’s trajectory with small perturbations is enough to make safety-critical objects undetectable or detected with incorrect positions. Moreover, polynomial trajectory perturbation is developed to achieve a temporally-smooth and highly-imperceptible attack. Extensive experiments on 3D object detection have shown that such attacks not only lower the performance of the state-of-the-art detectors effectively, but also transfer to other detectors, raising a red flag for the community. The code is available on https://ai4ce.github.io/FLAT/. Yiming Li 0003, Congcong Wen, Felix Juefei-Xu, Chen Feng 0002 |
ICCV | 2 |
| 2021 | Topology Constrained Shape CorrespondenceabstractTo better address the deformation and structural variation challenges inherently present in 3D shapes, researchers have shifted their focus from designing handcrafted point descriptors to learning point descriptors and their correspondences in a data-driven manner. Recent studies have developed deep neural networks for robust point descriptor and shape correspondence learning in consideration of local structural information. In this article, we developed a novel shape correspondence learning network, called TC-NET, which further enhances performance by encouraging the topological consistency between the embedding feature space and the input shape space. Specifically, in this article, we first calculate the topology-associated edge weights to represent the topological structure of each point. Then, in order to preserve this topological structure in high-dimensional feature space, a structural regularization term is defined to minimize the topology-consistent feature reconstruction loss (Topo-Loss) during the correspondence learning process. Our proposed method achieved state-of-the-art performance on three shape correspondence benchmark datasets. In addition, the proposed topology preservation concept can be easily generalized to other learning-based shape analysis tasks to regularize the topological structure of high-dimensional feature spaces. Xiang Li 0046, Congcong Wen, Lingjing Wang, Yi Fang 0006 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | A spatial co-location mining algorithm that includes adaptive proximity improvements and distant instance referencesabstractSpatial co-location pattern mining is employed to identify a group of spatial types whose instances are frequently located in spatial proximity. Current co-location mining methods have two limitations: (1) it is difficult to set an appropriate proximity threshold to identify close instances in an unknown region, and (2) such methods neglect the effects of the distance values between instances and long-distance instance effects on pattern significance. This paper proposes a novel maximal co-location algorithm to address these problems. To remove the first constraint, the algorithm uses Voronoi diagrams to extract the most related instance pairs of different types and their normalized distances, from which two distance-separating parameters are adaptively extracted using a statistical method. To remove the second constraint, the algorithm employs a reward-based verification based on distance-separating parameters to identify the prevalent patterns. Our experiments with both synthetic data and real data from Beijing, China, demonstrate that the algorithm can identify many interesting patterns that are neglected by traditional co-location methods. Xiaojing Yao, Liujia Chen, Congcong Wen, Tianhe Chi, Wenhao Yu 0001 |
Int. J. Geogr. Inf. Sci. | 3 |