EDBT 2026 Demo / reviewers in the wild / expert
Shuaihang Yuan
dblp:257/3707
· DBLP profile ↗
22ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0002-7092-7966ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 11 since 2021Systems, architecture and hardware · 7 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon TasksabstractPolicy learning focuses on devising strategies for agents in embodied artificial intelligence systems to perform optimal actions based on their perceived states. One of the key challenges in policy learning involves handling complex, long-horizon tasks that require managing extensive sequences of actions and observations with multiple modes. Wavelet analysis offers significant advantages in signal processing, notably in decomposing signals at multiple scales to capture both global trends and fine-grained details. In this work, we introduce a novel wavelet policy learning framework that utilizes wavelet transformations to enhance policy learning. Our approach leverages learnable multi-scale wavelet decomposition to facilitate detailed observation analysis and robust action planning over extended sequences. We detail the design and implementation of our wavelet policy, which incorporates lifting schemes for effective multi-resolution analysis and action generation. This framework is evaluated across multiple complex scenarios, including robotic manipulation, self-driving, and multi-robot collaboration, demonstrating the effectiveness of our method in improving the precision and reliability of the learned policy. Hao Huang 0003, Shuaihang Yuan, Geeta Chandra Raju Bethala, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
ICCV | 2 |
| 2025 | MultiTalk: Introspective and Extrospective Dialogue for Human-Environment-LLM AlignmentabstractLLMs have shown promising results in task planning due to their strong natural language understanding and reasoning capabilities. However, issues such as hallucinations, ambiguities in human instructions, environmental constraints, and limitations in the executing agent's capabilities often lead to flawed or incomplete plans. This paper proposes MultiTalk, an LLM-based task planning methodology that addresses these issues through a framework of introspective and extrospective dialogue loops. This approach helps ground generated plans in the context of the environment and the agent's capabilities, while also resolving uncertainties and ambiguities in the given task. These loops are enabled by specialized systems designed to extract and predict task-specific states, and flag mismatches or misalignments among the human user, the LLM agent, and the environment. Effective feedback pathways between these systems and the LLM planner foster meaningful dialogue. The efficacy of this methodology is demonstrated through its application to robotic manipulation tasks. Experiments and ablations highlight the robustness and reliability of our method, and comparisons with baselines further illustrate the superiority of MultiTalk in task planning for embodied agents. Project Website: https://llm-multitalk.github.io/ Venkata Naren Devarakonda, Ali Umut Kaypak, Shuaihang Yuan, Prashanth Krishnamurthy, Yi Fang 0006, Farshad Khorrami |
ICRA | 3 |
| 2025 | Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language ModelsabstractRobotic navigation plays a pivotal role in a wide range of real-world applications. While traditional navigation systems focus on efficiency and obstacle avoidance, their inability to model complex human behaviors in shared spaces has underscored the growing need for socially aware navigation. In this work, we explore a novel paradigm of socially aware robot navigation empowered by large language models (LLMs), and propose HSAC-LLM, a hybrid framework that seamlessly integrates deep reinforcement learning with the reasoning and communication capabilities of LLMs. Unlike prior approaches that passively predict pedestrian trajectories or issue pre-scripted alerts, HSAC-LLM enables bidirectional natural language interaction, allowing robots to proactively engage in dialogue with pedestrians to resolve potential conflicts and negotiate path decisions. Extensive evaluations across 2D simulations, Gazebo environments, and real-world deployments demonstrate that HSAC-LLM consistently outperforms state-of-the-art DRL baselines under our proposed socially aware navigation metric, which covers safety, efficiency, and human comfort. By bridging linguistic reasoning and interactive motion planning, our results highlight the potential of LLM-augmented agents for robust, adaptive, and human-aligned navigation in real-world settings. Project page: https://hsacllm.github.io/. Congcong Wen, Geeta Chandra Raju Bethala, Shuaihang Yuan, Hao Huang 0003, Mengyu Wang 0001, Yu-Shen Liu, Anthony Tzes, Yi Fang 0006 |
IROS | 4 |
| 2025 | Retrieval-Specific View Learning for Sketch-to-Shape RetrievalabstractSketch-based 3D shape retrieval (SBSR) can be approached by learning domain-invariant descriptors or ranking metrics from sketches and 2D view images of 3D shapes rendered through numerous viewpoints. However, determining the most appropriate viewpoints that convey discriminative geometric features to benefit the task of SBSR became an essential yet not fully explored area. Existing works extract 3D features from multi-view images observed through pre-defined viewpoints to match 2D sketches. Those methods, however, fail to dynamically select viewpoints by considering the SBSR task. In this work, we introduce a fully differentiable viewpoint learning paradigm driven by the downstream SBSR task, which supports the task-aware and sketch-dependent dynamic viewpoint determination process. We naturally integrate this task-specific and sketch-dependent viewpoint learning process into a meta-learning framework to develop a novel Dynamic Viewer (DV) module for SBSR. DV module comprises a Meta View Learner (MVL) block and a View Generator (VG) block. Specifically, as the first part of the DV module, the MVL block learns to initiate the necessary network parameters of the VG block. Then, the VG block that serves as the second part learns the best viewpoints to render 2D images. To learn the optimal viewpoints for SBSR, we further introduce a view mining loss that aims to maximize the similarity of feature-level information among rendered 2D views and the query sketch. Further, we adopt a variational autoencoder (VAE) to retrieve 3D shapes by setting the newly rendered images and query sketch as inputs. As evidenced by the comprehensive experimental results conducted on popular SBSR datasets, the proposed framework has been demonstrated to outperform recent methods in both category-level sketch-based and fine-grained SBSR. Shuaihang Yuan, Congcong Wen, Yu-Shen Liu, Yi Fang 0006 |
IEEE Trans. Multim. | 1 |
| 2024 | FairCLIP: Harnessing Fairness in Vision-Language LearningabstractFairness is a critical concern in deep learning, especially in healthcare, where these models influence diagnoses and treatment decisions. Although fairness has been investigated in the vision-only domain, the fairness of medical vision-language (VL) models remains unexplored due to the scarcity of medical VL datasets for studying fairness. To bridge this research gap, we introduce the first fair vision-language medical dataset (Harvard-FairVLMed) that provides detailed demographic attributes, ground-truth labels, and clinical notes to facilitate an in-depth examination of fairness within VL foundation models. Using Harvard-FairVLMed, we conduct a comprehensive fairness analysis of two widely-used VL models (CLIP and BLIP2), pre-trained on both natural and medical domains, across four different protected attributes. Our results highlight significant biases in all VL models, with Asian, Male, Non-Hispanic, and Spanish being the preferred subgroups across the protected attributes of race, gender, ethnicity, and language, respectively. In order to alleviate these biases, we propose FairCLIP an optimal-transport-based approach that achieves a favorable trade-off between performance and fairness by reducing the Sinkhorn distance between the overall sample distribution and the distributions corresponding to each demographic group. As the first VL dataset of its kind, Harvard-FairVLMed holds the potential to catalyze advancements in the development of machine learning models that are both ethically aware and clinically effective. Our dataset and code are available at https://ophai.hms.harvard.edu/datasets/harvard-fairvlmed10k. Yan Luo 0002, Min Shi 0001, Muhammad Osama Khan, Muhammad Muneeb Afzal, Hao Huang 0003, Shuaihang Yuan, Yu Tian 0001, Luo Song, Ava Kouhana, Tobias Elze, Yi Fang 0006, Mengyu Wang 0001 |
CVPR | 6 |
| 2024 | Optimizing Personalized Robot Actions with Ranking of Trajectories
Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
ICPR (29) | 3 |
| 2024 | Reliable Semantic Understanding for Real World Zero-Shot Object Goal Navigation
Halil Utku Unlu, Shuaihang Yuan, Congcong Wen, Hao Huang 0003, Anthony Tzes, Yi Fang 0006 |
ICPR (30) | 2 |
| 2024 | Zero-Shot Object Navigation with Vision-Language Models Reasoning
Congcong Wen, Yisiyuan Huang, Hao Huang 0003, Yanjia Huang, Shuaihang Yuan, Yu-Shen Liu, Yi Fang 0006 |
ICPR (18) | 5 |
| 2024 | Goal-Driven Transformer for Robot Behavior Learning from Play Data
Congcong Wen, Jiazhao Liang, Shuaihang Yuan, Hao Huang 0003, Yu-Shen Liu, Yi Fang 0006 |
ICPR (30) | 3 |
| 2024 | Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
Shuaihang Yuan, Halil Utku Unlu, Hao Huang 0003, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
ICPR (30) | 1 |
| 2024 | Noisy Few-shot 3D Point Cloud Scene Segmentationabstract3D scene semantic segmentation plays a crucial role in robotics by enabling robots to understand and interpret their environment in a detailed and context-aware manner, facilitating tasks such as navigation, object manipulation, and interaction within complex spaces. A preponderance of methodology predominantly adopts a fully supervised framework for 3D point cloud scene semantic segmentation. Such paradigms exhibit an intrinsic dependency on extensive labeled datasets, presenting challenges in acquisition and exhibiting incapacity to segment novel classes, especially when the training data are contaminated by noisy samples. To address these limitations, this study introduces a novel few-shot segmentation approach to robustly segment 3D point cloud scenes with noisy labels using a meta-learning scheme. Specifically, we first build a multi-prototype graph and then suppress samples with noisy labels based on the graph structure. A subgraph bagging scheme is then proposed to conduct semi-supervised transductive learning to propagate labels. To optimize the graph structure to learn discriminative prototype features, we design a triplet contrastive loss to increase the compactness of these subgraphs. We evaluated our method on two widely used 3D point cloud scene segmentation benchmarks within few-shot (i.e., 2/3-way 5-shot) segmentation settings with noisy samples. Experimental results demonstrate the improvement of our method over the compared baselines, illustrating the robustness of our method in few-shot 3D scene segmentation against noisy samples. The code is available at: https://github.com/hhuang-code/Noisy_Fewshot_Segmentation. Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
ICRA | 2 |
| 2024 | Weakly Scene Segmentation Using Efficient TransformerabstractCurrent methods for large-scale point cloud scene semantic segmentation rely on manually annotated dense point-wise labels, which are costly, labor-intensive, and prone to errors. Consequently, gathering point cloud scenes with billions of labeled points is impractical in real-world scenarios. In this paper, we introduce a novel weak supervision approach to semantically segment large-scale indoor scenes, requiring only 1‰ of the points to be labeled. Specifically, we develop an efficient point neighbor Transformer to capture the geometry of local point cloud patches. To address the quadratic complexity of self-attention computation in Transformers, particularly for large-scale point clouds, we propose approximating the self-attention matrix using low-rank and sparse decomposition. Building on the point neighbor Transformer as foundational blocks, we design a Low-rank Sparse Transformer Network (LST-Net) for weakly supervised large-scale point cloud scene semantic segmentation. Experimental results on two commonly used indoor point cloud scene segmentation benchmarks demonstrate that our model achieves performance comparable to those of both weakly supervised and fully supervised methods. Our code can be found in https://github.com/hhuang-code/LST-Net. Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
IROS | 2 |
| 2024 | ChatMap: A Wearable Platform Based on the Multi-modal Foundation Model to Augment Spatial Cognition for People with Blindness and Low VisionabstractSpatial cognition refers to the ability to gain knowledge about their surroundings and utilize this information to identify their location, acquire resources, and navigate their way back to familiar places. People with blindness and low vision (pBLV) face significant challenges with spatial cognition due to the reliance on visual input. Without the full range of visual cues, pBLV individuals often find it difficult to grasp a comprehensive understanding of their environment, leading to obstacles in scene recognition and precise object localization, especially in unfamiliar environments. This limitation extends to their ability to independently detect and avoid potential tripping hazards, making navigation and interaction with their environment more challenging. In this paper, we present a pioneering wearable platform tailored to enhance the spatial cognition of pBLV through the integration of multi-modal foundation model. The proposed platform integrates a wearable camera with audio module and leverages the advanced capabilities of vision language foundation model (i.e., GPT-4 and GPT-4V), for the nuanced processing of visual and textual data. Specifically, we employ vision language models to bridge the gap between visual information and the proprioception of visually impaired users, offering more intelligible guidance by aligning visual data with the natural perception of space and movement. Then we apply prompt engineering to guide the large language model to act as an assistant tailored specifically for pBLV users to produce accurate answers. Another innovation in our model is the incorporation of a chain of thought reasoning process, which enhances the accuracy and interpretability of the model, facilitating the generation of more precise responses to complex user inquiries across diverse environmental contexts. To assess the practical impact of our proposed wearable platform, we carried out a series of real-world experiments across three tasks that are commonly challenging for people with blindness and low vision: risk assessment, object localization, and scene recognition. Additionally, through an ablation study conducted on the VizWiz dataset, we rigorously assess the contribution of each individual module, substantiating the integral role in the model’s overall performance. Alexey Magay, Hao Huang 0003, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
IROS | 4 |
| 2024 | GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance GuidanceabstractZero-Shot Object Goal Navigation (ZS-OGN) enables robots to navigate toward objects of unseen categories without prior training. Traditional approaches often leverage categorical semantic information for navigation guidance, which struggles when only partial objects are observed or detailed and functional representations of the environment are lacking. To resolve the above two issues, we propose \textit{Geometric-part and Affordance Maps} (GAMap), a novel method that integrates object parts and affordance attributes for navigation guidance. Our method includes a multi-scale scoring approach to capture geometric-part and affordance attributes of objects at different scales. Comprehensive experiments conducted on the HM3D and Gibson benchmark datasets demonstrate improvements in Success Rates and Success weighted by Path Length, underscoring the efficacy of our geometric-part and affordance-guided navigation approach in enhancing robot autonomy and versatility, without any additional task-specific training or fine-tuning with the semantics of unseen objects and/or the locomotions of the robot. Shuaihang Yuan, Hao Huang 0003, Congcong Wen, Anthony Tzes, Yi Fang 0006 |
NeurIPS | 1 |
| 2024 | A single 3D shape wavelet-based generative model
Hao Huang 0003, Shuaihang Yuan, Zheng Peng 0013, Congcong Wen, Yi Fang 0006 |
Comput. Graph. | 2 |
| 2024 | Learning to learn point signature for 3D shape geometry
Hao Huang 0003, Lingjing Wang, Xiang Li 0046, Shuaihang Yuan, Congcong Wen, Yi Fang 0006 |
Pattern Recognit. Lett. | 4 |
| 2022 | Meta-Det3D: Learn to Learn Few-Shot 3D Object Detection
Shuaihang Yuan, Xiang Li 0046, Hao Huang 0003, Yi Fang 0006 |
ACCV (1) | 1 |
| 2022 | SHREC'22 track: Sketch-based 3D shape retrieval in the wild
Jie Qin 0004, Shuaihang Yuan, Jiaxin Chen 0002, Boulbaba Ben Amor, Yi Fang 0006, Nhat Hoang-Xuan, Chi-Bien Chu, Khoi-Nguyen Nguyen-Ngoc, Thien-Tri Cao, Nhat-Khang Ngô, Tuan-Luc Huynh, Hai-Dang Nguyen, Minh-Triet Tran, Haoyang Luo, Jianning Wang, Zheng Zhang 0006, Zihao Xin, Yang Wang 0023, Haiqin Chen, Qunying Zhou |
Comput. Graph. | 2 |
| 2020 | Unsupervised Deep Shape Descriptor With Point Distribution LearningabstractDeep learning models have achieved great success in supervised shape descriptor learning for 3D shape retrieval, classification, and correspondence. However, the unsupervised shape descriptor calculated via deep learning is less studied than that of supervised ones due to the design challenges of unsupervised neural network architecture. This paper proposes a novel probabilistic framework for the learning of unsupervised deep shape descriptors with point distribution learning. In our approach, we firstly associate each point with a Gaussian, and the point clouds are modeled as the distribution of the points. We then use deep neural networks (DNNs) to model a maximum likelihood estimation process that is traditionally solved with an iterative Expectation-Maximization (EM) process. Our key novelty is that ``training'' these DNNs with unsupervised self-correspondence L2 distance loss will elegantly reveal the statically significant deep shape descriptor representation for the distribution of the point clouds. We have conducted experiments over various 3D datasets. Qualitative and quantitative comparisons demonstrate that our proposed method achieves superior classification performance over existing unsupervised 3D shape descriptors. In addition, we verified the following attractive properties of our shape descriptor through experiments: multi-scale shape representation, robustness to shape rotation, and robustness to noise. Yi Shi 0008, Mengchen Xu, Shuaihang Yuan, Yi Fang 0006 |
CVPR | 3 |
| 2020 | 3DMotion-Net: Learning Continuous Flow Function for 3D Motion PredictionabstractThis paper deals with predicting future 3D motions of 3D object scans from the previous two consecutive frames. Previous methods mostly focus on sparse motion prediction in the form of skeletons. While in this paper, we focus on predicting dense 3D motions in the form of 3D point clouds. To approach this problem, we propose a self-supervised approach that leverages the power of the deep neural network to learn a continuous flow function of 3D point clouds that can predict temporally consistent future motions and naturally bring out the correspondences among consecutive point clouds at the same time. More specifically, in our approach, to eliminate the unsolved and challenging process of defining a discrete point convolution on 3D point cloud sequences to encode spatial and temporal information, we introduce a learnable latent code to represent the temporal-aware shape descriptor, which is optimized during the model training. Moreover, a temporally consistent motion Morpher is proposed to learn a continuous flow field which deforms a 3D scan from the current frame to the next frame. We perform extensive experiments on D-FAUST, SCAPE, and TOSCA benchmark data sets. The results demonstrate that our approach is capable of handling temporally inconsistent input and produces consistent future 3D motion while requiring no ground truth supervision. Shuaihang Yuan, Xiang Li 0046, Anthony Tzes, Yi Fang 0006 |
IROS | 1 |
| 2020 | ROSS: Robust Learning of One-Shot 3D Shape Segmentationabstract3D shape segmentation is a fundamental computer vision task that partitions the object into labeled semantic parts. Recent approaches to 3D shape segmentation learning heavily rely on high-quality labeled training datasets. This limits their use in applications to handle the large scale unannotated datasets. In this paper, we proposed a novel semi-supervised approach, named Robust Learning of OneShot 3D Shape Segmentation (ROSS), which only requires one single exemplar labeled shape for training. The proposed ROSS can generalize its ability from a one-shot training process to predict the segmentation for previously unseen 3D shape models. The proposed ROSS is composed of three major modules for 3D shape segmentation as follows. The global shape descriptor generator is the first module that utilizes the proposed reference weighted convolution to learn a 3D shape descriptor. The second module is a part-aware shape descriptor constructor that can generate weighted descriptors from a learned 3D shape descriptor according to semantic parts without supervision. The shape morphing with label transferring works as the last module. It morphs the exemplar shape and then transfers labels from the transformed exemplar shape to the target shape. The extensive experimental results on 3D mesh datasets demonstrate the ROSS is robust to noise and incomplete shapes and it can be applied to unannotated datasets. The experiment shows the proposed ROSS can achieve comparable performance with the supervised method. Shuaihang Yuan, Yi Fang 0006 |
WACV | 1 |
| 2019 | An assistive low-vision platform that augments spatial cognition through proprioceptive guidance: Point-to-Tell-and-TouchabstractSpatial cognition, as gained through the sense of vision, is one of the most important capabilities of human beings. However, for the visually impaired (VI), lack of this perceptual capability poses great challenges in their life. Therefore, we have designed Point-to-Tell-and-Touch, a wearable system with an ergonomic human-machine interface, for assisting the VI with active environmental exploration, with a particular focus on spatial intelligence and navigation to objects of interest in an alien environment. Our key idea is to link visual signals, as decoded synthetically, to the VI's proprioception for more intelligible guidance, in addition to vision-to-audio assistance, i.e., finger pose, as indicated by pointing, is used as “proprioceptive laser pointer” to target an object in that line of sight. The whole system consists of two features, Point-to-Tell and Point-to-Touch, both of which can work independently or cooperatively. The Point-to-Tell feature contains a camera with a novel one-stage neural network tailored for blind-centered object detection and recognition, and a headphone telling the VI the semantic label and distance from the pointed object. the Point-to-Touch, the second feature, leverages a vibrating wrist band to create a haptic feedback tool that supplements the initial vectorial guidance provided by the first stage (hand pose being direction and the distance being the extent, offered through audio cues). Both platform features utilize proprioception or joint position sense. Through hand pose, the VI end user knows where he or she is pointing relative to their egocentric coordinate system and we are able to use this foundation to build spatial intelligence. Our successful indoor experiments demonstrate the proposed system to be effective and reliable in helping the VI gain spatial cognition and explore the world in a more intuitive way. Wenjun Gui, Shuaihang Yuan, John-Ross Rizzo, Lakshay Sharma, Chen Feng 0002, Anthony Tzes, Yi Fang 0006 |
IROS | 3 |