Xiangtong Yao

dblp:248/3613 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0003-2556-3072ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging
abstract
Recent advancements in foundation models, such as the Segment Anything Model (SAM), have significantly impacted medical image segmentation, especially in retinal imaging, where precise segmentation is vital for diagnosis. Despite this progress, current methods face critical challenges: 1) modality ambiguity in textual disease descriptions, 2) a continued reliance on manual prompting for SAM-based workflows, and 3) a lack of a unified framework, with most methods being modalityand task-specific. To overcome these hurdles, we propose CLIP-unified Auto-Prompt Segmentation (CLAPS), a novel method for unified segmentation across diverse tasks and modalities in retinal imaging. Our approach begins by pre-training a CLIP-based image encoder on a large, multi-modal retinal dataset to handle data scarcity and distribution imbalance. We then leverage GroundingDINO to automatically generate spatial bounding box prompts by detecting local lesions. To unify tasks and resolve ambiguity, we use text prompts enhanced with a unique “modality signature” for each imaging modality. Ultimately, these automated textual and spatial prompts guide SAM to execute precise segmentation, creating a fully automated and unified pipeline. Extensive experiments on 12 diverse datasets across 11 critical segmentation categories show that CLAPS achieves performance on par with specialized expert models while surpassing existing benchmarks across most metrics, demonstrating its broad generalizability as a foundation model.
Yinzheng Zhao, Junjie Yang 0001, Xiangtong Yao, Quanmin Liang, Shahrooz Faghih Roohi, Kai Huang 0001, Nassir Navab, M. Ali Nasseri
BIBM4
2025 UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation
abstract
Significant advancements in AI-driven multimodal medical image diagnosis have led to substantial improvements in ophthalmic disease identification in recent years. However, acquiring paired multimodal ophthalmic images remains prohibitively expensive. While fundus photography is simple and cost-effective, the limited availability of OCT data and inherent modality imbalance hinder further progress. Conventional approaches that rely solely on fundus or textual features often fail to capture fine-grained spatial information, as each imaging modality provides distinct cues about lesion predilection sites. In this study, we propose a novel unpaired multimodal framework UOPSL that utilizes extensive OCT-derived spatial priors to dynamically identify predilection sites, enhancing fundus imagebased disease recognition. Our approach bridges unpaired fundus and OCTs via extended disease text descriptions. Initially, we employ contrastive learning on a large corpus of unpaired OCT and fundus images while simultaneously learning the predilection sites matrix in the OCT latent space. Through extensive optimization, this matrix captures lesion localization patterns within the OCT feature space. During the fine-tuning or inference phase of the downstream classification task based solely on fundus images, where paired OCT data is unavailable, we eliminate OCT input and utilize the predilection sites matrix to assist in fundus image classification learning. Extensive experiments conducted on 9 diverse datasets across 28 critical categories demonstrate that our framework outperforms existing benchmarks.
Yinzheng Zhao, Junjie Yang 0001, Xiangtong Yao, Quanmin Liang, Daniel Zapp, Kai Huang 0001, Nassir Navab, M. Ali Nasseri
BIBM4
2025 Predicting the Road Ahead: A Knowledge Graph Based Foundation Model for Scene Understanding in Autonomous Driving
Stefan Schimid, Yicong Li 0013, Lavdim Halilaj, Xiangtong Yao, Wei Cao 0015
ESWC (1)5
2025 Whisker-Based Active Tactile Perception for Contour Reconstruction
Yixuan Dang, Qinyang Xu, Yu Zhang 0182, Xiangtong Yao, Liding Zhang, Zhenshan Bing, Florian Röhrbein, Alois C. Knoll
ICRA4
2025 Adaptive Safety-Critical Control for High-Order Systems: A Real-Time Gaussian Process Approach
abstract
This paper proposes a novel adaptive fast variational sparse Gaussian process (AFVSGP) framework to ensure real-time safety for high-order systems under model uncertainties and dynamic obstacle environments. The framework effectively addresses the challenge of maintaining real-time safety guarantees during unknown trajectory transitions in nonstationary environments. To achieve this, the proposed framework incorporates three key innovations. First, a specialized kernel function is embedded within the VSGP algorithm to decouple control inputs from uncertainties while preserving the convexity of posterior-based safety constraints. Second, an adaptive online incremental learning mechanism is introduced, integrating forgetting capabilities with dynamic reconstruction rules for training datasets and inducing sets, thereby accelerating inference convergence and enabling compact uncertainty prediction with reduced computational complexity. Third, a high-order control barrier function (HOCBF)-based safety filter is developed to synthesize safe control inputs by leveraging the proposed learning model, thereby establishing rigorous probabilistic bounds on the satisfaction of safety specifications. The effectiveness of the proposed framework is validated through both simulation and real-world obstacle avoidance experiments on a 7-DOF Franka robot. The video is available at: https://www.youtube.com/watch?v=2tCKYM_79S8.
Yu Zhang 0182, Long Wen 0003, Zhenshan Bing, Xiangtong Yao, Linghuan Kong, Wei He 0001, Alois C. Knoll
IEEE Trans Autom. Sci. Eng.4
2024 Real-Time Adaptive Safety-Critical Control with Gaussian Processes in High-Order Uncertain Models
abstract
This paper presents an adaptive online learning framework for systems with uncertain parameters to ensure safety-critical control in non-stationary environments. Our approach consists of two phases. The initial phase is centered on a novel sparse Gaussian process (GP) framework. We first integrate a forgetting factor to refine a variational sparse GP algorithm, thus enhancing its adaptability. Subsequently, the hyperparameters of the Gaussian model are trained with a specially compound kernel, and the Gaussian model’s online inferential capability and computational efficiency are strengthened by updating a solitary inducing point derived from newly samples, in conjunction with the learned hyperparameters. In the second phase, we propose a safety filter based on high order control barrier functions (HOCBFs), synergized with the previously trained learning model. By leveraging the compound kernel from the first phase, we effectively address the inherent limitations of GPs in handling high-dimensional problems for real-time applications. The derived controller ensures a rigorous lower bound on the probability of satisfying the safety specification. Finally, the efficacy of our proposed algorithm is demonstrated through real-time obstacle avoidance experiments executed using both simulation platform and a real-world 7-DOF robot.
Yu Zhang 0182, Long Wen 0003, Xiangtong Yao, Zhenshan Bing, Linghuan Kong, Wei He 0001, Alois C. Knoll
ICRA3
2024 Online Efficient Safety-Critical Control for Mobile Robots in Unknown Dynamic Multi-Obstacle Environments
abstract
This paper proposes a LiDAR-based goal-seeking and exploration framework, addressing the efficiency of online obstacle avoidance in unstructured environments populated with static and moving obstacles. This framework addresses two significant challenges associated with traditional dynamic control barrier functions (D-CBFs): their online construction and the diminished real-time performance caused by utilizing multiple D-CBFs. To tackle the first challenge, the framework’s perception component begins with clustering point clouds via the DBSCAN algorithm, followed by encapsulating these clusters with the minimum bounding ellipses (MBEs) algorithm to create elliptical representations. By comparing the current state of MBEs with those stored from previous moments, the differentiation between static and dynamic obstacles is realized, and the Kalman filter is utilized to predict the movements of the latter. Such analysis facilitates the D-CBF’s online construction for each MBE. To tackle the second challenge, we introduce buffer zones, generating Type-II D-CBFs online for each identified obstacle. Utilizing these buffer zones as activation areas substantially reduces the number of D-CBFs that need to be activated. Upon entering these buffer zones, the system prioritizes safety, autonomously navigating safe paths, and hence referred to as the exploration mode. Exiting these buffer zones triggers the system’s transition to goal-seeking mode. We demonstrate that the system’s states under this framework achieve safety and asymptotic stabilization. Experimental results in simulated and real-world environments have validated our framework’s capability, allowing a LiDAR-equipped mobile robot to efficiently and safely reach the desired location within dynamic environments containing multiple obstacles. Video and code are available: https://zyzhang4.wixsite.com/iros2024.
Yu Zhang 0182, Guangyao Tian, Long Wen 0003, Xiangtong Yao, Liding Zhang, Zhenshan Bing, Wei He 0001, Alois C. Knoll
IROS4
2023 Meta-Reinforcement Learning Based on Self-Supervised Task Representation Learning
abstract
Meta-reinforcement learning enables artificial agents to learn from related training tasks and adapt to new tasks efficiently with minimal interaction data. However, most existing research is still limited to narrow task distributions that are parametric and stationary, and does not consider out-of-distribution tasks during the evaluation, thus, restricting its application. In this paper, we propose MoSS, a context-based Meta-reinforcement learning algorithm based on Self-Supervised task representation learning to address this challenge. We extend meta-RL to broad non-parametric task distributions which have never been explored before, and also achieve state-of-the-art results in non-stationary and out-of-distribution tasks. Specifically, MoSS consists of a task inference module and a policy module. We utilize the Gaussian mixture model for task representation to imitate the parametric and non-parametric task variations. Additionally, our online adaptation strategy enables the agent to react at the first sight of a task change, thus being applicable in non-stationary tasks. MoSS also exhibits strong generalization robustness in out-of-distributions tasks which benefits from the reliable and robust task representation. The policy is built on top of an off-policy RL algorithm and the entire network is trained completely off-policy to ensure high sample efficiency. On MuJoCo and Meta-World benchmarks, MoSS outperforms prior works in terms of asymptotic performance, sample efficiency (3-50x faster), adaptation efficiency, and generalization robustness on broad and diverse task distributions.
Mingyang Wang 0003, Zhenshan Bing, Xiangtong Yao, Shuai Wang 0007, Kai Huang 0001, Hang Su 0001, Chenguang Yang 0001, Alois C. Knoll
AAAI3
2023 Meta-Reinforcement Learning via Language Instructions
abstract
Although deep reinforcement learning has recently been very successful at learning complex behaviors, it requires a tremendous amount of data to learn a task. One of the fundamental reasons causing this limitation lies in the nature of the trial-and-error learning paradigm of reinforcement learning, where the agent communicates with the environment and pro-gresses in the learning only relying on the reward signal. This is implicit and rather insufficient to learn a task well. On the con-trary, humans are usually taught new skills via natural language instructions. Utilizing language instructions for robotic motion control to improve the adaptability is a recently emerged topic and challenging. In this paper, we present a meta-RL algorithm that addresses the challenge of learning skills with language instructions in multiple manipulation tasks. On the one hand, our algorithm utilizes the language instructions to shape its in-terpretation of the task, on the other hand, it still learns to solve task in a trial-and-error process. We evaluate our algorithm on the robotic manipulation benchmark (Meta-World) and it significantly outperforms state-of-the-art methods in terms of training and testing task success rates. Codes are available at https://tumi6robot.wixsite.com/million.
Zhenshan Bing, Alexander W. Koch, Xiangtong Yao, Kai Huang 0001, Alois C. Knoll
ICRA3
2023 Learning from Symmetry: Meta-Reinforcement Learning with Symmetrical Behaviors and Language Instructions
abstract
Meta-reinforcement learning (meta-RL) is a promising approach that enables the agent to learn new tasks quickly. However, most meta-RL algorithms show poor generalization in multi-task scenarios due to the insufficient task information provided only by rewards. Language-conditioned meta-RL improves the generalization capability by matching language instructions with the agent's behaviors. While both behaviors and language instructions have symmetry, which can speed up human learning of new knowledge. Thus, combining symmetry and language instructions into meta-RL can help improve the algorithm's generalization and learning efficiency. We propose a dual-MDP meta-reinforcement learning method that enables learning new tasks efficiently with symmetrical behav-iors and language instructions. We evaluate our method in mul-tiple challenging manipulation tasks, and experimental results show that our method can greatly improve the generalization and learning efficiency of meta-reinforcement learning. Videos are available at https://tumi6robot.wixsite.com/symmetry/.
Xiangtong Yao, Zhenshan Bing, Genghang Zhuang, Kejia Chen 0005, Kai Huang 0001, Alois C. Knoll
IROS1
2023 An Energy-Efficient Lane-Keeping System Using 3D LiDAR Based on Spiking Neural Network
abstract
Lane keeping, as a fundamental functionality of autonomous navigation, remains a challenging task for autonomous robots and vehicles. Recently, spiking neural networks (SNNs) have gained attention and research interest due to their biological plausibility and application potential on neuromorphic processors. SNNs have also been successfully deployed on robots to solve autonomous navigation problems. However, lane keeping with a LiDAR sensor is still an open problem for SNNs. In this work, we propose an end-to-end approach based on an SNN to solve the lane-keeping problem using a 3D LiDAR sensor. For the first time, we explore the capability of the proposed SNN controller to perceive the LiDAR input and exploit the features to perform reward-based feedback learning. To ensure the effectiveness of the controller, the proposed method is deployed and evaluated on two high-fidelity simulators. The experimental results demonstrate the high applicability and performance in different scenarios. Furthermore, experiments show that the SNN is capable of performing lane keeping in a simulated urban environment with only 18 control neurons and 32 synapse connections, producing on average only a 17cm deviation from lane center, which is 4.3 % of the lane width.
Genghang Zhuang, Zhenshan Bing, Xiangtong Yao, Yuhong Huang, Kai Huang 0001, Alois C. Knoll
IROS4
2023 Toward Intelligent Sensing: Optimizing Lidar Beam Distribution for Autonomous Driving
abstract
LiDAR (Light Detection And Ranging) sensors have been widely used in autonomous vehicles as the main sensors. According to the specification details of the widely used 3D LiDAR products in the market, the distribution of vertical beam channels is set according to a uniform angular resolution, which is not ideally efficient for specific autonomous tasks. In this paper, we propose a novel approach to find the optimized angular distribution of the vertical beam channels for different application scenarios and installation configurations. The experimental results in a study case suggest that concerning the vehicle detection task, the optimized LiDARs perform almost two times better than the ones with the same number of channels in terms of the detection range, and have perception performances close to the LiDARs with double channels in the long distance.
Genghang Zhuang, Zhenshan Bing, Xiangtong Yao, Yuhong Huang, Kai Huang 0001, Alois C. Knoll
IEEE Trans. Intell. Transp. Syst.3
2019 LiDAR Based Navigable Region Detection for Unmanned Surface Vehicles
abstract
Detection of the navigable regions for the unmanned surface vehicles (USVs) sailing on the narrow rivers is very important. Existing detection methods mostly depend on the cameras, which is sensitive to environments and cannot provide reliable navigable regions for sailing. In this paper, we propose a scheme to process 3D LiDAR data to achieve an accurate and robust navigable regions detection. We conduct field experiments in a narrow river in different scenarios to prove the performance of the proposed scheme, which reaches on average 93.8% precision and 92.7% recall.
Xiangtong Yao, Yunxiao Shan, Jieling Li, Donghui Ma, Kai Huang 0001
IROS1
2019 An Adaptive Path Tracking Controller Based on Reinforcement Learning with Urban Driving Application
abstract
Urban driving requires the autonomous vehicles to drive with smooth control and track the planned path accurately. However, most of the existing path-tracking controllers pay more attention to the tracking errors than the smoothness because of the difficulties to balance them. This paper proposes a learning-based method to achieve the trade-off between the smooth control and the tracking-error control. An Reinforcement Learning algorithm, which is called Proximal Policy Optimization, is used to train a neural model to tune the weights of a designed controller PP_PID (Pure-Pursuit_Proportional Integral Derivative). The successfully trained model will adaptively select the optimal weights for the Pure-Pursuit and PID to guarantee the control smoothness and accuracy. Finally, the proposed controller will be tested in two path tracking scenarios. The results show that the proposed controller can change the weights adaptively to maintain a balance in the tracking error and lateral acceleration under the 35km/h.
Longsheng Chen, Yuanpeng Chen, Xiangtong Yao, Yunxiao Shan, Long Chen 0005
IV3