Jun Ma 0008

dblp:91/4845-8 · DBLP profile ↗
← Back
72ranked-venue papers
8as first author
66since 2021 · last 2026
0000-0002-9405-8232ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 2 first-author · 36 since 2021Systems, architecture and hardware · 31 · 1 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DRM-Net: Explicit Residual Modelling with Subaquatic Multi-Scale Context Fusion for Underwater Image Enhancement
abstract
Clear and high-quality underwater images are essential for marine applications, including autonomous navigation, ecological monitoring, and infrastructure inspection. However, underwater images typically suffer from severe colour distortion, low contrast, and diminished structural visibility due to wavelength-dependent attenuation, scattering, and uneven illumination conditions. Recent deep learning-based underwater image enhancement (UIE) methods primarily adopt end-to-end frameworks, directly regressing enhanced images from degraded inputs. While these approaches have achieved significant progress, they often lack explicit modeling of the degradation process, leading to limited interpretability and suboptimal recovery of fine-grained details. To address these limitations, we propose DRM-Net, an explicit residual learning framework for UIE. Rather than estimating the enhanced image directly, DRM-Net first predicts a pixel-wise Degradation Residual Map (DRM) in the perceptually uniform CIELab colour space. This map explicitly quantifies local colour, contrast, and structural degradations, thereby enabling the network to precisely reconstruct missing visual information. Furthermore, we design a lightweight Subaquatic Multi-Scale Context Fusion module, which utilizes parallel atrous convolutions with softmax-weighted feature aggregation, significantly enhancing robustness against spatially heterogeneous scattering. Trained jointly with pixel-wise DRM and VGG-based perceptual losses, DRM-Net achieves superior colour fidelity, perceptual realism, and structural detail recovery. Comprehensive experiments conducted on multiple benchmarks demonstrate that our proposed approach attains competitive quantitative results and superior qualitative visual performance compared to state-of-the-art UIE methods, while maintaining low computational overhead, making it particularly suitable for resource-constrained underwater robotic systems.
Chang Huang, Zhexin Zhou, Jun Ma 0008, Jiatong Shen, Peixuan Xiong, Huayong Yang, Kaishun Wu
AAAI3
2026 Toward Real-World High-Precision Image Matting and Segmentation
abstract
High-precision scene parsing tasks, including image matting and dichotomous segmentation, aim to accurately predict masks with extremely fine details (such as hair). Most existing methods focus on salient, single foreground objects. While interactive methods allow for target adjustment, their class-agnostic design restricts generalization across different categories. Furthermore, the scarcity of high-quality annotation has led to a reliance on inharmonious synthetic data, resulting in poor generalization to real-world scenarios. To this end, we propose a Foreground Consistent Learning model, dubbed as FCLM, to address the aforementioned issues. Specifically, we first introduce a Depth-Aware Distillation strategy where we transfer the depth-related knowledge for better foreground representation. Considering the data dilemma, we term the processing of synthetic data as domain adaptation problem where we propose a domain-invariant learning strategy to focus on foreground learning. To support interactive prediction, we contribute an Object-Oriented Decoder that can receive both visual and language prompts to predict the referring target. Experimental results show that our method quantitatively and qualitatively outperforms state-of-the-art methods.
Haipeng Zhou, Zhaohu Xing, Hongqiu Wang, Jun Ma 0008, Ping Li 0016, Lei Zhu 0003
AAAI4
2026 nuPlan-R: A Closed-Loop Planning Benchmark for Autonomous Driving via Reactive Multi-Agent Simulation
abstract
Recent advances in closed-loop planning benchmarks have significantly improved the evaluation of autonomous vehicles. However, existing benchmarks still rely on rule-based reactive agents such as the Intelligent Driver Model (IDM), which lack behavioral diversity and fail to capture realistic human interactions, leading to oversimplified traffic dynamics. To address these limitations, we present nuPlan-R, a new reactive closed-loop planning benchmark that integrates learning-based reactive multi-agent simulation into the nuPlan framework. Our benchmark replaces the rule-based IDM agents with noise-decoupled diffusion-based reactive agents and introduces an interaction-aware agent selection mechanism to ensure both realism and computational efficiency. Furthermore, we extend the benchmark with two additional metrics to enable a more comprehensive assessment of planning performance. Extensive experiments demonstrate that our reactive agent model produces more realistic, diverse, and human-like traffic behaviors, leading to a benchmark environment that better reflects real-world interactive driving. We further reimplement a collection of rule-based, learning-based, and hybrid planning approaches within our nuPlan-R benchmark, providing a clearer reflection of planner performance in complex interactive scenarios and better highlighting the advantages of learning-based planners in handling complex and dynamic scenarios. These results establish nuPlan-R as a new standard for fair, reactive, and realistic closed-loop planning evaluation. We will open-source the code for the new benchmark. We have open-sourced our framework at https://github.com/Pemixing/nuPlan-R.
Mingxing Peng, Ruoyu Yao, Xusen Guo, Jun Ma 0008
IV4
2026 Occlusion-Aware Contingency Safety-Critical Planning for Autonomous Driving
abstract
Ensuring safe driving while maintaining travel efficiency for autonomous vehicles (AVs) in dynamic and occluded environments is a critical challenge. This article proposes an occlusion-aware contingency safety-critical planning approach for real-time autonomous driving. Leveraging reachability analysis for risk assessment, forward reachable sets (FRSs) of phantom vehicles (PVs) are used to derive risk-aware dynamic velocity boundaries. These velocity boundaries are incorporated into a biconvex nonlinear programming (NLP) formulation that formally enforces safety using spatiotemporal barrier constraints, while simultaneously optimizing exploration and fallback trajectories within a receding horizon planning framework. To enable real-time computation and coordination between trajectories, we employ the consensus alternating direction method of multipliers (ADMMs) to decompose the biconvex NLP problem into low-dimensional convex subproblems. The effectiveness of the proposed approach is validated through simulations and real-world experiments in occluded intersections. Experimental results demonstrate enhanced safety and improved travel efficiency, enabling real-time safe trajectory generation in dynamic occluded intersections under varying obstacle conditions. The project page is available at: https://zack4417.github.io/oacp-website/.
Lei Zheng 0007, Rui Yang 0025, Minzhe Zheng, Zengqi Peng, Michael Yu Wang, Jun Ma 0008
IEEE Trans. Cybern.6
2026 SCSV: Spatial-Temporal Consistent Dynamic 3D Scene Generation From Sparse Views
abstract
Generating dynamic scenes from images has gained increasing attention. Existing methods have two major limitations: 1) they can hardly handle sparse images which exhibit limited geometry constraints and insufficient motion; 2) they struggle to maintain spatial-temporal consistency when rendering multi-view videos. To address these limitations, we propose SCSV, a spatial-temporal consistent dynamic scene generation method from sparse views. Our method consists of two stages: scene reconstruction and scene expansion, both of which decouple background and foreground. In the scene reconstruction stage, we first interpolate a set of images between the input images based on a video generation model, followed by the optimization of the scene Gaussian from the interpolated and input images. To improve the spatial-temporal consistency of the reconstructed scene, we propose an uncertainty-aware Gaussian training approach, which introduces adaptive weights of images and pixels. In the scene expansion stage, for background, we render novel views and refine them with a geometry-aware diffusion process. These refined images are then used to incrementally add the Gaussians. As to foreground, we generate human motion according to previous motion, enabling temporal coherent generation of motion. To further enhance the physical plausibility, we integrate the expanded foreground into the background using a gravity-aware alignment. Experiments on NeuMan, Bonn, and EMDB datasets demonstrate that our SCSV achieves superior performance compared to state-of-the-art methods. The code will be released upon acceptance.
Shunbo Zhou, Jun Ma 0008, Hesheng Wang 0001, Haoang Li
IEEE Trans. Image Process.6
2026 Safe and Real-Time Consistent Planning for Autonomous Vehicles in Partially Observed Environments via Parallel Consensus Optimization
abstract
Ensuring safety and driving consistency is a significant challenge for autonomous vehicles operating in partially observed environments. This work introduces a consistent parallel trajectory optimization (CPTO) approach to enable safe and consistent driving in dense obstacle environments with perception uncertainties. Utilizing discrete-time barrier function theory, we develop a consensus safety barrier module that ensures reliable safety coverage within the spatiotemporal trajectory space across potential obstacle configurations. Following this, a bi-convex parallel trajectory optimization problem is derived that facilitates decomposition into a series of low-dimensional quadratic programming problems to accelerate computation. By leveraging the consensus alternating direction method of multipliers (ADMM) for parallel optimization, each generated candidate trajectory corresponds to a possible environment configuration while sharing a common consensus trajectory segment. This ensures driving safety and consistency when executing the consensus trajectory segment for the ego vehicle in real time. We validate our CPTO framework through extensive comparisons with state-of-the-art baselines across multiple driving tasks in partially observable environments. Our results demonstrate improved safety and consistency using both synthetic and real-world traffic datasets.
Lei Zheng 0007, Rui Yang 0025, Minzhe Zheng, Michael Yu Wang, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.5
2026 Fast and Scalable Game-Theoretic Trajectory Planning With Intentional Uncertainties
abstract
Trajectory planning involving multi-agent interactions has been a long-standing challenge in the field of robotics, primarily burdened by the inherent yet intricate interactions among agents. While game-theoretic methods are widely acknowledged for their effectiveness in managing multi-agent interactions, significant impediments persist when it comes to accommodating the intentional uncertainties of agents. In the context of intentional uncertainties, the heavy computational burdens associated with existing game-theoretic methods are induced, leading to inefficiencies and poor scalability. In this paper, we propose a novel game-theoretic interactive trajectory planning method to effectively address the intentional uncertainties of agents, and it demonstrates both high efficiency and enhanced scalability. As the underpinning basis, we model the interactions between agents under intentional uncertainties as a static Bayesian game, and we show that its agent-form equivalence can be represented as a potential game under certain assumptions. The existence and attainability of the optimal interactive trajectories are illustrated, as the corresponding static Bayesian Nash equilibrium can be attained by optimizing a unified optimization problem. Additionally, we present a distributed algorithm based on the dual consensus alternating direction method of multipliers (ADMM) tailored to the parallel solving of the problem, thereby significantly improving the scalability. The attendant outcomes from simulations and experiments demonstrate that the proposed method is effective across a range of scenarios characterized by general forms of intentional uncertainties. Its scalability surpasses that of existing centralized and decentralized baselines, allowing for real-time interactive trajectory planning in uncertain game settings. The source code will be available onhttps://github.com/zhuangdf/Potential-Bayesian-Game-release.
Zhenmin Huang, Yusen Xie, Benshan Ma, Shaojie Shen, Jun Ma 0008
IEEE Trans. Robotics5
2025 A Simple Data Augmentation for Feature Distribution Skewed Federated Learning
abstract
Federated Learning (FL) facilitates collaborative learning among multiple clients in a distributed manner and ensures the security of privacy. However, its performance inevitably degrades with non-Independent and Identically Distributed (non-IID) data. In this paper, we focus on the feature distribution skewed FL scenario, a common non-IID situation in real-world applications where data from different clients exhibit varying underlying distributions. This variation leads to feature shift, which is a key issue of this scenario. While previous works have made notable progress, few pay attention to the data itself, i.e., the root of this issue. The primary goal of this paper is to mitigate feature shift from the perspective of data. To this end, we propose a simple yet remarkably effective input-level data augmentation method, namely FedRDN, which randomly injects the statistical information of the local distribution from the entire federation into the client’s data. This is beneficial to improve the generalization of local feature representations, thereby mitigating feature shift. Moreover, our FedRDN is a plug-and-play component, which can be seamlessly integrated into the data augmentation flow with only a few lines of code. Extensive experiments on several datasets show that the performance of various representative FL methods can be further improved by integrating our FedRDN, demonstrating its effectiveness, strong compatibility and generalizability. Code is available at https://github.com/IAMJackYan/FedRDN.
Yunlu Yan, Huazhu Fu, Yuexiang Li, Jinheng Xie, Jun Ma 0008, Guang Yang 0006, Lei Zhu 0003
CVPR5
2025 MF-BERT: A Siamese Pre-training Framework for Motion Forecasting
abstract
Accurately predicting the future motions of traffic agents is essential for autonomous systems. Despite the significant success of existing motion forecasting methods based on supervised learning, they still exhibit two main limitations. First, when annotated data for a scene is limited, these methods often fail to achieve the expected accuracy. Second, they typically rely on complex architectures and extensive prior knowledge to improve performance. To overcome these challenges, we propose MF-BERT, a novel framework that adapts the concept of BERT to motion forecasting, inspired by advancements in the self-supervised pre-training paradigm. During pre-training, we design a siamese sequence modeling task with an asymmetric mask strategy to capture complex behavior patterns of agents. During fine-tuning, the pre-trained representation module initializes the feature encoder of the motion forecasting model, and a multimodal trajectory decoder generates all possible predictions. Experimental results demonstrate the superiority of MF-BERT over state-of-the-art methods.
Jianxin Shi 0004, Jun Ma 0008, Tianyu Wo
ICASSP4
2025 GS-LIVM: Real-Time Photo-Realistic LiDAR-Inertial-Visual Mapping with Gaussian Splatting
abstract
In this paper, we introduce GS-LIVM, a real-time photo-realistic LiDAR-Inertial-Visual mapping framework with Gaussian Splatting tailored for outdoor scenes. Compared to existing methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), our approach enables real-time photo-realistic mapping while ensuring high-quality image rendering in large-scale unbounded outdoor environments. In this work, Gaussian Process Regression (GPR) is employed to mitigate the issues resulting from sparse and unevenly distributed LiDAR observations. The voxel-based 3D Gaussians map representation facilitates real-time dense mapping in large outdoor environments with acceleration governed by custom CUDA kernels. Moreover, the overall framework is designed in a covariance-centered manner, where the estimated covariance is used to initialize the scale and rotation of 3D Gaussians, as well as update the parameters of the GPR. We evaluate our algorithm on several outdoor datasets, and the results demonstrate that our method achieves state-of-the-art performance in terms of mapping efficiency and rendering quality. The source code is available on GitHub.
Yusen Xie, Zhenmin Huang, Jin Wu 0002, Jun Ma 0008
ICCV4
2025 ScNet: Scene-Consistency Network Learning for Multi-Agent Motion Forecasting
abstract
Predicting the motion of traffic agents is a fundamental challenge in autonomous driving, essential for safe and efficient ego-vehicle planning. Traditional methods typically focus on marginal forecasting, where the trajectory of each agent is predicted separately, leading to inconsistencies in scene-level predictions. To address this issue, we propose a scene-consistency network, named ScNet, which jointly predicts the trajectories of multiple agents in a single feedforward pass, ensuring consistency across all predictions. Our method leverages dual independently initialized student models that interact through cross-network contrastive learning at the global feature level, enhancing robustness and scene consistency in the learned representations. To further improve scene coherence, we incorporate a scene-guided strategy that refines these representations. Additionally, we employ a lightweight, anchor-free decoder that generates predictions for all agents, aligning the forecasts with real-world dynamics. Experiments show significant improvements in multi-world prediction metrics across complex environments. Code and models will be publicly available.
Jianxin Shi 0004, Yusen Xie, Fali Wang, Jun Ma 0008, Tianyu Wo
ICME6
2025 Robot Navigation in Unknown and Cluttered Workspace with Dynamical System Modulation in Starshaped Roadmap
abstract
Compared to conventional decomposition methods that use ellipses or polygons to represent free space, starshaped representation can better capture the natural distribution of sensor data, thereby exploiting a larger portion of traversable space. This paper introduces a novel motion planning and control framework for navigating robots in unknown and cluttered environments using a dynamically constructed starshaped roadmap. Our approach generates a starshaped representation of the surrounding free space from real-time sensor data using piece-wise polynomials. Additionally, an incremental roadmap maintaining the connectivity information is constructed, and a searching algorithm efficiently selects short-term goals on this roadmap. Importantly, this framework addresses dead-end situations with a graph updating mechanism. To ensure safe and efficient movement within the starshaped roadmap, we propose a reactive controller based on Dynamic System Modulation (DSM). This controller facilitates smooth motion within starshaped regions and their intersections, avoiding conservative and short-sighted behaviors and allowing the system to handle intricate obstacle configurations in unknown and cluttered environments. Comprehensive evaluations in both simulations and real-world experiments show that the proposed method achieves higher success rates and reduced travel times compared to other methods. It effectively manages intricate obstacle configurations, avoiding conservative and myopic behaviors. The source code will be released on website11Available at: github.com/kkkkkaiai/starshaped_roadmap.
Kai Chen 0006, Haichao Liu 0003, Yulin Li 0001, Jianghua Duan, Lei Zhu 0003, Jun Ma 0008
ICRA6
2025 Scene-Aware Explainable Multimodal Trajectory Prediction
abstract
Advancements in intelligent technologies have significantly improved navigation in complex traffic environments by enhancing environment perception and trajectory prediction for automated vehicles. However, current research often overlooks the joint reasoning of scenario agents and lacks explainability in trajectory prediction models, limiting their practical use in real-world situations. To address this, we introduce the Explainable Conditional Diffusion-based Multimodal Trajectory Prediction (DMTP) model, which is designed to elucidate the environmental factors influencing predictions and reveal the underlying mechanisms. Our model integrates a modified conditional diffusion approach to capture multimodal trajectory patterns and employs a revised Shapley Value model to assess the significance of global and scenario-specific features. Experiments using the Waymo Open Motion Dataset demonstrate that our explainable model excels in identifying critical inputs and significantly outperforms baseline models in accuracy. Moreover, the factors identified align with the human driving experience, underscoring the model's effectiveness in learning accurate predictions. Code is available in our open-source repository: https://github.com/ocean-luna/Explainable-Prediction.
Junlan Chen, Yangfan He, Jun Ma 0008
ICRA7
2025 Task-Oriented Pre-Training for Drivable Area Detection
abstract
Pre-training techniques play a crucial role in deep learning, enhancing models' performance across a variety of tasks. By initially training on large datasets and subsequently fine-tuning on task-specific data, pre-training provides a solid foundation for models, improving generalization abilities and accelerating convergence rates. This approach has seen significant success in the fields of natural language processing and computer vision. However, traditional pre-training methods necessitate large datasets and substantial computational resources, and they can only learn shared features through prolonged training and struggle to capture deeper, task-specific features. In this paper, we propose a task-oriented pre-training method that begins with generating redundant segmentation proposals using the Segment Anything (SAM) model. We then introduce a Specific Category Enhancement Fine-tuning (SCEF) strategy for fine-tuning the Contrastive Language-Image Pre-training (CLIP) model to select proposals most closely related to the drivable area from those generated by SAM. This approach can generate a lot of coarse training data for pre-training models, which are further fine-tuned using manually annotated data, thereby improving model's performance. Comprehensive experiments conducted on the KITTI road dataset demonstrate that our task-oriented pre-training method achieves an all-around performance improvement compared to models without pre-training (as shown in Fig. 1). Moreover, our pre-training method not only surpasses traditional pre-training approach but also achieves the best performance compared to state-of-the-art self-training methods. The open-source project can be found at https://sites.google.com/view/task-oriented-pre-training.
Fulong Ma, Guoyang Zhao, Weiqing Qi, Ming Liu 0001, Jun Ma 0008
ICRA5
2025 FisheyeDepth: A Real Scale Self-Supervised Depth Estimation Model for Fisheye Camera
abstract
Accurate depth estimation is crucial for 3D scene comprehension in robotics and autonomous vehicles. Fisheye cameras, known for their wide field of view, have inherent geometric benefits. However, their use in depth estimation is restricted by a scarcity of ground truth data and image distortions. We present FisheyeDepth, a self-supervised depth estimation model tailored for fisheye cameras. We incorporate a fisheye camera model into the projection and reprojection stages during training to handle image distortions, thereby improving depth estimation accuracy and training stability. Furthermore, we incorporate real-scale pose information into the geometric projection between consecutive frames, replacing the poses estimated by the conventional pose network. Essentially, this method offers the necessary physical depth for robotic tasks, and also streamlines the training and inference procedures. Additionally, we devise a multi-channel output strategy to improve robustness by adaptively fusing features at various scales, which reduces the noise from real pose data. We demonstrate the superior performance and robustness of our model in fisheye image depth estimation through evaluations on public datasets and real-world scenarios. The project website is available at: https://github.com/guoyangzhaolFisheyeDepth.
Guoyang Zhao, Yuxuan Liu 0008, Weiqing Qi, Fulong Ma, Ming Liu 0001, Jun Ma 0008
ICRA6
2025 TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition
abstract
Traffic sign is a critical map feature for navigation and traffic control. Nevertheless, current methods for traffic sign recognition rely on traditional deep learning models, which typically suffer from significant performance degradation considering the variations in data distribution across different regions. In this paper, we propose TSCLIP, a robust fine-tuning approach with the contrastive language-image pre-training (CLIP) model for worldwide cross-regional traffic sign recognition. We first curate a cross-regional traffic sign benchmark dataset by combining data from ten different sources. Then, we propose a prompt engineering scheme tailored to the characteristics of traffic signs, which involves specific scene descriptions and corresponding rules to generate targeted text descriptions. During the TSCLIP fine-tuning process, we implement adaptive dynamic weight ensembling (ADWE) to seamlessly incorporate outcomes from each training iteration with the zero-shot CLIP model. This approach ensures that the model retains its ability to generalize while acquiring new knowledge about traffic signs. To the best knowledge of authors, TSCLIP is the first contrastive language-image model used for the worldwide cross-regional traffic sign recognition task. The project website is available at: https://github.com/guoyangzhao/TSCLIP.
Guoyang Zhao, Fulong Ma, Weiqing Qi, Yuxuan Liu 0008, Ming Liu 0001, Jun Ma 0008
ICRA7
2025 Interactive Navigation for Legged Manipulators with Learned Arm-Pushing Controller
abstract
Interactive navigation is crucial in scenarios where proactively interacting with objects can yield shorter paths, thus significantly improving traversal efficiency. Existing methods primarily focus on using the robot body to relocate obstacles during navigation. However, they prove ineffective in narrow or constrained spaces where the robot’s dimensions restrict its manipulation capabilities. This paper introduces a novel interactive navigation framework for legged manipulators, featuring an active arm-pushing mechanism that enables the robot to reposition movable obstacles in space-constrained environments. To this end, we develop a reinforcement learning-based arm-pushing controller with a two-stage reward strategy for object manipulation. Specifically, this strategy first directs the manipulator to a designated pushing zone to achieve a kinematically feasible contact configuration. Then, the end effector is guided to maintain its position at appropriate contact points for stable object displacement while preventing toppling. The simulations validate the robustness of the arm-pushing controller, showing that the two-stage reward strategy improves policy convergence and long-term performance. Real-World experiments further demonstrate the effectiveness of the proposed navigation framework, which achieves shorter paths and reduced traversal time. The open-source project can be found at https://zhihaibi.github.io/interactive-push.github.io/.
Zhihai Bi, Kai Chen 0006, Chunxin Zheng, Yulin Li 0001, Haoang Li, Jun Ma 0008
IROS6
2025 Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models
abstract
Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing for precise prediction of action trajectories. However, diffusion models typically rely on large parameter UNet backbones as policy networks, which can be challenging to deploy on resource-constrained devices. Recently, the Mamba model has emerged as a promising solution for efficient modeling, offering low computational complexity and strong performance in sequence modeling. In this work, we propose the Mamba Policy, a lighter but stronger policy that reduces the parameter count by over 80% compared to the original policy network while achieving superior performance. Specifically, we introduce the XMamba Block, which effectively integrates input information with conditional features and leverages a combination of Mamba and Attention mechanisms for deep feature extraction. Extensive experiments demonstrate that the Mamba Policy excels on the Adroit, Dexart, and MetaWorld datasets, requiring significantly fewer computational resources. Additionally, we highlight the Mamba Policy’s enhanced robustness in long-horizon scenarios compared to baseline methods and explore the performance of various Mamba variants within the Mamba Policy framework. Real-world experiments are also conducted to further validate its effectiveness. Our open-source project page can be found at https://andycao1125.github.io/mamba_policy/.
Jiahang Cao, Qiang Zhang 0029, Jingkai Sun, Hao Cheng 0015, Yulin Li 0001, Jun Ma 0008, Kun Wu 0001, Yecheng Shao, Yijie Guo, Renjing Xu
IROS7
2025 DuLoc: Life-Long Dual-Layer Localization in Changing and Dynamic Expansive Scenarios
abstract
LiDAR-based localization serves as a critical component in autonomous systems, yet existing approaches face persistent challenges in balancing repeatability, accuracy, and environmental adaptability. Traditional point cloud registration methods relying solely on offline maps often exhibit limited robustness against long-term environmental changes, leading to localization drift and reliability degradation in dynamic real-world scenarios. To address these challenges, this paper proposes DuLoc, a robust and accurate localization method that tightly couples LiDAR-inertial odometry with offline map-based localization, incorporating a constant-velocity motion model to mitigate outlier noise in real-world scenarios. Specifically, we develop a LiDAR-based localization framework that seamlessly integrates a prior global map with dynamic real-time local maps, enabling robust localization in unbounded and changing environments. Extensive real-world experiments in ultra unbounded port that involve 2,856 hours of operational data across 32 Intelligent Guided Vehicles (IGVs) are conducted and reported in this study. The results attained demonstrate that our system outperforms other state-of-the-art LiDAR localization systems in large-scale changing outdoor environments.
Haoxuan Jiang, Peicong Qian, Yusen Xie, Xiaocong Li, Ming Liu 0001, Jun Ma 0008
IROS6
2025 RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation
abstract
This paper introduces RoboDexVLM, an innovative framework for robot task planning and grasp detection tailored for a collaborative manipulator equipped with a dexterous hand. Previous methods focus on simplified and limited manipulation tasks, which often neglect the complexities associated with grasping a diverse array of objects in a long-horizon manner. In contrast, our proposed framework utilizes a dexterous hand capable of grasping objects of varying shapes and sizes while executing tasks based on natural language commands. The proposed approach has the following core components: First, a robust task planner with a task-level recovery mechanism that leverages vision-language models (VLMs) is designed, which enables the system to interpret and execute open-vocabulary commands for long sequence tasks. Second, a language-guided dexterous grasp perception algorithm is presented based on robot kinematics and formal methods, tailored for zero-shot dexterous manipulation with diverse objects and commands. Comprehensive experimental results validate the effectiveness, adaptability, and robustness of RoboDexVLM in handling long-horizon scenarios and performing dexterous grasping. These results highlight the framework’s ability to operate in complex environments, showcasing its potential for open-vocabulary dexterous manipulation. Our open-source project page can be found at https://henryhcliu.github.io/robodexvlm.
Haichao Liu 0003, Sikai Guo, Pengfei Mai, Jiahang Cao, Haoang Li, Jun Ma 0008
IROS6
2025 LMMCoDrive: Cooperative Driving with Large Multimodal Models
abstract
To address the intricate challenges of cooperative scheduling and motion planning in Autonomous Mobility-on-Demand (AMoD) systems, this paper introduces LMMCoDrive, a novel cooperative driving framework that leverages a Large Multimodal Model (LMM) to improve traffic efficiency and passenger experience in dynamic urban environments. This framework seamlessly integrates scheduling and motion planning processes to ensure the effective operation of Cooperative Autonomous Vehicles (CAVs). The spatial relationship between CAVs and passenger requests is abstracted into a Bird’s-Eye View (BEV) image to fully exploit the potential of the multimodal understanding ability of LMMs. Besides, trajectories are cautiously refined for each CAV while ensuring collision avoidance through safety constraints. A decentralized optimization strategy, facilitated by the Alternating Direction Method of Multipliers (ADMM) within the LMM framework, is proposed to drive the graph evolution of CAVs. Simulation results in diverse urban scenarios demonstrate the pivotal role and significant impact of LMM in optimizing CAV scheduling and seamlessly serving a decentralized cooperative optimization process for each CAV. This marks a substantial stride towards practical, efficient, and safe AMoD systems that are poised to revolutionize urban transportation. The code is available at https://github.com/henryhcliu/LMMCoDrive.
Haichao Liu 0003, Ruoyu Yao, Zhenmin Huang, Shaojie Shen, Jun Ma 0008
IROS5
2025 Annotation-Free Curb Detection Leveraging Altitude Difference Image
abstract
Road curbs are considered as one of the crucial and ubiquitous traffic features, which are essential for ensuring the safety of autonomous vehicles. Current methods for detecting curbs primarily rely on camera imagery or LiDAR point clouds. Image-based methods are vulnerable to fluctuations in lighting conditions and exhibit poor robustness, while methods based on point clouds circumvent the issues associated with lighting variations. However, it is the typical case that significant processing delays are encountered due to the voluminous amount of 3D points contained in each frame of the point cloud data. Furthermore, the inherently unstructured characteristics of point clouds poses challenges for integrating the latest deep learning advancements into point cloud data applications. To address these issues, this work proposes an annotation-free curb detection method leveraging Altitude Difference Image (ADI) (as shown in Fig. 1), which effectively mitigates the aforementioned challenges. Given that methods based on deep learning generally demand extensive, manually annotated datasets, which are both expensive and labor-intensive to create, we present an Automatic Curb Annotator (ACA) module. This module utilizes a deterministic curb detection algorithm to automatically generate a vast quantity of training data. Consequently, it facilitates the training of the curb detection model without necessitating any manual annotation of data. Finally, by incorporating a post-processing module, we manage to achieve state-of-the-art results on the KITTI 3D curb dataset [1] with considerably reduced processing delays compared to existing methods, which underscores the effectiveness of our approach in curb detection tasks. Our code and data will be open-sourced at: https://sites.google.com/view/adi-curb-detection.
Fulong Ma, Yuxuan Liu 0008, Ming Liu 0001, Jun Ma 0008
IROS6
2025 PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
abstract
Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in VLA models with increased chunking sizes. This reduces the inference efficiency. Therefore, accelerating VLA integrated with action chunking is an urgent need. To tackle this problem, we propose PD-VLA, the first parallel decoding framework for VLA models integrated with action chunking. Our framework reformulates autoregressive decoding as a nonlinear system solved by parallel fixed-point iterations. This approach preserves model performance with mathematical guarantees while significantly improving decoding speed. In addition, it enables training-free acceleration without architectural changes, as well as seamless synergy with existing acceleration techniques. Extensive simulations validate that our PD-VLA maintains competitive success rates while achieving 2.52× execution frequency on manipulators (with 7 degrees of freedom) compared with the fundamental VLA model. Furthermore, we experimentally identify the most effective settings for acceleration. Finally, real-world experiments validate its high applicability across different tasks.
Wenxuan Song, Pengxiang Ding, Han Zhao 0008, Zhide Zhong, ZongYuan Ge, Jun Ma 0008, Haoang Li
IROS11
2025 GDTS: Goal-Guided Diffusion Model with Tree Sampling for Multi-Modal Pedestrian Trajectory Prediction
abstract
Accurate prediction of pedestrian trajectories is crucial for improving the safety of autonomous driving. However, this task is generally nontrivial due to the inherent stochasticity of human motion, which naturally requires the predictor to generate multi-modal prediction. Previous works leverage various generative methods, such as GAN and VAE, for pedestrian trajectory prediction. Nevertheless, these methods may suffer from mode collapse and relatively low-quality results. The denoising diffusion probabilistic model (DDPM) has recently been applied to trajectory prediction due to its simple training process and powerful reconstruction ability. However, current diffusion-based methods do not fully utilize input information and usually require many denoising iterations that lead to a long inference time or an additional network for initialization. To address these challenges and facilitate the use of diffusion models in multi-modal trajectory prediction, we propose GDTS, a novel Goal-Guided Diffusion Model with Tree Sampling for multi-modal trajectory prediction. Considering the "goal-driven" characteristics of human motion, GDTS leverages goal estimation to guide the generation of the diffusion network. A two-stage tree sampling algorithm is presented, which leverages common features to reduce the inference time and improve accuracy for multi-modal prediction. Experimental results demonstrate that our proposed framework achieves comparable state-of-the-art performance with real-time inference speed in public datasets.
Sheng Wang 0017, Lei Zhu 0003, Ming Liu 0001, Jun Ma 0008
IROS5
2025 AKF-LIO: LiDAR-Inertial Odometry with Gaussian Map by Adaptive Kalman Filter
abstract
Existing LiDAR-Inertial Odometry (LIO) systems typically use sensor-specific or environment-dependent measurement covariances during state estimation, leading to laborious parameter tuning and suboptimal performance in challenging conditions (e.g., sensor degeneracy and noisy observations). Therefore, we propose an Adaptive Kalman Filter (AKF) framework that dynamically estimates time-varying noise covariances of LiDAR and Inertial Measurement Unit (IMU) measurements, enabling context-aware confidence weighting between sensors. During LiDAR degeneracy, the system prioritizes IMU data while suppressing contributions from unreliable inputs like moving objects or noisy point clouds. Furthermore, a compact Gaussian-based map representation is introduced to model environmental planarity and spatial noise. A correlated registration strategy ensures accurate plane normal estimation via pseudo-merge, even in unstructured environments like forests. Extensive experiments validate the robustness of the proposed system across diverse environments, including dynamic scenes and geometrically degraded scenarios. Our method achieves reliable localization results across all MARS-LVIG sequences and ranks 8th on the KITTI Odometry Benchmark. The code will be released at https://github.com/xpxie/AKF-LIO.git.
Xupeng Xie, Ruoyu Geng, Jun Ma 0008, Boyu Zhou
IROS3
2025 DACA-Net: A Degradation-Aware Conditional Diffusion Network for Underwater Image Enhancement
abstract
Underwater images typically suffer from severe colour distortions, low visibility, and reduced structural clarity due to complex optical effects such as scattering and absorption, which greatly degrade their visual quality and limit the performance of downstream visual perception tasks. Existing enhancement methods often struggle to adaptively handle diverse degradation conditions and fail to leverage underwater-specific physical priors effectively. In this paper, we propose a degradation-aware conditional diffusion model to enhance underwater images adaptively and robustly. Given a degraded underwater image as input, we first predict its degradation level using a lightweight dual-stream convolutional network, generating a continuous degradation score as semantic guidance. Based on this score, we introduce a novel conditional diffusion-based restoration network with a Swin UNet backbone, enabling adaptive noise scheduling and hierarchical feature refinement. To incorporate underwater-specific physical priors, we further propose a degradation-guided adaptive feature fusion module and a hybrid loss function that combines perceptual consistency, histogram matching, and feature-level contrast. Comprehensive experiments on benchmark datasets demonstrate that our method effectively restores underwater images with superior colour fidelity, perceptual quality, and structural details. Compared with SOTA approaches, our framework achieves significant improvements in both quantitative metrics and qualitative visual assessments.
Chang Huang, Jiahang Cao, Jun Ma 0008, Kieren Yu, Cong Li 0005, Huayong Yang, Kaishun Wu
ACM Multimedia3
2025 HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
abstract
While Retrieval-Augmented Generation (RAG) augments Large Language Models (LLMs) with external knowledge, conventional single-agent RAG remains fundamentally limited in resolving complex queries demanding coordinated reasoning across heterogeneous data ecosystems. We present HM-RAG, a novel Hierarchical Multi-agent Multimodal RAG framework that pioneers collaborative intelligence for dynamic knowledge synthesis across structured, unstructured, and graph-based data. The framework is composed of a three-tiered architecture with specialized agents: a Decomposition Agent that dissects complex queries into contextually coherent sub-tasks via semantic-aware query rewriting and schema-guided context augmentation; Multi-source Retrieval Agents that carry out parallel, modality-specific retrieval using plug-and-play modules designed for vector, graph, and web-based databases; and a Decision Agent that uses consistency voting to integrate multi-source answers and resolve discrepancies in retrieval results through Expert Model Refinement. This architecture attains comprehensive query understanding by combining textual, graph-relational, and web-derived evidence, resulting in a remarkable 12.95% improvement in answer accuracy and a 3.56% boost in question classification accuracy over baseline RAG systems on the ScienceQA and CrisisMMD benchmarks. Notably, HM-RAG establishes state-of-the-art results in zero-shot settings on both datasets. Its modular architecture ensures seamless integration of new data modalities while maintaining strict data governance, marking a significant advancement in addressing the critical challenges of multimodal reasoning and knowledge synthesis in RAG systems.
Ruoyu Yao, Siyuan Meng, Ding Wang 0001, Jun Ma 0008
ACM Multimedia7
2025 Tractor Semi-Trailer Off-Tracking and Stability Approximate Bi-Level Policy Optimization
abstract
Trajectory tracking control of tractor semi-trailer vehicles poses significant challenges due to inherent off-tracking behavior and roll instability risks. While existing approaches have demonstrated effectiveness, they often rely on computationally intensive numerical solvers and require time-consuming manual tuning of cost function weights. This paper presents an approximate bi-level policy optimization (ABPO) framework that simultaneously optimizes the cost function and synthesizes an explicit control policy to minimize off-tracking while reducing computational complexity. The proposed framework employs a hierarchical structure: the upper level updates cost weights based on the trailer’s stability trajectory, while the lower level derives an approximate optimal policy by solving the tractor’s control problem. By leveraging Pontryagin’s Maximum Principle (PMP), we have developed a novel method to analytically compute cost weight gradients through differentiation of the PMP conditions. This enables the formulation of a related optimal control problem (OCP) whose solutions directly yield gradients for cost parameter updates. The ABPO framework achieves automatic weight coefficient adjustment, enhances trajectory tracking accuracy for both tractor and trailer units, and significantly reduces computational burden. Simulation and experimental validation across 4 classical scenarios demonstrates that the learned policy reduces rearward amplification by 17.82%, lateral tracking errors by 84.15%, and rollover by 64.19%, respectively. Notably, the control policy computation requires less than 10 ms, making it suitable for real-time applications. The source code for the algorithms described in this paper is publicly available at https://github.com/TroyResearch/ABPO.git.
Fawang Zhang, Jingliang Duan, Hui Liu 0001, Xingyu Cao, Shida Nie, Congshuai Guo, Yujia Xie, Jun Ma 0008, Shangli Wang
IEEE Trans Autom. Sci. Eng.8
2025 UDMC: Unified Decision-Making and Control Framework for Urban Autonomous Driving With Motion Prediction of Traffic Participants
abstract
Current autonomous driving systems often struggle to balance decision-making and motion control while ensuring safety and traffic rule compliance, especially in complex urban environments. Existing methods may fall short due to separate handling of these functionalities, leading to inefficiencies and safety compromises. To address these challenges, we introduce UDMC, an interpretable and unified Level 4 autonomous driving framework. UDMC integrates decision-making and motion control into a single optimal control problem (OCP), considering the dynamic interactions with surrounding vehicles, pedestrians, road lanes, and traffic signals. By employing innovative potential functions to model traffic participants and regulations, and incorporating a specialized motion prediction module, our framework enhances on-road safety and rule adherence. The integrated design allows for real-time execution of flexible maneuvers suited to diverse driving scenarios. High-fidelity simulations conducted in CARLA exemplify the framework’s computational efficiency, robustness, and safety, resulting in superior driving performance when compared against various baseline models. Our open-source project is available athttps://github.com/henryhcliu/udmc_carla.git.
Haichao Liu 0003, Kai Chen 0006, Yulin Li 0001, Zhenmin Huang, Ming Liu 0001, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.6
2025 Monocular 3D Lane Detection for Autonomous Driving: Recent Achievements, Challenges, and Outlooks
abstract
3D lane detection is essential in autonomous driving (AD) as it extracts structural and traffic information from the road in 3D space, aiding autonomous vehicles in logical, safe, and comfortable path planning and motion control. Given the cost of sensors and the advantages of visual data in color information, 3D lane detection based on monocular vision is an important research direction in the realm of AD that increasingly gains attention in both industry and academia. Nevertheless, recent advancements in visual perception seem inadequate for the development of fully reliable 3D lane detection algorithms, which also hampers the progress of vision-based fully autonomous vehicles. We believe that it still leaves an open and interesting problem for improvement in 3D lane detection algorithms for autonomous vehicles using visual sensors, and significant enhancements are essentially required. This review summarizes and analyzes the current state of achievements in the field of 3D lane detection research. It covers all current monocular-based 3D lane detection processes, discusses the performance of these cutting-edge algorithms, analyzes the time complexity of various algorithms, and highlights the main achievements and limitations of ongoing research efforts. The survey also includes a comprehensive discussion of available 3D lane detection datasets and the challenges that researchers encounter but have not yet resolved. Finally, our work outlines future research directions and invites researchers and practitioners to join this exciting field.
Fulong Ma, Weiqing Qi, Guoyang Zhao, Linwei Zheng, Sheng Wang 0017, Yuxuan Liu 0008, Ming Liu 0001, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.8
2025 CurbNet: Curb Detection Framework Based on LiDAR Point Cloud Segmentation
abstract
Curb detection is a crucial function in intelligent driving, essential for determining drivable areas on the road. However, the complexity of road environments makes curb detection challenging. This paper introduces CurbNet, a novel framework for curb detection utilizing point cloud segmentation. To address the lack of comprehensive curb datasets with 3D annotations, we have developed the 3D-Curb dataset based on SemanticKITTI, currently the largest and most diverse collection of curb point clouds. Recognizing that the primary characteristic of curbs is height variation, our approach leverages spatially rich 3D point clouds for training. To tackle the challenges posed by the uneven distribution of curb features on the xy-plane and their dependence on high-frequency features along the z-axis, we introduce the Multi-Scale and Channel Attention (MSCA) module, a customized solution designed to optimize detection performance. Additionally, we propose an adaptive weighted loss function group specifically formulated to counteract the imbalance in the distribution of curb point clouds relative to other categories. Extensive experiments conducted on 2 major datasets demonstrate that our method surpasses existing benchmarks set by leading curb detection and point cloud segmentation models. Through the post-processing refinement of the detection results, we have significantly reduced noise in curb detection, thereby improving precision by 4.5 points. Similarly, our tolerance experiments also achieve state-of-the-art results. Furthermore, real-world experiments and dataset analyses mutually validate each other, reinforcing CurbNet’s superior detection capability and robust generalizability. The project website is available at:https://github.com/guoyangzhao/CurbNet/.
Guoyang Zhao, Fulong Ma, Weiqing Qi, Yuxuan Liu 0008, Ming Liu 0001, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.6
2025 Barrier-Enhanced Parallel Homotopic Trajectory Optimization for Safety-Critical Autonomous Driving
abstract
Enforcing safety while preventing overly conservative behaviors is essential for autonomous vehicles to achieve high task performance. In this paper, we propose a barrier-enhanced parallel homotopic trajectory optimization (BPHTO) approach with the over-relaxed alternating direction method of multipliers (ADMM) for real-time integrated decision-making and planning. To facilitate safety interactions between the ego vehicle (EV) and surrounding vehicles, a spatiotemporal safety module exhibiting bi-convexity is developed on the basis of barrier function. Varying barrier coefficients are adopted for different time steps in a planning horizon to account for the motion uncertainties of surrounding HVs and mitigate conservative behaviors. Additionally, we exploit the discrete characteristics of driving maneuvers to initialize nominal behavior-oriented free-end homotopic trajectories based on reachability analysis, and each trajectory is locally constrained to a specific driving maneuver while sharing the same task objectives. By leveraging the bi-convexity of the safety module and the kinematics of the EV, we formulate the BPHTO as a bi-convex optimization problem. Then constraint transcription and the over-relaxed ADMM are employed to streamline the optimization process, such that multiple trajectories are generated in real time with feasibility guarantees. Through a series of experiments, the proposed development demonstrates improved task accuracy, stability, and consistency in various traffic scenarios using synthetic and real-world traffic datasets.
Lei Zheng 0007, Rui Yang 0025, Michael Yu Wang, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.4
2025 Diffeomorphism-Transformed Iterative Linear Quadratic Regulator for Constrained Motion Planning in Autonomous Driving
abstract
Ensuring safe driving and real-time execution is a crucial requirement in the motion planning process for autonomous vehicles. Hence, there is a compelling demand for advanced motion planning algorithms that exhibit effective management of inequality constraints and exceptional computational performance. This paper investigates a diffeomorphism-transformed iterative linear quadratic regulator (DTiLQR) algorithm for addressing constrained motion planning problems in autonomous vehicles with nonlinear dynamics and multiple inequality constraints. With regard to the state and input constraints, a novel state-and-input diffeomorphism is proposed to transform the constrained state/input space into an unconstrained one. Subsequently, these inequality constraints are systematically incorporated into the vehicle dynamics, thereby leading to the newly constructed system in this context. Then, we reformulate and incorporate the obstacle avoidance constraint into the objective function using state diffeomorphism and logarithmic barrier function. With this, the original optimization problem is converted to the unconstrained counterpart, adhering only to the constructed system dynamics. In this sense, featuring a streamlined single-loop architecture (which is essentially different from the dual-loop algorithmic design of existing constrained iLQR algorithms), DTiLQR is used to solve the optimization problem effectively while maintaining motion performance and constraint satisfaction for the resulting optimal trajectory. Ultimately, case studies across various driving situations showcase the effectiveness and exceptional computational efficiency of the proposed DTiLQR algorithm.
Zicheng Zhu, Haichao Liu 0003, Jingliang Duan, Han Zhao 0007, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.6
2024 Parallel Optimization with Hard Safety Constraints for Cooperative Planning of Connected Autonomous Vehicles
abstract
The development of connected autonomous vehicles (CAVs) facilitates the enhancement of traffic efficiency in complicated scenarios. Difficulties remain unsolved in developing an effective and efficient coordination strategy for CAVs. In this paper, we formulate the cooperative autonomous driving task of CAVs as an optimal control problem with safety conditions enforced as hard constraints, and propose a computationally-efficient parallel optimization framework to generate strategies for CAVs with the travel efficiency improved and the hard safety constraints satisfied. Specifically, all constraints involved are addressed appropriately with convex approximation, such that the convexity property of the reformulated optimization problem is exhibited. Then, a parallel optimization algorithm is presented to solve the reformulated optimization problem, with an embodied iterative nearest neighbor search strategy to determine the optimal passing sequence. It is noteworthy that the travel efficiency is enhanced and the computation burden is considerably alleviated with the proposed innovation development. We also examine the proposed method in CARLA simulator and perform thorough comparisons to demonstrate the effectiveness and efficiency of the proposed approach.
Zhenmin Huang, Haichao Liu 0003, Shaojie Shen, Jun Ma 0008
ICRA4
2024 Chance-Aware Lane Change with High-Level Model Predictive Control Through Curriculum Reinforcement Learning
abstract
Lane change in dense traffic typically requires the recognition of an appropriate opportunity for maneuvers, which remains a challenging problem in self-driving. In this work, we propose a chance-aware lane-change strategy with high-level model predictive control (MPC) through curriculum reinforcement learning (CRL). In our proposed framework, full-state references and regulatory factors concerning the relative importance of each cost term in the embodied MPC are generated by a neural policy. Furthermore, effective curricula are designed and integrated into an episodic reinforcement learning (RL) framework with policy transfer and enhancement, to improve the convergence speed and ensure a high-quality policy. The proposed framework is deployed and evaluated in numerical simulations of dense and dynamic traffic. It is noteworthy that, given a narrow chance, the proposed approach generates high-quality lane-change maneuvers such that the vehicle merges into the traffic flow with a high success rate of 96%. Finally, our framework is validated in the high-fidelity simulator under dense traffic, demonstrating satisfactory practicality and generalizability.
Yulin Li 0001, Zengqi Peng, Hakim Ghazzai, Jun Ma 0008
ICRA5
2024 Analysis and Design for Inductively Coupled Plasma Power Source with Multi-level Power Mode
abstract
The inductively coupled plasma (ICP) source is widely used in the semiconductor manufacturing industry. The main challenge is delivering a high frequency and low harmonic current with multi-level power modulation. The multi-level power mode is required for advanced manufacturing processes and it modulates the output power repeatedly, the energy storage components charge and discharge rapidly, leading to very large transient power losses on switches, especially for the lagging bridge which loses ZVS operation. This paper discussed the matching network design for the ICP source and the steady-state analysis for its phase-shift modulation. Furthermore, the transient power loss mechanism during the multi-level power modulation is elaborated by using multiple harmonic approximation. Finally, the design guidance for ICP source working at multi-level power mode is discussed to achieve high efficiency and high reliability. A 2kW prototyping system is built to validate the proposed analysis and design for the ICP source.
Chenyue Chen, Jun Ma 0008, Ming Liu 0001
IECON2
2024 Every Dataset Counts: Scaling up Monocular 3D Object Detection with Joint Datasets Training
abstract
Monocular 3D object detection is essential for autonomous driving. However, current monocular 3D detection algorithms rely on expensive 3D labels from LiDAR scans, making it difficult to use in new datasets and unfamiliar environments. This study explores training a monocular 3D object detection model using a mix of 3D and 2D datasets. The proposed framework includes a robust monocular 3D model that can adapt to different camera settings, a selective-training strategy to handle varying class annotations in datasets, and a pseudo 3D training method using 2D labels to improve detection ability in scenes with only 2D labels (as shown in Fig. 1). By utilizing this framework, we can train models on a combination of 3D and 2D datasets to improve generalization and performance on new datasets with only 2D labels. Extensive experiments on KITTI, nuScenes, ONCE, Cityscapes, and BDD100K datasets showcase the scalability of our proposed approach. Here is our project page: https://sites.google.com/view/fmaafmono3d.
Fulong Ma, Xiaoyang Yan, Guoyang Zhao, Yuxuan Liu 0008, Jun Ma 0008, Ming Liu 0001
IROS6
2024 Reward-Driven Automated Curriculum Learning for Interaction-Aware Self-Driving at Unsignalized Intersections
abstract
In this work, we present a reward-driven automated curriculum reinforcement learning approach for interaction-aware self-driving at unsignalized intersections, taking into account the uncertainties associated with surrounding vehicles (SVs). These uncertainties encompass the uncertainty of SVs’ driving intention and also the quantity of SVs. To deal with this problem, the curriculum set is specifically designed to accommodate a progressively increasing number of SVs. By implementing an automated curriculum selection mechanism, the importance weights are rationally allocated across various curricula, thereby facilitating improved sample efficiency and training outcomes. Furthermore, the reward function is meticulously designed to guide the agent towards effective policy exploration. Thus the proposed framework could proactively address the above uncertainties at unsignalized intersections by employing the automated curriculum learning technique that progressively increases task difficulty, and this ensures safe self-driving through effective interaction with SVs. Comparative experiments are conducted in Highway_Env, and the results indicate that our approach achieves the highest task success rate, attains strong robustness to initialization parameters of the curriculum selection module, and exhibits superior adaptability to diverse situational configurations at unsignalized intersections. Furthermore, the effectiveness of the proposed method is validated using the high-fidelity CARLA simulator.
Zengqi Peng, Lei Zheng 0007, Jun Ma 0008
IROS5
2024 Arm-Constrained Curriculum Learning for Loco-Manipulation of a Wheel-Legged Robot
abstract
Incorporating a robotic manipulator into a wheellegged robot enhances its agility and expands its potential for practical applications. However, the presence of potential instability and uncertainties presents additional challenges for control objectives. In this paper, we introduce an arm-constrained curriculum learning architecture to tackle the issues introduced by adding the manipulator. Firstly, we develop an arm-constrained reinforcement learning algorithm to ensure safety and reliability in control performance after equipping the manipulator. Additionally, to address discrepancies in reward settings between the arm and the base, we propose a reward-aware curriculum learning method. The policy is first trained in Isaac gym and transferred to the physical robot to complete grasping tasks, including the door-opening task, fan-twitching task and the relay-baton-picking and following task. The results demonstrate that our proposed approach effectively controls the arm-equipped wheel-legged robot to master grasping abilities including the dynamic grasping skills, allowing it to chase and catch a moving object while in motion. Please refer to our website (https://acodedog.github.io/wheel-legged-loco-manipulation/) for the code and supplemental videos.
Yufei Jia, Haizhou Zhao, Jinni Zhou, Jun Ma 0008, Guyue Zhou
IROS8
2024 MCGMapper: Light-Weight Incremental Structure from Motion and Visual Localization with Planar Markers and Camera Groups
abstract
Structure from Motion (SfM) and visual localization in indoor texture-less scenes and industrial scenarios present prevalent yet challenging research topics. Existing SfM methods designed for natural scenes typically yield low accuracy or map-building failures due to insufficient robust feature extraction in such settings. Visual markers, with their artificially designed features, can effectively address these issues. Nonetheless, existing marker-assisted SfM methods encounter problems like slow running speed and difficulties in convergence; and also, they are governed by the strong assumption of unique marker size. In this paper, we propose a novel SfM framework that utilizes planar markers and multiple cameras with known extrinsics to capture the surrounding environment and reconstruct the marker map. In our algorithm, the initial poses of markers and cameras are calculated with Perspective-n-Points (PnP) in the front-end, while bundle adjustment methods customized for markers and camera groups are designed in the back-end to optimize the 6-DOF pose directly. Our algorithm facilitates the reconstruction of large scenes with different marker sizes, and its accuracy and speed of map building are shown to surpass existing methods. Our approach is suitable for a wide range of scenarios, including laboratories, basements, warehouses, and other industrial settings. Furthermore, we incorporate representative scenarios into simulations and also supply our datasets with pose labels to address the scarcity of quantitative ground-truth datasets in this research field. The datasets and source code are available on GitHub1.
Yusen Xie, Zhenmin Huang, Kai Chen 0006, Lei Zhu 0003, Jun Ma 0008
IROS5
2024 A Generic Trajectory Planning Method for Constrained All-Wheel-Steering Robots
abstract
This paper presents a generic trajectory planning method for wheeled robots with fixed steering axes while the steering angle of each wheel is constrained. In the existing literatures, All-Wheel-Steering (AWS) robots, incorporating modes such as rotation-free translation maneuvers, in-situ rotational maneuvers, and proportional steering, exhibit inefficient performance due to time-consuming mode switches. This inefficiency arises from wheel rotation constraints and inter-wheel cooperation requirements. The direct application of a holonomic moving strategy can lead to significant slip angles or even structural failure. Additionally, the limited steering range of AWS wheeled robots exacerbates non-linearity characteristics, thereby complicating control processes. To address these challenges, we developed a novel planning method termed Constrained AWS (C-AWS), which integrates second-order discrete search with predictive control techniques. Experimental results demonstrate that our method adeptly generates feasible and smooth trajectories for C-AWS while adhering to steering angle constraints. Code and video can be found at https://github.com/Rex-sys-hk/AWSPlanning.
Ren Xin, Hongji Liu, Yingbing Chen, Jie Cheng 0008, Sheng Wang 0017, Jun Ma 0008, Ming Liu 0001
IROS6
2024 Incremental Learning-Based Real-Time Trajectory Prediction for Autonomous Driving via Sparse Gaussian Process Regression
abstract
In the context of spatial-temporal autonomous driving, the accurate and real-time trajectory prediction of the surrounding vehicle (SV) is crucial. This paper aims to design an efficient, accurate, and interpretable unimodal trajectory prediction approach. To achieve this objective, we employ Sparse Gaussian Process Regression (SGPR), which enables large dataset learning and efficient inference of future trajectories. This approach ensures accurate predictions while maintaining high computational efficiency. To further enhance the robustness of the prediction module, we propose the translation and rotation transformation strategy, which effectively simplifies the prediction problem. Additionally, we utilize an instant evaluation algorithm to assess the prediction performance and maintain a streaming dataset for incremental learning, capable of adapting to dynamic driving environments. In our experimental evaluation, we compare our proposed trajectory prediction approach with a series of existing methods. The results demonstrate that our work achieves superior prediction accuracy while requiring less inference time. It is noteworthy that, the proposed SGPR-based trajectory prediction approach with rotation equivalence is able to swiftly infer and incrementally learn from dynamic environments, which makes it a promising tool for enhancing safety and efficiency in autonomous driving systems.
Haichao Liu 0003, Kai Chen 0006, Jun Ma 0008
IV3
2024 Timeline and Boundary Guided Diffusion Network for Video Shadow Detection
abstract
Video Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at \url{https://github.com/haipengzhou856/TBGDiff}.
Haipeng Zhou, Hongqiu Wang, Tian Ye 0001, Zhaohu Xing, Jun Ma 0008, Ping Li 0016, Qiong Wang 0001, Lei Zhu 0003
ACM Multimedia5
2024 Adaptive robust control for fuzzy underactuated mechanical systems: A Stackelberg game-theoretic optimization approach
Yuanjie Xian, Jun Ma 0008, Abdullah Al Mamun 0002, Tong Heng Lee
Inf. Sci.2
2024 Data-Driven Linear Quadratic Optimization for Controller Synthesis With Structural Constraints
abstract
For various typical cases and situations where the formulation results in an optimal control problem, the linear quadratic regulator (LQR) approach and its variants continue to be highly attractive. In certain scenarios, it can happen that some prescribed structural constraints on the gain matrix would arise. Consequently then, the algebraic Riccati equation (ARE) is no longer applicable in a straightforward way to obtain the optimal solution. This work presents a rather effective alternative optimization approach based on gradient projection. The utilized gradient is obtained through a data-driven methodology, and then projected onto applicable constrained hyperplanes. Essentially, this projection gradient determines a direction of progression and computation for the gain matrix update with a decreasing functional cost; and then the gain matrix is further refined in an iterative framework. With this formulation, a data-driven optimization algorithm is summarized for controller synthesis with structural constraints. This data-driven approach has the key advantage that it avoids the necessity of precise modeling which is always required in the classical model-based counterpart; and thus the approach can additionally accommodate various model uncertainties. Illustrative examples are also provided in the work to validate the theoretical results.
Jun Ma 0008, Zilong Cheng, Xiaocong Li, Masayoshi Tomizuka, Tong Heng Lee
IEEE Trans. Cybern.1
2024 Game-Theoretic Optimization Toward Diffeomorphism-Based Robust Control of Fuzzy Dynamical Systems With State and Input Constraints
abstract
This work investigates a game-theoretic optimization approach towards robust control of uncertain dynamical systems with state and input constraints. The uncertainty involved is possibly rapidly time-varying but bounded within a prescribed fuzzy set. For this, the associated fuzzy dynamical system is appropriately established and constructed based on fuzzy set theory. To cope with the bounded state and input constraints, a novel state-and-input diffeomorphism technique is proposed, where a transformed system is formulated such that the prescribed inequality constraints are innovatively merged into the stabilization and trajectory tracking problems. Furthermore, a diffeomorphism-based robust control (DBRC) strategy is developed to ensure the uniform boundedness (UB) and uniform ultimate boundedness (UUB) of the transformed system. Under this proposed control architecture, the constraint satisfaction of the original system is thus always analytically ensured based on the rigorous properties of the diffeomorphism technique. The resulting control parameter optimization problem then has to take into account the multiple considerations (and compromise) amongst the factors of the steady-state performance; the finite convergence time; and the control effort. For this, a two-player Nash game is formulated and solved in an effective manner. The Nash equilibrium is obtained and the existence of the solution is also proved theoretically. With this methodology, and with the resulting attainment of the desired Nash equilibrium, the attendant outcome of superior system performance is achieved. Finally, numerical simulations on a steer-by-wire (SBW) system demonstrate the effectiveness of the proposed approach.
Zicheng Zhu, Jun Ma 0008, Hao Sun 0008, Han Zhao 0007, Tong Heng Lee
IEEE Trans. Fuzzy Syst.2
2024 Stackelberg Game-Based Control Design for Fuzzy Underactuated Mechanical Systems With Inequality Constraints
abstract
A Stackelberg game-based design for an adaptive robust control for the fuzzy uncertain underactuated mechanical systems (UMSs) is proposed. The emphasis is on fuzzy-based uncertainty and inequality constraint. The uncertainty is time varying and bounded within a prescribed fuzzy set. For the inequality constraint, we creatively have it merge into constraint-following performance by a diffeomorphism technique. An adaptive robust control strategy is then proposed. Deterministic performance is guaranteed provided the control design parameters are within feasible regions. To further enhance the performance, we introduce a two-player Stackelberg game setting. The optimal choice of design parameters can be solved. The feasibility of this design is demonstrated on an autonomous wheeled mobile robot (AWMR), which is confined in a bounded space.
Zicheng Zhu, Han Zhao 0007, Yuanjie Xian, Ye-Hwa Chen, Hao Sun 0008, Jun Ma 0008
IEEE Trans. Syst. Man Cybern. Syst.6
2023 Incremental few-shot learning via implanting and consolidating
Haiyue Zhu, Jun Ma 0008, Cheng Xiang 0001, Prahlad Vadakkepat
Neurocomputing3
2023 Decentralized iLQR for Cooperative Trajectory Planning of Connected Autonomous Vehicles via Dual Consensus ADMM
abstract
Cooperative trajectory planning of connected autonomous vehicles (CAVs) generally admits strong nonlinearity and non-convexity, rendering great difficulties in finding the optimal solution. Existing methods typically suffer from low computational efficiency and poor scalability, which hinder the appropriate applications in large-scale scenarios involving an increasing number of vehicles. To tackle this problem, we propose a novel decentralized iterative linear quadratic regulator (iLQR) algorithm by leveraging the dual consensus alternating direction method of multipliers (ADMM). First, the original non-convex optimization problem is reformulated into a series of convex optimization problems through iterative neighbourhood approximation. Then, the dual of each convex optimization problem is shown to have a consensus structure, which facilitates the use of consensus ADMM to solve for the dual solution in a fully decentralized and parallel architecture. Finally, the primal solution corresponding to the trajectory of each vehicle is recovered by solving a linear quadratic regulator (LQR) problem iteratively, and a novel trajectory update strategy is proposed to ensure the dynamic feasibility of vehicles. With the proposed development, the computation burden is significantly alleviated such that real-time performance is attainable. Two traffic scenarios are presented to validate the proposed algorithm, and thorough comparisons between our proposed method and baseline methods (including centralized iLQR, IPOPT, and SQP) are conducted to demonstrate the scalability of the proposed approach.
Zhenmin Huang, Shaojie Shen, Jun Ma 0008
IEEE Trans. Intell. Transp. Syst.3
2023 Policy Iteration Based Approximate Dynamic Programming Toward Autonomous Driving in Constrained Dynamic Environment
abstract
In the area of autonomous driving, it typically brings great difficulty in solving the motion planning problem since the vehicle model is nonlinear and the driving scenarios are complex. Particularly, most of the existing methods cannot be generalized to dynamically changing scenarios with varying surrounding vehicles. To address this problem, this development here investigates the framework of integrated decision and control. As part of the modules, static path planning determines the reference candidates ahead, and then the optimal path-tracking controller realizes the specific autonomous driving task. An innovative and effective constrained finite-horizon approximate dynamic programming (ADP) algorithm is herein presented to generate the desired control policy for effective path tracking. With the generalized policy neural network that maps from the state to the control input, the proposed algorithm preserves the high effectiveness for the motion planning problem towards changing driving environments with varying surrounding vehicles. Moreover, the algorithm attains the noteworthy advantage of alleviating the typically heavy computational loads with the mode of offline training and online execution. As a result of the utilization of multi-layer neural networks in conjunction with the actor-critic framework, the constrained ADP method is capable of handling complex and multidimensional scenarios. Finally, various simulations have been carried out to show that the constrained ADP algorithm is effective.
Ziyu Lin, Jun Ma 0008, Jingliang Duan, Shengbo Eben Li, Haitong Ma, Bo Cheng 0003, Tong Heng Lee
IEEE Trans. Intell. Transp. Syst.2
2023 Policy-Iteration-Based Finite-Horizon Approximate Dynamic Programming for Continuous-Time Nonlinear Optimal Control
abstract
The Hamilton-Jacobi-Bellman (HJB) equation serves as the necessary and sufficient condition for the optimal solution to the continuous-time (CT) optimal control problem (OCP). Compared with the infinite-horizon HJB equation, the solving of the finite-horizon (FH) HJB equation has been a long-standing challenge, because the partial time derivative of the value function is involved as an additional unknown term. To address this problem, this study first-time bridges the link between the partial time derivative and the terminal-time utility function, and thus it facilitates the use of the policy iteration (PI) technique to solve the CT FH OCPs. Based on this key finding, the FH approximate dynamic programming (ADP) algorithm is proposed leveraging an actor-critic framework. It is shown that the algorithm exhibits important properties in terms of convergence and optimality. Rather importantly, with the use of multilayer neural networks (NNs) in the actor-critic architecture, the algorithm is suitable for CT FH OCPs toward more general nonlinear and complex systems. Finally, the effectiveness of the proposed algorithm is demonstrated by conducting a series of simulations on both a linear quadratic regulator (LQR) problem and a nonlinear vehicle tracking problem.
Ziyu Lin, Jingliang Duan, Shengbo Eben Li, Haitong Ma, Jie Li 0042, Jianyu Chen 0002, Bo Cheng 0003, Jun Ma 0008
IEEE Trans. Neural Networks Learn. Syst.8
2023 Local Learning Enabled Iterative Linear Quadratic Regulator for Constrained Trajectory Planning
abstract
Trajectory planning is one of the indispensable and critical components in robotics and autonomous systems. As an efficient indirect method to deal with the nonlinear system dynamics in trajectory planning tasks over the unconstrained state and control space, the iterative linear quadratic regulator (iLQR) has demonstrated noteworthy outcomes. In this article, a local-learning-enabled constrained iLQR algorithm is herein presented for trajectory planning based on hybrid dynamic optimization and machine learning. Rather importantly, this algorithm attains the key advantage of circumventing the requirement of system identification, and the trajectory planning task is achieved with a simultaneous refinement of the optimal policy and the neural network system in an iterative framework. The neural network can be designed to represent the local system model with a simple architecture, and thus it leads to a sample-efficient training pipeline. In addition, in this learning paradigm, the constraints of the general form that are typically encountered in trajectory planning tasks are preserved. Several illustrative examples on trajectory planning are scheduled as part of the test itinerary to demonstrate the effectiveness and significance of this work.
Jun Ma 0008, Zilong Cheng, Ziyu Lin, Frank L. Lewis, Tong Heng Lee
IEEE Trans. Neural Networks Learn. Syst.1
2023 On Symmetric Gauss-Seidel ADMM Algorithm for H∞ Guaranteed Cost Control With Convex Parameterization
abstract
This article involves the innovative development of a symmetric Gauss–Seidel ADMM algorithm to solve the$\mathcal {H}_{\infty }$guaranteed cost control problem. In the presence of parametric uncertainties, the$\mathcal {H}_{\infty }$guaranteed cost control problem generally leads to the large-scale optimization. This is due to the exponential growth of the number of the extreme systems involved with respect to the number of parametric uncertainties. In this work, through a variant of the Youla–Kucera parameterization, the stabilizing controllers are parameterized in a convex set; yielding the outcome that the$\mathcal {H}_{\infty }$guaranteed cost control problem is converted to a convex optimization problem. Based on an appropriate reformulation using the Schur complement, it then renders possible the use of the ADMM algorithm with symmetric Gauss–Seidel backward and forward sweeps. Significantly, this approach alleviates the often-times prohibitively heavy computational burden typical in many$\mathcal H_{\infty }$optimization problems while exhibiting good convergence guarantees, which is particularly essential for the related large-scale optimization procedures involved. With this approach, the desired robust stability is ensured, and the disturbance attenuation is maintained at the minimum level in the presence of parametric uncertainties. Rather importantly too, with the attained effectiveness, the methodology thus evidently possesses extensive applicability in various important controller synthesis problems, such as decentralized control, sparse control, and output feedback control problems.
Jun Ma 0008, Zilong Cheng, Masayoshi Tomizuka, Tong Heng Lee
IEEE Trans. Syst. Man Cybern. Syst.1
2023 Robust Fixed-Order Controller Design for Uncertain Systems With Generalized Common Lyapunov Strictly Positive Realness Characterization
abstract
This article investigates the design of a robust fixed-order controller for single-input–single-output (SISO) polytopic systems with interval uncertainties, with the aim that the closed-loop stability is appropriately ensured and the performance specifications on sensitivity shaping are conformed in a specific finite frequency range. Utilizing the notion of generalized common Lyapunov strictly positive realness (CL-SPRness), the equivalence between strictly positive realness (SPRness) and strictly bounded realness (SBRness) is established; and then, the specifications on robust stability and performance are transformed into the SPRness of newly constructed systems and further characterized in the framework of linear matrix inequality (LMI) conditions. The proposed methodology avoids the tedious yet mandatory evaluations of the specifications on all vertices of the uncertain polytopic system in an explicit form. Instead, solving five LMIs exclusively suffices for ensuring the robust stability and performance regardless of the number of vertices, and thus, the typically heavy computational burden is considerably alleviated. It is also noteworthy that the proposed methodology additionally provides the necessary and sufficient conditions for this robust controller design with the consideration of a prescribed finite frequency range, and therefore, significantly less conservatism is attained in the system performance.
Jun Ma 0008, Haiyue Zhu, Xiaocong Li, Clarence W. de Silva, Tong Heng Lee
IEEE Trans. Syst. Man Cybern. Syst.1
2022 Incremental Few-Shot Object Detection for Robotics
abstract
Incremental few-shot learning is highly expected for practical robotics applications. On one hand, robot is desired to learn new tasks quickly and flexibly using only few annotated training samples; on the other hand, such new additional tasks should be learned in a continuous and incremental manner without forgetting the previous learned knowledge dramatically. In this work, we propose a novel Class-Incremental Few- Shot Object Detection (CI-FSOD) framework that enables deep object detection network to perform effective continual learning from just few-shot samples without re-accessing the previous training data. We achieve this by equipping the widely-used Faster-RCNN detector with three elegant components. Firstly, to best preserve performance on the pre-trained base classes, we propose a novel Dual-Embedding-Space (DES) architecture which decouples the representation learning of base and novel categories into different spaces. Secondly, to mitigate the catastrophic forgetting on the accumulated novel classes, we propose a Sequential Model Fusion (SMF) method, which is able to achieve long-term memory without additional storage cost. Thirdly, to promote inter-task class separation in feature space, we propose a novel regularization technique that extends the classification boundary further away from the previous classes to avoid misclassification. Overall, our framework is simple yet effective and outperforms the previous SOTA with a significant margin of 2.4 points in AP performance.
Haiyue Zhu, Sichao Tian, Jun Ma 0008, Chek Sing Teo, Cheng Xiang 0001, Prahlad Vadakkepat, Tong Heng Lee
ICRA5
2022 Weight Imprinting Classification-Based Force Grasping With a Variable-Stiffness Robotic Gripper
abstract
Universal grasping for a diverse range of objects is a challenging problem in robotics, especially in the presence of mixed properties with fragile/rigid and heavy/light. Toward universal grasping, this article presents a practical and systematic grasping control framework that enables a variable stiffness gripper to handle the objects with diverse properties using a category-aware force regulation approach, termed classification-based force grasping. Under this framework, a convolutional neural network (CNN) is employed to classify the category of the grasping object, and a grasping force is determined based on the classified category through a database that records a predefined force magnitude per category. Sequentially, the gripper can be adjusted to a force-optimized stiffness, which facilitates the achievement of an accurate grasping force regulation in a large range. Technically, two novel enabling modules are developed for grasping classification and execution, respectively. First, a novel weight imprinting technique based on center-guided feature embedding is proposed for object classification. It enables the CNN to efficiently handle novel object categories using only a few samples even without retraining/fine-tuning. Second, a vision-based grasping force sensing module is developed, which takes advantage of the specifically designed variable-stiffness gripper. Its grasping force can be estimated from the deflection angle of finger flexure by the vision so that the contact force can be sensed and regulated. Remarkably, only single-source vision information is needed for both of the above modules without any additional force sensor. Experiments are conducted extensively to evaluate the performance of the proposed force grasping approach.Note to Practitioners—Robotic grasping often needs to handle novel categories of objects. As a result, frequent retraining of the classification neural network is a pain point, which is tedious and prone to overfitting with only a few samples. In this work, metric learning is introduced for grasping classification where a novel kind of weight imprinting classification is proposed to handle the novel classes by better feature embedding and directly setting the classifier weights without retraining or fine-tuning. Together with the benefits from the variable stiffness feature of the gripper, the proposed vision-based force grasping approach can handle a wide range of objects from fragile to heavy, and the grasping force is controllable from 0.2 N onward to the motor limitation. The controllable grasping force resolution of the proposed vision grasping is better than 0.05 N, the accuracy of the grasping force is evaluated from 0.2 to 12 N, and the evaluated grasping objects are from extremely fragile potato chips and eggshell to heavy flange and metal block.
Haiyue Zhu, Xiong Li 0001, Xiaocong Li, Jun Ma 0008, Chek Sing Teo, Tat Joo Teo, Wei Lin 0002
IEEE Trans Autom. Sci. Eng.5
2022 Fuzzy-Based Controller Synthesis and Optimization for Underactuated Mechanical Systems With Nonholonomic Servo Constraints
abstract
This article investigates the trajectory tracking problem of underactuated mechanical systems (UMSs) with companion nonholonomic servo constraints and uncertainties. For such motion tasks, the existing approaches in the literature attempt unrealistically to furnish a reliable closed-form solution, rendering it difficult to have high-quality tracking performance with theoretical support. In addition, the uncertainties typically pose substantial difficulty in the controller synthesis. Here, by invoking the methodology of fuzzy sets, the uncertainties in the UMSs are elegantly represented; and with this, the formulation becomes such that a closer link between the uncertain dynamical model of the UMSs and the real world is established. The reference trajectories are regarded appropriately as servo constraints, and subsequently an adaptive robust controller is designed to accomplish the trajectory tracking task from a specific viewpoint of servo constraint tracking. As supported by rigorous proofs, the closed-form solution to the proposed controller is obtained with guaranteed Lyapunov stability. Leveraging on the closed-form solution, the global optimizer to the controller gain parameter can be determined, which is shown to exhibit several important properties including existence and uniqueness. Finally, a numerical example is presented to demonstrate the effectiveness of the designed method.
Jun Ma 0008, Hao Sun 0008, Shengchao Zhen, Han Zhao 0007, Abdullah Al Mamun 0002, Tong Heng Lee
IEEE Trans. Fuzzy Syst.2
2022 Alternating Direction Method of Multipliers for Constrained Iterative LQR in Autonomous Driving
abstract
In the context of autonomous driving, the iterative linear quadratic regulator (iLQR) is known to be an efficient approach to deal with the nonlinear vehicle model in motion planning problems. Particularly, the constrained iLQR algorithm has shown noteworthy advantageous outcomes of computation efficiency in achieving motion planning tasks under general constraints of different types. However, the constrained iLQR methodology requires a feasible trajectory at the first iteration as a prerequisite when the logarithmic barrier function is used. Also, the methodology leaves open the possibility for incorporation of fast, efficient, and effective optimization methods (i.e., fast-solvers) to further speed up the optimization process such that the requirements of real-time implementation can be successfully fulfilled. In this paper, a well-defined and commonly-encountered motion planning problem is formulated under nonlinear vehicle dynamics and various constraints, and the alternating direction method of multipliers (ADMM) is utilized to determine the optimal control actions leveraging the iLQR. With this development, the approach is able to circumvent the feasibility requirement of the trajectory at the first iteration. An illustrative example of motion planning for autonomous vehicles is then investigated with different driving scenarios taken into consideration, and a noteworthy achievement of high computation efficiency is attained with the proposed development. Comparing with the constrained iLQR algorithm based on the logarithmic barrier function, our proposed method reduces the average computation time by 31.93%, 38.52%, and 44.57% in the three scenarios; compared with the optimization solver IPOPT, our proposed method reduces the average computation time by 46.02%, 53.26%, and 88.43% in the three scenarios. As a result, real-time computation and implementation can be realized through our proposed framework, and thus it provides additional safety to the on-road driving tasks.
Jun Ma 0008, Zilong Cheng, Masayoshi Tomizuka, Tong Heng Lee
IEEE Trans. Intell. Transp. Syst.1
2022 Semi-Definite Relaxation-Based ADMM for Cooperative Planning and Control of Connected Autonomous Vehicles
abstract
This paper investigates the cooperative planning and control problem for multiple connected autonomous vehicles (CAVs) in different scenarios. In the existing literature, most of the methods suffer from significant problems in computational efficiency. Furthermore, as the optimization problem is nonlinear and nonconvex, it typically poses great difficulty in determining the optimal solution. To address this issue, this work proposes a novel and completely parallel computation framework by leveraging the alternating direction method of multipliers (ADMM). The nonlinear and nonconvex optimization problem in the autonomous driving problem can be divided into two manageable sub-problems; and the resulting sub-problems can be solved by using effective optimization methods in a parallel framework. Here, the differential dynamic programming (DDP) algorithm is capable of addressing the nonlinearity of the system dynamics rather effectively; and the nonconvex coupling constraints with small dimensions can be resolved by invoking the notion of semi-definite relaxation (SDR), which can also be solved in a very short time. Due to the parallel computation and efficient relaxation of nonconvex constraints, our proposed approach effectively realizes real-time implementation; and thus extra assurance of driving safety is provided. In addition, two transportation scenarios for multiple CAVs are used to illustrate the effectiveness and efficiency of the proposed method.
Zilong Cheng, Jun Ma 0008, Sunan Huang 0001, Frank L. Lewis, Tong Heng Lee
IEEE Trans. Intell. Transp. Syst.3
2022 On Robust Stability and Performance With a Fixed-Order Controller Design for Uncertain Systems
abstract
Typically, it is desirable to design a control system that is not only robustly stable in the presence of parametric uncertainties but also guarantees an adequate level of system performance. However, most of the existing methods need to take all extreme models over an uncertain domain into consideration, which then results in costly computation. Also, since these approaches attempt rather unrealistically to guarantee the system performance over a full frequency range, a conservative design is always admitted. Here, taking a specific viewpoint of robust stability and performance under a stated restricted frequency range (which is applicable in rather many real-world situations), this article provides an essential basis for the design of a fixed-order controller for a system with bounded parametric uncertainties, which avoids the tedious but necessary evaluations of the specifications on all the extreme models in an explicit manner. A Hurwitz polynomial is used in the design and the robust stability is characterized by the notion of positive realness, such that the required robust stability condition is then successfully constructed. Also, the robust performance criteria in terms of sensitivity shaping under different frequency ranges are constructed based on an approach of bounded realness analysis. Furthermore, the conditions for robust stability and performance are expressed in the framework of linear matrix inequality (LMI) constraints, and thus can be efficiently solved. Comparative simulations are provided to demonstrate the effectiveness and efficiency of the proposed approach.
Jun Ma 0008, Haiyue Zhu, Masayoshi Tomizuka, Tong Heng Lee
IEEE Trans. Syst. Man Cybern. Syst.1
2021 Adaptive Iterative Sliding Mode Control: Development, Synthesis, and Application of a Flexure-Joint Biaxial Gantry Stage
abstract
In this work, an adaptive iterative sliding mode control method is proposed for multi-axis mechatronic systems. Commonly, the multi-axis mechatronic systems are applied in high-speed and high-precision contouring tasks. For such contouring tasks, the multi-axis coordination is a main issue. As an inevitable challenge, several factors affect the multi-axis coordinate and diminish the contouring performance. Also, some special mechanical structure brings strong coupling to the system, which makes the system identification rather difficult. To solve these problems, this work proposes a learning-based totally model-free control approach for contouring tasks in application to such multi-axis motion stages. With this approach, all the coupling, disturbance, nonlinearity, and other unknown dynamics are regarded as lumped uncertainties in each axis. As a result, these uncertainties can be attenuated and compensated by the proposed controller. To analyze the contouring performance, a case study of a flexure-linked dual-drive H-gantry system is investigated to illustrate the effectiveness of the proposed method.
Jun Ma 0008, Zilong Cheng, Xiaocong Li, Tong Heng Lee
IECON2
2021 Parallel Collaborative Motion Planning with Alternating Direction Method of Multipliers
abstract
Collaborative motion planning for multi-agent systems is a challenging problem because of the existence of highly nonlinear and nonconvex constraints. Such difficulties also lead to inavoidable computational inefficiency, which significantly prohibits applying the existing collaborative motion planning algorithms to complex scenarios. This paper proposes a parallel computational algorithm to achieve collaborative motion planning efficiently, considering the nonlinear dynamics model and the nonconvex collision-avoidance constraints. Specifically, the alternating direction method of multipliers (ADMM) framework is elegantly incorporated to separate the large-scale cooperative nonconvex planning problem as two tractable and manageable subproblems, where the two subproblems handle the dynamics constraints and collision-free constraints, respectively. In the proposed approach, the differential dynamic programming (DDP) method is utilized to effectively solve the nonlinear subproblem with the dynamics constraints; meanwhile, the interior point (IPOPT) method is employed to address the nonconvex subproblem derived from the collision-avoidance constraints. Finally, two simulation scenarios are successfully implemented to illustrate the effectiveness of the proposed algorithm.
Zilong Cheng, Jun Ma 0008, Lin Zhao 0009, Cheng Xiang 0001, Tong Heng Lee
IECON3
2021 sGS-sPALM for Optimal Decentralized Control: A Distributed Optimization Approach
abstract
A distributed optimization algorithm for a decentralized control problem for uncertain systems is investigated in this paper. Based on ℋ2formulation, the optimal control problem under parameter uncertainties can be reformulated and solved in parameter space. Besides, the stabilizing controller gains of the decentralized control system with parameter uncertainties can be parameterized in a convex set; thus, the decentralized control problem can be reformulated as a conic optimization problem, which can be solved by using the symmetric Gauss-Seidel (sGS) semi-proximal augmented Lagrangian method (sPALM) efficiently. Then, a comprehensive analysis is provided to employ the sGS-sPALM to find the optimal solution of the decentralized control problem under parameter uncertainties. Robust performance and robust stability can be guaranteed using this methodology while satisfying the sparsity constraints resulting from the decentralized structure. Two examples are used to illustrate the effectiveness of the proposed method.
Jun Ma 0008, Zilong Cheng, Tong Heng Lee
IECON2
2021 Towards Adaptive Robust Control and Optimization for Constrained Uncertain Under-Actuated Mechanical Systems
abstract
For a specific class of under-actuated mechanical systems, non-holonomic servo constraints and model uncertainties are usually encountered. For such systems, this paper investigates the design of an adaptive robust controller with parameter optimization. A tighter link between the fuzzy set theory and the control of UMSs is bridged appropriately. Based on the UMSs with fuzzy information, an adaptive robust control method is then designed, and an analytical solution of the control input is determined, even if the servo constraints are non-holonomic. Furthermore, a concomitant parameter in the designed controller is analyzed, and a feasible controller admitting the optimal performance can be determined by minimizing a predefined performance index, such that the deterministic system performance can be ensured to be at a satisfying level. As supported by rigorous proofs, the existence and the uniqueness of the global solution to the optimization problem are presented. Finally, a numerical experiment is implemented to demonstrate the effectiveness of the proposed control design methodology.
Jun Ma 0008, Zilong Cheng, Han Zhao 0007, Abdullah Al Mamun 0002, Tong Heng Lee
SMC2
2021 Robust Control of a Two-Degree-of-Freedom Flexure-Based Nanopositioner for Planar Scanning Tasks
abstract
A two-degree-of-freedom (2-DoF) flexure-based nanopositioner is investigated for the planar scanning tasks, and a robust controller design scheme based on the convex inner approximation method is proposed. In practice, a flexure-based mechanism is usually represented by a second-order dynamic model. However, the second-order dynamic model cannot precisely fit the real system dynamics, and the model mismatch renders it difficult to achieve satisfying system performance in applications. Such a mismatch includes the parameter uncertainties caused by inaccurate model identification, different motion conditions, as well as high-order resonances. Note that if the controller is not well designed, the high-order resonances can be frequently activated, especially when the system input variation is significant. Therefore, to deal with the above impediments, a novel scheme for the robust controller design is proposed, with the variation of system input considered. In the proposed scheme, a subset of gains that can stabilize the closed-loop system is characterized elegantly via an inner approximation method considering the model uncertainties, and the formulated optimization problem regarding the determination of the controller parameters can be efficiently solved. Furthermore, the proposed scheme guarantees the performance regarding the H2-norm level and limits the H∞-norm level in a designated range. Finally, numerical optimization and comparative experiments are carried out, and the results evidently show the effectiveness of the proposed method.
Zilong Cheng, Jun Ma 0008, Xiaocong Li, Haiyue Zhu, Tong Heng Lee
SMC2
2021 Trajectory Generation by Chance-Constrained Nonlinear MPC With Probabilistic Prediction
abstract
Continued great efforts have been dedicated toward high-quality trajectory generation based on optimization methods; however, most of them do not suitably and effectively consider the situation with moving obstacles; and more particularly, the future position of these moving obstacles in the presence of uncertainty within some possible prescribed prediction horizon. To cater to this rather major shortcoming, this work shows how a variational Bayesian Gaussian mixture model (vBGMM) framework can be employed to predict the future trajectory of moving obstacles; and then with this methodology, a trajectory generation framework is proposed which will efficiently and effectively address trajectory generation in the presence of moving obstacles, and incorporate the presence of uncertainty within a prediction horizon. In this work, the full predictive conditional probability density function (PDF) with mean and covariance is obtained and, thus, a future trajectory with uncertainty is formulated as a collision region represented by a confidence ellipsoid. To avoid the collision region, chance constraints are imposed to restrict the collision probability, and subsequently, a nonlinear model predictive control problem is constructed with these chance constraints. It is shown that the proposed approach is able to predict the future position of the moving obstacles effectively; and, thus, based on the environmental information of the probabilistic prediction, it is also shown that the timing of collision avoidance can be earlier than the method without prediction. The tracking error and distance to obstacles of the trajectory with prediction are smaller compared with the method without prediction.
Jun Ma 0008, Zilong Cheng, Sunan Huang 0001, Shuzhi Sam Ge, Tong Heng Lee
IEEE Trans. Cybern.2
2020 Learning-Based Controller Optimization for Repetitive Robotic Tasks
abstract
Dynamic control for robotic automation tasks is traditionally designed and optimized with a model-based approach, and the performance relies heavily upon accurate system modeling. However, modeling the true dynamics of increasingly complex robotic systems is an extremely challenging task and it often renders the automation system to operate in a non-optimal condition. Notably, many industrial robotic applications involve repetitive motions and constantly generate a large amount of motion data under the non-optimal condition. These motion data contain rich information, and therefore an intelligent automation system should be able to learn from these non-optimal motion data to drive the system to operate optimally in a data-driven manner. In this paper, we propose a learning-based controller optimization algorithm for repetitive robotic tasks. To achieve this, a multi-objective cost function is designed to take into consideration both the trajectory tracking accuracy and smoothness, and then a data-driven approach is developed to estimate the gradient and Hessian based on the motion data for optimization without relying on the dynamic model. Experiments based on a magnetically-levitated nanopositioning system are conducted to demonstrate the effectiveness and practical appeals of the proposed algorithm in repetitive robotic automation tasks.
Xiaocong Li, Haiyue Zhu, Jun Ma 0008, Tat Joo Teo, Chek Sing Teo, Masayoshi Tomizuka, Tong Heng Lee
IROS3
2020 Grasping Detection Network with Uncertainty Estimation for Confidence-Driven Semi-Supervised Domain Adaptation
abstract
Data-efficient domain adaptation with only a few labelled data is desired for many robotic applications, e.g., in grasping detection, the inference skill learned from a grasping dataset is not universal enough to directly apply on various other daily/industrial applications. This paper presents an approach enabling the easy domain adaptation through a novel grasping detection network with confidence-driven semi-supervised learning, where these two components deeply interact with each other. The proposed grasping detection network specially provides a prediction uncertainty estimation mechanism by leveraging on Feature Pyramid Network (FPN), and the mean-teacher semi-supervised learning utilizes such uncertainty information to emphasizing the consistency loss only for those unlabelled data with high confidence, which we referred it as the confidence-driven mean teacher. This approach largely prevents the student model to learn the incorrect/harmful information from the consistency loss, which speeds up the learning progress and improves the model accuracy. Our results show that the proposed network can achieve high success rate on the Cornell grasping dataset, and for domain adaptation with very limited data, the confidence- driven mean teacher outperforms the original mean teacher and direct training by more than 10% in evaluation loss especially for avoiding the overfitting and model diverging.
Haiyue Zhu, Fengjun Bai, Xiaocong Li, Jun Ma 0008, Chek Sing Teo, Pey Yuen Tao, Wei Lin 0002
IROS6
2019 Data-Driven Tuning Method for LQR Based Optimal PID Controller
abstract
Data-driven control methods for modern controller design are becoming popular recently. However, the traditional Proportional-Integral-Derivative (PID) controller is still the most widely used controller to the industrial preference. To tune the parameters of the PID controller, optimal PID tuning approaches such as solving the Riccati equation of the Linear Quadratic Regulator (LQR) provide the optimal solution. The disadvantages of the LQR are that an accurate model of the system is required, and the high-order system must be reduced to the second-order system so that the Riccati equation can be solved. In this paper, a novel data-driven method is proposed to cope with these problems. For the system which is difficult to be identified accurately, the proposed data-driven method can skip the procedure of system identification and tune the parameters of the PID controller directly with the experimental data instead of solving the Riccati equation. This data-driven tuning method also ensures that the parameters of the PID controller for the high-order system are optimized without using the reduced-order model of the system. Simulations are conducted on a tray indexing system with the second-order model and the full-order model demonstrating high applicability and accuracy of the proposed method.
Zilong Cheng, Xiaocong Li, Jun Ma 0008, Chek Sing Teo, Kok Kiong Tan, Tong Heng Lee
IECON3
2019 Robust Decentralized Controller Synthesis in Flexure-Linked H-Gantry by Iterative Linear Programming
abstract
The dual-drive H-gantry is widely used for high-speed, high-precision Cartesian motion. Compared with the conventional rigid-linked design, the flexure-linked counterpart is able to prevent the damage of joints for its smaller interaxial coupling force. However, there are still barriers to further push up its precision, such as parametric uncertainties due to the inaccurate dynamical model, the possible induced vibration during high-speed motion, and the decentralized control structure required by industries. To maintain the tracking precision of carriages and minimize the vibration of the end effector, we aim to optimize parameters in decentralized controllers with choices of flexure pieces. We find that such decentralized feedback structure yields some uncontrollable but stabilizable states in the closed-loop system, and no direct solution from solving the algebraic Riccati equation is available in this case. Such structural constraint, together with constraints due to stability requirement and model uncertainties facilitates us to formulate an H2guaranteed cost control problem within a projected convex domain. From here, efficient numerical procedures are developed to obtain the global optimum by iterative linear programming. The real-time experiment validates the optimality and the robustness of the proposed method.
Jun Ma 0008, Si-Lu Chen 0001, Wenyu Liang, Chek Sing Teo, Arthur Tay, Abdullah Al Mamun 0002, Kok Kiong Tan
IEEE Trans. Ind. Informatics1
2018 Intelligent Motion Control of Ultrasonic Motor for an Ear Surgical Device
abstract
Ultrasonic motors (USMs) are widely used in many applications that precise and fast motions are required, such as precision machines and medical devices, etc. In this paper, a USM -driven ear surgical device is introduced. To address the control challenges of the USM, an intelligent controller consisting of a cerebellar model articulation controller (CMAC) and a sliding mode compensator (SMC) is designed. Therein, the CMAC serves as the main controller due to its fast learning ability while the SMC is used to eliminate the approximation error between the CMAC and the desired perfect controller. Several experiments are carried out to validate the effectiveness of the proposed control scheme, and the results show that the proposed control scheme is able to achieve good tracking performance and guaranteed robustness.
Wenyu Liang, Sunan Huang 0001, Jun Ma 0008, Kok Kiong Tan
IECON3
2016 Optimal decentralized control approach toward integrated design of controller and jerk-decoupling cartridge
abstract
Linear direct feed drives are widely used in machine tools, but an abrupt counter force from the secondary part will induce the jerk to the metro frame contacted with the linear motor and cause the vibration of auxiliary devices on it. The jerk-decoupling cartridge (JDC) provides a buffer to reduce such an impact. To systematically take care of both the tracking error and the jerk induced to the metro frame, this paper presents an integrated design approach to determine parameters in the JDC and the position controller of the feed drive. The initial formulated non-convex optimization problem is converted to convex constrained gradient optimization problem and linear step searching problem. Thus, fast convergence of parameters is achieved within first few iterations. Through a series of simulation, the effectiveness of proposed methodology is verified.
Jun Ma 0008, Si-Lu Chen 0001, Chek Sing Teo, Chun Jeng Kong, Arthur Tay, Wei Lin 0002, Abdullah Al Mamun 0002
IECON1