VLDB 2026 Research / reviewers in the wild / expert
Xueqian Wang 0001
dblp:43/3563-1
· DBLP profile ↗
114ranked-venue papers
1as first author
97since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 73 · 64 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 17 since 2021Systems, architecture and hardware · 19 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 14 since 2021Human-computer interaction and ubiquitous computing · 19 · 14 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Surrogate-Assisted Evolutionary Multi-Agent Reinforcement Learning with Adaptive Fitness EvaluationabstractDeep Multi-Agent Reinforcement Learning (MARL) excels in co-operative tasks but often struggles with local optima in high - dimensional joint action spaces. In contrast, Evolutionary Algorithms (EAs) offer robust global exploration capabilities. Although hybrid approaches seek to combine the strengths of both paradigms, they typically face a critical bottleneck: the prohibitive sample cost of evaluating large populations via environment rollouts. To address this challenge, we propose Surrogate-assisted Evolutionary Multi-Agent Reinforcement Learning (SEMARL), a unified framework that synergizes gradient-based refinement with surrogate-assisted evolutionary search. SEMARL employs a cooperative co-evolutionary architecture to maintain diverse agent policies and injects gradient-refined parameters into the population to accelerate convergence. Crucially, we leverage the centralized critic from the gradient learner as a computationally efficient surrogate for fitness estimation. To prevent misleading guidance from an inaccurate critic, we introduce an adaptive reliability control mechanism based on temporal difference (TD) error, which dynamically regulates the surrogate's influence. Experiments on the Multi-Agent MuJoCo benchmark demonstrate that SEMARL significantly outperforms other algorithms, achieving superior asymptotic performance with substantially higher sample efficiency. Cong Yu 0018, Zaihui Yang, Haoyu Wang 0018, Junbo Tan, Yongzhe Chang, Tiantian Zhang 0002, Xueqian Wang 0001 |
GECCO | 9 |
| 2026 | Count Every Rotation and Every Rotation Counts: Exploring Drone Dynamics via Propeller SensingabstractAs drone-based applications proliferate, paramount contactless sensing of airborne drones from the ground becomes indispensable. This work demonstrates concentrating on propeller rotational speed will substantially improve drone sensing performance and proposes an event-camera-based solution, EventPro. EventPro features two components: Count Every Rotation achieves accurate, real-time propeller speed estimation by mitigating ultra-high sensitivity of event cameras to environmental noise. Every Rotation Counts leverages these speeds to infer both internal and external drone dynamics. Extensive evaluations in real-world drone delivery scenarios show that EventPro achieves a sensing latency of 3 ms and a rotational speed estimation error of merely 0.23%. Additionally, EventPro infers drone flight commands with 96.5% precision and improves drone tracking accuracy by over 22% when combined with other sensing modalities. Demo: https://eventpro25.github.io/EventPro/. Xuecheng Chen, Jingao Xu, Wenhua Ding, Haoyang Wang 0012, Xinyu Luo, Ruiyang Duan, Xueqian Wang 0001, Yunhao Liu 0001, Xinlei Chen |
SenSys | 8 |
| 2026 | GE-adapter: A general and efficient adapter for enhanced video editing with pretrained text-to-image diffusion models
Yangfan He, Kun Li 0014, Jianhui Wang 0001, Binxu Li, Tianyu Shi 0003, Miao Zhang 0010, Xueqian Wang 0001 |
Expert Syst. Appl. | 11 |
| 2026 | TCSTNet: A text-driven color style transfer network for low-light image enhancement
Tianyi Zeng, Miao Zhang 0010, Zimo Zeng, Junfeng Jiao, Yuantao Wang, Yangfan He, Junbo Tan, Christian G. Claudel, Xueqian Wang 0001 |
Expert Syst. Appl. | 13 |
| 2026 | LowLightReward: A unified framework for low-light enhancement across spatial, channel, and aesthetic domains
Miao Zhang 0010, Haoyue Han, Yuantao Wang, Chenghe Yang, Hanning Liu, Junbo Tan, Xueqian Wang 0001 |
Neurocomputing | 11 |
| 2026 | Manifold-aware triple cooperative multi-population differential evolution with reinforcement learning for irregular 3D UAV path planning
Yunhui Zhang, Guanglong Du, Ziwei Wang 0001, Xueqian Wang 0001, Cuifeng Du, Quanlong Guan, Xiaojian Qiu |
Knowl. Based Syst. | 5 |
| 2026 | Tacit mechanism: Bridging pre-training of individuality to multi-agent adversarial coordination
Shiqing Yao, Jiajun Chai, Haixin Yu, Yongzhe Chang, Tiantian Zhang 0002, Yuanheng Zhu, Xueqian Wang 0001 |
Neural Networks | 7 |
| 2026 | Decentralized Partial Model Personalization With Guaranteed Nonconvex ConvergenceabstractCompared with general Federated Learning (FL), Decentralized FL (DFL) has diminished central communication burdens and lower risks of disruption. To be compatible with real-world scenarios, some existing works produce several local personalized models rather than a universal model for all edge devices. However, they still suffer from inferior performance due to the full model aggregation in heterogeneous datasets. Therefore, we propose a DFL framework DFedMDC through partial model personalization, which can adapt to resource-heterogeneous environments (e.g., the Internet of Things (IoT)). Specifically, it personalizes the “right” components of local models and trains the shared modeluiand personal modelvialternately in each clienti. To further accelerate the convergence process, we propose DFedSMDC with properly directed noise perturbation into the gradient update. Theoretically, we provide convergence analysis of both algorithms in the general non-convex setting. It can shed light on how vital factors affect the convergence rate, such as these alternate updates of partial modelsui,vi, data heterogeneity δ2as well as various communication topologies (characterized by the spectral gap 1 − λ). Empirically, we confirm the state-of-the-art (SOTA) superiority of the proposed methods on several real-world datasets with various data distributions, relative to both SOTA personalized FL (PFL) and DFL baselines. Yingqi Liu, Zihao Lin 0003, Xueqian Wang 0001, Li Shen 0008, Xiaochun Cao, Dacheng Tao |
IEEE Trans. Netw. | 5 |
| 2026 | An Efficient Solution Method for Workspace Boundary of Serpentine Manipulators Based on a Unified Kinematics Model
Deshan Meng, Taowen Guo, Runhui Xiang, Junbo Tan, Xueqian Wang 0001, Bin Liang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2025 | Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence MinimizationabstractDirect Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to f-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of text-to-image alignment paradigm under f-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on text-image alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that text-to-image alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications. Bo Xia, Yongzhe Chang, Xueqian Wang 0001 |
AAAI | 4 |
| 2025 | Identical Human Preference Alignment Paradigm for Text-to-Image ModelsabstractImplicit reward mechanism of Direct Preference Optimization (DPO) has facilitated its recent applications beyond large language models (LLMs), notably in aligning text-to-image models with human preferences. While promising results have been achieved with algorithms such as Diffusion-DPO, their reliance on the assumptions of the Bradley-Terry model could potentially lead to significant overfitting. In this paper, we propose the Step Identical Preference Alignment (SIPA) method, departing text-to-image alignment from the assumptions of Bradley-Terry preference model. We assess the performance of four models, Diffusion-DPO, SPO, and SIPA, alongside the original model, on the HPS-V2 test set, which focus on three key aspects: text-image alignment, human value alignment, and generation diversity. Experimental results show that SIPA matches or outperforms existing SOTA alignment methods, and even exceeds the original model in terms of generation diversity, which compellingly demonstrates SIPA’s superiority in mitigating alignment overfitting. Bo Xia, Yongzhe Chang, Xueqian Wang 0001 |
ICASSP | 5 |
| 2025 | Positive Enhanced Preference Alignment for Text-to-Image ModelsabstractDirect Preference Optimization (DPO) has recently expanded its successful application beyond aligning large language models (LLMs), further targeting the alignment of text-to-image models with human preferences. However, traditional DPO approach would inadvertently result in a simultaneous reduction of sampling probabilities for preferred and dispreferred items during the alignment process, thereby potentially diminishing model's generative capacity. In this paper, we firstly undertake a revisit of DPO by grounding our analysis in the framework of contrastive loss. It reveals that DPO only emphasizes the part quantifying dissimilarity between items, while overlooking aspects pertinent to positive items. Hence, we propose the Positive Enhanced Preference Alignment (PEPA). Three enhancement strategies are introduced herein, and after comprehensive empirical evaluation, we recommend implementation of enhancing the log probability of preferred ratio in practice applications, which is distinguished by both stability and effectiveness. Experimental assessments are carried out on the HPS-V2 test set, with results demonstrating that PEPA outperforms or matches current state-of-the-art alignment techniques, thus highlighting PEPA's exceptional practical efficacy. Bo Xia, Yongzhe Chang, Xueqian Wang 0001 |
ICASSP | 5 |
| 2025 | FOSP: Fine-tuning Offline Safe Policy through World ModelsabstractOffline Safe Reinforcement Learning (RL) seeks to address safety constraints by learning from static datasets and restricting exploration. However, these approaches heavily rely on the dataset and struggle to generalize to unseen scenarios safely. In this paper, we aim to improve safety during the deployment of vision-based robotic tasks through online fine-tuning an offline pretrained policy. To facilitate effective fine-tuning, we introduce model-based RL, which is known for its data efficiency. Specifically, our method employs in-sample optimization to improve offline training efficiency while incorporating reachability guidance to ensure safety. After obtaining an offline safe policy, a safe policy expansion approach is leveraged for online fine-tuning. The performance of our method is validated on simulation benchmarks with five vision-only tasks and through real-world robot deployment using limited data. It demonstrates that our approach significantly improves the generalization of offline policies to unseen safety-constrained scenarios. To the best of our knowledge, this is the first work to explore offline-to-online RL for safe generalization tasks. The videos are available at https://sunlighted.github.io/fosp_web/. Yucheng Xin, Silang Wu, Longxiang He, Zichen Yan, Junbo Tan, Xueqian Wang 0001 |
ICLR | 7 |
| 2025 | Entropy-based Activation Function Optimization: A Method on Searching Better Activation FunctionsabstractThe success of artificial neural networks (ANNs) hinges greatly on the judicious selection of an activation function, introducing non-linearity into network and enabling them to model sophisticated relationships in data. However, the search of activation functions has largely relied on empirical knowledge in the past, lacking theoretical guidance, which has hindered the identification of more effective activation functions. In this work, we offer a proper solution to such issue. Firstly, we theoretically demonstrate the existence of the worst activation function with boundary conditions (WAFBC) from the perspective of information entropy. Furthermore, inspired by the Taylor expansion form of information entropy functional, we propose the Entropy-based Activation Function Optimization (EAFO) methodology. EAFO methodology presents a novel perspective for designing static activation functions in deep neural networks and the potential of dynamically optimizing activation during iterative training. Utilizing EAFO methodology, we derive a novel activation function from ReLU, known as Correction Regularized ReLU (CRReLU). Experiments conducted with vision transformer and its variants on CIFAR-10, CIFAR-100 and ImageNet-1K datasets demonstrate the superiority of CRReLU over existing corrections of ReLU. Extensive empirical studies on task of large language model (LLM) fine-tuning, CRReLU exhibits superior performance compared to GELU, suggesting its broader potential for practical applications. Bo Xia, Pu Chang, Zibin Dong, Yifu Yuan, Yongzhe Chang, Xueqian Wang 0001 |
ICLR | 8 |
| 2025 | Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision TransformerabstractDespite recent advancements in offline multi-task reinforcement learning (MTRL) have harnessed the powerful capabilities of the Transformer architecture, most approaches focus on a limited number of tasks, with scaling to extremely massive tasks remaining a formidable challenge. In this paper, we first revisit the key impact of task numbers on current MTRL method, and further reveal that naively expanding the parameters proves insufficient to counteract the performance degradation as the number of tasks escalates. Building upon these insights, we propose M3DT, a novel mixture-of-experts (MoE) framework that tackles task scalability by further unlocking the model’s parameter scalability. Specifically, we enhance both the architecture and the optimization of the agent, where we strengthen the Decision Transformer (DT) backbone with MoE to reduce task load on parameter subsets, and introduce a three-stage training mechanism to facilitate efficient training with optimal performance. Experimental results show that, by increasing the number of experts, M3DT not only consistently enhances its performance as model expansion on the fixed task numbers, but also exhibits remarkable task scalability, successfully extending to 160 tasks with superior performance. Yilun Kong, Guozheng Ma, Haoyu Wang 0018, Li Shen 0008, Xueqian Wang 0001, Dacheng Tao |
ICML | 6 |
| 2025 | Safety Reasoning with GuidelinesabstractTraining safe LLMs remains a critical challenge. The most widely used method, Refusal Training (RT), struggles to generalize against various Out-of-Distribution (OOD) jailbreaking attacks. Although various advanced methods have been proposed to address this issue, we instead question whether OOD attacks inherently surpass the capability of vanilla RT. Evaluations using Best-of-N (BoN) reveal significant safety improvements as N increases, indicating models possess adequate latent safety knowledge but RT fails to consistently elicit it under OOD scenarios. Further domain adaptation analysis reveals that direct RT causes reliance on superficial shortcuts, resulting in non-generalizable representation mappings. Inspired by our findings, we propose training model to perform safety reasoning for each query. Specifically, we synthesize reasoning supervision aligned with specified guidelines that reflect diverse perspectives on safety knowledge. This encourages model to engage in deeper reasoning, explicitly eliciting and utilizing latent safety knowledge for each query. Extensive experiments show that our method significantly improves model generalization against OOD attacks. Haoyu Wang 0018, Zeyu Qin, Li Shen 0008, Xueqian Wang 0001, Dacheng Tao, Minhao Cheng |
ICML | 4 |
| 2025 | D3-ARM: High-Dynamic, Dexterous and Fully Decoupled Cable-Driven Robotic ArmabstractCable transmission enables motors of robotic arm to operate lightweight and low-inertia joints remotely in various environments, but it also creates issues with motion coupling and cable routing that can reduce arm's control precision and performance. In this paper, we present a novel motion decoupling mechanism with low-friction to align the cables and efficiently transmit the motor's power. By arranging these mechanisms at the joints, we fabricate a fully decoupled and lightweight cable-driven robotic arm called D3-Arm with all the electrical components be placed at the base. Its 776 mm length moving part boasts six degrees of freedom (DOF) and only 1.6 kg weights. To address the issue of cable slack, a cable-pretension mechanism is integrated to enhance the stability of long-distance cable transmission. Through a series of comprehensive tests, D3-Arm demonstrated 1.29 mm average positioning error and 2.0 kg payload capacity, proving the practicality of the proposed decoupling mechanisms in cable-driven robotic arm. Jianle Xu, Shoujie Li, Huayue Liang, Yanbo Chen 0001, Chongkun Xia, Xueqian Wang 0001 |
ICRA | 7 |
| 2025 | Learning Pre-Trained Tacit Behavior for Efficient Multi-Agent Adversarial Coordination
Shiqing Yao, Jiajun Chai, Haixin Yu, Yongzhe Chang, Yuanheng Zhu, Xueqian Wang 0001 |
AAMAS | 6 |
| 2025 | DeepMF: Deep Motion Factorization for Closed-Loop Safety-Critical Driving Scenario SimulationabstractSafety-critical traffic scenarios are of great practical relevance to evaluating the robustness of autonomous driving (AD) systems. Given that these long-tail events are extremely rare in real-world traffic data, there is a growing body of work dedicated to the automatic traffic scenario generation. However, nearly all existing algorithms for generating safety-critical scenarios rely on snippets of previously recorded traffic events, transforming normal traffic flow into accident-prone situations directly. In other words, safety-critical traffic scenario generation is hindsight and not applicable to newly encountered and open-ended traffic events. In this paper, we propose the Deep Motion Factorization (DeepMF) framework, which extends static safety-critical driving scenario generation to closed-loop and interactive adversarial traffic simulation. DeepMF casts safety-critical traffic simulation as a Bayesian factorization that includes the assignment of hazardous traffic participants, the motion prediction of selected opponents, the reaction estimation of autonomous vehicle (AV) and the probability estimation of the accident occur. All the aforementioned terms are calculated using decoupled deep neural networks, with inputs limited to the current observation and historical states. Consequently, DeepMF can effectively and efficiently simulate safety-critical traffic scenarios at any triggered time and for any duration by maximizing the compounded posterior probability of traffic risk. Extensive experiments demonstrate that DeepMF excels in terms of risk management, flexibility, and diversity, showcasing outstanding performance in simulating a wide range of realistic, high-risk traffic scenarios. Linrui Zhang, Bo Xia, Xueqian Wang 0001, Houde Liu |
IJCNN | 4 |
| 2025 | Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement LearningabstractDiffusion probability models have shown significant promise in offline reinforcement learning by directly modeling trajectory sequences. However, existing approaches primarily focus on time-domain features while overlooking frequency-domain features, leading to frequency shift and degraded performance according to our observation. In this paper, we investigate the RL problem from a new perspective of the frequency domain. We first observe that time-domain-only approaches inadvertently introduce shifts in the low-frequency components of the frequency domain, which results in trajectory instability and degraded performance. To address this issue, we propose Wavelet Fourier Diffuser (WFDiffuser), a novel diffusion-based RL framework that integrates Discrete Wavelet Transform to decompose trajectories into low- and high-frequency components. To further enhance diffusion modeling for each component, WFDiffuser employs Short-Time Fourier Transform and cross attention mechanisms to extract frequency-domain features and facilitate cross-frequency interaction. Extensive experiment results on the D4RL benchmark demonstrate that WFDiffuser effectively mitigates frequency shift, leading to smoother, more stable trajectories and improved decision-making performance over existing methods. Yifu Luo, Yongzhe Chang, Xueqian Wang 0001 |
IJCNN | 3 |
| 2025 | AirTouch: A Low-Cost Versatile Visuotactile Feedback System for Enhanced Robotic TeleoperationabstractVision-based teleoperation systems are widely used due to their cost-effectiveness and intuitive operation. However, these systems often suffer from challenges such as hand occlusions, environmental variability, and the lack of tactile feedback, limiting their precision and applicability in complex tasks. To address these limitations, we present Air-Touch, a novel, low-cost visuotactile teleoperation system that integrates air pressure-based tactile feedback with lightweight hand pose estimation. AirTouch features an inflatable tactile bubble that provides adjustable feedback through closed-loop pneumatic control, enhancing the operator’s sense of interaction with remote environments. The system’s robust hand-tracking algorithm ensures accurate control even under dynamic and occlusion-prone conditions, while its hardware design eliminates the need for wearable devices, enabling intuitive operation. AirTouch supports a wide range of robotic end-effectors, including dexterous hands, parallel grippers, and suction cups, demonstrating versatility across multiple platforms. Extensive experiments validate AirTouch’s performance, achieving high precision in hand pose estimation and a 91% success rate in complex teleoperation tasks, all with a hardware cost as low as $39. These results highlight AirTouch as a scalable and practical solution for enhancing robotic teleoperation across industrial, medical, and hazardous scenarios. Shoujie Li, Xingting Li, Ken Jiankun Zheng, Xueqian Wang 0001, Wenbo Ding 0001 |
IROS | 6 |
| 2025 | Scalable MARL for Cooperative Exploration with Dynamic Robot Populations via Graph-Based Information AggregationabstractThis study addresses the challenge of multi-robot cooperative exploration under limited local observations in environments with dynamic robot populations. To achieve efficient area coverage within constrained timeframes, we propose the Multi-Robot Informative Planner (MIP), a novel reinforcement learning (RL)-based planning module. The core component of MIP is the Neighborhood Information Aggregator, which employs a graph neural network (GNN) to integrate local neighborhood information for each robot. Our design enhances sample efficiency by minimizing information requirements while ensuring scalability across environments with varying robot numbers. To generate high-quality, expressive neighborhood feature representations, we utilize Graphical Mutual Information (GMI) to maximize the correlation between neighboring robots’ input features and their high-level hidden representations. Furthermore, MIP incorporates the Spatial-Neighborhood Transformer, which captures spatial features and inter-robot interactions through spatial self-attention mechanisms. These components collectively form the Multi-Robot Neural Informative Mapping (MRNIM) framework, outperforming traditional benchmarks in Habitat simulator. Xiaoqi Ren, Guanglong Du, Zhuoyao Wang 0001, Xueqian Wang 0001, Quanlong Guan, Xiaojian Qiu |
IROS | 5 |
| 2025 | MuxHand: A Cost-Effective and Compact Dexterous Robotic Hand Using Time-Division Multiplexing MechanismabstractThe number of motors directly influences the dexterity, size, and cost of a robotic hand. In this paper, we present MuxHand, a robotic hand that utilizes a time-division multiplexing motor (TDMM) mechanism. This system enables independent control of 9 cables with just 4 motors, significantly reducing both cost and size while maintaining high dexterity. To enhance stability and smoothness during grasping and manipulation tasks, we integrate magnetic joints into the three 3D-printed fingers. These joints provide impact resistance, resetting capabilities. The three fingers together have a total of 30 degrees of freedom (DOF), 18 of which are passive DOF, allowing the hand to conform closely to the surface of an object during grasping. We conduct a series of experiments to assess the performance parameters of MuxHand, including its grasping and manipulation capabilities. The results show that the TDMM mechanism precisely controls each cable connected to the finger joints, enabling robust grasping and dexterous manipulation. Furthermore, compared to the traditional approach of assigning a motor to each active DOF, the cost is reduced by 42.06%. The maximum load of a single finger reaches 7.0 kg, the maximum load at the finger joint root is 12.0 kg, the maximum driving force at the joint root is 5.0 kg, and the maximum fingertip force is 10.0 N. Jianle Xu, Shoujie Li, Houde Liu, Xueqian Wang 0001, Wenbo Ding 0001, Chongkun Xia |
IROS | 5 |
| 2025 | PlanVWM: Autonomous Vehicle Planning Method Based on Vectorized World Model ModelingabstractWith the increasing demand for handling complex scenarios in autonomous driving, data-driven planning methods based on imitation learning have attracted significant attention. In this context, this work proposes the PlanVMN method, aiming to enhance the model's robustness to data and its causal inference ability in planning. Based on vectorized scene inputs, this method integrates the world model into the core architecture of the planning model and deduces the temporal features of the scene based on the actions of the ego vehicle. Meanwhile, we introduce an attention-based history enhancement component within the world model, which remarkably improves the planning performance of the ego vehicle. Experiments show that the model has achieved outstanding results in the open-loop and closed-loop tests of nuPlan. Notably, thanks to the generative ability endowed by the unique world model architecture, the model can still exhibit excellent performance when facing blind area, sensor misfuctioning or detection instability, providing strong support for planning in complex scenarios of autonomous driving. Yuying Chen, Ziqing Gu, Siyuan Cheng 0012, Xueqian Wang 0001 |
IV | 8 |
| 2025 | Twin Co-Adaptive Dialogue for Progressive Image GenerationabstractModern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in user prompts. In this work, we present Twin-Co, a framework that leverages synchronized, co-adaptive dialogue to progressively refine image generation. Instead of a static generation process, Twin-Co employs a dynamic, iterative workflow where an intelligent dialogue agent continuously interacts with the user. Initially, a base image is generated from the user's prompt. Then, through a series of synchronized dialogue exchanges, the system adapts and optimizes the image according to evolving user feedback. The co-adaptive process allows the system to progressively narrow down ambiguities and better align with user intent. Experiments demonstrate that Twin-Co not only enhances user experience by reducing trial-and-error iterations but also improves the quality of the generated images, streamlining creative process across various applications. Jianhui Wang 0001, Yangfan He, Yan Zhong 0001, Xinyuan Song 0002, Jiayi Su, Yuheng Feng, Hongyang He, Wenyu Zhu, Xinhang Yuan, Miao Zhang 0010, Tianyu Shi 0003, Xueqian Wang 0001 |
ACM Multimedia | 15 |
| 2025 | Robust Policy Expansion for Offline-to-Online RL under Diverse Data CorruptionabstractPretraining a policy on offline data followed by fine-tuning through online interactions, known as Offline-to-Online Reinforcement Learning (O2O RL), has emerged as a promising paradigm for real-world RL deployment. However, both offline datasets and online interactions in practical environments are often noisy or even maliciously corrupted, severely degrading the performance of O2O RL. Existing works primarily focus on mitigating the conservatism of offline policies via online exploration, while the robustness of O2O RL under data corruption, including states, actions, rewards, and dynamics, is still unexplored. In this work, we observe that data corruption induces heavy-tailed behavior in the policy, thereby substantially degrading the efficiency of online exploration. To address this issue, we incorporate Inverse Probability Weighted (IPW) into the online exploration policy to alleviate heavy-tailedness, and propose a novel, simple yet effective method termed $\textbf{RPEX}$: $\textbf{R}$obust $\textbf{P}$olicy $\textbf{EX}$pansion. Extensive experimental results on D4RL datasets demonstrate that RPEX achieves SOTA O2O performance across a wide range of data corruption scenarios. Longxiang He, Deheng Ye, Junbo Tan, Xueqian Wang 0001, Li Shen 0008 |
NeurIPS | 4 |
| 2025 | Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image GenerationabstractReinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO, the first method to incorporate Group Relative Policy Optimization (GRPO)-based RL into this overlooked paradigm. Our core insight is to redefine the transition probability, which is different from current approaches, and formulate the unmasking process as a multi-step decision-making problem. To further enhance our method, we explore several useful strategies, including removing the Kullback–Leibler constraint, applying the reduction strategy, and filtering out low-quality samples. Using Mask-GRPO, we improve a base model, Show-o, with substantial improvements on standard T2I benchmarks and preference alignment, outperforming existing state-of-the-art approaches. Yifu Luo, Xinhao Hu, Keyu Fan, Bo Xia, Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001 |
NeurIPS | 9 |
| 2025 | Lifelong Safety Alignment for Language ModelsabstractLLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment. While many existing defenses focus on known types of attacks, it is more critical to prepare LLMs for *unseen* attacks that may arise during deployment. To address this, we propose a **lifelong safety alignment** framework that enables LLMs to continuously adapt to new and evolving jailbreaking strategies. Our framework introduces a competitive setup between two components: a **Meta-Attacker**, trained to actively discover novel jailbreaking strategies, and a **Defender**, trained to resist them. To effectively warm up the Meta-Attacker, we first leverage the GPT-4o API to extract key insights from a large collection of jailbreak-related research papers. Through iterative training, the first iteration Meta-Attacker achieves a 73% attack success rate (ASR) on RR and a 57% transfer ASR on LAT using only *single-turn* attacks. Meanwhile, the Defender progressively improves its robustness and ultimately reduces the Meta-Attacker's success rate to just 7%, enabling safer and more reliable deployment of LLMs in open-ended environments. Haoyu Wang 0018, Zeyu Qin, Xueqian Wang 0001, Tianyu Pang |
NeurIPS | 6 |
| 2025 | Prompt-Guided Region-Adaptive Enhancement for Aesthetic Low-Light Imaging
Miao Zhang 0010, Pengyu Zeng, Xueqian Wang 0001 |
PRCV (9) | 7 |
| 2025 | A Universal Vehicle-Trailer Navigation System with Neural Kinematics and Online Residual LearningabstractAutonomous navigation of vehicle-trailer systems is crucial in environments like airports, supermarkets, and concert venues, where various types of trailers are needed to navigate with different payloads and conditions. However, accurately modeling such systems remains challenging, especially for trailers with castor wheels. In this work, we propose a novel universal vehicle-trailer navigation system that integrates a hybrid nominal kinematic model—combining classical nonholonomic constraints for vehicles and neural network-based trailer kinematics—with a lightweight online residual learning module to correct real-time modeling discrepancies and disturbances. Additionally, we develop a model predictive control framework with a weighted model combination strategy that improves long-horizon prediction accuracy and ensures safer motion planning. Our approach is validated through extensive real-world experiments involving multiple trailer types and varying payload conditions, demonstrating robust performance without manual tuning or trailer-specific calibration. Yanbo Chen 0001, Yunzhe Tan, Yaojia Wang, Zhengzhe Xu, Junbo Tan, Xueqian Wang 0001 |
SMC | 6 |
| 2025 | SafeSim: An Open-Source Platform for Safety-Critical Driving Scenario Simulation and Curriculum-based Adversarial TrainingabstractWe present SafeSim, a comprehensive benchmarking platform and a unified end-to-end framework for safety-critical driving scenario simulation and curriculum-based adversarial training. SafeSim integrates 8 classic adversarial environment generation algorithms, enabling the generation at any trigger moment and for any duration based on various naturalistic traffic datasets. During the automatic vehicle (AV) training process, SafeSim dynamically adjusts the risk-level of scenarios based on the current capability of AV, progressively enhancing its ability to handle accident-prone situations. The platform features rich interfaces and extensibility, providing a complete workflow and tool-chain for risk scenario generation, risk assessment, AV training, and algorithm evaluation. Additionally, SafeSim includes multiple baseline results, serving as a standardized benchmark for future research. Linrui Zhang, Jiuzhou Lin, Xueqian Wang 0001, Houde Liu |
SMC | 7 |
| 2025 | Data-Driven MPC with Data Selection for Flexible Cable-Driven Robotic ArmsabstractFlexible cable-driven robotic arms (FCRAs) offer dexterous and compliant motion. Still, the inherent properties of cables, such as resilience, hysteresis, and friction, often lead to particular difficulties in modeling and control. This paper proposes a model predictive control (MPC) method that relies exclusively on input-output data, without a physical model, to improve the control accuracy of FCRAs. First, we develop an implicit model based on input-output data and integrate it into an MPC optimization framework. Second, a data selection algorithm (DSA) is introduced to filter the data that best characterize the system, thereby reducing the solution time per step to approximately 4 ms, which is an improvement of nearly 80%. Lastly, the influence of hyperparameters on tracking error is investigated through simulation. The proposed method has been validated on a real FCRA platform, including five-point positioning accuracy tests, a five-point response tracking test, and trajectory tracking for letter drawing. The results demonstrate that the average positioning accuracy is approximately 2.070 mm. Moreover, compared to the PID method with an average tracking error of 1.418°, the proposed method achieves an average tracking error of 0.541°. Huayue Liang, Yanbo Chen 0001, Hongyang Cheng, Yanzhao Yu, Shoujie Li, Junbo Tan, Xueqian Wang 0001, Long Zeng 0001 |
SMC | 7 |
| 2025 | A Hybrid Force-Position Strategy for Shape Control of Deformable Linear Objects With Graph Attention NetworksabstractManipulating deformable linear objects (DLOs) such as wires and cables is crucial in various applications like electronics assembly and medical surgeries. However, it faces challenges due to DLOs’ infinite degrees of freedom, complex nonlinear dynamics, and the underactuated nature of the system. To address these issues, this paper proposes a hybrid force-position strategy for DLO shape control. The framework, combining both force and position representations of DLO, integrates state trajectory planning in the force space and Model Predictive Control (MPC) in the position space. We present a dynamics model with an explicit action encoder, a property extractor and a graph processor based on Graph Attention Networks. The model is used in the MPC to enhance prediction accuracy. Results from both simulations and real-world experiments demonstrate the effectiveness of our approach in achieving efficient and stable shape control of DLOs. Codes and videos are available at https://sites.google.com/view/dlom. Yanzhao Yu, Junbo Tan, Xueqian Wang 0001 |
SMC | 4 |
| 2025 | Global-Guided Focal Neural Radiance Field for Large-Scale Scene RenderingabstractNeural radiance fields (NeRF) have recently been applied to render large-scale scenes. However, their limited model capacity typically results in blurred rendering results. Existing large-scale NeRFs primarily address this limitation by partitioning the scene into blocks, which are subsequently handled by separate sub-NeRFs. These sub-NeRFs, trained from scratch and processed independently, lead to inconsistencies in geometry and appearance across the scene. Consequently, the rendering quality fails to exhibit significant improvement despite the expansion of model capacity. In this work, we present global-guided focal neural radiance field (GF-NeRF) that achieves high-fidelity rendering of large-scale scenes. Our proposed GF-NeRF utilizes a two-stage (Global and Focal) architecture and a global-guided training strategy. The global stage obtains a continuous representation of the entire scene while the focal stage decomposes the scene into multiple blocks and further processes them with distinct sub-encoders. Leveraging this two-stage architecture, sub-encoders only need fine-tuning based on the global encoder, thus reducing training complexity in the focal stage while maintaining scene-wide consistency. Spatial information and error information from the global stage also benefit the sub-encoders to focus on crucial areas and effectively capture more details of large-scale scenes. No-tably, our approach does not rely on any prior knowledge about the target scene, attributing GF-NeRF adaptable to various large-scale scene types, including street-view and aerial-view scenes. We demonstrate that our method achieves high-fidelity, natural rendering results on various types of large-scale datasets. Our project page: https://shaomq2187.github.io/GF-NeRF/ Mingqi Shao, Mu Xu, Xueqian Wang 0001 |
WACV | 7 |
| 2025 | CCMA: A framework for cascading cooperative multi-agent in autonomous driving merging using Large Language Models
Miao Zhang 0010, Zhenlong Fang, Xueqian Wang 0001, Tianyu Shi 0003 |
Expert Syst. Appl. | 5 |
| 2025 | A Comprehensive Survey of Data Augmentation in Visual Reinforcement Learning
Guozheng Ma, Zhen Wang 0030, Zhecheng Yuan, Xueqian Wang 0001, Bo Yuan 0003, Dacheng Tao |
Int. J. Comput. Vis. | 4 |
| 2025 | OMR-diffusion: Optimizing multi-round enhanced training in diffusion models for improved intent understanding
Kun Li 0014, Jianhui Wang 0001, Yangfan He, Miao Zhang 0010, Xueqian Wang 0001 |
Neurocomputing | 5 |
| 2025 | SAGE: Self-evolving Agents with Reflective and Memory-augmented Abilities
Xuechen Liang, Meiling Tao, Yinghui Xia, Jianhui Wang 0001, Kun Li 0014, Yangfan He, Jingsong Yang, Tianyu Shi 0003, Yuantao Wang, Miao Zhang 0010, Xueqian Wang 0001 |
Neurocomputing | 12 |
| 2025 | PaddingFlow: Improving normalizing flows with padding-dimensional noise
Qinglong Meng, Chongkun Xia, Xueqian Wang 0001, Bin Liang 0001 |
Neurocomputing | 3 |
| 2025 | MDANet: A multi-stage domain adaptation framework for generalizable low-light image enhancement
Jianhui Wang 0001, Yangfan He, Kun Li 0014, Miao Zhang 0010, Tianyu Shi 0003, Xueqian Wang 0001 |
Neurocomputing | 9 |
| 2025 | Enhancing intent understanding for ambiguous prompt: A human-machine co-adaption strategy
Yangfan He, Jianhui Wang 0001, Kun Li 0014, Li Sun 0010, Miao Zhang 0010, Xueqian Wang 0001 |
Neurocomputing | 8 |
| 2025 | TSCnet: A text-driven semantic-level controllable framework for customized low-light image enhancement
Miao Zhang 0010, Pengyu Zeng, Yiqing Shen 0003, Xueqian Wang 0001 |
Neurocomputing | 6 |
| 2025 | A delay-robust method for enhanced real-time reinforcement learning
Bo Xia, Bo Yuan 0003, Zhiheng Li 0001, Bin Liang 0001, Xueqian Wang 0001 |
Neural Networks | 6 |
| 2025 | Toward the Flatter Landscape and Better Generalization in Federated Learning Under Client-Level Differential PrivacyabstractTo defend the inference attacks and mitigate the sensitive information leakages in Federated Learning (FL), client-level Differentially Private FL (DPFL) is the de-facto standard for privacy protection by clipping local updates and adding random noise. However, existing DPFL methods tend to make a sharp loss landscape and have poor weight perturbation robustness, resulting in severe performance degradation. To alleviate these issues, we propose a novel DPFL algorithm named DP-FedSAM, which leverages gradient perturbation to mitigate the negative impact of DP. Specifically, DP-FedSAM integrates Sharpness Aware Minimization (SAM) optimizer to generate local flatness models with improved stability and weight perturbation robustness, which results in the small norm of local updates and robustness to DP noise, thereby improving the performance. To further reduce the magnitude of random noise while achieving better performance, we propose DP-FedSAM-$\operatorname{top}_{k}$topk by adopting the local update sparsification technique. From the theoretical perspective, we present the convergence analysis to investigate how our algorithms mitigate the performance degradation induced by DP. Meanwhile, we give rigorous privacy guarantees with Rényi DP, the sensitivity analysis of local updates, and generalization analysis. At last, we empirically confirm that our algorithms achieve state-of-the-art (SOTA) performance compared with existing SOTA baselines in DPFL. Kang Wei 0004, Li Shen 0008, Yingqi Liu, Xueqian Wang 0001, Bo Yuan 0003, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | CVaR-Constrained Policy Optimization for Safe Reinforcement LearningabstractCurrent constrained reinforcement learning (RL) methods guarantee constraint satisfaction only in expectation, which is inadequate for safety-critical decision problems. Since a constraint satisfied in expectation remains a high probability of exceeding the cost threshold, solving constrained RL problems with high probabilities of satisfaction is critical for RL safety. In this work, we consider the safety criterion as a constraint on the conditional value-at-risk (CVaR) of cumulative costs, and propose the CVaR-constrained policy optimization algorithm (CVaR-CPO) to maximize the expected return while ensuring agents pay attention to the upper tail of constraint costs. According to the bound on the CVaR-related performance between two policies, we first reformulate the CVaR-constrained problem in augmented state space using the state extension procedure and the trust-region method. CVaR-CPO then derives the optimal update policy by applying the Lagrangian method to the constrained optimization problem. In addition, CVaR-CPO utilizes the distribution of constraint costs to provide an efficient quantile-based estimation of the CVaR-related value function. We conduct experiments on constrained control tasks to show that the proposed method can produce behaviors that satisfy safety constraints, and achieve comparable performance to most safe RL (SRL) methods. Shu Leng, Xiaoteng Ma, Qihan Liu, Xueqian Wang 0001, Bin Liang 0001, Yu Liu 0036, Jun Yang 0028 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Agent-Based Space Teleoperation: Mitigating Time Delays With Deep Reinforcement LearningabstractSpace teleoperation significantly extends human reach in space missions. However, traditional approaches are constrained by factors, such as the reliance on accurate dynamic models and the risk of operator fatigue during prolonged tasks. Additionally, while data-driven intelligent approaches reduce the need for prior knowledge, they have yet to adequately address the time delay issues inherent in these systems. To overcome these challenges, we introduce the belief state actor-critic (BSAC) method, the first deep reinforcement learning approach tailored for space teleoperation capture tasks within a bilateral control framework. We first establish a generalized agent-based architecture for space teleoperation, shifting decision-making from human operators to autonomous agents. Following a comprehensive analysis of the time delay challenges, we propose the BSAC algorithm, which integrates state augmentation and belief state techniques to mitigate the effects of delays in teleoperated Markov decision processes. Extensive experiments are conducted on the MuJoCo simulation platform, modeling a real hardware system across various scenarios. The learned policies are then successfully transferred and validated in a real-world setup, demonstrating the effectiveness and robustness of BSAC. In summary, our results support the feasibility of agent-based frameworks capable of overcoming time delay challenges in space teleoperation. Bo Xia, Xianru Tian, Bo Yuan 0003, Chunju Yang, Zhiheng Li 0001, Bin Liang 0001, Xueqian Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2024 | Decentralized Directed Collaboration for Personalized Federated LearningabstractPersonalized Federated Learning (PFL) is proposed to find the greatest personalized models for each client. To avoid the central failure and communication bottleneck in the server-based FL, we concentrate on the Decentralized Personalized Federated Learning (DPFL) that performs distributed model training in a Peer-to-Peer (P2P) manner. Most personalized works in DPFL are based on undi-rected and symmetric topologies, however, the data, computation and communication resources heterogeneity result in large variances in the personalized models, which lead the undirected aggregation to suboptimal personalized per-formance and unguaranteed convergence. To address these issues, we propose a directed collaboration DPFL framework by incorporating stochastic gradient push and partial model personalized, called Decentralized Federated Partial Gradient Push (DFedPGP). It personalizes the linear clas-sifier in the modern deep model to customize the local solution and learns a consensus representation in a fully de-centralized manner. Clients only share gradients with a subset of neighbors based on the directed and asymmetric topologies, which guarantees flexible choices for resource efficiency and better convergence. Theoretically, we show that the proposed DFedPGP achieves a superior conver-gence rate of O (1/√T) in the general non-convex setting, and prove the tighter connectivity among clients will speed up the convergence. The proposed method achieves state-of-the-art (SOTA) accuracy in both data and computation heterogeneity scenarios, demonstrating the efficiency of the directed collaboration and partial gradient push. Yingqi Liu, Baoyuan Wu, Qinglun Li, Xueqian Wang 0001, Li Shen 0008 |
CVPR | 5 |
| 2024 | Interpretable Data Fusion for Distributed Learning: A Representative Approach via Gradient MatchingabstractThis paper introduces a representative-based approach for distributed learning that transforms multiple raw data points into a virtual representation. Unlike traditional distributed learning methods such as Federated Learning, which do not offer human interpretability, our method makes complex machine learning processes accessible and comprehensible. It achieves this by condensing extensive datasets into digestible formats, thus fostering intuitive human-machine interactions. Additionally, this approach maintains privacy and communication efficiency, and it matches the training performance of models using raw data. Simulation results show that our approach is competitive with or outperforms traditional Federated Learning in accuracy and convergence, especially in scenarios with complex models and a higher number of clients. This framework marks a step forward in integrating human intuition with machine intelligence, which potentially enhances human-machine learning interfaces and collaborative efforts. Mengchen Fan, Baocheng Geng, Keren Li, Xueqian Wang 0001, Pramod K. Varshney |
FUSION | 4 |
| 2024 | Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training StagesabstractPlasticity, the ability of a neural network to evolve with new data, is crucial for high-performance and sample-efficient visual reinforcement learning (VRL). Although methods like resetting and regularization can potentially mitigate plasticity loss, the influences of various components within the VRL framework on the agent's plasticity are still poorly understood. In this work, we conduct a systematic empirical exploration focusing on three primary underexplored facets and derive the following insightful conclusions: (1) data augmentation is essential in maintaining plasticity; (2) the critic's plasticity loss serves as the principal bottleneck impeding efficient training; and (3) without timely intervention to recover critic's plasticity in the early stages, its loss becomes catastrophic. These insights suggest a novel strategy to address the high replay ratio (RR) dilemma, where exacerbated plasticity loss hinders the potential improvements of sample efficiency brought by increased reuse frequency. Rather than setting a static RR for the entire training process, we propose Adaptive RR, which dynamically adjusts the RR based on the critic’s plasticity level. Extensive evaluations indicate that Adaptive RR not only avoids catastrophic plasticity loss in the early stages but also benefits from more frequent reuse in later phases, resulting in superior sample efficiency. Guozheng Ma, Sen Zhang 0006, Zixuan Liu 0002, Zhen Wang 0030, Yixin Chen 0001, Li Shen 0008, Xueqian Wang 0001, Dacheng Tao |
ICLR | 8 |
| 2024 | Offline Goal-Conditioned Reinforcement Learning for Safety-Critical Tasks with Recovery PolicyabstractOffline goal-conditioned reinforcement learning (GCRL) aims at solving goal-reaching tasks with sparse rewards from an offline dataset. While prior work has demonstrated various approaches for agents to learn near-optimal policies, these methods encounter limitations when dealing with diverse constraints in complex environments, such as safety constraints. Some of these approaches prioritize goal attainment without considering safety, while others excessively focus on safety at the expense of training efficiency. In this paper, we study the problem of constrained offline GCRL and propose a new method called Recovery-based Supervised Learning (RbSL) to accomplish safety-critical tasks with various goals. To evaluate the method performance, we build a benchmark based on the robot-fetching environment with a randomly positioned obstacle and use expert or random policies to generate an offline dataset. We compare RbSL with three offline GCRL algorithms and one offline safe RL algorithm. As a result, our method outperforms the existing state-of-the-art methods to a large extent. Furthermore, we validate the practicality and effectiveness of RbSL by deploying it on a real Panda manipulator. Code is available at https://github.com/Sunlighted/RbSL.git. Zichen Yan, Renhao Lu, Junbo Tan, Xueqian Wang 0001 |
ICRA | 5 |
| 2024 | Learning Language-Conditioned Deformable Object Manipulation with Graph DynamicsabstractMulti-task learning of deformable object manipulation is a challenging problem in robot manipulation. Most previous works address this problem in a goal-conditioned way and adapt goal images to specify different tasks, which limits the multi-task learning performance and can not generalize to new tasks. Thus, we adapt language instruction to specify deformable object manipulation tasks and propose a learning framework. We first design a unified Transformer-based architecture to understand multi-modal data and output picking and placing action. Besides, we have applied the visible connectivity graph to tackle nonlinear dynamics and complex configuration of the deformable object. Both simulated and real experiments have demonstrated that the proposed method is effective and can generalize to unseen instructions and tasks. Compared with the state-of-the-art method, our method achieves higher success rates (87.2% on average) and has a 75.6% shorter inference time. We also demonstrate that our method performs well in real-world experiments. Supplementary videos can be found at https://sites.google.com/view/language-deformable. Yuhong Deng, Kai Mo, Chongkun Xia, Xueqian Wang 0001 |
ICRA | 4 |
| 2024 | Seamless Robot Teleoperation: Intuitive Control through Hand Gestures and Neural Network DecodingabstractRobotic teleoperation has enabled remote interaction with hazardous environments, overcoming spatial constraints on human perception and manipulation. Most teleoperation systems rely on task-dependent interfaces to generate human instructions. This can lead to barriers in familiarizing the robot’s workspace and thus increase the training time for less experienced users. In order to address these problems, we introduce a novel hand gestures based robot teleoperation method, eliminating the need for specialized controlling devices. Leveraging hand landmark detection and a neural network-based decoding algorithm, the system interprets hand movements to control robot velocity, offering a user-friendly solution to communicating with the robot. Our trained model achieves an F2 score of 0.994 and outperforms algorithms in the collected dataset. Furthermore, the proposed method has been validated on a real-world Franka robot, achieving success rates of 100%, 80%, and 86.7% across three manipulation tasks. Haolin Fei, Shijie Lee, Ziwei Wang 0001, Liucheng Guo, Darren Williams, Stefano Tedeschi 0002, Xueqian Wang 0001 |
IJCNN | 7 |
| 2024 | D3D: Conditional Diffusion Model for Decision-Making Under Random Frame DroppingabstractThe occurrence of frame drops due to issues such as corrupted communications or malfunctioning sensors presents a significant challenge to an agent’s decision-making, especially in remote control scenarios. Classical reinforcement learning (RL) usually assumes a continuous data stream without frame drops and relies heavily on online interactions, which is time-consuming, resource-intensive, and often impractical in certain scenarios. Consequently, the performance of RL may deteriorate significantly in face of non-negligible frame drops. To tackle this challenge caused by frame dropping, We propose Conditional Diffusion Model for Decision-Making under Random Frame Dropping (D3D), an offline algorithm that can effectively enhance performance robustness in frame dropping scenarios. D3D addresses this issue through a two-phase approach: 1) During the policy generation phase, D3D adopts a return-conditional diffusion model for decision making rather than the temporal difference learning, whose policy is derived using offline datasets of return-labeled trajectories without information loss. 2) When frame dropping occurs during evaluation, D3D seamlessly substitutes the missing state with its corresponding prediction in the horizon made by the diffusion model. Extensive experiments are conducted on MuJoCo and Adroit tasks to validate D3D’s robustness and efficiency. The results demonstrate that D3D consistently outperforms state-of-the-art RL algorithms, especially excelling on tasks featuring severe drop rates. Bo Xia, Yifu Luo, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001 |
RO-MAN | 6 |
| 2024 | Solving time-delay issues in reinforcement learning via transformers
Bo Xia, Zaihui Yang, Minzhi Xie, Yongzhe Chang, Bo Yuan 0003, Zhiheng Li 0001, Xueqian Wang 0001, Bin Liang 0001 |
Appl. Intell. | 7 |
| 2024 | Integrated Design for Active Fault Diagnosis and Control: A Decomposition-Composition MethodabstractActive fault diagnosis (AFD) designs inputs to excite the system to obtain more operation information for fault diagnosis. Since both AFD and control are required to design systems’ inputs, individually designing inputs for fault diagnosis will restrict the systems’ control performance. In order to achieve satisfactory performance for both AFD and control, this paper presents a novel set-based input design method for simultaneous AFD and control. The new method adopts a decomposition-composition idea to integrate AFD and control for better control performance during the AFD stage. Particularly, the proposed method first considers each single system mode separately and establishes an optimization problem, combining the AFD and control objectives, to design the optimal input. Then, all possible system modes are comprehensively considered to select the final input from all inputs designed under the single mode by an adaptive selection strategy. Meanwhile, the proposed method realizes AFD by maximizing the separation trend of all output sets online, without considering the strict set separation conditions. Hence, the method in this paper has a simple mathematical form and lower computational complexity. At the end, simulation results based on two different examples are presented to verify the effectiveness and portability of the proposed method.Note to Practitioners—Set-based AFD methods have two main features. The first one is that they only require the bounds of system uncertainties, such as modelling errors, process disturbances and measurement noise, and do not require their specific distributions and values. This requirement can be easily satisfied by varieties of engineering systems. The second one is that they design inputs to actively excite the system to obtain more system operation information for fault diagnosis. This feature enables AFD to achieve high fault diagnosis sensitivity and detect more faults, such as incipient faults, small faults, etc. Thus, when AFD methods are used in active fault-tolerant control (FTC) systems, the whole FTC performance can be improved. However, set-based AFD has conflicting requirements on input design with control, while there only exist few works in the literature on the integrated design of set-based AFD and control. This paper proposes a novel integrated design method for AFD and control, which can achieve satisfactory performance for both AFD and control with low computational complexity. Besides, although this paper only considers discrete linear time-invariant (LTI) systems, the proposed method can be extended to more complex systems such as linear parameter-varying (LPV) systems, linear time-varying systems, etc. Moreover, based on LPV modelling techniques, equilibrium linearization techniques, etc., the proposed method can be used for fault diagnosis and FTC of some nonlinear systems as well. Therefore, the proposed method has advantages in online fault diagnosis and FTC applications of engineering systems such as unmanned aerial vehicles, unmanned ships, driverless automobiles, space systems, robots, process industries, etc., which own values and potential to improve safety and reliability of varieties of engineering systems. Yushuai Wang, Feng Xu 0006, Juntian Qu, Xueqian Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Efficient Federated Learning With Enhanced Privacy via Lottery Ticket Pruning in Edge ComputingabstractFederated learning (FL) can train collaboratively with several mobile terminals (MTs), which faces critical challenges in communication, resource, and privacy. Existing privacy-preserving methods usually adopt instance-level differential privacy (DP), which provides a rigorous privacy guarantee but with several bottlenecks: performance degradation, transmission overhead, and resource constraints. Therefore, we propose Fed-LTP, an efficient and privacy-enhanced FL framework withLotteryTicketHypothesis (LTH) and zero-concentrated DP(zCDP). It generates a pruned global model on the server side and conducts sparse-to-sparse training from scratch with zCDP on the client side. On the server side, two pruning schemes are proposed: (i) the weight-based pruning (LTH) determines the pruned global model structure; (ii) the iterative pruning further shrinks the size of the pruned model. Meanwhile, the performance of Fed-LTP is boosted via model validation based on the Laplace mechanism. On the client side, we use sparse-to-sparse training to solve the resource-constraints issue and provide tighter privacy analysis to reduce the privacy budget. We evaluate the effectiveness of Fed-LTP on several real-world datasets in both independent and identically distributed (IID) and non-IID settings. The results confirm the superiority of Fed-LTP over state-of-the-art (SOTA) methods in communication, computation, and memory efficiencies while realizing a better utility-privacy trade-off. Kang Wei 0004, Li Shen 0008, Jun Li 0004, Xueqian Wang 0001, Bo Yuan 0003, Song Guo 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Polarimetric Inverse Rendering for Transparent Shapes ReconstructionabstractThe acquisition of transparent 3D shapes will facilitate many multimedia and computer vision tasks, such as game/movie production and virtual enrioment applications. In this work, we propose a novel method for detailed reconstruction of transparent objects by exploiting polarimetric cues. Most of existing transparent shapes reconstruction methods usually lack sufficient constraints and suffer from the over-smooth problem. Hence, we introduce polarization information as a complementary cue. Specifically, we employ the implicit representation for object's geometry with a neural network, while the polarization render is capable of differentiably rendering the object's polarization images from given illumination configuration. However, direct comparison of rendered polarization images to the real-world captured images will have additional errors due to the transmission in the transparent object. To make the polarimetric cues technically feasible on transparent shapes reconstruction, the concept of reflection percentage which represents proportion of the reflection component is introduced as the weight of the polarization loss. Based on controllable environment setup, we build a polarization dataset containing several solid and smooth transparent objects to verify our method. Experimental results show that our method is capable of recovering detailed shapes and improving reconstruction quality of transparent objects. Mingqi Shao, Chongkun Xia, Dongxu Duan, Xueqian Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Dynamics-Adaptive Continual Reinforcement Learning via Progressive ContextualizationabstractA key challenge of continual reinforcement learning (CRL) in dynamic environments is to promptly adapt the reinforcement learning (RL) agent's behavior as the environment changes over its lifetime while minimizing the catastrophic forgetting of the learned information. To address this challenge, in this article, we propose DaCoRL, that is, dynamics-adaptive continual RL. DaCoRL learns a context-conditioned policy using progressive contextualization, which incrementally clusters a stream of stationary tasks in the dynamic environment into a series of contexts and opts for an expandable multihead neural network to approximate the policy. Specifically, we define a set of tasks with similar dynamics as an environmental context and formalize context inference as a procedure of online Bayesian infinite Gaussian mixture clustering on environment features, resorting to online Bayesian inference to infer the posterior distribution over contexts. Under the assumption of a Chinese restaurant process (CRP) prior, this technique can accurately classify the current task as a previously seen context or instantiate a new context as needed without relying on any external indicator to signal environmental changes in advance. Furthermore, we employ an expandable multihead neural network whose output layer is synchronously expanded with the newly instantiated context and a knowledge distillation regularization term for retaining the performance on learned tasks. As a general framework that can be coupled with various deep RL algorithms, DaCoRL features consistent superiority over existing methods in terms of stability, overall performance, and generalization ability, as verified by extensive experiments on several robot navigation and MuJoCo locomotion tasks. Tiantian Zhang 0002, Zichuan Lin, Deheng Ye, Qiang Fu 0016, Wei Yang 0032, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003, Xiu Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Evaluating Model-Free Reinforcement Learning toward Safety-Critical TasksabstractSafety comes first in many real-world applications involving autonomous agents. Despite a large number of reinforcement learning (RL) methods focusing on safety-critical tasks, there is still a lack of high-quality evaluation of those algorithms that adheres to safety constraints at each decision step under complex and unknown dynamics. In this paper, we revisit prior work in this scope from the perspective of state-wise safe RL and categorize them as projection-based, recovery-based, and optimization-based approaches, respectively. Furthermore, we propose Unrolling Safety Layer (USL), a joint method that combines safety optimization and safety projection. This novel technique explicitly enforces hard constraints via the deep unrolling architecture and enjoys structural advantages in navigating the trade-off between reward improvement and constraint satisfaction. To facilitate further research in this area, we reproduce related algorithms in a unified pipeline and incorporate them into SafeRL-Kit, a toolkit that provides off-the-shelf interfaces and evaluation utilities for safety-critical tasks. We then perform a comparative study of the involved algorithms on six benchmarks ranging from robotic control to autonomous driving. The empirical results provide an insight into their applicability and robustness in learning zero-cost-return policies without task-dependent handcrafting. The project page is available at https://sites.google.com/view/saferlkit. Linrui Zhang, Li Shen 0008, Bo Yuan 0003, Xueqian Wang 0001, Dacheng Tao |
AAAI | 5 |
| 2023 | Make Landscape Flatter in Differentially Private Federated LearningabstractTo defend the inference attacks and mitigate the sensitive information leakages in Federated Learning (FL), clientlevel Differentially Private FL (DPFL) is the de-facto standard for privacy protection by clipping local updates and adding random noise. However, existing DPFL methods tend to make a sharper loss landscape and have poorer weight perturbation robustness, resulting in severe performance degradation. To alleviate these issues, we propose a novel DPFL algorithm named DP-FedSAM, which leverages gradient perturbation to mitigate the negative impact of DP. Specifically, DP-FedSAM integrates Sharpness Aware Minimization (SAM) optimizer to generate local flatness models with better stability and weight perturbation robustness, which results in the small norm of local updates and robustness to DP noise, thereby improving the performance. From the theoretical perspective, we analyze in detail how DP-FedSAM mitigates the performance degradation induced by DP. Meanwhile, we give rigorous privacy guarantees with Rényi DP and present the sensitivity analysis of local updates. At last, we empirically confirm that our algorithm achieves state-of-the-art (SOTA) performance compared with existing SOTA baselines in DPFL. Yingqi Liu, Kang Wei 0004, Li Shen 0008, Xueqian Wang 0001, Dacheng Tao |
CVPR | 5 |
| 2023 | Addressing Delays in Reinforcement Learning via Delayed Adversarial Imitation Learning
Minzhi Xie, Bo Xia, Yalou Yu, Xueqian Wang 0001, Yongzhe Chang |
ICANN (3) | 4 |
| 2023 | Volumetric 3D Reconstruction with Window-Wise Global Feature AggregationabstractVolumetric 3D reconstruction methods have shown great performance in reconstructing indoor scenarios from monocular videos. However, as such approaches utilize discrete feature voxels to encode the observed scenes, the global feature interaction within and across different voxels is ignored, leading to imperfect reconstructions. To solve this problem, we propose a novel volumetric 3D reconstruction method named VolGARecon. The core portion of VolGARecon includes two parts: first, we use an MLP-based weighted fusion module (WFM) to unproject the extracted features to each voxel, which considers the visibility and is capable to reduce the noise caused by occlusion; second, a 3D transformer module (3DTR) is used to perform window-wise global feature interaction in a local sliding window, which strengthens the feature expression in 3D space and benefits estimating more complete and spatially coherent 3D models. In addition, we propose a multi-dimensional hybrid loss (MHL) that incorporates the 3D supervision in classical volumetric methods and the 2D supervision in novel view synthesis works. Extensive experiments show our method achieves superior performance on multiple datasets. Shihao Ren, Yikang Ding, Jinli Liao, Xinghui Li, Wensen Feng, Xueqian Wang 0001 |
ICASSP | 7 |
| 2023 | Transparent Shape from a Single View Polarization ImageabstractThis paper presents a learning-based method for transparent surface estimation from a single view polarization image. Existing shape from polarization(SfP) methods have the difficulty in estimating transparent shape since the inherent transmission interference heavily reduces the reliability of physics-based prior. To address this challenge, we propose the concept of physics-based prior confidence, which is inspired by the characteristic that the transmission component in the polarization image has more noise than reflection. The confidence is used to determine the contribution of the interfered physics-based prior. Then, we build a network(TransSfP) with multi-branch architecture to avoid the destruction of relationships between different hierarchical inputs. To train and test our method, we construct a dataset for transparent shape from polarization with paired polarization images and ground-truth normal maps. Extensive experiments and comparisons demonstrate the superior accuracy of our method. Our cdataset and code are publicly available at https://github.com/shaomq2187/TransSfP Mingqi Shao, Chongkun Xia, Zhendong Yang, Junnan Huang, Xueqian Wang 0001 |
ICCV | 5 |
| 2023 | Curriculum-based Co-design of Morphology and Control of Voxel-based Soft Robots
Haobo Fu, Qiang Fu 0016, Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001 |
ICLR | 7 |
| 2023 | Improving the Model Consistency of Decentralized Federated LearningabstractTo mitigate the privacy leakages and communication burdens of Federated Learning (FL), decentralized FL (DFL) discards the central server and each client only communicates with its neighbors in a decentralized communication network. However, existing DFL suffers from high inconsistency among local clients, which results in severe distribution shift and inferior performance compared with centralized FL (CFL), especially on heterogeneous data or sparse communication topologies. To alleviate this issue, we propose two DFL algorithms named DFedSAM and DFedSAM-MGS to improve the performance of DFL. Specifically, DFedSAM leverages gradient perturbation to generate local flat models via Sharpness Aware Minimization (SAM), which searches for models with uniformly low loss values. DFedSAM-MGS further boosts DFedSAM by adopting Multiple Gossip Steps (MGS) for better model consistency, which accelerates the aggregation of local flat models and better balances communication complexity and generalization. Theoretically, we present improved convergence rates $\small \mathcal{O}\big(\frac{1}{\sqrt{KT}}+\frac{1}{T}+\frac{1}{K^{1/2}T^{3/2}(1-\lambda)^2}\big)$ and $\small \mathcal{O}\big(\frac{1}{\sqrt{KT}}+\frac{1}{T}+\frac{\lambda^Q+1}{K^{1/2}T^{3/2}(1-\lambda^Q)^2}\big)$ in non-convex setting for DFedSAM and DFedSAM-MGS, respectively, where $1-\lambda$ is the spectral gap of gossip matrix and $Q$ is the number of MGS. Empirically, our methods can achieve competitive performance compared with CFL methods and outperform existing DFL methods. Li Shen 0008, Kang Wei 0004, Bo Yuan 0003, Xueqian Wang 0001, Dacheng Tao |
ICML | 6 |
| 2023 | Quadruped Guidance Robot for the Visually Impaired: A Comfort-Based ApproachabstractGuidance robots that can guide people and avoid various obstacles, could potentially be owned by more visually impaired people at a fairly low cost. Most of the previous guidance robots for the visually impaired ignored the human response behavior and comfort, treating the human as an appendage dragged by the robot, which can lead to imprecise guidance of the human and sudden changes in the traction force experienced by the human. In this paper, we propose a novel quadruped guidance robot system with a comfort-based concept. We design a controllable traction device that can adjust the length and force between human and robot to ensure comfort. To allow the human to be guided safely and comfortably to the target position in complex environments, our proposed human motion planner can plan the traction force with the force-based human motion model. To track the planned force, we also propose a robot motion planner that can generate the specific robot motion command and design the force control device. Our system has been deployed on Unitree Laikago quadrupedal platform and validated in real-world scenarios. (Video11Video demonstration: https://youtu.be/gd-RcYOqGuo.) Yanbo Chen 0001, Zhengzhe Xu, Zhuozhu Jian, Gengpan Tang, Liyunong Yang, Anxing Xiao, Xueqian Wang 0001, Bin Liang 0001 |
ICRA | 7 |
| 2023 | Dynamic Control Barrier Function-based Model Predictive Control to Safety-Critical Obstacle-Avoidance of Mobile RobotabstractThis paper presents an efficient and safe method to avoid static and dynamic obstacles based on LiDAR. First, point cloud is used to generate a real-time local grid map for obstacle detection. Then, obstacles are clustered by DBSCAN algorithm and enclosed with minimum bounding ellipses (MBEs). In addition, data association is conducted to match each MBE with the obstacle in the current frame. Considering MBE as an observation, Kalman filter (KF) is used to estimate and predict the motion state of the obstacle. In this way, the trajectory of each obstacle in the forward time domain can be parameterized as a set of ellipses. Due to the uncertainty of the MBE, the semi-major and semi-minor axes of the parameterized ellipse are extended to ensure safety. We extend the traditional Control Barrier Function (CBF) and propose Dynamic Control Barrier Function (D-CBF). We combine D-CBF with Model Predictive Control (MPC) to implement safety-critical dynamic obstacle avoidance. Experiments in simulated and real scenarios are conducted to verify the effectiveness of our algorithm. The source code is released for the reference of the community11Code: https://github.com/jianzhuozhuTHU/MPC-D-CBF.. Zhuozhu Jian, Zihong Yan, Xuanang Lei, Zihong Lu, Bin Lan, Xueqian Wang 0001, Bin Liang 0001 |
ICRA | 6 |
| 2023 | USEEK: Unsupervised SE(3)-Equivariant 3D Keypoints for Generalizable ManipulationabstractCan a robot manipulate intra-category unseen objects in arbitrary poses with the help of a mere demonstration of grasping pose on a single object instance? In this paper, we try to address this intriguing challenge by using USEEK, an unsupervised SE(3)-equivariant keypoints method that enjoys alignment across instances in a category, to perform generaliz-able manipulation. USEEK follows a teacher-student structure to decouple the unsupervised keypoint discovery and SE(3)-equivariant keypoint detection. With USEEK in hand, the robot can infer the category-level task-relevant object frames in an efficient and explainable manner, enabling manipulation of any intra-category objects from and to any poses. Through extensive experiments, we demonstrate that the keypoints produced by USEEK possess rich semantics, thus successfully transferring the functional knowledge from the demonstration object to the novel ones. Compared with other object representations for manipulation, USEEK is more adaptive in the face of large intra-category shape variance, more robust with limited demonstrations, and more efficient at inference time. Project website: https://sites.google.com/view/useek/. Zhengrong Xue, Zhecheng Yuan, Jiashun Wang, Xueqian Wang 0001, Yang Gao 0029, Huazhe Xu |
ICRA | 4 |
| 2023 | Extended PID Controller for Nonminimum Phase Systems with Application to a Hypersonic VehicleabstractIn our previous work, we proposed the extended PID (EPID) controller, which is a state-space extension of traditional PID control. Compared to PID control, EPID is more suitable for multi-input-multi-output (MIMO) and higher-order systems. In this paper, we further extend EPID to nonminimum phase systems and investigate its performance limitation. EPID uses feedback of all state tracking errors. But for nonminimum phase systems, the reference trajectories for the internal states are unknown (assuming we do not have the system model), making us decide to take out the internal states from the integral part to avoid an unbounded input, which results in a slightly different controller form. Besides, in previous study, we found an important property of EPID is that it can achieve accurate tracking/rejecting for time-varying references/disturbances by using a high integral gain. However, when applied to nonminimum phase systems, we found that the integral gain cannot be set too high, otherwise the closed-loop system will be unstable, which indicates an inherent performance limitation. To verify this, simulation results are provided by applying EPID to a hypersonic vehicle model and a cart pole system. Linqi Ye, Xueqian Wang 0001, Bin Liang 0001 |
IECON | 2 |
| 2023 | Uncertainty-Aware Data Augmentation for Offline Reinforcement LearningabstractOne of the key challenges in Offline Reinforcement Learning is that it cannot conduct further environment exploration and performs poorly in terms of out-of-distribution generalizations. Data augmentation is commonly used to solve the issue of limited coverage of the full state-action space in static offline dataset. However, the existing data augmentation methods for proprioceptive observation suffer from the dilemma where the data coverage is often limited by tight constraints, while aggressive methods may exacerbate the performance. At the heart of this phenomenon are the diverged action distribution and the high uncertainty of the value function. In this paper, we propose to extend the static offline datasets during training by adding gradient-based perturbation to the state and utilizing the estimated uncertainty of the value function to constrain the range of the gradient. The estimated uncertainty of the value function works as a guidance to adjust the range of augmentation automatically, ensuring the adaptability and reliability of the state perturbation. The proposed algorithm Uncertainty-Aware Data Augmentation(UADA), is plugged into various standard offline RL algorithms and evaluated on several offline rein-forcement learning tasks. The empirical results confirm that UADA substantially improves the performance and achieves better model stability compared with the original algorithms. Yunjie Su, Yilun Kong, Xueqian Wang 0001 |
IJCNN | 3 |
| 2023 | Visuotactile Sensor Enabled Pneumatic Device Towards Compliant Oropharyngeal Swab SamplingabstractManual oropharyngeal (OP) swab sampling is an intensive and risky task. In this article, a novel OP swab sampling device of low cost and high compliance is designed by combining the visuotactile sensor and the pneumatic actuator-based gripper. Here, a concave visuotactile sensor called CoTac is first proposed to address the problems of high cost and poor reliability of traditional multi-axis force sensors. Besides, by imitating the doctor's fingers, a soft pneumatic actuator with a rigid skeleton structure is designed, which is demonstrated to be reliable and safe via finite element modeling and experiments. Furthermore, we propose a sampling method that adopts a compliant control algorithm based on the adaptive virtual force to enhance the safety and compliance of the swab sampling process. The effectiveness of the device has been verified through sampling experiments as well as in vivo tests, indicating great application potential. The cost of the device is around 30 US dollars and the total weight of the functional part is less than 0.1 kg, allowing the device to be rapidly deployed on various robotic arms. Shoujie Li, Mingshan He, Wenbo Ding 0001, Linqi Ye, Xueqian Wang 0001, Junbo Tan, Jinqiu Yuan, Xiao-Ping Zhang 0002 |
IROS | 5 |
| 2023 | Learning Better with Less: Effective Augmentation for Sample-Efficient Visual Reinforcement LearningabstractData augmentation (DA) is a crucial technique for enhancing the sample efficiency of visual reinforcement learning (RL) algorithms.
Notably, employing simple observation transformations alone can yield outstanding performance without extra auxiliary representation tasks or pre-trained encoders. However, it remains unclear which attributes of DA account for its effectiveness in achieving sample-efficient visual RL. To investigate this issue and further explore the potential of DA, this work conducts comprehensive experiments to assess the impact of DA's attributes on its efficacy and provides the following insights and improvements: (1) For individual DA operations, we reveal that both ample spatial diversity and slight hardness are indispensable. Building on this finding, we introduce Random PadResize (Rand PR), a new DA operation that offers abundant spatial diversity with minimal hardness. (2) For multi-type DA fusion schemes, the increased DA hardness and unstable data distribution result in the current fusion schemes being unable to achieve higher sample efficiency than their corresponding individual operations. Taking the non-stationary nature of RL into account, we propose a RL-tailored multi-type DA fusion scheme called Cycling Augmentation (CycAug), which performs periodic cycles of different DA operations to increase type diversity while maintaining data distribution consistency. Extensive evaluations on the DeepMind Control suite and CARLA driving simulator demonstrate that our methods achieve superior sample efficiency compared with the prior state-of-the-art methods. Guozheng Ma, Linrui Zhang, Haoyu Wang 0018, Zilin Wang 0002, Zhen Wang 0030, Li Shen 0008, Xueqian Wang 0001, Dacheng Tao |
NeurIPS | 8 |
| 2023 | Overcoming Delayed Feedback via Overlook Decision MakingabstractReinforcement learning is one of the most general paradigms to solve sequential decision making issues on the assumption that the action selection and environmental feedback are instantaneous, however, unfortunately this assumption is rarely true with regard to such ubiquitous delays in real-world system which could degrade the performance of reinforcement learning algorithms. The most common solution to solve a fixed delay problem is to design a forward dynamic model which is used to predict the newest state by recursively iterating over long steps so that a predicted state can be got and it would be taken as the agent's observation to make the newest decision. However, there exists cumulative errors during the iterative process which make long-term prediction inaccurate and further affect agent's decision. Motivated by the goal to reduce cumulative errors, we propose a new algorithm named Multi-step Prediction model with Delayed Observation(MPDO), aiming at accurately predicting future state at longer horizons for better decision making. Our approach includes two parts: a multi-step prediction model and a strategy training based on proximal policy optimization algorithms(PPO). Our model only needs a small amount of data to conduct dynamic modeling quickly, and the accuracy of prediction and iteration speed are higher than traditional methods. Experiments on Gym and MuJoCo show that MPDO achieves higher performance in such different tasks with different delays compared with other state-of-the-art methods, which verify our method's effectiveness. Yalou Yu, Bo Xia, Minzhi Xie, Xueqian Wang 0001, Zhiheng Li 0001, Yongzhe Chang |
SMC | 4 |
| 2023 | EPO-S: A Constrained RL Method to Enhance UAV Safety with Spatial RepresentationabstractPath planning and collision avoidance are critical components of UAV control algorithms that play a crucial role in executing UAV missions. As scenarios become increasingly complex, the traditional control methods just ain't cutting it to meet the requirements. Reinforcement learning is an emerging decision-making control algorithm that attempts to address these issues as an alternative to traditional methods and has made significant advances. Unfortunately, standard RL approaches only aim to maximize rewards, however balancing task performance and safety in completing UAV tasks poses a challenge since these two objectives sometimes conflict, leading to a trade-off often difficult to manage. This paper proposes three techniques to address this problem. First, we model the path planning and collision avoidance issue in a constrained RL framework, eliminating the need for complex reward engineering. Second, we expand our previous work in the UAV setting and introduce an exact penalty optimization (EPO) algorithm to provide stricter constraint guarantees. We also propose a novel spatial information representation method for the UAV scenario to help UAVs better understand environmental information. The experimental results demonstrate the effectiveness of the EPO and spatial representation modules proposed in this paper, through a significant reduction in collisions as well as a strong improvement in the rate of reaching the destination. Linrui Zhang, Zaihui Yang, Haoyu Wang 0018, Xueqian Wang 0001, Yongzhe Chang |
SMC | 5 |
| 2023 | Catastrophic Interference in Reinforcement Learning: A Solution Based on Context Division and Knowledge DistillationabstractThe powerful learning ability of deep neural networks enables reinforcement learning (RL) agents to learn competent control policies directly from continuous environments. In theory, to achieve stable performance, neural networks assume identically and independently distributed (i.i.d.) inputs, which unfortunately does not hold in the general RL paradigm where the training data are temporally correlated and nonstationary. This issue may lead to the phenomenon of "catastrophic interference" and the collapse in performance. In this article, we present interference-aware deep Q-learning (IQ) to mitigate catastrophic interference in single-task deep RL. Specifically, we resort to online clustering to achieve on-the-fly context division, together with a multihead network and a knowledge distillation regularization term for preserving the policy of learned contexts. Built upon deep Q networks (DQNs), IQ consistently boosts the stability and performance when compared to existing methods, verified with extensive experiments on classic control and Atari tasks. The code is publicly available at https://github.com/ Sweety-dm/Interference-aware-Deep-Q-learning. Tiantian Zhang 0002, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Optimization Design Method of Tendon-Sheath Transmission Path Under Curvature ConstraintabstractThe application requirements of the tendon-sheath mechanism in the field of precision machinery are becoming increasingly extensive. However, the contact friction between the tendon and sheath seriously affects the transmission accuracy. In the case of unavoidable friction, optimizing the tendon transmission path to reduce tension loss and elastic deformation has become an important research direction. In this article, the influence law of the tendon transmission path on the tension and displacement transmission is obtained using the two parameters related to the curvature of the transmission path: total bending angle and equivalent tendon length. Then, based on the optimal control theory and minimum principle, the different transmission path solutions of the minimum tension loss, the minimum tendon deformation, and the coupling of tension and displacement are obtained; the numerical optimization method verifies the correctness of the proposed theory. Finally, an optimal design of a tendon-constrained synchronous rotation mechanism for the manipulator is carried out, and the linkage performance is greatly improved by optimizing the transmission path. Weining Lu, Yu Liu 0036, Deshan Meng, Xueqian Wang 0001, Bin Liang 0001 |
IEEE Trans. Robotics | 5 |
| 2023 | Visual-Tactile Fusion for Transparent Object Grasping in Complex BackgroundsabstractThe grasping of transparent objects is challenging but of significance to robots. In this article, a visual–tactile fusion framework for transparent object grasping in complex backgrounds is proposed, which synergizes the advantages of vision and touch, and greatly improves the grasping efficiency of transparent objects. First, we propose a multiscene synthetic grasping dataset named SimTrans12 K together with a Gaussian-mask annotation method. Next, based on the TaTa gripper, we propose a grasping network named transparent object-grasping convolutional neural network for grasping position detection, which shows good performance in both synthetic and real scenes. Inspired by human grasping, a tactile calibration method and a visual–tactile fusion classification method are designed, which improve the grasping success rate by 36.7% compared with direct grasping and the classification accuracy by 39.1%. Furthermore, the tactile height sensing module and the tactile position exploration module are added to solve the problem of grasping transparent objects in irregular and visually undetectable scenes. The experimental results demonstrate the validity of the framework. Shoujie Li, Haixin Yu, Wenbo Ding 0001, Houde Liu, Linqi Ye, Chongkun Xia, Xueqian Wang 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Robotics | 7 |
| 2022 | Orientation to Pose: Continuum Robots Shape Reconstruction Based on the Multi-Attitude Solving ApproachabstractContinuum robots are typically slender and flexible with infinite freedoms in theory, which poses a challenge for their control and application. The shape reconstruction of continuum robots is vital to realize closed-loop control. This paper proposes a novel general real-time shape reconstruction framework of continuum robots based on the piecewise polynomial curvature (PPC) kinematics model. We illustrate the coupling between orientation and position at any given location of the continuum robots. Further, the coupling relation could be bridged by the PPC kinematics. Therefore, we propose to estimate the shape through multi-attitude solving, using the off-the-shelf orientation sensors, e.g., IMUs, mounted on certain locations. The approach gives a valuable framework to real-time shape reconstruction of continuum robots, which is general, accurate and convenient. The accuracy of our approach is verified in the experiments of distinct physical prototypes. Hejie Xu, Hongji Shang, Xueqian Wang 0001, Houde Liu, Bin Liang 0001 |
ICRA | 4 |
| 2022 | TaTa: A Universal Jamming Gripper with High-Quality Tactile Perception and Its Application to Underwater ManipulationabstractLarge-area and high-precision tactile sensing information can not only improve the stability of robot grasping but also compensate for the lack of visual information in specific environments such as turbid underwater, dimness, and smoke. In this paper, we devise a universal jamming gripper with high-quality tactile sensing capability. The gripper adopts the particle jamming mechanism for grasping, and simultaneously uses a built-in camera to detect the deformation of its surface to obtain tactile information. To make the inside of the gripper transparent, glass beads and liquid with the same refractive index are applied as the internal filling. Besides, special treatments are taken to improve the tactile perception resolution of the gripper. The design perfectly merges visual-based tactile sensing into the traditional universal jamming gripper without changing its original gripping performance, making it possible for simultaneous grasping and sensing. To verify the tactile perception and grasping ability of the gripper in specific environments, we design two underwater experiments for grasping and pipe leak detection based on tactile information. Both have achieved a success rate not less than 95%, which demonstrates the effectiveness of the proposed gripper for manipulation in low visibility environments. Shoujie Li, Xianghui Yin, Chongkun Xia, Linqi Ye, Xueqian Wang 0001, Bin Liang 0001 |
ICRA | 5 |
| 2022 | Don't Touch What Matters: Task-Aware Lipschitz Data Augmentation for Visual Reinforcement LearningabstractOne of the key challenges in visual Reinforcement Learning (RL) is to learn policies that can generalize to unseen environments. Recently, data augmentation techniques aiming at enhancing data diversity have demonstrated proven performance in improving the generalization ability of learned policies. However, due to the sensitivity of RL training, naively applying data augmentation, which transforms each pixel in a task-agnostic manner, may suffer from instability and damage the sample efficiency, thus further exacerbating the generalization performance. At the heart of this phenomenon is the diverged action distribution and high-variance value estimation in the face of augmented images. To alleviate this issue, we propose Task-aware Lipschitz Data Augmentation (TLDA) for visual RL, which explicitly identifies the task-correlated pixels with large Lipschitz constants, and only augments the task-irrelevant pixels for stability. We verify the effectiveness of our approach on DeepMind Control suite, CARLA and DeepMind Manipulation tasks. The extensive empirical results show that TLDA improves both sample efficiency and generalization; it outperforms previous state-of-the-art methods across 3 different visual control benchmarks. Zhecheng Yuan, Guozheng Ma, Yao Mu 0001, Bo Xia, Bo Yuan 0003, Xueqian Wang 0001, Ping Luo 0002, Huazhe Xu |
IJCAI | 6 |
| 2022 | Penalized Proximal Policy Optimization for Safe Reinforcement LearningabstractSafe reinforcement learning aims to learn the optimal policy while satisfying safety constraints, which is essential in real-world applications. However, current algorithms still struggle for efficient policy updates with hard constraint satisfaction. In this paper, we propose Penalized Proximal Policy Optimization (P3O), which solves the cumbersome constrained policy iteration via a single minimization of an equivalent unconstrained problem. Specifically, P3O utilizes a simple yet effective penalty approach to eliminate cost constraints and removes the trust-region constraint by the clipped surrogate objective. We theoretically prove the exactness of the penalized method with a finite penalty factor and provide a worst-case analysis for approximate error when evaluated on sample trajectories. Moreover, we extend P3O to more challenging multi-constraint and multi-agent scenarios which are less studied in previous work. Extensive experiments show that P3O outperforms state-of-the-art algorithms with respect to both reward improvement and constraint satisfaction on a set of constrained locomotive tasks. Linrui Zhang, Li Shen 0008, Long Yang 0004, Shixiang Chen, Xueqian Wang 0001, Bo Yuan 0003, Dacheng Tao |
IJCAI | 5 |
| 2022 | Deep Reinforcement Learning Based on Local GNN for Goal-Conditioned Deformable Object RearrangingabstractObject rearranging is one of the most common deformable manipulation tasks, where the robot needs to rearrange a deformable object into a goal configuration. Previous studies focus on designing an expert system for each specific task by model-based or data-driven approaches and the application scenarios are therefore limited. Some research has been attempting to design a general framework to obtain more advanced manipulation capabilities for deformable rearranging tasks, with lots of progress achieved in simulation. However, transferring from simulation to reality is difficult due to the limitation of the end-to-end CNN architecture. To address these challenges, we design a local GNN (Graph Neural Network) based learning method, which utilizes two representation graphs to encode keypoints detected from images. Self-attention is applied for graph updating and cross-attention is applied for generating manipulation actions. Extensive experiments have been conducted to demonstrate that our framework is effective in multiple 1-D (rope, rope ring) and 2-D (cloth) rearranging tasks in simulation and can be easily transferred to a real robot by fine-tuning a keypoint detector. Yuhong Deng, Chongkun Xia, Xueqian Wang 0001, Lipeng Chen |
IROS | 3 |
| 2022 | PUTN: A Plane-fitting based Uneven Terrain Navigation FrameworkabstractAutonomous navigation of ground robots has been widely used in indoor structured 2D environments, but there are still many challenges in outdoor 3D unstructured environments, especially in rough, uneven terrains. This paper proposed a plane-fitting based uneven terrain navigation framework (PUTN) to solve this problem. The implementation of PUTN is divided into three steps. First, based on Rapidly-exploring Random Trees (RRT), an improved sample-based algorithm called Plane Fitting RRT*(PF- RRT*) is proposed to obtain a sparse trajectory. Each sampling point corresponds to a custom traversability index and a fitted plane on the point cloud. These planes are connected in series to form a traversable “strip”. Second, Gaussian Process Regression is used to generate traversability of the dense trajectory interpolated from the sparse trajectory, and the sampling tree is used as the training set. Finally, local planning is performed using nonlinear model predictive control (NMPC). By adding the traversability index and uncertainty to the cost function, and adding obstacles generated by the real-time point cloud to the constraint function, a safe motion planning algorithm with smooth speed and strong robustness is available. Experiments in real scenarios are conducted to verify the effectiveness of the method. The source code is released for the reference of the community11Source code: https://github.com/jianzhuozhuTHU/putn.. Zhuozhu Jian, Zihong Lu, Bin Lan, Anxing Xiao, Xueqian Wang 0001, Bin Liang 0001 |
IROS | 6 |
| 2022 | Safety Correction from Baseline: Towards the Risk-aware Policy in Robotics via Dual-agent Reinforcement LearningabstractLearning a risk-aware policy is essential but rather challenging in unstructured robotic tasks. Safe reinforcement learning methods open up new possibilities to tackle this problem. However, the conservative policy updates make it intractable to achieve sufficient exploration and desirable performance in complex, sample-expensive environments. In this paper, we propose a dual-agent safe reinforcement learning strategy consisting of a baseline and a safe agent. Such a decoupled framework enables high flexibility, data efficiency and risk-awareness for RL-based control. Concretely, the baseline agent is responsible for maximizing rewards under standard RL settings. Thus, it is compatible with off-the-shelf training techniques of unconstrained optimization, exploration and exploitation. On the other hand, the safe agent mimics the baseline agent for policy improvement and learns to fulfill safety constraints via off-policy RL tuning. In contrast to training from scratch, safe policy correction requires significantly fewer interactions to obtain a near-optimal policy. The dual policies can be optimized synchronously via a shared replay buffer, or leveraging the pre-trained model or the non-learning-based controller as a fixed baseline agent. Experimental results show that our approach can learn feasible skills without prior knowledge as well as deriving risk-averse counterparts from pre-trained unsafe policies. The proposed method outperforms the state-of-the-art safe RL algorithms on difficult robot locomotion and manipulation tasks with respect to both safety constraint satisfaction and sample efficiency. Linrui Zhang, Zichen Yan, Li Shen 0008, Shoujie Li, Xueqian Wang 0001, Dacheng Tao |
IROS | 5 |
| 2022 | Pre-Trained Image Encoder for Generalizable Visual Reinforcement LearningabstractLearning generalizable policies that can adapt to unseen environments remains challenging in visual Reinforcement Learning (RL). Existing approaches try to acquire a robust representation via diversifying the appearances of in-domain observations for better generalization. Limited by the specific observations of the environment, these methods ignore the possibility of exploring diverse real-world image datasets. In this paper, we investigate how a visual RL agent would benefit from the off-the-shelf visual representations. Surprisingly, we find that the early layers in an ImageNet pre-trained ResNet model could provide rather generalizable representations for visual RL. Hence, we propose Pre-trained Image Encoder for Generalizable visual reinforcement learning (PIE-G), a simple yet effective framework that can generalize to the unseen visual scenarios in a zero-shot manner. Extensive experiments are conducted on DMControl Generalization Benchmark, DMControl Manipulation Tasks, Drawer World, and CARLA to verify the effectiveness of PIE-G. Empirical evidence suggests PIE-G improves sample efficiency and significantly outperforms previous state-of-the-art methods in terms of generalization performance. In particular, PIE-G boasts a 55% generalization performance gain on average in the challenging video background setting. Project Page: https://sites.google.com/view/pie-g/home. Zhecheng Yuan, Zhengrong Xue, Xueqian Wang 0001, Yi Wu 0013, Yang Gao 0029, Huazhe Xu |
NeurIPS | 4 |
| 2022 | Input Enhanced Logarithmic Factorization Network for CTR Prediction
Xianzhuang Li, Zhen Wang 0030, Xuesong Wu 0003, Bo Yuan 0003, Xueqian Wang 0001 |
PAKDD (3) | 5 |
| 2022 | Value Penalized Q-Learning for Recommender SystemsabstractScaling reinforcement learning (RL) to recommender systems (RS) is promising since maximizing the expected cumulative rewards for RL agents meets the objective of RS, i.e., improving customers' long-term satisfaction. A key approach to this goal is offline RL, which aims to learn policies from logged data rather than expensive online interactions. In this paper, we propose Value Penalized Q-learning (VPQ), a novel uncertainty-based offline RL algorithm that penalizes the unstable Q-values in the regression target using uncertainty-aware weights, achieving the conservative Q-function without the need of estimating the behavior policy, suitable for RS with a large number of items. Experiments on two real-world datasets show the proposed method serves as a gain plug-in for existing RS models. Chengqian Gao, Ke Xu 0002, Kuangqi Zhou, Lanqing Li, Xueqian Wang 0001, Bo Yuan 0008, Peilin Zhao |
SIGIR | 5 |
| 2022 | Graph-Transporter: A Graph-based Learning Method for Goal-Conditioned Deformable Object Rearranging TaskabstractRearranging deformable objects is a long-standing challenge in robotic manipulation for the high dimensionality of configuration space and the complex dynamics of deformable objects. We present a novel framework, Graph-Transporter, for goal-conditioned deformable object rearranging tasks. To tackle the challenge of complex configuration space and dynamics, we represent the configuration space of a deformable object with a graph structure and the graph features are encoded by a graph convolution network. Our framework adopts an architecture based on Fully Convolutional Network (FCN) to output pixel-wise pick-and-place actions from only visual input. Extensive experiments have been conducted to validate the effectiveness of the graph representation of deformable object configuration. The experimental results also demonstrate that our framework is effective and general in handling goal-conditioned deformable object rearranging tasks. Yuhong Deng, Chongkun Xia, Xueqian Wang 0001, Lipeng Chen |
SMC | 3 |
| 2022 | A surrogate-assisted controller for expensive evolutionary reinforcement learning
Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003 |
Inf. Sci. | 4 |
| 2022 | A Multimodal Fusion Fatigue Driving Detection Method Based on Heart Rate and PERCLOSabstractExisting visual-based fatigue detection methods usually monitor drivers’ fatigue by capturing their facial features, including eyelid movements, yawn frequency and head pose. However, these approaches typically do not take drivers’ biological signals into consideration. An accurate model for fatigue detection requires combining both facial behavior and biological data. This paper proposes a novel non-intrusive method for driver multimodal fusion fatigue detection by extracting eyelid features and heart rate signals from the RGB video. The multimodal feature fusion method could significantly increase the accuracy of fatigue detection. Specifically, we established two fatigue detection models based on heart rate and the PERCLOS value respectively with one-dimensional Convolutional Neural Network (1D CNN), where the PERCLOS refers to the percentage of eyelid closure over the pupil. Finally, the outputs of the two models are weighted to achieve the multimodal fusion fatigue detection. Simulation results show that our method yield better performance than traditional methods. Guanglong Du, Linlin Zhang 0011, Kang Su, Xueqian Wang 0001, Shaohua Teng, Peter Xiaoping Liu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Soft-CCD Algorithm for Inverse Kinematics of Soft Continuum ManipulatorsabstractTo date, soft robots have been increasingly designed and analyzed, especially, Soft Continuum Manipulators (SCMs). Due to dexterous deformability, their Inverse Kinematics (IK) is still difficult to solve. Cyclic Coordinate Descent (CCD) algorithm is one of the classical optimization algorithms to solve IK of rigid manipulators with prismatic or rotational joints. However, it cannot be directly extrapolated to SCMs with configuration space parameters such as center arc, bending angle, and torsion angle. Here, we modified the CCD algorithm from a new view and proposed several tricks to set constraints. Numerical and experimental results show that the soft-CCD algorithm can quickly and accurately generate solutions for IK of SCMs. This study provides the necessary kinematic foundation for fault tolerance, obstacle avoidance, trajectory planning, and other further explorations of SCMs. Deshan Meng, Xueqian Wang 0001, Bin Liang 0001 |
IROS | 4 |
| 2021 | Design of a Tactile Sensing Robotic Gripper and Its Grasping MethodabstractAlthough computer vision has the advantages of long detection distance and large amount of information, it also has certain limitations for complex scenes such as dimness, reflections, and smoke. In order to solve the problem of robot grasping in these scenes, we designed a novel gripper that can search, identify and grasp objects based on tactile information. The gripper can effectively grasp the objects in real life, and can sense the shape and posture of the objects through the touch. We proposed a lifting finger structure that allows the gripper to switch between sensing and grasping modes. We applied visual-tactile detection methods to obtain tactile information and propose a feature extraction algorithm based on U-net. We designed a method of grasping the center of mass of the object contour, and the success rate of the grasping can reach 85%. In addition, we also designed experiments to show the feasibility of object searching and grasping by tactile information when visual information is not available. Shoujie Li, Linqi Ye, Chongkun Xia, Xueqian Wang 0001, Bin Liang 0001 |
SMC | 4 |
| 2021 | Reference Governor-Based Control for Active Rollover Avoidance of Mobile RobotsabstractRollover is a potential dangerous factor for mobile robots to accomplish a task. However, to our best knowledge, there still lacks a systematic research on the rollover mechanism and active rollover prevention control of mobile robots in the literature. This paper aims to propose a general control framework for rollover prevention of high-speed wheeled mobile robots. First, the lateral dynamics of the robot is modelled and the Load Transfer Ratio (LTR) is used as an index to measure the rollover level. Second, an optimal algorithm-based reference governor (RG) is developed, by which the wheel speed command that satisfies the constraint is induced, retaining the actual LTR within the threshold safety value. In addition, an integral sliding mode (ISM) wheel speed tracking controller is proposed. Lastly, simulations results show that for both trajectory tracking and path following cases, the controlled robot avoids possible rollover successfully. Chuan Yan, Xueqian Wang 0001, Jinchuan Zheng, Bin Liang 0001 |
SMC | 3 |
| 2021 | Symmetry in Biped WalkingabstractSymmetry in running was observed by Marc Raibert and was applied to simplify the control of dynamic legged systems. In this paper, we show that symmetry also exists in biped walking and investigate it using two simplified 2D models, that are, the inverted pendulum (IP) model and the linear inverted pendulum (LIP) model, both leading to similar conclusions. To characterize the symmetry in biped walking, the concept of acceleration factor is proposed. Symmetry occurs when the acceleration factor is zero, which results in an unchanged mid-stance velocity. And an important property of symmetry is that the n-step reachable region and the n-step controllable region are exactly the same. This means that if we can achieve speed B from A in n steps, then we can also achieve speed A from B in n steps. Symmetry in walking helps us to better understand human walking and also provides an intuitive way to control robotic walking. As an example, we propose a feedforward controller and a feedback controller, respectively, which can regulate the walking speed very effectively. This work provides us some new insights to view biped walking. Linqi Ye, Xueqian Wang 0001, Houde Liu, Bin Liang 0001 |
SMC | 2 |
| 2021 | Optimal Bounded Inversion for Nonminimum Phase Nonhyperbolic SystemsabstractAccurate tracking control of nonminimum phase systems relies on the calculation of the ideal internal dynamics (IID). Traditional IID calculation methods fail when applied to nonminimum phase nonhyperbolic systems (systems with nonhyperbolic zero dynamics). Recently, we propose the optimal bounded inversion method for IID calculation, which obtains IID by solving a trajectory optimization problem. In this paper, we extend our previous result and show that optimal bounded inversion can also deal with nonminimum phase nonhyperbolic systems. More than that, it is also possible to achieve different control goals by setting different cost functions. Particularly, three cases are investigated in this paper. The first uses minimal initial value deviation as the cost function, resulting in "T-IID" which can achieve accurate output tracking. The second applies minimal terminal value as the cost function, resulting in "S-IID" which leads to a final rest for the system. The last combines "T-IID" and "S-IID" to achieve a compound goal. The effectiveness is verified through Matlab simulations of a two-cart inverted-pendulum system. Linqi Ye, Deshan Meng, Xueqian Wang 0001, Bin Liang 0001 |
SMC | 4 |
| 2021 | Admissibility Analysis and Robust ${H_\infty }$ Control for T-S Fuzzy Descriptor Systems With Structured Parametric UncertaintiesabstractThis article considers admissibility analysis and robust${H_\infty }$control for continuous-time Takagi–Sugeno fuzzy descriptor systems with norm-bounded uncertainties in all parametric matrices. The system under consideration contains singular derivative matrices and different membership functions, thus it generalizes other related forms. First, uncertainties in derivative matrices are divided into two cases, i.e., one is expressed by a constant matrix left multiplied by an invertible uncertain matrix and the other is produced by its dual form. Then, admissible conditions and${H_\infty }$performance for the system with the first case of uncertainties are derived based on a new augmented system. As for the second case, the admissibility analysis is converted into the first case by an equivalent companion system then solved as well. All conditions are cast into strict linear matrix inequalities. Besides, due to the introduction of a new nonquadratic fuzzy Lyapunov function and slack decision variables, the proposed methods are less conservative than related ones. Finally, simulation examples are provided to illustrate improvements and effectiveness of the main results. Jiabao He 0001, Feng Xu 0006, Xueqian Wang 0001, Bin Liang 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | Tracking Control of a Linear Motor Positioner Based on Barrier Function Adaptive Sliding ModeabstractThe tracking performance of linear motor (LM) positioners is subject to payload uncertainty and external time-varying disturbances. Conventional robust controllers typically employ a high control gain that is substantially greater than the known a priori upper bound of the disturbance to ensure tracking error convergence. The main disadvantage of those controllers lies in that when the disturbance decreases, the control input is often overly generated, which then results in undesired control chattering effect or even actuator saturation. To overcome this problem, this article develops a robust tracking controller based on barrier function adaptive sliding mode (BFASM) for the LM positioners. The main benefits of BFASM are twofold: first, the controller is designed without the need for any disturbance information; second, its control gain is adaptively adjusted in terms of the amplitude of disturbance and, thus, leads to decreased control input when the disturbance becomes small. Furthermore, a modified barrier function (MBF) is proposed for applications with actuator saturation. It is proved that both the BFASM and MBF-based controllers can ensure the convergence of the tracking error into a prespecified neighborhood of zero in finite time. Experimental results on a real LM positioner demonstrate the superior properties of the developed controllers in comparison with two existing robust control schemes. Jinchuan Zheng, Hai Wang 0004, Xueqian Wang 0001, Renquan Lu, Zhihong Man |
IEEE Trans. Ind. Informatics | 4 |
| 2020 | Admissibility Analysis and Robust Stabilization via State Feedback for Uncertain T-S Fuzzy Descriptor SystemsabstractThis paper considers the admissibility and robust stabilization via state feedback for continuous-time T-S fuzzy descriptor systems (TSFDS) with a class of uncertainties. First, the admissibility of the nominal system without uncertainties is investigated. An equivalent augmented system is presented to deal with different and singular derivative matrices. Then, admissible conditions for the open-loop and close-loop systems are both derived based on a non-quadratic fuzzy Lyapunov function. Second, the admissibility and robust stabilization of TSFDS with uncertainties in all matrices are investigated. The uncertainty in each derivative matrix is equivalently expressed by a constant matrix left multiplied by an invertible uncertain matrix so that a similar augmented system can be constructed. Then admissible conditions are derived. This paper generalizes existing related results since we consider a wider class of TSFDS with different derivative matrices and different membership functions in each subsystem. All conditions are expressed as strict linear matrix inequalities (LMIs). Finally, a simulation example is provided to show effectiveness of the proposed results. Jiabao He 0001, Feng Xu 0006, Xueqian Wang 0001, Bin Liang 0001 |
FUZZ-IEEE | 3 |
| 2020 | Multi-task Control for a Quadruped Robot with Changeable Leg ConfigurationabstractThis paper proposes a multi-task control strategy for a quadruped robot named THU-QUAD II. The mechanical design of the robot ensures a wide range of motion for all joints, which allows it to stand and walk like a mammal as well as sprawl to the ground and crawl like a reptile. Five basic leg configurations are defined for the robot, including four mammal-type configurations with bidirectional knees and one sprawling-type configuration. A multi-task control framework is developed by combining configuration selection and gait planning. According to the locomotion environments, the robot can nimbly switch between different configurations, which gives it more flexibility when facing different tasks. For the mammal-type configuration, a parametric climbing gait is designed to traverse structural terrain. For the sprawling-type configuration, a crawling gait is designed to achieve robust locomotion on uneven terrain. Simulations and experiments show that the robot is capable to move on multiple challenging terrains, including doorsills, stairs, slopes, sand and stones. This paper demonstrates that even some challenging locomotion tasks can be achieved in a rather simple way without using complicated control algorithms, which suggests us to rethink about the leg configurations in designing quadruped robots. Linqi Ye, Houde Liu, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003 |
IROS | 3 |
| 2020 | Approximate Piecewise Constant Curvature Equivalent Model and Their Application to Continuum Robot Configuration EstimationabstractThe continuum robot has attracted more attention for its flexibility. Continuum robot kinematics models are the basis for further perception, planning, and control. The design and research of continuum robots are usually based on the assumption of piecewise constant curvature (PCC). However, due to the influence of friction, etc., the actual motion of the continuum robot is approximate piecewise constant curvature (APCC). To address this, we present a kinematic equivalent model for continuum robots, i.e. APCC 2L-5R. Using classical rigid linkages to replace the original model in kinematic, the APCC 2L-5R model effectively reduces complexity and improves numerical stability. Furthermore, based on the model, the configuration self-estimation of the continuum robot is realized by monocular cameras installed at the end of each approximate constant curvature segment. The potential of APCC 2L-5R in perception, planning, and control of continuum robots remains to be explored. Houde Liu, Xueqian Wang 0001, Bin Liang 0001 |
SMC | 3 |
| 2020 | Conservatism Comparison of State Estimation Error and Residual in Multiple Actuator Faults DetectionabstractThis paper focuses on analyzing and comparing the performance of two robust fault detection (FD) criteria for discrete-time linear parameter varying (LPV) systems with bounded uncertainties, namely the state estimation error-based criterion and the classical residual-based criterion. First, a new FD criterion for the detection of multiple multiplicative actuator faults is proposed by testing consistency between the state estimation errors and the healthy state estimation error sets on-line. Then, a guaranteed FD condition is established based on set-separation of healthy and faulty invariant sets of state estimation error. Moreover, the generalized minimum detectable fault (MDF) for multiple actuator faults is defined and computed in order to characterize the performance of the two FD criteria. Finally, a proof is provided to compare the conservatism of the FD criterion using state estimation errors with the classical one based on residuals. At the end of this paper, a numerical example is used to illustrate the effectiveness of the obtained results. Bo Min, Junbo Tan, Xueqian Wang 0001, Jun Yang 0028, Bin Liang 0001 |
SMC | 3 |
| 2020 | A Static Gait Generation for Quadruped Robots with Optimized Walking Speed*abstractTraversing at a high speed while maintaining stability is important for the application of quadruped robots. Prior works mainly concentrated on optimizing the stability margin of quadruped robots when walking through a variety of terrains. However, the problem of improving quadruped robots' walking velocity with static gait is less concerned in their works. In this paper, the static gait planning problem is considered under the assumption that a set of irregular footholds on the rough terrain is given, and two approaches are proposed to improve the walking speed. The first one is a distance optimization algorithm, which can minimize the moving distance of the center of gravity (COG) in the stance phases based on the stability and the kinematic constraint. The other is a velocity optimization algorithm, which enables the body and the feet to move at the highest velocity with the joint angular velocity limit. The joint application of these two optimization algorithms significantly improves the walking speed of the quadruped robot. Simulation results in V-REP are presented to demonstrate the effectiveness of the proposed approaches in improving the walking speed. Compared with the traditional gait planning techniques, one that moves the robot with the optimal stability margin, and the other that moves the robot without optimizing the velocity, our algorithms increase the average walking velocity by 81.6% and 32.8%, respectively. Linqi Ye, Xueqian Wang 0001, Nong Cheng, Houde Liu, Bin Liang 0001 |
SMC | 3 |
| 2019 | Distributed Detection of Generalized Gaussian Sparse Signals with One-Bit Measurements (Poster)
Xueqian Wang 0001, Gang Li 0008, Pramod K. Varshney |
FUSION | 1 |
| 2019 | A 3D Static Modeling Method and Experimental Verification of Continuum Robots Based on Pseudo-Rigid Body TheoryabstractContinuum robots composed of elastic backbones have a broad application prospect in the narrow and restricted environment because they overcome the disadvantages of traditional articulated robots, such as being bulky and inflexible. Statics plays an important role in the planning and control of the continuum robot composed of the elastic backbone. Pseudo-Rigid Body (PRB) theory has shown great potential in the description of flexible body statics. The PRB 3R model accurately describes the large deformation of the flexible body and has high computational efficiency. However, PRB 3R models mostly focus on the planar static modeling, and there are few applications in three-dimensional (3D) statics. In this paper, a 3D static modeling method of cable-driven continuum robot based on PRB 3R theory is proposed. By introducing the equilibrium constraint equations of resultant force/moment and bending plane normal of the elastic backbone, the state of the continuum robot is determined. The 3D static equations established by the proposed method take into account the comprehensive effects of the elastic force, external force, gravity and friction. A static verification experiment system of the cable-driven continuum robot is designed to verify the proposed method. The accuracy of the proposed method is verified by comparison with experimental data. The maximum position error between simulation and experimental results is 7.6%. Shaoping Huang, Deshan Meng, Xueqian Wang 0001, Bin Liang 0001, Weining Lu |
IROS | 3 |
| 2019 | Modeling and Control of Free-Floating Space Manipulator Using the T-S Fuzzy Descriptor System ApproachabstractIn this paper, a Takagi-Sugeno (T-S) fuzzy descriptor approach for control of a two-link free-floating space manipulator (FFSM) is proposed. The T-S fuzzy descriptor model of the FFSM is first derived from its nonlinear dynamic model, which makes more sense in reality since it avoids the use of joint acceleration measurement and the inversion of inertia matrix. And some nonlinear terms are considered as uncertainties to balance the complexity and accuracy of the model. Then a robust controller based on the Lyapunov stability theory is designed and reformulated as a linear matrix inequality (LMI) optimization problem which can be efficiently solved with the solver SeduMi. Finally, simulation results are carried out with the SimMechanics to demonstrate the effectiveness of the proposed approach. Jiabao He 0001, Feng Xu 0006, Xueqian Wang 0001, Jun Yang 0028, Bin Liang 0001 |
SMC | 3 |
| 2019 | Singularity-Free Trajectory Planning of Free-Floating Multiarm Space Robots for Keeping the Base Inertially StabilizedabstractIn a multiarm space robotic system, one or more manipulators can be used to stabilize the base through counteracting the disturbance caused by other manipulators performing on-orbital tasks. However, singularities are inevitably present in the traditional methods based on differential kinematics solutions. In this paper, we propose a singularity-free trajectory planning method to simultaneously keep the attitude and centroid position of the base stabilized in inertial space; the balance arms are also designed. First, we derive the coupling motion equations of a free-floating multiarm space robotic system. Then, the singularity problems are theoretically analyzed, and the theoretical basis for singularity-free trajectory planning is established. Second, we decompose the six degrees of freedom pose (attitude and position) stabilization problem into two 3DOF subproblems related to attitude and position balancing. We then design two robotic arms: 1) a position balance arm and 2) an attitude balance arm, to maintain the base centroid position and attitude, respectively. Third, we plan the coordinated trajectories of the two balance arms according to holonomic and nonholonomic constraints. As long as the desired motion is not beyond its balance ability, the reasonable joint variables can always be determined without encountering a singularity problem. Finally, the proposed methods are verified using simulations of typical on-orbital missions, including joint trajectory tracking and target capturing. Wenfu Xu, Deshan Meng, Houde Liu, Xueqian Wang 0001, Bin Liang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2018 | BRoPH: An efficient and compact binary descriptor for 3D point clouds
Xueqian Wang 0001, Tao Zhang 0006, Bin Liang 0001, Jingyan Song, Houde Liu |
Pattern Recognit. | 2 |
| 2018 | Mixed Active/Passive Robust Fault Detection and Isolation Using Set-Theoretic Unknown Input ObserversabstractThis paper proposes a robust fault detection and isolation (FDI) approach that combines active and passive robust FDI approaches. Standard active FDI approaches obtain robustness by using the unknown input observer (UIO) to decouple unknown inputs from residuals. Differently, standard passive FDI approaches achieve robustness by using the set theory to bound the effect of uncertain factors (disturbances and noises). In this paper, we combine the UIO-based and the set-based approaches to produce a mixed robust FDI, which can mitigate the disadvantages and exert the advantages of the two robust FDI approaches. In order to emphasize the role of set theory, the UIO design based on the set theory is named as the set-theoretic UIO (SUIO). A quadrotor subsystem is used to illustrate the effectiveness of the proposed FDI approach. Feng Xu 0006, Junbo Tan, Xueqian Wang 0001, Vicenç Puig, Bin Liang 0001, Bo Yuan 0003 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2016 | Ubiquitous Robot: A New Paradigm for Intelligence
Tiantian Zhang 0002, Bo Yuan 0003, Yinghao Ren, Houde Liu, Xueqian Wang 0001 |
IDEAL | 6 |
| 2014 | Measurement of relative pose between two non-cooperative spacecrafts based on graph cut theoryabstractIn final approach of rendezvous between a space robot and a non-cooperative target, due to the light and the folds of heat cladding materials on the target surface, the edge of target cannot be accurately extracted by classic Canny algorithm and we are unable to complete measurement of relative position and attitude. To solve this problem, the measurement method of relative position and attitude between two non-cooperative spacecrafts based on graph cut and edge information algorithm is proposed. A circular feature of the target is chosen as the recognition and measurement object. Firstly, the edge of the circular feature on the target is accurately extracted by graph cut and edge information algorithm. Secondly, the edges are ellipse fitting. Lastly, the relative position and attitude of target is obtained by fitting ellipse parameters of binocular cameras. The simulation results show that the precision of this method is better than Canny algorithm's and it can better meet the mission requirements. Bin Liang 0001, Xiaodong Du, Xueqian Wang 0001 |
ICARCV | 4 |
| 2014 | Autonomous path planning and experiment study of free-floating space robot for spinning satellite capturingabstractRobotic systems are expected to play an increasingly important role in future space activities with the development of space technology. The robotic on-orbital service, whose key is the capturing technology, becomes research hot in recent years. This paper focuses on the guidance of a robot manipulator to capture a spinning satellite with unknown dynamics parameters. In capturing a spinning satellite, a reference trajectory for control of the manipulator is generated with time delay due to the processing time of the target motion estimator and the manipulator controller. Consequently, the control system shows a poor performance and the end-effector sometimes fails to capture the target satellite. To solve this problem, the motion characteristics and motion prediction of the spinning satellite is analyzed Firstly, and using Unscented Kaiman Filter (UKF) to predict its movement. Then, a method of autonomous path planning of a free-floating space robot for target capturing is proposed, which is based on motion prediction and speed compensation. Finally, a ground experiment system is set up based on the concept of dynamic emulation and kinematic equivalence. With the experiment system, the autonomous target capturing experiments are conducted. The experiment results validate the proposed algorithm. Houde Liu, Bin Liang 0001, Xueqian Wang 0001 |
ICARCV | 3 |
| 2014 | On the autonomous target capturing of flexible-base space robotic systemabstractAutonomous target capturing is the key for space robot to perform on-orbital servicing tasks. To meet the requirement of complex and long-term task, large flexible appendages, such as solar paddles and antenna reflectors are usually mounted on the base of a space robot. Due to the structure vibration, it is very challenging to capture a free-floating target satellite. In this paper, we derived the kinematics equations and proposed the autonomous target capturing method for free-floating flexible-base space robots. The kinematics equation established the mapping from the base velocities, joint rates and elastic motion to the end-effector velocities. Based on this equation, we designed resolved motion rate control with vibration compensation for the space manipulator. Another contribution of this paper is that we modeled the dynamic coupling between the rigid movement of the end-effector and the flexible vibration of the solar paddles. Based on this model, we analyzed the coupling effect which was very important for the design of the manipulator and determining the trajectory planning and control strategy. At last, a simulation system was created and simulation studies of the proposed methods were carried out. The simulation results verify the proposed methods. Deshan Meng, Bin Liang 0001, Wenfu Xu, Xueqian Wang 0001, Houde Liu |
ICARCV | 4 |
| 2012 | A semi-physical simulation system for binocular vision guided rendezvousabstractAutonomous rendezvous in close range requires adequate ground simulations due to its significant difficulties and risks. In this paper, a novel semi-physical simulation system for binocular vision guided rendezvous is established. In this system, virtual three-dimensional models of the spacecrafts and the scene are created using computer graphic technology. Accordingly, images of the binocular cameras on board chaser (servicer) spacecraft are generated and displayed on the liquid crystal displays (LCDs). As the physical component in the simulation loop, two industrial cameras photograph the virtual images on the LCDs so that real camera noise is involved. In order to perform the closed-loop simulation, image acquisition, image processing, pose measurement, chaser guidance, navigation and control, and the system's dynamic motion are conducted. Through the combination of “virtual environment” and “physical environment”, the simulation system can successfully demonstrate binocular vision guided rendezvous. Simulation data is capable to verify the key algorithms during close range rendezvous. Changing the object model and dynamic model, this system can be applied to other vision-related researches. Xiaodong Du, Bin Liang 0001, Wenfu Xu, Xueqian Wang 0001, Xuehai Gao |
ICARCV | 4 |
| 2012 | Development of ground experiment system for space robot performing fine manipulationabstractRobotic systems are expected to play an increasingly important role in future space activities with the development of space technology. One broad area of application is in the servicing, construction, and maintenance of satellites and large space structures in orbit. Fine manipulation technology is very important for space robot to perform there tasks, since it must ensure safe and reliable interaction with objects or environment. In order to assure the task is accomplished successfully, ground experimentations are required for verifying key planning and control algorithms before the space robot is launched. In this paper, based on the concept of a hybrid approach combining the mathematical model with the physical model, a ground experiment system is set up, which is composed of two industrial robots, global and hand-eye visual equipments, six-axis force/momentum sensors, guide rail and four computers. Many control approaches of fine manipulation, such as compliance control, impedance control, hybrid force/position control, intelligent control, and so on, can be verified using this system. As an example, contour curves tracking experiment based on compliance control strategy is performed. Experiment results show that the ground system is very useful for verifying dexterous manipulation technology of space robot. Houde Liu, Bin Liang 0001, Wenfu Xu, Xueqian Wang 0001 |
ICARCV | 4 |