VLDB 2026 Research / reviewers in the wild / expert
Junbo Tan
dblp:192/2867
· DBLP profile ↗
14ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-8956-5408ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Surrogate-Assisted Evolutionary Multi-Agent Reinforcement Learning with Adaptive Fitness EvaluationabstractDeep Multi-Agent Reinforcement Learning (MARL) excels in co-operative tasks but often struggles with local optima in high - dimensional joint action spaces. In contrast, Evolutionary Algorithms (EAs) offer robust global exploration capabilities. Although hybrid approaches seek to combine the strengths of both paradigms, they typically face a critical bottleneck: the prohibitive sample cost of evaluating large populations via environment rollouts. To address this challenge, we propose Surrogate-assisted Evolutionary Multi-Agent Reinforcement Learning (SEMARL), a unified framework that synergizes gradient-based refinement with surrogate-assisted evolutionary search. SEMARL employs a cooperative co-evolutionary architecture to maintain diverse agent policies and injects gradient-refined parameters into the population to accelerate convergence. Crucially, we leverage the centralized critic from the gradient learner as a computationally efficient surrogate for fitness estimation. To prevent misleading guidance from an inaccurate critic, we introduce an adaptive reliability control mechanism based on temporal difference (TD) error, which dynamically regulates the surrogate's influence. Experiments on the Multi-Agent MuJoCo benchmark demonstrate that SEMARL significantly outperforms other algorithms, achieving superior asymptotic performance with substantially higher sample efficiency. Cong Yu 0018, Zaihui Yang, Haoyu Wang 0018, Junbo Tan, Yongzhe Chang, Tiantian Zhang 0002, Xueqian Wang 0001 |
GECCO | 6 |
| 2026 | TCSTNet: A text-driven color style transfer network for low-light image enhancement
Tianyi Zeng, Miao Zhang 0010, Zimo Zeng, Junfeng Jiao, Yuantao Wang, Yangfan He, Junbo Tan, Christian G. Claudel, Xueqian Wang 0001 |
Expert Syst. Appl. | 11 |
| 2026 | LowLightReward: A unified framework for low-light enhancement across spatial, channel, and aesthetic domains
Miao Zhang 0010, Haoyue Han, Yuantao Wang, Chenghe Yang, Hanning Liu, Junbo Tan, Xueqian Wang 0001 |
Neurocomputing | 10 |
| 2026 | An Efficient Solution Method for Workspace Boundary of Serpentine Manipulators Based on a Unified Kinematics Model
Deshan Meng, Taowen Guo, Runhui Xiang, Junbo Tan, Xueqian Wang 0001, Bin Liang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2025 | FOSP: Fine-tuning Offline Safe Policy through World ModelsabstractOffline Safe Reinforcement Learning (RL) seeks to address safety constraints by learning from static datasets and restricting exploration. However, these approaches heavily rely on the dataset and struggle to generalize to unseen scenarios safely. In this paper, we aim to improve safety during the deployment of vision-based robotic tasks through online fine-tuning an offline pretrained policy. To facilitate effective fine-tuning, we introduce model-based RL, which is known for its data efficiency. Specifically, our method employs in-sample optimization to improve offline training efficiency while incorporating reachability guidance to ensure safety. After obtaining an offline safe policy, a safe policy expansion approach is leveraged for online fine-tuning. The performance of our method is validated on simulation benchmarks with five vision-only tasks and through real-world robot deployment using limited data. It demonstrates that our approach significantly improves the generalization of offline policies to unseen safety-constrained scenarios. To the best of our knowledge, this is the first work to explore offline-to-online RL for safe generalization tasks. The videos are available at https://sunlighted.github.io/fosp_web/. Yucheng Xin, Silang Wu, Longxiang He, Zichen Yan, Junbo Tan, Xueqian Wang 0001 |
ICLR | 6 |
| 2025 | Robust Policy Expansion for Offline-to-Online RL under Diverse Data CorruptionabstractPretraining a policy on offline data followed by fine-tuning through online interactions, known as Offline-to-Online Reinforcement Learning (O2O RL), has emerged as a promising paradigm for real-world RL deployment. However, both offline datasets and online interactions in practical environments are often noisy or even maliciously corrupted, severely degrading the performance of O2O RL. Existing works primarily focus on mitigating the conservatism of offline policies via online exploration, while the robustness of O2O RL under data corruption, including states, actions, rewards, and dynamics, is still unexplored. In this work, we observe that data corruption induces heavy-tailed behavior in the policy, thereby substantially degrading the efficiency of online exploration. To address this issue, we incorporate Inverse Probability Weighted (IPW) into the online exploration policy to alleviate heavy-tailedness, and propose a novel, simple yet effective method termed $\textbf{RPEX}$: $\textbf{R}$obust $\textbf{P}$olicy $\textbf{EX}$pansion. Extensive experimental results on D4RL datasets demonstrate that RPEX achieves SOTA O2O performance across a wide range of data corruption scenarios. Longxiang He, Deheng Ye, Junbo Tan, Xueqian Wang 0001, Li Shen 0008 |
NeurIPS | 3 |
| 2025 | A Universal Vehicle-Trailer Navigation System with Neural Kinematics and Online Residual LearningabstractAutonomous navigation of vehicle-trailer systems is crucial in environments like airports, supermarkets, and concert venues, where various types of trailers are needed to navigate with different payloads and conditions. However, accurately modeling such systems remains challenging, especially for trailers with castor wheels. In this work, we propose a novel universal vehicle-trailer navigation system that integrates a hybrid nominal kinematic model—combining classical nonholonomic constraints for vehicles and neural network-based trailer kinematics—with a lightweight online residual learning module to correct real-time modeling discrepancies and disturbances. Additionally, we develop a model predictive control framework with a weighted model combination strategy that improves long-horizon prediction accuracy and ensures safer motion planning. Our approach is validated through extensive real-world experiments involving multiple trailer types and varying payload conditions, demonstrating robust performance without manual tuning or trailer-specific calibration. Yanbo Chen 0001, Yunzhe Tan, Yaojia Wang, Zhengzhe Xu, Junbo Tan, Xueqian Wang 0001 |
SMC | 5 |
| 2025 | Data-Driven MPC with Data Selection for Flexible Cable-Driven Robotic ArmsabstractFlexible cable-driven robotic arms (FCRAs) offer dexterous and compliant motion. Still, the inherent properties of cables, such as resilience, hysteresis, and friction, often lead to particular difficulties in modeling and control. This paper proposes a model predictive control (MPC) method that relies exclusively on input-output data, without a physical model, to improve the control accuracy of FCRAs. First, we develop an implicit model based on input-output data and integrate it into an MPC optimization framework. Second, a data selection algorithm (DSA) is introduced to filter the data that best characterize the system, thereby reducing the solution time per step to approximately 4 ms, which is an improvement of nearly 80%. Lastly, the influence of hyperparameters on tracking error is investigated through simulation. The proposed method has been validated on a real FCRA platform, including five-point positioning accuracy tests, a five-point response tracking test, and trajectory tracking for letter drawing. The results demonstrate that the average positioning accuracy is approximately 2.070 mm. Moreover, compared to the PID method with an average tracking error of 1.418°, the proposed method achieves an average tracking error of 0.541°. Huayue Liang, Yanbo Chen 0001, Hongyang Cheng, Yanzhao Yu, Shoujie Li, Junbo Tan, Xueqian Wang 0001, Long Zeng 0001 |
SMC | 6 |
| 2025 | A Hybrid Force-Position Strategy for Shape Control of Deformable Linear Objects With Graph Attention NetworksabstractManipulating deformable linear objects (DLOs) such as wires and cables is crucial in various applications like electronics assembly and medical surgeries. However, it faces challenges due to DLOs’ infinite degrees of freedom, complex nonlinear dynamics, and the underactuated nature of the system. To address these issues, this paper proposes a hybrid force-position strategy for DLO shape control. The framework, combining both force and position representations of DLO, integrates state trajectory planning in the force space and Model Predictive Control (MPC) in the position space. We present a dynamics model with an explicit action encoder, a property extractor and a graph processor based on Graph Attention Networks. The model is used in the MPC to enhance prediction accuracy. Results from both simulations and real-world experiments demonstrate the effectiveness of our approach in achieving efficient and stable shape control of DLOs. Codes and videos are available at https://sites.google.com/view/dlom. Yanzhao Yu, Junbo Tan, Xueqian Wang 0001 |
SMC | 3 |
| 2024 | Offline Goal-Conditioned Reinforcement Learning for Safety-Critical Tasks with Recovery PolicyabstractOffline goal-conditioned reinforcement learning (GCRL) aims at solving goal-reaching tasks with sparse rewards from an offline dataset. While prior work has demonstrated various approaches for agents to learn near-optimal policies, these methods encounter limitations when dealing with diverse constraints in complex environments, such as safety constraints. Some of these approaches prioritize goal attainment without considering safety, while others excessively focus on safety at the expense of training efficiency. In this paper, we study the problem of constrained offline GCRL and propose a new method called Recovery-based Supervised Learning (RbSL) to accomplish safety-critical tasks with various goals. To evaluate the method performance, we build a benchmark based on the robot-fetching environment with a randomly positioned obstacle and use expert or random policies to generate an offline dataset. We compare RbSL with three offline GCRL algorithms and one offline safe RL algorithm. As a result, our method outperforms the existing state-of-the-art methods to a large extent. Furthermore, we validate the practicality and effectiveness of RbSL by deploying it on a real Panda manipulator. Code is available at https://github.com/Sunlighted/RbSL.git. Zichen Yan, Renhao Lu, Junbo Tan, Xueqian Wang 0001 |
ICRA | 4 |
| 2024 | Demo Abstract: Range-SLAM: UWB based Realtime Indoor Location and MappingabstractSimultaneous localization and mapping (SLAM) systems frequently employ LiDAR and cameras as essential sensing components. However, these sensors are proved to be unreliable in environments with poor visibility or reflective surfaces. And UWB (Ultra Wide Band) sensor with a longer wavelength shows better potential to achieve perception tasks. However, since UWB sensors can only obtain distance information from the anchors, it is difficult to densely construct the geometric structure of the environment. In this paper, We propose Range-SLAM, a method based on received signal strength indicator (RSSI) recognition and binary filtering to complete the mapping task and enhance positioning based on the map, and only require UWB as external perception sensor. Real-world experiments are conducted and prove the effectiveness, real-time performance and robustness of the Range-SLAM algorithm. Zhuozhu Jian, Junbo Tan, Lunfei Liang, Houde Liu, Xinlei Chen |
IPSN | 3 |
| 2023 | Visuotactile Sensor Enabled Pneumatic Device Towards Compliant Oropharyngeal Swab SamplingabstractManual oropharyngeal (OP) swab sampling is an intensive and risky task. In this article, a novel OP swab sampling device of low cost and high compliance is designed by combining the visuotactile sensor and the pneumatic actuator-based gripper. Here, a concave visuotactile sensor called CoTac is first proposed to address the problems of high cost and poor reliability of traditional multi-axis force sensors. Besides, by imitating the doctor's fingers, a soft pneumatic actuator with a rigid skeleton structure is designed, which is demonstrated to be reliable and safe via finite element modeling and experiments. Furthermore, we propose a sampling method that adopts a compliant control algorithm based on the adaptive virtual force to enhance the safety and compliance of the swab sampling process. The effectiveness of the device has been verified through sampling experiments as well as in vivo tests, indicating great application potential. The cost of the device is around 30 US dollars and the total weight of the functional part is less than 0.1 kg, allowing the device to be rapidly deployed on various robotic arms. Shoujie Li, Mingshan He, Wenbo Ding 0001, Linqi Ye, Xueqian Wang 0001, Junbo Tan, Jinqiu Yuan, Xiao-Ping Zhang 0002 |
IROS | 6 |
| 2020 | Conservatism Comparison of State Estimation Error and Residual in Multiple Actuator Faults DetectionabstractThis paper focuses on analyzing and comparing the performance of two robust fault detection (FD) criteria for discrete-time linear parameter varying (LPV) systems with bounded uncertainties, namely the state estimation error-based criterion and the classical residual-based criterion. First, a new FD criterion for the detection of multiple multiplicative actuator faults is proposed by testing consistency between the state estimation errors and the healthy state estimation error sets on-line. Then, a guaranteed FD condition is established based on set-separation of healthy and faulty invariant sets of state estimation error. Moreover, the generalized minimum detectable fault (MDF) for multiple actuator faults is defined and computed in order to characterize the performance of the two FD criteria. Finally, a proof is provided to compare the conservatism of the FD criterion using state estimation errors with the classical one based on residuals. At the end of this paper, a numerical example is used to illustrate the effectiveness of the obtained results. Bo Min, Junbo Tan, Xueqian Wang 0001, Jun Yang 0028, Bin Liang 0001 |
SMC | 2 |
| 2018 | Mixed Active/Passive Robust Fault Detection and Isolation Using Set-Theoretic Unknown Input ObserversabstractThis paper proposes a robust fault detection and isolation (FDI) approach that combines active and passive robust FDI approaches. Standard active FDI approaches obtain robustness by using the unknown input observer (UIO) to decouple unknown inputs from residuals. Differently, standard passive FDI approaches achieve robustness by using the set theory to bound the effect of uncertain factors (disturbances and noises). In this paper, we combine the UIO-based and the set-based approaches to produce a mixed robust FDI, which can mitigate the disadvantages and exert the advantages of the two robust FDI approaches. In order to emphasize the role of set theory, the UIO design based on the set theory is named as the set-theoretic UIO (SUIO). A quadrotor subsystem is used to illustrate the effectiveness of the proposed FDI approach. Feng Xu 0006, Junbo Tan, Xueqian Wang 0001, Vicenç Puig, Bin Liang 0001, Bo Yuan 0003 |
IEEE Trans Autom. Sci. Eng. | 2 |