EDBT 2026 Demo / reviewers in the wild / expert
Mingyu Cai
dblp:263/6237
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Online Caching With Delayed Hits in Multi-Server Edge NetworksabstractEdge caching is a critical application scenario in edge networks. By storing diverse files on edge servers and dynamically fetching new files from the cloud, edge networks can provide low-latency file access services for mobile users. In practice, the file fetching latency is non-negligible. Consecutive requests for the same missing file during the fetching phase introduce additional latency (referred to as delayed hits). Existing studies either ignore the delayed hits when making caching decisions or are not applicable to multi-server edge networks. In this paper, we investigate the online caching problem with delayed hits in the multi-server edge networks and prove its hardness. The objective is to minimize the total file access latency. To solve the proposed problem, we propose Cadle, which makes caching decisions based on the latency of different file access operations and weights of files in an online manner, without relying on any prior knowledge of future requests. We prove the competitive ratio of Cadle.We also conduct extensive experiments on the real-world dataset to verify the performance of Cadle. The experimental results show that Cadle reduces the total file access latency by at least 31.8% on average, and improves the hit ratio by at least 25.9% on average compared with state-of-the-art approaches. Xin He 0010, Mingyu Cai, Meng Li 0010, Haipeng Dai 0001, Jian Zhou 0009, Fu Xiao 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | A Unified Framework to Learn Collision-Free Loco-Manipulation via Adversarial Motion PriorsabstractDesigning a whole-body controller for loco-manipulation in unstructured real-world environments remains a formidable challenge. Previous approaches have primarily focused on extending the workspace of robotic arms while maintaining quadrupedal landing postures. However, these methods fail to fully exploit the mobility of legged robots. To address these limitations, we propose a unified framework for collision-free loco-manipulation in real-world applications. The framework comprises two key modules: (1) a Loco-manipulation Motion Prior, which generates loco-manipulation skill trajectories via Trajectory Optimization (TO), and (2) a Collision-free Manipulation module using a Model Predictive Path Integral (MPPI)-based trajectory generator and a vector-based trajectory follower. Extensive experiments have been conducted in both simulation and real-world scenarios to evaluate our framework’s tracking accuracy, whole-body coordination, and workspace expansion capabilities. Supplementary and videos are available at: https://sites.google.com/view/loco-mani-amp/ Huayang Yin, Tangyu Qian, Guanchen Lu, Mingyu Cai, Zhen Kan |
IROS | 5 |
| 2025 | Fast Motion Planning in Dynamic Environments With Extended Predicate-Based Temporal LogicabstractFormal languages effectively outline robots’ task specifications, yet current temporal logic struggles to balance semantic expression with solution speed. To address this challenge, we propose extended predicate-based temporal logic (E-pTL), augmenting conventional linear temporal logic with more expressive atomic predicates to reflect the time, space, and order attributes of task and represent complex tasks with dynamic propositions, time windows, and even relative time window. This approach, blending automata-based task abstraction and extended predicates, offers enhanced expressiveness and conciseness for intricate specifications. To cope with E-pTL, we introduce a novel planning framework, the Planning Decision Tree (PDT). PDT incrementally builds a tree through automata and system state searches, recording potential task plans. The proposed pruning method can reduce the exploration space. This method swiftly handles complex temporal tasks defined by E-pTL. Rigorous analysis confirms PDT-based planning’s feasibility (ensuring satisfactory planning aligned with task specifications) and completeness (guaranteeing a feasible solution if available). Moreover, PDT-based planning proves efficient, with solution times approximately linearly proportional to automaton states squared. Extensive simulations and experiments validate its effectiveness and efficiency. Note to Practitioners—Temporal logic, as a formal language, has been widely applied to describe the task of robotic systems. However, linear temporal logic (LTL) struggles with the explicit time constraints, while other temporal logics, e.g., signal temporal logic (STL), cannot compactly represent task progress via an automaton. Consequently, current solutions for complex task with temporal logic specification heavily lean on optimization-based approaches, which are time-consuming and impractical for real-time systems. This work presents a novel approach for fast motion planning in dynamic environments with extended predicate-based temporal logic (E-pTL). The proposed E-pTL not only enhances semantic expression but also can be checked by an automaton, simplifying progress checks. However, existing automata-based methods encounter limitations in dynamic tasks. Graph search methods require replanning at each time step and lack completeness guarantees, while sampling methods can’t efficiently handle the feasibility of each atomic proposition. In contrast, the developed planning decision tree (PDT)-based framework incrementally searches for automata and validates propositions without a transition system, significantly reducing the search space. The suggested pruning method further trims the search space without compromising feasibility. Overall, PDT enables rapid, reliable planning, outperforming optimization-based methods like ILP/MILP and proving well-suited for real-time applications. Mingyu Cai, Zhangli Zhou, Zhen Kan |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Local Observation Based Reactive Temporal Logic Planning of Human-Robot SystemsabstractHuman-robot collaboration plays an important role in intelligent manufacturing. However, the main challenge is how the robot can make online reactive changes to the plan based on the observed human behavior to ensure the completion of user-defined tasks. Such a challenge is further exacerbated if eye-in-hand manipulation is considered since the local field of the camera view cannot capture global observations. Different from existing planning approaches that separate the perception and planning modules, and make strong assumptions about perception abilities, we develop a framework of real-time local reactive planning that enables the robot to quickly adapt its actions if necessary through its limited perception of surroundings using an eye-in-hand camera. Specifically, we develop a locally observable transition system (LOTS) and interpretably express the task using linear temporal logic (LTL). To improve the grasping performance using local visual perception, we propose a high-resolution grasp network (HRG-Net) that achieves state-of-the-art results on multiple datasets (99.50% in Cornell and 97.50% in Jacquard and 96% in Graspnet-1Billion) for the task. A physical experiment using a 7DoF Franka Emika Panda robot demonstrates the effectiveness of the reactive planning framework.Note to Practitioners—Intelligent manufacturing often requires the human operator to work collaboratively with the robot in a shared workspace. Due to possible (assistive or non-assistive) interference of human operators, it is highly desired that the robot can perceive human behaviors and react properly to ensure task accomplishment. Hence, this work is particularly motivated to develop a reactive planning framework that relies on real-time local visual perception (i.e., eye-in-hand camera) to quickly react to its dynamic surroundings and replan its motion when necessary. In future work, rather than using the observed human behavior, we will investigate how to predict human intentions to further improve human-robot collaboration. Zhangli Zhou, Mingyu Cai, Hao Wang 0161, Zhijun Li 0001, Zhen Kan |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Visual Object Tracking With Multi-Frame Distractor SuppressionabstractWith the rapid development of CNN or Transformer, the present mainstream approaches regard an image patch as the reference of the target to perform tracking, which is known as template matching-based trackers. However, most existing template matching-based trackers only consider the per-frame localization accuracy, neglecting the potential distractor (similar object) dependencies among multiple video frames, which poses a fundamental challenge in template matching-based tracking. In this work, we propose a novel comprehensive framework with multi-frame distractor suppression for visual object tracking (MFDSTrack), which explicitly models the temporal history of both the target object and potential distractors. Specifically, we utilize a universal target candidate generation module to detect target candidates (both target and distractors), providing a holistic view of the scene. In addition, a temporal and distractor-aware association module is designed to suppress multi-frame distractors by adopting a simple encoder-decoder Transformer architecture. The encoder accepts inputs of target candidates’ history, while the decoder takes current target candidate queries and the output of the encoder as inputs to associate current target candidate queries with historical trajectories. We extensively evaluate our trackers, MFDSTrack-SD, MFDSTrack-OS, MFDSTrack-GRM, and MFDSTrack-LT on the LaSOT,${\mathrm {LaSOT}}_{ext}$, TrackingNet, GOT-10k, UAV123, NFS, and OTB100 benchmark. Extensive experiments show that our methods outperform previous state-of-the-art trackers on seven tracking benchmarks. Mingyu Cai, Zhixuan Bai, Tao Zhuo, Hongming Zhang 0002, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | A Two-Level Control Algorithm for Autonomous Driving in Urban EnvironmentsabstractWe propose a two-level hierarchical architecture for controlling an autonomous vehicle (ego) in complex urban driving environments. This approach ensures both collision avoidance and adherence to traffic rules while maintaining real-time performance. At the top level of our framework, we use a simple dynamic model for ego and a simplified representation of the environment to formulate a Model Predictive Control (MPC) problem. The traffic rules are represented by Signal Temporal Logic (STL) formulas and incorporated as mixed integer-linear constraints within the MPC optimization. The top level MPC solution is then simulated at the bottom level, which employs detailed models of both ego dynamics and the environment. If a collision or traffic rule violation occurs, the bottom level provides feedback to the top level in the form of correction constraints, which are mixed integer-linear constraints affecting the state and control input of ego. This closed-loop feedback from the bottom level helps address discrepancies between the simplified models used in the MPC and the complex real-world models. We assess the effectiveness and runtime performance of our method by comparing it with existing approaches, through simulations of various urban driving scenarios in the CARLA simulator. Erfan Aasi, Mingyu Cai, Cristian Ioan Vasile, Calin Belta |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Hierarchical Deep Learning for Intention Estimation of Teleoperation Manipulation in Assembly TasksabstractIn human-robot collaboration, shared control presents an opportunity to teleoperate robotic manipulation to improve the efficiency of manufacturing and assembly processes. Robots are expected to assist in executing the user’s intentions. To this end, robust and prompt intention estimation is needed, relying on behavioral observations. The framework presents an intention estimation technique at hierarchical levels i.e., low-level actions and high-level tasks, by incorporating multi-scale hierarchical information in neural networks. Technically, we employ hierarchical dependency loss to boost overall accuracy. Furthermore, we propose a multi-window method that assigns proper hierarchical prediction windows of input data. An analysis of the predictive power with various inputs demonstrates the predominance of the deep hierarchical model in the sense of prediction accuracy and early intention identification. We implement the algorithm on a virtual reality (VR) setup to teleoperate robotic hands in a simulation with various assembly tasks to show the effectiveness of online estimation. Video demonstration is available at: https://youtu.be/CMYDgcI4j1g. Mingyu Cai, Karankumar Patel, Soshi Iba, Songpo Li |
ICRA | 1 |
| 2024 | LEEPS: Learning End-to-End Legged Perceptive Parkour Skills on Challenging TerrainsabstractEmpowering legged robots with agile maneuvers is a great challenge. While existing works have proposed diverse control-based and learning-based methods, it remains an open problem to endow robots with animal-like perception and athleticism. Towards this goal, we develop an End-to-End Legged Perceptive Parkour Skill Learning (LEEPS) framework to train quadruped robots to master parkour skills in complex environments. In particular, LEEPS incorporates a vision-based perception module equipped with multi-layered scans, supplying robots with comprehensive, precise, and adaptable information about their surroundings. Leveraging such visual data, a position-based task formulation liberates the robot from velocity constraints and directs it toward the target using innovative reward mechanisms. The resulting controller empowers an affordable quadruped robot to successfully traverse previously challenging and unprecedented obstacles. We evaluate LEEPS on various challenging tasks, which demonstrate its effectiveness, robustness, and generalizability. Supplementary and videos are available at: https://sites.google.com/view/leeps Tangyu Qian, Hao Zhang 0127, Zhangli Zhou, Hao Wang 0161, Mingyu Cai, Zhen Kan |
IROS | 5 |
| 2024 | Model-free reinforcement learning for motion planning of autonomous agents with complex tasks in partially observable environments
Mingyu Cai, Zhen Kan, Shaoping Xiao |
Auton. Agents Multi Agent Syst. | 2 |
| 2021 | Reinforcement Learning Based Temporal Logic Control with Maximum Probabilistic SatisfactionabstractThis paper presents a model-free reinforcement learning (RL) algorithm to synthesize a control policy that maximizes the satisfaction probability of complex tasks, which are expressed by linear temporal logic (LTL) specifications. Due to the consideration of environment and motion uncertainties, we model the robot motion as a probabilistic labeled Markov decision process (PL-MDP) with unknown transition probabilities and probabilistic labeling functions. The LTL task specification is converted to a limit deterministic generalized Büchi automaton (LDGBA) with several accepting sets to maintain dense rewards during learning. The novelty of applying LDGBA is to construct an embedded LDGBA (E-LDGBA) by designing a synchronous tracking-frontier function, which enables the record of non-visited accepting sets of LDGBA at each round of the repeated visiting pattern, to overcome the difficulties of directly applying conventional LDGBA. With appropriate dependent reward and discount functions, rigorous analysis shows that any method, which optimizes the expected discount return of the RL-based approach, is guaranteed to find the optimal policy to maximize the satisfaction probability of the LTL specifications. A model-free RL-based motion planning strategy is developed to generate the optimal policy in this paper. The effectiveness of the RL-based control synthesis is demonstrated via simulation and experimental results. Mingyu Cai, Shaoping Xiao, Baoluo Li, Zhiliang Li, Zhen Kan |
ICRA | 1 |