EDBT 2026 Demo / reviewers in the wild / expert
Yanbing Mao
dblp:141/4975
· DBLP profile ↗
9ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0002-7233-4179ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 46% Motion planning and robot control · 28% Transfer learning and domain adaptation · 12% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
safe reinforcement learning |
1.6 | 2 | 2025 | Real-DRL: Teach and Learn at Runtime · NeurIPS 2025 Physics-Regulated Deep Reinforcement Learning: Invariant Embeddings · ICLR 2024 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.9 | 1 | 2025 | Real-DRL: Teach and Learn at Runtime · NeurIPS 2025 |
Robotics › Motion planning and robot control
robot control |
0.9 | 1 | 2025 | Real-DRL: Teach and Learn at Runtime · NeurIPS 2025 |
Robotics › Motion planning and robot control › robot control › safe control
safety-critical robot control |
0.9 | 1 | 2025 | Real-DRL: Teach and Learn at Runtime · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
0.9 | 1 | 2025 | Real-DRL: Teach and Learn at Runtime · NeurIPS 2025 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.8 | 1 | 2024 | Physics-Regulated Deep Reinforcement Learning: Invariant Embeddings · ICLR 2024 |
Machine learning › Deep learning architectures and training
physics-informed neural network |
0.8 | 1 | 2024 | Physics-Regulated Deep Reinforcement Learning: Invariant Embeddings · ICLR 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.2 | 1 | 2024 | Physics-Regulated Deep Reinforcement Learning: Invariant Embeddings · ICLR 2024 |
Robotics › Motion planning and robot control
safety guarantees |
0.2 | 1 | 2024 | Physics-Regulated Deep Reinforcement Learning: Invariant Embeddings · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
teaching-to-learn · 0.9physics-model-based teacher · 0.9batch sampling · 0.9residual action policy · 0.8neural network editing · 0.8invariant embeddings · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-DRL: Teach and Learn at RuntimeabstractThis paper introduces the Real-DRL framework for safety-critical autonomous systems, enabling runtime learning of a deep reinforcement learning (DRL) agent to develop safe and high-performance action policies in real plants while prioritizing safety. The Real-DRL consists of three interactive components: a DRL-Student, a PHY-Teacher, and a Trigger. The DRL-Student is a DRL agent that innovates in the dual self-learning and teaching-to-learn paradigm and the safety-status-dependent batch sampling. On the other hand, PHY-Teacher is a physics-model-based design of action policies that focuses solely on safety-critical functions. PHY-Teacher is novel in its real-time patch for two key missions: i) fostering the teaching-to-learn paradigm for DRL-Student and ii) backing up the safety of real plants. The Trigger manages the interaction between the DRL-Student and the PHY-Teacher. Powered by the three interactive components, the Real-DRL can effectively address safety challenges that arise from the unknown unknowns and the Sim2Real gap. Additionally, Real-DRL notably features i) assured safety, ii) automatic hierarchy learning (i.e., safety-first learning and then high-performance learning), and iii) safety-informed batch sampling to address the experience imbalance caused by corner cases. Experiments with a real quadruped robot, a quadruped robot in Nvidia Isaac Gym, and a cart-pole system, along with comparisons and ablation studies, demonstrate the Real-DRL's effectiveness and unique features. Yanbing Mao, Yihao Cai, Lui Sha |
NeurIPS | 1 |
| 2025 | Phy-Taylor: Partially Physics-Knowledge-Enhanced Deep Neural Networks via NN EditingabstractPurely data-driven deep neural networks (DNNs) applied to physical engineering systems can infer relations that violate physics laws, thus leading to unexpected consequences. To address this challenge, we propose a physics-knowledge-enhanced DNN framework called Phy-Taylor, accelerating learning-compliant representations with physics knowledge. The Phy-Taylor framework makes two key contributions; it introduces a new architectural physics-compatible neural network (PhN) and features a novel compliance mechanism, which we call physics-guided neural network (NN) editing. The PhN aims to directly capture nonlinear physical quantities, such as kinetic energy, electrical power, and aerodynamic drag force. To do so, the PhN augments NN layers with two key components: 1) monomials of the Taylor series for capturing physical quantities and 2) a suppressor for mitigating the influence of noise. The NN editing mechanism further modifies network links and activation functions consistently with physics knowledge. As an extension, we also propose a self-correcting Phy-Taylor framework for safety-critical control of autonomous systems, which introduces two additional capabilities: 1) safety relationship learning and 2) automatic output correction when safety violations occur. Through experiments, we show that Phy-Taylor features considerably fewer parameters and a remarkably accelerated training process while offering enhanced model robustness and accuracy. Yanbing Mao, Yuliang Gu, Lui Sha, Huajie Shao, Qixin Wang 0001, Tarek F. Abdelzaher |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Physics-Regulated Deep Reinforcement Learning: Invariant EmbeddingsabstractThis paper proposes the Phy-DRL: a physics-regulated deep reinforcement learning (DRL) framework for safety-critical autonomous systems. The Phy-DRL has three distinguished invariant-embedding designs: i) residual action policy (i.e., integrating data-driven-DRL action policy and physics-model-based action policy), ii) automatically constructed safety-embedded reward, and iii) physics-model-guided neural network (NN) editing, including link editing and activation editing. Theoretically, the Phy-DRL exhibits 1) a mathematically provable safety guarantee and 2) strict compliance of critic and actor networks with physics knowledge about the action-value function and action policy. Finally, we evaluate the Phy-DRL on a cart-pole system and a quadruped robot. The experiments validate our theoretical results and demonstrate that Phy-DRL features guaranteed safety compared to purely data-driven DRL and solely model-based design while offering remarkably fewer learning parameters and fast training towards safety guarantee. Hongpeng Cao, Yanbing Mao, Lui Sha, Marco Caccamo |
ICLR | 2 |
| 2024 | Social System Inference From Noisy ObservationsabstractThis article studies social system inference from a single noisy trajectory of public evolving opinions, wherein observation noise leads to the statistical dependence of samples on time and coordinates. We first propose a cyber-social system that comprises individuals in a social network and a set of information sources in a cyber layer, whose opinion dynamics explicitly takes the asymmetric cognitive bias including confirmation bias and negativity bias and the process noise into account. Based on the proposed cyber-social model, we then study the sample complexity of least-square auto-regressive model estimation, which governs the length of a single observed trajectory that is sufficient for the identified model to achieve the prescribed levels of accuracy and confidence (PAC). Building on the identified social model, we then investigate social inference, with a particular focus on the weighted network topology and the model parameters of asymmetric cognitive bias. Finally, the theoretical results and the effectiveness of the proposed inference framework are validated by the U.S. Senate Member Ideology data. Yanbing Mao, Naira Hovakimyan, Tarek F. Abdelzaher, Evangelos A. Theodorou |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Cost Function Learning in Memorized Social Networks With Cognitive Behavioral AsymmetryabstractThis article investigates the cost function learning in social information networks, wherein human memory and cognitive bias are explicitly taken into account. We first propose a model for social information-diffusion dynamics, with a focus on the systematic modeling of asymmetric cognitive bias represented by confirmation bias and novelty bias. Building on the dynamics model, we then propose the M3IRL—a memorized model and maximum-entropy-based inverse reinforcement learning—for learning cost functions. Compared with the existing model-free IRLs, the characteristics of M3IRL are significantly different here: no dependence on the Markov decision process principle, the need for only a single finite-time trajectory sample, and bounded decision variables. Finally, the effectiveness of the proposed social information-diffusion model and the M3IRL algorithm is validated by the online social media data. Yanbing Mao, Jinning Li 0001, Naira Hovakimyan, Tarek F. Abdelzaher, Christian Lebiere |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Sℒ1-Simplex: Safe Velocity Regulation of Self-Driving Vehicles in Dynamic and Unforeseen EnvironmentsabstractThis article proposes a novel extension of the Simplex architecture with model switching and model learning to achieve safe velocity regulation of self-driving vehicles in dynamic and unforeseen environments. To guarantee the reliability of autonomous vehicles, an ℒ 1 adaptive controller that compensates for uncertainties and disturbances is employed by the Simplex architecture as a verified high-assurance controller (HAC) to tolerate concurrent software and physical failures. Meanwhile, the safe switching controller is incorporated into the HAC for safe velocity regulation in the dynamic (prepared) environments, through the integration of the traction control system and anti-lock braking system. Due to the high dependence of vehicle dynamics on the driving environments, the HAC leverages the finite-time model learning to timely learn and update the vehicle model for ℒ 1 adaptive controller, when any deviation from the safety envelope or the uncertainty measurement threshold occurs in the unforeseen driving environments. With the integration of ℒ 1 adaptive controller, safe switching controller and finite-time model learning, the vehicle’s angular and longitudinal velocities can asymptotically track the provided references in the dynamic and unforeseen driving environments, while the wheel slips are restricted to safety envelopes to prevent slipping and sliding. Finally, the effectiveness of the proposed Simplex architecture for safe velocity regulation is validated by the AutoRally platform. Yanbing Mao, Yuliang Gu, Naira Hovakimyan, Lui Sha, Petros G. Voulgaris |
ACM Trans. Cyber Phys. Syst. | 1 |
| 2020 | Asymptotic Frequency Synchronization of Kuramoto Model by Step ForceabstractThis paper explains how network topology acts as a control variable for the asymptotic frequency synchronization of the Kuramoto model with finite oscillators. We first investigate the stability of asymptotic frequency synchronization in the Kuramoto model with step force generated by topology switching. We derive a sufficient condition for the asymptotic synchronization. Interestingly, the stability implies that with no constraint on the phase differences or the coupling strength, the asymptotic frequency synchronization can be achieved under certain topology-switching signals. Then, based on the stability of asymptotic frequency synchronization, two triggered topology-switching algorithms are proposed. The step force together with the topology-switching algorithm work as a new frequency synchronization algorithm. Compared with the existing results, the merit of the proposed frequency synchronization algorithms is that they have no constraint on the magnitude of the coupling strength or the phase differences. Simulations are provided to verify the effectiveness of the proposed algorithms. Yanbing Mao |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2017 | Finite-Time Stabilization of Discrete-Time Switched Nonlinear Systems Without Stable Subsystems via Optimal Switching Signal DesignabstractThis paper investigates the finite-time exponential stability analysis and stabilization problem of discrete-time switched nonlinear systems without stable subsystems. In the stability analysis, the Takagi-Sugeno (T-S) fuzzy model is employed to approximate nonlinear subsystems. With two level functions, namely, crisp switching functions and local fuzzy weighting functions, we introduce switched fuzzy systems with approximation errors, which inherently contain both the features of the switched systems and T-S fuzzy systems. By constructing the “decreasing-jump” piecewise Lyapunov-like functions and minimum dwell time technique, a finite-time exponential stability of switched fuzzy systems with bounded approximation errors is obtained. Then, based on the finite-time exponential stability, a multiobjective evolution algorithm (nondominated sorting genetic algorithm, NSGA-II), which considers two conflicting objectives, such as the average convergence error and the average switching cost, is proposed to generate tradeoff switching sequences to stabilize the discrete-time switched nonlinear systems over a finite-time interval. A numerical example and a practical example are provided to illustrate the effectiveness of the stability and the algorithm, respectively. Yanbing Mao, Hongbin Zhang 0002 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2014 | The Exponential Stability and Asynchronous Stabilization of a Class of Switched Nonlinear System Via the T-S Fuzzy ModelabstractIn this paper, we investigate the problem of exponential stability and asynchronous stabilization for a class of switched nonlinear systems. The Takagi and Sugeno (T-S) fuzzy model is employed to approximate the subnonlinear dynamic systems. With two-level functions, namely, crisp switching functions and local fuzzy weighting functions, we introduce continuous-time switched fuzzy systems, which inherently contain the features of the switched hybrid systems and T-S fuzzy systems. By the use of delicately constructed piecewise Lyapunov-like functions (PLFs) and minimum dwell time method, we obtain the exponential stability of the switched fuzzy systems, which allows us to have stable and unstable nonlinear subsystems. In practice, for the control problem, it inevitably takes some time to identify the system modes and apply the matched controller, the asynchronous phenomena between the system modes switching and the controllers switching generally exists. Based on the result of stability, the fuzzy state feedback controller under asynchronous switching is proposed for switched fuzzy systems. In addition, the lower bound of minimum dwell time can be obtained using convex optimization such that the switched fuzzy system can be exponentially stabilized if its minimum dwell time is larger than the bound. The stability results and control laws of the switched fuzzy systems are formulated in the form of linear matrix inequalities that are numerically feasible. Finally, two illustrated numerical examples are presented to show the effectiveness of the obtained theoretical results. Yanbing Mao, Hongbin Zhang 0002 |
IEEE Trans. Fuzzy Syst. | 1 |