VLDB 2026 Research / reviewers in the wild / expert
Peide Huang
dblp:295/8645
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Autonomous driving · 32% Reinforcement learning · 32% Trustworthy machine learning · 11% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Autonomous driving
safety-critical scenario generation |
0.9 | 1 | 2025 | CaDRE: Controllable and Diverse Generation of Safety-Critical Driving Scenarios Using Real-World Trajectories · ICRA 2025 |
Robotics › Autonomous driving › autonomous vehicle testing
simulation-based testing |
0.9 | 1 | 2025 | CaDRE: Controllable and Diverse Generation of Safety-Critical Driving Scenarios Using Real-World Trajectories · ICRA 2025 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.6 | 1 | 2022 | Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training · IJCAI 2022 |
Machine learning › Reinforcement learning
curriculum reinforcement learning |
0.6 | 1 | 2022 | Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain Adaptation · NeurIPS 2022 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › continual domain adaptation
gradual domain adaptation |
0.6 | 1 | 2022 | Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain Adaptation · NeurIPS 2022 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.6 | 1 | 2022 | Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training · IJCAI 2022 |
Machine learning › Reinforcement learning
robust reinforcement learning |
0.6 | 1 | 2022 | Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training · IJCAI 2022 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
stackelberg game |
0.6 | 1 | 2022 | Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training · IJCAI 2022 |
Machine learning › Optimization for machine learning
optimal transport |
0.2 | 1 | 2022 | Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain Adaptation · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
domain knowledge integration · 0.9black-box optimization · 0.9wasserstein barycenter · 0.6stackelberg policy gradient · 0.6optimal transport · 0.6geodesic interpolation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CaDRE: Controllable and Diverse Generation of Safety-Critical Driving Scenarios Using Real-World TrajectoriesabstractSimulation is an indispensable tool in the development and testing of autonomous vehicles (AVs), offering an efficient and safe alternative to road testing. An outstanding challenge with simulation-based testing is the generation of safety-critical scenarios, which are essential to ensure that AVs can handle rare but potentially fatal situations. This paper addresses this challenge by introducing a novel framework CaDRE, to generate realistic, diverse, and controllable safetycritical scenarios. Our approach optimizes for both the quality and diversity of scenarios by employing a unique formulation and algorithm that integrates real-world scenarios, domain knowledge, and black-box optimization. We validate the effectiveness of our framework through extensive testing in three representative types of traffic scenarios. The results demonstrate superior performance in generating diverse and highquality scenarios with greater sample efficiency than existing reinforcement learning (RL) and sampling-based methods. Peide Huang, Wenhao Ding, Benjamin Stoler, Jonathan Francis, Bingqing Chen, Ding Zhao |
ICRA | 1 |
| 2024 | In-Home Gait Abnormality Detection Through Footstep-Induced Floor Vibration Sensing and Person-Invariant Contrastive LearningabstractDetecting gait abnormalities is crucial for assessing fall risks and early identification of neuromusculoskeletal disorders such as Parkinson's and stroke. Traditional assessments in gait clinics are infrequent and pose barriers, particularly for disadvantaged populations. Previous efforts have explored sensor-based approaches for in-home gait assessments, yet they face limitations such as visual obstructions (cameras), limited coverage (pressure mats), and the need for device carrying (wearables and insoles). To overcome these limitations, we introduce an in-home gait abnormality detection system using footstep-induced floor vibrations, enabling low-cost, non-intrusive, device-free gait health monitoring. The main research challenge is the high uncertainty in floor vibrations due to gait variations among people, making it challenging to develop a generalizable model for new patients. To address this, we analyze time-frequency-domain features of floor vibration data during specific gait phases and develop a feature transformation method through contrastive learning to address the between-people gait variation challenge. Our method transforms the features from vibrations to an embedding space where samples from different people stay close to each other (robust to people variation) while normal and abnormal gait samples are far apart (sensitive to gait abnormalities). Then, gait abnormalities are detected by a downstream classifier after feature transformation. We evaluated our approach through a real-world walking experiment with 21 participants and achieved an 85% to 95% mean accuracy in detecting various gait abnormalities. This novel method overcomes prior limitations in in-home gait assessments, offering accessible gait abnormality detection without the need for intrusive devices or labels for new patients. Yiwen Dong 0001, Sung Eun Kim, Kornel Schadl, Peide Huang, Wenhao Ding, Jessica Rose, Hae Young Noh |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Group Distributionally Robust Reinforcement Learning with Hierarchical Latent VariablesabstractOne key challenge for multi-task Reinforcement learning (RL) in practice is the absence of task specifications. Robust RL has been applied to deal with task ambiguity but may result in over-conservative policies. To balance the worst-case (robustness) and average performance, we propose Group Distributionally Robust Markov Decision Process (GDR-MDP), a flexible hierarchical MDP formulation that encodes task groups via a latent mixture model. GDR-MDP identifies the optimal policy that maximizes the expected return under the worst-possible qualified belief over task groups within an ambiguity set. We rigorously show that GDR-MDP’s hierarchical structure improves distributional robustness by adding regularization to the worst possible outcomes. We then develop deep RL algorithms for GDR-MDP for both value-based and policy-based RL methods. Extensive experiments on Box2D control tasks, MuJoCo benchmarks, and Google football platforms show that our algorithms outperform classic robust training algorithms across diverse environments in terms of robustness under belief uncertainties. Demos are available on our project page (https://sites.google.com/view/gdr-rl/home). Mengdi Xu, Peide Huang, Yaru Niu, Visak Kumar, Jielin Qiu, Kuan-Hui Lee, Xuewei Qi, Henry Lam, Bo Li 0026, Ding Zhao |
AISTATS | 2 |
| 2023 | Cardiac Disease Diagnosis on Imbalanced Electrocardiography Data Through Optimal Transport AugmentationabstractIn this paper, we focus on a new method of data augmentation to solve the data imbalance problem within imbalanced ECG datasets to improve the robustness and accuracy of heart disease detection. By using Optimal Transport, we augment the ECG disease data from normal ECG beats to balance the data among different categories. We build a Multi-Feature Transformer (MF-Transformer) as our classification model, where different features are extracted from both time and frequency domains to diagnose various heart conditions. Our results demonstrate 1) the classification models’ ability to make competitive predictions on five ECG categories; 2) improvements in accuracy and robustness reflecting the effectiveness of our data augmentation method. Jielin Qiu, Mengdi Xu, Peide Huang, Michael A. Rosenberg, Douglas Weber, Emerson Liu, Ding Zhao |
ICASSP | 4 |
| 2022 | Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial TrainingabstractRobust Reinforcement Learning (RL) focuses on improving performances under model errors or adversarial attacks, which facilitates the real-life deployment of RL agents. Robust Adversarial Reinforcement Learning (RARL) is one of the most popular frameworks for robust RL. However, most of the existing literature models RARL as a zero-sum simultaneous game with Nash equilibrium as the solution concept, which could overlook the sequential nature of RL deployments, produce overly conservative agents, and induce training instability. In this paper, we introduce a novel hierarchical formulation of robust RL -- a general-sum Stackelberg game model called RRL-Stack -- to formalize the sequential nature and provide extra flexibility for robust training. We develop the Stackelberg Policy Gradient algorithm to solve RRL-Stack, leveraging the Stackelberg learning dynamics by considering the adversary's response. Our method generates challenging yet solvable adversarial environments which benefit RL agents' robust learning. Our algorithm demonstrates better training stability and robustness against different testing conditions in the single-agent robotics control and multi-agent highway merging tasks. Peide Huang, Mengdi Xu, Fei Fang 0001, Ding Zhao |
IJCAI | 1 |
| 2022 | Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event SamplingabstractEvaluating rare but high-stakes events is one of the main challenges in obtaining reliable reinforcement learning policies, especially in large or infinite state/action spaces where limited scalability dictates a prohibitively large number of testing iterations. On the other hand, a biased or inaccurate policy evaluation in a safety-critical system could potentially cause unexpected catastrophic failures during deployment. This paper proposes the Accelerated Policy Evaluation (APE) method, which simultaneously uncovers rare events and estimates the rare event probability in Markov decision processes. The APE method treats the environment nature as an adversarial agent and learns towards, through adaptive importance sampling, the zero-variance sampling distribution for the policy evaluation. Moreover, APE is scalable to large discrete or continuous spaces by incorporating function approximators. We investigate the convergence property of APE in the tabular setting. Our empirical studies show that APE can estimate the rare event probability with a smaller bias while only using orders of magnitude fewer samples than baselines in multi-agent and single-agent environments. Mengdi Xu, Peide Huang, Fengpei Li, Xuewei Qi, Kentaro Oguchi 0001, Henry Lam, Ding Zhao |
IROS | 2 |
| 2022 | Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain AdaptationabstractCurriculum Reinforcement Learning (CRL) aims to create a sequence of tasks, starting from easy ones and gradually learning towards difficult tasks. In this work, we focus on the idea of framing CRL as interpolations between a source (auxiliary) and a target task distribution. Although existing studies have shown the great potential of this idea, it remains unclear how to formally quantify and generate the movement between task distributions. Inspired by the insights from gradual domain adaptation in semi-supervised learning, we create a natural curriculum by breaking down the potentially large task distributional shift in CRL into smaller shifts. We propose GRADIENT which formulates CRL as an optimal transport problem with a tailored distance metric between tasks. Specifically, we generate a sequence of task distributions as a geodesic interpolation between the source and target distributions, which are actually the Wasserstein barycenter. Different from many existing methods, our algorithm considers a task-dependent contextual distance metric and is capable of handling nonparametric distributions in both continuous and discrete context settings. In addition, we theoretically show that GRADIENT enables smooth transfer between subsequent stages in the curriculum under certain conditions. We conduct extensive experiments in locomotion and manipulation tasks and show that our proposed GRADIENT achieves higher performance than baselines in terms of learning efficiency and asymptotic performance. Peide Huang, Mengdi Xu, Laixi Shi, Fei Fang 0001, Ding Zhao |
NeurIPS | 1 |