Peide Huang

dblp:295/8645 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Autonomous driving · 32% Reinforcement learning · 32% Trustworthy machine learning · 11%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
safety-critical scenario generation
0.912025
CaDRE: Controllable and Diverse Generation of Safety-Critical Driving Scenarios Using Real-World Trajectories · ICRA 2025
Robotics › Autonomous driving › autonomous vehicle testing
simulation-based testing
0.912025
CaDRE: Controllable and Diverse Generation of Safety-Critical Driving Scenarios Using Real-World Trajectories · ICRA 2025
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.612022
Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training · IJCAI 2022
Machine learning › Reinforcement learning
curriculum reinforcement learning
0.612022
Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain Adaptation · NeurIPS 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation › continual domain adaptation
gradual domain adaptation
0.612022
Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain Adaptation · NeurIPS 2022
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.612022
Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training · IJCAI 2022
Machine learning › Reinforcement learning
robust reinforcement learning
0.612022
Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training · IJCAI 2022
Knowledge, reasoning and agents › Multi-agent systems › game theory
stackelberg game
0.612022
Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training · IJCAI 2022
Machine learning › Optimization for machine learning
optimal transport
0.212022
Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain Adaptation · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

domain knowledge integration · 0.9black-box optimization · 0.9wasserstein barycenter · 0.6stackelberg policy gradient · 0.6optimal transport · 0.6geodesic interpolation · 0.6
YearPublicationVenuePosition
2025 CaDRE: Controllable and Diverse Generation of Safety-Critical Driving Scenarios Using Real-World Trajectories
abstract
Simulation is an indispensable tool in the development and testing of autonomous vehicles (AVs), offering an efficient and safe alternative to road testing. An outstanding challenge with simulation-based testing is the generation of safety-critical scenarios, which are essential to ensure that AVs can handle rare but potentially fatal situations. This paper addresses this challenge by introducing a novel framework CaDRE, to generate realistic, diverse, and controllable safetycritical scenarios. Our approach optimizes for both the quality and diversity of scenarios by employing a unique formulation and algorithm that integrates real-world scenarios, domain knowledge, and black-box optimization. We validate the effectiveness of our framework through extensive testing in three representative types of traffic scenarios. The results demonstrate superior performance in generating diverse and highquality scenarios with greater sample efficiency than existing reinforcement learning (RL) and sampling-based methods.
Peide Huang, Wenhao Ding, Benjamin Stoler, Jonathan Francis, Bingqing Chen, Ding Zhao
ICRA1
2024 In-Home Gait Abnormality Detection Through Footstep-Induced Floor Vibration Sensing and Person-Invariant Contrastive Learning
abstract
Detecting gait abnormalities is crucial for assessing fall risks and early identification of neuromusculoskeletal disorders such as Parkinson's and stroke. Traditional assessments in gait clinics are infrequent and pose barriers, particularly for disadvantaged populations. Previous efforts have explored sensor-based approaches for in-home gait assessments, yet they face limitations such as visual obstructions (cameras), limited coverage (pressure mats), and the need for device carrying (wearables and insoles). To overcome these limitations, we introduce an in-home gait abnormality detection system using footstep-induced floor vibrations, enabling low-cost, non-intrusive, device-free gait health monitoring. The main research challenge is the high uncertainty in floor vibrations due to gait variations among people, making it challenging to develop a generalizable model for new patients. To address this, we analyze time-frequency-domain features of floor vibration data during specific gait phases and develop a feature transformation method through contrastive learning to address the between-people gait variation challenge. Our method transforms the features from vibrations to an embedding space where samples from different people stay close to each other (robust to people variation) while normal and abnormal gait samples are far apart (sensitive to gait abnormalities). Then, gait abnormalities are detected by a downstream classifier after feature transformation. We evaluated our approach through a real-world walking experiment with 21 participants and achieved an 85% to 95% mean accuracy in detecting various gait abnormalities. This novel method overcomes prior limitations in in-home gait assessments, offering accessible gait abnormality detection without the need for intrusive devices or labels for new patients.
Yiwen Dong 0001, Sung Eun Kim, Kornel Schadl, Peide Huang, Wenhao Ding, Jessica Rose, Hae Young Noh
IEEE J. Biomed. Health Informatics4
2023 Group Distributionally Robust Reinforcement Learning with Hierarchical Latent Variables
abstract
One key challenge for multi-task Reinforcement learning (RL) in practice is the absence of task specifications. Robust RL has been applied to deal with task ambiguity but may result in over-conservative policies. To balance the worst-case (robustness) and average performance, we propose Group Distributionally Robust Markov Decision Process (GDR-MDP), a flexible hierarchical MDP formulation that encodes task groups via a latent mixture model. GDR-MDP identifies the optimal policy that maximizes the expected return under the worst-possible qualified belief over task groups within an ambiguity set. We rigorously show that GDR-MDP’s hierarchical structure improves distributional robustness by adding regularization to the worst possible outcomes. We then develop deep RL algorithms for GDR-MDP for both value-based and policy-based RL methods. Extensive experiments on Box2D control tasks, MuJoCo benchmarks, and Google football platforms show that our algorithms outperform classic robust training algorithms across diverse environments in terms of robustness under belief uncertainties. Demos are available on our project page (https://sites.google.com/view/gdr-rl/home).
Mengdi Xu, Peide Huang, Yaru Niu, Visak Kumar, Jielin Qiu, Kuan-Hui Lee, Xuewei Qi, Henry Lam, Bo Li 0026, Ding Zhao
AISTATS2
2023 Cardiac Disease Diagnosis on Imbalanced Electrocardiography Data Through Optimal Transport Augmentation
abstract
In this paper, we focus on a new method of data augmentation to solve the data imbalance problem within imbalanced ECG datasets to improve the robustness and accuracy of heart disease detection. By using Optimal Transport, we augment the ECG disease data from normal ECG beats to balance the data among different categories. We build a Multi-Feature Transformer (MF-Transformer) as our classification model, where different features are extracted from both time and frequency domains to diagnose various heart conditions. Our results demonstrate 1) the classification models’ ability to make competitive predictions on five ECG categories; 2) improvements in accuracy and robustness reflecting the effectiveness of our data augmentation method.
Jielin Qiu, Mengdi Xu, Peide Huang, Michael A. Rosenberg, Douglas Weber, Emerson Liu, Ding Zhao
ICASSP4
2022 Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training
abstract
Robust Reinforcement Learning (RL) focuses on improving performances under model errors or adversarial attacks, which facilitates the real-life deployment of RL agents. Robust Adversarial Reinforcement Learning (RARL) is one of the most popular frameworks for robust RL. However, most of the existing literature models RARL as a zero-sum simultaneous game with Nash equilibrium as the solution concept, which could overlook the sequential nature of RL deployments, produce overly conservative agents, and induce training instability. In this paper, we introduce a novel hierarchical formulation of robust RL -- a general-sum Stackelberg game model called RRL-Stack -- to formalize the sequential nature and provide extra flexibility for robust training. We develop the Stackelberg Policy Gradient algorithm to solve RRL-Stack, leveraging the Stackelberg learning dynamics by considering the adversary's response. Our method generates challenging yet solvable adversarial environments which benefit RL agents' robust learning. Our algorithm demonstrates better training stability and robustness against different testing conditions in the single-agent robotics control and multi-agent highway merging tasks.
Peide Huang, Mengdi Xu, Fei Fang 0001, Ding Zhao
IJCAI1
2022 Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event Sampling
abstract
Evaluating rare but high-stakes events is one of the main challenges in obtaining reliable reinforcement learning policies, especially in large or infinite state/action spaces where limited scalability dictates a prohibitively large number of testing iterations. On the other hand, a biased or inaccurate policy evaluation in a safety-critical system could potentially cause unexpected catastrophic failures during deployment. This paper proposes the Accelerated Policy Evaluation (APE) method, which simultaneously uncovers rare events and estimates the rare event probability in Markov decision processes. The APE method treats the environment nature as an adversarial agent and learns towards, through adaptive importance sampling, the zero-variance sampling distribution for the policy evaluation. Moreover, APE is scalable to large discrete or continuous spaces by incorporating function approximators. We investigate the convergence property of APE in the tabular setting. Our empirical studies show that APE can estimate the rare event probability with a smaller bias while only using orders of magnitude fewer samples than baselines in multi-agent and single-agent environments.
Mengdi Xu, Peide Huang, Fengpei Li, Xuewei Qi, Kentaro Oguchi 0001, Henry Lam, Ding Zhao
IROS2
2022 Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain Adaptation
abstract
Curriculum Reinforcement Learning (CRL) aims to create a sequence of tasks, starting from easy ones and gradually learning towards difficult tasks. In this work, we focus on the idea of framing CRL as interpolations between a source (auxiliary) and a target task distribution. Although existing studies have shown the great potential of this idea, it remains unclear how to formally quantify and generate the movement between task distributions. Inspired by the insights from gradual domain adaptation in semi-supervised learning, we create a natural curriculum by breaking down the potentially large task distributional shift in CRL into smaller shifts. We propose GRADIENT which formulates CRL as an optimal transport problem with a tailored distance metric between tasks. Specifically, we generate a sequence of task distributions as a geodesic interpolation between the source and target distributions, which are actually the Wasserstein barycenter. Different from many existing methods, our algorithm considers a task-dependent contextual distance metric and is capable of handling nonparametric distributions in both continuous and discrete context settings. In addition, we theoretically show that GRADIENT enables smooth transfer between subsequent stages in the curriculum under certain conditions. We conduct extensive experiments in locomotion and manipulation tasks and show that our proposed GRADIENT achieves higher performance than baselines in terms of learning efficiency and asymptotic performance.
Peide Huang, Mengdi Xu, Laixi Shi, Fei Fang 0001, Ding Zhao
NeurIPS1