EDBT 2026 Demo / reviewers in the wild / expert
Zhiyu Huang
dblp:08/10083
· DBLP profile ↗
34ranked-venue papers
12as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 9 since 2021Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Computer networks · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented ReasoningabstractUnderstanding and adhering to traffic regulations is essential for autonomous vehicles to ensure safety and trustworthiness. However, traffic regulations are complex, context-dependent, and differ between regions, posing a major challenge to conventional rule-based decision-making approaches. We present an interpretable, regulation-aware decision-making framework, DriveReg, which enables autonomous vehicles to understand and adhere to region-specific traffic laws and safety guidelines. The framework integrates a Retrieval Augmented Generation (RAG)-based Traffic Regulation Retrieval Agent, which retrieves relevant rules from regulatory documents based on the current situation, and a Large Language Model (LLM)-powered Reasoning Agent that evaluates actions for legal compliance and safety. Our design emphasizes interpretability to enhance transparency and trustworthiness. To support systematic evaluation, we introduce DriveReg Scenarios Dataset, a comprehensive dataset of driving scenarios across Boston, Singapore, and Los Angeles, with both hypothesized text-based cases and real-world driving data, specifically constructed and annotated to evaluate models’ capacity for regulation understanding and reasoning. We validate our framework on the DriveReg Scenarios Dataset and real-world deployment, demonstrating strong performance and robustness across diverse environments. Tianhui Cai, Zewei Zhou, Haoxuan Ma, Seth Z. Zhao, Zhiwen Wu, Xu Han 0014, Zhiyu Huang, Jiaqi Ma 0003 |
AAAI | 8 |
| 2026 | Fluid Antenna System-Assisted Physical Layer Secret Key GenerationabstractThis paper investigates physical-layer key generation (PLKG) in multi-antenna base station systems, by leveraging a fluid antenna system (FAS) to dynamically customize radio environments. Without requiring additional nodes or extensive radio frequency (RF) chains, the FAS effectively enables adaptive antenna port selection by exploiting channel spatial correlation to enhance the secret key rate (SKR) at legitimate nodes. To comprehensively evaluate the performance of the FAS in PLKG, we propose an FAS-assisted PLKG model that integrates transmit beamforming and sparse port selection under independent and identically distributed (i.i.d.) and spatially correlated channel models, respectively. Specifically, the PLKG utilizes reciprocal channel probing to derive an approximate SKR expression based on the mutual information between legitimate channel estimates, explicitly accounting for the Eve’s channel observation under spatially correlated channel scenarios. Nonconvex optimization problems for these scenarios are formulated to maximize the SKR subject to transmit power constraints and sparse port activation. We propose an iterative algorithm by capitalizing on successive convex approximation and Cauchy-Schwarz inequality to obtain a locally optimal solution. A reweighted ℓ1-norm-based algorithm is applied to advocate for the sparse port activation of FAS-assisted PLKG. To approximate the optimal activated ports obtained by exhaustive search, a low-complexity sliding window-based port selection is proposed to substitute reweighted ℓ1-norm method based on Rayleigh-quotient analysis. Simulation results demonstrate that the FAS-assisted PLKG scheme significantly outperforms fixed antenna-assisted PLKG schemes in both environments. It is shown that the FAS achieves higher SKR with fewer RF chains through dynamic sparse port selection, which effectively reduces the resource overhead. Also, the sliding window approach closely approximates the globally optimal port selection compared to the reweighted ℓ1-norm method, rendering it suitable for practical deployments. Zhiyu Huang, Guyue Li, Hao Xu 0003, Derrick Wing Kwan Ng |
IEEE J. Sel. Areas Commun. | 1 |
| 2026 | Versatile Behavior Diffusion for Generalized Traffic Agent SimulationabstractExisting traffic simulation models often fall short in capturing the intricacies of real-world scenarios, particularly the interactive behaviors among multiple traffic participants, thereby limiting their utility in the evaluation and validation of autonomous driving systems. We introduce Versatile Behavior Diffusion (VBD), a novel traffic scenario generation framework based on diffusion generative models that synthesizes scene-consistent, realistic, and controllable multi-agent interactions. VBD achieves strong performance in closed-loop traffic simulation, generating scene-consistent agent behaviors that reflect complex agent interactions. A key capability of VBD is inference-time scenario editing through multi-step refinement, guided by behavior priors and model-based optimization objectives, enabling flexible and controllable behavior generation. Despite being trained on real-world traffic datasets with only normal conditions, we introduce conflict-prior and game-theoretic guidance approaches. These approaches enable the generation of interactive, customizable, or long-tail safety-critical scenarios, which are essential for comprehensive testing and validation of autonomous driving systems. Extensive experiments validate the effectiveness and versatility of VBD and highlight its promise as a foundational tool for advancing traffic simulation and autonomous vehicle development. Project website:https://sites.google.com/view/versatile-behavior-diffusion Zhiyu Huang, Zixu Zhang, Ameya Vaidya, Yuxiao Chen 0001, Jaime Fernández Fisac, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and PredictionabstractVehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception, which fuses agents' information across different spatial locations but ignores temporal cues and temporal tasks (e.g., temporal perception and prediction). In this paper, we focus on the spatio-temporal fusion in V2X scenarios and design one-step and multi-step communication strategies (when to transmit) as well as examine their integration with three fusion strategies - early, late, and intermediate (what to transmit), providing comprehensive benchmarks with 11 fusion models (how to fuse). Furthermore, we propose V2XPnP, a novel intermediate fusion framework within one-step communication for end-to-end perception and prediction. Our framework employs a unified Transformer-based architecture to effectively model complex spatio-temporal relationships across multiple agents, frames, and high-definition maps. Moreover, we introduce the V2XPnP Sequential Dataset that supports all V2X collaboration modes and addresses the limitations of existing real-world datasets, which are restricted to single-frame or single-mode cooperation. Extensive experiments demonstrate that our framework outperforms state-of-the-art methods in both perception and prediction tasks. Zewei Zhou, Hao Xiang 0001, Zhaoliang Zheng, Seth Z. Zhao, Mingyue Lei, Tianhui Cai, Johnson Liu, Maheswari Bajji, Xin Xia 0007, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
ICCV | 12 |
| 2025 | TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and PredictionabstractEnd-to-end training of multi-agent systems offers significant advantages in improving multi-task performance. However, training such models remains challenging and requires extensive manual design and monitoring. In this work, we introduce TurboTrain, a novel and efficient training framework for multi-agent perception and prediction. TurboTrain comprises two key components: a multi-agent spatiotemporal pretraining scheme based on masked reconstruction learning and a balanced multi-task learning strategy based on gradient conflict suppression. By streamlining the training process, our framework eliminates the need for manually designing and tuning complex multi-stage training pipelines, substantially reducing training time and improving performance. We evaluate TurboTrain on a real-world cooperative driving dataset, V2XPnP-Seq, and demonstrate that it further improves the performance of state-of-the-art multi-agent perception and prediction models. Our results highlight that pretraining effectively captures spatiotemporal multi-agent features and significantly benefits downstream tasks. Moreover, the proposed balanced multi-task learning strategy enhances detection and prediction. Zewei Zhou, Seth Z. Zhao, Tianhui Cai, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
ICCV | 4 |
| 2025 | Gen-Drive: Enhancing Diffusion Generative Driving Policies with Reward Modeling and Reinforcement Learning Fine-TuningabstractAutonomous driving necessitates the ability to reason about future interactions between traffic agents and to make informed evaluations for planning. This paper introduces the Gen-Drive framework, which shifts from the traditional prediction and deterministic planning framework to a generation-then-evaluation planning paradigm. The framework employs a behavior diffusion model as a scene generator to produce diverse possible future scenarios, thereby enhancing the capability for joint interaction reasoning. To facilitate decision-making, we propose a scene evaluator (reward) model, trained with pairwise preference data collected through VLM assistance, thereby reducing human workload and enhancing scalability. Furthermore, we utilize an RL fine-tuning framework to improve the generation quality of the diffusion model, rendering it more effective for planning tasks. We conduct training and closed-loop planning tests on the nuPlan dataset, and the results demonstrate that employing such a generation-then-evaluation strategy outperforms other learning-based approaches. Additionally, the fine-tuned generative driving policy shows significant enhancements in planning performance. We further demonstrate that utilizing our learned reward model for evaluation or RL fine-tuning leads to better planning performance compared to relying on human-designed rewards. Project website: https://mczhi.github.io/GenDrive. Zhiyu Huang, Xinshuo Weng, Maximilian Igl, Yuxiao Chen 0008, Boris Ivanovic, Marco Pavone 0001, Chen Lv 0001 |
ICRA | 1 |
| 2025 | AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningabstractRecent advancements in Vision-Language-Action (VLA) models have shown promise for end-to-end autonomous driving by leveraging world knowledge and reasoning capabilities. However, current VLA models often struggle with physically infeasible action outputs, complex model structures, or unnecessarily long reasoning. In this paper, we propose AutoVLA, a novel VLA model that unifies reasoning and action generation within a single autoregressive generation model for end-to-end autonomous driving. AutoVLA performs semantic reasoning and trajectory planning directly from raw visual inputs and language instructions. We tokenize continuous trajectories into discrete, feasible actions, enabling direct integration into the language model. For training, we employ supervised fine-tuning to equip the model with dual thinking modes: fast thinking (trajectory-only) and slow thinking (enhanced with chain-of-thought reasoning). To further enhance planning performance and efficiency, we introduce a reinforcement fine-tuning method based on Group Relative Policy Optimization (GRPO), reducing unnecessary reasoning in straightforward scenarios. Extensive experiments across real-world and simulated datasets and benchmarks, including nuPlan, nuScenes, Waymo, and CARLA, demonstrate the competitive performance of AutoVLA in both open-loop and closed-loop settings. Qualitative results showcase the adaptive reasoning and accurate planning capabilities of AutoVLA in diverse scenarios. Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Zhiyu Huang, Bolei Zhou, Jiaqi Ma 0003 |
NeurIPS | 5 |
| 2025 | Online 3D Trajectory and Transmit Power Optimization for Securing UAV-Assisted Full-Duplex Communication NetworkabstractIn this paper, we investigate a full-duplex (FD) UAV-assisted multi-user system under multiple malicious jammers and propose a robust online scheme for secure communication with mobile downlink and uplink users. A random mobility model is adopted to simulate user movement, and the problem is formulated as a two-stage online optimization framework comprising a present-point and a prediction-point problem. To tackle the non-convexity, we develop inner-approximation algorithms using successive convex approximation (SCA) and the S-procedure. Simulation results verify the effectiveness of the proposed method. Zhiyu Huang, Yi Wang 0011, Ali A. Nasir, Zhichao Sheng |
VTC2025-Fall | 1 |
| 2025 | Hybrid-Prediction Integrated Planning for Autonomous DrivingabstractAutonomous driving systems require a comprehensive understanding and accurate prediction of the surrounding environment to facilitate informed decision-making in complex scenarios. Recent advances in learning-based systems have highlighted the importance of integrating prediction and planning. However, this integration poses significant alignment challenges through consistency between prediction patterns, to interaction between future prediction and planning. To address these challenges, we introduce a Hybrid-Prediction integrated Planning (HPP) framework, which operates through three novel modules collaboratively. First, we introduce marginal-conditioned occupancy prediction to align joint occupancy with agent-specific motion forecasting. Our proposed MS-OccFormer module achieves spatial-temporal alignment with motion predictions across multiple granularities. Second, we propose a game-theoretic motion predictor, GTFormer, to model the interactive dynamics among agents based on their joint predictive awareness. Third, hybrid prediction patterns are concurrently integrated into the Ego Planner and optimized by prediction guidance. The HPP framework establishes state-of-the-art performance on the nuScenes dataset, demonstrating superior accuracy and safety in end-to-end configurations. Moreover, HPP's interactive open-loop and closed-loop planning performance are demonstrated on the Waymo Open Motion Dataset (WOMD) and CARLA benchmark, outperforming existing integrated pipelines by achieving enhanced consistency between prediction and planning. Zhiyu Huang, Wenhui Huang 0001, Haohan Yang, Xiaoyu Mo, Chen Lv 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | DTPP: Differentiable Joint Conditional Prediction and Cost Evaluation for Tree Policy Planning in Autonomous DrivingabstractMotion prediction and cost evaluation are vital components in the decision-making system of autonomous vehicles. However, existing methods often ignore the importance of cost learning and treat them as separate modules. In this study, we employ a tree-structured policy planner and propose a differentiable joint training framework for both ego-conditioned prediction and cost models, resulting in a direct improvement of the final planning performance. For conditional prediction, we introduce a query-centric Transformer model that performs efficient ego-conditioned motion prediction. For planning cost, we propose a learnable context-aware cost function with latent interaction features, facilitating differentiable joint learning. We validate our proposed approach using the real-world nuPlan dataset and its associated planning test platform. Our framework not only matches state-of-the-art planning methods but outperforms other learning-based methods in planning quality, while operating more efficiently in terms of runtime. We show that joint training delivers significantly better performance than separate training of the two modules. Additionally, we find that tree-structured policy planning outperforms the conventional single-stage planning approach. Code is available: https://github.com/MCZhi/DTPP. Zhiyu Huang, Péter Karkus, Boris Ivanovic, Yuxiao Chen 0008, Marco Pavone 0001, Chen Lv 0001 |
ICRA | 1 |
| 2024 | Scalable Traffic Simulation for Autonomous Driving via Multi-Agent Goal Assignment and Autoregressive Goal-Directed PlanningabstractSimulation provides a fast, cost-effective, and secure environment for developing autonomous driving systems. However, mitigating the gap between simulation and reality is a challenging task as it demands a behavior simulation method that is human-like, diverse, controllable, socially consistent, and scalable. This work proposes a data-driven traffic agent simulation method to address the aforementioned challenges. Our approach centers around a graph-based scene representation and an encoding method, dividing the simulation into two stages: Multi-Agent Goal assignment (MAG) and Goal-Directed Planning (GDP). Firstly, we create joint goal sets for all agents involved in the scenario. Subsequently, we assign target centerlines (TCLs) to each agent based on their predicted goals. To account for any potential mismatch between the predicted joint goal sets and the road structure, we further align the goals of each agent with their respective assigned TCLs. These on-TCL goals serve as inputs for our interactive autoregressive Goal-Directed Planner (AR-GDP), constituting the second stage of our method that generates roll-outs for simulations. Evaluation results on the leaderboard of the Waymo Open Sim Agents Challenge (WOSAC) 2023 show the competitiveness of the proposed method. Xiaoyu Mo, Zhiyu Huang, Jianwu Fang, Jianru Xue, Chen Lv 0001 |
IV | 3 |
| 2024 | NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and BenchmarkingabstractBenchmarking vision-based driving policies is challenging. On one hand, open-loop evaluation with real data is easy, but these results do not reflect closed-loop performance. On the other, closed-loop evaluation is possible in simulation, but is hard to scale due to its significant computational demands. Further, the simulators available today exhibit a large domain gap to real data. This has resulted in an inability to draw clear conclusions from the rapidly growing body of research on end-to-end autonomous driving. In this paper, we present NAVSIM, a middle ground between these evaluation paradigms, where we use large datasets in combination with a non-reactive simulator to enable large-scale real-world benchmarking. Specifically, we gather simulation-based metrics, such as progress and time to collision, by unrolling bird's eye view abstractions of the test scenes for a short simulation horizon. Our simulation is non-reactive, i.e., the evaluated policy and environment do not influence each other. As we demonstrate empirically, this decoupling allows open-loop metric computation while being better aligned with closed-loop evaluations than traditional displacement errors. NAVSIM enabled a new competition held at CVPR 2024, where 143 teams submitted 463 entries, resulting in several new insights. On a large set of challenging scenarios, we observe that simple methods with moderate compute requirements such as TransFuser can match recent large-scale end-to-end driving architectures such as UniAD. Our modular framework can potentially be extended with new datasets, data curation strategies, and metrics, and will be continually maintained to host future challenges. Our code is available at https://github.com/autonomousvision/navsim. Daniel Dauner, Marcel Hallgarten, Tianyu Li 0004, Xinshuo Weng, Zhiyu Huang, Zetong Yang, Hongyang Li 0001, Igor Gilitschenski, Boris Ivanovic, Marco Pavone 0001, Andreas Geiger 0001, Kashyap Chitta |
NeurIPS | 5 |
| 2024 | Fear-Neuro-Inspired Reinforcement Learning for Safe Autonomous DrivingabstractEnsuring safety and achieving human-level driving performance remain challenges for autonomous vehicles, especially in safety-critical situations. As a key component of artificial intelligence, reinforcement learning is promising and has shown great potential in many complex tasks; however, its lack of safety guarantees limits its real-world applicability. Hence, further advancing reinforcement learning, especially from the safety perspective, is of great importance for autonomous driving. As revealed by cognitive neuroscientists, the amygdala of the brain can elicit defensive responses against threats or hazards, which is crucial for survival in and adaptation to risky environments. Drawing inspiration from this scientific discovery, we present a fear-neuro-inspired reinforcement learning framework to realize safe autonomous driving through modeling the amygdala functionality. This new technique facilitates an agent to learn defensive behaviors and achieve safe decision making with fewer safety violations. Through experimental tests, we show that the proposed approach enables the autonomous driving agent to attain state-of-the-art performance compared to the baseline agents and perform comparably to 30 certified human drivers, across various safety-critical scenarios. The results demonstrate the feasibility and effectiveness of our framework while also shedding light on the crucial role of simulating the amygdala function in the application of reinforcement learning to safety-critical autonomous driving domains. Xiangkun He, Jingda Wu, Zhiyu Huang, Zhongxu Hu, Jun Wang 0012, Alberto L. Sangiovanni-Vincentelli, Chen Lv 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Safety-Aware Human-in-the-Loop Reinforcement Learning With Shared Control for Autonomous DrivingabstractThe learning from intervention (LfI) approach has been proven effective in improving the performance of RL algorithms; nevertheless, existing methodologies in this domain tend to operate under the assumption that human guidance is invariably devoid of risk, thereby possibly leading to oscillations or even divergence in RL training as a result of improper demonstrations. In this paper, we propose a safety-aware human-in-the-loop reinforcement learning (SafeHIL-RL) approach to bridge the abovementioned gap. We first present a safety assessment module based on the artificial potential field (APF) model that incorporates dynamic information of the environment under the Frenet coordinate system, which we call the Frenet-based dynamic potential field (FDPF), for evaluating the real-time safety throughout the intervention process. Subsequently, we propose a curriculum guidance mechanism inspired by the pedagogical principle of whole-to-part patterns in human education. The curriculum guidance facilitates the RL agent’s early acquisition of comprehensive global information through continual guidance while also allowing for fine-tuning local behavior through intermittent human guidance through a human-AI shared control strategy. Consequently, our approach enables a safe, robust, and efficient reinforcement learning process independent of the quality of guidance human participants provide. The proposed method is validated in two highway autonomous driving scenarios under highly dynamic traffic flows (https://github.com/OscarHuangWind/Safe-Human-in-the-Loop-RL). The experiments’ results confirm the superiority and generalization capability of our approach when compared to other state-of-the-art (SOTA) baselines, as well as the effectiveness of the curriculum guidance. Wenhui Huang 0001, Zhiyu Huang, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Map-Adaptive Multimodal Trajectory Prediction via Intention-Aware Unimodal Trajectory PredictorsabstractAutonomous vehicles necessitate prediction of future motions of surrounding traffic participants for safe navigation. However, prediction is challenging due to complex road structures and the multimodality of driving behaviors. Recent approaches usually output a fixed number of predicted trajectories, thus can hardly generalize to situations with more options and match modalities in different situations. They also rely on complex loss design and time-consuming training. This work proposes a novel map-adaptive multimodal trajectory predictor that links driving modalities, driver’s intentions, and a vehicle’s candidate centerlines (CCLs) together, rendering the predictor map-adaptive and the multimodality explainable. The predictor is derived by training an intention-aware unimodal trajectory predictor, which consists of aCCL-based goal predictorand agoal-directed trajectory completer, and aCCL scorerfor estimating the possibilities of a target vehicle (TV) choosing a CCL to follow. This decomposed approach simplifies the training process and reduces the computational resources required, thereby rendering it a faster and more cost-effective alternative. Additionally, the proposed predictor has demonstrated comparable or even superior performance to traditional multimodal predictors in specific applications. Overall, the unimodal predictor presents a promising approach for practical machine learning applications, particularly when computational resources are limited. Xiaoyu Mo, Zhiyu Huang, Xiuxian Li, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Differentiable Integrated Motion Prediction and Planning With Learnable Cost Function for Autonomous DrivingabstractPredicting the future states of surrounding traffic participants and planning a safe, smooth, and socially compliant trajectory accordingly are crucial for autonomous vehicles (AVs). There are two major issues with the current autonomous driving system: the prediction module is often separated from the planning module, and the cost function for planning is hard to specify and tune. To tackle these issues, we propose a differentiable integrated prediction and planning (DIPP) framework that can also learn the cost function from data. Specifically, our framework uses a differentiable nonlinear optimizer as the motion planner, which takes as input the predicted trajectories of surrounding agents given by the neural network and optimizes the trajectory for the AV, enabling all operations to be differentiable, including the cost function weights. The proposed framework is trained on a large-scale real-world driving dataset to imitate human driving trajectories in the entire driving scene and validated in both open-loop and closed-loop manners. The open-loop testing results reveal that the proposed method outperforms the baseline methods across a variety of metrics and delivers planning-centric prediction results, allowing the planning module to output trajectories close to those of human drivers. In closed-loop testing, the proposed method outperforms various baseline methods, showing the ability to handle complex urban driving scenarios and robustness against the distributional shift. Importantly, we find that joint training of planning and prediction modules achieves better performance than planning with a separate trained prediction module in both open-loop and closed-loop tests. Moreover, the ablation study indicates that the learnable components in the framework are essential to ensure planning stability and performance. Code and Supplementary Videos are available at https://mczhi.github.io/DIPP/. Zhiyu Huang, Jingda Wu, Chen Lv 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Prioritized Experience-Based Reinforcement Learning With Human Guidance for Autonomous DrivingabstractReinforcement learning (RL) requires skillful definition and remarkable computational efforts to solve optimization and control problems, which could impair its prospect. Introducing human guidance into RL is a promising way to improve learning performance. In this article, a comprehensive human guidance-based RL framework is established. A novel prioritized experience replay mechanism that adapts to human guidance in the RL process is proposed to boost the efficiency and performance of the RL algorithm. To relieve the heavy workload on human participants, a behavior model is established based on an incremental online learning method to mimic human actions. We design two challenging autonomous driving tasks for evaluating the proposed algorithm. Experiments are conducted to access the training and testing performance and learning mechanism of the proposed algorithm. Comparative results against the state-of-the-art methods suggest the advantages of our algorithm in terms of learning efficiency, performance, and robustness. Jingda Wu, Zhiyu Huang, Wenhui Huang 0001, Chen Lv 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | GameFormer: Game-theoretic Modeling and Learning of Transformer-based Interactive Prediction and Planning for Autonomous DrivingabstractAutonomous vehicles operating in complex real-world environments require accurate predictions of interactive behaviors between traffic participants. This paper tackles the interaction prediction problem by formulating it with hierarchical game theory and proposing the GameFormer model for its implementation. The model incorporates a Transformer encoder, which effectively models the relationships between scene elements, alongside a novel hierarchical Transformer decoder structure. At each decoding level, the decoder utilizes the prediction outcomes from the previous level, in addition to the shared environmental context, to iteratively refine the interaction process. Moreover, we propose a learning process that regulates an agent’s behavior at the current level to respond to other agents’ behaviors from the preceding level. Through comprehensive experiments on large-scale real-world driving datasets, we demonstrate the state-of-the-art accuracy of our model on the Waymo interaction prediction task. Additionally, we validate the model’s capacity to jointly reason about the motion plan of the ego agent and the behaviors of multiple agents in both open-loop and closed-loop planning tests, outperforming various baseline methods. Furthermore, we evaluate the efficacy of our model on the nuPlan planning benchmark, where it achieves leading performance. Project website: https://mczhi.github.io/GameFormer/ Zhiyu Huang |
ICCV | 1 |
| 2023 | Multi-modal Hierarchical Transformer for Occupancy Flow Field Prediction in Autonomous DrivingabstractForecasting the future states of surrounding traffic participants is a crucial capability for autonomous vehicles. The recently proposed occupancy flow field prediction introduces a scalable and effective representation to jointly predict surrounding agents' future motions in a scene. However, the challenging part is to model the underlying social interactions among traffic agents and the relations between occupancy and flow. Therefore, this paper proposes a novel Multi-modal Hierarchical Transformer network that fuses the vectorized (agent motion) and visual (scene flow, map, and occupancy) modalities and jointly predicts the flow and occupancy of the scene. Specifically, visual and vector features from sensory data are encoded through a multi-stage Transformer module and then a late-fusion Transformer module with temporal pixel-wise attention. Importantly, a flow-guided multi-head self-attention (FG-MSA) module is designed to better aggregate the information on occupancy and flow and model the mathematical relations between them. The proposed method is comprehensively validated on the Waymo Open Motion Dataset and compared against several state-of-the-art models. The results reveal that our model with much more compact architecture and data inputs than other methods can achieve comparable performance. We also demonstrate the effectiveness of incorporating vectorized agent motion features and the proposed FG-MSA module. Compared to the ablated model without the FG-MSA module, which won 2ndplace in the 2022 Waymo Occupancy and Flow Prediction Challenge, the current model shows better separability for flow and occupancy and further performance improvements. Zhiyu Huang |
ICRA | 2 |
| 2023 | UAV-Assisted Downlink-and-Uplink Communication in the Presence of Multiple Malicious JammersabstractThis paper investigates the unmanned aerial vehicle (UAV)-assisted communication network with multiple downlink users (DLUs) and uplink users (ULUs) in the presence of multiple malicious jammers. To guarantee fairness among the users and their uplink and downlink communication throughput, we aim to maximize the minimum average throughput by jointly optimizing the scheduling of ULUs/DLUs, three dimensional (3D) trajectory and the UAV transmission power. Although the optimization problem is computationally intractable due to its non-convexity, we develop an iterative algorithm based on the block coordinate descend approach and the successive convex approximation technique to solve the problem efficiently. Numerical outcomes show that our proposed algorithm can improve throughput significantly over several benchmark schemes. Zhiyu Huang, Zhichao Sheng, Ali A. Nasir, Antonino Masaracchia |
WCNC | 1 |
| 2023 | Human-Guided Reinforcement Learning With Sim-to-Real Transfer for Autonomous NavigationabstractReinforcement learning (RL) is a promising approach in unmanned ground vehicles (UGVs) applications, but limited computing resource makes it challenging to deploy a well-behaved RL strategy with sophisticated neural networks. Meanwhile, the training of RL on navigation tasks is difficult, which requires a carefully-designed reward function and a large number of interactions, yet RL navigation can still fail due to many corner cases. This shows the limited intelligence of current RL methods, thereby prompting us to rethink combining RL with human intelligence. In this paper, a human-guided RL framework is proposed to improve RL performance both during learning in the simulator and deployment in the real world. The framework allows humans to intervene in RL's control progress and provide demonstrations as needed, thereby improving RL's capabilities. An innovative human-guided RL algorithm is proposed that utilizes a series of mechanisms to improve the effectiveness of human guidance, including human-guided learning objective, prioritized human experience replay, and human intervention-based reward shaping. Our RL method is trained in simulation and then transferred to the real world, and we develop a denoised representation for domain adaptation to mitigate the simulation-to-real gap. Our method is validated through simulations and real-world experiments to navigate UGVs in diverse and dynamic environments based only on tiny neural networks and image inputs. Our method performs better in goal-reaching and safety than existing learning- and model-based navigation approaches and is robust to changes in input features and ego kinetics. Furthermore, our method allows small-scale human demonstrations to be used to improve the trained RL agent and learn expected behaviors online. Jingda Wu, Yanxin Zhou, Haohan Yang, Zhiyu Huang, Chen Lv 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Conditional Predictive Behavior Planning With Inverse Reinforcement Learning for Human-Like Autonomous DrivingabstractMaking safe and human-like decisions is an essential capability of autonomous driving systems, and learning-based behavior planning presents a promising pathway toward achieving this objective. Distinguished from existing learning-based methods that directly output decisions, this work introduces a predictive behavior planning framework that learns to predict and evaluate from human driving data. This framework consists of three components: a behavior generation module that produces a diverse set of candidate behaviors in the form of trajectory proposals, a conditional motion prediction network that predicts future trajectories of other agents based on each proposal, and a scoring module that evaluates the candidate plans using maximum entropy inverse reinforcement learning (IRL). We validate the proposed framework on a large-scale real-world urban driving dataset through comprehensive experiments. The results show that the conditional prediction model can predict distinct and reasonable future trajectories given different trajectory proposals and the IRL-based scoring module can select plans that are close to human driving. The proposed framework outperforms other baseline methods in terms of similarity to human driving trajectories. Additionally, we find that the conditional prediction model improves both prediction and planning performance compared to the non-conditional model. Lastly, we note that the learning of the scoring module is crucial for aligning the evaluations with human drivers. Zhiyu Huang, Jingda Wu, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Efficient Deep Reinforcement Learning With Imitative Expert Priors for Autonomous DrivingabstractDeep reinforcement learning (DRL) is a promising way to achieve human-like autonomous driving. However, the low sample efficiency and difficulty of designing reward functions for DRL would hinder its applications in practice. In light of this, this article proposes a novel framework to incorporate human prior knowledge in DRL, in order to improve the sample efficiency and save the effort of designing sophisticated reward functions. Our framework consists of three ingredients, namely, expert demonstration, policy derivation, and RL. In the expert demonstration step, a human expert demonstrates their execution of the task, and their behaviors are stored as state-action pairs. In the policy derivation step, the imitative expert policy is derived using behavioral cloning and uncertainty estimation relying on the demonstration data. In the RL step, the imitative expert policy is utilized to guide the learning of the DRL agent by regularizing the KL divergence between the DRL agent's policy and the imitative expert policy. To validate the proposed method in autonomous driving applications, two simulated urban driving scenarios (unprotected left turn and roundabout) are designed. The strengths of our proposed method are manifested by the training results as our method can not only achieve the best performance but also significantly improve the sample efficiency in comparison with the baseline algorithms (particularly 60% improvement compared with soft actor-critic). In testing conditions, the agent trained by our method obtains the highest success rate and shows diverse and human-like driving behaviors as demonstrated by the human expert. We also find that using the imitative expert policy trained with the ensemble method that estimates both policy and model uncertainties, as well as increasing the training sample size, can result in better training and testing performance, especially for more difficult tasks. As a result, the proposed method has shown its potential to facilitate the applications of DRL-enabled human-like autonomous driving systems in practice. The code and supplementary videos are also provided. [https://mczhi.github.io/Expert-Prior-RL/]. Zhiyu Huang, Jingda Wu, Chen Lv 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Multi-modal Motion Prediction with Transformer-based Neural Network for Autonomous DrivingabstractPredicting the behaviors of other agents on the road is critical for autonomous driving to ensure safety and efficiency. However, the challenging part is how to represent the social interactions between agents and output different possible trajectories with interpretability. In this paper, we introduce a neural prediction framework based on the Transformer structure to model the relationship among the interacting agents and extract the attention of the target agent on the map waypoints. Specifically, we organize the interacting agents into a graph and utilize the multi-head attention Transformer encoder to extract the relations between them. To address the multi-modality of motion prediction, we propose a multi-modal attention Transformer encoder, which modifies the multi-head attention mechanism to multi-modal attention, and each predicted trajectory is conditioned on an independent attention mode. The proposed model is validated on the Argoverse motion forecasting dataset and shows state-of-the-art prediction accuracy while maintaining a small model size and a simple training process. We also demonstrate that the multi-modal attention module can automatically identify different modes of the target agent's attention on the map, which improves the interpretability of the model. Zhiyu Huang, Xiaoyu Mo |
ICRA | 1 |
| 2022 | Improved Deep Reinforcement Learning with Expert Demonstrations for Urban Autonomous DrivingabstractLearning-based approaches, such as reinforcement learning (RL) and imitation learning (IL), have indicated superiority over rule-based approaches in complex urban autonomous driving environments, showing great potential to make intelligent decisions. However, current RL and IL approaches still have their own drawbacks, such as low data efficiency for RL and poor generalization capability for IL. In light of this, this paper proposes a novel learning-based method that combines deep reinforcement learning and imitation learning from expert demonstrations, which is applied to longitudinal vehicle motion control in autonomous driving scenarios. Our proposed method employs the soft actor-critic structure and modifies the learning process of the policy network to incorporate both the goals of maximizing reward and imitating the expert. Moreover, an adaptive prioritized experience replay is designed to sample experience from both the agent’s self-exploration and expert demonstration, in order to improve sample efficiency. The proposed method is validated in a simulated urban roundabout scenario and compared with various prevailing RL and IL baseline approaches. The results manifest that the proposed method has a faster training speed, as well as better performance in navigating safely and time-efficiently. Zhiyu Huang, Jingda Wu |
IV | 2 |
| 2022 | Driving Behavior Modeling Using Naturalistic Human Driving Data With Inverse Reinforcement LearningabstractDriving behavior modeling is of great importance for designing safe, smart, and personalized autonomous driving systems. In this paper, an internal reward function-based driving model that emulates the human’s decision-making mechanism is utilized. To infer the reward function parameters from naturalistic human driving data, we propose a structural assumption about human driving behavior that focuses on discrete latent driving intentions. It converts the continuous behavior modeling problem to a discrete setting and thus makes maximum entropy inverse reinforcement learning (IRL) tractable to learn reward functions. Specifically, a polynomial trajectory sampler is adopted to generate candidate trajectories considering high-level intentions and approximate the partition function in the maximum entropy IRL framework. An environment model considering interactive behaviors among the ego and surrounding vehicles is built to better estimate the generated trajectories. The proposed method is applied to learn personalized reward functions for individual human drivers from the NGSIM highway driving dataset. The qualitative results demonstrate that the learned reward functions are able to explicitly express the preferences of different drivers and interpret their decisions. The quantitative results reveal that the learned reward functions are robust, which is manifested by only a marginal decline in proximity to the human driving trajectories when applying the reward function in the testing conditions. For the testing performance, the personalized modeling method outperforms the general modeling approach, significantly reducing the modeling errors in human likeness (a custom metric to gauge accuracy), and these two methods deliver better results compared to other baseline methods. Moreover, it is found that predicting the response actions of surrounding vehicles and incorporating their potential decelerations caused by the ego vehicle are critical in estimating the generated trajectories, and the accuracy of personalized planning using the learned reward functions relies on the accuracy of the forecasting model. Zhiyu Huang, Jingda Wu, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Multi-Agent Trajectory Prediction With Heterogeneous Edge-Enhanced Graph Attention NetworkabstractSimultaneous trajectory prediction for multiple heterogeneous traffic participants is essential for safe and efficient operation of connected automated vehicles under complex driving situations. Two main challenges for this task are to handle the varying number of heterogeneous target agents and jointly consider multiple factors that would affect their future motions. This is because different kinds of agents have different motion patterns, and their behaviors are jointly affected by their individual dynamics, their interactions with surrounding agents, as well as the traffic infrastructures. A trajectory prediction method handling these challenges will benefit the downstream decision-making and planning modules of autonomous vehicles. To meet these challenges, we propose a three-channel framework together with a novel Heterogeneous Edge-enhanced graph ATtention network (HEAT). Our framework is able to deal with the heterogeneity of the target agents and traffic participants involved. Specifically, agents’ dynamics are extracted from their historical states using type-specific encoders. The inter-agent interactions are represented with a directed edge-featured heterogeneous graph and processed by the designed HEAT network to extract interaction features. Besides, the map features are shared across all agents by introducing a selective gate-mechanism. And finally, the trajectories of multiple agents are predicted simultaneously. Validations using both urban and highway driving datasets show that the proposed model can realize simultaneous trajectory predictions for multiple agents under complex traffic situations, and achieve state-of-the-art performance with respect to prediction accuracy. The achieved final displacement error (FDE@3sec) is 0.66 meter under urban driving, demonstrating the feasibility and effectiveness of the proposed approach. Xiaoyu Mo, Zhiyu Huang, Yang Xing 0002, Chen Lv 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Triplet Cross-Fusion Learning for Unpaired Image Denoising in Optical Coherence TomographyabstractOptical coherence tomography (OCT) is a widely-used modality in clinical imaging, which suffers from the speckle noise inevitably. Deep learning has proven its superior capability in OCT image denoising, while the difficulty of acquiring a large number of well-registered OCT image pairs limits the developments of paired learning methods. To solve this problem, some unpaired learning methods have been proposed, where the denoising networks can be trained with unpaired OCT data. However, majority of them are modified from the cycleGAN framework. These cycleGAN-based methods train at least two generators and two discriminators, while only one generator is needed for the inference. The dual-generator and dual-discriminator structures of cycleGAN-based methods demand a large amount of computing resource, which may be redundant for OCT denoising tasks. In this work, we propose a novel triplet cross-fusion learning (TCFL) strategy for unpaired OCT image denoising. The model complexity of our strategy is much lower than those of the cycleGAN-based methods. During training, the clean components and the noise components from the triplet of three unpaired images are cross-fused, helping the network extract more speckle noise information to improve the denoising accuracy. Furthermore, the TCFL-based network which is trained with triplets can deal with limited training data scenarios. The results demonstrate that the TCFL strategy outperforms state-of-the-art unpaired methods both qualitatively and quantitatively, and even achieves denoising performance comparable with paired methods. Code is available at: https://github.com/gengmufeng/TCFL-OCT. Mufeng Geng, Xiangxi Meng 0001, Lei Zhu 0012, Mengdi Gao, Zhiyu Huang, Bin Qiu, Yibao Zhang, Qiushi Ren, Yanye Lu |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Improved Short-Term Speed Prediction Using Spatiotemporal-Vision-Based Deep Neural Network for Intelligent Fuel Cell VehiclesabstractIn this article, an improved short-term speed prediction method is proposed to predict short-term future speed and analyze future energy consumption of intelligent fuel cell vehicles. The short-term future speed is predicted by the proposed Inflated 3-D Inception long short-term memory (LSTM) network, which takes the spatiotemporal-vision information and vehicle motion states. Specifically, the spatiotemporal-vision-based deep neural network utilizes image sequences captured by a front-facing camera as environmental information and historical speed series as motion information to improve the prediction accuracy. Then, a case study of the proposed speed prediction method, with rule-based energy management strategy to calculate future energy consumption, is presented. The simulation results show that short-term speed prediction based on the Inflated 3-D Inception LSTM network can achieve high accuracy of speed prediction in various traffic densities, as well as low prediction errors of future energy consumption including the hydrogen consumption and state-of-charge attenuation. Yuanzhi Zhang 0001, Zhiyu Huang, Caizhi Z. Zhang, Chen Lv 0001, Chenghao Deng, Dong Hao, Jinrui Chen, Hongxu Ran |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Weakly Supervised Deep Learning-Based Optical Coherence Tomography AngiographyabstractOptical coherence tomography angiography (OCTA) is a promising imaging modality for microvasculature studies. Deep learning networks have been widely applied in the field of OCTA reconstruction, benefiting from its powerful mapping capability among images. However, these existing deep learning-based methods depend on high-quality labels, which are hard to acquire considering imaging hardware limitations and practical data acquisition conditions. In this article, we proposed an unprecedented weakly supervised deep learning-based pipeline for OCTA reconstruction task, in the absence of high-quality training labels. The proposed pipeline was investigated on an in vivo animal dataset and a human eye dataset by a cross-validation strategy. Compared with supervised learning approaches, the proposed approach demonstrated similar or even better performance in the OCTA reconstruction task. These investigations indicate that the proposed weakly supervised learning strategy is well capable of performing OCTA reconstruction, and has a certain potential towards clinical applications. Zhiyu Huang, Bin Qiu, Xiangxi Meng 0001, Yunfei You, Mufeng Geng, Gangjun Liu, Chuanqing Zhou, Andreas K. Maier, Qiushi Ren, Yanye Lu |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Multi-Scale Driver Behaviors Reasoning System for Intelligent Vehicles Based on a Joint Deep Learning FrameworkabstractThe mutual understanding between driver and vehicle is critically important to the design of intelligent vehicles and customized interaction interface. In this study, a deep learning-based joint driver behavior reasoning system toward multi-scale and multi-tasks behavior recognition is proposed. Specifically, a multi-scale driver behavior recognition system is designed to recognize both the driver's physical and mental states based on a deep encoder-decoder framework. The system jointly recognizes three driver behaviors, namely, mirror-checking, lane change intention, and emotions based on the shared encoder network. The encoder network is designed based on a deep convolutional neural network (CNN), and several decoders for different driver states estimation are proposed with fully connected (FC), and long short-term memory (LSTM) based recurrent neural networks (RNN), respectively. The proposed framework can be used as a solution to exploit the relationship between different driver states for intelligent vehicles towards an efficient driver-side understanding. The testing results on the Brain4Car dataset show accurate performance and outperform existing methods on driver postures, intention, and emotion recognition. Yang Xing 0002, Zhongxu Hu, Zhiyu Huang, Chen Lv 0001, Dongpu Cao, Efstathios Velenis |
SMC | 3 |
| 2017 | Optimal planning of distributed photovoltaic considering three-phase unbalanced degreeabstractThe optimal planning of distributed photovoltaic system is conducive to reducing the adverse impact of high-permeability random access on power distribution systems. In this paper, a optimization planning model of locating and sizing of distributed photovoltaic power generation system is established to improve the permeability of PV system on the premise of the safe operation of the power system, which aims at the biggest system voltage stability margin, the smallest system loss and the highest permeability of PV system, while considers the system three-phase unbalance constraint. Finally, the multi-objective particle swarm optimization algorithm based on mixed integer programming is applied to solve the optimal algorithm on the IEEE 33-node distribution system. The obtained distributed PV optimal plan is verified the rationality and practicability of the proposed planning model. Chengli Zheng, Zijian Huang 0008, Zhiyu Huang |
IECON | 5 |
| 2012 | Exploiting constructive interference for scalable flooding in wireless networksabstractExploiting constructive interference in wireless networks is an emerging trend for it allows multiple senders transmit an identical packet simultaneously. Constructive interference based flooding can realize millisecond network flooding latency and sub-microsecond time synchronization accuracy, require no network state information and adapt to topology changes. However, constructive interference has a precondition to function, namely, the maximum temporal displacement Δ of concurrent packet transmissions should be less than a given hardware constrained threshold. We disclose that constructive interference based flooding suffers the scalability problem. The packet reception performances of intermediate nodes degrade significantly as the density or the size of the network increases. We theoretically show that constructive interference based flooding has a packet reception ratio (PRR) lower bound (95:4%) in the grid topology. For a general topology, we propose the spine constructive interference based flooding (SCIF) protocol. With little overhead, SCIF floods the entire network much more reliably than Glossy [1] in high density or large-scale networks. Extensive simulations illustrate that the PRR of SCIF keeps stable above 96% as the network size grows from 400 to 4000 while the PRR of Glossy is only 26% when the size of the network is 4000. We also propose to use waveform analysis to explain the root cause of constructive interference, which is mainly examined in simulations and experiments. We further derive the closed-form PRR formula and define interference gain factor (IGF) to quantitatively measure constructive interference. Yuan He 0004, Xufei Mao, Yunhao Liu 0001, Zhiyu Huang, Xiang-Yang Li 0001 |
INFOCOM | 5 |
| 2012 | Direct multi-hop time synchronization with constructive interferenceabstractMulti-hop time synchronization in wireless sensor networks (WSNs) is often time-consuming and error-prone due to random time-stamp delays for MAC layer access and unstable clocks of intermediate nodes. Constructive interference (CI), a recently discovered physical layer phenomenon, allows multiple nodes transmit and forward an identical packet simultaneously. By leveraging CI, we propose direct multi-hop (DMH) time synchronization by directly utilizing the time-stamps from the sink node instead of intermediate nodes, which avoids the error caused by the unstable clock of intermediate nodes. DMH doesn't need decode the flooding time synchronization beacons. Moreover, DMH explores the linear regression technique in CI based time synchronization to counterbalance the clock drifts due to clock skews. Gaofeng Pan, Zhiyu Huang |
IPSN | 3 |