Borong Zhang

dblp:336/4144 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 54% Language models and text generation · 26% Trustworthy machine learning · 13%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
safe reinforcement learning
2.442025
OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research · J. Mach. Learn. Res. 2024
SafeDreamer: Safe Reinforcement Learning with World Models · ICLR 2024
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023
Machine learning › Trustworthy machine learning › AI safety
safety alignment
1.722025
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning · NeurIPS 2025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Natural language and speech › Language models and text generation
alignment
1.622025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Machine learning › Reinforcement learning
model-based reinforcement learning
1.122026
SafeDreamer: Safe Reinforcement Learning with World Models · ICLR 2024
Latent State-Predictive Exploration for Deep Reinforcement Learning · AAAI 2026
Machine learning › Reinforcement learning
exploration
1.012026
Latent State-Predictive Exploration for Deep Reinforcement Learning · AAAI 2026
Machine learning › Reinforcement learning › exploration
novelty-based exploration
1.012026
Latent State-Predictive Exploration for Deep Reinforcement Learning · AAAI 2026
Machine learning › Reinforcement learning
constrained reinforcement learning
0.912025
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning · NeurIPS 2025
Robotics › Robot manipulation
mobile manipulation
0.912025
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning · NeurIPS 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
0.912025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.812024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model
0.812024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.812024
SafeDreamer: Safe Reinforcement Learning with World Models · ICLR 2024
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments
0.712023
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process
0.312025
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning · NeurIPS 2025
Machine learning › Trustworthy machine learning
robustness
0.312025
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference · ACL (1) 2025
Machine learning › Transfer learning and domain adaptation › foundation model adaptation
model-agnostic adaptation
0.212024
Aligner: Efficient Alignment by Learning to Correct · NeurIPS 2024
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.212023
Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

state encoder · 1.0self-predictive representation learning · 1.0diffusion model · 1.0safe reinforcement learning · 0.9reinforcement learning from human feedback · 0.9min-max optimization · 0.9constrained markov decision process · 0.9world model planning · 0.8lagrangian method · 0.8dreamer · 0.8
YearPublicationVenuePosition
2026 Latent State-Predictive Exploration for Deep Reinforcement Learning
abstract
Reinforcement learning (RL) has achieved promising results in continuous control tasks, where efficient exploration of the state space is crucial for success. However, many recent RL approaches still struggle with sample inefficiency and insufficient exploration for long-horizon tasks, particularly in environments characterized by high-dimensional and complex state spaces. To address these challenges, we propose a novel exploration framework, Latent State Predictive Exploration (LSPE). The core idea behind LSPE is to endow the agent with a form of ``foresight" to enhance exploration in long-horizon settings. Specifically, LSPE employs a state encoder to learn compact latent representations from high-dimensional visual observations, effectively filtering out irrelevant or noisy information. To further enrich and stabilize these representations, we incorporate a diffusion-based self-predictive module that enforces temporal consistency by predicting future states, thereby improving both exploration and downstream predictive control. Additionally, we introduce an Exploration Reward Function (ERF) that explicitly encourages the agent to visit novel latent states. This reward signal promotes more efficient and scalable exploration in complex environments. We evaluate LSPE across a diverse set of challenging long-horizon navigation and manipulation tasks, spanning simulation environments such as Habitat and Robosuite, as well as deployment on a real robot in a **physical indoor environment**. Experimental results show that LSPE substantially enhances exploration efficiency and scales effectively to complex, high-dimensional tasks.
Kaiyan Zhao, Borong Zhang, Yan Li 0122, Leong Hou U
AAAI3
2025 PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
abstract
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen 0008, Josef Dai, Boren Zheng, Tianyi Qiu, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang 0001
ACL (1)3
2025 SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
abstract
Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-world deployment, including the risk of harm to the environment, the robot itself, and humans. *How can safety constraints be explicitly integrated into VLAs?* We address this by exploring an integrated safety approach (ISA), systematically **modeling** safety requirements, then actively **eliciting** diverse unsafe behaviors, effectively **constraining** VLA policies via safe reinforcement learning, and rigorously **assuring** their safety through targeted evaluations. Leveraging the constrained Markov decision process (CMDP) paradigm, ISA optimizes VLAs from a min-max perspective against elicited safety risks. Thus, policies aligned through this comprehensive approach achieve the following key features: (I) effective **safety-performance trade-offs**, reducing the cumulative cost of safety violations by 83.58\% compared to the state-of-the-art method, while also maintaining task success rate (+3.85\%). (II) strong **safety assurance**, with the ability to mitigate long-tail risks and handle extreme failure scenarios. (III) robust **generalization** of learned safety behaviors to various out-of-distribution perturbations. The effectiveness is evaluated on long-horizon mobile manipulation tasks.
Borong Zhang, Jiaming Ji, Yingshan Lei, Juntao Dai, Yuanpei Chen, Yaodong Yang 0001
NeurIPS1
2024 SafeDreamer: Safe Reinforcement Learning with World Models
abstract
The deployment of Reinforcement Learning (RL) in real-world applications is constrained by its failure to satisfy safety criteria. Existing Safe Reinforcement Learning (SafeRL) methods, which rely on cost functions to enforce safety, often fail to achieve zero-cost performance in complex scenarios, especially vision-only tasks. These limitations are primarily due to model inaccuracies and inadequate sample efficiency. The integration of the world model has proven effective in mitigating these shortcomings. In this work, we introduce SafeDreamer, a novel algorithm incorporating Lagrangian-based methods into world model planning processes within the superior Dreamer framework. Our method achieves nearly zero-cost performance on various tasks, spanning low-dimensional and vision-only input, within the Safety-Gymnasium benchmark, showcasing its efficacy in balancing performance and safety in RL tasks. Further details can be found in the code repository: https://github.com/PKU-Alignment/SafeDreamer.
Weidong Huang 0008, Jiaming Ji, Chunhe Xia, Borong Zhang, Yaodong Yang 0001
ICLR4
2024 Aligner: Efficient Alignment by Learning to Correct
abstract
With the rapid development of large language models (LLMs) and ever-evolving practical requirements, finding an efficient and effective alignment method has never been more critical. However, the tension between the complexity of current alignment methods and the need for rapid iteration in deployment scenarios necessitates the development of a model-agnostic alignment approach that can operate under these constraints. In this paper, we introduce Aligner, a novel and simple alignment paradigm that learns the correctional residuals between preferred and dispreferred answers using a small model. Designed as a model-agnostic, plug-and-play module, Aligner can be directly applied to various open-source and API-based models with only one-off training, making it suitable for rapid iteration. Notably, Aligner can be applied to any powerful, large-scale upstream models. Moreover, it can even iteratively bootstrap the upstream models using corrected responses as synthetic human preference data, breaking through the model's performance ceiling. Our experiments demonstrate performance improvements by deploying the same Aligner model across 11 different LLMs, evaluated on the 3H dimensions (helpfulness, harmlessness, and honesty). Specifically, Aligner-7B has achieved an average improvement of 68.9% in helpfulness and 22.8% in harmlessness across the tested LLMs while also effectively reducing hallucination. In the Alpaca-Eval leaderboard, stacking Aligner-2B on GPT-4 Turbo improved its LC Win Rate from 55.0% to 58.3%, surpassing GPT-4 Omni's 57.5% Win Rate (community report).
Jiaming Ji, Boyuan Chen 0008, Hantao Lou, Donghai Hong, Borong Zhang, Xuehai Pan, Tianyi Qiu, Juntao Dai, Yaodong Yang 0001
NeurIPS5
2024 OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research
abstract
AI systems empowered by reinforcement learning (RL) algorithms harbor the immense potential to catalyze societal advancement, yet their deployment is often impeded by significant safety concerns. Particularly in safety-critical applications, researchers have raised concerns about unintended harms or unsafe behaviors of unaligned RL agents. The philosophy of safe reinforcement learning (SafeRL) is to align RL agents with harmless intentions and safe behavioral patterns. In SafeRL, agents learn to develop optimal policies by receiving feedback from the environment, while also fulfilling the requirement of minimizing the risk of unintended harm or unsafe behavior. However, due to the intricate nature of SafeRL algorithm implementation, combining methodologies across various domains presents a formidable challenge. This had led to an absence of a cohesive and efficacious learning framework within the contemporary SafeRL research milieu. In this work, we introduce a foundational framework designed to expedite SafeRL research endeavors. Our comprehensive framework encompasses an array of algorithms spanning different RL domains and places heavy emphasis on safety elements. Our efforts are to make the SafeRL-related research process more streamlined and efficient, therefore facilitating further research in AI safety.
Jiaming Ji, Borong Zhang, Juntao Dai, Xuehai Pan, Ruiyang Sun, Weidong Huang 0008, Yiran Geng, Mickel Liu, Yaodong Yang 0001
J. Mach. Learn. Res.3
2023 Safety Gymnasium: A Unified Safe Reinforcement Learning Benchmark
abstract
Artificial intelligence (AI) systems possess significant potential to drive societal progress. However, their deployment often faces obstacles due to substantial safety concerns. Safe reinforcement learning (SafeRL) emerges as a solution to optimize policies while simultaneously adhering to multiple constraints, thereby addressing the challenge of integrating reinforcement learning in safety-critical scenarios. In this paper, we present an environment suite called Safety-Gymnasium, which encompasses safety-critical tasks in both single and multi-agent scenarios, accepting vector and vision-only input. Additionally, we offer a library of algorithms named Safe Policy Optimization (SafePO), comprising 16 state-of-the-art SafeRL algorithms. This comprehensive library can serve as a validation tool for the research community. By introducing this benchmark, we aim to facilitate the evaluation and comparison of safety performance, thus fostering the development of reinforcement learning for safer, more reliable, and responsible real-world applications. The website of this project can be accessed at https://sites.google.com/view/safety-gymnasium.
Jiaming Ji, Borong Zhang, Xuehai Pan, Weidong Huang 0008, Ruiyang Sun, Yiran Geng, Yifan Zhong, Josef Dai, Yaodong Yang 0001
NeurIPS2