Junlin Wu 0001

dblp:188/8292-1 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0006-1037-1827ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Trustworthy machine learning · 44% Reinforcement learning · 33% Motion planning and robot control · 17%
Theoretical computer science
3 papers
Algorithmic game theory and mechanism design · 78% Computational complexity · 17% Mathematical optimization · 6%
Network and information security
2 papers
Security and privacy of machine learning · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
poisoning attack
1.622025
Preference Poisoning Attacks on Reward Model Learning · SP 2025
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models · ACL (1) 2024
Machine learning › Trustworthy machine learning
verification
1.422024
Verified Safe Reinforcement Learning for Neural Network Dynamic Models · NeurIPS 2024
Neural Lyapunov Control for Discrete-Time Systems · NeurIPS 2023
Machine learning › Reinforcement learning
safe reinforcement learning
1.122026
Verified Safe Reinforcement Learning for Neural Network Dynamic Models · NeurIPS 2024
Learning Vision-Based Neural Network Controllers with Semi-Probabilistic Safety Guarantees · AAAI 2026
Machine learning › Trustworthy machine learning › verification
formal verification of neural networks
1.012026
Learning Vision-Based Neural Network Controllers with Semi-Probabilistic Safety Guarantees · AAAI 2026
Robotics › Motion planning and robot control › robot control › safe control
safe learning-based control
1.012026
Learning Vision-Based Neural Network Controllers with Semi-Probabilistic Safety Guarantees · AAAI 2026
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.022024
Axioms for AI Alignment from Human Feedback · NeurIPS 2024
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models · ACL (1) 2024
Machine learning › Reinforcement learning
reward learning
0.812024
Axioms for AI Alignment from Human Feedback · NeurIPS 2024
Robotics › Motion planning and robot control › robot control
safe control
0.812024
Verified Safe Reinforcement Learning for Neural Network Dynamic Models · NeurIPS 2024
Security and privacy of machine learning › adversarial attack
backdoor attack
0.812024
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models · ACL (1) 2024
Security and privacy of machine learning › poisoning attack
reward poisoning
0.812024
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models · ACL (1) 2024
Algorithmic game theory and mechanism design › social choice
preference aggregation
0.812024
Axioms for AI Alignment from Human Feedback · NeurIPS 2024
Algorithmic game theory and mechanism design
social choice
0.812024
Axioms for AI Alignment from Human Feedback · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
neural network verification
0.712023
Exact Verification of ReLU Neural Control Barrier Functions · NeurIPS 2023
Machine learning › Trustworthy machine learning › AI safety › safety assurance
safety verification
0.712023
Exact Verification of ReLU Neural Control Barrier Functions · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.612022
Robust Deep Reinforcement Learning through Bootstrapped Opportunistic Curriculum · ICML 2022
Machine learning › Learning paradigms
curriculum learning
0.612022
Robust Deep Reinforcement Learning through Bootstrapped Opportunistic Curriculum · ICML 2022
Machine learning › Reinforcement learning
robust reinforcement learning
0.612022
Robust Deep Reinforcement Learning through Bootstrapped Opportunistic Curriculum · ICML 2022
Algorithmic game theory and mechanism design › social choice › computational social choice
election control
0.612022
Manipulating Elections by Changing Voter Perceptions · IJCAI 2022
Mathematical optimization › discrete optimization
mixed integer linear programming
0.212023
Neural Lyapunov Control for Discrete-Time Systems · NeurIPS 2023
Machine learning › Trustworthy machine learning › adversarial machine learning
robustness to adversarial attacks
0.212022
Robust Deep Reinforcement Learning through Bootstrapped Opportunistic Curriculum · ICML 2022

Methods — techniques the papers use, named apart from their topics

curriculum learning · 1.8reward poisoning · 1.5bradley-terry-luce model · 1.5backdoor attack · 1.5axiomatic analysis · 1.5reachability analysis · 1.0distribution-free tail bound · 1.0conditional generative network · 1.0rank-by-distance methods · 0.9gradient-based attack · 0.9maximum likelihood estimation · 0.8incremental verification · 0.8gradient-based learning · 0.8neural lyapunov function · 0.7mixed integer linear programming · 0.7gradient-based counterexample search · 0.7spatial voting · 0.6complexity analysis · 0.6
YearPublicationVenuePosition
2026 Learning Vision-Based Neural Network Controllers with Semi-Probabilistic Safety Guarantees
abstract
Ensuring safety in autonomous systems with vision-based control remains a critical challenge due to the high dimensionality of image inputs and the fact that the relationship between true system state and its visual manifestation is unknown. Existing methods for learning-based control in such settings typically lack formal safety guarantees. To address this challenge, we introduce a novel semi-probabilistic verification framework that integrates reachability analysis with conditional generative networks and distribution-free tail bounds to enable efficient and scalable verification of vision-based neural network controllers. Next, we develop a gradient-based training approach that employs a novel safety loss function, safety-aware data-sampling strategy to efficiently select and store critical training examples, and curriculum learning, to efficiently synthesize safe controllers in the semi-probabilistic framework. Empirical evaluations in X-Plane 11 airplane landing simulation, CARLA-simulated autonomous lane following, F1Tenth vehicle lane following in a physical visually-rich miniature environment, and Airsim-simulated drone navigation and obstacle avoidance demonstrate the effectiveness of our method in achieving formal safety guarantees while maintaining strong nominal performance.
Xinhang Ma, Junlin Wu 0001, Hussein Sibai, Yiannis Kantaros, Yevgeniy Vorobeychik
AAAI2
2025 Preference Poisoning Attacks on Reward Model Learning
abstract
Learning reward models from pairwise comparisons is a fundamental component in a number of domains, in-cluding autonomous control, conversational agents, and rec-ommendation systems, as part of a broad goal of aligning automated decisions with user preferences. These approaches entail collecting preference information from people, with feedback often provided anonymously. Since preferences are subjective, there is no gold standard to compare against; yet, reliance of high-impact systems on preference learning creates a strong motivation for malicious actors to skew data collected in this fashion to their ends. We investigate the nature and extent of this vulnerability by considering an attacker who can flip a small subset of preference comparisons to either promote or demote a target outcome. We propose two classes of algorithmic approaches for these attacks: a gradient-based framework, and several variants of rank-by-distance methods. Next, we evaluate the efficacy of best attacks in both these classes in successfully achieving malicious goals on datasets from three domains: autonomous control, recommendation system, and textual prompt-response preference learning. We find that the best attacks are often highly successful, achieving in the most extreme case 100% success rate with only 0.3% of the data poisoned. However, which attack is best can vary significantly across domains. In addition, we observe that the simpler and more scalable rank-by-distance approaches are often competitive with, and on occasion significantly outper-form, gradient-based methods. Finally, we show that state-of-the-art defenses against other classes of poisoning attacks exhibit limited efficacy in our setting.
Junlin Wu 0001, Jiongxiao Wang, Chaowei Xiao, Ning Zhang 0017, Yevgeniy Vorobeychik
SP1
2024 RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
abstract
Reinforcement Learning with Human Feedback (RLHF) is a methodology designed to align Large Language Models (LLMs) with human preferences, playing an important role in LLMs alignment.Despite its advantages, RLHF relies on human annotators to rank the text, which can introduce potential security vulnerabilities if any adversarial annotator (i.e., attackers) manipulates the ranking score by upranking any malicious text to steer the LLM adversarially.To assess the red-teaming of RLHF against human preference data poisoning, we propose RankPoison, a poisoning attack method on candidates' selection of preference rank flipping to reach certain malicious behaviors (e.g., generating longer sequences, which can increase the computational cost).With poisoned dataset generated by RankPoison, we can perform poisoning attacks on LLMs to generate longer tokens without hurting the original safety alignment performance.Moreover, applying RankPoison, we also successfully implement a backdoor attack where LLMs can generate longer answers under questions with the trigger word.Our findings highlight critical security challenges in RLHF, underscoring the necessity for more robust alignment methods for LLMs.
Jiongxiao Wang, Junlin Wu 0001, Muhao Chen 0001, Yevgeniy Vorobeychik, Chaowei Xiao
ACL (1)2
2024 Verified Safe Reinforcement Learning for Neural Network Dynamic Models
abstract
Learning reliably safe autonomous control is one of the core problems in trustworthy autonomy. However, training a controller that can be formally verified to be safe remains a major challenge. We introduce a novel approach for learning verified safe control policies in nonlinear neural dynamical systems while maximizing overall performance. Our approach aims to achieve safety in the sense of finite-horizon reachability proofs, and is comprised of three key parts. The first is a novel curriculum learning scheme that iteratively increases the verified safe horizon. The second leverages the iterative nature of gradient-based learning to leverage incremental verification, reusing information from prior verification runs. Finally, we learn multiple verified initial-state-dependent controllers, an idea that is especially valuable for more complex domains where learning a single universal verified safe controller is extremely challenging. Our experiments on five safe control problems demonstrate that our trained controllers can achieve verified safety over horizons that are as much as an order of magnitude longer than state-of-the-art baselines, while maintaining high reward, as well as a perfect safety record over entire episodes. Our code is available at https://github.com/jlwu002/VSRL.
Junlin Wu 0001, Yevgeniy Vorobeychik
NeurIPS1
2024 Axioms for AI Alignment from Human Feedback
abstract
In the context of reinforcement learning from human feedback (RLHF), the reward function is generally derived from maximum likelihood estimation of a random utility model based on pairwise comparisons made by humans. The problem of learning a reward function is one of preference aggregation that, we argue, largely falls within the scope of social choice theory. From this perspective, we can evaluate different aggregation methods via established axioms, examining whether these methods meet or fail well-known standards. We demonstrate that both the Bradley-Terry-Luce Model and its broad generalizations fail to meet basic axioms. In response, we develop novel rules for learning reward functions with strong axiomatic guarantees. A key innovation from the standpoint of social choice is that our problem has a *linear* structure, which greatly restricts the space of feasible rules and leads to a new paradigm that we call *linear social choice*.
Luise Ge, Daniel Halpern 0002, Evi Micha, Ariel D. Procaccia, Itai Shapira, Yevgeniy Vorobeychik, Junlin Wu 0001
NeurIPS7
2023 Neural Lyapunov Control for Discrete-Time Systems
abstract
While ensuring stability for linear systems is well understood, it remains a major challenge for nonlinear systems. A general approach in such cases is to compute a combination of a Lyapunov function and an associated control policy. However, finding Lyapunov functions for general nonlinear systems is a challenging task. To address this challenge, several methods have been proposed that represent Lyapunov functions using neural networks. However, such approaches either focus on continuous-time systems, or highly restricted classes of nonlinear dynamics. We propose the first approach for learning neural Lyapunov control in a broad class of discrete-time systems. Three key ingredients enable us to effectively learn provably stable control policies. The first is a novel mixed-integer linear programming approach for verifying the discrete-time Lyapunov stability conditions, leveraging the particular structure of these conditions. The second is a novel approach for computing verified sublevel sets. The third is a heuristic gradient-based method for quickly finding counterexamples to significantly speed up Lyapunov function learning. Our experiments on four standard benchmarks demonstrate that our approach significantly outperforms state-of-the-art baselines. For example, on the path tracking benchmark, we outperform recent neural Lyapunov control baselines by an order of magnitude in both running time and the size of the region of attraction, and on two of the four benchmarks (cartpole and PVTOL), ours is the first automated approach to return a provably stable controller. Our code is available at: https://github.com/jlwu002/nlc_discrete.
Junlin Wu 0001, Andrew Clark 0001, Yiannis Kantaros, Yevgeniy Vorobeychik
NeurIPS1
2023 Exact Verification of ReLU Neural Control Barrier Functions
abstract
Control Barrier Functions (CBFs) are a popular approach for safe control of nonlinear systems. In CBF-based control, the desired safety properties of the system are mapped to nonnegativity of a CBF, and the control input is chosen to ensure that the CBF remains nonnegative for all time. Recently, machine learning methods that represent CBFs as neural networks (neural control barrier functions, or NCBFs) have shown great promise due to the universal representability of neural networks. However, verifying that a learned CBF guarantees safety remains a challenging research problem. This paper presents novel exact conditions and algorithms for verifying safety of feedforward NCBFs with ReLU activation functions. The key challenge in doing so is that, due to the piecewise linearity of the ReLU function, the NCBF will be nondifferentiable at certain points, thus invalidating traditional safety verification methods that assume a smooth barrier function. We resolve this issue by leveraging a generalization of Nagumo's theorem for proving invariance of sets with nonsmooth boundaries to derive necessary and sufficient conditions for safety. Based on this condition, we propose an algorithm for safety verification of NCBFs that first decomposes the NCBF into piecewise linear segments and then solves a nonlinear program to verify safety of each segment as well as the intersections of the linear segments. We mitigate the complexity by only considering the boundary of the safe region and by pruning the segments with Interval Bound Propagation (IBP) and linear relaxation. We evaluate our approach through numerical studies with comparison to state-of-the-art SMT-based methods. Our code is available at https://github.com/HongchaoZhang-HZ/exactverif-reluncbf-nips23.
Junlin Wu 0001, Yevgeniy Vorobeychik, Andrew Clark 0001
NeurIPS2
2022 Robust Deep Reinforcement Learning through Bootstrapped Opportunistic Curriculum
abstract
Despite considerable advances in deep reinforcement learning, it has been shown to be highly vulnerable to adversarial perturbations to state observations. Recent efforts that have attempted to improve adversarial robustness of reinforcement learning can nevertheless tolerate only very small perturbations, and remain fragile as perturbation size increases. We propose Bootstrapped Opportunistic Adversarial Curriculum Learning (BCL), a novel flexible adversarial curriculum learning framework for robust reinforcement learning. Our framework combines two ideas: conservatively bootstrapping each curriculum phase with highest quality solutions obtained from multiple runs of the previous phase, and opportunistically skipping forward in the curriculum. In our experiments we show that the proposed BCL framework enables dramatic improvements in robustness of learned policies to adversarial perturbations. The greatest improvement is for Pong, where our framework yields robustness to perturbations of up to 25/255; in contrast, the best existing approach can only tolerate adversarial noise up to 5/255. Our code is available at: https://github.com/jlwu002/BCL.
Junlin Wu 0001, Yevgeniy Vorobeychik
ICML1
2022 Manipulating Elections by Changing Voter Perceptions
abstract
The integrity of elections is central to democratic systems. However, a myriad of malicious actors aspire to influence election outcomes for financial or political benefit. A common means to such ends is by manipulating perceptions of the voting public about select candidates, for example, through misinformation. We present a formal model of the impact of perception manipulation on election outcomes in the framework of spatial voting theory, in which the preferences of voters over candidates are generated based on their relative distance in the space of issues. We show that controlling elections in this model is, in general, NP-hard, whether issues are binary or real-valued. However, we demonstrate that critical to intractability is the diversity of opinions on issues exhibited by the voting public. When voter views lack diversity, and we can instead group them into a small number of categories---for example, as a result of political polarization---the election control problem can be solved in polynomial time in the number of issues and candidates for arbitrary scoring rules.
Junlin Wu 0001, Andrew Estornell, Lecheng Kong, Yevgeniy Vorobeychik
IJCAI1