Zaiyan Xu

dblp:326/5171 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-6194-912XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 100%
Artificial intelligence
1 paper
Reinforcement learning · 87% Optimization for machine learning · 13%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
hardware verification and test
1.012026
ReFuzz: Reusing Tests for Processor Fuzzing with Contextual Bandits · NDSS 2026
Electronic design automation › hardware verification and test › processor verification
processor fuzzing
1.012026
ReFuzz: Reusing Tests for Processor Fuzzing with Contextual Bandits · NDSS 2026
Electronic design automation › hardware verification and test
processor verification
1.012026
ReFuzz: Reusing Tests for Processor Fuzzing with Contextual Bandits · NDSS 2026
Machine learning › Reinforcement learning
offline reinforcement learning
0.612022
Robust Reinforcement Learning using Offline Data · NeurIPS 2022
Machine learning › Reinforcement learning
robust reinforcement learning
0.612022
Robust Reinforcement Learning using Offline Data · NeurIPS 2022
Machine learning › Optimization for machine learning
convergence analysis
0.212022
Robust Reinforcement Learning using Offline Data · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

test reuse · 1.0fuzzing · 1.0contextual bandit · 1.0max-min optimization · 0.6fitted q-iteration · 0.6
YearPublicationVenuePosition
2026 ReFuzz: Reusing Tests for Processor Fuzzing with Contextual Bandits
Chen Chen 0125, Zaiyan Xu, Mohamadreza Rostami, Dileep M. Kalathil, Ahmad-Reza Sadeghi, Jeyavijayan Rajendran
NDSS2
2025 Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
abstract
A major challenge in aligning large language models (LLMs) with human preferences is the issue of distribution shift. LLM alignment algorithms rely on static preference datasets, assuming that they accurately represent real-world user preferences. However, user preferences vary significantly across geographical regions, demographics, linguistic patterns, and evolving cultural trends. This preference distribution shift leads to catastrophic alignment failures in many real-world applications. We address this problem using the principled framework of distributionally robust optimization, and develop two novel distributionally robust direct preference optimization (DPO) algorithms, namely, Wasserstein DPO (WDPO) and Kullback–Leibler DPO (KLDPO). We characterize the sample complexity of learning the optimal policy parameters for WDPO and KLDPO. Moreover, we propose scalable gradient descent-style learning algorithms by developing suitable approximations for the challenging minimax loss functions of WDPO and KLDPO. Our empirical experiments using benchmark data sets and LLMs demonstrate the superior performance of WDPO and KLDPO in substantially improving the alignment when there is a preference distribution shift.
Zaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep M. Kalathil, Rahul Jain 0002, Deepak Ramachandran
NeurIPS1
2023 Improved Sample Complexity Bounds for Distributionally Robust Reinforcement Learning
abstract
We consider the problem of learning a control policy that is robust against the parameter mismatches between the training environment and testing environment. We formulate this as a distributionally robust reinforcement learning (DR-RL) problem where the objective is to learn the policy which maximizes the value function against the worst possible stochastic model of the environment in an uncertainty set. We focus on the tabular episodic learning setting where the algorithm has access to a generative model of the nominal (training) environment around which the uncertainty set is defined. We propose the Robust Phased Value Learning (RPVL) algorithm to solve this problem for the uncertainty sets specified by four different divergences: total variation, chi-square, Kullback-Leibler, and Wasserstein. We show that our algorithm achieves $\tilde{\mathcal{O}}(|\mathcal{S}||\mathcal{A}| H^{5})$ sample complexity, which is uniformly better than the existing results by a factor of $|\mathcal{S}|$, where $|\mathcal{S}|$ is number of states, $|\mathcal{A}|$ is the number of actions, and $H$ is the horizon length. We also provide the first-ever sample complexity result for the Wasserstein uncertainty set. Finally, we demonstrate the performance of our algorithm using simulation experiments.
Zaiyan Xu, Kishan Panaganti, Dileep M. Kalathil
AISTATS1
2022 Robust Reinforcement Learning using Offline Data
abstract
The goal of robust reinforcement learning (RL) is to learn a policy that is robust against the uncertainty in model parameters. Parameter uncertainty commonly occurs in many real-world RL applications due to simulator modeling errors, changes in the real-world system dynamics over time, and adversarial disturbances. Robust RL is typically formulated as a max-min problem, where the objective is to learn the policy that maximizes the value against the worst possible models that lie in an uncertainty set. In this work, we propose a robust RL algorithm called Robust Fitted Q-Iteration (RFQI), which uses only an offline dataset to learn the optimal robust policy. Robust RL with offline data is significantly more challenging than its non-robust counterpart because of the minimization over all models present in the robust Bellman operator. This poses challenges in offline data collection, optimization over the models, and unbiased estimation. In this work, we propose a systematic approach to overcome these challenges, resulting in our RFQI algorithm. We prove that RFQI learns a near-optimal robust policy under standard assumptions and demonstrate its superior performance on standard benchmark problems.
Kishan Panaganti, Zaiyan Xu, Dileep M. Kalathil, Mohammad Ghavamzadeh
NeurIPS2