EDBT 2026 Demo / reviewers in the wild / expert
Yanchao Sun
dblp:132/6840
· DBLP profile ↗
38ranked-venue papers
15as first author
30since 2021 · last 2026
0000-0001-7392-4472ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 8 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Instructed Diffuser With Temporal Condition Guidance for Offline Reinforcement LearningabstractRecentworks have shown the potential of diffusion models in computer vision and natural language processing. Apart from the classical supervised learning fields, diffusion models have also shown strong competitiveness in reinforcement learning (RL) by formulating decision-making as sequential generation. However, incorporating temporal information of sequential data and utilizing it to guide diffusion models to perform better generation is still an open challenge. In this paper, we take one step forward to investigate controllable generation with temporal conditions that are refined from temporal information. We observe the importance of temporal conditions in sequential generation in sufficient scenarios and provide a comprehensive discussion and comparison of different temporal conditions. Based on the observations, we propose an effective temporally-conditional diffusion model coined Temporally-Composable Diffuser (TCD), which extracts temporal information from interaction sequences and explicitly guides generation with temporal conditions. Specifically, we separate the sequences into three parts according to time expansion and identify historical, immediate, and prospective conditions accordingly. Each condition preserves non-overlapping temporal information of sequences, enabling more controllable generation when we jointly use them to guide the diffuser. Finally, we conduct extensive experiments and analysis to reveal the favorable applicability of TCD in offline RL tasks, where our method reaches or matches the best performance compared with prior SOTA baselines. Jifeng Hu, Yanchao Sun, Sili Huang, Siyuan Guo 0001, Hechang Chen, Li Shen 0008, Lichao Sun 0001, Yi Chang 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement LearningabstractOffline reinforcement learning (RL) provides a promising solution to learning an agent fully relying on a data-driven paradigm. However, constrained by the limited quality of the offline dataset, its performance is often suboptimal. Therefore, it is desired to further finetune the agent via extra online interactions before deployment. Unfortunately, offline-to-online RL can be challenging due to two main challenges: constrained exploratory behavior and state-action distribution shift. In view of this, we propose a simple unified uncertainty-guided (SUNG) framework, which naturally unifies the solution to both challenges with the tool of uncertainty. Specifically, SUNG quantifies uncertainty via a variational autoencoder (VAE)-based state-action visitation density estimator. To facilitate efficient exploration, SUNG presents a practical optimistic exploration strategy to select informative actions with both high value and high uncertainty. Moreover, SUNG develops an adaptive exploitation method by applying conservative offline RL objectives to high-uncertainty samples and standard online RL objectives to low-uncertainty samples to smoothly bridge offline and online stages. SUNG achieves state-of-the-art online finetuning performance when combined with different offline RL methods, across various environments and datasets in the D4RL benchmark. Codes are made publicly available in https://github.com/guosyjlu/ACBR. Siyuan Guo 0001, Yanchao Sun, Jifeng Hu, Sili Huang, Hechang Chen, Haiyin Piao, Lichao Sun 0001, Yi Chang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes TheoryabstractLifelong reinforcement learning (RL) has been developed as a paradigm for extending single-task RL to more realistic, dynamic settings. In lifelong RL, the "life" of an RL agent is modeled as a stream of tasks drawn from a task distribution. We propose EPIC (Empirical PAC-Bayes that Improves Continuously), a novel algorithm designed for lifelong RL using PAC-Bayes theory. EPIC learns a shared policy distribution, referred to as the world policy, which enables rapid adaptation to new tasks while retaining valuable knowledge from previous experiences. Our theoretical analysis establishes a relationship between the algorithm’s generalization performance and the number of prior tasks preserved in memory. We also derive the sample complexity of EPIC in terms of RL regret. Extensive experiments on a variety of environments demonstrate that EPIC significantly outperforms existing methods in lifelong RL, offering both theoretical guarantees and practical efficacy through the use of the world policy. Zhi Zhang 0012, Chris Chow, Yasi Zhang, Yanchao Sun, Eric Hanchen Jiang, Han Liu 0001, Furong Huang, Yuchen Cui, Oscar Hernan Madrid Padilla |
AISTATS | 4 |
| 2025 | TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated WeightsabstractDirect Preference Optimization (DPO) has been widely adopted for preference alignment of Large Language Models (LLMs) due to its simplicity and effectiveness.
However, DPO is derived as a bandit problem in which the whole response is treated as a single arm, ignoring the importance differences between tokens, which may affect optimization efficiency and make it difficult to achieve optimal results.
In this work, we propose that the optimal data for DPO has equal expected rewards for each token in winning and losing responses, as there is no difference in token importance.
However, since the optimal dataset is unavailable in practice, we propose using the original dataset for importance sampling to achieve unbiased optimization.
Accordingly, we propose a token-level importance sampling DPO objective named TIS-DPO that assigns importance weights to each token based on its reward.
Inspired by previous works, we estimate the token importance weights using the difference in prediction probabilities from a pair of contrastive LLMs. We explore three methods to construct these contrastive LLMs: (1) guiding the original LLM with contrastive prompts, (2) training two separate LLMs using winning and losing responses, and (3) performing forward and reverse DPO training with winning and losing responses.
Experiments show that TIS-DPO significantly outperforms various baseline methods on harmlessness and helpfulness alignment and summarization tasks. We also visualize the estimated weights, demonstrating their ability to identify key token positions. Aiwei Liu, Haoping Bai, Zhiyun Lu, Yanchao Sun, Xiang Kong, Xiaoming Simon Wang, Jiulong Shan, Albin Madappally Jose, Xiaojiang Liu, Lijie Wen 0001, Philip S. Yu |
ICLR | 4 |
| 2025 | Safety Guaranteed Robust Multi-Agent Reinforcement Learning with Hierarchical Control for Connected and Automated VehiclesabstractWe address the problem of coordination and control of Connected and Automated Vehicles (CAVs) in the presence of imperfect observations in mixed traffic environment. A commonly used approach is learning-based decision-making, such as reinforcement learning (RL). However, most existing safe RL methods suffer from two limitations: (i) they assume accurate state information, and (ii) safety is generally defined over the expectation of the trajectories. It remains challenging to design optimal coordination between multi-agents while ensuring hard safety constraints under system state uncertainties (e.g., those that arise from noisy sensor measurements, communication, or state estimation methods) at every time step. We propose a safety guaranteed hierarchical coordination and control scheme called Safe-RMM to address the challenge. Specifically, the high-level coordination policy of CAVs in mixed traffic environment is trained by the Robust Multi-Agent Proximal Policy Optimization (RMAPPO) method. Though trained without uncertainty, our method leverages a worst-case Q network to ensure the model's robust performances when state uncertainties are present during testing. The low-level controller is implemented using model predictive control (MPC) with robust Control Barrier Functions (CBFs) to guarantee safety through their forward invariance property. We compare our method with baselines in different road networks in the CARLA simulator. Results show that our method provides the best evaluated safety and efficiency in challenging mixed traffic environments with uncertainties. H. M. Sabbir Ahmad, Ehsan Sabouni, Yanchao Sun, Furong Huang, Wenchao Li 0001, Fei Miao |
ICRA | 4 |
| 2025 | Checklists Are Better Than Reward Models For Aligning Language ModelsabstractLanguage models must be adapted to understand and follow user instructions. Reinforcement learning is widely used to facilitate this —typically using fixed criteria such as "helpfulness" and "harmfulness". In our work, we instead propose using flexible, instruction-specific criteria as a means of broadening the impact that reinforcement learning can have in eliciting instruction following. We propose "Reinforcement Learning from Checklist Feedback" (RLCF). From instructions, we extract checklists and evaluate how well responses satisfy each item—using both AI judges and specialized verifier programs—then combine these scores to compute rewards for RL. We compare RLCF with other alignment methods on top of a strong instruction following model (Qwen2.5-7B-Instruct) on five widely-studied benchmarks — RLCF is the only method to help on every benchmark, including a 4-point boost in hard satisfaction rate on FollowBench, a 6-point increase on InFoBench, and a 3-point rise in win rate on Arena-Hard. We show that RLCF can also be used off-policy to improve Llama 3.1 8B Instruct and OLMo 2 7B Instruct. These results establish rubrics as a key tool for improving language models' support of queries that express a multitude of needs. We release our our dataset of rubrics (WildChecklists), models, and code to the public. Vijay Viswanathan 0002, Yanchao Sun, Xiang Kong, Graham Neubig, Sherry Tongshuang Wu |
NeurIPS | 2 |
| 2024 | Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled PerturbationsabstractDeploying reinforcement learning (RL) systems requires robustness to uncertainty and model misspecification, yet prior robust RL methods typically only study noise introduced independently across time. However, practical sources of uncertainty are usually coupled across time.
We formally introduce temporally-coupled perturbations, presenting a novel challenge for existing robust RL methods. To tackle this challenge, we propose GRAD, a novel game-theoretic approach that treats the temporally-coupled robust RL problem as a partially-observable two-player zero-sum game. By finding an approximate equilibrium within this game, GRAD optimizes for general robustness against temporally-coupled perturbations. Experiments on continuous control tasks demonstrate that, compared with prior methods, our approach achieves a higher degree of robustness to various types of attacks on different attack domains, both in settings with temporally-coupled perturbations and decoupled perturbations. Yongyuan Liang, Yanchao Sun, Ruijie Zheng, Benjamin Eysenbach, Tuomas Sandholm, Furong Huang, Stephen McAleer |
ICLR | 2 |
| 2024 | Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RLabstractMost existing works focus on direct perturbations to the victim's state/action or the underlying transition dynamics to demonstrate the vulnerability of reinforcement learning agents to adversarial attacks.
However, such direct manipulations may not be always realizable.
In this paper, we consider a multi-agent setting where a well-trained victim agent $\nu$ is exploited by an attacker controlling another
agent $\alpha$ with an \textit{adversarial policy}. Previous models do not account for the possibility that the attacker may only have partial control over
$\alpha$ or that the attack may produce easily detectable ``abnormal'' behaviors. Furthermore, there is a lack of provably efficient defenses against these adversarial policies.
To address these limitations, we introduce a generalized attack framework that has the flexibility to model to what extent the adversary is able to control the agent, and allows the attacker to regulate the state distribution shift and produce stealthier adversarial policies. Moreover, we offer a provably efficient defense with polynomial convergence to the most robust victim policy through adversarial training with timescale separation.
This stands in sharp contrast to supervised learning, where adversarial training typically provides only \textit{empirical} defenses.
Using the Robosumo competition experiments, we show that our generalized attack formulation results in much stealthier adversarial policies when maintaining the same winning rate as baselines.
Additionally, our adversarial training approach yields stable learning dynamics and less exploitable victim policies. Souradip Chakraborty, Yanchao Sun, Furong Huang |
ICLR | 3 |
| 2024 | Beyond Worst-case Attacks: Robust RL with Adaptive Defense via Non-dominated PoliciesabstractIn light of the burgeoning success of reinforcement learning (RL) in diverse real-world applications, considerable focus has been directed towards ensuring RL policies are robust to adversarial attacks during test time. Current approaches largely revolve around solving a minimax problem to prepare for potential worst-case scenarios. While effective against strong attacks, these methods often compromise performance in the absence of attacks or the presence of only weak attacks. To address this, we study policy robustness under the well-accepted state-adversarial attack model, extending our focus beyond merely worst-case attacks. We first formalize this task at test time as a regret minimization problem and establish its intrinsic difficulty in achieving sublinear regret when the baseline policy is from a general continuous policy class, $\Pi$. This finding prompts us to \textit{refine} the baseline policy class $\Pi$ prior to test time, aiming for efficient adaptation within a compact, finite policy class $\tilde{\Pi}$, which can resort to an adversarial bandit subroutine. In light of the importance of a finite and compact $\tilde{\Pi}$, we propose a novel training-time algorithm to iteratively discover \textit{non-dominated policies}, forming a near-optimal and minimal $\tilde{\Pi}$, thereby ensuring both robustness and test-time efficiency. Empirical validation on the Mujoco corroborates the superiority of our approach in terms of natural and robust performance, as well as adaptability to various attack scenarios. Chenghao Deng, Yanchao Sun, Yongyuan Liang, Furong Huang |
ICLR | 3 |
| 2024 | COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RLabstractDyna-style model-based reinforcement learning contains two phases: model rollouts to generate sample for policy learning and real environment exploration using current policy for dynamics model learning. However, due to the complex real-world environment, it is inevitable to learn an imperfect dynamics model with model prediction error, which can further mislead policy learning and result in sub-optimal solutions. In this paper, we propose $\texttt{COPlanner}$, a planning-driven framework for model-based methods to address the inaccurately learned dynamics model problem with conservative model rollouts and optimistic environment exploration. $\texttt{COPlanner}$ leverages an uncertainty-aware policy-guided model predictive control (UP-MPC) component to plan for multi-step uncertainty estimation. This estimated uncertainty then serves as a penalty during model rollouts and as a bonus during real environment exploration respectively, to choose actions. Consequently, $\texttt{COPlanner}$ can avoid model uncertain regions through conservative model rollouts, thereby alleviating the influence of model error. Simultaneously, it explores high-reward model uncertain regions to reduce model error actively through optimistic real environment exploration. $\texttt{COPlanner}$ is a plug-and-play framework that can be applied to any dyna-style model-based methods. Experimental results on a series of proprioceptive and visual continuous control tasks demonstrate that both sample efficiency and asymptotic performance of strong model-based methods are significantly improved combined with $\texttt{COPlanner}$. Ruijie Zheng, Yanchao Sun, Ruonan Jia, Wichayaporn Wongkamjan, Huazhe Xu, Furong Huang |
ICLR | 3 |
| 2024 | Adapting Static Fairness to Sequential Decision-Making: Bias Mitigation Strategies towards Equal Long-term Benefit RateabstractDecisions made by machine learning models can have lasting impacts, making long-term fairness a critical consideration. It has been observed that ignoring the long-term effect and directly applying fairness criterion in static settings can actually worsen bias over time. To address biases in sequential decision-making, we introduce a long-term fairness concept named Equal Long-term Benefit Rate (ELBERT). This concept is seamlessly integrated into a Markov Decision Process (MDP) to consider the future effects of actions on long-term fairness, thus providing a unified framework for fair sequential decision-making problems. ELBERT effectively addresses the temporal discrimination issues found in previous long-term fairness notions. Additionally, we demonstrate that the policy gradient of Long-term Benefit Rate can be analytically simplified to standard policy gradients. This simplification makes conventional policy optimization methods viable for reducing bias, leading to our bias mitigation approach ELBERT-PO. Extensive experiments across various diverse sequential decision-making environments consistently reveal that ELBERT-PO significantly diminishes bias while maintaining high utility. Code is available at https://github.com/umd-huang-lab/ELBERT. Yuancheng Xu, Chenghao Deng, Yanchao Sun, Ruijie Zheng, Jieyu Zhao 0001, Furong Huang |
ICML | 3 |
| 2024 | Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language ModelsabstractVision-Language Models (VLMs) excel in generating textual responses from visual inputs, but their versatility raises security concerns. This study takes the first step in exposing VLMs’ susceptibility to data poisoning attacks that can manipulate responses to innocuous, everyday prompts. We introduce Shadowcast, a stealthy data poisoning attack where poison samples are visually indistinguishable from benign images with matching texts. Shadowcast demonstrates effectiveness in two attack types. The first is a traditional Label Attack, tricking VLMs into misidentifying class labels, such as confusing Donald Trump for Joe Biden. The second is a novel Persuasion Attack, leveraging VLMs’ text generation capabilities to craft persuasive and seemingly rational narratives for misinformation, such as portraying junk food as healthy. We show that Shadowcast effectively achieves the attacker’s intentions using as few as 50 poison samples. Crucially, the poisoned samples demonstrate transferability across different VLM architectures, posing a significant concern in black-box settings. Moreover, Shadowcast remains potent under realistic conditions involving various text prompts, training data augmentation, and image compression techniques. This work reveals how poisoned VLMs can disseminate convincing yet deceptive misinformation to everyday, benign users, emphasizing the importance of data integrity for responsible VLM deployments. Our code is available at: https://github.com/umd-huang-lab/VLM-Poisoning. Yuancheng Xu, Jiarui Yao, Manli Shu, Yanchao Sun, Zichu Wu, Ning Yu 0006, Tom Goldstein, Furong Huang |
NeurIPS | 4 |
| 2024 | Provable Unrestricted Adversarial Training Without Compromise With GeneralizabilityabstractAdversarial training (AT) is widely considered as the most promising strategy to defend against adversarial attacks and has drawn increasing interest from researchers. However, the existing AT methods still suffer from two challenges. First, they are unable to handle unrestricted adversarial examples (UAEs), which are built from scratch, as opposed to restricted adversarial examples (RAEs), which are created by adding perturbations bound by an$l_{p}$norm to observed examples. Second, the existing AT methods often achieve adversarial robustness at the expense of standard generalizability (i.e., the accuracy on natural examples) because they make a tradeoff between them. To overcome these challenges, we propose a unique viewpoint that understands UAEs as imperceptibly perturbed unobserved examples. Also, we find that the tradeoff results from the separation of the distributions of adversarial examples and natural examples. Based on these ideas, we propose a novel AT approach called Provable Unrestricted Adversarial Training (PUAT), which can provide a target classifier with comprehensive adversarial robustness against both UAE and RAE, and simultaneously improve its standard generalizability. Particularly, PUAT utilizes partially labeled data to achieve effective UAE generation by accurately capturing the natural data distribution through a novel augmented triple-GAN. At the same time, PUAT extends the traditional AT by introducing the supervised loss of the target classifier into the adversarial loss and achieves the alignment between the UAE distribution, the natural data distribution, and the distribution learned by the classifier, with the collaboration of the augmented triple-GAN. Finally, the solid theoretical analysis and extensive experiments conducted on widely-used benchmarks demonstrate the superiority of PUAT. Lilin Zhang, Ning Yang 0001, Yanchao Sun, Philip S. Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Is Imitation All You Need? Generalized Decision-Making with Dual-Phase TrainingabstractWe introduce DualMind, a generalist agent designed to tackle various decision-making tasks that addresses challenges posed by current methods, such as overfitting behaviors and dependence on task-specific fine-tuning. DualMind uses a novel "Dual-phase" training strategy that emulates how humans learn to act in the world. The model first learns fundamental common knowledge through a self-supervised objective tailored for control tasks and then learns how to make decisions based on different contexts through imitating behaviors conditioned on given prompts. DualMind can handle tasks across domains, scenes, and embodiments using just a single set of model weights and can execute zero-shot prompting without requiring task-specific finetuning. We evaluate DualMind on MetaWorld [40] and Habitat [31] through extensive experiments and demonstrate its superior generalizability compared to previous techniques, outperforming other generalist agents by over 50% and 70% on Habitat and MetaWorld, respectively. On the 45 tasks in MetaWorld, DualMind achieves over 30 tasks at a 90% success rate. Our source code is available at https://github.com/yunyikristy/DualMind. Yao Wei 0002, Yanchao Sun, Ruijie Zheng, Sai Vemprala, Rogerio Bonatti, Ratnesh Madaan, Zhongjie Ba, Ashish Kapoor |
ICCV | 2 |
| 2023 | SMART: Self-supervised Multi-task pretrAining with contRol Transformers
Yanchao Sun, Ratnesh Madaan, Rogerio Bonatti, Furong Huang, Ashish Kapoor |
ICLR | 1 |
| 2023 | Certifiably Robust Policy Learning against Adversarial Multi-Agent Communication
Yanchao Sun, Ruijie Zheng, Parisa Hassanzadeh, Yongyuan Liang, Soheil Feizi, Sumitra Ganesh, Furong Huang |
ICLR | 1 |
| 2023 | Exploring and Exploiting Decision Boundary Dynamics for Adversarial Robustness
Yuancheng Xu, Yanchao Sun, Micah Goldblum, Tom Goldstein, Furong Huang |
ICLR | 2 |
| 2023 | Learning Generalizable Agents via Saliency-guided Features DecorrelationabstractIn visual-based Reinforcement Learning (RL), agents often struggle to generalize well to environmental variations in the state space that were not observed during training. The variations can arise in both task-irrelevant features, such as background noise, and task-relevant features, such as robot configurations, that are related to the optimal decisions. To achieve generalization in both situations, agents are required to accurately understand the impact of changed features on the decisions, i.e., establishing the true associations between changed features and decisions in the policy model. However, due to the inherent correlations among features in the state space, the associations between features and decisions become entangled, making it difficult for the policy to distinguish them. To this end, we propose Saliency-Guided Features Decorrelation (SGFD) to eliminate these correlations through sample reweighting. Concretely, SGFD consists of two core techniques: Random Fourier Functions (RFF) and the saliency map. RFF is utilized to estimate the complex non-linear correlations in high-dimensional images, while the saliency map is designed to identify the changed features. Under the guidance of the saliency map, SGFD employs sample reweighting to minimize the estimated correlations related to changed features, thereby achieving decorrelation in visual RL tasks. Our experimental results demonstrate that SGFD can generalize well on a wide range of test environments and significantly outperforms state-of-the-art methods in handling both task-irrelevant variations and task-relevant variations. Sili Huang, Yanchao Sun, Jifeng Hu, Siyuan Guo 0001, Hechang Chen, Yi Chang 0001, Lichao Sun 0001, Bo Yang 0002 |
NeurIPS | 2 |
| 2023 | TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement Learning
Ruijie Zheng, Yanchao Sun, Jieyu Zhao 0001, Huazhe Xu, Hal Daumé III, Furong Huang |
NeurIPS | 3 |
| 2022 | DDoS Attack Detection Combining Time Series-based Multi-dimensional Sketch and Machine LearningabstractMachine learning-based DDoS attack detection methods are mostly implemented at the packet level with expensive computational time costs, and the space cost of those sketch-based detection methods is uncertain. This paper proposes a two-stage DDoS attack detection algorithm combining time series-based multi-dimensional sketch and machine learning technologies. Besides packet numbers, total lengths, and protocols, we construct the time series-based multi-dimensional sketch with limited space cost by storing elephant flow information with the Boyer-Moore voting algorithm and hash index. For the first stage of detection, we adopt CNN to generate sketch-level DDoS attack detection results from the time series-based multi-dimensional sketch. For the sketch with potential DDoS attacks, we use RNN with flow information extracted from the sketch to implement flow-level DDoS attack detection in the second stage. Experimental results show that not only is the detection accuracy of our proposed method much close to that of packet-level DDoS attack detection methods based on machine learning, but also the computational time cost of our method is much smaller with regard to the number of machine learning operations. Yanchao Sun, Yuanfeng Han, Mingsong Chen 0001, Shui Yu 0001, Yimin Xu |
APNOMS | 1 |
| 2022 | Who Is the Strongest Enemy? Towards Optimal and Efficient Evasion Attacks in Deep RL
Yanchao Sun, Ruijie Zheng, Yongyuan Liang, Furong Huang |
ICLR | 1 |
| 2022 | Transfer RL across Observation Feature Spaces via Model-Based Regularization
Yanchao Sun, Ruijie Zheng, Andrew E. Cohen, Furong Huang |
ICLR | 1 |
| 2022 | Distributional Reward Estimation for Effective Multi-agent Deep Reinforcement LearningabstractMulti-agent reinforcement learning has drawn increasing attention in practice, e.g., robotics and automatic driving, as it can explore optimal policies using samples generated by interacting with the environment. However, high reward uncertainty still remains a problem when we want to train a satisfactory model, because obtaining high-quality reward feedback is usually expensive and even infeasible. To handle this issue, previous methods mainly focus on passive reward correction. At the same time, recent active reward estimation methods have proven to be a recipe for reducing the effect of reward uncertainty. In this paper, we propose a novel Distributional Reward Estimation framework for effective Multi-Agent Reinforcement Learning (DRE-MARL). Our main idea is to design the multi-action-branch reward estimation and policy-weighted reward aggregation for stabilized training. Specifically, we design the multi-action-branch reward estimation to model reward distributions on all action branches. Then we utilize reward aggregation to obtain stable updating signals during training. Our intuition is that consideration of all possible consequences of actions could be useful for learning policies. The superiority of the DRE-MARL is demonstrated using benchmark multi-agent scenarios, compared with the SOTA baselines in terms of both effectiveness and robustness. Jifeng Hu, Yanchao Sun, Hechang Chen, Sili Huang, Haiyin Piao, Yi Chang 0001, Lichao Sun 0001 |
NeurIPS | 2 |
| 2022 | Efficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement LearningabstractRecent studies reveal that a well-trained deep reinforcement learning (RL) policy can be particularly vulnerable to adversarial perturbations on input observations. Therefore, it is crucial to train RL agents that are robust against any attacks with a bounded budget. Existing robust training methods in deep RL either treat correlated steps separately, ignoring the robustness of long-term rewards, or train the agents and RL-based attacker together, doubling the computational burden and sample complexity of the training process. In this work, we propose a strong and efficient robust training framework for RL, named Worst-case-aware Robust RL (WocaR-RL) that directly estimates and optimizes the worst-case reward of a policy under bounded l_p attacks without requiring extra samples for learning an attacker. Experiments on multiple environments show that WocaR-RL achieves state-of-the-art performance under various strong attacks, and obtains significantly higher training efficiency than prior state-of-the-art robust training methods. The code of this work is available at https://github.com/umd-huang-lab/WocaR-RL. Yongyuan Liang, Yanchao Sun, Ruijie Zheng, Furong Huang |
NeurIPS | 2 |
| 2022 | Adversarial Auto-Augment with Label Preservation: A Representation Learning Principle Guided ApproachabstractData augmentation is a critical contributing factor to the success of deep learning but heavily relies on prior domain knowledge which is not always available. Recent works on automatic data augmentation learn a policy to form a sequence of augmentation operations, which are still pre-defined and restricted to limited options. In this paper, we show that a prior-free autonomous data augmentation's objective can be derived from a representation learning principle that aims to preserve the minimum sufficient information of the labels. Given an example, the objective aims at creating a distant ``hard positive example'' as the augmentation, while still preserving the original label. We then propose a practical surrogate to the objective that can be optimized efficiently and integrated seamlessly into existing methods for a broad class of machine learning tasks, e.g., supervised, semi-supervised, and noisy-label learning. Unlike previous works, our method does not require training an extra generative model but instead leverages the intermediate layer representations of the end-task model for generating data augmentations. In experiments, we show that our method consistently brings non-trivial improvements to the three aforementioned learning tasks from both efficiency and final performance, either or not combined with pre-defined augmentations, e.g., on medical images when domain knowledge is unavailable and the existing augmentation techniques perform poorly. Code will be released publicly. Yanchao Sun, Jiahao Su, Fengxiang He, Xinmei Tian 0001, Furong Huang, Tianyi Zhou 0001, Dacheng Tao |
NeurIPS | 2 |
| 2022 | Distributed adaptive neural network constraint containment control for the benthic autonomous underwater vehicles
Yanchao Sun, Yutong Du, Hongde Qin |
Neurocomputing | 1 |
| 2021 | Research on the Model of Academic Status Based on Deep LearningabstractFor universities, the quality of talent training is particularly important. This article attempts to use machine learning methods to build a predictive model, using undergraduates’ entrance scores and the scores obtained in the first two years of university as the data source, With the increase of training times, the prediction accuracy of the model proposed in this paper is better than other algorithms, which can explain the method proposed in this paper. Has a certain universality and promotion. Yanchao Sun, Minzheng Jia |
ICIS | 1 |
| 2021 | A New Wolves Intelligent Optimization AlgorithmabstractThis paper proposes a new wolf pack intelligent optimization algorithm, which is based on an adaptive shrinking grid search chaotic wolf optimization algorithm with an adaptive standard deviation update amount. Theoretical research and experimental results show that compared with traditional genetic algorithm, particle swarm algorithm, wolf pack optimization algorithm based on leadership strategy, and chaotic wolf pack optimization algorithm, the algorithm proposed in this paper has better global optimization accuracy under the same conditions. Faster convergence speed and higher robustness. Yanchao Sun, Yuqing Wu |
ICIS | 1 |
| 2021 | TempLe: Learning Template of Transitions for Sample Efficient Multi-task RLabstractTransferring knowledge among various environments is important for efficiently learning multiple tasks online. Most existing methods directly use the previously learned models or previously learned optimal policies to learn new tasks. However, these methods may be inefficient when the underlying models or optimal policies are substantially different across tasks. In this paper, we propose Template Learning (TempLe), a PAC-MDP method for multi-task reinforcement learning that could be applied to tasks with varying state/action space without prior knowledge of inter-task mappings. TempLe gains sample efficiency by extracting similarities of the transition dynamics across tasks even when their underlying models or optimal policies have limited commonalities. We present two algorithms for an ``online'' and a ``finite-model'' setting respectively. We prove that our proposed TempLe algorithms achieve much lower sample complexity than single-task learners or state-of-the-art multi-task methods. We show via systematically designed experiments that our TempLe method universally outperforms the state-of-the-art multi-task methods (PAC-MDP or not) in various settings and regimes. Yanchao Sun, Furong Huang |
AAAI | 1 |
| 2021 | Vulnerability-Aware Poisoning Mechanism for Online RL with Unknown Dynamics
Yanchao Sun, Furong Huang |
ICLR | 1 |
| 2020 | Understanding Generalization in Deep Learning via Tensor MethodsabstractDeep neural networks generalize well on unseen data though the number of parameters often far exceeds the number of training examples. Recently proposed complexity measures have provided insights to understanding the generalizability in neural networks from perspectives of PAC-Bayes, robustness, overparametrization, compression and so on. In this work, we advance the understanding of the relations between the network’s architecture and its generalizability from the compression perspective. Using tensor analysis, we propose a series of intuitive, data-dependent and easily-measurable properties that tightly characterize the compressibility and generalizability of neural networks; thus, in practice, our generalization bound outperforms the previous compression-based ones, especially for neural networks using tensors as their weight kernels (e.g. CNNs). Moreover, these intuitive measurements provide further insights into designing neural network architectures with properties favorable for better/guaranteed generalizability. Our experimental results demonstrate that through the proposed measurable properties, our generalization error bound matches the trend of the test error well. Our theoretical analysis further provides justifications for the empirical success and limitations of some widely-used tensor-based compression approaches. We also discover the improvements to the compressibility and robustness of current neural networks when incorporating tensor operations via our proposed layer-wise structure. Yanchao Sun, Jiahao Su, Taiji Suzuki, Furong Huang |
AISTATS | 2 |
| 2020 | Adaptive Fault-Tolerant Prescribed-Time Control for Teleoperation Systems With Position Error ConstraintsabstractIn this article, we present an adaptive prescribed-time control method for a class of nonlinear telerobotic systems with actuator faults and position error constraints. Extended from prescribed-time stability, practically prescribed-time stability (PPTS) is proposed for the first time aiming at stability analysis and control synthesis of nonlinear systems with disturbance and uncertainty. We show that, under the control scheme in the framework of PPTS, the system states are guaranteed to converge to a user-defined set (physically realizable) within user-defined settling time (physically realizable). Based on PPTS, an adaptive fault-tolerant controller is developed by integrating a novel exponential-type barrier Lyapunov function. Rigorous stability analysis based on back-stepping approach proves that, under the proposed control strategy, synchronization errors converge to a user-defined residual-set within predefined settling time and never exceed the prescribed range. Universal performance indexes, including the settling time, residual-set, accuracy, and overshoot, can be user-defined and only dependent on fewer user-defined parameters. Simulation results illustrate the effectiveness of the developed control scheme. Ziwei Wang 0001, Bin Liang 0001, Yanchao Sun, Tao Zhang 0006 |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Research on Undergraduate Academic Prediction Model Based on Deep LearningabstractAcademic prediction is an important management means to strengthen the construction of study style and improve the quality of talent training in Colleges and universities. In view of the current colleges and universities lack of scientific and effective methods in the undergraduate academic prediction, based on deep learning of undergraduate academic prediction model, the model mainly includes the construction of feature reconstruction and prediction model of the two modules, the feature extraction module, including the selection of students, characteristics of format conversion and dimension recombination; model the convolutional neural network the characteristics of the input after the reorganization of the improved training and validation, the model parameters were optimized, and then use the model to forecast. Taking the data from 2007 to 2010 students in a university in Beijing as training set, the model is trained, and 2011 level students are selected as the validation set to optimize the model parameters, and the 2012 level student data set is used to evaluate the model. The results show that the undergraduate academic prediction model based on CNN neural network has higher prediction accuracy. And then help colleges and universities to optimize the teaching management methods and improve the quality of teaching. Yanchao Sun, Minzheng Jia |
ICIS | 1 |
| 2018 | An Improved Deep Belief Network Prediction MethodabstractDeep belief network applying unsupervised methods of greedy layer training, from the training set automatic feature extraction value, will cause the error by layer transfer, thus affecting the accuracy of the model prediction, in order to solve this problem, proposed using conjugate gradient algorithm in gradient descent can accelerate the convergence of ideas, improvement of the restricted Boltzmann machine network algorithm in the depth of confidence, first from the complexity of the algorithm and the reconstruction error analysis of improved model differences and advantages, and classify the verification on the MNIST data set, and a detailed analysis of the feasibility of the improved model and efficiency, the experimental results shows the feature extraction ability improved deep belief network model has better and classification results. Yanchao Sun, Minzheng Jia |
ICIS | 1 |
| 2017 | Collaborative Inference of Coexisting Information DiffusionsabstractThe purpose of diffusion history inference is to reconstruct the missing traces of information diffusion according to incomplete observations. Existing methods, however, often focus only on single diffusion trace, while in a real-world social network, there often coexist multiple information diffusions. In this paper, we propose a novel approach called Collaborative Inference Model (CIM) for the problem of the inference of coexisting information diffusions. CIM can holistically model multiple information diffusions without any prior assumption of diffusion models, and collaboratively infer the histories of the coexisting information diffusions via low-rank approximation with a fusion of heterogeneous constraints generated from additional data sources. We also propose an optimized algorithm called Time Window based Parallel Decomposition Algorithm (TWPDA) to speed up the inference without compromise on the accuracy. Extensive experiments are conducted on real-world datasets to verify the effectiveness and efficiency of CIM and TWPDA. Yanchao Sun, Cong Qian, Ning Yang 0001, Philip S. Yu |
ICDM | 1 |
| 2014 | Dynamic Output Feedback Guaranteed Cost Control for T-S Fuzzy Systems with Uncertainties and Time Delays
Guangfu Ma, Yanchao Sun, Chuanjiang Li |
ICIC (2) | 2 |
| 2014 | On Spacecraft Relative Orbital Motion Based on Main-Flying Direction Method
Yanchao Sun, Huixiang Ling, Chuanjiang Li, Guangfu Ma |
ICIC (2) | 1 |
| 2013 | PEEC Modeling for Linear and Platy Structures with Efficient Capacitance Calculations
Yanchao Sun, Xinwei Song |
ICIC (2) | 1 |