Zishun Yu

dblp:320/4542 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 64% Efficient and distributed learning · 18% Language models and text generation · 7%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
imitation learning
1.012026
Language Model Distillation: A Temporal Difference Imitation Learning Perspective · AAAI 2026
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
LLM distillation
1.012026
Language Model Distillation: A Temporal Difference Imitation Learning Perspective · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
Language Model Distillation: A Temporal Difference Imitation Learning Perspective · AAAI 2026
Machine learning › Reinforcement learning
temporal difference learning
1.012026
Language Model Distillation: A Temporal Difference Imitation Learning Perspective · AAAI 2026
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.912025
Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning · AAAI 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
0.912025
Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning · AAAI 2025
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning · AAAI 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization · ICML 2025
Machine learning › Reinforcement learning
deep reinforcement learning
0.812024
B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis · ICLR 2024
Machine learning › Reinforcement learning
value-based reinforcement learning
0.812024
B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis · ICLR 2024
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning
0.712023
Actor-Critic Alignment for Offline-to-Online Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning
value function
0.712023
Actor-Critic Alignment for Offline-to-Online Reinforcement Learning · ICML 2023
Machine learning › Trustworthy machine learning › robustness
certified robustness
0.612022
Certifying Robust Graph Classification under Orthogonal Gromov-Wasserstein Threats · NeurIPS 2022
Machine learning › Graph learning
graph neural network
0.612022
Certifying Robust Graph Classification under Orthogonal Gromov-Wasserstein Threats · NeurIPS 2022
Machine learning › Efficient and distributed learning › model compression
large language model compression
0.312026
Language Model Distillation: A Temporal Difference Imitation Learning Perspective · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems
agent communication
0.312025
Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning · AAAI 2025
Mathematical optimization › continuous optimization
convex optimization
0.212022
Certifying Robust Graph Classification under Orthogonal Gromov-Wasserstein Threats · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

temporal difference learning · 1.0inverse reinforcement learning · 1.0behavior cloning · 1.0utility maximization · 0.9transformer · 0.9self-consistency · 0.9graph modeling · 0.9factor-based attention · 0.9encoder-decoder architecture · 0.9deep reinforcement learning · 0.8orthogonal gromov-wasserstein discrepancy · 0.6fenchel biconjugate · 0.6convex optimization · 0.6
YearPublicationVenuePosition
2026 Language Model Distillation: A Temporal Difference Imitation Learning Perspective
abstract
Large language models have led to significant progress across many NLP tasks, although their massive sizes often incur substantial computational costs. Distillation has become a common practice to compress these large and highly capable models into smaller, more efficient ones. Many existing language model distillation methods can be viewed as behavior cloning from the perspective of imitation learning or inverse reinforcement learning. This viewpoint has inspired subsequent studies that leverage (inverse) reinforcement learning techniques, including variations of behavior cloning and temporal difference learning methods. Rather than proposing yet another specific temporal difference method, we introduce a general framework for temporal difference-based distillation by exploiting the distributional sparsity of the teacher model. Specifically, it is often observed that language models assign most probability mass to a small subset of tokens. Motivated by this observation, we design a temporal difference learning framework that operates on a reduced action space (a subset of vocabulary), and demonstrate how practical algorithms can be derived and the resulting performance improvements.
Zishun Yu, Shangzhe Li
AAAI1
2025 Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning
abstract
In multi-agent reinforcement learning, a commonly considered paradigm is centralized training with decentralized execution. However, in this framework, decentralized execution restricts the development of coordinated policies due to the local observation limitation. In this paper, we consider the cooperation among neighboring agents during execution and formulate their interactions as a graph. Thus, we introduce a novel encoder-decoder architecture named Factor-based Multi-Agent Transformer (f-MAT) that utilizes a transformer to enable communication between neighboring agents during both training and execution. By dividing agents into different overlapping groups and representing each group with a factor, f-MAT achieves efficient message passing and parallel action generation through factor-based attention layers. Empirical results in networked systems such as traffic scheduling and power control demonstrate that f-MAT achieves superior performance compared to strong baselines, thereby paving the way for handling complex collaborative problems.
Wenzhe Fan 0001, Zishun Yu, Chengdong Ma, Changye Li 0003, Yaodong Yang 0001
AAAI2
2025 Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
abstract
Solving mathematics problems has been an intriguing capability of large language models, and many efforts have been made to improve reasoning by extending reasoning length, such as through self-correction and extensive long chain-of-thoughts. While promising in problem-solving, advanced long reasoning chain models exhibit an undesired single-modal behavior, where trivial questions require unnecessarily tedious long chains of thought. In this work, we propose a way to allow models to be aware of inference budgets by formulating it as utility maximization with respect to an inference budget constraint, hence naming our algorithm Inference Budget-Constrained Policy Optimization (IBPO). In a nutshell, models fine-tuned through IBPO learn to ``understand'' the difficulty of queries and allocate inference budgets to harder ones. With different inference budgets, our best models are able to have a $4.14$\% and $5.74$\% absolute improvement ($8.08$\% and $11.2$\% relative improvement) on MATH500 using $2.16$x and $4.32$x inference budgets respectively, relative to LLaMA3.1 8B Instruct. These improvements are approximately $2$x those of self-consistency under the same budgets.
Zishun Yu, Tengyu Xu, Di Jin 0005, Karthik Abinav Sankararaman, Zhouhao Zeng, Eryk Helenowski, Sinong Wang, Hao Ma 0001
ICML1
2024 Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
abstract
Reinforcement learning generalizes multi-armed bandit problems with additional difficulties of a longer planning horizon and unknown transition kernel. We explore a black-box reduction from discounted infinite-horizon tabular reinforcement learning to multi-armed bandits, where, specifically, an independent bandit learner is placed in each state. We show that, under ergodicity and fast mixing assumptions, any slowly changing adversarial bandit algorithm achieving optimal regret in the adversarial bandit setting can also attain optimal expected regret in infinite-horizon discounted Markov decision processes, with respect to the number of rounds $T$. Furthermore, we examine our reduction using a specific instance of the exponential-weight algorithm.
Ian A. Kash, Lev Reyzin, Zishun Yu
ALT3
2024 B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis
Zishun Yu, Yunzhe Tao, Liyu Chen, Hongxia Yang
ICLR1
2024 Offline Reward Perturbation Boosts Distributional Shift in Online RL
abstract
Offline-to-online reinforcement learning has recently been shown effective in reducing the online sample complexity by first training from offline collected data. However, this additional data source may also invite new poisoning attacks that target offline training. In this work, we reveal such vulnerabilities in critic-regularized offline RL by proposing a novel data poisoning attack method, which is stealthy in the sense that the performance during the offline training remains intact, but the online fine-tuning stage will suffer a significant performance drop. Our method leverages the techniques from bi-level optimization to promote the over-estimation/distribution shift under offline-to-online reinforcement learning. Experiments on four environments confirm the satisfaction of the new stealthiness requirement, and can be effective in attacking with only a small budget and without having white-box access to the victim model.
Zishun Yu, Siteng Kang
UAI1
2023 Actor-Critic Alignment for Offline-to-Online Reinforcement Learning
abstract
Deep offline reinforcement learning has recently demonstrated considerable promises in leveraging offline datasets, providing high-quality models that significantly reduce the online interactions required for fine-tuning. However, such a benefit is often diminished due to the marked state-action distribution shift, which causes significant bootstrap error and wipes out the good initial policy. Existing solutions resort to constraining the policy shift or balancing the sample replay based on their online-ness. However, they require online estimation of distribution divergence or density ratio. To avoid such complications, we propose deviating from existing actor-critic approaches that directly transfer the state-action value functions. Instead, we post-process them by aligning with the offline learned policy, so that the $Q$-values for actions outside the offline policy are also tamed. As a result, the online fine-tuning can be simply performed as in the standard actor-critic algorithms. We show empirically that the proposed method improves the performance of the fine-tuned robotic agents on various simulated tasks.
Zishun Yu
ICML1
2022 Certifying Robust Graph Classification under Orthogonal Gromov-Wasserstein Threats
abstract
Graph classifiers are vulnerable to topological attacks. Although certificates of robustness have been recently developed, their threat model only counts local and global edge perturbations, which effectively ignores important graph structures such as isomorphism. To address this issue, we propose measuring the perturbation with the orthogonal Gromov-Wasserstein discrepancy, and building its Fenchel biconjugate to facilitate convex optimization. Our key insight is drawn from the matching loss whose root connects two variables via a monotone operator, and it yields a tight outer convex approximation for resistance distance on graph nodes. When applied to graph classification by graph convolutional networks, both our certificate and attack algorithm are demonstrated effective.
Zishun Yu
NeurIPS2
2022 Orthogonal Gromov-Wasserstein discrepancy with efficient lower bound
abstract
Comparing structured data from possibly different metric-measure spaces is a fundamental task in machine learning, with applications in, e.g., graph classification. The Gromov-Wasserstein (GW) discrepancy formulates a coupling between the structured data based on optimal transportation, tackling the incomparability between different structures by aligning the intra-relational geometries. Although efficient local solvers such as conditional gradient and Sinkhorn are available, the inherent non-convexity still prevents a tractable evaluation, and the existing lower bounds are not tight enough for practical use. To address this issue, we take inspirations from the connection with the quadratic assignment problem, and propose the orthogonal Gromov-Wasserstein (OGW) discrepancy as a surrogate of GW. It admits an efficient and closed-form lower bound with O(n^3) complexity, and directly extends to the fused Gromov-Wasserstein distance, incorporating node features into the coupling. Extensive experiments on both the synthetic and real-world datasets show the tightness of our lower bounds, and both OGW and its lower bounds efficiently deliver accurate predictions and satisfactory barycenters for graph sets.
Zishun Yu
UAI2
2022 Deep Reinforcement Learning With Graph Representation for Vehicle Repositioning
abstract
On-demand ride services, e.g., taxi companies, ridesourcing/transportation network companies (TNCs), play a vital role in nowadays transportation modes. The relocation/repositioning of idle vehicles is an important operational problem for service providers. With the recent success of deep learning, a large and growing body of literature emerged on learning-based transportation problems, including the relocation problem. In this work, we study the idle vehicle relocation problem to achieve better drivers’ profits and customers’ satisfaction. In detail, we formulate the learning problem on real-world road networks while recent repositioning works under a deep reinforcement learning scheme formulate the environment as grids. Meanwhile, a graph formulation enables learning with Graph Neural Network (GNN), which shows promising results in recent transportation studies. We empirically show that preserving the road network structure and learning the data representation with GNN contribute to improving the repositioning decision policy. Our simulation experiments are conducted on two real-world datasets, in New York City and Austin, respectively. The results show our formulation’s advantages in terms of reductions of order rejection rates and passenger waiting time, and increasing of vehicle occupancy rates and driver revenues.
Zishun Yu
IEEE Trans. Intell. Transp. Syst.1