Zhanyi Sun

dblp:324/2470 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 37% Reinforcement learning · 32% Motion planning and robot control · 17%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.222023
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design · HPCA 2023
SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning · ECCV (11) 2022
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning
0.912025
Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution · NeurIPS 2025
Robotics › Robot manipulation
diffusion policy
0.912025
Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution · NeurIPS 2025
Robotics › Motion planning and robot control › robot learning › visuomotor learning
visuomotor policy learning
0.912025
Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
preference-based reward learning
0.812024
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback · ICML 2024
Machine learning › Reinforcement learning
reward design
0.812024
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback · ICML 2024
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
attention head pruning
0.712023
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design · HPCA 2023
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.712023
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design · HPCA 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
transformer accelerator
0.712023
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design · HPCA 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
vision transformer accelerator
0.712023
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design · HPCA 2023
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis
0.612022
SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning · ECCV (11) 2022
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.612022
SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning · ECCV (11) 2022
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
one-shot neural architecture search
0.612022
SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning · ECCV (11) 2022
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.312025
Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution · NeurIPS 2025
Computer vision › Vision and language › vision-language model
vision-language foundation model
0.212024
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback · ICML 2024
Hardware accelerators and domain-specific architectures
algorithm-hardware co-design
0.212023
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design · HPCA 2023
Robotics › Motion planning and robot control › motion planning › geometric motion planning
high-dimensional motion planning
0.212022
Human-Guided Motion Planning in Partially Observable Environments · ICRA 2022

Methods — techniques the papers use, named apart from their topics

autoencoder-based data movement reduction · 1.3attention map pruning · 1.3diffusion model · 0.9control barrier functions · 0.9vision-language foundation model · 0.8preference learning · 0.8parameter pruning · 0.6motion-planner-guided discrete task model · 0.6inverse reinforcement learning · 0.6architecture search · 0.6
YearPublicationVenuePosition
2025 Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
abstract
Visuomotor policies trained via behavior cloning are vulnerable to covariate shift, where small deviations from expert trajectories can compound into failure. Common strategies to mitigate this issue involve expanding the training distribution through human-in-the-loop corrections or synthetic data augmentation. However, these approaches are often labor-intensive, rely on strong task assumptions, or compromise the quality of imitation. We introduce Latent Policy Barrier, a framework for robust visuomotor policy learning. Inspired by Control Barrier Functions, LPB treats the latent embeddings of expert demonstrations as an implicit barrier separating safe, in-distribution states from unsafe, out-of-distribution (OOD) ones. Our approach decouples the role of precise expert imitation and OOD recovery into two separate modules: a base diffusion policy solely on expert data, and a dynamics model trained on both expert and suboptimal policy rollout data. At inference time, the dynamics model predicts future latent states and optimizes them to stay within the expert distribution. Both simulated and real-world experiments show that LPB improves both policy robustness and data efficiency, enabling reliable manipulation from limited expert data and without additional human correction or annotation. More details are on our anonymous project website https://latentpolicybarrier.github.io.
Zhanyi Sun, Shuran Song
NeurIPS1
2024 RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
abstract
Reward engineering has long been a challenge in Reinforcement Learning (RL) research, as it often requires extensive human effort and iterative processes of trial-and-error to design effective reward functions. In this paper, we propose RL-VLM-F, a method that automatically generates reward functions for agents to learn new tasks, using only a text description of the task goal and the agent's visual observations, by leveraging feedbacks from vision language foundation models (VLMs). The key to our approach is to query these models to give preferences over pairs of the agent's image observations based on the text description of the task goal, and then learn a reward function from the preference labels, rather than directly prompting these models to output a raw reward score, which can be noisy and inconsistent. We demonstrate that RL-VLM-F successfully produces effective rewards and policies across various domains — including classic control, as well as manipulation of rigid, articulated, and deformable objects — without the need for human supervision, outperforming prior methods that use large pretrained models for reward generation under the same assumptions. Videos can be found on our project website: https://rlvlmf2024.github.io/
Yufei Wang 0007, Zhanyi Sun, Jesse Zhang, Zhou Xian, Erdem Biyik, David Held, Zackory Erickson
ICML2
2023 ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design
abstract
Vision Transformers (ViTs) have achieved state-of-the-art performance on various vision tasks. However, ViTs’ self-attention module is still arguably a major bottleneck, limiting their achievable hardware efficiency and more extensive applications to resource constrained platforms. Meanwhile, existing accelerators dedicated to NLP Transformers are not optimal for ViTs. This is because there is a large difference between ViTs and Transformers for natural language processing (NLP) tasks: ViTs have a relatively fixed number of input tokens, whose attention maps can be pruned by up to 90% even with fixed sparse patterns, without severely hurting the model accuracy (e.g.,=50%). To this end, we propose a dedicated algorithm and accelerator co-design framework dubbed ViTCoD for accelerating ViTs. Specifically, on the algorithm level, ViTCoD prunes and polarizes the attention maps to have either denser or sparser fixed patterns for regularizing two levels of workloads without hurting the accuracy, largely reducing the attention computations while leaving room for alleviating the remaining dominant data movements; on top of that, we further integrate a lightweight and learnable auto-encoder module to enable trading the dominant high-cost data movements for lower-cost computations. On the hardware level, we develop a dedicated accelerator to simultaneously coordinate the aforementioned enforced denser and sparser workloads for boosted hardware utilization, while integrating on-chip encoder and decoder engines to leverage ViTCoD’s algorithm pipeline for much reduced data movements. Extensive experiments and ablation studies validate that ViTCoD largely reduces the dominant data movement costs, achieving speedups of up to 235.3×, 142.9×, 86.0×, 10.1×, and 6.8× over general computing platforms CPUs, EdgeGPUs, GPUs, and prior-art Transformer accelerators SpAtten and Sanger under an attention sparsity of 90%, respectively. Our code implementation is available at https://github.com/GATECH-EIC/ViTCoD.
Haoran You, Zhanyi Sun, Huihong Shi, Zhongzhi Yu, Yang Zhao 0013, Yongan Zhang, Chaojian Li, Baopu Li, Yingyan (Celine) Lin
HPCA2
2022 SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning
Haoran You, Baopu Li, Zhanyi Sun, Xu Ouyang, Yingyan (Celine) Lin
ECCV (11)3
2022 Human-Guided Motion Planning in Partially Observable Environments
abstract
Motion planning is a core problem in robotics, with a range of existing methods aimed to address its diverse set of challenges. However, most existing methods rely on complete knowledge of the robot environment; an assumption that seldom holds true due to inherent limitations of robot perception. To enable tractable motion planning for high-DOF robots under partial observability, we introduce BLIND, an algorithm that leverages human guidance. BLIND utilizes inverse reinforcement learning to derive motion-level guidance from human critiques. The algorithm overcomes the computational challenge of reward learning for high-DOF robots by projecting the robot's continuous configuration space to a motion-planner-guided discrete task model. The learned reward is in turn used as guidance to generate robot motion using a novel motion planner. We demonstrate BLIND using the Fetch robot and perform two simulation experiments with partial observability. Our experiments demonstrate that, despite the challenge of partial observability and high dimensionality, BLIND is capable of generating safe robot motion and outperforms baselines on metrics of teaching efficiency, success rate, and path quality.
Carlos Quintero-Peña, Constantinos Chamzas, Zhanyi Sun, Vaibhav V. Unhelkar, Lydia E. Kavraki
ICRA3