Yunzhe Tao

dblp:196/4687 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0001-5819-5304ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Reinforcement learning · 28% Trustworthy machine learning · 14% Vision and language · 14%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
deep reinforcement learning
1.222024
B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis · ICLR 2024
DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning · ICRA 2020
Machine learning › Representation and self-supervised learning
redundancy reduction
0.812024
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction · ACL (1) 2024
Machine learning › Reinforcement learning
value-based reinforcement learning
0.812024
B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis · ICLR 2024
Computer vision › Video understanding and tracking
video annotation
0.812024
DeVAn: Dense Video Annotation for Video-Language Models · ACL (1) 2024
Computer vision › Vision and language
video-language model
0.812024
DeVAn: Dense Video Annotation for Video-Language Models · ACL (1) 2024
Machine learning › Graph learning
graph neural network
0.712023
DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation · WSDM 2023
Recommender systems
diversified recommendation
0.712023
DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation · WSDM 2023
Recommender systems
graph-based recommendation
0.712023
DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation · WSDM 2023
Machine learning › Reinforcement learning
policy learning
0.512021
REPAINT: Knowledge Transfer in Deep Reinforcement Learning · ICML 2021
Robotics › Autonomous driving › autonomous ground vehicle
autonomous racing
0.412020
DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning · ICRA 2020
Machine learning › Trustworthy machine learning › uncertainty estimation
model uncertainty
0.412020
Robust Multi-Agent Reinforcement Learning with Model Uncertainty · NeurIPS 2020
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412020
Robust Multi-Agent Reinforcement Learning with Model Uncertainty · NeurIPS 2020
Machine learning › Trustworthy machine learning
robustness
0.412020
Robust Multi-Agent Reinforcement Learning with Model Uncertainty · NeurIPS 2020
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.312018
A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization · IJCAI 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
bayesian nonparametric regression
0.312018
Explaining Deep Learning Models - A Bayesian Non-parametric Approach · NeurIPS 2018
Machine learning › Trustworthy machine learning › interpretability › model explanation
global explanation
0.312018
Explaining Deep Learning Models - A Bayesian Non-parametric Approach · NeurIPS 2018
Machine learning › Trustworthy machine learning
interpretability
0.312018
Explaining Deep Learning Models - A Bayesian Non-parametric Approach · NeurIPS 2018
Machine learning › Deep learning architectures and training › attention mechanism › attention network
non-local neural network
0.312018
Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal Modeling · NeurIPS 2018
Machine learning › Learning theory
spectral analysis
0.312018
Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal Modeling · NeurIPS 2018
Natural language and speech › Language models and text generation
text summarization
0.312018
A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization · IJCAI 2018
Natural language and speech › Language models and text generation › text summarization › controllable summarization
topic-aware summarization
0.312018
A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization · IJCAI 2018
Recommender systems › beyond-accuracy recommendation
long-tail recommendation
0.212023
DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation · WSDM 2023
Machine learning › Reinforcement learning › off-policy reinforcement learning › experience replay
experience selection
0.112021
REPAINT: Knowledge Transfer in Deep Reinforcement Learning · ICML 2021
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.112021
REPAINT: Knowledge Transfer in Deep Reinforcement Learning · ICML 2021
Robotics › Motion planning and robot control
path planning
0.112020
DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning · ICRA 2020
Robotics › Motion planning and robot control
robot control
0.112020
DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning · ICRA 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › markov processes
markov jump processes
0.112018
Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal Modeling · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

submodular neighbor selection · 1.3loss reweighting · 1.3layer attention · 1.3redundancy reduction · 0.8dense video annotation · 0.8deep reinforcement learning · 0.8representation transfer · 0.5advantage-based experience selection · 0.5sim-to-real transfer · 0.4model-free reinforcement learning · 0.4
YearPublicationVenuePosition
2024 Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
abstract
Yiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang, Soroush Vosoughi, Hongxia Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yiren Jian, Tingkai Liu, Yunzhe Tao, Soroush Vosoughi, Hongxia Yang
ACL (1)3
2024 DeVAn: Dense Video Annotation for Video-Language Models
abstract
Tingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fang, Ding Zhou, Huaibo Huang, Ran He, Hongxia Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan, Huaibo Huang, Ran He 0001, Hongxia Yang
ACL (1)2
2024 B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis
Zishun Yu, Yunzhe Tao, Liyu Chen, Hongxia Yang
ICLR2
2023 DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation
abstract
Graph Neural Network (GNN) based recommender systems have been attracting more and more attention in recent years due to their excellent performance in accuracy. Representing user-item interactions as a bipartite graph, a GNN model generates user and item representations by aggregating embeddings of their neighbors. However, such an aggregation procedure often accumulates information purely based on the graph structure, overlooking the redundancy of the aggregated neighbors and resulting in poor diversity of the recommended list. In this paper, we propose diversifying GNN-based recommender systems by directly improving the embedding generation procedure. Particularly, we utilize the following three modules: submodular neighbor selection to find a subset of diverse neighbors to aggregate for each GNN node, layer attention to assign attention weights for each layer, and loss reweighting to focus on the learning of items belonging to long-tail categories. Blending the three modules into GNN, we present DGRec (Diversified GNN-based Recommender System) for diversified recommendation. Experiments on real-world datasets demonstrate that the proposed method can achieve the best diversity while keeping the accuracy comparable to state-of-the-art GNN-based recommender systems. We open source DGRec at https://github.com/YangLiangwei/DGRec.
Liangwei Yang, Shengjie Wang 0001, Yunzhe Tao, Jiankai Sun, Xiaolong Liu 0012, Philip S. Yu, Taiqing Wang
WSDM3
2021 REPAINT: Knowledge Transfer in Deep Reinforcement Learning
abstract
Accelerating learning processes for complex tasks by leveraging previously learned tasks has been one of the most challenging problems in reinforcement learning, especially when the similarity between source and target tasks is low. This work proposes REPresentation And INstance Transfer (REPAINT) algorithm for knowledge transfer in deep reinforcement learning. REPAINT not only transfers the representation of a pre-trained teacher policy in the on-policy learning, but also uses an advantage-based experience selection approach to transfer useful samples collected following the teacher policy in the off-policy learning. Our experimental results on several benchmark tasks show that REPAINT significantly reduces the total training time in generic cases of task similarity. In particular, when the source tasks are dissimilar to, or sub-tasks of, the target tasks, REPAINT outperforms other baselines in both training-time reduction and asymptotic performance of return scores.
Yunzhe Tao, Sahika Genc, Jonathan Chung 0001, Tao Sun 0008, Sunil Mallya
ICML1
2020 DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning
abstract
DeepRacer is a platform for end-to-end experimentation with RL and can be used to systematically investigate the key challenges in developing intelligent control systems. Using the platform, we demonstrate how a 1/18th scale car can learn to drive autonomously using RL with a monocular camera. It is trained in simulation with no additional tuning in the physical world and demonstrates: 1) formulation and solution of a robust reinforcement learning algorithm, 2) narrowing the reality gap through joint perception and dynamics, 3) distributed on-demand compute architecture for training optimal policies, and 4) a robust evaluation method to identify when to stop training. It is the first successful large-scale deployment of deep reinforcement learning on a robotic control agent that uses only raw camera images as observations and a model-free learning method to perform robust path planning. We open source our code and video demo on GitHub2.
Bharathan Balaji, Sunil Mallya, Sahika Genc, Leo Dirac, Vineet Khare, Gourav Roy, Tao Sun 0008, Yunzhe Tao, Brian Townsend, Eddie Calleja, Sunil Muralidhara, Dhanasekar Karuppasamy
ICRA9
2020 Robust Multi-Agent Reinforcement Learning with Model Uncertainty
abstract
In this work, we study the problem of multi-agent reinforcement learning (MARL) with model uncertainty, which is referred to as robust MARL. This is naturally motivated by some multi-agent applications where each agent may not have perfectly accurate knowledge of the model, e.g., all the reward functions of other agents. Little a priori work on MARL has accounted for such uncertainties, neither in problem formulation nor in algorithm design. In contrast, we model the problem as a robust Markov game, where the goal of all agents is to find policies such that no agent has the incentive to deviate, i.e., reach some equilibrium point, which is also robust to the possible uncertainty of the MARL model. We first introduce the solution concept of robust Nash equilibrium in our setting, and develop a Q-learning algorithm to find such equilibrium policies, with convergence guarantees under certain conditions. In order to handle possibly enormous state-action spaces in practice, we then derive the policy gradients for robust MARL, and develop an actor-critic algorithm with function approximation. Our experiments demonstrate that the proposed algorithm outperforms several baseline MARL methods that do not account for the model uncertainty, in several standard but uncertain cooperative and competitive MARL environments.
Kaiqing Zhang, Tao Sun 0008, Yunzhe Tao, Sahika Genc, Sunil Mallya, Tamer Basar
NeurIPS3
2018 A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization
abstract
In this paper, we propose a deep learning approach to tackle the automatic summarization tasks by incorporating topic information into the convolutional sequence-to-sequence (ConvS2S) model and using self-critical sequence training (SCST) for optimization. Through jointly attending to topics and word-level alignment, our approach can improve coherence, diversity, and informativeness of generated summaries via a biased probability generation mechanism. On the other hand, reinforcement training, like SCST, directly optimizes the proposed model with respect to the non-differentiable metric ROUGE, which also avoids the exposure bias during inference. We carry out the experimental evaluation with state-of-the-art methods over the Gigaword, DUC-2004, and LCSTS datasets. The empirical results demonstrate the superiority of our proposed method in the abstractive summarization.
Li Wang 0092, Junlin Yao, Yunzhe Tao, Wei Liu 0005, Qiang Du 0001
IJCAI3
2018 Explaining Deep Learning Models - A Bayesian Non-parametric Approach
abstract
Understanding and interpreting how machine learning (ML) models make decisions have been a big challenge. While recent research has proposed various technical approaches to provide some clues as to how an ML model makes individual predictions, they cannot provide users with an ability to inspect a model as a complete entity. In this work, we propose a novel technical approach that augments a Bayesian non-parametric regression mixture model with multiple elastic nets. Using the enhanced mixture model, we can extract generalizable insights for a target model through a global approximation. To demonstrate the utility of our approach, we evaluate it on different ML models in the context of image recognition. The empirical results indicate that our proposed approach not only outperforms the state-of-the-art techniques in explaining individual decisions but also provides users with an ability to discover the vulnerabilities of the target ML models.
Wenbo Guo 0002, Sui Huang, Yunzhe Tao, Xinyu Xing 0001, Lin Lin 0003
NeurIPS3
2018 Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal Modeling
abstract
Nonlocal neural networks have been proposed and shown to be effective in several computer vision tasks, where the nonlocal operations can directly capture long-range dependencies in the feature space. In this paper, we study the nature of diffusion and damping effect of nonlocal networks by doing spectrum analysis on the weight matrices of the well-trained networks, and then propose a new formulation of the nonlocal block. The new block not only learns the nonlocal interactions but also has stable dynamics, thus allowing deeper nonlocal structures. Moreover, we interpret our formulation from the general nonlocal modeling perspective, where we make connections between the proposed nonlocal network and other nonlocal models, such as nonlocal diffusion process and Markov jump process.
Yunzhe Tao, Qiang Du 0001, Wei Liu 0005
NeurIPS1