VLDB 2026 Research / reviewers in the wild / expert
Yunzhe Tao
dblp:196/4687
· DBLP profile ↗
10ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0001-5819-5304ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Reinforcement learning · 28% Trustworthy machine learning · 14% Vision and language · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
deep reinforcement learning |
1.2 | 2 | 2024 | B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis · ICLR 2024 DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning · ICRA 2020 |
Machine learning › Representation and self-supervised learning
redundancy reduction |
0.8 | 1 | 2024 | Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction · ACL (1) 2024 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
0.8 | 1 | 2024 | B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis · ICLR 2024 |
Computer vision › Video understanding and tracking
video annotation |
0.8 | 1 | 2024 | DeVAn: Dense Video Annotation for Video-Language Models · ACL (1) 2024 |
Computer vision › Vision and language
video-language model |
0.8 | 1 | 2024 | DeVAn: Dense Video Annotation for Video-Language Models · ACL (1) 2024 |
Machine learning › Graph learning
graph neural network |
0.7 | 1 | 2023 | DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation · WSDM 2023 |
Recommender systems
diversified recommendation |
0.7 | 1 | 2023 | DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation · WSDM 2023 |
Recommender systems
graph-based recommendation |
0.7 | 1 | 2023 | DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation · WSDM 2023 |
Machine learning › Reinforcement learning
policy learning |
0.5 | 1 | 2021 | REPAINT: Knowledge Transfer in Deep Reinforcement Learning · ICML 2021 |
Robotics › Autonomous driving › autonomous ground vehicle
autonomous racing |
0.4 | 1 | 2020 | DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning · ICRA 2020 |
Machine learning › Trustworthy machine learning › uncertainty estimation
model uncertainty |
0.4 | 1 | 2020 | Robust Multi-Agent Reinforcement Learning with Model Uncertainty · NeurIPS 2020 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.4 | 1 | 2020 | Robust Multi-Agent Reinforcement Learning with Model Uncertainty · NeurIPS 2020 |
Machine learning › Trustworthy machine learning
robustness |
0.4 | 1 | 2020 | Robust Multi-Agent Reinforcement Learning with Model Uncertainty · NeurIPS 2020 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.3 | 1 | 2018 | A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization · IJCAI 2018 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
bayesian nonparametric regression |
0.3 | 1 | 2018 | Explaining Deep Learning Models - A Bayesian Non-parametric Approach · NeurIPS 2018 |
Machine learning › Trustworthy machine learning › interpretability › model explanation
global explanation |
0.3 | 1 | 2018 | Explaining Deep Learning Models - A Bayesian Non-parametric Approach · NeurIPS 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2018 | Explaining Deep Learning Models - A Bayesian Non-parametric Approach · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › attention mechanism › attention network
non-local neural network |
0.3 | 1 | 2018 | Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal Modeling · NeurIPS 2018 |
Machine learning › Learning theory
spectral analysis |
0.3 | 1 | 2018 | Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal Modeling · NeurIPS 2018 |
Natural language and speech › Language models and text generation
text summarization |
0.3 | 1 | 2018 | A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization · IJCAI 2018 |
Natural language and speech › Language models and text generation › text summarization › controllable summarization
topic-aware summarization |
0.3 | 1 | 2018 | A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization · IJCAI 2018 |
Recommender systems › beyond-accuracy recommendation
long-tail recommendation |
0.2 | 1 | 2023 | DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation · WSDM 2023 |
Machine learning › Reinforcement learning › off-policy reinforcement learning › experience replay
experience selection |
0.1 | 1 | 2021 | REPAINT: Knowledge Transfer in Deep Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.1 | 1 | 2021 | REPAINT: Knowledge Transfer in Deep Reinforcement Learning · ICML 2021 |
Robotics › Motion planning and robot control
path planning |
0.1 | 1 | 2020 | DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning · ICRA 2020 |
Robotics › Motion planning and robot control
robot control |
0.1 | 1 | 2020 | DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning · ICRA 2020 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › markov processes
markov jump processes |
0.1 | 1 | 2018 | Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal Modeling · NeurIPS 2018 |
Methods — techniques the papers use, named apart from their topics
submodular neighbor selection · 1.3loss reweighting · 1.3layer attention · 1.3redundancy reduction · 0.8dense video annotation · 0.8deep reinforcement learning · 0.8representation transfer · 0.5advantage-based experience selection · 0.5sim-to-real transfer · 0.4model-free reinforcement learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Expedited Training of Visual Conditioned Language Generation via Redundancy ReductionabstractYiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang, Soroush Vosoughi, Hongxia Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yiren Jian, Tingkai Liu, Yunzhe Tao, Soroush Vosoughi, Hongxia Yang |
ACL (1) | 3 |
| 2024 | DeVAn: Dense Video Annotation for Video-Language ModelsabstractTingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fang, Ding Zhou, Huaibo Huang, Ran He, Hongxia Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Tingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan, Huaibo Huang, Ran He 0001, Hongxia Yang |
ACL (1) | 2 |
| 2024 | B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis
Zishun Yu, Yunzhe Tao, Liyu Chen, Hongxia Yang |
ICLR | 2 |
| 2023 | DGRec: Graph Neural Network for Recommendation with Diversified Embedding GenerationabstractGraph Neural Network (GNN) based recommender systems have been attracting more and more attention in recent years due to their excellent performance in accuracy. Representing user-item interactions as a bipartite graph, a GNN model generates user and item representations by aggregating embeddings of their neighbors. However, such an aggregation procedure often accumulates information purely based on the graph structure, overlooking the redundancy of the aggregated neighbors and resulting in poor diversity of the recommended list. In this paper, we propose diversifying GNN-based recommender systems by directly improving the embedding generation procedure. Particularly, we utilize the following three modules: submodular neighbor selection to find a subset of diverse neighbors to aggregate for each GNN node, layer attention to assign attention weights for each layer, and loss reweighting to focus on the learning of items belonging to long-tail categories. Blending the three modules into GNN, we present DGRec (Diversified GNN-based Recommender System) for diversified recommendation. Experiments on real-world datasets demonstrate that the proposed method can achieve the best diversity while keeping the accuracy comparable to state-of-the-art GNN-based recommender systems. We open source DGRec at https://github.com/YangLiangwei/DGRec. Liangwei Yang, Shengjie Wang 0001, Yunzhe Tao, Jiankai Sun, Xiaolong Liu 0012, Philip S. Yu, Taiqing Wang |
WSDM | 3 |
| 2021 | REPAINT: Knowledge Transfer in Deep Reinforcement LearningabstractAccelerating learning processes for complex tasks by leveraging previously learned tasks has been one of the most challenging problems in reinforcement learning, especially when the similarity between source and target tasks is low. This work proposes REPresentation And INstance Transfer (REPAINT) algorithm for knowledge transfer in deep reinforcement learning. REPAINT not only transfers the representation of a pre-trained teacher policy in the on-policy learning, but also uses an advantage-based experience selection approach to transfer useful samples collected following the teacher policy in the off-policy learning. Our experimental results on several benchmark tasks show that REPAINT significantly reduces the total training time in generic cases of task similarity. In particular, when the source tasks are dissimilar to, or sub-tasks of, the target tasks, REPAINT outperforms other baselines in both training-time reduction and asymptotic performance of return scores. Yunzhe Tao, Sahika Genc, Jonathan Chung 0001, Tao Sun 0008, Sunil Mallya |
ICML | 1 |
| 2020 | DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement LearningabstractDeepRacer is a platform for end-to-end experimentation with RL and can be used to systematically investigate the key challenges in developing intelligent control systems. Using the platform, we demonstrate how a 1/18th scale car can learn to drive autonomously using RL with a monocular camera. It is trained in simulation with no additional tuning in the physical world and demonstrates: 1) formulation and solution of a robust reinforcement learning algorithm, 2) narrowing the reality gap through joint perception and dynamics, 3) distributed on-demand compute architecture for training optimal policies, and 4) a robust evaluation method to identify when to stop training. It is the first successful large-scale deployment of deep reinforcement learning on a robotic control agent that uses only raw camera images as observations and a model-free learning method to perform robust path planning. We open source our code and video demo on GitHub2. Bharathan Balaji, Sunil Mallya, Sahika Genc, Leo Dirac, Vineet Khare, Gourav Roy, Tao Sun 0008, Yunzhe Tao, Brian Townsend, Eddie Calleja, Sunil Muralidhara, Dhanasekar Karuppasamy |
ICRA | 9 |
| 2020 | Robust Multi-Agent Reinforcement Learning with Model UncertaintyabstractIn this work, we study the problem of multi-agent reinforcement learning (MARL) with model uncertainty, which is referred to as robust MARL. This is naturally motivated by some multi-agent applications where each agent may not have perfectly accurate knowledge of the model, e.g., all the reward functions of other agents. Little a priori work on MARL has accounted for such uncertainties, neither in problem formulation nor in algorithm design. In contrast, we model the problem as a robust Markov game, where the goal of all agents is to find policies such that no agent has the incentive to deviate, i.e., reach some equilibrium point, which is also robust to the possible uncertainty of the MARL model. We first introduce the solution concept of robust Nash equilibrium in our setting, and develop a Q-learning algorithm to find such equilibrium policies, with convergence guarantees under certain conditions. In order to handle possibly enormous state-action spaces in practice, we then derive the policy gradients for robust MARL, and develop an actor-critic algorithm with function approximation. Our experiments demonstrate that the proposed algorithm outperforms several baseline MARL methods that do not account for the model uncertainty, in several standard but uncertain cooperative and competitive MARL environments. Kaiqing Zhang, Tao Sun 0008, Yunzhe Tao, Sahika Genc, Sunil Mallya, Tamer Basar |
NeurIPS | 3 |
| 2018 | A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text SummarizationabstractIn this paper, we propose a deep learning approach to tackle the automatic summarization tasks by incorporating topic information into the convolutional sequence-to-sequence (ConvS2S) model and using self-critical sequence training (SCST) for optimization. Through jointly attending to topics and word-level alignment, our approach can improve coherence, diversity, and informativeness of generated summaries via a biased probability generation mechanism. On the other hand, reinforcement training, like SCST, directly optimizes the proposed model with respect to the non-differentiable metric ROUGE, which also avoids the exposure bias during inference. We carry out the experimental evaluation with state-of-the-art methods over the Gigaword, DUC-2004, and LCSTS datasets. The empirical results demonstrate the superiority of our proposed method in the abstractive summarization. Li Wang 0092, Junlin Yao, Yunzhe Tao, Wei Liu 0005, Qiang Du 0001 |
IJCAI | 3 |
| 2018 | Explaining Deep Learning Models - A Bayesian Non-parametric ApproachabstractUnderstanding and interpreting how machine learning (ML) models make decisions have been a big challenge. While recent research has proposed various technical approaches to provide some clues as to how an ML model makes individual predictions, they cannot provide users with an ability to inspect a model as a complete entity. In this work, we propose a novel technical approach that augments a Bayesian non-parametric regression mixture model with multiple elastic nets. Using the enhanced mixture model, we can extract generalizable insights for a target model through a global approximation. To demonstrate the utility of our approach, we evaluate it on different ML models in the context of image recognition. The empirical results indicate that our proposed approach not only outperforms the state-of-the-art techniques in explaining individual decisions but also provides users with an ability to discover the vulnerabilities of the target ML models. Wenbo Guo 0002, Sui Huang, Yunzhe Tao, Xinyu Xing 0001, Lin Lin 0003 |
NeurIPS | 3 |
| 2018 | Nonlocal Neural Networks, Nonlocal Diffusion and Nonlocal ModelingabstractNonlocal neural networks have been proposed and shown to be effective in several computer vision tasks, where the nonlocal operations can directly capture long-range dependencies in the feature space. In this paper, we study the nature of diffusion and damping effect of nonlocal networks by doing spectrum analysis on the weight matrices of the well-trained networks, and then propose a new formulation of the nonlocal block. The new block not only learns the nonlocal interactions but also has stable dynamics, thus allowing deeper nonlocal structures. Moreover, we interpret our formulation from the general nonlocal modeling perspective, where we make connections between the proposed nonlocal network and other nonlocal models, such as nonlocal diffusion process and Markov jump process. Yunzhe Tao, Qiang Du 0001, Wei Liu 0005 |
NeurIPS | 1 |