VLDB 2026 Research / reviewers in the wild / expert
Tiantian Zhang 0002
dblp:92/6866-2 · also Tian-Tian Zhang 0002
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-0204-4758ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 48% Reinforcement learning · 21% Robot manipulation · 16% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
masked generative modeling |
0.9 | 1 | 2025 | Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation · NeurIPS 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation · NeurIPS 2025 |
Robotics › Motion planning and robot control › design optimization
co-design of morphology and control |
0.7 | 1 | 2023 | Curriculum-based Co-design of Morphology and Control of Voxel-based Soft Robots · ICLR 2023 |
Robotics › Robot manipulation
soft robotics |
0.7 | 1 | 2023 | Curriculum-based Co-design of Morphology and Control of Voxel-based Soft Robots · ICLR 2023 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
masked generative model · 0.9group relative policy optimization · 0.9evolutionary algorithm · 0.7curriculum learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Surrogate-Assisted Evolutionary Multi-Agent Reinforcement Learning with Adaptive Fitness EvaluationabstractDeep Multi-Agent Reinforcement Learning (MARL) excels in co-operative tasks but often struggles with local optima in high - dimensional joint action spaces. In contrast, Evolutionary Algorithms (EAs) offer robust global exploration capabilities. Although hybrid approaches seek to combine the strengths of both paradigms, they typically face a critical bottleneck: the prohibitive sample cost of evaluating large populations via environment rollouts. To address this challenge, we propose Surrogate-assisted Evolutionary Multi-Agent Reinforcement Learning (SEMARL), a unified framework that synergizes gradient-based refinement with surrogate-assisted evolutionary search. SEMARL employs a cooperative co-evolutionary architecture to maintain diverse agent policies and injects gradient-refined parameters into the population to accelerate convergence. Crucially, we leverage the centralized critic from the gradient learner as a computationally efficient surrogate for fitness estimation. To prevent misleading guidance from an inaccurate critic, we introduce an adaptive reliability control mechanism based on temporal difference (TD) error, which dynamically regulates the surrogate's influence. Experiments on the Multi-Agent MuJoCo benchmark demonstrate that SEMARL significantly outperforms other algorithms, achieving superior asymptotic performance with substantially higher sample efficiency. Cong Yu 0018, Zaihui Yang, Haoyu Wang 0018, Junbo Tan, Yongzhe Chang, Tiantian Zhang 0002, Xueqian Wang 0001 |
GECCO | 8 |
| 2026 | Tacit mechanism: Bridging pre-training of individuality to multi-agent adversarial coordination
Shiqing Yao, Jiajun Chai, Haixin Yu, Yongzhe Chang, Tiantian Zhang 0002, Yuanheng Zhu, Xueqian Wang 0001 |
Neural Networks | 5 |
| 2025 | Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image GenerationabstractReinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO, the first method to incorporate Group Relative Policy Optimization (GRPO)-based RL into this overlooked paradigm. Our core insight is to redefine the transition probability, which is different from current approaches, and formulate the unmasking process as a multi-step decision-making problem. To further enhance our method, we explore several useful strategies, including removing the Kullback–Leibler constraint, applying the reduction strategy, and filtering out low-quality samples. Using Mask-GRPO, we improve a base model, Show-o, with substantial improvements on standard T2I benchmarks and preference alignment, outperforming existing state-of-the-art approaches. Yifu Luo, Xinhao Hu, Keyu Fan, Bo Xia, Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001 |
NeurIPS | 7 |
| 2024 | Dynamics-Adaptive Continual Reinforcement Learning via Progressive ContextualizationabstractA key challenge of continual reinforcement learning (CRL) in dynamic environments is to promptly adapt the reinforcement learning (RL) agent's behavior as the environment changes over its lifetime while minimizing the catastrophic forgetting of the learned information. To address this challenge, in this article, we propose DaCoRL, that is, dynamics-adaptive continual RL. DaCoRL learns a context-conditioned policy using progressive contextualization, which incrementally clusters a stream of stationary tasks in the dynamic environment into a series of contexts and opts for an expandable multihead neural network to approximate the policy. Specifically, we define a set of tasks with similar dynamics as an environmental context and formalize context inference as a procedure of online Bayesian infinite Gaussian mixture clustering on environment features, resorting to online Bayesian inference to infer the posterior distribution over contexts. Under the assumption of a Chinese restaurant process (CRP) prior, this technique can accurately classify the current task as a previously seen context or instantiate a new context as needed without relying on any external indicator to signal environmental changes in advance. Furthermore, we employ an expandable multihead neural network whose output layer is synchronously expanded with the newly instantiated context and a knowledge distillation regularization term for retaining the performance on learned tasks. As a general framework that can be coupled with various deep RL algorithms, DaCoRL features consistent superiority over existing methods in terms of stability, overall performance, and generalization ability, as verified by extensive experiments on several robot navigation and MuJoCo locomotion tasks. Tiantian Zhang 0002, Zichuan Lin, Deheng Ye, Qiang Fu 0016, Wei Yang 0032, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003, Xiu Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Curriculum-based Co-design of Morphology and Control of Voxel-based Soft Robots
Haobo Fu, Qiang Fu 0016, Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001 |
ICLR | 5 |
| 2023 | Catastrophic Interference in Reinforcement Learning: A Solution Based on Context Division and Knowledge DistillationabstractThe powerful learning ability of deep neural networks enables reinforcement learning (RL) agents to learn competent control policies directly from continuous environments. In theory, to achieve stable performance, neural networks assume identically and independently distributed (i.i.d.) inputs, which unfortunately does not hold in the general RL paradigm where the training data are temporally correlated and nonstationary. This issue may lead to the phenomenon of "catastrophic interference" and the collapse in performance. In this article, we present interference-aware deep Q-learning (IQ) to mitigate catastrophic interference in single-task deep RL. Specifically, we resort to online clustering to achieve on-the-fly context division, together with a multihead network and a knowledge distillation regularization term for preserving the policy of learned contexts. Built upon deep Q networks (DQNs), IQ consistently boosts the stability and performance when compared to existing methods, verified with extensive experiments on classic control and Atari tasks. The code is publicly available at https://github.com/ Sweety-dm/Interference-aware-Deep-Q-learning. Tiantian Zhang 0002, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | A surrogate-assisted controller for expensive evolutionary reinforcement learning
Tiantian Zhang 0002, Yongzhe Chang, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003 |
Inf. Sci. | 2 |
| 2016 | Visualizing MOOC User Behaviors: A Case Study on XuetangX
Tiantian Zhang 0002, Bo Yuan 0003 |
IDEAL | 1 |
| 2016 | Ubiquitous Robot: A New Paradigm for Intelligence
Tiantian Zhang 0002, Bo Yuan 0003, Yinghao Ren, Houde Liu, Xueqian Wang 0001 |
IDEAL | 1 |