Zheng Tian 0002

dblp:17/2752-2 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0008-0622-8512ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
abstract
Motion retrieval is crucial for motion acquisition, offering superior precision, realism, controllability, and editability compared to motion generation. Existing approaches leverage contrastive learning to construct a unified embedding space for motion retrieval from text or visual modality. However, these methods lack a more intuitive and user-friendly interaction mode and often overlook the sequential representation of most modalities for improved retrieval performance. To address these limitations, we propose a framework that aligns four modalities—text, audio, video, and motion—within a fine-grained joint embedding space, incorporating audio for the first time in motion retrieval to enhance user immersion and convenience. This fine-grained space is achieved through a sequence-level contrastive learning approach, which captures critical details across modalities for better alignment. To evaluate our framework, we augment existing text-motion datasets with synthetic but diverse audio recordings, creating two multi-modal motion retrieval datasets. Experimental results demonstrate superior performance over state-of-the-art methods across multiple sub-tasks, including an 10.16% improvement in R@10 for text-to-motion retrieval and a 25.43% improvement in R@1 for video-to-motion retrieval on the HumanML3D dataset. Furthermore, our results show that our 4-modal framework significantly outperforms its 3-modal counterpart, underscoring the potential of multi-modal motion retrieval for advancing motion acquisition.
Shiyao Yu, Zi-An Wang, Kangning Yin, Zheng Tian 0002, Weixin Si, Shihao Zou
IEEE Trans. Multim.4
2024 Tri-Modal Motion Retrieval by Learning a Joint Embedding Space
abstract
Information retrieval is an ever-evolving and crucial re-search domain. The substantial demand for high-quality human motion data especially in online acquirement has led to a surge in human motion research works. Prior works have mainly concentrated on dual-modality learning, such as text and motion tasks, but three-modality learning has been rarely explored. Intuitively, an extra introduced modality can enrich a model's application scenario, and more importantly, an adequate choice of the extra modality can also act as an intermediary and enhance the alignment between the other two disparate modalities. In this work, we introduce LAVIMO (LAnguage-VIdeo-MOtion alignment), a novel framework for three-modality learning integrating human-centric videos as an additional modality, thereby ef-fectively bridging the gap between text and motion. More-over, our approach leverages a specially designed attention mechanism to foster enhanced alignment and synergistic effects among text, video, and motion modalities. Empirically, our results on the HumanML3D and KIT-ML datasets show that LAVIMO achieves state-of-the-art performance in various motion-related cross-modal retrieval tasks, in-cluding text-to-motion, motion-to-text, video-to-motion and motion-to-video. Our project webpage can be found in https://lavimo2023.github.io/LAVIMO/.
Kangning Yin, Shihao Zou, Yuxuan Ge, Zheng Tian 0002
CVPR4
2024 RACon: Retrieval-Augmented Simulated Character Locomotion Control
abstract
In computer animation, driving a simulated character with lifelike motion is challenging. Current generative models, though able to generalize to diverse motions, often pose challenges to the responsiveness of end-user control. To address these issues, we introduce RACon: Retrieval-Augmented Simulated Character Locomotion Control. Our end-to-end hierarchical reinforcement learning method utilizes a retriever and a motion controller. The retriever searches motion experts from a user-specified database in a task-oriented fashion, which boosts the responsiveness to the user’s control. The selected motion experts and the manipulation signal are then transferred to the controller to drive the simulated character. In addition, a retrieval-augmented discriminator is designed to stabilize the training process. Our method surpasses existing techniques in both quality and quantity in locomotion control, as demonstrated in our empirical study. Moreover, by switching extensive databases for retrieval, it can adapt to distinctive motion types at run time. We will release our code upon acceptance.
Yuxuan Mu, Shihao Zou, Kangning Yin, Zheng Tian 0002, Li Cheng 0001, Weinan Zhang 0001, Jun Wang 0012
ICME4
2024 Language and Sketching: An LLM-driven Interactive Multimodal Multitask Robot Navigation Framework
abstract
The socially-aware navigation system has evolved to adeptly avoid various obstacles while performing multiple tasks, such as point-to-point navigation, human-following, and -guiding. However, a prominent gap persists: in Human-Robot Interaction (HRI), the procedure of communicating commands to robots demands intricate mathematical formulations. Furthermore, the transition between tasks does not quite possess the intuitive control and user-centric interactivity that one would desire. In this work, we propose an LLM-driven interactive multimodal multitask robot navigation framework, termed LIM2N, to solve the above new challenge in the navigation field. We achieve this by first introducing a multimodal interaction framework where language and hand-drawn inputs can serve as navigation constraints and control objectives. Next, a reinforcement learning agent is built to handle multiple tasks with the received information. Crucially, LIM2N creates smooth cooperation among the reasoning of multimodal input, multitask planning, and adaptation and processing of the intelligent sensing modules in the complicated system. Detailed experiments are conducted in both simulation and the real world demonstrating that LIM2N has solid user needs understanding, alongside an enhanced interactive experience.
Weiqin Zu, Wenbin Song, Ruiqing Chen, Ze Guo, Fanglei Sun, Zheng Tian 0002, Wei Pan 0004, Jun Wang 0012
ICRA6
2024 Off-Agent Trust Region Policy Optimization
Ruiqing Chen, Yali Du 0001, Yifan Zhong, Zheng Tian 0002, Fanglei Sun, Yaodong Yang 0001
IJCAI5
2024 Cross-Utterance Conditioned VAE for Speech Generation
abstract
Speech synthesis systems powered by neural networks hold promise for multimedia production, but frequently face issues with producing expressive speech and seamless editing. In response, we present the Cross-Utterance Conditioned Variational Autoencoder speech synthesis (CUC-VAE S2) framework to enhance prosody and ensure natural speech generation. This framework leverages the powerful representational capabilities of pre-trained language models and the re-expression abilities of variational autoencoders (VAEs). The core component of the CUC-VAE S2 framework is the cross-utterance CVAE, which extracts acoustic, speaker, and textual features from surrounding sentences to generate context-sensitive prosodic features, more accurately emulating human prosody generation. We further propose two practical algorithms tailored for distinct speech synthesis applications: CUC-VAE TTS for text-to-speech and CUC-VAE SE for speech editing. The CUC-VAE TTS is a direct application of the framework, designed to generate audio with contextual prosody derived from surrounding texts. On the other hand, the CUC-VAE SE algorithm leverages real mel spectrogram sampling conditioned on contextual information, producing audio that closely mirrors real sound and thereby facilitating flexible speech editing based on text such as deletion, insertion, and replacement. Experimental results on the LibriTTS datasets demonstrate that our proposed models significantly enhance speech synthesis and editing, producing more natural and expressive speech.
Yang Li 0116, Guangzhi Sun, Weiqin Zu, Zheng Tian 0002, Ying Wen 0001, Wei Pan 0004, Chao Zhang 0031, Jun Wang 0012, Yang Yang 0001, Fanglei Sun
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 Self-Supervised MAFENN for Classifying Low-Labeled Distorted Images Over Mobile Fading Channels
abstract
Image distortion during wireless transmission presents a significant challenge for real-world artificial intelligence (AI) applications. Recent methods have attempted to address this issue by integrating neural networks into the wireless transmission system. However, these approaches often require a large volume of labeled training data, which can be expensive and time-consuming to collect. To address this issue, we propose a novel approach,Self-SupervisedMulti-AgentFeedbackEnabledNeuralNetworks (S2MAFENN). S2MAFENN is designed to improve the efficiency of labeled data in wireless image transmission. It incorporates a Feedbacker agent that emulates the error correction mechanisms observed in primate brains and employs self-supervised contrastive learning to extract representations from unlabeled distorted images independently. From a theoretical perspective, we model the training process of S2MAFENN as a three-player Stackelberg game and provide evidence that S2MAFENN can achieve exponential convergence rates. We then empirically validate our approach by assessing the representations learned through S2MAFENN. We use varied labeled CIFAR10 and CIFAR100 data to simulate real image transmissions over the Rayleigh fading and 5G channels. Our results show that S2MAFENN matches or even surpasses the performance of state-of-the-art self-supervised training methods, even when only 50% of labels are used. Moreover, S2MAFENN yields average accuracy gains of 5.11%, 5.8%, and 4.58% with only 0.1, 0.2, and 0.5 of the labels transmitted over the 5G channel, respectively. For the downstream task of semantic segmentation over the 5G channel, S2MAFENN exhibits significant advancements on the ADE20K dataset. It achieves enhancements of approximately 7% and 8.7% in Mean IoU and DICE metrics, respectively, surpassing the performance of current state-of-the-art methods.
Yang Li 0116, Fanglei Sun, Jingchen Hu, Fan Wu 0006, Kai Li 0022, Ying Wen 0001, Zheng Tian 0002, Yaodong Yang 0001, Jiangcheng Zhu, Jun Wang 0012, Yang Yang 0001
IEEE Trans. Mob. Comput.8
2023 ROMO: Retrieval-enhanced Offline Model-based Optimization
abstract
Data-driven black-box model-based optimization (MBO) problems arise in a great number of practical application scenarios, where the goal is to find a design over the whole space maximizing a black-box target function based on a static offline dataset. In this work, we consider a more general but challenging MBO setting, named constrained MBO (CoMBO), where only part of the design space can be optimized while the rest is constrained by the environment. A new challenge arising from CoMBO is that most observed designs that satisfy the constraints are mediocre in evaluation. Therefore, we focus on optimizing these mediocre designs in the offline dataset while maintaining the given constraints rather than further boosting the best observed design in the traditional MBO setting. We propose retrieval-enhanced offline model-based optimization (ROMO), a new derivable forward approach that retrieves the offline dataset and aggregates relevant samples to provide a trusted prediction, and use it for gradient-based optimization. ROMO is simple to implement and outperforms state-of-the-art approaches in the CoMBO setting. Empirically, we conduct experiments on a synthetic Hartmann (3D) function dataset, an industrial CIO dataset, and a suite of modified tasks in the Design-Bench benchmark. Results show that ROMO performs well in a wide range of constrained optimization tasks.
Mingcheng Chen, Hulei Fan, Hongqiao Gao, Yong Yu 0001, Zheng Tian 0002
DAI7
2023 Order Matters: Agent-by-agent Policy Optimization
Xihuai Wang, Zheng Tian 0002, Ziyu Wan, Ying Wen 0001, Jun Wang 0012, Weinan Zhang 0001
ICLR2
2023 Sim-to-Real Transfer for Quadrupedal Locomotion via Terrain Transformer
abstract
Deep reinforcement learning has recently emerged as an appealing alternative for legged locomotion over multiple terrains by training a policy in physical simulation and then transferring it to the real world (i.e., sim-to-real transfer). Despite considerable progress, the capacity and scalability of traditional neural networks are still limited, which may hinder their applications in more complex environments. In contrast, the Transformer architecture has shown its superiority in a wide range of large-scale sequence modeling tasks, including natural language processing and decision-making problems. In this paper, we propose Terrain Transformer (TERT), a high-capacity Transformer model for quadrupedal locomotion control on various terrains. Furthermore, to better leverage Transformer in sim-to-real scenarios, we present a novel two-stage training framework consisting of an offline pretraining stage and an online correction stage, which can naturally integrate Transformer with privileged training. Extensive experiments in simulation demonstrate that TERT outperforms state-of-the-art baselines on different terrains in terms of return, energy consumption and control smoothness. In further real-world validation, TERT successfully traverses nine challenging terrains, including sand pit and stair down, which can not be accomplished by strong baselines.
Hang Lai, Weinan Zhang 0001, Xialin He, Zheng Tian 0002, Yong Yu 0001, Jun Wang 0012
ICRA5
2023 Multi-embodiment Legged Robot Control as a Sequence Modeling Problem
abstract
Robots are traditionally bounded by a fixed embodiment during their operational lifetime, which limits their ability to adapt to their surroundings. Co-optimizing control and morphology of a robot, however, is often inefficient due to the complex interplay between the controller and morphology. In this paper, we propose a learning-based control method that can inherently take morphology into consideration such that once the control policy is trained in the simulator, it can be easily deployed to real robots with different embodiments. In particular, we present the Embodiment-aware Transformer (EAT), an architecture that casts this control problem as conditional sequence modeling. EAT outputs the optimal actions by leveraging a causally masked Transformer. By conditioning an autoregressive model on the desired robot embodiment, past states, and actions, our EAT model can generate future actions that best fit the current robot embodiment. Experimental results show that EAT can outperform all other alternatives in embodiment-varying tasks, and succeed in an example of real-world evolution tasks: stepping down a stair through updating the morphology alone. We hope that EAT will inspire a new push toward real-world evolution across many domains, where algorithms like EAT can blaze a trail by bridging the field of evolutionary robotics and big data sequence modeling.
Weinan Zhang 0001, Hang Lai, Zheng Tian 0002, Laurent Kneip, Jun Wang 0012
ICRA4
2022 A Game-Theoretic Approach to Multi-agent Trust Region Optimization
Ying Wen 0001, Yaodong Yang 0001, Minne Li, Zheng Tian 0002, Xu Chen 0017, Jun Wang 0012
DAI5
2022 M2N: Mesh Movement Networks for PDE Solvers
abstract
Numerical Partial Differential Equation (PDE) solvers often require discretizing the physical domain by using a mesh. Mesh movement methods provide the capability to improve the accuracy of the numerical solution without introducing extra computational burden to the PDE solver, by increasing mesh resolution where the solution is not well-resolved, whilst reducing unnecessary resolution elsewhere. However, sophisticated mesh movement methods, such as the Monge-Ampère method, generally require the solution of auxiliary equations. These solutions can be extremely expensive to compute when the mesh needs to be adapted frequently. In this paper, we propose to the best of our knowledge the first learning-based end-to-end mesh movement framework for PDE solvers. Key requirements of learning-based mesh movement methods are: alleviating mesh tangling, boundary consistency, and generalization to mesh with different resolutions. To achieve these goals, we introduce the neural spline model and the graph attention network (GAT) into our models respectively. While the Neural-Spline based model provides more flexibility for large mesh deformation, the GAT based model can handle domains with more complicated shapes and is better at performing delicate local deformation. We validate our methods on stationary and time-dependent, linear and non-linear equations, as well as regularly and irregularly shaped domains. Compared to the traditional Monge-Ampère method, our approach can greatly accelerate the mesh adaptation process by three to four orders of magnitude, whilst achieving comparable numerical error reduction.
Wenbin Song, Joseph G. Wallwork, Junpeng Gao, Zheng Tian 0002, Fanglei Sun, Matthew D. Piggott, Zuoqiang Shi, Jun Wang 0012
NeurIPS5
2020 Learning to Model Opponent Learning (Student Abstract)
abstract
Multi-Agent Reinforcement Learning (MARL) considers settings in which a set of coexisting agents interact with one another and their environment. The adaptation and learning of other agents induces non-stationarity in the environment dynamics. This poses a great challenge for value function-based algorithms whose convergence usually relies on the assumption of a stationary environment. Policy search algorithms also struggle in multi-agent settings as the partial observability resulting from an opponent's actions not being known introduces high variance to policy training. Modelling an agent's opponent(s) is often pursued as a means of resolving the issues arising from the coexistence of learning opponents. An opponent model provides an agent with some ability to reason about other agents to aid its own decision making. Most prior works learn an opponent model by assuming the opponent is employing a stationary policy or switching between a set of stationary policies. Such an approach can reduce the variance of training signals for policy search algorithms. However, in the multi-agent setting, agents have an incentive to continually adapt and learn. This means that the assumptions concerning opponent stationarity are unrealistic. In this work, we develop a novel approach to modelling an opponent's learning dynamics which we term Learning to Model Opponent Learning (LeMOL). We show our structured opponent model is more accurate and stable than naive behaviour cloning baselines. We further show that opponent modelling can improve the performance of algorithmic agents in multi-agent settings.
Ian Davies, Zheng Tian 0002, Jun Wang 0012
AAAI2
2020 Learning to Communicate Implicitly by Actions
abstract
In situations where explicit communication is limited, human collaborators act by learning to: (i) infer meaning behind their partner's actions, and (ii) convey private information about the state to their partner implicitly through actions. The first component of this learning process has been well-studied in multi-agent systems, whereas the second — which is equally crucial for successful collaboration — has not. To mimic both components mentioned above, thereby completing the learning process, we introduce a novel algorithm: Policy Belief Learning (PBL). PBL uses a belief module to model the other agent's private information and a policy module to form a distribution over actions informed by the belief module. Furthermore, to encourage communication by actions, we propose a novel auxiliary reward which incentivizes one agent to help its partner to make correct inferences about its private information. The auxiliary reward for communication is integrated into the learning of the policy module. We evaluate our approach on a set of environments including a matrix game, particle environment and the non-competitive bidding problem from contract bridge. We show empirically that this auxiliary reward is effective and easy to generalize. These results demonstrate that our PBL algorithm can produce strong pairs of agents in collaborative games where explicit communication is disabled.
Zheng Tian 0002, Shihao Zou, Ian Davies, Tim Warr, Lisheng Wu, Haitham Bou-Ammar, Jun Wang 0012
AAAI1
2019 A Regularized Opponent Model with Maximum Entropy Objective
abstract
In a single-agent setting, reinforcement learning (RL) tasks can be cast into an inference problem by introducing a binary random variable o, which stands for the "optimality". In this paper, we redefine the binary random variable o in multi-agent setting and formalize multi-agent reinforcement learning (MARL) as probabilistic inference. We derive a variational lower bound of the likelihood of achieving the optimality and name it as Regularized Opponent Model with Maximum Entropy Objective (ROMMEO). From ROMMEO, we present a novel perspective on opponent modeling and show how it can improve the performance of training agents theoretically and empirically in cooperative games. To optimize ROMMEO, we first introduce a tabular Q-iteration method ROMMEO-Q with proof of convergence. We extend the exact algorithm to complex environments by proposing an approximate version, ROMMEO-AC. We evaluate these two algorithms on the challenging iterated matrix game and differential game respectively and show that they can outperform strong MARL baselines.
Zheng Tian 0002, Ying Wen 0001, Zhichen Gong, Faiz Punakkath, Shihao Zou, Jun Wang 0012
IJCAI1
2017 Thinking Fast and Slow with Deep Learning and Tree Search
abstract
Sequential decision making problems, such as structured prediction, robotic control, and game playing, require a combination of planning policies and generalisation of those plans. In this paper, we present Expert Iteration (ExIt), a novel reinforcement learning algorithm which decomposes the problem into separate planning and generalisation tasks. Planning new policies is performed by tree search, while a deep neural network generalises those plans. Subsequently, tree search is improved by using the neural network policy to guide search, increasing the strength of new plans. In contrast, standard deep Reinforcement Learning algorithms rely on a neural network not only to generalise plans, but to discover them too. We show that ExIt outperforms REINFORCE for training a neural network to play the board game Hex, and our final tree search agent, trained tabula rasa, defeats MoHex1.0, the most recent Olympiad Champion player to be publicly released.
Thomas W. Anthony 0001, Zheng Tian 0002, David Barber
NIPS2