Tianying Ji

dblp:124/2199 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
13since 2021 · last 2025
0009-0006-8949-5880ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scrutinize What We Ignore: Reining In Task Representation Shift Of Context-Based Offline Meta Reinforcement Learning
abstract
Offline meta reinforcement learning (OMRL) has emerged as a promising approach for interaction avoidance and strong generalization performance by leveraging pre-collected data and meta-learning techniques. Previous context-based approaches predominantly rely on the intuition that alternating optimization between the context encoder and the policy can lead to performance improvements, as long as the context encoder follows the principle of maximizing the mutual information between the task variable $M$ and its latent representation $Z$ ($I(Z;M)$) while the policy adopts the standard offline reinforcement learning (RL) algorithms conditioning on the learned task representation. Despite promising results, the theoretical justification of performance improvements for such intuition remains underexplored. Inspired by the return discrepancy scheme in the model-based RL field, we find that the previous optimization framework can be linked with the general RL objective of maximizing the expected return, thereby explaining performance improvements. Furthermore, after scrutinizing this optimization framework, we observe that the condition for monotonic performance improvements does not consider the variation of the task representation. When these variations are considered, the previously established condition may no longer be sufficient to ensure monotonicity, thereby impairing the optimization process. We name this issue \underline{task representation shift} and theoretically prove that the monotonic performance improvements can be guaranteed with appropriate context encoder updates. We use different settings to rein in the task representation shift on three widely adopted training objectives concerning maximizing $I(Z;M)$ across different data qualities. Empirical results show that reining in the task representation shift can indeed improve performance. Our work opens up a new avenue for OMRL, leading to a better understanding between the task representation and performance improvements.
Tianying Ji, Jinhang Liu, Anqi Guo, Junqiao Zhao, Lanqing Li
ICLR3
2025 H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps
abstract
Solving real-world complex tasks using reinforcement learning (RL) without high-fidelity simulation environments or large amounts of offline data can be quite challenging. Online RL agents trained in imperfect simulation environments can suffer from severe sim-to-real issues. Offline RL approaches although bypass the need for simulators, often pose demanding requirements on the size and quality of the offline datasets. The recently emerged hybrid offline-and-online RL provides an attractive framework that enables joint use of limited offline data and imperfect simulator for transferable policy learning. In this paper, we develop a new algorithm, called$\mathrm{H} 2 \mathrm{O}+$, which offers great flexibility to bridge various choices of offline and online learning methods, while also accounting for dynamics gaps between the real and simulation environments. Through extensive simulation and real-world robotics experiments, we demonstrate superior performance and flexibility of$\mathbf{H 2 O}+$over advanced cross-domain online and offline RL algorithms.
Tianying Ji, Bingqi Liu, Haocheng Zhao, Jianying Zheng, Guyue Zhou, Jianming Hu, Xianyuan Zhan
ICRA2
2025 Goal-Conditioned Hierarchical Reinforcement Learning With High-Level Model Approximation
abstract
Hierarchical reinforcement learning (HRL) exhibits remarkable potential in addressing large-scale and long-horizon complex tasks. However, a fundamental challenge, which arises from the inherently entangled nature of hierarchical policies, has not been understood well, consequently compromising the training stability and exploration efficiency of HRL. In this article, we propose a novel HRL algorithm, high-level model approximation (HLMA), presenting both theoretical foundations and practical implementations. In HLMA, a Planner constructs an innovative high-level dynamic model to predict the -step transition of the Controller in a subtask. This allows for the estimation of the evolving performance of the Controller. At low level, we leverage the initial state of each subtask, transforming absolute states into relative deviations by a designed operator as Controller input. This approach facilitates the reuse of subtask domain knowledge, enhancing data efficiency. With this designed structure, we establish the local convergence of each component within HLMA and subsequently derive regret bounds to ensure global convergence. Abundant experiments conducted on complex locomotion and navigation tasks demonstrate that HLMA surpasses other state-of-the-art single-level RL and HRL algorithms in terms of sample efficiency and asymptotic performance. In addition, thorough ablation studies validate the effectiveness of each component of HLMA.
Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Huaping Liu 0001, Jianwei Zhang 0001, Mingxuan Jing, Wenbing Huang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization
abstract
Visual reinforcement learning (RL) has shown promise in continuous control tasks. Despite its progress, current algorithms are still unsatisfactory in virtually every aspect of the performance such as sample efficiency, asymptotic performance, and their robustness to the choice of random seeds. In this paper, we identify a major shortcoming in existing visual RL methods that is the agents often exhibit sustained inactivity during early training, thereby limiting their ability to explore effectively. Expanding upon this crucial observation, we additionally unveil a significant correlation between the agents' inclination towards motorically inactive exploration and the absence of neuronal activity within their policy networks. To quantify this inactivity, we adopt dormant ratio as a metric to measure inactivity in the RL agent's network. Empirically, we also recognize that the dormant ratio can act as a standalone indicator of an agent's activity level, regardless of the received reward signals. Leveraging the aforementioned insights, we introduce DrM, a method that uses three core mechanisms to guide agents' exploration-exploitation trade-offs by actively minimizing the dormant ratio. Experiments demonstrate that DrM achieves significant improvements in sample efficiency and asymptotic performance with no broken seeds (76 seeds in total) across three continuous control benchmark environments, including DeepMind Control Suite, MetaWorld, and Adroit. Most importantly, DrM is the first model-free algorithm that consistently solves tasks in both the Dog and Manipulator domains from the DeepMind Control Suite as well as three dexterous hand manipulation tasks without demonstrations in Adroit, all based on pixel observations.
Guowei Xu 0001, Ruijie Zheng, Yongyuan Liang, Zhecheng Yuan, Tianying Ji, Yu Luo 0021, Xiaoyu Liu 0003, Pu Hua, Shuzhen Li, Yanjie Ze, Hal Daumé III, Furong Huang, Huazhe Xu
ICLR6
2024 ACE: Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
abstract
The varying significance of distinct primitive behaviors during the policy learning process has been overlooked by prior model-free RL algorithms. Leveraging this insight, we explore the causal relationship between different action dimensions and rewards to evaluate the significance of various primitive behaviors during training. We introduce a causality-aware entropy term that effectively identifies and prioritizes actions with high potential impacts for efficient exploration. Furthermore, to prevent excessive focus on specific primitive behaviors, we analyze the gradient dormancy phenomenon and introduce a dormancy-guided reset mechanism to further enhance the efficacy of our method. Our proposed algorithm, ACE: Off-policy Actor-critic with Causality-aware Entropy regularization, demonstrates a substantial performance advantage across 29 diverse continuous control tasks spanning 7 domains compared to model-free RL baselines, which underscores the effectiveness, versatility, and efficient sample efficiency of our approach. Benchmark results and videos are available at https://ace-rl.github.io/.
Tianying Ji, Yongyuan Liang, Yan Zeng 0002, Yu Luo 0021, Guowei Xu 0001, Ruijie Zheng, Furong Huang, Fuchun Sun 0001, Huazhe Xu
ICML1
2024 Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-Critic
abstract
Learning high-quality $Q$-value functions plays a key role in the success of many modern off-policy deep reinforcement learning (RL) algorithms. Previous works primarily focus on addressing the value overestimation issue, an outcome of adopting function approximators and off-policy learning. Deviating from the common viewpoint, we observe that $Q$-values are often underestimated in the latter stage of the RL training process, potentially hindering policy learning and reducing sample efficiency. We find that such a long-neglected phenomenon is often related to the use of inferior actions from the current policy in Bellman updates as compared to the more optimal action samples in the replay buffer. We propose the Blended Exploitation and Exploration (BEE) operator, a simple yet effective approach that updates $Q$-value using both historical best-performing actions and the current policy. Based on BEE, the resulting practical algorithm BAC outperforms state-of-the-art methods in over 50 continuous control tasks and achieves strong performance in failure-prone scenarios and real-world robot tasks. Benchmark results and videos are available at https://jity16.github.io/BEE/.
Tianying Ji, Yu Luo 0021, Fuchun Sun 0001, Xianyuan Zhan, Jianwei Zhang 0001, Huazhe Xu
ICML1
2024 OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
abstract
Training reinforcement learning policies using environment interaction data collected from varying policies or dynamics presents a fundamental challenge. Existing works often overlook the distribution discrepancies induced by policy or dynamics shifts, or rely on specialized algorithms with task priors, thus often resulting in suboptimal policy performances and high learning variances. In this paper, we identify a unified strategy for online RL policy learning under diverse settings of policy and dynamics shifts: transition occupancy matching. In light of this, we introduce a surrogate policy learning objective by considering the transition occupancy discrepancies and then cast it into a tractable min-max optimization problem through dual reformulation. Our method, dubbed Occupancy-Matching Policy Optimization (OMPO), features a specialized actor-critic structure equipped with a distribution discriminator and a small-size local buffer. We conduct extensive experiments based on the OpenAI Gym, Meta-World, and Panda Robots environments, encompassing policy shifts under stationary and non-stationary dynamics, as well as domain adaption. The results demonstrate that OMPO outperforms the specialized baselines from different categories in all settings. We also find that OMPO exhibits particularly strong performance when combined with domain randomization, highlighting its potential in RL-based robotics applications.
Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Jianwei Zhang 0001, Huazhe Xu, Xianyuan Zhan
ICML2
2024 Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL
abstract
Off-policy reinforcement learning (RL) has achieved notable success in tackling many complex real-world tasks, by leveraging previously collected data for policy learning. However, most existing off-policy RL algorithms fail to maximally exploit the information in the replay buffer, limiting sample efficiency and policy performance. In this work, we discover that concurrently training an offline RL policy based on the shared online replay buffer can sometimes outperform the original online learning policy, though the occurrence of such performance gains remains uncertain. This motivates a new possibility of harnessing the emergent outperforming offline optimal policy to improve online policy learning. Based on this insight, we present Offline-Boosted Actor-Critic (OBAC), a model-free online RL framework that elegantly identifies the outperforming offline policy through value comparison, and uses it as an adaptive constraint to guarantee stronger policy learning performance. Our experiments demonstrate that OBAC outperforms other popular model-free RL baselines and rivals advanced model-based RL methods in terms of sample efficiency and asymptotic performance across **53** tasks spanning **6** task suites.
Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Jianwei Zhang 0001, Huazhe Xu, Xianyuan Zhan
ICML2
2024 Smooth Computation without Input Delay: Robust Tube-Based Model Predictive Control for Robot Manipulator Planning
abstract
Model Predictive Control (MPC) has exhibited remarkable capabilities in optimizing objectives and meeting constraints. However, the substantial computational burden associated with solving the Optimal Control Problem (OCP) at each triggering instant introduces significant delays between state sampling and control application. These delays limit the practicality of MPC in resource-constrained systems when engaging in complex tasks. The intuition to address this issue in this paper is that by predicting the successor state, the controller can solve the OCP one time step ahead of time thus avoiding the delay of the next action. To this end, we compute deviations between real and nominal system states, predicting forthcoming real states as initial conditions for the imminent OCP solution. Anticipatory computation stores optimal control based on current nominal states, thus mitigating the delay effects. Additionally, we establish an upper bound for linearization error, effectively linearizing the nonlinear system, reducing OCP complexity, and enhancing response speed. We provide empirical validation through two numerical simulations and corresponding real-world robot tasks, demonstrating significant performance improvements and augmented response speed (up to 90%) resulting from the seamless integration of our proposed approach compared to conventional time-triggered MPC strategies.
Yu Luo 0021, Qie Sima, Tianying Ji, Fuchun Sun 0001, Huaping Liu 0001, Jianwei Zhang 0001
ICRA3
2024 Robust tube-based MPC with smooth computation for dexterous robot manipulation
Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Qie Sima, Huaping Liu 0001, Mingxuan Jing, Jianwei Zhang 0001
Sci. China Inf. Sci.2
2022 Efficient Feature Compression for the Object Tracking Task
abstract
In object tracking systems, often clients capture video, encode it and transmit it to a server that performs the actual machine task. In this paper we propose an alternative architecture, where we instead transmit features to the server. Specifically, we partition the Joint Detection and Embedding (JDE) person tracking network into client and server side sub-networks and code the intermediate tensors i.e. features. The features are compressed for transmission using a Deep Neural Network (DNN) we design and train specifically for carrying out the tracking task. The DNN uses trainable non-uniform quantizers, conditional probability estimators, hierarchical coding; concepts that have been used in the past for neural networks based image and video compression. Additionally, the DNN includes a novel parameterized dual-path layer that comprises of an autoencoder in one path and a convolution layer in the other. The tensor output by each path is added before being consumed by subsequent layers. The parameter value for this dual-path layer controls the output channel count and correspondingly the bitrate of transmitted bitstream. We demonstrate that our model improves coding efficiency by 43.67% over state-of-the-art Versatile Video Coding standard that codes the source video in pixel domain.
Robert Henzel, Kiran M. Misra, Tianying Ji
ICIP3
2022 Video Feature Compression for Machine Tasks
abstract
We consider the problem of transmitting video from a remote device to a cloud-based classification system in a bandwidth limited network. Our focus is on developing an end-to-end system that extracts features from the video data and compresses these features for transmission. In this paper, we consider approaches that operate on each video frame independently as well as exploiting the temporal correlation between frames. In both cases, the transmitted features can be used for object detection and instance segmentation tasks using existing, pre-trained networks. Results show the efficacy of the approach with improvements in coding efficiency ranging from 46.3% to 92.8% when compared to compressing the video data using state-of-the-art video compression standards.
Kiran M. Misra, Tianying Ji, C. Andrew Segall, Frank Bossen
ICME2
2022 When to Update Your Model: Constrained Model-based Reinforcement Learning
abstract
Designing and analyzing model-based RL (MBRL) algorithms with guaranteed monotonic improvement has been challenging, mainly due to the interdependence between policy optimization and model learning. Existing discrepancy bounds generally ignore the impacts of model shifts, and their corresponding algorithms are prone to degrade performance by drastic model updating. In this work, we first propose a novel and general theoretical scheme for a non-decreasing performance guarantee of MBRL. Our follow-up derived bounds reveal the relationship between model shifts and performance improvement. These discoveries encourage us to formulate a constrained lower-bound optimization problem to permit the monotonicity of MBRL. A further example demonstrates that learning models from a dynamically-varying number of explorations benefit the eventual returns. Motivated by these analyses, we design a simple but effective algorithm CMLO (Constrained Model-shift Lower-bound Optimization), by introducing an event-triggered mechanism that flexibly determines when to update the model. Experiments show that CMLO surpasses other state-of-the-art methods and produces a boost when various policy optimization methods are employed.
Tianying Ji, Yu Luo 0021, Fuchun Sun 0001, Mingxuan Jing, Fengxiang He, Wenbing Huang 0001
NeurIPS1
2019 BDAC: A Behavior-aware Dynamic Adaptive Configuration on DHCP in Wireless LANs
abstract
DHCP is widely used to dynamically allocate IP addresses to the devices on local area networks, but the explosive increases of WiFi devices and their frequent mobility pose great challenges on DHCP performance in wireless LANs. In this paper, by analyzing large scale real network traces, we observe that the dynamic WiFi user behavior (e.g., online time pattern and spatio-temporal mobility pattern) leads to the poor DHCP performance. The IP pools in some VLANs have been exhausted in rush hours although the total IP utilization in WLAN is only 24%. Therefore, we have to configure IP lease times and IP pools dynamically and make sure that they are adaptive to the WiFi user behavior. In order to achieve this goal, we characterize and model the user behavior across online time pattern and spatiotemporal mobility pattern. Then we propose BDAC, a behaviour-aware dynamic adaptive configuration, which is combined of two strategies: adaptive IP lease time configuration and dynamic IP pool configuration. The former is to set adaptive lease times across user roles and area types based on online time pattern to reclaim IP addresses in time and reduce the peak IP usage, while the latter dynamically migrates the IP addresses across VLANs based on spatio-temporal mobility correlation to save the IP addresses. Using the real network traces of a different week, we conduct experiments to evaluate the performance of BDAC. Results show that BDAC can save up to 60% of IP addresses and the actual IP utilization rises from 24% to 59%. Furthermore, BDAC maintains high IP utilization when the number of VLANs in a WLAN increases.
Congcong Miao, Jilong Wang 0001, Tianying Ji, Hui Wang 0011, Chao Xu 0015, Fengyuan Ren
ICNP3
2013 Transform coefficient coding design for AVS2 video coding standard
abstract
AVS2 is a next-generation audio and video coding standard currently under development by the Audio Video Coding Standard Workgroup of China. In this paper, a coefficient-group based transform coefficient coding design for AVS2 video coding standard is presented, which includes two main coding tools, namely, two-level coefficient coding and intra-mode based context design. The two-level coefficient coding scheme allows accurate coefficient position information to be used in the context model design and improves the coding efficiency. It also helps increase the entropy coding throughput and facilitate parallel implementation. The intra-mode based context design further improves coding performance by utilizing the intra-prediction mode information in the context model. The two coding tools combined provide consistent rate-distortion performance gains under standard test conditions. Both tools were adopted into the AVS2 working draft. Furthermore, an improved rate-distortion optimized quantization algorithm is designed based on the proposed scheme, which significantly reduces the encoder complexity.
Tianying Ji, Dake He
VCIP3
2012 Transform Coefficient Coding in HEVC
abstract
This paper describes transform coefficient coding in the draft international standard of High Efficiency Video Coding (HEVC) specification and the driving motivations behind its design. Transform coefficient coding in HEVC encompasses the scanning patterns and coding methods for the last significant coefficient, significance map, coefficient levels, and sign data. Special attention is paid to the new methods of last significant coefficient coding, multilevel significance maps, high-throughput binarization, and sign data hiding. Experimental results are provided to evaluate the performance of transform coefficient coding in HEVC.
Joel Sole, Rajan L. Joshi, Tianying Ji, Marta Karczewicz, Gordon Clare, Félix Henry, Alberto Duenas
IEEE Trans. Circuits Syst. Video Technol.4