Peng Liu 0008

dblp:21/6121-8 · DBLP profile ↗
← Back
56ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0001-6568-1335ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-Tuning
abstract
Humanoid robots are promising to learn a diverse set of human-like locomotion behaviors, including standing up, walking, running, and jumping. However, existing methods predominantly require training independent policies for each skill, yielding behavior-specific controllers that exhibit limited generalization and brittle performance when deployed on irregular terrains and in diverse situations. To address this challenge, we propose Adaptive Humanoid Control (AHC) that adopts a two-stage framework to learn an adaptive humanoid locomotion controller across different skills and terrains. Specifically, we first train several primary locomotion policies and perform a multi-behavior distillation process to obtain a basic multi-behavior controller, facilitating adaptive behavior switching based on the environment. Then, we perform reinforced fine-tuning by collecting online feedback in performing adaptive behaviors on more diverse terrains, enhancing terrain adaptability for the adaptive behavior controller. We conduct experiments in both simulation and real-world experiments in Unitree G1 robots. The results show that our method exhibits strong adaptability across various situations and terrains.
Yingnan Zhao 0002, Xinmiao Wang, Dan Lu 0004, Qilong Han, Peng Liu 0008, Chenjia Bai
AAAI7
2026 A knowledge-guided multimodal network for video summarization
Xiaoyan Tian, Peng Liu 0008, Fenglei Ni
Expert Syst. Appl.4
2026 Improving multi-instance learning with hierarchical attention and frequency-domain hard sample distillation
Ting Xiao 0002, Minqian Sun, Yiqing Xia, Hai Yang 0002, Zhe Wang 0002, Peng Liu 0008
Knowl. Based Syst.6
2026 Curriculum reinforcement learning with measurable task representation learning
Yongyan Wen, Siyuan Li 0003, Mingjian Fu 0001, Yiqin Yang, Peng Liu 0008
Neural Networks6
2026 SDGScenes: User-intent driven indoor scene generation via semantic dependency graph
abstract
3D indoor scene generation aims to generate scenes that are physically plausible, consistent with common sense, and well-aligned with user intent. However, existing methods struggle to effectively capture user intent, as coarse-grained instruction methods yield plausible but intent-missing layouts, while fine-grained instruction methods reflect user intent but rely on manually defined relationships that burden users and compromise physical plausibility in complex scenes. To address this challenge, we propose SDGScenes, a novel indoor scene generation framework that automatically infers and synthesizes complete scenes from user intent and commonsense knowledge. Our approach firstly encodes scene requirements using Semantic Dependency Graph (SDG), a representation that captures both user intent module and commonsense module relationships. Sequentially, guided by the SDG, a Vision-Language Model (VLM) infers spatial constraints through commonsense reasoning. Finally, an optimization solver is applied to optimize object placement based on SDG-guided spatial constraints and physical constraints, including collision avoidance, boundary compliance, and reachability. Both quantitative and qualitative experimental results demonstrate that SDGScenes outperforms state-of-the-art methods in satisfying user intent.
Jicong Ao, Peng Liu 0008, Chenjia Bai
Pattern Recognit.5
2026 An Imitative Reinforcement Learning Framework for Pursuit-Lock-Launch Missions
abstract
Unmanned combat aerial vehicle (UCAV) within-visual-range (WVR) engagement, referring to a fight between two or more UCAVs at close quarters, plays a decisive role on the aerial battlefields. With the development of artificial intelligence, WVR engagement progressively advances toward intelligent and autonomous modes. However, autonomous WVR engagement policy learning is hindered by challenges such as weak exploration capabilities, low learning efficiency, and unrealistic simulated environments. To overcome these challenges, we propose a novel imitative reinforcement learning framework, which efficiently leverages expert data while enabling autonomous exploration. The proposed framework not only enhances learning efficiency through expert imitation but also ensures adaptability to dynamic environments via autonomous exploration with reinforcement learning. Therefore, the proposed framework can learn a successful policy of “pursuit-lock-launch” for UCAVs. To support data-driven learning, we establish an environment based on the Harfang3D sandbox. The extensive experimental results indicate that the proposed framework excels in this multistage task and significantly outperforms state-of-the-art reinforcement learning and imitation learning methods. Thanks to the ability of imitating experts and autonomous exploration, our framework can quickly learn the critical knowledge in complex aerial combat tasks, achieving up to a 100% success rate and demonstrating excellent robustness.
Siyuan Li 0003, Rongchang Zuo, Bofei Liu, Yaoyu He, Peng Liu 0008, Yingnan Zhao 0002
ACM Trans. Auton. Adapt. Syst.5
2026 Unsupervised Skill Discovery Through Skill Regions Differentiation
abstract
Unsupervised reinforcement learning (RL) aims to discover diverse behaviors that can accelerate the learning of downstream tasks. Previous methods typically focus on entropy-based exploration or empowerment-driven skill learning. However, entropy-based exploration struggles in large-scale state spaces (e.g., images), and empowerment-based methods with mutual information (MI) estimations have limitations in state exploration. To address these challenges, we propose a novel skill discovery objective that maximizes the deviation of the state density of one skill from the explored regions of other skills, encouraging inter-skill state diversity similar to the initial MI objective. For state-density estimation, we construct a novel conditional autoencoder with soft modularization for different skill policies in high-dimensional space. Meanwhile, to incentivize intra-skill exploration, we formulate an intrinsic reward based on the learned autoencoder that resembles count-based exploration in a compact latent space. Through extensive experiments in challenging state and image-based tasks, we find our method learns meaningful skills and achieves superior performance in various downstream tasks.
Ting Xiao 0002, Jiakun Zheng, Rushuai Yang, Qiaosheng Zhang 0002, Peng Liu 0008, Zhe Wang 0002, Chenjia Bai
IEEE Trans. Neural Networks Learn. Syst.6
2025 Safe Planner: Empowering Safety Awareness in Large Pre-Trained Models for Robot Task Planning
abstract
Robot task planning is an important problem for autonomous robots in long-horizon challenging tasks. As large pre-trained models have demonstrated superior planning ability, recent research investigates utilizing large models to achieve autonomous planning for robots in diverse tasks. However, since the large models are pre-trained with Internet data and lack the knowledge of real task scenes, large models as planners may make unsafe decisions that hurt the robots and the surrounding environments. To solve this challenge, we propose a novel Safe Planner framework, which empowers safety awareness in large pre-trained models to accomplish safe and executable planning. In this framework, we develop a safety prediction module to guide the high-level large model planner, and this safety module trained in a simulator can be effectively transferred to real-world tasks. The proposed Safe Planner framework is evaluated on both simulated environments and real robots. The experiment results demonstrate that Safe Planner not only achieves state-of-the-art task success rates, but also substantially improves safety during task execution.
Siyuan Li 0003, Lingfei Cui, Jiani Lu, Qinqin Xiao, Xirui Yang, Peng Liu 0008, Kewu Sun
AAAI7
2025 SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control Tasks
abstract
Deep reinforcement learning (DRL) has achieved remarkable success in various domains, yet its reliance on neural networks results in a lack of transparency, which limits its practical applications in safety-critical and human-agent interaction domains. Decision trees, known for their notable explainability, have emerged as a promising alternative to neural networks. However, decision trees often struggle in long-horizon continuous control tasks with high-dimensional observation space due to their limited expressiveness. To address this challenge, we propose SkillTree, a novel hierarchical framework that reduces the complex continuous action space of challenging control tasks into discrete skill space. By integrating the differentiable decision tree within the high-level policy, SkillTree generates discrete skill embeddings that guide low-level policy execution. Furthermore, through distillation, we obtain a simplified decision tree model that improves performance while further reducing complexity. Experiment results validate SkillTree’s effectiveness across various robotic manipulation tasks, providing clear skill-level insights into the decision-making process. The proposed approach not only achieves performance comparable to neural network based methods in complex long-horizon control tasks but also significantly enhances the transparency and explainability of the decision-making process.
Yongyan Wen, Siyuan Li 0003, Rongchang Zuo, Hangyu Mao, Peng Liu 0008
AAAI6
2025 Radiology Report Generation via Multi-objective Preference Optimization
abstract
Automatic Radiology Report Generation (RRG) is an important topic for alleviating the substantial workload of radiologists. Existing RRG approaches rely on supervised regression based on different architectures or additional knowledge injection, while the generated report may not align optimally with radiologists’ preferences. Especially, since the preferences of radiologists are inherently heterogeneous and multi-dimensional, e.g., some may prioritize report fluency, while others emphasize clinical accuracy. To address this problem, we propose a new RRG method via Multi-objective Preference Optimization (MPO) to align the pre-trained RRG model with multiple human preferences, which can be formulated by multi-dimensional reward functions and optimized by multi-objective reinforcement learning (RL). Specifically, we use a preference vector to represent the weight of preferences and use it as a condition for the RRG model. Then, a linearly weighed reward is obtained via a dot product between the preference vector and multi-dimensional reward. Next, the RRG model is optimized to align with the preference vector by optimizing such a reward via RL. In the training stage, we randomly sample diverse preference vectors from the preference space and align the model by optimizing the weighted multi-objective rewards, which leads to an optimal policy on the entire preference space. When inference, our model can generate reports aligned with specific preferences without further fine-tuning. Extensive experiments on two public datasets show the proposed method can generate reports that cater to different preferences in a single model and achieve state-of-the-art performance.
Ting Xiao 0002, Lei Shi 0004, Peng Liu 0008, Zhe Wang 0002, Chenjia Bai
AAAI3
2025 Learning Adaptive Spatial-temporal Structured Correlation Filters for UAV Object Tracking
abstract
In visual object tracking via unmanned aerial vehicle (UAV), discriminative correlation filtering (DCF) is one of the major methods owing to circulant samples which can be utilized not only for computing economically but also to hasten the optimization of filters. The universal DCF methods are seen as ridge regression models with some kinds of regularizations, which may likely result in tracking drift. Inspired by structured SVM, a generic framework that combines the squared hinge loss with spatial-temporal regularizations is advocated in this letter to distinguish the feature of targets from surrounding backgrounds and thus improve the robustness in tracking. Meanwhile, the standard squared norm penalty is turned into a group lasso penalty in the multichannel framework which enables filters with modest channel selection. The proposed adaptive spatial-temporal structured correlation filtering (ASTSCF) method has attained competitive results on major UAV tracking benchmarks.
Wei Zhao 0008, Peng Liu 0008, Xianglong Tang
ICASSP3
2025 ML-NC-TTT: Multi-layer Noise Contrastive Learning for Test-Time Training
Cangning Fan, Peng Liu 0008, Wei Zhao 0008, Qiquan Quan
PRCV (1)2
2025 Combining long and short spatiotemporal reasoning for deep reinforcement learning
Peng Liu 0008, Chenjia Bai
Neurocomputing2
2025 Auxiliary Reward Generation With Transition Distance Representation Learning
abstract
Reinforcement learning (RL) has shown strengths in challenging sequential decision-making problems. The reward function in RL is crucial to the learning performance, as it quantifies the degree of task completion. In real-world problems, the rewards are predominantly human-designed, which requires laborious tuning, and is susceptible to human cognitive biases. To achieve automatic auxiliary reward generation, we propose a novel representation learning approach that can measure the “transition distance” between states. Building upon these representations, we introduce an auxiliary reward generation technique for both single-task and skill-chaining scenarios without the need for human knowledge. Furthermore, we theoretically show that the proposed auxiliary rewards maintain the policy invariance property, i.e., the generated rewards will not hurt the policy optimality under the original rewards. In the experiment section, we evaluate the proposed approach in both online and offline learning settings in a wide range of tasks, including robot manipulation and locomotion. The experiment results demonstrate the effectiveness of measuring the transition distance and the induced improvement by auxiliary rewards, which promotes better learning efficiency and increases convergent stability. Beyond that, we demonstrate that the learned manipulation policy with the auxiliary rewards in a simulator can be transferred to the real robot, as shown inhttps://sites.google.com/view/transition-distance-rp/tdrp. Note to Practitioners—The motivation for this paper arises from the need for a technique that enhances robot skill-learning efficiency and performance in both single-task and skill-chaining scenarios. Our research primarily focuses on robot arm manipulation tasks. To accelerate the policy learning process and improve policy performance for executing these tasks, we introduce an auxiliary reward generation technique for both single-task and skill-chaining scenarios without requiring human expertise. This technique leverages the proposed novel representation learning approach, which can measure the “transition distance” between states. During each policy training round, the robot receives a dense reshaped reward created by our approach. Using the policy trained by our method, we successfully control a real Franka Panda robot arm to complete various manipulation tasks.
Siyuan Li 0003, Shijie Han, Yingnan Zhao 0002, Yiqin Yang, Qianchuan Zhao, Peng Liu 0008
IEEE Trans Autom. Sci. Eng.6
2024 Robust Visual Imitation Learning with Inverse Dynamics Representations
abstract
Imitation learning (IL) has achieved considerable success in solving complex sequential decision-making problems. However, current IL methods mainly assume that the environment for learning policies is the same as the environment for collecting expert datasets. Therefore, these methods may fail to work when there are slight differences between the learning and expert environments, especially for challenging problems with high-dimensional image observations. However, in real-world scenarios, it is rare to have the chance to collect expert trajectories precisely in the target learning environment. To address this challenge, we propose a novel robust imitation learning approach, where we develop an inverse dynamics state representation learning objective to align the expert environment and the learning environment. With the abstract state representation, we design an effective reward function, which thoroughly measures the similarity between behavior data and expert data not only element-wise, but also from the trajectory level. We conduct extensive experiments to evaluate the proposed approach under various visual perturbations and in diverse visual control tasks. Our approach can achieve a near-expert performance in most environments, and significantly outperforms the state-of-the-art visual IL methods and robust IL methods.
Siyuan Li 0003, Rongchang Zuo, Kewu Sun, Lingfei Cui, Jishiyu Ding, Peng Liu 0008
AAAI7
2024 IOB: integrating optimization transfer and behavior transfer for multi-policy reuse
Siyuan Li 0003, Hao Li 0069, Jin Zhang 0016, Zhen Wang 0004, Peng Liu 0008, Chongjie Zhang
Auton. Agents Multi Agent Syst.5
2024 Curriculum adaptation method based on graph neural networks for universal domain adaptation
Cangning Fan, Peng Liu 0008, Wei Zhao 0008
Expert Syst. Appl.2
2024 Monotonic Quantile Network for Worst-Case Offline Reinforcement Learning
abstract
A key challenge in offline reinforcement learning (RL) is how to ensure the learned offline policy is safe, especially in safety-critical domains. In this article, we focus on learning a distributional value function in offline RL and optimizing a worst-case criterion of returns. However, optimizing a distributional value function in offline RL can be hard, since the crossing quantile issue is serious, and the distribution shift problem needs to be addressed. To this end, we propose monotonic quantile network (MQN) with conservative quantile regression (CQR) for risk-averse policy learning. First, we propose an MQN to learn the distribution over returns with non-crossing guarantees of the quantiles. Then, we perform CQR by penalizing the quantile estimation for out-of-distribution (OOD) actions to address the distribution shift in offline RL. Finally, we learn a worst-case policy by optimizing the conditional value-at-risk (CVaR) of the distributional value function. Furthermore, we provide theoretical analysis of the fixed-point convergence in our method. We conduct experiments in both risk-neutral and risk-sensitive offline settings, and the results show that our method obtains safe and conservative behaviors in robotic locomotion tasks.
Chenjia Bai, Ting Xiao 0002, Zhoufan Zhu, Lingxiao Wang 0003, Animesh Garg, Bin He 0003, Peng Liu 0008, Zhaoran Wang 0001
IEEE Trans. Neural Networks Learn. Syst.8
2024 Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
abstract
Deep reinforcement learning (DRL) and deep multiagent reinforcement learning (MARL) have achieved significant success across a wide range of domains, including game artificial intelligence (AI), autonomous vehicles, and robotics. However, DRL and deep MARL agents are widely known to be sample inefficient that millions of interactions are usually needed even for relatively simple problem settings, thus preventing the wide application and deployment in real-industry scenarios. One bottleneck challenge behind is the well-known exploration problem, i.e., how efficiently exploring the environment and collecting informative experiences that could benefit policy learning toward the optimal ones. This problem becomes more challenging in complex environments with sparse rewards, noisy distractions, long horizons, and nonstationary co-learners. In this article, we conduct a comprehensive survey on existing exploration methods for both single-agent RL and multiagent RL. We start the survey by identifying several key challenges to efficient exploration. Then, we provide a systematic survey of existing approaches by classifying them into two major categories: uncertainty-oriented exploration and intrinsic motivation-oriented exploration. Beyond the above two main branches, we also include other notable exploration methods with different ideas and techniques. In addition to algorithmic analysis, we provide a comprehensive and unified empirical comparison of different exploration methods for DRL on a set of commonly used benchmarks. According to our algorithmic and empirical investigation, we finally summarize the open problems of exploration in DRL and deep MARL and point out a few future directions.
Jianye Hao, Tianpei Yang, Hongyao Tang, Chenjia Bai, Jinyi Liu 0002, Zhaopeng Meng, Peng Liu 0008, Zhen Wang 0004
IEEE Trans. Neural Networks Learn. Syst.7
2024 A coupling method of learning structured support correlation filters for visual tracking
Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
Vis. Comput.1
2024 Correction: A coupling method of learning structured support correlation filters for visual tracking
Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
Vis. Comput.1
2023 Behavior Contrastive Learning for Unsupervised Skill Discovery
abstract
In reinforcement learning, unsupervised skill discovery aims to learn diverse skills without extrinsic rewards. Previous methods discover skills by maximizing the mutual information (MI) between states and skills. However, such an MI objective tends to learn simple and static skills and may hinder exploration. In this paper, we propose a novel unsupervised skill discovery method through contrastive learning among behaviors, which makes the agent produce similar behaviors for the same skill and diverse behaviors for different skills. Under mild assumptions, our objective maximizes the MI between different behaviors based on the same skill, which serves as an upper bound of the previous MI objective. Meanwhile, our method implicitly increases the state entropy to obtain better state coverage. We evaluate our method on challenging mazes and continuous control tasks. The results show that our method generates diverse and far-reaching skills, and also obtains competitive performance in downstream tasks compared to the state-of-the-art methods.
Rushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li 0003, Bin Zhao 0001, Zhen Wang 0004, Peng Liu 0008, Xuelong Li 0001
ICML7
2023 Classifying ambiguous identities in hidden-role Stochastic games with multi-agent reinforcement learning
Shijie Han, Siyuan Li 0003, Bo An 0001, Wei Zhao 0008, Peng Liu 0008
Auton. Agents Multi Agent Syst.5
2023 Channel-Weighted Structured Correlation Filters for UAV Tracking
abstract
In the visual object tracking via unmanned aerial vehicle (UAV) the correlation filtering (CF) is one of the mainstream methods for the reason that optimizing filters by circulant samples facilitates the calculation. The prevalent discriminative CF (DCF) frameworks credited to a ridge regression model followed by various regularizations focus on regressing to a fixed Gaussian label, which maybe easily induce over-fitting. To these concerns, we integrate the hinge-squared loss (HSL) of structured SVM with the CF model so as to attain a dynamical label which regresses the difference between target and background samples and thus enhances the robustness in tracking. Moreover, we assume weighted channels to augment the discrimination of hinge-squared loss, thus indicating it has as good extensibility as the ridge regression in usual CF frameworks. The proposed method channel-weighted structured correlation filtering (CWSCF) has achieved excellent performance compared with other methods in mainstream UAV tracking benchmarks.
Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
IEEE Geosci. Remote. Sens. Lett.1
2023 Variational Diversity Maximization for Hierarchical Skill Discovery
Yingnan Zhao 0002, Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
Neural Process. Lett.2
2023 Addressing Hindsight Bias in Multigoal Reinforcement Learning
abstract
Multigoal reinforcement learning (RL) extends the typical RL with goal-conditional value functions and policies. One efficient multigoal RL algorithm is the hindsight experience replay (HER). By treating a hindsight goal from failed experiences as the original goal, HER enables the agent to receive rewards frequently. However, a key assumption of HER is that the hindsight goals do not change the likelihood of the sampled transitions and trajectories used in training, which is not the fact according to our analysis. More specifically, we show that using hindsight goals changes such a likelihood and results in a biased learning objective for multigoal RL. We analyze the hindsight bias due to this use of hindsight goals and propose the bias-corrected HER (BHER), an efficient algorithm that corrects the hindsight bias in training. We further show that BHER outperforms several state-of-the-art multigoal RL approaches in challenging robotics tasks.
Chenjia Bai, Lingxiao Wang 0003, Yixin Wang 0002, Zhaoran Wang 0001, Chenyao Bai, Peng Liu 0008
IEEE Trans. Cybern.7
2023 Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
abstract
Efficient exploration remains a challenging problem in reinforcement learning, especially for tasks where extrinsic rewards from environments are sparse or even totally disregarded. Significant advances based on intrinsic motivation show promising results in simple environments but often get stuck in environments with multimodal and stochastic dynamics. In this work, we propose a variational dynamic model based on the conditional variational inference to model the multimodality and stochasticity. We consider the environmental state-action transition as a conditional generative process by generating the next-state prediction under the condition of the current state, action, and latent variable, which provides a better understanding of the dynamics and leads to a better performance in exploration. We derive an upper bound of the negative log likelihood of the environmental transition and use such an upper bound as the intrinsic reward for exploration, which allows the agent to learn skills by self-supervised exploration without observing extrinsic rewards. We evaluate the proposed method on several image-based simulation tasks and a real robotic manipulating task. Our method outperforms several state-of-the-art environment model-based exploration approaches.
Chenjia Bai, Peng Liu 0008, Kaiyu Liu, Lingxiao Wang 0003, Yingnan Zhao 0002, Lei Han 0001, Zhaoran Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning
Chenjia Bai, Lingxiao Wang 0003, Zhuoran Yang, Zhihong Deng 0002, Animesh Garg, Peng Liu 0008, Zhaoran Wang 0001
ICLR6
2022 Anatomical prior based vertebra modelling for reappearance of human spines
Qinghua Huang, Cui Yang, Qifeng Deng, Peng Liu 0008, Maoqing Fu, Le Li 0003, Xuelong Li 0001
Neurocomputing6
2021 Principled Exploration via Optimistic Bootstrapping and Backward Induction
abstract
One principled approach for provably efficient exploration is incorporating the upper confidence bound (UCB) into the value function as a bonus. However, UCB is specified to deal with linear and tabular settings and is incompatible with Deep Reinforcement Learning (DRL). In this paper, we propose a principled exploration method for DRL through Optimistic Bootstrapping and Backward Induction (OB2I). OB2I constructs a general-purpose UCB-bonus through non-parametric bootstrap in DRL. The UCB-bonus estimates the epistemic uncertainty of state-action pairs for optimistic exploration. We build theoretical connections between the proposed UCB-bonus and the LSVI-UCB in linear setting. We propagate future uncertainty in a time-consistent manner through episodic backward update, which exploits the theoretical advantage and empirically improves the sample-efficiency. Our experiments in MNIST maze and Atari suit suggest that OB2I outperforms several state-of-the-art exploration approaches.
Chenjia Bai, Lingxiao Wang 0003, Lei Han 0001, Jianye Hao, Animesh Garg, Peng Liu 0008, Zhaoran Wang 0001
ICML6
2021 Dynamic Bottleneck for Robust Self-Supervised Exploration
abstract
Exploration methods based on pseudo-count of transitions or curiosity of dynamics have achieved promising results in solving reinforcement learning with sparse rewards. However, such methods are usually sensitive to environmental dynamics-irrelevant information, e.g., white-noise. To handle such dynamics-irrelevant information, we propose a Dynamic Bottleneck (DB) model, which attains a dynamics-relevant representation based on the information-bottleneck principle. Based on the DB model, we further propose DB-bonus, which encourages the agent to explore state-action pairs with high information gain. We establish theoretical connections between the proposed DB-bonus, the upper confidence bound (UCB) for linear case, and the visiting count for tabular case. We evaluate the proposed method on Atari suits with dynamics-irrelevant noises. Our experiments show that exploration with DB bonus outperforms several state-of-the-art exploration methods in noisy environments.
Chenjia Bai, Lingxiao Wang 0003, Lei Han 0001, Animesh Garg, Jianye Hao, Peng Liu 0008, Zhaoran Wang 0001
NeurIPS6
2020 Correlation-Guided Attention for Corner Detection Based Visual Tracking
abstract
Accurate bounding box estimation has recently attracted much attention in the tracking community because traditional multi-scale search strategies cannot estimate tight bounding boxes in many challenging scenarios involving changes to the target. A tracker capable of detecting target corners can flexibly adapt to such changes, but existing corner detection based tracking methods have not achieved adequate success. We analyze the reasons for their failure and propose a state-of-the-art tracker that performs correlation-guided attentional corner detection in two stages. First, a region of interest (RoI) is obtained by employing an efficient Siamese network to distinguish the target from the background. Second, a pixel-wise correlation-guided spatial attention module and a channel-wise correlation-guided channel attention module exploit the relationship between the target template and the RoI to highlight corner regions and enhance features of the RoI for corner detection. The correlation-guided attention modules improve the accuracy of corner detection, thus enabling accurate bounding box estimation. When trained on large-scale datasets using a novel RoI augmentation strategy, the performance of the proposed tracker, running at a high speed of 70 FPS, is comparable with that of state-of-the-art trackers in meeting five challenging performance benchmarks.
Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
CVPR2
2020 Importance-weighted conditional adversarial network for unsupervised domain adaptation
Peng Liu 0008, Ting Xiao 0002, Cangning Fan, Wei Zhao 0008, Xianglong Tang, Hongwei Liu 0002
Expert Syst. Appl.1
2020 Domain adaptation based on domain-invariant and class-distinguishable feature learning using multiple adversarial networks
Cangning Fan, Peng Liu 0008, Ting Xiao 0002, Wei Zhao 0008, Xianglong Tang
Neurocomputing2
2020 Generating attentive goals for prioritized hindsight reinforcement learning
Peng Liu 0008, Chenjia Bai, Yingnan Zhao 0002, Chenyao Bai, Wei Zhao 0008, Xianglong Tang
Knowl. Based Syst.1
2020 Obtaining accurate estimated action values in categorical distributional reinforcement learning
Yingnan Zhao 0002, Peng Liu 0008, Chenjia Bai, Wei Zhao 0008, Xianglong Tang
Knowl. Based Syst.2
2020 Joint Channel Reliability and Correlation Filters Learning for Visual Tracking
abstract
Multi-channel discriminative correlation filter (DCF) tracking methods have exhibited superior performance on several benchmarks. However, existing methods usually treat each channel of the features equally, whereas they pay less attention to the contribution of different channels. Different channels exhibit variant properties in the tracking process. A DCF learned with equally important channels is likely to be contaminated by the unreliable ones, which results in model degradation. To address this problem, we propose a new formulation for jointly learning the channel reliability and the correlation filters. The formulation is generic, and it can be combined with existing techniques in the DCF framework to further improve the performance. Our method can adaptively increase the impact of reliable channels and down-weight the corrupted ones. To solve the joint learning problem, we propose an optimization strategy that alternates between the correlation filters and the channel weights. Further, we prove the upper bound of the objective function and solve the channel weights efficiently. The joint learning strategy makes the correlation filters more discriminative and the channel weights more accurate. To verify the joint formulation, we propose a tracker based on the proposed formulation and the techniques used in the ECO tracker. We conduct extensive experiments to evaluate the proposed tracker on three benchmarks. The experimental results show that our formulation is effective and efficient, and that it performs favorably against other state-of-the-art trackers.
Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
IEEE Trans. Circuits Syst. Video Technol.2
2020 Visual Tracking by Structurally Optimizing Pre-Trained CNN
abstract
In this paper, we propose a novel channel pruning method for convolutional neural network (CNN)-based trackers. Pre-trained CNNs are widely used in visual tracking to obtain high-level representations of targets. However, most pre-trained CNNs are trained for other tasks (e.g., VGGNet is trained for image classification), and they require a considerable amount of time to generate features. First, we introduce a dimensionality reduction method considering the information amount and tracking errors to obtain good low-dimensionality features from the last convolutional layer for tracking. Then, a backward channel selection method is proposed to select representative channels layer by layer. In this process, we aim to minimize the target changes and maximize the loss of the background or other objects. Finally, we reconstruct the neural network weights to reduce the information loss of the target with one-shot learning. Experimental results on challenging benchmarks show that the proposed channel pruning method can enhance the tracking performance and reduce the computational requirements.
Chang Liu 0011, Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
IEEE Trans. Circuits Syst. Video Technol.2
2019 An exploratory rollout policy for imagination-augmented agents
Peng Liu 0008, Yingnan Zhao 0002, Wei Zhao 0008, Xianglong Tang, Zichan Yang
Appl. Intell.1
2019 Guided goal generation for hindsight multi-goal reinforcement learning
Chenjia Bai, Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
Neurocomputing2
2019 Structure preservation and distribution alignment in discriminative transfer subspace learning
Ting Xiao 0002, Peng Liu 0008, Wei Zhao 0008, Hongwei Liu 0002, Xianglong Tang
Neurocomputing2
2019 Multi-level context-adaptive correlation tracking
Peng Liu 0008, Chang Liu 0011, Wei Zhao 0008, Xianglong Tang
Pattern Recognit.1
2018 Spatial-temporal adaptive feature weighted correlation filter for visual tracking
Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
Signal Process. Image Commun.2
2018 Robust Tracking and Redetection: Collaboratively Modeling the Target and Its Context
abstract
Robust object tracking and redetection require stably predicting the trajectory of the target object and recovering from tracking failure by quickly redetecting it when it is lost during long-term tracking. The locations of the target and the background are calculated relative to the region occupied by the object. The effect of tracking can be enhanced by isolating the target and the background, modeling and tracking them, respectively, and integrating their tracking results. In this study, we propose an approach that builds motion models for the target and its context. Tracking results from a target tracker and a context tracker are integrated through linear fusion to predict the position of the target. A kernelized correlation filter tracker is used to track the target in the predicted position. When the target is lost, it can be quickly recovered by searching in the given field of view using a target model built and updated through observation models that are constructed prior to the loss of the target. Our approach is not sensitive to the segmentation of the target and the context. The motion models and observation models of the target and the context work together in the tracking process, whereas the target model alone is involved in redetection. Experiments to test our proposed approach, which simultaneously models the target and its context, showed that it can effectively enhance the robustness of long-term tracking.
Chang Liu 0011, Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
IEEE Trans. Multim.2
2017 Making the torch lighter: Areinforced active sampling framework for image classification
abstract
In this paper, we aim to construct a more reasonable and effective active sampling model, named as reinforcement uncertainty sampling with bag-of-visual-words (RUSB). Compared with traditional active sampling strategy based on uncertainty, both certainty metric and sample post-processing are introduced for better performance. The certainty metric is measured by the bag-of-visual-words (BoVW) classification model in order to entirely evaluate samples, and the post-processing module is driven by the Q-learning method to construct a compact and efficient training set for the BoVW module. The performance of BoVW is used to initialize and determine the status of the post-processing module during the process of iteration. Meanwhile, the weight of the measurement is associated with each iteration instead of being set manually. Experimental results on real world datasets show the effectiveness of the proposed framework.
Peng Liu 0008, Zhipeng Ye, Xianglong Tang, Wei Zhao 0008
ICIP1
2017 RGB-D Object Recognition Using the Knowledge Transferred from Relevant RGB Images
Depeng Gao, Rui Wu 0002, Jiafeng Liu, Qingcheng Huang, Xianglong Tang, Peng Liu 0008
ICONIP (6)6
2017 Extended Kernelized Correlation Tracking with Target Enhancement and Sample Selection
abstract
In this paper, we address the problem of fast motion and bound effect about the popular high-speed correlation filters-based trackers. Such trackers are facing with the contradiction between extended detection region and reduced precision. To improve the robustness against fast motion, we firstly propose a tracker with extended region. In addition, in order for adapting different region sizes and target sizes, we introduce the target enhancement strategy to increase the effect of the target in learning a discriminative regression. Furthermore, a novel sample selection mechanism is established to drop the error samples generated by the circular structure of correlation filters. Our approach enlarges the detection region, improves the tracking accuracy and preserves the significant kernel structure of the correlation filters. Moreover, extensive experimental results in a recent benchmark datasets show that our proposed method have a promising performance compared to the state-of-art methods.
Peng Liu 0008, Chang Liu 0011, Wei Zhao 0008, Xianglong Tang
ICTAI1
2017 Abnormal crowd motion detection using double sparse representation
Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
Neurocomputing1
2016 Practice makes perfect: An adaptive active learning framework for image classification
Zhipeng Ye, Peng Liu 0008, Jiafeng Liu, Xianglong Tang, Wei Zhao 0008
Neurocomputing2
2015 Knowledge as action: A cognitive framework for indoor scene classification
abstract
Indoor scene classification is an important topic in computer vision, which is challenging due to the variability of decoration. Human vision system, on the other hand, is marvelous in adaptively recognizing scene categories with excellent performance and can be used for reference. Although bio-inspired computer vision algorithms have proven their effectiveness in classification applications, nowadays few researches on indoor scene classification algorithms attempt to model human vision system, restricting further improvement of performance and making it difficult to achieve adaptive scene understanding. To deal with this problem, in this paper we attempt to model the human vision system and achieve scene classification according to the cognitive theory, by dividing the problem into low-level objection annotation and high-level knowledge inference respectively on a macro perspective. Inspired by the biotical perception principle, a novel cognitive hybrid motivation framework is proposed, including empirical based annotation and inference over knowledge base, which is a simple yet effective framework based on techniques of object detection and classification. For a given indoor scene, objects of indoor scene are first annotated, then knowledge base is utilized to infer the category, reducing the effect of variable background. Environmental context is also utilized to assist classification. The proposed framework are evaluated on popular indoor scene dataset, and its effectiveness is proved by experimental results.
Rui Wu 0002, Zhipeng Ye, Peng Liu 0008, Xianglong Tang, Wei Zhao 0008
ICIP3
2015 May the torcher light our way: A negative-accelerated active learning framework for image classification
abstract
Uncertainty sampling is one of the most widely used strategy for pool-based active learning, however, there exists the problem that selected images do not reflect the desired training distribution and need additional labeling cost. To deal with this problem, from aspects of image classification and visual perception, we improve the traditional entropy-based sampling strategy by introducing bag-of-visual-words classification method and negative-accelerated learning principle from Rescorla-Wagner perceptive model. Differs from previous researches that treated sampling and classifying process separately, under the unified negative-accelerated learning model, we combine the two processes as a uniform model, named as negative-accelerated uncertainty sampling strategy with BoVW (NUSB) by proposing a new evolving sample selection measure, which takes category distribution into consideration. Classifier is trained to provide category distribution for the sampling process, reducing additional cost of annotation. Also, transfer test is utilized to prevent over-fitting and further evaluate the performance of different sampling strategies. Experimental results on real world datasets show that our active sampling framework outperforms both baseline active sampling strategies and state-of-the-art active learning based image classification method.
Zhipeng Ye, Peng Liu 0008, Xianglong Tang, Wei Zhao 0008
ICIP2
2014 Forecasting Crowd State in Video by an Improved Lattice Boltzmann Model
Peng Liu 0008, Wei Zhao 0008, Xianglong Tang
ICONIP (3)2
2013 Removal of dynamic weather conditions based on variable time window
abstract
Dynamic weather conditions, which mainly include rain and snow, make prevailing algorithms for many applications of outdoor video analysis and computer vision lapse. To remove dynamic weather conditions, the authors propose a pixel‐wise framework combining a detection method with a removal approach. Dynamic weather conditions are detected by a strategy‐driven state transition, which integrates static initialisation using K ‐means clustering with dynamic maintenance of Gaussian mixture model. Moreover, a variable time window is presented for removal of rain and snow. Each component of the framework is addressed using detailed descriptions of corresponding algorithms. Experiments demonstrate the effectiveness of the method on detection and removal of dynamic weather conditions.
Xudong Zhao 0001, Peng Liu 0008, Jiafeng Liu, Xianglong Tang
IET Comput. Vis.2
2011 Adaptive background estimation of outdoor illumination variations for foreground detection
abstract
A background estimation system, which integrates pixel-level features with a region-level one and combines short-term and long-term analysis of videos in outdoor illumination variations, is proposed for accurate foreground detection. Firstly, we discuss autocorrelation-based features for identification of the presence of foreground and outdoor illumination variations in short-term sequences, and propose an adaptive threshold learning approach insensitive to inner-pixel fast illumination variation based on histograms of intensity differences between successive frames. Then, we employ a pixel-wise rapid autoregressive model against gradual illumination change for background estimation in long-term sequence. Finally, we devise a texture measure to eliminate the regional effect of fast illumination variation. The effectiveness of our system is demonstrated using experiments on foreground detection in videos with various illumination changes.
Xudong Zhao 0001, Peng Liu 0008, Jiafeng Liu, Xianglong Tang
VCIP2
2011 A time, space and color-based classification of different weather conditions
abstract
Evaluation of different weather conditions provides a first step support for many different applications of outdoor video analysis and computer vision. In this paper, a simple but effective classification method on visual effects of different weather conditions is proposed. Due to the complex manifestations of weather conditions, we firstly provide a two-stage classification scheme. Then, we extract spatio-temporal and chromatic features to represent different weather situations. Using these features, we develop a classifier based on an experiential decision binary tree associated with C-SVM. The experimental results of classification on our newly-built video dataset indicate the effectiveness of our method.
Xudong Zhao 0001, Peng Liu 0008, Jiafeng Liu, Xianglong Tang
VCIP2
2006 An MLP-orthogonal Gaussian mixture model hybrid model for Chinese bank check printed numeral recognition
Hui Zhu 0013, Xianglong Tang, Peng Liu 0008
Int. J. Document Anal. Recognit.3