Peixi Peng

dblp:119/8511 · DBLP profile ↗
← Back
63ranked-venue papers
10as first author
44since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 6 first-author · 27 since 2021Artificial intelligence and machine learning · 38 · 8 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BulletTime4D: Towards High Spatio-Temporal Resolution Dynamic Scene Rendering via Spike-Guided Stereo Vision
abstract
High spatio‑temporal resolution novel‑view scene rendering is crucial for applications such as sports analysis and scientific experiments. However, existing Dynamic Scene Rendering (DSR) approaches typically rely on conventional RGB cameras with limited frame rates, making it difficult to achieve high spatio‑temporal resolution. In this paper, we present BulletTime4D, a high spatio‑temporal resolution DSR framework, which is the first trial to integrate a spike camera with binocular RGB cameras for dynamic scene reconstruction. Specifically, we first develop a hybrid camera prototype and build a real‑world dynamic scene reconstruction dataset. Then, BulletTime4D presents a multi‑timescale deformation representation by combining low‑frequency spatio‑temporal features with high‑frequency inter‑frame motion features. Finally, a rendering network is designed capable of projecting 4D Gaussians into the spike domain for spike rendering, and a cross‑domain supervision strategy is proposed to achieve high‑frame‑rate texture and color rendering. The results show that BulletTime4D outperforms state‑of‑the‑art methods on both simulated and real‑world datasets. In addition, BulletTime4D can synthesize 300 FPS novel‑view renderings using stereo RGB cameras at 30 FPS and a single spike camera.
Yiqian Chang, Haoran Xu 0004, Qinghong Ye, Jianing Li 0001, Xuan Wang 0002, Wei Zhang 0161, Peixi Peng
AAAI7
2026 Perceiving the Knowledge Boundary: Uncertainty-Guided Exploration and Imagination for World Models
abstract
World-model-based reinforcement learning achieves high sample efficiency by learning from imagined rollouts. However, its success critically depends on the accuracy of the learned world model, which is prone to producing unrealistic or hallucinated rollouts when queried beyond its domain of competence. These flawed predictions can trap the agent in a vicious cycle: by misleading exploration toward implausible or uninformative regions, they degrade the quality of collected data, which in turn corrupts policy learning with inaccurate rollouts. To break this cycle, we introduce the notion of a knowledge boundary—the region within which the world model provides reliable predictions—and propose a unified framework that both identifies and leverages this boundary. Concretely, we approximate the boundary using model uncertainty, quantified via disagreement across an ensemble of lightweight predictors, which serves as a practical proxy. This uncertainty signal is used in two complementary ways: as an intrinsic reward to guide exploration toward under-explored yet learnable regions, and as a dynamic filter to exclude unreliable imagined rollouts from policy optimization. Extensive experiments across diverse benchmarks—including CARLA, DeepMind Control Suite, Atari, and MemoryMaze—demonstrate that our approach consistently outperforms prior state-of-the-art methods.
Zhenxian Liu, Peixi Peng, Yangru Huang, Yonghong Tian 0001
AAAI2
2026 LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
abstract
The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings. To fill this gap, this paper proposes the first omnimodal benchmark for interactive livestream videos, LiViBench. It features a diverse set of 24 tasks, highlighting the perceptual, reasoning, and livestream-specific challenges. To efficiently construct the dataset, we design a standardized semi-automatic annotation workflow that incorporates the human-in-the-loop at multiple stages. The workflow leverages multiple MLLMs to form a multi-agent system for comprehensive video description and uses a seed-question-driven method to construct high-quality annotations. All interactive videos in the benchmark include audio, speech, and real-time comments modalities. To enhance models' understanding of interactive videos, we design tailored two-stage instruction-tuning and propose a Video-to-Comment Retrieval (VCR) module to improve the model's ability to utilize real-time comments. Based on these advancements, we develop LiVi-LLM-7B, an MLLM with enhanced knowledge of interactive livestreams. Experiments show that our model outperforms larger open-source models with up to 72B parameters, narrows the gap with leading proprietary models on LiViBench, and achieves enhanced performance on general video benchmarks, including VideoMME, LongVideoBench, MLVU, and VideoEval-Pro.
Langling Huang, Zhirong Wu, Xuhong Xia, Peixi Peng
AAAI7
2026 Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World Model
abstract
Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error accumulations when predicting the long-term future, which limits practical applications. Recent studies utilize the Diffusion Transformer (DiT) as the backbone of driving world models to improve learning flexibility. However, these models are always trained on short video clips, and multiple roll-out generations struggle to produce consistent and reasonable long videos due to the training-inference gap. To this end, we propose several solutions to build a simple yet effective long-term driving world model. First, we hierarchically decouple world model learning into large motion learning and bidirectional continuous motion learning. Then, considering the continuity of driving scenes, we propose a simple distillation method where fine-grained video flows are self-supervised signals for coarse-grained flows. The distillation is designed to improve the coherence of infinite video generation. The coarse-grained and fine-grained modules are coordinated to generate long-term and temporally coherent videos. On NuScenes, compared with the state-of-the-art front-view models, our model improves FVD by 27% and reduces inference time by 85% for the video task of generating 110+ frames.
Zhirong Wu, Peixi Peng
AAAI3
2026 COVR: Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
abstract
Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language models (VLMs) can assist RL, they often focus on knowledge distillation from the VLM to RL, overlooking the potential of RL-generated interaction data to enhance the VLM. To address this, we propose COVR, a collaborative optimization framework that enables the mutual enhancement of the VLM and RL policies. Specifically, COVR fine-tunes the VLM with RL-generated data to enhance the semantic reasoning ability consistent with the target task, and uses the enhanced VLM to further guide policy learning via action priors. To improve fine-tuning efficiency, we introduce two key modules: (1) an Exploration-Driven Dynamic Filter module that preserves valuable exploration samples using adaptive thresholds based on the degree of exploration, and (2) a Return-Aware Adaptive Loss Weight module that improves the stability of training by quantifying the inconsistency of sampling actions via return signals of RL. We further design a progressive fine-tuning strategy to reduce resource consumption. Extensive experiments show that COVR achieves strong performance across various challenging visual control tasks.
Canming Xia, Peixi Peng, Guang Tan, Haoran Xu 0004, Zhenxian Liu, Luntong Li
AAAI2
2026 SAH-NeRF: Enhancing NeRF on Novel View Synthesis With an SNN-ANN Hybrid Framework
abstract
Neural Radiance Field (NeRF) utilizes Artificial Neural Networks (ANNs) to map 3D points and their corresponding 2D viewing directions to colors and densities. However, this approach encounters several challenges due to the inherent characteristics of ANNs. Firstly, ANNs tend to extract smooth features across sampled points, which makes it difficult to accurately represent the non-smooth variations between object surfaces and the surrounding air. Secondly, ANNs compute densities and colors independently for each point, failing to account for the sequential dependencies among points along the same ray. To tackle these issues, Spiking Neural Networks (SNNs) are introduced, which are better at processing sequential information and non-smooth representations. In this work, we propose a novel hybrid NeRF framework called SAH-NeRF that combines ANNs with SNNs. By harnessing ANNs’ robust representation capabilities alongside SNNs’ strengths in handling non-smooth distributions, our method could significantly improve the performance of three ANN-based NeRFs, surpassing state-of-the-art methods including 3D Gaussian Splatting. Notably, our SAH-NeRF could meanwhile enhance novel view synthesis and reduce energy consumption.
Yiqian Chang, Peixi Peng, Zhaokun Zhou, Xuan Wang 0002, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Parallel Multi-Attention and Gated Fusion for Visual Question Localized Answering in Surgical Scenes
abstract
Surgical Visual Question Localized Answering (Surgical-VQLA) is an emerging task that supports surgical education by generating accurate answers and localizing relevant anatomical regions based on visual content and textual queries. This task requires precise spatial reasoning and tight semantic alignment across modalities, which remain challenging for current models due to limited spatial sensitivity and insufficient semantic integration. Mitigating these limitations, we propose EndoVisLoc, a dedicated framework that enhances visual-textual interaction through structured attention and gated fusion. Specifically, we design a Parallel Multi Attention Module (PMAM) to capture different visual features, improving the perception of anatomical structures. We further develop a Dynamic Gated Fusion Module (DGFM) to adaptively inject semantic priors into visual features via gated control, facilitating robust cross-modal fusion. Finally, we introduce a Hierarchical Classifier Head (HCH) to refine the fused representations and jointly optimize answer prediction and spatial localization. Extensive experiments on the EndoVis-18-VQLA and EndoVis-17-VQLA datasets demonstrate the superior performance of EndoVisLoc, surpassing the state-of-the-art OTAS model by +5.72% ACC, +5.73% F-score, and +1.82% mIoU on EndoVis-18-VQLA and by +1.59% ACC, +2.48% F-score, and +0.33% mIoU on EndoVis-17-VQLA. These results confirm the consistent advantage of EndoVisLoc in both answer accuracy and precise anatomical localization.
Peixi Peng, Wanshu Fan, Zhongbin Han
IEEE J. Biomed. Health Informatics3
2026 SERF: Spatiotemporal-Aware Event-RGB Fusion for Steering Angle Prediction
abstract
Existing end-to-end methods for steering angle prediction (SAP) primarily rely on RGB imagery from conventional cameras as input; however, they suffer from limitations such as poor performance in low-light conditions and motion blur. Recently, event cameras have garnered attention as complementary to RGB imagery, providing advantages such as high dynamic range and low latency. Nevertheless, earlier SAP methods that integrate event and RGB data may not fully exploit the spatio-temporal characteristics of events, resulting in performance degradation in low-light scenarios affected by noise interference. To address this limitation, we present a novel spatiotemporal-aware event-RGB fusion method for SAP, referred to as SERF, which aims to enhance the accuracy of event-based SAP. Specifically, SERF introduces three key components: 1) An innovative multi-layer Interaction Module based on attention mechanisms to fuse the multi-frame data, enabling more fine-grained feature processing; 2) a dynamic spatiotemporal mask mechanism, focusing RGB’s attention on spatially proximate events while diminishing the influence of temporally distant events, thereby reducing the impact of noise; and 3) a Memory Module that utilizes learnable tokens to accumulate essential latent fusion features through dynamic feature consolidation. Extensive experiments conducted on a variety of real-world and simulated datasets demonstrate the superior performance of SERF compared to the state-of-the-art methods. The experiments also validate the advantages of SERF in terms of inference performance, meeting the real-time requirements for actual deployment.
Canming Xia, Peixi Peng, Haoran Xu 0004, Guang Tan, Luntong Li, Yonghong Tian 0001
IEEE Trans. Intell. Transp. Syst.2
2026 KPGS: Toward Real-World Complex Dynamic Scene Rendering With Keyframe-Driven Predictable Gaussian Splatting
abstract
Rendering complex dynamic scenes offers the advantage of observing and understanding the real world. However, existing Dynamic Scene Rendering (DSR) methods remain challenged by suboptimal reconstruction fidelity. These limitations stem from relying on a single, unified deformation model, which struggles to capture complex motions involving multiple sub-motions and abrupt geometric transitions. While temporal decomposition methods could alleviate such shortcomings, they introduce the additional challenge of ignoring motion correlations and increasing storage requirements. To address these issues, we introduce Keyframe-driven Predictable Gaussian Splatting (KPGS)-an efficient framework for high-fidelity complex dynamic scene rendering. First, we present a patch-wise HSV clustering for extracting keyframes. Second, a prediction network based on the Transformer is utilized to calculate the deformable Gaussians at discrete keyframe times via voxelization. Third, we propose an inter-frame deformation network and a mutual supervision between adjacent segments to maintain the temporal continuity. Extensive experiments on our newly built dataset (MotionGS), as well as public benchmarks HyperNeRF and Neu3D, demonstrate that KPGS could achieve a higher average view synthesis performance than SOTA approaches, while maintaining a balance between storage cost and performance. More details of the demo and dataset are available at KPGS Supplementary.
Yiqian Chang, Haoran Xu 0004, Jianing Li 0001, Xuan Wang 0002, Yonghong Tian 0001, Peixi Peng
IEEE Trans. Vis. Comput. Graph.6
2025 Visual Reinforcement Learning with Residual Action
abstract
Learning control policy from continuous action space by visual observations is a fundamental and challenging task in reinforcement learning (RL). An essential problem is how to accurately map the high-dimensional images to the optimal actions by the policy network. Traditional decision-making modules output actions solely based on the current observation, while the distributions of optimal actions are dependent on specific tasks and cannot be known priorly, which increases the learning difficulty. To make the learning easier, we analyze the action characteristics in several control tasks, and propose Reinforcement Learning with Residual Action (ResAct) to explicitly model the adjustments of actions based on the differences between adjacent observations, rather than learning actions directly from observations. The method just redefines the output of the policy network, and doesn’t introduce any prior assumption to constrain or simplify the vanilla control problem. Extensive experiments on DeepMind Control Suite and CARLA demonstrate that the method could improve different RL baselines significantly, and achieve state-of-the-art performance.
Zhenxian Liu, Peixi Peng, Yonghong Tian 0001
AAAI2
2025 Towards Building Human-like Smart Agents in Modern 3D Video Games (Student Abstract)
abstract
In recent years, reinforcement learning has been widely applied in the field of games. However, most studies focus on assisting agents to achieve victory, with less attention paid to whether the agents exhibit human-like characteristics. In order to build human-like agents with high performance, we propose a method for learning the strategies of human players in modern three-dimensional video games. Our method utilizes a hierarchical framework, learning basic behaviors and intentions of human players at the lower level through imitation learning, and generalized policies at the high level through reinforcement learning. Compared with other existing methods, our method demonstrates significant advantages in learning human-like strategies in complex environments.
Zhihang Sun, Shuhan Qi, Xinhao Huang, Xinyu Xiao, Jiajia Zhang 0001, Xuan Wang 0002, Peixi Peng
AAAI7
2025 Exploiting Continuous Motion Clues for Vision-Based Occupancy Prediction
abstract
Occupancy networks aim to reconstruct the surroundings with occupied semantic voxels. However, frequent object occlusions often occur in dynamic real-world scenarios, which cannot be captured by independent frames. Most existing occupancy networks generate results without explicitly considering past occupancy states and continuous visual changes over time, limiting their temporal accuracy. We tackle it by treating the task from a new continuous updating perspective, which considers historical data and continuous motion clues. We propose a new approach termed Continuous Motion clue exploitation for Occupancy Prediction (CMOP), which incorporates three key designs: (i) Propagator: which forecasts future occupancy states based on historical data; (ii) Tracker: which updates the occupancy on a per-frame basis using dynamic visual motion information; and (iii) Fuser: which aggregates results from the Propagator and Tracker into more robust and accurate occupancy results. Experiments on several benchmarks demonstrate that CMOP outperforms state-of-the-art baselines.
Haoran Xu 0004, Peixi Peng, Guang Tan, Yaokun Li, Shuaixian Wang, Luntong Li
AAAI2
2025 VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement Learning
abstract
Vision-based Reinforcement Learning (VRL) attempts to establish associations between visual inputs and optimal actions through interactions with the environment. Given the high-dimensional and complex nature of visual data, it becomes essential to learn a policy based on high-quality state representation. To this end, existing VRL methods primarily rely on interaction-collected data, combined with selfsupervised auxiliary tasks. However, two key challenges remain: limited data samples and a lack of task-relevant semantic constraints. To tackle these challenges, we propose DGC, a method that Distills Guidance from Visual Language Models (VLMs) alongside self-supervised learning into a Compact VRL agent. Notably, we leverage the state representation capabilities of VLMs, rather than their decision-making abilities. Within DGC, a novel promptingreasoning pipeline is designed to convert historical observations and actions into usable supervision signals, enabling semantic understanding within the compact visual encoder. By leveraging these distilled semantic representations, the VRL agent achieves significant improvements in sample efficiency. Extensive experiments on the Carla benchmark demonstrate our state-of-the-art performance.
Haoran Xu 0004, Peixi Peng, Guang Tan, Yiqian Chang, Luntong Li, Yonghong Tian 0001
CVPR2
2025 CASA: Class-Agnostic Shared Attributes in Vision-Language Models for Efficient Incremental Object Detection
abstract
Incremental object detection is fundamentally challenged by catastrophic forgetting. A major factor contributing to this issue is background shift, where background categories in sequential tasks may overlap with either previously learned or future unseen classes. To address this, we propose a novel method called Class-Agnostic Shared Attribute Base (CASA) that encourages the model to learn category-agnostic attributes shared across incremental classes. Our approach leverages an LLM to generate candidate textual attributes, selects the most relevant ones based on the current training data, and records their importance in an assignment matrix. For subsequent tasks, the retained attributes are frozen, and new attributes are selected from the remaining candidates, ensuring both knowledge retention and adaptability. Extensive experiments on the COCO dataset demonstrate the state-of-the-art performance of our method.
Mingyi Guo, Zhiyuan Yan 0002, Zongying Lin, Peixi Peng, Yonghong Tian 0001
ICME5
2025 When Every Millisecond Counts: Real-Time Anomaly Detection via the Multimodal Asynchronous Hybrid Network
abstract
Anomaly detection is essential for the safety and reliability of autonomous driving systems. Current methods often focus on detection accuracy but neglect response time, which is critical in time-sensitive driving scenarios. In this paper, we introduce real-time anomaly detection for autonomous driving, prioritizing both minimal response time and high accuracy. We propose a novel multimodal asynchronous hybrid network that combines event streams from event cameras with image data from RGB cameras. Our network utilizes the high temporal resolution of event cameras through an asynchronous Graph Neural Network and integrates it with spatial features extracted by a CNN from RGB images. This combination effectively captures both the temporal dynamics and spatial details of the driving environment, enabling swift and precise anomaly detection. Extensive experiments on benchmark datasets show that our approach outperforms existing methods in both accuracy and response time, achieving millisecond-level real-time performance.
Peixi Peng, Yangru Huang, Yifan Zhao 0002, Yongxing Dai, Yonghong Tian 0001
ICML3
2025 Multi-faceted Complementary Learning for Incomplete Multi-view Multi-label Classification
Xinyu Xiao, Peixi Peng, Qiang Wang 0022, Shuhan Qi
ACM Multimedia2
2025 Spike4DGS: Towards High-Speed Dynamic Scene Rendering with 4D Gaussian Splatting via a Spike Camera Array
abstract
Spike camera with high temporal resolution offers a new perspective on high-speed dynamic scene rendering. Most existing rendering methods rely on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) for static scenes using a monocular spike camera. However, these methods struggle with dynamic motion, while a single camera suffers from limited spatial coverage, making it challenging to reconstruct fine details in high-speed scenes. To address these problems, we propose Spike4DGS, the first high-speed dynamic scene rendering framework with 4D Gaussian Splatting using spike camera arrays. Technically, we first build a multi-view spike camera array to validate our solution, then establish both synthetic and real-world multi-view spike-based reconstruction datasets. Then, we design a multi-view spike-based dense initialization module that obtains dense point clouds and camera poses from continuous spike streams. Finally, we propose a spike-pixel synergy constraint supervision to optimize Spike4DGS, incorporating both rendered image quality loss and dynamic spatiotemporal spike loss. The results show that our Spike4DGS outperforms state-of-the-art methods in terms of novel view rendering quality on both synthetic and real-world datasets. More details are available at https://github.com/Qinghongye/Spike4DGS.
Qinghong Ye, Yiqian Chang, Jianing Li 0001, Haoran Xu 0004, Xuan Wang 0002, Wei Zhang 0161, Yonghong Tian 0001, Peixi Peng
NeurIPS8
2025 Optimizing Spectrum Sharing in Low-Altitude Intelligent Networks: A Model-Based Reinforcement Learning Approach
abstract
The Low-Altitude Intelligent Network (LAIN) leverages cutting-edge Artificial Intelligence (AI) technologies to enable fast and reliable communications for Uncrewed Aerial Vehicles (UAVs), supporting the rapid growth of various low-altitude economy applications. However, as UAV deployments increase, the scarcity of spectrum resources has emerged as a critical bottleneck that severely limits the quality of service in LAIN. While Reinforcement Learning (RL) has shown promise in optimizing spectrum sharing and improving spectral efficiency of UAVs, existing RL-based methods suffer from high training costs and require extensive interactions with the environment to converge to near-optimal solutions. To address these challenges, we propose a novel model-based RL framework for dynamic spectrum sharing of UAVs in LAIN. Our approach employs a world model to learn compact state representations and reconstruct environment dynamics, significantly enhancing the sample efficiency in complex scenarios. Extensive simulations demonstrate that our approach drastically reduces the number of environmental interactions required for UAVs to achieve near-optimal spectrum allocations, lowering training costs while maintaining high performance.
Tianle Li, Peixi Peng
TrustCom3
2025 Fully Spiking Actor Network With Intralayer Connections for Reinforcement Learning
abstract
With the help of special neuromorphic hardware, spiking neural networks (SNNs) are expected to realize artificial intelligence (AI) with less energy consumption. It provides a promising energy-efficient way for realistic control tasks by combining SNNs with deep reinforcement learning (DRL). In this article, we focus on the task where the agent needs to learn multidimensional deterministic policies to control, which is very common in real scenarios. Recently, the surrogate gradient method has been utilized for training multilayer SNNs, which allows SNNs to achieve comparable performance with the corresponding deep networks in this task. Most existing spike-based reinforcement learning (RL) methods take the firing rate as the output of SNNs, and convert it to represent continuous action space (i.e., the deterministic policy) through a fully connected (FC) layer. However, the decimal characteristic of the firing rate brings the floating-point matrix operations to the FC layer, making the whole SNN unable to deploy on the neuromorphic hardware directly. To develop a fully spiking actor network (SAN) without any floating-point matrix operations, we draw inspiration from the nonspiking interneurons found in insects and employ the membrane voltage of the nonspiking neurons to represent the action. Before the nonspiking neurons, multiple population neurons are introduced to decode different dimensions of actions. Since each population is used to decode a dimension of action, we argue that the neurons in each population should be connected in time domain and space domain. Hence, the intralayer connections are used in output populations to enhance the representation capacity. This mechanism exists extensively in animals and has been demonstrated effectively. Finally, we propose a fully SAN with intralayer connections (ILC-SAN). Extensive experimental results demonstrate that the proposed method outperforms the state-of-the-art performance on continuous control tasks from OpenAI gym. Moreover, we estimate the theoretical energy consumption when deploying ILC-SAN on neuromorphic chips to illustrate its high energy efficiency.
Peixi Peng, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Adaptive Discovering and Merging for Incremental Novel Class Discovery
abstract
One important desideratum of lifelong learning aims to discover novel classes from unlabelled data in a continuous manner. The central challenge is twofold: discovering and learning novel classes while mitigating the issue of catastrophic forgetting of established knowledge. To this end, we introduce a new paradigm called Adaptive Discovering and Merging (ADM) to discover novel categories adaptively in the incremental stage and integrate novel knowledge into the model without affecting the original knowledge. To discover novel classes adaptively, we decouple representation learning and novel class discovery, and use Triple Comparison (TC) and Probability Regularization (PR) to constrain the probability discrepancy and diversity for adaptive category assignment. To merge the learned novel knowledge adaptively, we propose a hybrid structure with base and novel branches named Adaptive Model Merging (AMM), which reduces the interference of the novel branch on the old classes to preserve the previous knowledge, and merges the novel branch to the base model without performance loss and parameter growth. Extensive experiments on several datasets show that ADM significantly outperforms existing class-incremental Novel Class Discovery (class-iNCD) approaches. Moreover, our AMM also benefits the class-incremental Learning (class-IL) task by alleviating the catastrophic forgetting problem. The source code is included in the supplementary materials.
Peixi Peng, Yangru Huang, Mengyue Geng, Yonghong Tian 0001
AAAI2
2024 Density-Adaptive Model Based on Motif Matrix for Multi-Agent Trajectory Prediction
abstract
Multi-agent trajectory prediction is essential in autonomous driving, risk avoidance, and traffic flow control. However, the heterogeneous traffic density on interactions, which caused by physical laws, social norms and so on, is often overlooked in existing methods. When the density varies, the number of agents involved in interactions and the corresponding interaction probability change dynami-cally. To tackle this issue, we propose a new method, called Density-Adaptive Model based on Motif Matrix for Multi-Agent Trajectory Prediction (DAMM), to gain insights into multi-agent systems. Here we leverage the motif matrix to represent dynamic connectivity in a higher-order pattern, and distill the interaction information from the perspectives of the spatial and the temporal dimensions. Specifically, in spatial dimension, we utilize multi-scale feature fusion to adaptively select the optimal range of neighbors participating in interactions for each time slot. In temporal dimension, we extract the temporal interaction features and adapt a pyramidal pooling layer to generate the interaction probability for each agent. Experimental results demonstrate that our approach surpasses state-of-the-art methods on autonomous driving dataset.
Di Wen 0005, Haoran Xu 0004, Zhaocheng He, Zhe Wu 0006, Guang Tan, Peixi Peng
CVPR6
2024 DMR: Decomposed Multi-Modality Representations for Frames and Events Fusion in Visual Reinforcement Learning
abstract
We explore visual reinforcement learning (RL) using two complementary visual modalities: frame-based RGB cam-era and event-based Dynamic Vision Sensor (DVS). Ex-isting multi-modality visual RL methods often encounter challenges in effectively extracting task-relevant information from multiple modalities while suppressing the in-creased noise, only using indirect reward signals instead of pixel-level supervision. To tackle this, we propose a Decomposed Multi-Modality Representation (DMR) framework for visual RL. It explicitly decomposes the inputs into three distinct components: combined task-relevant features (co-features), RGB-specific noise, and DVS-specific noise. The co-features represent the full information from both modalities that is relevant to the RL task; the two noise components, each constrained by a data reconstruction loss to avoid information leak, are contrasted with the co-features to maximize their difference. Extensive experiments demonstrate that, by explicitly separating the different types of information, our approach achieves substan-tially improved policy performance compared to state-of-the-art approaches.
Haoran Xu 0004, Peixi Peng, Guang Tan, Yuan Li 0014, Xinhai Xu, Yonghong Tian 0001
CVPR2
2024 Prior-Posterior Knowledge Prompting-and-Reasoning for Surgical Visual Question Localized-Answering
abstract
The Surgical Visual Question Localized-Answering (VQLA) aims to locate the specific instance area while responing the associated question, which has the potential to assist junior resident doctors in understanding the surgical process and offering decision support for surgeons. Yet, this task remains a challenging job for data-driven neural networks, due to the serious reliance on surgical scene information provided by posterior knowledge. Hence, we propose a prior-posterior knowledge prompting-and-Reasoning (PPKPR) method to imitate the operational mode of a surgeon. Surgeons systematically inspect each instance within the surgical scene and its associated question, subsequently relying on their accumulated work experience to understand and answer the questions. The PPKPR comprises three modules: prior-posterior multi-domain knowledge prompter (PPMP), prior-posterior instance knowledge prompter (PPIP), and posterior knowledge Reasoner (PKR). Specifically, PPMP aligns prior-posterior multi-domain knowledge, thus prompting model to alleviate the misinterpretations of the textual question. PPIP provides the prior instance knowledge, ensuring model to focus on correct areas in the visual scene. The prompted knowledge is refined by the PKR to reasoning the final answer. Experimental results demonstrate that our method performs favorably against the state-of-the-art methods on the EndoVis-18 and EndoVis-17 datasets.
Peixi Peng, Wanshu Fan, Wenfei Liu, Xin Yang 0011
IJCNN1
2024 Seek Commonality but Preserve Differences: Dissected Dynamics Modeling for Multi-modal Visual RL
abstract
Accurate environment dynamics modeling is crucial for obtaining effective state representations in visual reinforcement learning (RL) applications. However, when facing multiple input modalities, existing dynamics modeling methods (e.g., DeepMDP) usually stumble in addressing the complex and volatile relationship between different modalities. In this paper, we study the problem of efficient dynamics modeling for multi-modal visual RL. We find that under the existence of modality heterogeneity, modality-correlated and distinct features are equally important but play different roles in reflecting the evolution of environmental dynamics. Motivated by this fact, we propose Dissected Dynamics Modeling (DDM), a novel multi-modal dynamics modeling method for visual RL. Unlike existing methods, DDM explicitly distinguishes consistent and inconsistent information across modalities and treats them separately with a divide-and-conquer strategy. This is done by dispatching the features carrying different information into distinct dynamics modeling pathways, which naturally form a series of implicit regularizations along the learning trajectories. In addition, a reward predictive function is further introduced to filter task-irrelevant information in both modality-consistent and inconsistent features, ensuring information integrity while avoiding potential distractions. Extensive experiments show that DDM consistently achieves competitive performance in challenging multi-modal visual environments.
Yangru Huang, Peixi Peng, Yifan Zhao 0002, Yonghong Tian 0001
NeurIPS2
2024 Sensitivity Decouple Learning for Image Compression Artifacts Reduction
abstract
With the benefit of deep learning techniques, recent researches have made significant progress in image compression artifacts reduction. Despite their improved performances, prevailing methods only focus on learning a mapping from the compressed image to the original one but ignore the intrinsic attributes of the given compressed images, which greatly harms the performance of downstream parsing tasks. Different from these methods, we propose to decouple the intrinsic attributes into two complementary features for artifacts reduction, i.e., the compression-insensitive features to regularize the high-level semantic representations during training and the compression-sensitive features to be aware of the compression degree. To achieve this, we first employ adversarial training to regularize the compressed and original encoded features for retaining high-level semantics, and we then develop the compression quality-aware feature encoder for compression-sensitive features. Based on these dual complementary features, we propose a Dual Awareness Guidance Network (DAGN) to utilize these awareness features as transformation guidance during the decoding phase. In our proposed DAGN, we develop a cross-feature fusion module to maintain the consistency of compression-insensitive features by fusing compression-insensitive features into the artifacts reduction baseline. Our method achieves an average 2.06 dB PSNR gains on BSD500, outperforming state-of-the-art methods, and only requires 29.7 ms to process one image on BSD500. Besides, the experimental results on LIVE1 and LIU4K also demonstrate the efficiency, effectiveness, and superiority of the proposed method in terms of quantitative metrics, visual quality, and downstream machine vision tasks.
Li Ma 0009, Yifan Zhao 0002, Peixi Peng, Yonghong Tian 0001
IEEE Trans. Image Process.3
2024 Eye Gaze Guided Cross-Modal Alignment Network for Radiology Report Generation
abstract
The potential benefits of automatic radiology report generation, such as reducing misdiagnosis rates and enhancing clinical diagnosis efficiency, are significant. However, existing data-driven methods lack essential medical prior knowledge, which hampers their performance. Moreover, establishing global correspondences between radiology images and related reports, while achieving local alignments between images correlated with prior knowledge and text, remains a challenging task. To address these shortcomings, we introduce a novel Eye Gaze Guided Cross-modal Alignment Network (EGGCA-Net) for generating accurate medical reports. Our approach incorporates prior knowledge from radiologists' Eye Gaze Region (EGR) to refine the fidelity and comprehensibility of report generation. Specifically, we design a Dual Fine-Grained Branch (DFGB) and a Multi-Task Branch (MTB) to collaboratively ensure the alignment of visual and textual semantics across multiple levels. To establish fine-grained alignment between EGR-related images and sentences, we introduce the Sentence Fine-grained Prototype Module (SFPM) within DFGB to capture cross-modal information at different levels. Additionally, to learn the alignment of EGR-related image topics, we introduce the Multi-task Feature Fusion Module (MFFM) within MTB to refine the encoder output information. Finally, a specifically designed label matching mechanism is designed to generate reports that are consistent with the anticipated disease states. The experimental outcomes indicate that the introduced methodology surpasses previous advanced techniques, yielding enhanced performance on two extensively used benchmark datasets: Open-i and MIMIC-CXR.
Peixi Peng, Wanshu Fan, Wenfei Liu, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei
IEEE J. Biomed. Health Informatics1
2024 Cascaded Attention: Adaptive and Gated Graph Attention Network for Multiagent Reinforcement Learning
abstract
Modeling the interactive relationships of agents is critical to improving the collaborative capability of a multiagent system. Some methods model these by predefined rules. However, due to the nonstationary problem, the interactive relationship changes over time and cannot be well captured by rules. Other methods adopt a simple mechanism such as an attention network to select the neighbors the current agent should collaborate with. However, in large-scale multiagent systems, collaborative relationships are too complicated to be described by a simple attention network. We propose an adaptive and gated graph attention network (AGGAT), which models the interactive relationships between agents in a cascaded manner. In the AGGAT, we first propose a graph-based hard attention network that roughly filters irrelevant agents. Then, normal soft attention is adopted to decide the importance of each neighbor. Finally, gated attention further refines the collaborative relationship of agents. By using cascaded attention, the collaborative relationship of agents is precisely learned in a coarse-to-fine style. Extensive experiments are conducted on a variety of cooperative tasks. The results indicate that our proposed method outperforms state-of-the-art baselines.
Shuhan Qi, Xinhao Huang, Peixi Peng, Xuzhong Huang, Jiajia Zhang 0001, Xuan Wang 0002
IEEE Trans. Neural Networks Learn. Syst.3
2023 Learning with Fantasy: Semantic-Aware Virtual Contrastive Constraint for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) aims at learning to classify new classes continually from limited samples without forgetting the old classes. The mainstream framework tackling FSCIL is first to adopt the cross-entropy (CE) loss for training at the base session, then freeze the feature extractor to adapt to new classes. However, in this work, we find that the CE loss is not ideal for the base session training as it suffers poor class separation in terms of representations, which further degrades generalization to novel classes. One tempting method to mitigate this problem is to apply an additional naïve supervised contrastive learning (SCL) in the base session. Unfortunately, we find that although SCL can create a slightly better representation separation among different base classes, it still struggles to separate base classes and new classes. Inspired by the observations made, we propose Semantic-Aware Virtual Contrastive model (SAVC), a novel method that facilitates separation between new classes and base classes by introducing virtual classes to SCL. These virtual classes, which are generated via pre-defined transformations, not only act as placeholders for unseen classes in the representation space, but also provide diverse semantic information. By learning to recognize and contrast in the fantasy space fostered by virtual classes, our SAVC significantly boosts base class separation and novel class generalization, achieving new state-of-the-art performance on the three widely-used FSCIL benchmark datasets. Code is available at: https://github.com/zysong0113/SAVC.
Zeyin Song, Yifan Zhao 0002, Yujun Shi, Peixi Peng, Li Yuan 0007, Yonghong Tian 0001
CVPR4
2023 Simoun: Synergizing Interactive Motion-appearance Understanding for Vision-based Reinforcement Learning
abstract
Efficient motion and appearance modeling are critical for vision-based Reinforcement Learning (RL). However, existing methods struggle to reconcile motion and appearance information within the state representations learned from a single observation encoder. To address the problem, we present Synergizing Interactive Motion-appearance Understanding (Simoun), a unified framework for vision-based RL Given consecutive observation frames, Simoun deliberately and interactively learns both motion and appearance features through a dual-path network architecture. The learning process collaborates with a structural interactive module, which explores the latent motion-appearance structures from the two network paths to leverage their complementarity. To promote sample efficiency, we further design a consistency-guided curiosity module to encourage the exploration of under-learned observations. During training, the curiosity module provides intrinsic rewards according to the consistency of environmental temporal dynamics, which are deduced from both motion and appearance network paths. Experiments conducted on Deep-Mind control suite and CARLA automatic driving benchmarks demonstrate the effectiveness of Simoun, where it performs favorably against state-of-the-art methods.
Yangru Huang, Peixi Peng, Yifan Zhao 0002, Yunpeng Zhai, Haoran Xu 0004, Yonghong Tian 0001
ICCV2
2023 Stabilizing Visual Reinforcement Learning via Asymmetric Interactive Cooperation
abstract
Vision-based reinforcement learning (RL) depends on discriminative representation encoders to abstract the observation states. Despite the great success of increasing CNN parameters for many supervised computer vision tasks, reinforcement learning with temporal-difference (TD) losses cannot benefit from it in most complex environments. In this paper, we analyze that the training instability arises from the oscillating self-overfitting of the heavy-optimizable encoder. We argue that serious oscillation will occur to the parameters when enforced to fit the sensitive TD targets, causing uncertain drifting of the latent state space and thus transmitting these perturbations to the policy learning. To alleviate this phenomenon, we propose a novel asymmetric interactive cooperation approach with the interaction between a heavy-optimizable encoder and a supportive light-optimizable encoder, in which both their advantages are integrated including the highly discriminative capability as well as the training stability. We also present a greedy bootstrapping optimization to isolate the visual perturbations from policy learning, where representation and policy are trained sufficiently by turns. Finally, we demonstrate the effectiveness of our method in utilizing larger visual models by first-person highway driving task CARLA and Vizdoom environments.
Yunpeng Zhai, Peixi Peng, Yifan Zhao 0002, Yangru Huang, Yonghong Tian 0001
ICCV2
2023 Learning Sparse Neural Networks with Identity Layers
Mingjian Ni, Xiawu Zheng, Peixi Peng, Li Yuan 0007, Yonghong Tian 0001
ICIG (3)4
2023 Reinforcement Learning-Based Consensus Reaching in Large-Scale Social Networks
Shijun Guo, Haoran Xu 0004, Guangqiang Xie, Di Wen 0005, Yangru Huang, Peixi Peng
ICONIP (8)6
2023 Dynamic Belief for Decentralized Multi-Agent Cooperative Learning
abstract
Decentralized multi-agent cooperative learning is a practical task due to the partially observed setting both in training and execution. Every agent learns to cooperate without access to the observations and policies of others. However, the decentralized training of multi-agent is of great difficulty due to non-stationarity, especially when other agents' policies are also in learning during training. To overcome this, we propose to learn a dynamic policy belief for each agent to predict the current policies of other agents and accordingly condition the policy of its own. To quickly adapt to the development of others' policies, we introduce a historical context to learn the belief inference according to a few recent action histories of other agents and a latent variational inference to model their policies by a learned distribution. We evaluate our method on the StarCraft II micro management task (SMAC) and demonstrate its superior performance in the decentralized training settings and comparable results with the state-of-the-art CTDE methods.
Yunpeng Zhai, Peixi Peng, Yonghong Tian 0001
IJCAI2
2023 Hierarchical Adaptive Value Estimation for Multi-modal Visual Reinforcement Learning
abstract
Integrating RGB frames with alternative modality inputs is gaining increasing traction in many vision-based reinforcement learning (RL) applications. Existing multi-modal vision-based RL methods usually follow a Global Value Estimation (GVE) pipeline, which uses a fused modality feature to obtain a unified global environmental description. However, such a feature-level fusion paradigm with a single critic may fall short in policy learning as it tends to overlook the distinct values of each modality. To remedy this, this paper proposes a Local modality-customized Value Estimation (LVE) paradigm, which dynamically estimates the contribution and adjusts the importance weight of each modality from a value-level perspective. Furthermore, a task-contextual re-fusion process is developed to achieve a task-level re-balance of estimations from both feature and value levels. To this end, a Hierarchical Adaptive Value Estimation (HAVE) framework is formed, which adaptively coordinates the contributions of individual modalities as well as their collective efficacy. Agents trained by HAVE are able to exploit the unique characteristics of various modalities while capturing their intricate interactions, achieving substantially improved performance. We specifically highlight the potency of our approach within the challenging landscape of autonomous driving, utilizing the CARLA benchmark with neuromorphic event and depth data to demonstrate HAVE's capability and the effectiveness of its distinct components.
Yangru Huang, Peixi Peng, Yifan Zhao 0002, Haoran Xu 0004, Mengyue Geng, Yonghong Tian 0001
NeurIPS2
2023 Population-Based Evolutionary Gaming for Unsupervised Person Re-identification
Yunpeng Zhai, Peixi Peng, Mengxi Jia, Xuesong Gao, Yonghong Tian 0001
Int. J. Comput. Vis.2
2023 Picking Up Quantization Steps for Compressed Image Classification
abstract
The sensitivity of deep neural networks to compressed images hinders their usage in many real applications, which means classification networks may fail just after taking a screenshot and saving it as a compressed file. In this paper, we argue that neglected disposable coding parameters stored in compressed files could be picked up to reduce the sensitivity of deep neural networks to compressed images. Specifically, we resort to using one of the representative parameters, quantization steps, to facilitate image classification. Firstly, based on quantization steps, we propose a novel quantization aware confidence (QAC), which is utilized as sample weights to reduce the influence of quantization on network training. Secondly, we utilize quantization steps to alleviate the variance of feature distributions, where a quantization aware batch normalization (QABN) is proposed to replace batch normalization of classification networks. Extensive experiments show that the proposed method significantly improves the performance of classification networks on CIFAR-10, CIFAR-100, and ImageNet.
Li Ma 0009, Peixi Peng, Yifan Zhao 0002, Siwei Dong, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 MetaVIM: Meta Variationally Intrinsic Motivated Reinforcement Learning for Decentralized Traffic Signal Control
abstract
Traffic signal control aims to coordinate traffic signals across intersections to improve the traffic efficiency of a district or a city. Deep reinforcement learning (RL) has been applied to traffic signal control recently and demonstrated promising performance where each traffic signal is regarded as an agent. However, there are still several challenges that may limit its large-scale application in the real world. On the one hand, the policy of the current traffic signal is often heavily influenced by its neighbor agents, and the coordination between the agent and its neighbors needs to be considered. Hence, the control of a road network composed of multiple traffic signals is naturally modeled as a multi-agent system, and all agents’ policies need to be optimized simultaneously. On the other hand, once the policy function is conditioned on not only the current agent's observation but also the neighbors’, the policy function would be closely related to the training scenario and cause poor generalizability because the agents in various scenarios often have heterogeneous neighbors. To make the policy learned from a training scenario generalizable to new unseen scenarios, a novel Meta Variationally Intrinsic Motivated (MetaVIM) RL method is proposed to learn the decentralized policy for each intersection that considers neighbor information in a latent way. Specifically, we formulate the policy learning as a meta-learning problem over a set of related tasks, where each task corresponds to traffic signal control at an intersection whose neighbors are regarded as the unobserved part of the state. Then, a learned latent variable is introduced to represent the task's specific information and is further brought into the policy for learning. In addition, to make the policy learning stable, a novel intrinsic reward is designed to encourage each agent's received rewards and observation transition to be predictable only conditioned on its own history. Extensive experiments conducted on CityFlow demonstrate that the proposed method substantially outperforms existing approaches and shows superior generalizability.
Liwen Zhu 0003, Peixi Peng, Zongqing Lu 0002, Yonghong Tian 0001
IEEE Trans. Knowl. Data Eng.2
2022 Spectrum Random Masking for Generalization in Image-based Reinforcement Learning
abstract
Generalization in image-based reinforcement learning (RL) aims to learn a robust policy that could be applied directly on unseen visual environments, which is a challenging task since agents usually tend to overfit to their training environment. To handle this problem, a natural approach is to increase the data diversity by image based augmentations. However, different with most vision tasks such as classification and detection, RL tasks are not always invariant to spatial based augmentations due to the entanglement of environment dynamics and visual appearance. In this paper, we argue with two principles for augmentations in RL: First, the augmented observations should facilitate learning a universal policy, which is robust to various distribution shifts. Second, the augmented data should be invariant to the learning signals such as action and reward. Following these rules, we revisit image-based RL tasks from the view of frequency domain and propose a novel augmentation method, namely Spectrum Random Masking (SRM),which is able to help agents to learn the whole frequency spectrum of observation for coping with various distributions and compatible with the pre-collected action and reward corresponding to original observation. Extensive experiments conducted on DMControl Generalization Benchmark demonstrate the proposed SRM achieves the state-of-the-art performance with strong generalization potentials.
Yangru Huang, Peixi Peng, Yifan Zhao 0002, Yonghong Tian 0001
NeurIPS2
2022 Adversarial Reciprocal Points Learning for Open Set Recognition
abstract
Open set recognition (OSR), aiming to simultaneously classify the seen classes and identify the unseen classes as 'unknown', is essential for reliable machine learning. The key challenge of OSR is how to reduce the empirical classification risk on the labeled known data and the open space risk on the potential unknown data simultaneously. To handle the challenge, we formulate the open space risk problem from the perspective of multi-class integration, and model the unexploited extra-class space with a novel concept Reciprocal Point. Follow this, a novel learning framework, termed Adversarial Reciprocal Point Learning (ARPL), is proposed to minimize the overlap of known distribution and unknown distributions without loss of known classification accuracy. Specifically, each reciprocal point is learned by the extra-class space with the corresponding known category, and the confrontation among multiple known categories are employed to reduce the empirical classification risk. Then, an adversarial margin constraint is proposed to reduce the open space risk by limiting the latent open space constructed by reciprocal points. To further estimate the unknown distribution from open space, an instantiated adversarial enhancement method is designed to generate diverse and confusing training samples, based on the adversarial mechanism between the reciprocal points and known classes. This can effectively enhance the model distinguishability to the unknown classes. Extensive experimental results on various benchmark datasets indicate that the proposed method is significantly superior to other existing approaches and achieves state-of-the-art performance. The code is released on github.com/iCGY96/ARPL.
Peixi Peng, Xiangqian Wang 0001, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Self-Guided Adaptation: Progressive Representation Alignment for Domain Adaptive Object Detection
abstract
Unsupervised domain adaptation (UDA) has achieved unprecedented success in improving the cross-domain robustness of object detection models. However, existing UDA methods largely ignore the instantaneous data distribution and the sampling strategy during model learning, which could deteriorate the feature representation given large domain shift. In this work, we propose a Self-Guided Adaptation (SGA) model, targeting at aligning feature representation and transferring object detection models across domains while considering the instantaneous alignment difficulty. The core of SGA is to calculate “hardness” factors for sample pairs indicating domain distance in a kernel space. With the hardness factor, the proposed SGA adaptively indicates the importance of samples and assigns them different constrains. Indicated by these hardness factors, Self-Guided Progressive Sampling (SPS) is implemented in an “easy-to-hard” way during model adaptation. Using multi-stage convolutional features, SGA is further aggregated to fully align hierarchical representations of detection models. Extensive experiments on commonly-used benchmarks show that SGA improves the state-of-the-art methods with significant margins especially on large domain shift cases.
Zongxian Li, Peixi Peng, Qixiang Ye, Shijian Lu, Tiejun Huang 0001, Yonghong Tian 0001
IEEE Trans. Multim.4
2021 Reducing Image Compression Artifacts for Deep Neural Networks
abstract
Existing compression artifacts reduction methods aim to restore images on pixel-level, which can improve the human visual experience. However, in many applications, large-scale images are collected not for visual examination by humans. Instead, they are used for many high-level vision tasks usually by Deep Neural Networks (DNN). In this paper, we find that these methods have limited performance improvements to high-level tasks, even bring negative effects. Therefore, inspired by the teacher-student network framework, we propose a compression artifacts reduction framework (ARF) for DNN. In addition, we generalize our method to the unsupervised setting (U-ARF) where the corresponding original images are unavailable in training. Extensive experiments indicate the proposed methods can help DNNs improve performance on the highly compressed images significantly.
Li Ma 0009, Peixi Peng, Peiyin Xing, Yaowei Wang 0001, Yonghong Tian 0001
DCC2
2021 Allocating DNN Layers Computation Between Front-End Devices and The Cloud Server for Video Big Data Processing
abstract
With the development of intelligent hardware, front-end devices can also perform DNN computation. Moreover, the deep neural network can be divided into several layers. In this way, part of the computation of DNN models can be migrated to the front-end devices, which can alleviate the cloud burden and shorten the processing latency. This paper proposes a computation allocation algorithm of DNN between the front-end devices and the cloud server. In brief, we divide the DNN layers dynamically according to the current and the predicted future status of the processing system, by which we obtain a shorter end-to-end latency. The simulation results reveal that the overall latency reduction is more than 70% compared with traditional cloud-centered processing.
Peiyin Xing, Peixi Peng, Tiejun Huang 0001, Yonghong Tian 0001
ICASSP3
2021 Amplitude-Phase Recombination: Rethinking Robustness of Convolutional Neural Networks in Frequency Domain
abstract
Recently, the generalization behavior of Convolutional Neural Networks (CNN) is gradually transparent through explanation techniques with the frequency components decomposition. However, the importance of the phase spectrum of the image for a robust vision system is still ignored. In this paper, we notice that the CNN tends to converge at the local optimum which is closely related to the high-frequency components of the training images, while the amplitude spectrum is easily disturbed such as noises or common corruptions. In contrast, more empirical studies found that humans rely on more phase components to achieve robust recognition. This observation leads to more explanations of the CNN’s generalization behaviors in both robustness to common perturbations and out-of-distribution detection, and motivates a new perspective on data augmentation designed by re-combing the phase spectrum of the current image and the amplitude spectrum of the distracter image. That is, the generated samples force the CNN to pay more attention to the structured information from phase components and keep robust to the variation of the amplitude. Experiments on several image datasets indicate that the proposed method achieves state-of-the-art performances on multiple generalizations and calibration tasks, including adaptability for common corruptions and surface variations, out-of-distribution detection, and adversarial attack. The code is released on github/iCGY96/APR.
Peixi Peng, Li Ma 0009, Jia Li 0003, Lin Du 0010, Yonghong Tian 0001
ICCV2
2021 Model Latent Views With Multi-Center Metric Learning for Vehicle Re-Identification
abstract
Multi-view vehicle re-identification (Re-ID) aims to retrieve all images of a target vehicle from a large gallery where the vehicles are captured from non-overlapping cameras. However, the drastic variation in vehicle appearance under different viewpoints greatly affects the performance of the multi-view vehicle Re-ID model, so the key issue in multi-view vehicle Re-ID is learning an effective feature representation that is robust to both dramatic intra-class variability and small inter-class variability. To achieve this goal, we have proposed a multi-center metric learning framework for multi-view vehicle Re-ID. In our approach, we model latent views from vehicle visual appearance directly without any extra labels except ID. Firstly, we introduce several latent view clusters for a vehicle to model latent multi-view information and each view cluster has a learnable center. Then multi-view vehicle matching task can be transformed into two subproblems, cross-view matching and cross-target matching. Finally, an intra-class ranking loss with cross-view center constraint and a cross-class ranking loss with cross-vehicle center constraint are proposed to address the two subproblems, respectively. Extensive experimental evaluations on three widely used benchmarks show the superiority of the proposed framework in contrast to a series of existing state-of-the-arts.
Yi Jin 0001, Chenning Li, Yidong Li, Peixi Peng, George A. Giannopoulos
IEEE Trans. Intell. Transp. Syst.4
2020 Domain Adaptive Attention Learning for Unsupervised Person Re-Identification
abstract
Person re-identification (Re-ID) across multiple datasets is a challenging task due to two main reasons: the presence of large cross-dataset distinctions and the absence of annotated target instances. To address these two issues, this paper proposes a domain adaptive attention learning approach to reliably transfer discriminative representation from the labeled source domain to the unlabeled target domain. In this approach, a domain adaptive attention model is learned to separate the feature map into domain-shared part and domain-specific part. In this manner, the domain-shared part is used to capture transferable cues that can compensate cross-dataset distinctions and give positive contributions to the target task, while the domain-specific part aims to model the noisy information to avoid the negative transfer caused by domain diversity. A soft label loss is further employed to take full use of unlabeled target data by estimating pseudo labels. Extensive experiments on the Market-1501, DukeMTMC-reID and MSMT17 benchmarks demonstrate the proposed approach outperforms the state-of-the-arts.
Yangru Huang, Peixi Peng, Yi Jin 0001, Yidong Li, Junliang Xing
AAAI2
2020 Binary Representation and High Efficient Compression of 3D CNN Features for Action Recognition
abstract
A common framework of the action recognition is to collect the videos from different cameras into a cloud center firstly, and then perform the 3D CNN on the cloud server. Although directly, this framework will bring a huge burden to the cloud server and video transmission. To handle this challenge, the "front-cloud" collaborative processing architecture can be used. The most import issue is to compress the feature from 3D CNN effectively without significant loss of accuracy. We propose logarithmic quantization with a maximum value threshold and HEVC inter encoding for 3D CNN features. Experimental results on ResNet-50 and InceptionV1 show that the features can be represented by only 1 bit without significant loss of accuracy. The compression ratio of the quantized 1 bit features using HEVC inter coding can reach to 5000 times and the loss of accuracy is less than 1%.
Peiyin Xing, Peixi Peng, Yongsheng Liang 0001, Tiejun Huang 0001, Yonghong Tian 0001
DCC2
2020 Learning Open Set Network with Discriminative Reciprocal Points
Limeng Qiao, Yemin Shi 0001, Peixi Peng, Jia Li 0003, Tiejun Huang 0001, Shiliang Pu, Yonghong Tian 0001
ECCV (3)4
2020 Hybrid Learning for Multi-agent Cooperation with Sub-optimal Demonstrations
abstract
This paper aims to learn multi-agent cooperation where each agent performs its actions in a decentralized way. In this case, it is very challenging to learn decentralized policies when the rewards are global and sparse. Recently, learning from demonstrations (LfD) provides a promising way to handle this challenge. However, in many practical tasks, the available demonstrations are often sub-optimal. To learn better policies from these sub-optimal demonstrations, this paper follows a centralized learning and decentralized execution framework and proposes a novel hybrid learning method based on multi-agent actor-critic. At first, the expert trajectory returns generated from demonstration actions are used to pre-train the centralized critic network. Then, multi-agent decisions are made by best response dynamics based on the critic and used to train the decentralized actor networks. Finally, the demonstrations are updated by the actor networks, and the critic and actor networks are learned jointly by running the above two steps alliteratively. We evaluate the proposed approach on a real-time strategy combat game. Experimental results show that the approach outperforms many competing demonstration-based methods.
Peixi Peng, Junliang Xing, Lili Cao
IJCAI1
2020 Masked Face Recognition with Latent Part Detection
abstract
This paper focuses on a novel task named masked faces recognition (MFR), which aims to match masked faces with common faces and is important especially during the global outbreak of COVID-19. It is challenging to identify masked faces for two main reasons. Firstly, there is no large-scale training data and test data with ground truth for MFR. Collecting and annotating millions of masked faces is labor-consuming. Secondly, since most facial cues are occluded by mask, it is necessary to learn representations which are both discriminative and robust to mask wearing. To handle the first challenge, this paper collects two datasets designed for MFR: MFV with 400 pairs of 200 identities for verification, and MFI which contains 4,916 images of 669 identities for identification. As is known, a robust face recognition model needs images of millions of identities to train, and hundreds of identities is far from enough. Hence, MFV and MFI are only considered as test datasets to evaluate algorithms. Besides, a data augmentation method for training data is introduced to automatically generate synthetic masked face images from existing common face datasets. In addition, a novel latent part detection (LPD) model is proposed to locate the latent facial part which is robust to mask wearing, and the latent part is further used to extract discriminative features. The proposed LPD model is trained in an end-to-end manner and only utilizes the original and synthetic training data. Experimental results on MFV, MFI and synthetic masked LFW demonstrate that LPD model generalizes well on both realistic and synthetic masked data and outperforms other methods by a large margin.
Feifei Ding, Peixi Peng, Yangru Huang, Mengyue Geng, Yonghong Tian 0001
ACM Multimedia2
2020 Masked Face Recognition with Generative Data Augmentation and Domain Constrained Ranking
abstract
Masked faces recognition (MFR) aims to match a masked face with its corresponding full face, which is an important task especially during the global outbreak of COVID-19. However, most existing face recognition models generalize poorly in this case, and it is hard to train a robust MFR model due to two main reasons: 1) the absence of large scale training data as well as ground truth testing data, and 2) the presence of large intra-class variation between masked faces and full faces. To address the first challenge, this paper firstly contributes a new dataset denoted as MFSR, which consists of two parts. The first part contains 9,742 masked face images with mask region segmentation annotation. The second part contains 11,615 images of 1,004 identities, and each identity has masked and full face images with various orientations, lighting conditions and mask types. However, it is still not enough for training MFR models with deep learning. To obtain sufficient training data, based on the MFSR, we introduce a novel Identity Aware Mask GAN (IAMGAN) with segmentation guided multi-level identity preserve module to generate the synthetic masked face images from the full face images. In addition, to tackle the second challenge, a Domain Constrained Ranking (DCR) loss is proposed by adopting a center-based cross-domain ranking strategy. For each identity, two centers are designed which correspond to the full face images and the masked face images respectively. The DCR forces the feature of masked faces getting closer to its corresponding full face center and vice-versa. Experimental results on the MFSR dataset demonstrate the effectiveness of the proposed approaches.
Mengyue Geng, Peixi Peng, Yangru Huang, Yonghong Tian 0001
ACM Multimedia2
2020 Discriminative Spatial Feature Learning for Person Re-Identification
abstract
Person re-identification (ReID) aims to match detected pedestrian images from multiple non-overlapping cameras. Most existing methods employ a backbone CNN to extract a vectorized feature representation by performing some global pooling operations (such as global average pooling and global max pooling) on the 3D feature map (i.e., the output of the backbone CNN). Although simple and effective in some situations, the global pooling operation only focuses on the statistical properties and ignores the spatial distribution of the feature map. Hence, it can not distinguish two feature maps when they have similar response values located in totally different positions. To handle this challenge, a novel method is proposed to learn the discriminative spatial features. Firstly, a self-constrained spatial transformer network (SC-STN) is introduced to handle the misalignments caused by detection errors. Then, based on the prior knowledge that the spatial structure of a pedestrian often keeps robust in vertical orientation of images, a novel vertical convolution network (VCN) is proposed to extract the spatial feature in vertical. Extensive experimental evaluations on several benchmarks demonstrate that the proposed method achieves state-of-the-art performances by introducing only a few parameters to the backbone.
Peixi Peng, Yonghong Tian 0001, Yangru Huang, Xiangqian Wang 0001, Huilong An
ACM Multimedia1
2019 Unsupervised Person Re-identification Based on Clustering and Domain-Invariant Network
Yangru Huang, Yi Jin 0001, Peixi Peng, Congyan Lang, Yidong Li
ICIG (3)3
2019 Vehicle Re-Identification by Multi-Grain Learni
abstract
Vehicle re-identification (re-ID) is to identify the same vehicle captured by different cameras with non-overlapping views, which plays an important role in intelligent transportation system and traffic safety. Compared with face recognition and person re-ID tasks, it is difficult to train an effective vehicle re-ID model since different vehicles of the same vehicle model, such as Mercedes-Benz C300, may exhibit strong inter-class similarity. To handle this difficulty, we propose a multi-grain ranking loss with the auxiliary of vehicle model, which models the vehicle re-ID task as two sub-tasks with different granularities including matching vehicles in the same vehicle model and different vehicle models. In particular, we infer the vehicle model labels online for the unlabeled training samples by clustering. The experimental results on two benchmarks demonstrate the proposed method can achieve state-of-the-art performance.
Xiaoliang Yang, Congyan Lang, Peixi Peng, Junliang Xing
ICIP3
2019 Multi-View Learning for Vehicle Re-Identification
abstract
Vehicle re-identification (ReID) aims to identify a target vehicle in different cameras with non-overlapping views, and it plays an important role when the car licence plate recognition is unavailable or unreliable. Compared with face recognition and person ReID tasks, it is difficult to train an effective vehicle ReID model due to two reasons: the different views greatly affect the visual appearance of a vehicle, and different vehicles may exhibit fairly similar visual appearance when their images are captured from one unified single view. To handle these training difficulties, we introduce several latent groups to represent multiple views. Then, the vehicle ReID problem is modeled as two sub tasks, including matching vehicles in a same view and across different views. A fine-grain ranking loss and a relative coarse-grain ranking loss are proposed to each task respectively. Extensive experimental analyses and evaluations on two benchmarks demonstrate the proposed method can achieve state-of-the-art performance.
Weipeng Lin, Yidong Li, Xiaoliang Yang, Peixi Peng, Junliang Xing
ICME4
2019 Learning Deep Decentralized Policy Network by Collective Rewards for Real-Time Combat Game
abstract
The task of real-time combat game is to coordinate multiple units to defeat their enemies controlled by the given opponent in a real-time combat scenario. It is difficult to design a high-level Artificial Intelligence (AI) program for such a task due to its extremely large state-action space and real-time requirements. This paper formulates this task as a collective decentralized partially observable Markov decision process, and designs a Deep Decentralized Policy Network (DDPN) to model the polices. To train DDPN effectively, a novel two-stage learning algorithm is proposed which combines imitation learning from opponent and reinforcement learning by no-regret dynamics. Extensive experimental results on various combat scenarios indicate that proposed method can defeat different opponent models and significantly outperforms many state-of-the-art approaches.
Peixi Peng, Junliang Xing, Lili Cao, Lisen Mu, Chang Huang
IJCAI1
2019 Joint Learning of Dictionary and Convolutional Network for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition is to predict the presence of a set of attributes from a given image, and it plays an important role in video surveillance applications. Most existing works model the task as a multi-label classification problem. Although effective, they ignore the existence of correlations among attributes. In this work, to learn multiple attributes jointly, the attributes are modeled as a subspace and a dictionary is introduced to represent the subspace. Furthermore, to extract the convolutional features which are more suitable for attribute prediction, the dictionary is modeled as a network layer which is learned jointly with the convolutional network. Finally, a novel learning algorithm is proposed to optimize the dictionary and the convolutional network corporately. Extensive experimental analyses and evaluations on two largest pedestrian attribute benchmarks PETA and PA-100K demonstrate that the proposed method achieves state-of-the-art performance.
Yan Sha, Congyan Lang, Peixi Peng, Junliang Xing, Danxia Li
VCIP3
2018 Visual Tracking via Spatially Aligned Correlation Filters Network
Mengdan Zhang, Qiang Wang 0051, Junliang Xing, Peixi Peng, Weiming Hu 0004, Stephen J. Maybank
ECCV (3)5
2018 Joint Semantic and Latent Attribute Modelling for Cross-Class Transfer Learning
abstract
A number of vision problems such as zero-shot learning and person re-identification can be considered as cross-class transfer learning problems. As mid-level semantic properties shared cross different object classes, attributes have been studied extensively for knowledge transfer across classes. Most previous attribute learning methods focus only on human-defined/nameable semantic attributes, whilst ignoring the fact there also exist undefined/latent shareable visual properties, or latent attributes. These latent attributes can be either discriminative or non-discriminative parts depending on whether they can contribute to an object recognition task. In this work, we argue that learning the latent attributes jointly with user-defined semantic attributes not only leads to better representation but also helps semantic attribute prediction. A novel dictionary learning model is proposed which decomposes the dictionary space into three parts corresponding to semantic, latent discriminative and latent background attributes respectively. Such a joint attribute learning model is then extended by following a multi-task transfer learning framework to address a more challenging unsupervised domain adaptation problem, where annotations are only available on an auxiliary dataset and the target dataset is completely unlabelled. Extensive experiments show that the proposed models, though being linear and thus extremely efficient to compute, produce state-of-the-art results on both zero-shot learning and person re-identification.
Peixi Peng, Yonghong Tian 0001, Tao Xiang 0002, Yaowei Wang 0001, Massimiliano Pontil, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Unsupervised Cross-Dataset Transfer Learning for Person Re-identification
abstract
Most existing person re-identification (Re-ID) approaches follow a supervised learning framework, in which a large number of labelled matching pairs are required for training. This severely limits their scalability in realworld applications. To overcome this limitation, we develop a novel cross-dataset transfer learning approach to learn a discriminative representation. It is unsupervised in the sense that the target dataset is completely unlabelled. Specifically, we present an multi-task dictionary learning method which is able to learn a dataset-shared but target-data-biased representation. Experimental results on five benchmark datasets demonstrate that the method significantly outperforms the state-of-the-art.
Peixi Peng, Tao Xiang 0002, Yaowei Wang 0001, Massimiliano Pontil, Shaogang Gong, Tiejun Huang 0001, Yonghong Tian 0001
CVPR1
2016 Joint Learning of Semantic and Latent Attributes
Peixi Peng, Yonghong Tian 0001, Tao Xiang 0002, Yaowei Wang 0001, Tiejun Huang 0001
ECCV (4)1
2015 Robust multiple cameras pedestrian detection with multi-view Bayesian network
Peixi Peng, Yonghong Tian 0001, Yaowei Wang 0001, Jia Li 0003, Tiejun Huang 0001
Pattern Recognit.1
2012 Single and Multiple View Detection, Tracking and Video Analysis in Crowded Environments
abstract
In this paper, we present our detection, tracking and event recognition methods and the results for PETS 2012. First, ROIs (Regions of Interest) based on geometric constraints are utilized in single view detection to eliminate the negative influence of clutter environment. Then, an optimized observation model is applied to address the ID switching or tracking drifting problem in single view tracking. Third, we introduce the multi-view Bayesian network (MBN) to reduce the "phantom" phenomena which frequently happen in general multi-view detection tasks. At last, a motion-based event recognition method is proposed to handle the event recognition task. Experimental results on the PETS 2012 dataset indicate that our methods are very promising.
Teng Xu 0002, Peixi Peng, Xiaoyu Fang, Chi Su, Yaowei Wang 0001, Yonghong Tian 0001, Wei Zeng 0006, Tiejun Huang 0001
AVSS2
2012 Multi-camera Pedestrian Detection with Multi-view Bayesian Network Model
Peixi Peng, Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
BMVC1