VLDB 2026 Research / reviewers in the wild / expert
Peng Zhai
dblp:92/4002
· DBLP profile ↗
37ranked-venue papers
3as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 1 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based DeformationabstractJoint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately, making it difficult to accurately handle occlusions and transparency. Moreover, the deformed 3DGS still suffers from visual artifacts due to the sensitivity to the topology quality of the proxy mesh. These issues pose serious obstacles to the joint use of 3DGS and meshes, making it difficult to adapt 3DGS to conventional mesh-oriented graphics pipelines. We propose UniMGS, the first unified framework for rasterizing mesh and 3DGS in a single-pass anti-aliased manner, with a novel binding strategy for 3DGS deformation based on proxy mesh. Our key insight is to blend the colors of both triangle and Gaussian fragments by anti-aliased α-blending in a single pass, achieving visually coherent results with precise handling of occlusion and transparency. To improve the visual appearance of the deformed 3DGS, our Gaussian-centric binding strategy employs a proxy mesh and spatially associates Gaussians with the mesh faces, significantly reducing rendering artifacts. With these two components, UniMGS enables the visualization and manipulation of 3D objects represented by mesh or 3DGS within a unified framework, opening up new possibilities in embodied AI, virtual reality, and gaming. We will release our source code to facilitate future research. Zeyu Xiao 0001, Yimin Cong, Dongliang Kou, Zhenyi Wu, Dingkang Yang, Peng Zhai, Lihua Zhang 0002 |
AAAI | 8 |
| 2025 | MMPF: Multi-Modal Perception Framework for Abnormal Medical Condition DetectionabstractAs the global population ages and the incidence of chronic diseases increases, the demand for early detection of abnormal medical conditions is increasing. Traditional health monitoring methods often require significant resources and specialized personnel, limiting their widespread use. Leveraging advancements in AI technologies, this study proposes a non-invasive method for detecting abnormal medical conditions from image data. A multimodal perception framework is introduced, integrating features from various modalities, including facial expressions and body postures, to enhance detection accuracy. The framework employs a Cascaded Squeeze-Excitation (CSE) module, consisting of Adaptive and Multi-modal Squeeze-Excitation components, to capture complex feature dependencies and improve cross-modal performance. Extensive experiments demonstrate the effectiveness of this approach, showing improved performance over existing methods. In addition, a new dataset that encompasses a wide range of medical conditions has been released, providing a valuable resource for future research in this domain. Chuyi Zhong, Dingkang Yang, Peng Zhai, Lihua Zhang 0002 |
AAAI | 3 |
| 2025 | Music-Driven Legged Robots: Synchronized Walking to Rhythmic BeatsabstractWe address the challenge of effectively controlling the locomotion of legged robots by incorporating precise frequency and phase characteristics, which is often ignored in locomotion policies that do not account for the periodic nature of walking. We propose a hierarchical architecture that integrates a low-level phase tracker, oscillators, and a high-level phase modulator. This controller allows quadruped robots to walk in a natural manner that is synchronized with external musical rhythms. Our method generates diverse gaits across different frequencies and achieves real-time synchronization with music in the physical world. This research establishes a foundational framework for enabling real-time execution of accurate rhythmic motions in legged robots. The video and code are available at https://music-walker.github.io/. Taixian Hou, Xiaoyi Wei, Zhiyan Dong, Jiafu Yi, Peng Zhai, Lihua Zhang 0002 |
ICRA | 6 |
| 2025 | Continuous Control of Diverse Skills in Quadruped Robots Without Complete Expert DatasetsabstractLearning diverse skills for quadruped robots presents significant challenges, such as mastering complex transitions between different skills and handling tasks of varying difficulty. Existing imitation learning methods, while successful, rely on expensive datasets to reproduce expert behaviors. Inspired by introspective learning, we propose Progressive Adversarial Self-Imitation Skill Transition (PASIST), a novel method that eliminates the need for complete expert datasets. PASIST autonomously explores and selects high-quality trajectories based on predefined target poses instead of demonstrations, leveraging the Generative Adversarial Self-Imitation Learning (GASIL) framework. To further enhance learning, We develop a skill selection module to mitigate mode collapse by balancing the weights of skills with varying levels of difficulty. Through these methods, PASIST is able to reproduce skills corresponding to the target pose while achieving smooth and natural transitions between them. Evaluations on both simulation platforms and the Solo 8 robot confirm the effectiveness of PASIST, offering an efficient alternative to expert-driven learning. Jiaxin Tu, Xiaoyi Wei, Taixian Hou, Xiaofei Gao, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002 |
ICRA | 7 |
| 2025 | $\mathrm{A}^{3} \mathrm{E}^{2}$ Net: Integrating Waypoints, Lanes, and Traces for Multi-Modal Trajectory PredictionabstractDue to increasing complexity of traffic conditions, accurately predicting future trajectories of traffic agents is essential for driving efficiency and collision avoidance, making it vital to the decision-making of autonomous vehicles. However, it presents a significant challenge in trajectory prediction owing to the complex road topology and uncertainty of drivers' intention. In this paper, we introduce$\mathrm{A}^{3} \mathrm{E}^{2} \text{Net}$, a multi-modal trajectory prediction network that integrates waypoints, lanes, and motion traces, which collectively constitute the driving environment. It employs self-attention mechanism for waypoint representation, a Graph Neural Network (GNN) for road topology, and a spatiotemporal Transformer to encode motion dynamics. Subsequently, a cross-attention module fuses context information across granularities and captures agent-environment interactions. Finally, inspired by the outstanding performance of the autoregressive model in large language models (LLMs), we propose a goal-driven autoregression-based trajectory generation module to introduce both temporal logic and latent intentions, and thus enable precise and diverse multi-modal trajectory prediction (MTP). Evaluations on the Argoverse motion forecasting benchmark highlight our method's state-of-the-art performance and computational efficiency. Xiaoyi Wei, Peng Zhai, Lihua Zhang 0002 |
ICTAI | 4 |
| 2025 | UMSD: High Realism Motion Style Transfer via Unified Mamba-based DiffusionabstractMotion style transfer is a significant research area in computer vision, enabling the rapid switching of stylistic variations for the same motion in virtual digital humans. This dramatically enhances the richness and realism of motions, making it widely applicable in multimedia contexts such as film, gaming, and the Metaverse. However, most existing methods employ a two-stream structure, which often overlooks the intrinsic relationships between content and style motions, resulting in information loss and misalignment. Additionally, these methods struggle to capture temporal dependencies in long-range motion sequences, resulting in less natural outputs. To address these limitations, we propose a Unified Motion Style Diffusion (UMSD) Framework that simultaneously extracts features from content and style motions, achieving comprehensive information interaction. We also introduce the Motion Style Mamba (MSM) denoiser, which, for the first time in motion style transfer, leverages Mamba's powerful sequence modelling capability to produce more temporally coherent stylized motion sequences. Furthermore, we design a diffusion-based content consistency loss and a style consistency loss to ensure that the framework preserves content motion while effectively learning style motion features. Extensive experiments demonstrate that our approach outperforms State-Of-The-Art (SOTA) methods qualitatively and quantitatively, achieving more realistic and coherent motion style transfer. Ziyun Qian, Zeyu Xiao 0001, Xingliang Jin, Dingkang Yang, Mingcheng Li, Zhenyi Wu, Dongliang Kou, Peng Zhai, Lihua Zhang 0002 |
ACM Multimedia | 8 |
| 2025 | Diverse object placement with dual interaction
Xianhe Cheng, Peng Zhai, Dingkang Yang, Lihua Zhang 0002 |
Neurocomputing | 2 |
| 2025 | Role play: Learning adaptive role-specific strategies in multi-agent interactions
Weifan Long, Wen Wen 0005, Peng Zhai, Lihua Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2025 | Denoising Transformer for BEV 3D Object Detection via Multiview Multiscale Cross-Attention
Xukun Zhang, Xiaoyi Wei, Zhi Xu 0010, Youxing Wang, Peng Zhai, Lihua Zhang 0002 |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2024 | Large Vision-Language Models as Emotion Recognizers in Context Awareness
Yuxuan Lei, Dingkang Yang, Zhaoyu Chen 0001, Jiawei Chen 0012, Peng Zhai, Lihua Zhang 0002 |
ACML | 5 |
| 2024 | CPR-Coach: Recognizing Composite Error Actions Based on Single-Class TrainingabstractFine- grained medical action analysis plays a vital role in improving medical skill training efficiency, but it faces the problems of data and algorithm shortage. Cardiopul-monary Resuscitation (CPR) is an essential skill in emer-gency treatment. Currently, the assessment of CPR skills mainly depends on dummies and trainers, leading to high training costs and low efficiency. For the first time, this pa-per constructs a vision-based system to complete error action recognition and skill assessment in CPR. Specifically, we define 13 types of single-error actions and 74 types of composite error actions during external cardiac compres-sion and then develop a video dataset named CPR-Coach. By taking the CPR-Coach as a benchmark, this paper in-vestigates and compares the performance of existing action recognition models based on different data modalities. To solve the unavoidable “Single-class Training & Multi-class Testing” problem, we propose a human-cognition-inspired framework named ImagineNet to improve the model's multi-error recognition performance under restricted supervision. Extensive comparison and actual deployment experiments verify the effectiveness of the framework. We hope this work could bring new inspiration to the computer vision and medical skills training communities simultaneously. The dataset and the code are publicly available on https://github.com/Shunli-Wang/CPR-Coach. Shunli Wang 0001, Shuaibing Wang, Dingkang Yang, Mingcheng Li, Haopeng Kuang, Liuzhen Su, Peng Zhai, Lihua Zhang 0002 |
CVPR | 8 |
| 2024 | Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Peng Zhai, Ke Li 0015, Lihua Zhang 0002 |
ECCV (58) | 8 |
| 2024 | STGCN-DHD: Spatio-Temporal Graph Convolutional Network for EEG-Based Driving Hazard Detection
Jialong Liang, Weifan Long, Peng Zhai, Lihua Zhang 0002 |
ICONIP (4) | 4 |
| 2024 | Multi-Task Learning of Active Fault-Tolerant Controller for Leg Failures in Quadruped robotsabstractElectric quadruped robots used in outdoor exploration are susceptible to leg-related electrical or mechanical failures. Unexpected joint power loss and joint locking can immediately pose a falling threat. Typically, controllers lack the capability to actively sense the condition of their own joints and take proactive actions. Maintaining the original motion patterns could lead to disastrous consequences, as the controller may produce irrational output within a short period of time, further creating the risk of serious physical injuries. This paper presents a hierarchical fault-tolerant control scheme employing a multi-task training architecture capable of actively perceiving and overcoming two types of leg joint faults. The architecture simultaneously trains three joint task policies for health, power loss, and locking scenarios in parallel, introducing a symmetric reflection initialization technique to ensure rapid and stable gait skill transformations. Experiments demonstrate that the control scheme is robust in unexpected scenarios where a single leg experiences concurrent joint faults in two joints. Furthermore, the policy retains the robot’s planar mobility, enabling rough velocity tracking. Finally, zero-shot Sim2Real transfer is achieved on the real-world SOLO8 robot, countering both electrical and mechanical failures. Taixian Hou, Jiaxin Tu, Xiaofei Gao, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002 |
ICRA | 5 |
| 2024 | MDRPC: Music-Driven Robot Primitives ChoreographyabstractDance has been an important art form and means of communication since the dawn of human civilization. Equipping humanoid robots with the ability to perform smooth dance movements to music is a key research priority in artificial intelligence, robotics and human-computer interaction. However, existing kinematics-based dance generation methods often violate real-world physical laws as they do not consider physical constraints, leading to unrealistic movements. Additionally, due to the diversity and dynamic variability of input music, most existing physics-based methods, which rely on task-specific reward functions, face significant challenges in effectively handling music-driven dance generation tasks. To address these issues, we introduce MDRPC, the first physics-based, music-driven dance generation method for humanoid robots. Inspired by human choreographic principles, MDRPC is defined as a two-phase framework. The initial phase utilizes adversarial imitation learning to acquire a rich set of reusable dance primitives from a music-dance dataset. In the subsequent phase, these dance primitives are orchestrated under the guidance of musical theory and choreographic rules to generate complex humanoid dance sequences. Specifically, we propose beat alignment and dance diversity reward functions to synchronize motion rhythms with music beats and enhance the diversity of dance movements. We implement MDRPC on a simulated humanoid robot, and the results confirm that our method effectively controls the humanoid, enabling it to perform dance movements harmoniously synchronized with the music. Haiyang Guan, Xiaoyi Wei, Weifan Long, Dingkang Yang, Peng Zhai, Lihua Zhang 0002 |
ICTAI | 5 |
| 2024 | PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric ApplicationsabstractDeveloping intelligent pediatric consultation systems offers promising prospects for improving diagnostic efficiency, especially in China, where healthcare resources are scarce. Despite recent advances in Large Language Models (LLMs) for Chinese medicine, their performance is sub-optimal in pediatric applications due to inadequate instruction data and vulnerable training procedures.
To address the above issues, this paper builds PedCorpus, a high-quality dataset of over 300,000 multi-task instructions from pediatric textbooks, guidelines, and knowledge graph resources to fulfil diverse diagnostic demands. Upon well-designed PedCorpus, we propose PediatricsGPT, the first Chinese pediatric LLM assistant built on a systematic and robust training pipeline.
In the continuous pre-training phase, we introduce a hybrid instruction pre-training mechanism to mitigate the internal-injected knowledge inconsistency of LLMs for medical domain adaptation. Immediately, the full-parameter Supervised Fine-Tuning (SFT) is utilized to incorporate the general medical knowledge schema into the models. After that, we devise a direct following preference optimization to enhance the generation of pediatrician-like humanistic responses. In the parameter-efficient secondary SFT phase,
a mixture of universal-specific experts strategy is presented to resolve the competency conflict between medical generalist and pediatric expertise mastery. Extensive results based on the metrics, GPT-4, and doctor evaluations on distinct downstream tasks show that PediatricsGPT consistently outperforms previous Chinese medical LLMs. The project and data will be released at https://github.com/ydk122024/PediatricsGPT. Dingkang Yang, Jinjie Wei, Dongling Xiao, Shunli Wang 0001, Mingcheng Li, Shuaibing Wang, Jiawei Chen 0012, Qingyao Xu, Ke Li 0015, Peng Zhai, Lihua Zhang 0002 |
NeurIPS | 13 |
| 2024 | Expression guided medical condition detection via the Multi-Medical Condition Image Dataset
Chuyi Zhong, Dingkang Yang, Shunli Wang 0001, Peng Zhai, Lihua Zhang 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | CPR-CLIP: Multimodal Pre-Training for Composite Error Recognition in CPR TrainingabstractThe expensive cost of the medical skill training paradigm hinders the development of medical education, which has attracted widespread attention in the intelligent signal processing community. To address the issue of composite error action recognition in Cardiopulmonary Resuscitation (CPR) training, this letter proposes a multimodal pre-training framework named CPR-CLIP based on prompt engineering. Specifically, we design three prompts to fuse multiple errors naturally on the semantic level and then align linguistic and visual features via the contrastive pre-training loss. Extensive experiments verify the effectiveness of the CPR-CLIP. Ultimately, the CPR-CLIP is encapsulated to an electronic assistant, and four doctors are recruited for evaluation. Nearly four times efficiency improvement is observed in comparative experiments, which demonstrates the practicality of the system. We hope this work brings new insights to the intelligent medical skill training and signal processing communities simultaneously. Code is available onhttps://github.com/Shunli-Wang/CPR-CLIP. Shunli Wang 0001, Dingkang Yang, Peng Zhai, Lihua Zhang 0002 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic RepresentationsabstractUnderstanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial expressions, and auditory clues. Despite the impressive advancements of previous works via attention-based paradigms, the inherent temporal asynchrony and modality heterogeneity challenges remain in multimodal sequence fusion, causing adverse performance bottlenecks. To tackle these issues, we propose a Multimodal fusion approach for learning modality-Exclusive and modality-Agnostic representations (MEA) to refine multimodal features and leverage the complementarity across distinct modalities. On the one hand, MEA introduces a predictive self-attention module to capture reliable context dynamics within modalities and reinforce unique features over the modality-exclusive spaces. On the other hand, a hierarchical cross-modal attention module is designed to explore valuable element correlations among modalities over the modality-agnostic space. Meanwhile, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, we propose a decoupled graph fusion mechanism to enhance knowledge exchange across heterogeneous modalities and learn robust multimodal representations for downstream tasks. Numerous experiments are implemented on three multimodal datasets with asynchronous sequences. Systematic analyses show the necessity of our approach. Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang 0010, Peng Zhai, Song Wang 0002, Lihua Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Context De-Confounded Emotion RecognitionabstractContext-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representations from subjects and contexts. However, a long-overlooked issue is that a context bias in existing datasets leads to a significantly unbalanced distribution of emotional states among different context scenarios. Concretely, the harmful bias is a confounder that misleads existing models to learn spurious correlations based on conventional likelihood estimation, significantly limiting the models' performance. To tackle the issue, this paper provides a causality-based perspective to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a tailored causal graph. Then, we propose a Contextual Causal Intervention Module (CCIM) based on the backdoor adjustment to de-confound the confounder and exploit the true causal effect for model training. CCIM is plug-in and model-agnostic, which improves diverse state-of-the-art approaches by considerable margins. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our CCIM and the significance of causal insight. Dingkang Yang, Zhaoyu Chen 0001, Shunli Wang 0001, Mingcheng Li, Siao Liu, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002 |
CVPR | 10 |
| 2023 | D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object DetectionabstractAlthough CNN-based and Transformer-based detectors have made impressive improvements in 3D object detection, these two network paradigms suffer from the interference of insufficient receptive field and local detail weakening, which significantly limits the feature extraction performance of the backbone. In this paper, we propose to fuse convolution and transformer, and simultaneously considering the different contributions of non-empty voxels at different positions in 3D space to object detection, it is not consistent with applying standard convolution and transformer directly on voxels. Specifically, we design a novel deformable sparse transformer to perform long-range information interaction on fine-grained local detail semantics aggregated by focal sparse convolution, termed D-Conformer. D-Conformer learns valuable voxels with position-wise in sparse space and can be applied to most voxel-based detectors as a backbone. Extensive experiments demonstrate that our method achieves satisfactory detection results and outperforms state-of-the-art 3D detection methods by a large margin. Liuzhen Su, Xukun Zhang, Dingkang Yang, Shunli Wang 0001, Peng Zhai, Lihua Zhang 0002 |
ICASSP | 7 |
| 2023 | AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionabstractDriver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE. Dingkang Yang, Zhi Xu 0010, Shunli Wang 0001, Mingcheng Li, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Yan Wang 0068, Jing Liu 0050, Peixuan Zhang, Peng Zhai, Lihua Zhang 0002 |
ICCV | 14 |
| 2023 | Anatomical-Aware Point-Voxel Network for Couinaud Segmentation in Liver CT
Xukun Zhang, Yang Liu 0007, Sharib Ali, Minghao Han, Tao Liu 0050, Peng Zhai, Zhiming Cui 0001, Peixuan Zhang, Lihua Zhang 0002 |
MICCAI (3) | 8 |
| 2023 | How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent PerceptionabstractMulti-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay, and collaboration heterogeneity. To tackle these issues, we propose \textit{How2comm}, a collaborative perception framework that seeks a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we devise a mutual information-aware communication mechanism to maximally sustain the informative features shared by collaborators. The spatial-channel filtering is adopted to perform effective feature sparsification for efficient communication. Second, we present a flow-guided delay compensation strategy to predict future characteristics from collaborators and eliminate feature misalignment due to temporal asynchrony. Ultimately, a pragmatic collaboration transformer is introduced to integrate holistic spatial semantics and temporal context clues among agents. Our framework is thoroughly evaluated on several LiDAR-based collaborative detection datasets in real-world and simulated scenarios. Comprehensive experiments demonstrate the superiority of How2comm and the effectiveness of all its vital components. The code will be released at https://github.com/ydk122024/How2comm. Dingkang Yang, Kun Yang 0010, Jing Liu 0050, Zhi Xu 0010, Rongbin Yin, Peng Zhai, Lihua Zhang 0002 |
NeurIPS | 7 |
| 2023 | Target and source modality co-reinforcement for emotion understanding from asynchronous multimodal sequences
Dingkang Yang, Yang Liu 0246, Can Huang 0002, Mingcheng Li, Kun Yang 0010, Yan Wang 0068, Peng Zhai, Lihua Zhang 0002 |
Knowl. Based Syst. | 9 |
| 2022 | Robust Adversarial Reinforcement Learning with Dissipation Inequation ConstraintabstractRobust adversarial reinforcement learning is an effective method to train agents to manage uncertain disturbance and modeling errors in real environments. However, for systems that are sensitive to disturbances or those that are difficult to stabilize, it is easier to learn a powerful adversary than establish a stable control policy. An improper strong adversary can destabilize the system, introduce biases in the sampling process, make the learning process unstable, and even reduce the robustness of the policy. In this study, we consider the problem of ensuring system stability during training in the adversarial reinforcement learning architecture. The dissipative principle of robust H-infinity control is extended to the Markov Decision Process, and robust stability constraints are obtained based on L2 gain performance in the reinforcement learning system. Thus, we propose a dissipation-inequation-constraint-based adversarial reinforcement learning architecture. This architecture ensures the stability of the system during training by imposing constraints on the normal and adversarial agents. Theoretically, this architecture can be applied to a large family of deep reinforcement learning algorithms. Results of experiments in MuJoCo and GymFc environments show that our architecture effectively improves the robustness of the controller against environmental changes and adapts to more powerful adversaries. Results of the flight experiments on a real quadcopter indicate that our method can directly deploy the policy trained in the simulation environment to the real environment, and our controller outperforms the PID controller based on hardware-in-the-loop. Both our theoretical and empirical results provide new and critical outlooks on the adversarial reinforcement learning architecture from a rigorous robust control perspective. Peng Zhai, Zhiyan Dong, Lihua Zhang 0002, Shunli Wang 0001, Dingkang Yang |
AAAI | 1 |
| 2022 | Emotion Recognition for Multiple Context Awareness
Dingkang Yang, Shunli Wang 0001, Yang Liu 0246, Peng Zhai, Liuzhen Su, Mingcheng Li, Lihua Zhang 0002 |
ECCV (37) | 5 |
| 2022 | Curriculum Adversarial Training for Robust Reinforcement LearningabstractReinforcement learning with adversarial training is currently a key method for improving the robustness of DRL. However, in adversarial training, especially for unstable or disturbance-sensitive systems, the adversary always learns the policy significantly faster than the DRL agent and thus easily generates powerful perturbations. The agent cannot effectively adapt to the overly powerful adversary, which leads to unstable training and even failure to learn the robust policy. In this work, we propose a novel adversarial training method, called Curriculum Adversarial Training, inspired by the idea of curriculum learning. The method dynamically adjusts the strength of the adversary through natural curriculum learning for progressive adversarial training. Thus, the DRL system considers to reasonable learning rules, and the agent faces a suitable learning process. Furthermore, we adopt an advanced action space perturbation method with a most attractive ability as the adversary during training. The proposed method is compared with popular baseline methods through MuJoCo tasks. Experimental results show that our method can improve the robustness of the policy significantly and adapt to uncertain environment effectively. Junru Sheng, Peng Zhai, Zhiyan Dong, Xiaoyang Kang 0001, Chixiao Chen, Lihua Zhang 0002 |
IJCNN | 2 |
| 2022 | CA-SpaceNet: Counterfactual Analysis for 6D Pose Estimation in SpaceabstractReliable and stable 6D pose estimation of un-cooperative space objects plays an essential role in on-orbit servicing and debris removal missions. Considering that the pose estimator is sensitive to background interference, this paper proposes a counterfactual analysis framework named CA-SpaceNet to complete robust 6D pose estimation of the space-borne targets under complicated background. Specifically, conventional methods are adopted to extract the features of the whole image in the factual case. In the counterfactual case, a non-existent image without the target but only the background is imagined. Side effect caused by background interference is reduced by counterfactual analysis, which leads to unbiased prediction in final results. In addition, we also carry out low-bit-width quantization for CA-SpaceNet and deploy part of the framework to a Processing-In-Memory (PIM) accelerator on FPGA. Qualitative and quantitative results demonstrate the effectiveness and efficiency of our proposed method. To our best knowledge, this paper applies causal inference and network quantization to the 6D pose estimation of space-borne targets for the first time. The code is available at https://github.com/Shunli-Wang/CA-SpaceNet. Shunli Wang 0001, Shuaibing Wang, Bo Jiao 0003, Dingkang Yang, Liuzhen Su, Peng Zhai, Chixiao Chen, Lihua Zhang 0002 |
IROS | 6 |
| 2021 | Learning Associative Representation for Facial Expression RecognitionabstractThe main inherent challenges with the Facial Expression Recognition (FER) are high intra-class variations and high inter-class similarities, while existing methods pay little attention to the association within inter- and intra-class expressions. This paper introduces a novel Expression Associative Network (EAN) to learn association of facial expression, specifically, from two aspects: 1) associative topological relation over mini-batch is constructed by similarity matrix with an adjacent regularization, and 2) learning association of expressions with Graph Convolutional Network (GCN). Besides, an auxiliary module as invariant feature generator based on Generative Adversarial Networks (GAN) is designed to suppress pose variations, illumination changes, and occlusions. Results on public benchmarks achieve comparable or better performance compared with current state-of-the-art methods, with 90.07% on FERPlus, 86.36% on RAF-DB, and improve by 3.92% over SOTA on synthetic wrong labeling datasets. Yangtao Du, Dingkang Yang, Peng Zhai, Lihua Zhang 0002 |
ICIP | 3 |
| 2021 | TSA-Net: Tube Self-Attention Network for Action Quality AssessmentabstractIn recent years, assessing action quality from videos has attracted growing attention in computer vision community and human-computer interaction. Most existing approaches usually tackle this problem by directly migrating the model from action recognition tasks, which ignores the intrinsic differences within the feature map such as foreground and background information. To address this issue, we propose a Tube Self-Attention Network (TSA-Net) for action quality assessment (AQA). Specifically, we introduce a single object tracker into AQA and propose the Tube Self-Attention Module (TSA), which can efficiently generate rich spatio-temporal contextual information by adopting sparse feature interactions. The TSA module is embedded in existing video networks to form TSA-Net. Overall, our TSA-Net is with the following merits: 1) High computational efficiency, 2) High flexibility, and 3) The state-of-the-art performance. Extensive experiments are conducted on popular action quality assessment datasets including AQA-7 and MTL-AQA. Besides, a dataset named Fall Recognition in Figure Skating (FR-FS) is proposed to explore the basic action assessment in the figure skating scene. Our TSA-Net achieves the Spearman's Rank Correlation of 0.8476 and 0.9393 on AQA-7 and MTL-AQA, respectively, which are the new state-of-the-art results. The results on FR-FS also verify the effectiveness of the TSA-Net. The code and FR-FS dataset are publicly available at https://github.com/Shunli-Wang/TSA-Net. Shunli Wang 0001, Dingkang Yang, Peng Zhai, Chixiao Chen, Lihua Zhang 0002 |
ACM Multimedia | 3 |
| 2020 | An Efficient Authentication Protocol for Wireless Mesh Networks "In Prepress"
Peng Zhai, Jingsha He, Nafei Zhu |
J. Web Eng. | 1 |
| 2020 | Dynamic Evaluation of Recommendation Trust in Open Networks "In Prepress"
Guangmin Sun, Peng Zhai, Yuge Sun |
J. Web Eng. | 4 |
| 2018 | An underwater acoustic direct sequence spread spectrum communication system using dual spread spectrum codeabstractWith the goal of achieving high stability and reliability to support underwater point-to-point communications and code division multiple access (CDMA) based underwater networks, a direct sequence spread spectrum based underwater acoustic communication system using dual spread spectrum code is proposed. To solve the contradictions between the information data rate and the accuracy of Doppler estimation, channel estimation, and frame synchronization, a data frame structure based on dual spread spectrum code is designed. A long spread spectrum code is used as the training sequence, which can be used for data frame detection and synchronization, Doppler estimation, and channel estimation. A short spread spectrum code is used to modulate the effective information data. A delay cross-correlation algorithm is used for Doppler estimation, and a correlation algorithm is used for channel estimation. For underwater networking, each user is assigned a different pair of spread spectrum codes. Simulation results show that the system has a good anti-multipath, anti-interference, and anti-Doppler performance, the bit error rate can be smaller than 10 −6 when the signal-to-noise ratio is larger than −10 dB, the data rate can be as high as 355 bits/s, and the system can be used in the downlink of CDMA based networks. Jian-fen Li, Peng Zhai, Jiucai Jin, Zhichao Lv |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2017 | MetaComp: comprehensive analysis software for comparative meta-omics including comparative metagenomicsabstractBACKGROUND: During the past decade, the development of high throughput nucleic sequencing and mass spectrometry analysis techniques have enabled the characterization of microbial communities through metagenomics, metatranscriptomics, metaproteomics and metabolomics data. To reveal the diversity of microbial communities and interactions between living conditions and microbes, it is necessary to introduce comparative analysis based upon integration of all four types of data mentioned above. Comparative meta-omics, especially comparative metageomics, has been established as a routine process to highlight the significant differences in taxon composition and functional gene abundance among microbiota samples. Meanwhile, biologists are increasingly concerning about the correlations between meta-omics features and environmental factors, which may further decipher the adaptation strategy of a microbial community. RESULTS: We developed a graphical comprehensive analysis software named MetaComp comprising a series of statistical analysis approaches with visualized results for metagenomics and other meta-omics data comparison. This software is capable to read files generated by a variety of upstream programs. After data loading, analyses such as multivariate statistics, hypothesis testing of two-sample, multi-sample as well as two-group sample and a novel function-regression analysis of environmental factors are offered. Here, regression analysis regards meta-omic features as independent variable and environmental factors as dependent variables. Moreover, MetaComp is capable to automatically choose an appropriate two-group sample test based upon the traits of input abundance profiles. We further evaluate the performance of its choice, and exhibit applications for metagenomics, metaproteomics and metabolomics samples. CONCLUSION: MetaComp, an integrative software capable for applying to all meta-omics data, originally distills the influence of living environment on microbial community by regression analysis. Moreover, since the automatically chosen two-group sample test is verified to be outperformed, MetaComp is friendly to users without adequate statistical training. These improvements are aiming to overcome the new challenges under big data era for all meta-omics data. MetaComp is available at: http://cqb.pku.edu.cn/ZhuLab/MetaComp/ and https://github.com/pzhaipku/MetaComp/ . Peng Zhai, Longshu Yang, Huaiqiu Zhu |
BMC Bioinform. | 1 |
| 2015 | Crucial Data Selection Based on Random Weight Neural NetworkabstractTraining time is an important consideration in classification practices. Various algorithms like SVM usually suffer from high computational cost during the training process. In this paper, we proposed a hybrid classification algorithm, by combining RNN and SVM together. RNN is a rapid neural network with randomly generated input weights, which performs as a fast, light weighted data selector. Selected data by RNN are then used to train an alternative SVM. The definition of crucial data and corresponding selection method is defined for RNN in this paper. We then proposed an intuitive parameter selecting method, so that to use only one parameter in the training process to determine the crucial data margin. Experimental results conducted on artificially generated databases and several public benchmarks show that the proposed method does retrieve equivalent dataset, with significant reduction of the sample number. Hongcheng Jiang, Bin Zhao 0005, Peng Zhai |
SMC | 4 |
| 2004 | Segmentation and recognition of multi-attribute motion sequencesabstractIn this work, we focus on fast and efficient recognition of motions in multi-attribute continuous motion sequences. 3D motion capture data, animation motion data, and sensor data from gesture sensing devices are examples of multi-attribute continuous motion sequences. These sequences have multiple attributes rather than only one attribute as time series data has. Motions can have different rates and durations, and the resulting data can thus have different lengths. Also, motion data can have noises due to transitions between successive motions. Hence, traditional distance measuring approaches used for time series data (such as Euclidean distances or dynamic time-warped distances) are not suitable for recognition in multi-attribute motion sequences. Hence, we have defined a similarity measure based on the analysis of singular value decomposition (SVD) properties of similar multi-attribute motions. A five-phase algorithm has then been proposed that gives good pruning power by exploiting the proximity of continuous motion data. We experimented this algorithm with data from different sources: 3D motion capture devices, animation motions, and CyberGlove gesture sensing device. These experiments show that our algorithm can segment and recognize long motion streams with high accuracy and in real time without knowing beforehand the number of motions in a stream. Chuanjun Li, Peng Zhai, Si-Qing Zheng, B. Prabhakaran 0001 |
ACM Multimedia | 2 |