VLDB 2026 Research / reviewers in the wild / expert
Mengyuan Liu 0001
dblp:143/0160-1
· DBLP profile ↗
77ranked-venue papers
15as first author
66since 2021 · last 2026
0000-0002-6332-8316ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 9 first-author · 43 since 2021Artificial intelligence and machine learning · 35 · 6 first-author · 32 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Masked Clustering Prediction for Unsupervised Point Cloud Pre-trainingabstractVision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We propose MaskClu, a novel unsupervised pre-training method for ViTs on 3D point clouds that integrates masked point modeling with clustering-based learning. MaskClu is designed to reconstruct both cluster assignments and cluster centers from masked point clouds, thus encouraging the model to capture dense semantic information. Additionally, we introduce a global contrastive learning mechanism that enhances instance-level feature learning by contrasting different masked views of the same point cloud. By jointly optimizing these complementary objectives, i.e., dense semantic reconstruction, and instance-level contrastive learning. MaskClu enables ViTs to learn richer and more semantically meaningful representations from 3D point clouds. We validate the effectiveness of MaskClu via multiple 3D tasks, including part segmentation, semantic segmentation, object detection, and classification, setting new competitive results. Bin Ren 0005, Xiaoshui Huang, Mengyuan Liu 0001, Hong Liu 0008, Fabio Poiesi, Nicu Sebe, Guofeng Mei |
AAAI | 3 |
| 2026 | MP1: MeanFlow Tames Policy Learning in 1-step for Robotic ManipulationabstractIn robot manipulation, robot learning has become a prevailing approach. However, generative models within this field face a fundamental trade-off between the slow, iterative sampling of diffusion models and the architectural constraints of faster Flow-based methods, which often rely on explicit consistency losses. To address these limitations, we introduce MP1, which pairs 3D point-cloud inputs with the MeanFlow paradigm to generate action trajectories in one network function evaluation (1-NFE). By directly learning the interval-averaged velocity via the "MeanFlow Identity", our policy avoids any additional consistency constraints. This formulation eliminates numerical ODE-solver errors during inference, yielding more precise trajectories. MP1 further incorporates CFG for improved trajectory controllability while retaining 1-NFE inference without reintroducing structural constraints. Because subtle scene-context variations are critical for robot learning, especially in few-shot learning, we introduce a lightweight Dispersive Loss that repels state embeddings during training, boosting generalization without slowing inference. We validate our method on the Adroit and Meta-World benchmarks, as well as in real-world scenarios. Experimental results show MP1 achieves superior average task success rates, outperforming DP3 by 10.2% and FlowPolicy by 7.3%. Its average inference time is only 6.8 ms—19 times faster than DP3 and nearly 2 times faster than FlowPolicy. Juyi Sheng, Peiming Li, Mengyuan Liu 0001 |
AAAI | 4 |
| 2026 | Point-In-Context: Understanding Point Cloud via In-Context Learning
Mengyuan Liu 0001, Zhongbin Fang, Xia Li 0005, Joachim M. Buhmann, Deheng Ye, Xiangtai Li, Chen Change Loy |
Int. J. Comput. Vis. | 1 |
| 2026 | H2OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose TransformersabstractTransformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper, we present a hierarchical plug-and-play pruning-and-recovering framework, calledHierarchicalHourglassTokenizer (H2OT), for efficient transformer-based 3D human pose estimation from videos. H2OT begins with progressively pruning pose tokens of redundant frames and ends with recovering full-length sequences, resulting in a few pose tokens in the intermediate transformer blocks and thus improving the model efficiency. It works with two key modules, namely, a Token Pruning Module (TPM) and a Token Recovering Module (TRM). TPM dynamically selects a few representative tokens to eliminate the redundancy of video frames, while TRM restores the detailed spatio-temporal information based on the selected tokens, thereby expanding the network output to the original full-length temporal resolution for fast inference. Our method is general-purpose: it can be easily incorporated into common VPT models on bothseq2seqandseq2framepipelines while effectively accommodating different token pruning and recovery strategies. In addition, our H2OT reveals that maintaining the full pose sequence is unnecessary, and a few pose tokens of representative frames can achieve both high efficiency and estimation accuracy. Extensive experiments on multiple benchmark datasets demonstrate both the effectiveness and efficiency of the proposed method. Code and models are available athttps://github.com/NationalGAILab/HoT. Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Pichao Wang, Shijian Lu, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Heatmap Pooling Network for Action Recognition From RGB VideosabstractHuman action recognition (HAR) in videos has garnered widespread attention due to the rich information in RGB videos. Nevertheless, existing methods for extracting deep features from RGB videos face challenges such as information redundancy, susceptibility to noise and high storage costs. To address these issues and fully harness the useful information in videos, we propose a novel heatmap pooling network (HP-Net) for action recognition from videos, which extracts information-rich, robust and concise pooled features of the human body in videos through a feedback pooling module. The extracted pooled features demonstrate obvious performance advantages over the previously obtained pose data and heatmap features from videos. In addition, we design a spatial-motion co-learning module and a text refinement modulation module to integrate the extracted pooled features with other multimodal data, enabling more robust action recognition. Extensive experiments on several benchmarks namely NTU RGB+D 60, NTU RGB+D 120, Toyota-Smarthome and uncrewed aerial vehicles (UAV)-Human consistently verify the effectiveness of our HP-Net, which outperforms the existing human action recognition methods. Mengyuan Liu 0001, Yongkang Jiang, Bin He 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | SARL: Structure-aware representation learning for 3D human pose estimation from point clouds
Yong Wang 0053, Mengyuan Liu 0001 |
Pattern Recognit. | 4 |
| 2026 | Adaptive learning from noisy estimated depth maps benefits monocular RGB-based 3D human pose estimation
Mengyuan Liu 0001, Jingting Liu |
Pattern Recognit. | 1 |
| 2026 | MXPose: Multiplex interactive learning for multi-view 3D Human Pose Estimation
Wanruo Zhang, Zhou Guan, Mengyuan Liu 0001, Bin Ren 0005, Hong Liu 0008 |
Pattern Recognit. | 3 |
| 2026 | Improving Fine-Grained Understanding for Retrieval in Human Motion and TextabstractThis work focuses on human motion-text retrieval (MTR), a task recently proposed for motion understanding. Unlike traditional visual-text retrieval, human motion can be understood as the superposition of numerous atomic actions, and its description is also limited to human-centered themes. Considering this characteristic, directly mapping similar samples into a joint embedding space and conducting naive contrastive training is suboptimal, as it lacks cognition of fine-grained human language descriptions and fails to alleviate semantic conflicts between similar samples. To address this, we propose a meticulous Cross-perceptual Salience Mapping, highlighting fine-grained poses or words to provide more accurate similarity measurement. Additionally, a novel Drop-then-Contrast scheme is designed for MTR, discarding false negative samples from the negative set and mining the remaining sample for contrastive training to reduce violations they caused. Our framework, termed as improving fine-grained understanding for Retrieval in Human Motion and Text or Rehamot for short, outperforms previous works by a recall of 58.6% and 56.5% on HumanML3D and KIT-ML respectively (motion retrieval, R@10). Yong Wang 0053, Hongchang Jin, Mengyuan Liu 0001 |
IEEE Signal Process. Lett. | 5 |
| 2026 | PePNet: Pose-Enhanced Point Cloud Network for LiDAR-Based Human Action Recognition in Outdoor Long-Range ScenariosabstractWith potential applications in robotics and autonomous vehicles, LiDAR-based human action recognition (HAR) in outdoor long-range scenarios is challenging due to the degradation of point cloud density with distance and the simultaneous motion of humans and sensors. To address these issues, we propose the Pose-Enhanced Point Cloud Network (PePNet), a distance-aware framework for long-range HAR. As the core component, the Pose-Enhanced Point Cloud Block (PeP Block) integrates three modules: a Dynamic Enhancement Module that mitigates point cloud sparsity at long distances by generating supplementary points from motion cues, a Pose Prompter Module that introduces pose priors, and an Adaptive Point Selection Module that suppresses irrelevant body-part movements. We further design a Spatiotemporal Tube Embedding (ST-Tube), combined with the Mamba state space model, to capture long-range dependencies and complex motion dynamics. In addition, we construct Momo, a large-scale LiDAR-based HAR dataset that focuses on long-range (2-30 m) outdoor scenarios where sparse point clouds and simultaneous human-sensor motion pose prominent challenges, complementing existing benchmarks by providing a dedicated evaluation platform for long-range outdoor HAR. Experimental results show that PePNet achieves consistent performance gains over existing methods on Momo. Moreover, the proposed PeP Block can serve as a plug-and-play module to enhance other point cloud action recognition frameworks in long-range outdoor settings. The code is available at https://github.com/Shark0-0/PePNet. Mengyuan Liu 0001, Zhichao Deng, Peiming Li, Jun Liu 0036 |
IEEE Trans. Image Process. | 1 |
| 2026 | Adaptive Interaction Network for Human Motion Prediction During Human-Robot CollaborationabstractHuman motion prediction during human-robot collaboration is critical for achieving safe and efficient interactions in shared environments. Unlike previous studies that have primarily focused on humans and objects, we focus on the robot-aware human motion prediction task, which explicitly models the influence of robots on human motion. This task presents new challenges, such as capturing human-robot heterogeneity and modeling complex spatial interactions. To address these issues, we develop an adaptive interaction network (AINet) model that consists of two branches: a primary branch that predicts future human motion and an auxiliary branch that estimates robot trajectories. The two branches are jointly optimized and coupled via a Local-Global Spatial Interaction (LGSI) Module, which effectively captures fine-grained and global contextual dependencies between human and robot motion sequences. Then, we introduce an Adaptive Weighted Aggregation (AWA) Module to dynamically fuse motion features using context-dependent weights, thereby increasing adaptability across diverse scenarios. Furthermore, we adopt a coarse-to-fine prediction strategy, in which a coarse pseudo-trajectory is first predicted, followed by the refinement of detailed human poses based on that trajectory. Extensive experiments on three challenging datasets reveal that our proposed method achieves superior performance, validating its effectiveness in human-robot collaboration scenarios. Our code is available at: https://github.com/LyTingHub/AINet. Mengyuan Liu 0001, Yangting Lin, Qiongjie Cui |
IEEE Trans. Image Process. | 1 |
| 2026 | Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action RecognitionabstractRGB camera-based surveillance systems enable human action recognition for public safety and healthcare, yet raise serious privacy concerns. Existing methods rely on post-capture algorithms, which fail to protect privacy during data acquisition. We propose Lens Privacy Sealing (LPS), a simple hardware solution that physically obscures camera lenses with adjustable laminating film, providing pre-sensor privacy protection at minimal cost. Unlike software methods or expensive engineered optics, LPS achieves strong privacy through stochastic multi-layer scattering that is physically irreversible. We introduce the P3AR dataset for privacy-preserving action recognition, featuring both large-scale replay-captured (P3AR-NTU, 114K videos) and real-world collected (P3AR-PKU) subsets with privacy attribute annotations. To handle video degradation from LPS, we propose MSPNet, a single-stage framework incorporating Inter-Frame Noise Suppressor (IFNS) and Cross-Frame Semantic Aggregator (CFSA), enhanced by contrastive language-image pre-training for robust semantic extraction. Extensive experiments demonstrate that MSPNet with IFNS and CFSA nearly doubles action recognition accuracy compared to baseline methods while suppressing identity recognition to low levels. Comprehensive validation shows LPS achieves a superior privacy-utility trade-off compared to state-of-the-art hardware methods, resists reconstruction attacks including PSF inversion and data-driven recovery, and generalizes robustly across optical configurations and challenging environments. Code is available at https://github.com/wangzy01/MSPNet. Mengyuan Liu 0001, Peiming Li, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Double-Chain Graph Convolution Transformer for 3D Human Pose EstimationabstractReconstructing 3D poses from 2D poses lacking depth information is particularly challenging due to the complexity and diversity of human motion. The key is to effectively model the spatial constraints between joints to leverage their inherent dependencies. Thus, we propose a novel model, called Double-chain Graph Convolution Transformer (DC-GCT), to constrain the pose through a double-chain design consisting of local-to-global and global-to-local chains to obtain a complex representation more suitable for the current human pose. Specifically, we combine the advantages of GCN and Transformer and design a Local Constraint Module (LCM) based on GCN and a Global Constraint Module (GCM) based on self-attention mechanism as well as a Feature Interaction Module (FIM). The proposed method fully captures the multi-level dependencies between human body joints to optimize the modeling capability of the model. Moreover, we propose a method to use temporal information into the single-frame model by guiding the video sequence embedding through the joint embedding of the target frame, with negligible increase in computational cost. Experimental results demonstrate that DC-GCT achieves state-of-the-art performance on two challenging datasets (Human3.6 M and MPI-INF-3DHP). Notably, our model achieves state-of-the-art performance on all action categories in the Human3.6 M dataset using detected 2D poses from CPN, and our code is available at:https://github.com/KHB1698/DC-GCT. Hongbo Kang, Yong Wang 0053, Mengyuan Liu 0001, Doudou Wu, Wenming Yang |
IEEE Trans. Multim. | 3 |
| 2026 | Uncertainty-Aware Testing-Time Optimization for 3D Human Pose EstimationabstractAlthough data-driven methods have achieved success in 3D human pose estimation, they often suffer from domain gaps and exhibit limited generalization. In contrast, optimization-based methods excel in fine-tuning for specific cases but are generally inferior to data-driven methods in overall performance. We observe that previous optimization-based methods commonly rely on projection constraint, which only ensures alignment in 2D space, potentially leading to the overfitting problem. To address this, we propose an Uncertainty-Aware testing-time Optimization (UAO) framework, which keeps the prior information of pre-trained model and alleviates the overfitting problem using the uncertainty of joints. Specifically, during the training phase, we design an effective 2D-to-3D network for estimating the corresponding 3D pose while quantifying the uncertainty of each 3D joint. For optimization during testing, the proposed optimization framework freezes the pre-trained model and optimizes only a latent state. Projection loss is then employed to ensure the generated poses are well aligned in 2D space for high-quality optimization. Furthermore, we utilize the uncertainty of each joint to determine how much each joint is allowed for optimization. The effectiveness and superiority of the proposed framework are validated through extensive experiments on challenging datasets: Human3.6M, MPI-INF-3DHP, and 3DPW. Notably, our approach outperforms the previous best result by a large margin of 5.5% on Human3.6M. Ti Wang, Mengyuan Liu 0001, Hong Liu 0008, Bin Ren 0005, Yingxuan You, Wenhao Li 0002, Nicu Sebe, Xia Li 0005 |
IEEE Trans. Multim. | 2 |
| 2026 | STAR: Skeletal Token Alignment and Rearrangement for Interaction RecognitionabstractUnderstanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision. While most existing methods rely on skeleton sequences-effective in low-light and privacy-sensitive environments-they face two major challenges: 1) learning and effectively exploiting interaction cues from skeletal data, and 2) compensating for the lack of visual information absent in skeletons alone. To address these challenges, we propose skeletal token alignment and rearrangement (STAR) for human-robot and human-human interaction recognition. It learns interaction-specific skeleton features and enriches them using visual cues by aligning skeleton and RGB video representations in a shared latent space. Specifically, STAR consists of three key components. First, we design a skeleton encoder that captures fine-grained interdependencies using Entity Rearrangement (ER) and Interactive Spatiotemporal Tokens (ISTs). Second, we present Visual Interaction Encoding that introduces a Focus on Interactions (FoI) strategy to attend to spatiotemporal regions relevant to interactions in RGB videos. Finally, these representations are aligned via a contrastive learning objective, with a refinement head further refines predictions. During training, STAR leverages both skeleton and RGB video data to learn robust, discriminative interaction representations. At inference time, it operates on skeletons alone, retaining visual-informed benefits while preserving skeleton-only efficiency. Extensive experiments on Chico, HARPER, NTU Mutual 11 and 26 datasets consistently validate our approach by demonstrating superior performance over state-of-the-art methods. Our code is publicly available athttps://github.com/Necolizer/STAR. Yuhang Wen 0001, Mengyuan Liu 0001, Zixuan Tang, Junsong Yuan 0001, Beichen Ding |
IEEE Trans. Multim. | 2 |
| 2025 | Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language AlignmentabstractLearning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it difficult to compute the similarity between these two modalities accurately and efficiently. In this paper, we propose a novel framework called Asymmetric Visual Semantic Embedding (AVSE) to dynamically select features from various regions of images tailored to different textual inputs for similarity calculation. To capture information from different views in the image, we design a radial bias sampling module to sample image patches and obtain image features from various views, Furthermore, AVSE introduces a novel module for efficient computation of visual semantic similarity between asymmetric image and text embeddings. Central to this module is the presumption of foundational semantic units within the embeddings, denoted as ``meta-semantic embeddings." It segments all embeddings into meta-semantic embeddings with the same dimension and calculates visual semantic similarity by finding the optimal match of meta-semantic embeddings of two modalities. Our proposed AVSE model is extensively evaluated on the large-scale MS-COCO and Flickr30K datasets, demonstrating its superiority over recent state-of-the-art methods. Yang Liu 0264, Mengyuan Liu 0001, Shudong Huang, Jiancheng Lv 0001 |
AAAI | 2 |
| 2025 | Recognizing Actions From Robotic View for Natural Human-Robot Interaction
Peiming Li, Hong Liu 0008, Zhichao Deng, Can Wang 0006, Jun Liu 0036, Junsong Yuan 0001, Mengyuan Liu 0001 |
ICCV | 8 |
| 2025 | Recognizing Skeleton-Based Actions As PointsabstractRecent advances in skeleton-based action recognition have been primarily driven by Graph Convolutional Networks (GCNs) and skeleton transformers. While conventional approaches focus on modeling joint co-occurrences through skeletal connections, they overlook the inherent positional information in 3D coordinates. Although Hyper-Graphs partially address the limitation of pairwise aggregation in capturing higher-order kinematic dependencies, challenges remain in their topological definitions. To solve these problems, this paper proposes a skeleton-to-point network (Skeleton2Point) to model joints’ position relationships directly in three-dimensional space without fixed topology limitation, which is the first to regard skeleton recognition as point clouds. However, simply considering the raw 3D coordinates would result in the loss of the anatomical identity of each keypoint and its temporal position in the sequence. To address this limitation, we augment the three-dimensional spatial coordinates with two additional dimensions: the anatomical index of each keypoint and its corresponding frame number with a proposed Information Transform Module (ITM). This transformation extends the representation from a three-dimensional to a five-dimensional feature space. Furthermore, we propose a Cluster-Dispatch-Based Interaction module (CDI) to enhance the discrimination of local-global information. In comparison with existing methods on NTU-RGB+D 60 and NTU-RGB+D 120 datasets, Skeleton2Point has demonstrated state-of-the-art performance on both joint modality and stream fusion. Especially, on the challenging NTU-RGB+D 120 dataset under the X-Sub and X-Set setting, the accuracies reach 90.63% and 91.92%. Baiqiao Yin, Jiajun Wen 0003, Mengyuan Liu 0001 |
IROS | 7 |
| 2025 | GraphMLP: A graph MLP-like architecture for 3D human pose estimation
Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Tianyu Guo 0001, Ti Wang, Hao Tang 0005, Nicu Sebe |
Pattern Recognit. | 2 |
| 2025 | Frequency-Aware Self-Supervised Group Activity Recognition with skeleton sequences
Mengyuan Liu 0001, Hong Liu 0008, Jinyan Zhang, Peini Guo, Ruijia Fan |
Pattern Recognit. | 2 |
| 2025 | Learning Mutual Excitation for Hand-to-Hand and Human-to-Human Interaction RecognitionabstractRecognizing interactive actions, including hand-to-hand interaction and human-to-human interaction, has attracted increasing attention for various applications in the field of video analysis and human–robot interaction. Considering the success of graph convolution in modeling topology-aware features from skeleton data, recent methods commonly operate graph convolution on separate entities and use late fusion for interactive action recognition, which can barely model the mutual semantic relationships between pairwise entities. To this end, we propose a mutual excitation graph convolutional network (me-GCN) by stacking mutual excitation graph convolution (me-GC) layers. Specifically, me-GC uses a mutual topology excitation module to firstly extract adjacency matrices from individual entities and then adaptively model the mutual constraints between them. Moreover, me-GC extends the above idea and further uses a mutual feature excitation module to extract and merge deep features from pairwise entities. Compared with graph convolution, our proposed me-GC gradually learns mutual information in each layer and each stage of graph convolution operations. Extensive experiments on a challenging hand-to-hand interaction dataset, i.e., the Assembely101 dataset, and two large-scale human-to-human interaction datasets, i.e., NTU60-Interaction and NTU120-Interaction consistently verify the superiority of our proposed method, which outperforms the state-of-the-art GCN-based and Transformer-based methods. Mengyuan Liu 0001, Chen Chen 0015, Songtao Wu, Fanyang Meng, Hong Liu 0008 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2025 | NanoHTNet: Nano Human Topology Network for Efficient 3D Human Pose EstimationabstractThe widespread application of 3D human pose estimation (HPE) is limited by resource-constrained edge devices like Jetson Nano, requiring more efficient models. A key approach to enhancing efficiency involves designing networks based on the structural characteristics of input data. However, effectively utilizing the structural priors in human skeletal inputs remains challenging. To address this, we leverage both explicit and implicit spatio-temporal priors of the human body through innovative model design and a pre-training proxy task. First, we propose a Nano Human Topology Network (NanoHTNet), a tiny 3D HPE network with stacked Hierarchical Mixers to capture explicit features. Specifically, the spatial Hierarchical Mixer efficiently learns the human physical topology across multiple semantic levels, while the temporal Hierarchical Mixer with discrete cosine transform and low-pass filtering captures local instantaneous movements and global action coherence. Moreover, Efficient Temporal-Spatial Tokenization (ETST) is introduced to enhance spatio-temporal interaction and reduce computational complexity significantly. Second, PoseCLR is proposed as a general pre-training method based on contrastive learning for 3D HPE, aimed at extracting implicit representations of human topology. By aligning 2D poses from diverse viewpoints in the proxy task, PoseCLR aids 3D HPE encoders like NanoHTNet in more effectively capturing the high-dimensional features of the human body, leading to further performance improvements. Extensive experiments verify that NanoHTNet with PoseCLR outperforms other state-of-the-art methods in efficiency, making it ideal for deployment on edge devices like the Jetson Nano. Code and models are available at https://github.com/vefalun/NanoHTNet. Jialun Cai, Mengyuan Liu 0001, Hong Liu 0008, Shuheng Zhou 0001, Wenhao Li 0002 |
IEEE Trans. Image Process. | 2 |
| 2025 | HYRE: Hybrid Regressor for 3D Human Pose and Shape EstimationabstractRegression-based 3D human pose and shape estimation often fall into one of two different paradigms. Parametric approaches, which regress the parameters of a human body model, tend to produce physically plausible but image-mesh misalignment results. In contrast, non-parametric approaches directly regress human mesh vertices, resulting in pixel-aligned but unreasonable predictions. In this paper, we consider these two paradigms together for a better overall estimation. To this end, we propose a novel HYbrid REgressor (HYRE) that greatly benefits from the joint learning of both paradigms. The core of our HYRE is a hybrid intermediary across paradigms that provides complementary clues to each paradigm at the shared feature level and fuses their results at the part-based decision level, thereby bridging the gap between the two. We demonstrate the effectiveness of the proposed method through both quantitative and qualitative experimental analyses, resulting in improvements for each approach and ultimately leading to better hybrid results. Our experiments show that HYRE outperforms previous methods on challenging 3D human pose and shape benchmarks. Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Bin Ren 0005, Xia Li 0005, Yingxuan You, Nicu Sebe |
IEEE Trans. Image Process. | 2 |
| 2025 | PoseMoE: Mixture-of-Experts Network for Monocular 3D Human Pose EstimationabstractThe lifting-based methods have dominated monocular 3D human pose estimation by leveraging detected 2D poses as intermediate representations. The 2D component of the final 3D human pose benefits from the detected 2D poses, whereas its depth counterpart must be estimated from scratch. The lifting-based methods encode the detected 2D pose and unknown depth in an entangled feature space, explicitly introducing depth uncertainty to the detected 2D pose, thereby limiting overall estimation accuracy. This work reveals that the depth representation is pivotal for the estimation process. Specifically, when depth is in an initial, completely unknown state, jointly encoding depth features with 2D pose features is detrimental to the estimation process. In contrast, when depth is initially refined to a more dependable state via network-based estimation, encoding it together with 2D pose information is beneficial. To address this limitation, we present a Mixture-of-Experts network for monocular 3D pose estimation named PoseMoE. Our approach introduces: 1) A mixture-of-experts network where specialized expert modules refine the well-detected 2D pose features and learn the depth features. This mixture-of-experts design disentangles the feature encoding process for 2D pose and depth, therefore reducing the explicit influence of uncertain depth features on 2D pose features. 2) A cross-expert knowledge aggregation module is proposed to aggregate cross-expert spatio-temporal contextual information. This step enhances features through bidirectional mapping between 2D pose and depth. Extensive experiments show that our proposed PoseMoE outperforms the conventional lifting-based methods on three widely used datasets: Human3.6M, MPI-INF-3DHP, and 3DPW. Mengyuan Liu 0001, Jinyan Zhang, Wenhao Li 0002, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-supervised Learning
Bin Ren 0005, Guofeng Mei, Danda Pani Paudel, Weijie Wang 0002, Yawei Li 0001, Mengyuan Liu 0001, Rita Cucchiara, Luc Van Gool, Nicu Sebe |
ACCV (7) | 6 |
| 2024 | Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationabstractTransformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper, we present a plug-and-play pruning-and-recovering framework, called Hourglass Tokenizer (HoT), for efficient transformer-based 3D human pose estimation from videos. Our HoT begins with pruning pose tokens of re-dundant frames and ends with recovering full-length tokens, resulting in a few pose tokens in the intermediate transformer blocks and thus improving the model efficiency. To effectively achieve this, we propose a token pruning cluster (TPC) that dynamically selects a few representative tokens with high semantic diversity while eliminating the redundancy of video frames. In addition, we develop a token recovering attention (TRA) to restore the detailed spatio-temporal information based on the selected tokens, thereby expanding the network output to the original full-length temporal resolution for fast inference. Extensive experiments on two benchmark datasets (i.e., Human3.6M and MPI-INF-3DHP) demonstrate that our method can achieve both high efficiency and estimation accuracy compared to the original VPT models. For instance, applying to MotionBERT and MixSTE on Hu-man3.6M, our HoT can save nearly 50% FLOPs without sacrificing accuracy and nearly 40% FLOPs with only 0.2% accuracy drop, respectively. Code and models are available at https://github.com/NationalGAILab/HoT. Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Pichao Wang, Jialun Cai, Nicu Sebe |
CVPR | 2 |
| 2024 | Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context LearningabstractIn-context learning provides a new perspective for multi-task modeling for vision and NLP. Under this setting, the model can perceive tasks from prompts and accomplish them without any extra task-specific head predictions or model fine-tuning. However, skeleton sequence modeling via in-context learning remains unexplored. Directly applying existing in-context models from other areas onto skeleton sequences fails due to the similarity between inter-frame and cross-task poses, which makes it exceptionally hard to perceive the task correctly from a subtle context. To address this challenge, we propose Skeleton-in-Context (SiC), an effective framework for in-context skeleton sequence modeling. Our SiC is able to handle multiple skeleton-based tasks simultaneously after a single training process and accomplish each task from context according to the given prompt. It can further generalize to new, unseen tasks according to customized prompts. To facilitate context perception, we additionally propose a task-unified prompt, which adaptively learns tasks of different natures, such as partial joint-level generation, sequence-level prediction, or 2D-to-3D motion prediction. We conduct extensive experiments to evaluate the effectiveness of our SiC on multiple tasks, including motion prediction, pose estimation, joint completion, and future pose estimation. We also evaluate its generalization capability on unseen tasks such as motion-in-between. These experiments show that our model achieves state-of-the-art multi-task performance and even outperforms single-task methods on certain tasks. Xinshun Wang, Zhongbin Fang, Xia Li 0005, Xiangtai Li, Chen Chen 0001, Mengyuan Liu 0001 |
CVPR | 6 |
| 2024 | Diffusion-Based Pose Refinement and Multi-Hypothesis Generation for 3D Human Pose EstimationabstractPrevious probabilistic models for 3D Human Pose Estimation (3DHPE) aimed to enhance pose accuracy by generating multiple hypotheses. However, most of the hypotheses generated deviate substantially from the true pose. Compared to deterministic models, the excessive uncertainty in probabilistic models leads to weaker performance in single-hypothesis prediction. To address these two challenges, we propose a diffusion-based refinement framework called DRPose, which refines the output of deterministic models by reverse diffusion and achieves more suitable multi-hypothesis prediction for the current pose benchmark by multi-step refinement with multiple noises. To this end, we propose a Scalable Graph Convolution Transformer (SGCT) and a Pose Refinement Module (PRM) for denoising and refining. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets demonstrate that our method achieves state-of-the-art performance on both single and multi-hypothesis 3DHPE. Code is available at https://github.com/KHB1698/DRPose. Hongbo Kang, Yong Wang 0053, Mengyuan Liu 0001, Doudou Wu, Xinlin Yuan, Wenming Yang |
ICASSP | 3 |
| 2024 | Denoising Diffusion Probabilistic Models for Action-Conditioned 3D Motion GenerationabstractDiffusion-based generative models have proven to be highly effective in various domains of synthesis. In this work, we propose a conditional paradigm utilizing the denoising diffusion probabilistic model (DDPM) to address the challenge of realistic and diverse action-conditioned 3D skeleton-based motion generation. The proposed method leverages bidirectional Markov chains to generate samples by inferring the reversed Markov chain based on the learned distribution mapping during the forward diffusion process. To the best of our knowledge, our work is the first to employ DDPM to synthesize a variable number of motion sequences conditioned on a categorical action. The proposed method is evaluated on the NTU RGB+D dataset and the NTU RGB+D two-person dataset, showing significant improvements over state-of-the-art motion generation methods. Mengyi Zhao, Mengyuan Liu 0001, Bin Ren 0005, Shuling Dai, Nicu Sebe |
ICASSP | 2 |
| 2024 | VG4D: Vision-Language Model Goes 4D Video RecognitionabstractUnderstanding the real world through point cloud video is a crucial aspect of robotics and autonomous driving systems. However, prevailing methods for 4D point cloud recognition have limitations due to sensor resolution, which leads to a lack of detailed information. Recent advances have shown that Vision-Language Models (VLM) pre-trained on web-scale text-image datasets can learn fine-grained visual concepts that can be transferred to various downstream tasks. However, effectively integrating VLM into the domain of 4D point clouds remains an unresolved problem. In this work, we propose the Vision-Language Models Goes 4D (VG4D) framework to transfer VLM knowledge from visual-text pretrained models to a 4D point cloud network. Our approach involves aligning the 4D encoder’s representation with a VLM learning a shared visual and text space from training on large-scale image-text pairs. By transferring the knowledge of the VLM to the 4D encoder and combining the VLM, our VG4D achieves improved recognition performance. To enhance the 4D encoder, we modernize the classic dynamic point cloud backbone and propose an improved version of PSTNet, im-PSTNet, which can efficiently model point cloud videos. Experiments demonstrate that our method achieves state-of-the-art performance for action recognition on both NTU RGB+D 60 dataset and NTU RGB+D 120 dataset. Zhichao Deng, Xiangtai Li, Xia Li 0005, Yunhai Tong, Mengyuan Liu 0001 |
ICRA | 6 |
| 2024 | Multi-Modality Co-Learning for Efficient Skeleton-based Action RecognitionabstractSkeleton-based action recognition has garnered significant attention due to the utilization of concise and resilient skeletons. Nevertheless, the absence of detailed body information in skeletons restricts performance, while other multimodal methods require substantial inference resources and are inefficient when using multimodal data during both training and inference stages. To address this and fully harness the complementary multimodal features, we propose a novel multi-modality co-learning (MMCL) framework by leveraging the multimodal large language models (LLMs) as auxiliary networks for efficient skeleton-based action recognition, which engages in multi-modality co-learning during the training stage and keeps efficiency by employing only concise skeletons in inference. Our MMCL framework primarily consists of two modules. First, the Feature Alignment Module (FAM) extracts rich RGB features from video frames and aligns them with global skeleton features via contrastive learning. Second, the Feature Refinement Module (FRM) uses RGB images with temporal information and text instruction to generate instructive features based on the powerful generalization of multimodal LLMs. These instructive text features will further refine the classification scores and the refined scores will enhance the model's robustness and generalization in a manner similar to soft labels. Extensive experiments on NTU RGB+D, NTU RGB+D 120 and Northwestern-UCLA benchmarks consistently verify the effectiveness of our MMCL, which outperforms the existing skeleton-based action recognition methods. Meanwhile, experiments on UTD-MHAD and SYSU-Action datasets demonstrate the commendable generalization of our MMCL in zero-shot and domain-adaptive action recognition. Our code is publicly available at: https://github.com/liujf69/MMCL-Action. Chen Chen 0001, Mengyuan Liu 0001 |
ACM Multimedia | 3 |
| 2024 | CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionabstractSkeleton-based multi-entity action recognition is a challenging task aiming to identify interactive actions or group activities involving multiple diverse entities. Existing models for individuals often fall short in this task due to the inherent distribution discrepancies among entity skeletons, leading to suboptimal backbone optimization. To this end, we introduce a Convex Hull Adaptive Shift based multi-Entity action recognition method (CHASE), which mitigates inter-entity distribution gaps and unbiases subsequent backbones. Specifically, CHASE comprises a learnable parameterized network and an auxiliary objective. The parameterized network achieves plausible, sample-adaptive repositioning of skeleton sequences through two key components. First, the Implicit Convex Hull Constrained Adaptive Shift ensures that the new origin of the coordinate system is within the skeleton convex hull. Second, the Coefficient Learning Block provides a lightweight parameterization of the mapping from skeleton sequences to their specific coefficients in convex combinations. Moreover, to guide the optimization of this network for discrepancy minimization, we propose the Mini-batch Pair-wise Maximum Mean Discrepancy as the additional objective. CHASE operates as a sample-adaptive normalization method to mitigate inter-entity distribution discrepancies, thereby reducing data bias and improving the subsequent classifier's multi-entity action recognition performance. Extensive experiments on six datasets, including NTU Mutual 11/26, H2O, Assembly101, Collective Activity and Volleyball, consistently verify our approach by seamlessly adapting to single-entity backbones and boosting their performance in multi-entity scenarios. Our code is publicly available at https://github.com/Necolizer/CHASE . Yuhang Wen 0001, Mengyuan Liu 0001, Songtao Wu, Beichen Ding |
NeurIPS | 2 |
| 2024 | Sharing Key Semantics in Transformer Makes Efficient Image RestorationabstractImage Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergence of Vision Transformers (ViTs) has further propelled these advancements. When computing, the self-attention mechanism, a cornerstone of ViTs, tends to encompass all global cues, even those from semantically unrelated objects or regions. This inclusivity introduces computational inefficiencies, particularly noticeable with high input resolution, as it requires processing irrelevant information, thereby impeding efficiency. Additionally, for IR, it is commonly noted that small segments of a degraded image, particularly those closely aligned semantically, provide particularly relevant information to aid in the restoration process, as they contribute essential contextual cues crucial for accurate reconstruction. To address these challenges, we propose boosting IR's performance by sharing the key semantics via Transformer for IR (i.e., SemanIR) in this paper. Specifically, SemanIR initially constructs a sparse yet comprehensive key-semantic dictionary within each transformer stage by establishing essential semantic connections for every degraded patch. Subsequently, this dictionary is shared across all subsequent transformer blocks within the same stage. This strategy optimizes attention calculation within each block by focusing exclusively on semantically related components stored in the key-semantic dictionary. As a result, attention calculation achieves linear computational complexity within each window. Extensive experiments across 6 IR tasks confirm the proposed SemanIR's state-of-the-art performance, quantitatively and qualitatively showcasing advancements. The visual results, code, and trained models are available at: https://github.com/Amazingren/SemanIR. Bin Ren 0005, Yawei Li 0001, Jingyun Liang, Mengyuan Liu 0001, Rita Cucchiara, Luc Van Gool, Ming-Hsuan Yang 0001, Nicu Sebe |
NeurIPS | 5 |
| 2024 | ModelNet-O: A large-scale synthetic dataset for occlusion-aware point cloud classification
Zhongbin Fang, Xia Li 0005, Xiangtai Li, Mengyuan Liu 0001 |
Comput. Vis. Image Underst. | 5 |
| 2024 | A gated cross-domain collaborative network for underwater object detection
Linhui Dai, Hong Liu 0008, Pinhao Song, Mengyuan Liu 0001 |
Pattern Recognit. | 4 |
| 2024 | Improving self-supervised action recognition from extremely augmented skeleton sequences
Tianyu Guo 0001, Mengyuan Liu 0001, Hong Liu 0008, Wenhao Li 0002 |
Pattern Recognit. | 2 |
| 2024 | TFAN: Twin-Flow Axis Normalization for Human Motion PredictionabstractHuman motion prediction involves forecasting upcoming body poses from historically observed sequences. Presently, various methods focus on modeling the positions of skeletal joints to generate future movements. Beyond solely using the joint position information, this letter further explores the informative bone vector information between joints for human motion prediction. Therefore, we propose TFAN, which integrates joint position and bone vector information to achieve more precise predictions. Additionally, joint positions and bone vectors, both represented as 3D vectors, were frequently imprecisely modeled due to neglect of variations in data distribution across axes in prior methods. Instead, we introduce Axis Normalization, which standardizes each coordinate axis individually, enhancing the model's sensitivity to data distribution disparities. Through experimental evaluations on the Human3.6M, AMASS, and 3DPW datasets, we consistently demonstrate that TFAN outperforms other existing methods. Our code will be available athttps://github.com/Deante-dx/TFAN. Yong Wang 0053, Zongying Li, Mengyuan Liu 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | MLP: Motion Label Prior for Temporal Sentence Localization in Untrimmed 3D Human MotionsabstractIn this paper, we address the unexplored question of temporal sentence localization in human motions (TSLM), aiming to locate a target moment from a 3D human motion that semantically corresponds to a text query. Considering that 3D human motions are captured using specialized motion capture devices, motions with only a few joints lack complex scene information like objects and lighting. Due to this character, motion data has low contextual richness and semantic ambiguity between frames, which limits the accuracy of predictions made by current video localization frameworks extended to TSLM to only a rough level. To refine this, we devise two novel label-prior-assisted training schemes: one embed prior knowledge of foreground and background to highlight the localization chances of target moments, and the other forces the originally rough predictions to overlap with the more accurate predictions obtained from the flipped start/end prior label sequences during recovery training. We show that injecting label-prior knowledge into the model is crucial for improving performance at high IoU. In our constructed TSLM benchmark, our model termedMLPachieves a recall of 44.13 at [email protected] on the BABEL dataset and 71.17 on HumanML3D (Restore), outperforming prior works. Finally, we showcase the potential of our approach in corpus-level moment retrieval. Our source code is openly accessible athttps://github.com/eanson023/mlp. Mengyuan Liu 0001, Yong Wang 0053, Yang Liu 0264, Hong Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | CHAMP: A Large-Scale Dataset for Skeleton-Based Composite HumAn Motion PredictionabstractSkeleton-based human motion prediction task aims to forecast future skeleton frames conditioned by observed skeleton sequence. Different from previous methods that focus on human motion prediction for atomic actions, we observe that people are witnessed to perform composite actions which consist of atomic actions that simultaneously happen. Considering the large number of action types, it is more laborious to collect composite actions than atomic actions. This paper presents a practical composite human motion prediction task, whose training data just contains atomic actions meanwhile the test data contains both atomic actions and composite actions. To evaluate this task, we collect a large-scale Composite HumAn Motion Prediction (CHAMP) dataset, whose training data has 16 types of atomic actions and test data has 50 types of composite actions. Despite the success of previous human motion prediction methods using Graph Convolutional Networks (GCN), these methods achieve inferior performances on our CHAMP dataset due to the huge domain gap between the training and test data. To solve this problem, we present a composite human motion prediction framework containing three modules. First, a Composite Motion Synthesis (CMS) module is designed to generate synthesized composite human actions from atomic actions. Second, a Composite GCN module is presented to predict human motion by modeling different human body parts. Third, a human body partition policy network is used to choose the best partition strategy for both the CMS and Composite GCN modules. Extensive experiments on the CHAMP dataset verify the effectiveness of our framework which obviously outperforms GCN-based methods. Mengyuan Liu 0001, Xinshun Wang, Can Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Cross-Model Cross-Stream Learning for Self-Supervised Human Action RecognitionabstractConsidering the instance-level discriminative ability, contrastive learning methods, including MoCo and SimCLR, have been adapted from the original image representation learning task to solve the self-supervised skeleton-based action recognition task. These methods usually use multiple data streams (i.e., joint, motion, and bone) for ensemble learning, meanwhile, how to construct a discriminative feature space within a single stream and effectively aggregate the information from multiple streams remains an open problem. To this end, this article first applies a new contrastive learning method called bootstrap your own latent (BYOL) to learn from skeleton data, and then formulate SkeletonBYOL as a simple yet effective baseline for self-supervised skeleton-based action recognition. Inspired by SkeletonBYOL, this article further presents a cross-model and cross-stream (CMCS) framework. This framework combines cross-model adversarial learning (CMAL) and cross-stream collaborative learning (CSCL). Specifically, CMAL learns single-stream representation by cross-model adversarial loss to obtain more discriminative features. To aggregate and interact with multistream information, CSCL is designed by generating similarity pseudolabel of ensemble learning as supervision and guiding feature generation for individual streams. Extensive experiments on three datasets verify the complementary properties between CMAL and CSCL and also verify that the proposed method can achieve better results than state-of-the-art methods using various evaluation protocols. Mengyuan Liu 0001, Hong Liu 0008, Tianyu Guo 0001 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2024 | Dynamic Dense Graph Convolutional Network for Skeleton-Based Human Motion PredictionabstractGraph Convolutional Networks (GCN) which typically follows a neural message passing framework to model dependencies among skeletal joints has achieved high success in skeleton-based human motion prediction task. Nevertheless, how to construct a graph from a skeleton sequence and how to perform message passing on the graph are still open problems, which severely affect the performance of GCN. To solve both problems, this paper presents a Dynamic Dense Graph Convolutional Network (DD-GCN), which constructs a dense graph and implements an integrated dynamic message passing. More specifically, we construct a dense graph with 4D adjacency modeling as a comprehensive representation of motion sequence at different levels of abstraction. Based on the dense graph, we propose a dynamic message passing framework that learns dynamically from data to generate distinctive messages reflecting sample-specific relevance among nodes in the graph. Extensive experiments on benchmark Human 3.6M and CMU Mocap datasets verify the effectiveness of our DD-GCN which obviously outperforms state-of-the-art GCN-based methods, especially when using long-term and our proposed extremely long-term protocol. Xinshun Wang, Can Wang 0006, Yuan Gao 0008, Mengyuan Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Facial Prior Guided Micro-Expression GenerationabstractThis paper focuses on the facial micro-expression (FME) generation task, which has potential application in enlarging digital FME datasets, thereby alleviating the lack of training data with labels in existing micro-expression datasets. Despite obvious progress in the image animation task, FME generation remains challenging because existing image animation methods can hardly encode subtle and short-term facial motion information. To this end, we present a facial-prior-guided FME generation framework that takes advantage of facial priors for facial motion generation. Specifically, we first estimate the geometric locations of action units (AUs) with detected facial landmarks. We further calculate an adaptive weighted prior (AWP) map, which alleviates the estimation error of AUs while efficiently capturing subtle facial motion patterns. To achieve smooth and realistic synthesis results, we use our proposed facial prior module to guide motion representation and generation modules in mainstream image animation frameworks. Extensive experiments on three benchmark datasets consistently show that our proposed facial prior module can be adopted in image animation frameworks and significantly improve their performance on micro-expression generation. Moreover, we use the generation technique to enlarge existing datasets, thereby improving the performance of general action recognition backbones on the FME recognition task. Our code is available at https://github.com/sysu19351158/FPB-FOMM. Xinhua Xu, Youjun Zhao, Yuhang Wen 0001, Zixuan Tang, Mengyuan Liu 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | Temporal Decoupling Graph Convolutional Network for Skeleton-Based Gesture RecognitionabstractSkeleton-based gesture recognition methods have achieved high success using Graph Convolutional Network (GCN), which commonly uses an adjacency matrix to model the spatial topology of skeletons. However, previous methods use the same adjacency matrix for skeletons from different frames, which limits the flexibility of GCN to model temporal information. To solve this problem, we propose a Temporal Decoupling Graph Convolutional Network (TD-GCN), which applies different adjacency matrices for skeletons from different frames. The main steps of each convolution layer in our proposed TD-GCN are as follows. To extract deep spatiotemporal information from skeleton joints, we first extract high-level spatiotemporal features from skeleton data. Then, channel-dependent and temporal-dependent adjacency matrices corresponding to different channels and frames are calculated to capture the spatiotemporal dependencies between skeleton joints. Finally, to fuse topology information from neighbor skeleton joints, spatiotemporal features of skeleton joints are fused based on channel-dependent and temporal-dependent adjacency matrices. To the best of our knowledge, we are the first to use temporal-dependent adjacency matrices for temporal-sensitive topology learning from skeleton joints. The proposed TD-GCN effectively improves the modeling ability of GCN and achieves state-of-the-art results on gesture datasets including SHREC'17 Track and DHG-14/28. Xinshun Wang, Can Wang 0006, Yuan Gao 0008, Mengyuan Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Feature Completion Transformer for Occluded Person Re-IdentificationabstractOccluded person re-identification is a challenging problem due to the destruction of occluders in different camera views. Most existing paradigms focus on visible human body parts through some external models to reduce noise interference. However, the feature misalignment problem caused by discarded occlusions negatively affects the performance of the network. Different from most previous works that discard the occluded regions, we present Feature Completion Transformer (FCFormer) that reduces noise interference and complements missing features in occluded parts. Specifically, Occlusion Instance Augmentation is proposed to simulate real and diverse occlusion situations on the holistic image, which enlarges the occlusion samples in the training set and forms aligned occluded-holistic pairs. To reduce the interference of noise, a two-stream architecture is proposed to learn pairwise discriminative features from aligned image pairs, while obtaining self-aligned occluded-holistic feature level sample-label pairs without additional auxiliary models. To complement the features of occluded regions, a Feature Completion Decoder is designed to aggregate possible information from self-generated occluded features in a self-supervised manner. Further, in order to correlate the completion features with identity information, Feature Completion Consistency loss is introduced to enforce the distribution of the generated completion features to be consistent with the real holistic feature distribution. In addition, we propose the Cross Hard Triplet loss to further bridge the gap between completion features and extracting features under the same ID. Extensive experiments over five challenging datasets demonstrate that the proposed FCFormer achieves superior performance and outperforms the state-of-theart methods by significant margins on Occluded-Duke dataset. Mengyuan Liu 0001, Hong Liu 0008, Wenhao Li 0002, Miaoju Ban, Tianyu Guo 0001, Yidi Li 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Style-Agnostic Representation Learning for Visible-Infrared Person Re-IdentificationabstractOne main challenge of visible-infrared person re-identification (VI Re-ID) lies in the large style discrepancy between the heterogeneous data. We present a STyle-Agnostic Representation learning (STAR) framework that bridges the modality gaps at both data and feature levels in a progressive manner. At the data level, we present Cross Modality Blending (CMB), a powerful and parameter-free data augmentation scheme that smoothly synthesizes intermediate modalities by conducting identity-preserving patch exchange and smooth cross-modality blending. At the feature level, we explore the inter-modality feature alignment problem from a new perspective of the style-related feature statistics. Specifically, we design a plug-and-play Adaptive Style Normalization (ASN) module to discard the intrinsic style distractors without losing discriminative content via dual-level adaptive distribution normalization and discriminability compensation. Moreover, considering that an appropriate modality intermediary can convey relevant information on the inter-modality distribution shift, we propose Reciprocal Modality Bridging Learning (RMBL) to better steer the modality bridging process. Two lightweight modality transformation modules are designed in RMBL to model an appropriate intermediate space by manipulating high-order statistics under our shortest distance constraint. Meanwhile, intermediary-guided distribution alignment is reciprocally conducted to align heterogeneous features to the modality intermediary. Experiments on VI Re-ID benchmarks demonstrate the superiority and flexibility of STAR over state-of-the-art methods. Jianbing Wu, Hong Liu 0008, Wei Shi 0009, Mengyuan Liu 0001, Wenhao Li 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | BCAN: Bidirectional Correct Attention Network for Cross-Modal RetrievalabstractAs a fundamental topic in bridging the gap between vision and language, cross-modal retrieval purposes to obtain the correspondences' relationship between fragments, i.e., subregions in images and words in texts. Compared with earlier methods that focus on learning the visual semantic embedding from images and sentences to the shared embedding space, the existing methods tend to learn the correspondences between words and regions via cross-modal attention. However, such attention-based approaches invariably result in semantic misalignment between subfragments for two reasons: 1) without modeling the relationship between subfragments and the semantics of the entire images or sentences, it will be hard for such approaches to distinguish images or sentences with multiple same semantic fragments and 2) such approaches focus attention evenly on all subfragments, including nonvisual words and a lot of redundant regions, which also will face the problem of semantic misalignment. To solve these problems, this article proposes a bidirectional correct attention network (BCAN), which introduces a novel concept of the relevance between subfragments and the semantics of the entire images or sentences and designs a novel correct attention mechanism by modeling the local and global similarity between images and sentences to correct the attention weights focused on the wrong fragments. Specifically, we introduce a concept about the semantic relationship between subfragments and entire images or sentences and use this concept to solve the semantic misalignment from two aspects. In our correct attention mechanism, we design two independent units to correct the weight of attention focused on the wrong fragments. Global correct unit (GCU) with modeling the global similarity between images and sentences into the attention mechanism to solve the semantic misalignment problem caused by focusing attention on relevant subfragments in irrelevant pairs (RI) and the local correct unit (LCU) consider the difference in the attention weights between fragments among two steps to solve the semantic misalignment problem caused by focusing attention on irrelevant subfragments in relevant pairs (IR). Extensive experiments on large-scale MS-COCO and Flickr30K show that our proposed method outperforms all the attention-based methods and is competitive to the state-of-the-art. Our code and pretrained model are publicly available at: https://github.com/liuyyy111/BCAN. Yang Liu 0264, Hong Liu 0008, Huaqiu Wang, Fanyang Meng, Mengyuan Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Novel Motion Patterns Matter for Practical Skeleton-Based Action RecognitionabstractMost skeleton-based action recognition methods assume that the same type of action samples in the training set and the test set share similar motion patterns. However, action samples in real scenarios usually contain novel motion patterns which are not involved in the training set. As it is laborious to collect sufficient training samples to enumerate various types of novel motion patterns, this paper presents a practical skeleton-based action recognition task where the training set contains common motion patterns of action samples and the test set contains action samples that suffer from novel motion patterns. For this task, we present a Mask Graph Convolutional Network (Mask-GCN) to focus on learning action-specific skeleton joints that mainly convey action information meanwhile masking action-agnostic skeleton joints that convey rare action information and suffer more from novel motion patterns. Specifically, we design a policy network to learn layer-wise body masks to construct masked adjacency matrices, which guide a GCN-based backbone to learn stable yet informative action features from dynamic graph structure. Extensive experiments on our newly collected dataset verify that Mask-GCN outperforms most GCN-based methods when testing with various novel motion patterns. Mengyuan Liu 0001, Fanyang Meng, Chen Chen 0001, Songtao Wu |
AAAI | 1 |
| 2023 | Spatio-Temporal Graph Diffusion for Text-Driven Human Motion Generation
Chang Liu 0030, Mengyi Zhao, Bin Ren 0005, Mengyuan Liu 0001, Nicu Sebe |
BMVC | 4 |
| 2023 | PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose EstimationabstractRecently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relations of human joints in each video frame and human dynamics across frames with cascaded transformer layers and has achieved impressive performance. However, in real scenarios, the performance of PoseFormer and its follow-ups is limited by two factors: (a) The length of the input joint sequence; (b) The quality of 2D joint detection. Existing methods typically apply self-attention to all frames of the input sequence, causing a huge computational burden when the frame number is increased to obtain advanced estimation accuracy, and they are not robust to noise naturally brought by the limited capability of 2D joint detectors. In this paper, we propose PoseFormerV2, which exploits a compact representation of lengthy skeleton sequences in the frequency domain to efficiently scale up the receptive field and boost robustness to noisy 2D joint detection. With minimum modifications to PoseFormer, the proposed method effectively fuses features both in the time domain and frequency domain, enjoying a better speed-accuracy trade-off than its precursor. Extensive experiments on two benchmark datasets (i.e., Human3.6M and MPI-INF-3DHP) demonstrate that the proposed approach significantly outperforms the original PoseFormer and other transformer-based variants. Code is released at https://github.com/ QitaoZhao/PoseFormerV2. Qitao Zhao, Mengyuan Liu 0001, Pichao Wang, Chen Chen 0001 |
CVPR | 3 |
| 2023 | Multi-Stream Facial Adaptive Network for Expression Recognition from a Single ImageabstractFacial expression recognition from a single image has potential applications in fields including human-computer interaction and medical diagnosis. Most recent methods use deep neural networks to directly learn from a roughly cropped facial image which is usually detected from a whole image by face detection algorithms. We observe that unrelated surrounding regions in the rough facial image prevent deep neural networks from learning facial-related discriminate features. To solve this problem, we present a Facial Adaptive Network (FAN) which is able to adaptively select an interest region from the given facial image, thus suffering less from the effect of unrelated regions. Based on the selected interest region, we further apply the self-attention mechanism to learn discriminate facial features. Moreover, we introduce a multi-stream FAN (ms-FAN) that learns richer facial features from multiple interest regions that are selected from pose-augmented facial images. Extensive experiments on Oulu-CASIA, CK+, and RAF-DB datasets consistently verify the effect of our proposed MS-FAN by achieving comparable results with state-of-the-art methods. Our code is available at https://github.com/zhangbc12/DAtt-ViT. Baichuan Zhang, Fanyang Meng, Runwei Ding, Mengyuan Liu 0001 |
ICASSP | 4 |
| 2023 | Interactive Spatiotemporal Token Attention Network for Skeleton-Based General Interactive Action RecognitionabstractRecognizing interactive action plays an important role in human-robot interaction and collaboration. Previous methods use late fusion and co-attention mechanism to capture interactive relations, which have limited learning capability or inefficiency to adapt to more interacting entities. With assumption that priors of each entity are already known, they also lack evaluations on a more general setting addressing the diversity of subjects. To address these problems, we propose an Interactive Spatiotemporal Token Attention Network (ISTA-Net), which simultaneously model spatial, temporal, and interactive relations. Specifically, our network contains a tokenizer to partition Interactive Spatiotemporal Tokens (ISTs), which is a unified way to represent motions of multiple diverse entities. By extending the entity dimension, ISTs provide better interactive representations. To jointly learn along three dimensions in ISTs, multi-head self-attention blocks integrated with 3D convolutions are designed to capture inter-token correlations. When modeling correlations, a strict entity ordering is usually irrelevant for recognizing interactive actions. To this end, Entity Rearrangement is proposed to eliminate the orderliness in ISTs for interchangeable entities. Extensive experiments on four datasets verify the effectiveness of ISTA-Net by outperforming state-of-the-art methods. Our code is publicly available at https://github.com/Necolizer/ISTA-Net. Yuhang Wen 0001, Zixuan Tang, Yunsheng Pang, Beichen Ding, Mengyuan Liu 0001 |
IROS | 5 |
| 2023 | Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised LearningabstractMasked Autoencoders (MAE) have demonstrated promising performance in self-supervised learning for both 2D and 3D computer vision. Nevertheless, existing MAE-based methods still have certain drawbacks. Firstly, the functional decoupling between the encoder and decoder is incomplete, which limits the encoder's representation learning ability. Secondly, downstream tasks solely utilize the encoder, failing to fully leverage the knowledge acquired through the encoder-decoder architecture in the pre-text task. In this paper, we propose Point Regress AutoEncoder (Point-RAE), a new scheme for regressive autoencoders for point cloud self-supervised learning. The proposed method decouples functions between the decoder and the encoder by introducing a mask regressor, which predicts the masked patch representation from the visible patch representation encoded by the encoder and the decoder reconstructs the target from the predicted masked patch representation. By doing so, we minimize the impact of decoder updates on the representation space of the encoder. Moreover, we introduce an alignment constraint to ensure that the representations for masked patches, predicted from the encoded representations of visible patches, are aligned with the masked patch presentations computed from the encoder. To make full use of the knowledge learned in the pre-training stage, we design a new finetune mode for the proposed Point-RAE. Extensive experiments demonstrate that our approach is efficient during pre-training and generalizes well on various downstream tasks. Specifically, our pre-trained models achieve a high accuracy of 90.28% on the ScanObjectNN hardest split and 94.1% accuracy on ModelNet40, surpassing all the other self-supervised learning methods. Our code and pretrained model are public available at: https://github.com/liuyyy111/Point-RAE. Yang Liu 0264, Chen Chen 0001, Can Wang 0006, Xulin King, Mengyuan Liu 0001 |
ACM Multimedia | 5 |
| 2023 | Dynamic Compositional Graph Convolutional Network for Efficient Composite Human Motion PredictionabstractWith potential applications in fields including intelligent surveillance and human-robot interaction, the human motion prediction task has become a hot research topic and also has achieved high success, especially using the recent Graph Convolutional Network (GCN). Current human motion prediction task usually focuses on predicting human motions for atomic actions. Observing that atomic actions can happen at the same time and thus formulating the composite actions, we propose the composite human motion prediction task. To handle this task, we first present a Composite Action Generation (CAG) module to generate synthetic composite actions for training, thus avoiding the laborious work of collecting composite action samples. Moreover, we alleviate the effect of composite actions on demand for a more complicated model by presenting a Dynamic Compositional Graph Convolutional Network (DC-GCN). Extensive experiments on the Human3.6M dataset and our newly collected CHAMP dataset consistently verify the efficiency of our DC-GCN method, which achieves state-of-the-art motion prediction accuracies and meanwhile needs few extra computational costs than traditional GCN-based human motion methods. Fanyang Meng, Songtao Wu, Mengyuan Liu 0001 |
ACM Multimedia | 5 |
| 2023 | Learning Snippet-to-Motion Progression for Skeleton-based Human Motion PredictionabstractExisting Graph Convolutional Networks to achieve human motion prediction largely adopt a one-step scheme, which output the prediction straight from history input, failing to exploit human motion patterns. We observe that human motions have transitional patterns and can be split into snippets representative of each transition. Each snippet can be reconstructed from its starting and ending poses referred to as the transitional poses. We propose a snippet-to-motion multi-stage framework that breaks motion prediction into sub-tasks easier to accomplish. Each sub-task integrates three modules: transitional pose prediction, snippet reconstruction, and snippet-to-motion prediction. Specifically, we propose to first predict only the transitional poses. Then we use them to reconstruct the corresponding snippets, obtaining a close approximation to the true motion sequence. Finally we refine them to produce the final prediction output. To implement the network, we propose a novel unified graph modeling, which allows for direct and effective feature propagation compared to existing approaches which rely on separate space-time modeling. Extensive experiments on Human 3.6M, CMU Mocap and 3DPW datasets verify the effectiveness of our method which achieves state-of-the-art performance. Xinshun Wang, Qiongjie Cui, Chen Chen 0001, Mengyuan Liu 0001 |
MMAsia | 5 |
| 2023 | Graph-Guided MLP-Mixer for Skeleton-Based Human Motion PredictionabstractIn recent years, Graph Convolutional Networks (GCNs) have been widely used in human motion prediction, but their performance remains unsatisfactory. Recently, MLP-Mixer, initially developed for vision tasks, has been leveraged into human motion prediction as a promising alternative to GCNs, which achieves both better performance and better efficiency than GCNs. Xinshun Wang, Qiongjie Cui, Chen Chen 0001, Mengyuan Liu 0001 |
MMAsia | 5 |
| 2023 | Cross-Modal Retrieval for Motion and Text via DropTriple LossabstractCross-modal retrieval of image-text and video-text is a prominent research area in computer vision and natural language processing. However, there has been insufficient attention given to cross-modal retrieval between human motion and text, despite its wide-ranging applicability. To address this gap, we utilize a concise yet effective dual-unimodal transformer encoder for tackling this task. Recognizing that overlapping atomic actions in different human motion sequences can lead to semantic conflicts between samples, we explore a novel triplet loss function called DropTriple Loss. This loss function discards false negative samples from the negative sample set and focuses on mining remaining genuinely hard negative samples for triplet training, thereby reducing violations they cause. We evaluate our model and approach on the HumanML3D and KIT Motion-Language datasets. On the latest HumanML3D dataset, we achieve a recall of 62.9% for motion retrieval and 71.5% for text retrieval (both based on R@10). The source code for our approach is publicly available at https://github.com/eanson023/rehamot. Yang Liu 0264, Haoqiang Wang, Mengyuan Liu 0001, Hong Liu 0008 |
MMAsia | 5 |
| 2023 | Explore In-Context Learning for 3D Point Cloud UnderstandingabstractWith the rise of large-scale models trained on broad data, in-context learning has become a new learning paradigm that has demonstrated significant potential in natural language processing and computer vision tasks. Meanwhile, in-context learning is still largely unexplored in the 3D point cloud domain. Although masked modeling has been successfully applied for in-context learning in 2D vision, directly extending it to 3D point clouds remains a formidable challenge. In the case of point clouds, the tokens themselves are the point cloud positions (coordinates) that are masked during inference. Moreover, position embedding in previous works may inadvertently introduce information leakage. To address these challenges, we introduce a novel framework, named Point-In-Context, designed especially for in-context learning in 3D point clouds, where both inputs and outputs are modeled as coordinates for each task. Additionally, we propose the Joint Sampling module, carefully designed to work in tandem with the general point sampling operator, effectively resolving the aforementioned technical issues. We conduct extensive experiments to validate the versatility and adaptability of our proposed methods in handling a wide range of tasks. Furthermore, with a more effective prompt selection strategy, our framework surpasses the results of individually trained models. Zhongbin Fang, Xiangtai Li, Xia Li 0005, Joachim M. Buhmann, Chen Change Loy, Mengyuan Liu 0001 |
NeurIPS | 6 |
| 2023 | A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose EstimationabstractThe dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for improved accuracy, which incurs performance saturation, intractable computation and the non-causal problem. This can be attributed to their inherent inability to perceive spatial context as plain 2D joint coordinates carry no visual cues. To address this issue, we propose a straightforward yet powerful solution: leveraging the $\textit{readily available}$ intermediate visual representations produced by off-the-shelf (pre-trained) 2D pose detectors -- no finetuning on the 3D task is even needed. The key observation is that, while the pose detector learns to localize 2D joints, such representations (e.g., feature maps) implicitly encode the joint-centric spatial context thanks to the regional operations in backbone networks. We design a simple baseline named $\textbf{Context-Aware PoseFormer}$ to showcase its effectiveness. $\textit{Without access to any temporal information}$, the proposed method significantly outperforms its context-agnostic counterpart, PoseFormer, and other state-of-the-art methods using up to $\textit{hundreds of}$ video frames regarding both speed and precision. $\textit{Project page:}$ https://qitaozhao.github.io/ContextAware-PoseFormer Qitao Zhao, Mengyuan Liu 0001, Chen Chen 0015 |
NeurIPS | 3 |
| 2023 | Exploiting Temporal Contexts With Strided Transformer for 3D Human Pose EstimationabstractDespite the great progress in 3D human pose estimation from videos, it is still an open problem to take full advantage of a redundant 2D pose sequence to learn representative representations for generating one 3D pose. To this end, we propose an improved Transformer-based architecture, called Strided Transformer, which simply and effectively lifts a long sequence of 2D joint locations to a single 3D pose. Specifically, a Vanilla Transformer Encoder (VTE) is adopted to model long-range dependencies of 2D pose sequences. To reduce the redundancy of the sequence, fully-connected layers in the feed-forward network of VTE are replaced with strided convolutions to progressively shrink the sequence length and aggregate information from local contexts. The modified VTE is termed as Strided Transformer Encoder (STE), which is built upon the outputs of VTE. STE not only effectively aggregates long-range information to a single-vector representation in a hierarchical global and local fashion, but also significantly reduces the computation cost. Furthermore, a full-to-single supervision scheme is designed at both full sequence and single target frame scales applied to the outputs of VTE and STE, respectively. This scheme imposes extra temporal smoothness constraints in conjunction with the single target frame supervision and hence helps produce smoother and more accurate 3D poses. The proposed Strided Transformer is evaluated on two challenging benchmark datasets, Human3.6 M and HumanEva-I, and achieves state-of-the-art results with fewer parameters. Code and models are available athttps://github.com/Vegetebird/StridedTransformer-Pose3D. Wenhao Li 0002, Hong Liu 0008, Runwei Ding, Mengyuan Liu 0001, Pichao Wang, Wenming Yang |
IEEE Trans. Multim. | 4 |
| 2022 | Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action RecognitionabstractIn recent years, self-supervised representation learning for skeleton-based action recognition has been developed with the advance of contrastive learning methods. The existing contrastive learning methods use normal augmentations to construct similar positive samples, which limits the ability to explore novel movement patterns. In this paper, to make better use of the movement patterns introduced by extreme augmentations, a Contrastive Learning framework utilizing Abundant Information Mining for self-supervised action Representation (AimCLR) is proposed. First, the extreme augmentations and the Energy-based Attention-guided Drop Module (EADM) are proposed to obtain diverse positive samples, which bring novel movement patterns to improve the universality of the learned representations. Second, since directly using extreme augmentations may not be able to boost the performance due to the drastic changes in original identity, the Dual Distributional Divergence Minimization Loss (D3M Loss) is proposed to minimize the distribution divergence in a more gentle way. Third, the Nearest Neighbors Mining (NNM) is proposed to further expand positive samples to make the abundant information mining process more reasonable. Exhaustive experiments on NTU RGB+D 60, PKU-MMD, NTU RGB+D 120 datasets have verified that our AimCLR can significantly perform favorably against state-of-the-art methods under a variety of evaluation protocols with observed higher quality action representations. Our code is available at https://github.com/Levigty/AimCLR. Tianyu Guo 0001, Hong Liu 0008, Mengyuan Liu 0001, Runwei Ding |
AAAI | 4 |
| 2022 | IRANet: Identity-relevance aware representation for cloth-changing person re-identification
Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001 |
Image Vis. Comput. | 3 |
| 2022 | Image-to-video person re-identification using three-dimensional semantic appearance alignment and cross-modal interactive learning
Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001 |
Pattern Recognit. | 3 |
| 2022 | Spatial-Temporal Asynchronous Normalization for Unsupervised 3D Action Representation LearningabstractUnsupervised 3D action representation learning from skeleton sequences has attracted increasing attention in recent years. Existing methods have successfully applied autoencoder network to learn 3D action representation by reconstructing original skeleton sequence. However, these methods ignore motion cues thus suffer from distinguishing actions especially with similar shape information and slightly different motion information. Instead of reconstructing original skeleton sequence, we learn distinctive 3D action representation with autoencoder network by reconstructing normalized motion sequence extracted from original input. To obtain the normalized motion sequence, we specifically design a novel spatial-temporal asynchronous normalization (STAN) method, which normalizes original skeleton sequence in two steps. First, STAN reduces redundant temporal information and extracts motion sequence by subtracting mean value along the temporal dimension. Second, STAN further normalizes the motion sequence along the spatial dimension and generates normalized motion sequence that suffers less from the effect of different human body shapes. Extensive experiments on large scale NTU RGB+D 60 and NTU RGB+D 120 datasets verify the effectiveness of our proposed STAN method, which achieves comparative results with state-of-the-art methods, and also outperforms alternative normalization methods. Mengyuan Liu 0001, Youneng Bao, Yongsheng Liang 0001, Fanyang Meng |
IEEE Signal Process. Lett. | 1 |
| 2022 | Regularizing Visual Semantic Embedding With Contrastive Learning for Image-Text MatchingabstractLearning visual semantic embedding for image-text matching has achieved high success by using triplet loss to pull positive image-text pairs which share similar semantic meaning and to push negative image-text pairs which share different semantic meaning. Without modeling constraints from image-image or text-text pairs, the generated visual semantic embedding inevitably faces the problem of semantic misalignments among similar images or among similar texts. To solve this problem, we present a contrastive visual semantic embedding framework, named as ConVSE, which achieves intra-modal semantic alignment by contrastive learning from augmented image-image (or text-text) pairs and achieves inter-modal semantic alignment by applying hardest-negative-enhanced triplet loss on image-text pairs. To the best of our knowledge, we are the first to find that contrastive learning benefits visual semantic embedding. Extensive experiments on large scale MSCOCO and Flickr30K datasets verify the effectiveness of our proposed ConVSE by outperforming visual semantic embedding-based methods and achieving new state-of-the-arts. Our code and pretrained model are publicly available at: \url{https://github.com/liuyyy111/ConVSE}. Yang Liu 0264, Hong Liu 0008, Huaqiu Wang, Mengyuan Liu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Attend, Correct And Focus: A Bidirectional Correct Attention Network For Image-Text MatchingabstractImage-text matching task aims to learn the fine-grained correspondences between images and sentences. Existing methods use attention mechanism to learn the correspondences by attending to all fragments without considering the relationship between fragments and global semantics, which inevitably lead to semantic misalignment among irrelevant fragments. To this end, we propose a Bidirectional Correct Attention Network (BCAN), which leverages global similarities and local similarities to reassign the attention weight, to avoid such semantic misalignment. Specifically, we introduce a global correct unit to correct the attention focused on relevant fragments in irrelevant semantics. A local correct unit is used to correct the attention focused on irrelevant fragments in relevant semantics. Experiments on Flickr30K and MSCOCO datasets verify the effectiveness of our proposed BCAN by outperforming both previous attention-based methods and state-of-the-art methods. Code can be found at: https://github.com/liuyyy111/BCAN. Yang Liu 0264, Huaqiu Wang, Fanyang Meng, Mengyuan Liu 0001, Hong Liu 0008 |
ICIP | 4 |
| 2021 | Facial Prior Based First Order Motion Model for Micro-expression GenerationabstractSpotting facial micro-expression from videos finds various potential applications in fields including clinical diagnosis and interrogation, meanwhile this task is still difficult due to the limited scale of training data. To solve this problem, this paper tries to formulate a new task called micro-expression generation and then presents a strong baseline which combines the first order motion model with facial prior knowledge. Given a target face, we intend to drive the face to generate micro-expression videos according to the motion patterns of source videos. Specifically, our new model involves three modules. First, we extract facial prior features from a region focusing module. Second, we estimate facial motion using key points and local affine transformations with a motion prediction module. Third, expression generation module is used to drive the target face to generate videos. We train our model on public CASME II, SAMM and SMIC datasets and then use the model to generate new micro-expression videos for evaluation. Our model achieves the first place in the Facial Micro-Expression Challenge 2021 (MEGC2021), where our superior performance is verified by three experts with Facial Action Coding System certification. Source code is provided in https://github.com/Necolizer/Facial-Prior-Based-FOMM. Youjun Zhao, Yuhang Wen 0001, Zixuan Tang, Xinhua Xu, Mengyuan Liu 0001 |
ACM Multimedia | 6 |
| 2020 | Grouped Temporal Enhancement Module for Human Action RecognitionabstractTemporal information is a significant cue for recognizing human actions from videos. Different from 2D CNN which can only capture spatial information in an efficient way, 3D CNN is good at capturing both spatial and temporal information at the expense of high computational cost. Beyond both methods, this paper presents a Grouped Temporal Enhancement (GTE) module which even outperforms 3D CNN, meanwhile only needs similar low computational cost as 2D CNN. The GTE module firstly decomposes an input video into spatial and temporal groups along channel dimension, and then uses a learnable temporal shift (LTS) operation for efficient temporal modeling. Finally, a 2D convolution filter is used to enhance the ability of LTS for spatial modeling. Extensive experiments on three benchmark datasets validate the effect of our method. Hong Liu 0008, Bin Ren 0005, Mengyuan Liu 0001, Runwei Ding |
ICIP | 3 |
| 2020 | Identity-sensitive loss guided and instance feature boosted deep embedding for person searchabstractPerson search aims at detecting and re-identifying pedestrians from whole monitoring images, which is vital for intelligent surveillance. However, this task is still challenging due to the extremely few instances per training identity and inherent fine-grained differences among different identities. To this end, this work proposes an identity-sensitive loss guided and instance feature boosted pipeline to extract deep discriminative feature embedding for person search. First, a prior anchor pre-trained network (PAPN) is designed to obtain proper initial state for the whole deep person search training baseline. Second, a new loss function called instance enhancing loss (IEL) is proposed to learn identity-sensitive features by introducing unlabeled identity information. Specifically, the proposed IEL can selectively utilize unlabeled identities with similar appearances to labeled identities to train the person search network. Third, considering the intra-class compactness of features learned by center loss and contextual inter-class relations, two instance boosting strategies (Boosting) are used to learn more discriminative features. Extensive experiments on two benchmark datasets, namely CUHK-SYSU and PRW, demonstrate the effectiveness of our approach. Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001 |
Neurocomputing | 3 |
| 2019 | Joint Dynamic Pose Image and Space Time Reversal for Human Action Recognition from VideosabstractHuman action recognition aims to classify a given video according to which type of action it contains. Disturbance brought by clutter background and unrelated motions makes the task challenging for video frame-based methods. To solve this problem, this paper takes advantage of pose estimation to enhance the performances of video frame features. First, we present a pose feature called dynamic pose image (DPI), which describes human action as the aggregation of a sequence of joint estimation maps. Different from traditional pose features using sole joints, DPI suffers less from disturbance and provides richer information about human body shape and movements. Second, we present attention-based dynamic texture images (att-DTIs) as pose-guided video frame feature. Specifically, a video is treated as a space-time volume, and DTIs are obtained by observing the volume from different views. To alleviate the effect of disturbance on DTIs, we accumulate joint estimation maps as attention map, and extend DTIs to attention-based DTIs (att-DTIs). Finally, we fuse DPI and att-DTIs with multi-stream deep neural networks and late fusion scheme for action recognition. Experiments on NTU RGB+D, UTD-MHAD, and Penn-Action datasets show the effectiveness of DPI and att-DTIs, as well as the complementary property between them. Mengyuan Liu 0001, Fanyang Meng, Chen Chen 0001, Songtao Wu |
AAAI | 1 |
| 2019 | Self-Refining Deep Symmetry Enhanced Network for Rain RemovalabstractRain removal aims to remove the rain streaks on rain images. The state-of-the-art methods are mostly based on Convolutional Neural Network (CNN). However, as CNN is not equivariant to object rotation, these methods are unsuitable for dealing with the tilted rain streaks. To tackle this problem, we propose Deep Symmetry Enhanced Network (DSEN) that is able to explicitly extract the rotation equivariant features from rain images. In addition, we design a self-refining mechanism to remove the accumulated rain streaks in a coarse-to-fine manner. This mechanism reuses DSEN with a novel information link which passes the gradient flow to the higher stages. Extensive experiments on both synthetic and real-world rain images show that our self-refining DSEN yields the top performance. Hong Liu 0008, Hanrong Ye, Xia Li 0005, Wei Shi 0009, Mengyuan Liu 0001, Qianru Sun |
ICIP | 5 |
| 2019 | Sample Fusion Network: An End-to-End Data Augmentation Network for Skeleton-Based Human Action RecognitionabstractData augmentation is a widely used technique for enhancing the generalization ability of deep neural networks for skeleton-based human action recognition (HAR) tasks. Most existing data augmentation methods generate new samples by means of handcrafted transforms. However, these methods often cannot be trained and then are discarded during testing because of the lack of learnable parameters. To solve those problems, a novel type of data augmentation network called a sample fusion network (SFN) is proposed. Instead of using handcrafted transforms, an SFN generates new samples via a long short-term memory (LSTM) autoencoder (AE) network. Therefore, an SFN and HAR network can be cascaded together to form a combined network that can be trained in an end-to-end manner. Moreover, an adaptive weighting strategy is employed to improve the complementarity between a sample and the new sample generated from it by an SFN, thus allowing the SFN to more efficiently improve the performance of the HAR network during testing. The experimental results on various datasets verify that the proposed method outperforms state-of-the-art data augmentation methods. More importantly, the proposed SFN architecture is a general framework that can be integrated with various types of networks for HAR. For example, when a baseline HAR model with three LSTM layers and one fully connected (FC) layer was used, the classification accuracy was increased from 79.53% to 90.75% on the NTU RGB+D dataset using a cross-view protocol, thus outperforming most other methods. Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Juanhui Tu, Mengyuan Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Spatial-Temporal Data Augmentation Based on LSTM Autoencoder Network for Skeleton-Based Human Action RecognitionabstractData augmentation is known to be of crucial importance for the generalization of RNN-based methods of skeleton-based human action recognition. Traditional data augmentation methods artificially adopt various transformations merely in spatial domain, which lack effective temporal representation. This paper extends traditional Long Short-Term Memory (LSTM) and presents a novel LSTM autoencoder network (LSTM-AE) for spatial-temporal data augmentation. In the LSTM-AE, the LSTM network preserves the temporal information of skeleton sequences, and the autoencoder architecture can automatically eliminate irrelevant and redundant information. Meanwhile, a regularized cross-entropy loss is defined to guide the LSTM-AE to learn more suitable representations of skeleton data. Experimental results on the currently largest NTU RGB+D dataset and public SmartHome dataset verify that the proposed model outperforms the state-of-the-art methods, and can be integrated with most of the RNN-based action recognition models easily. Juanhui Tu, Hong Liu 0008, Fanyang Meng, Mengyuan Liu 0001, Runwei Ding |
ICIP | 4 |
| 2018 | Hierarchical Dropped Convolutional Neural Network for Speed Insensitive Human Action RecognitionabstractHuman action recognition using skeleton data has lots of potential applications in content-based action retrieval and intelligent surveillance, with wide usage of depth sensors and robust skeleton estimation algorithms. Previous methods describe spatial temporal skeleton joints as a compact color image and then use Convolutional Neural Network (CNN) to extract more discriminative deep features. However, these methods ignore the effect of speed variation, which is a common phenomenon and can bring severe intra-varieties to same types of actions. To solve this problem, this paper presents a novel hierarchical dropped CNN architecture, which is constructed in two stages. Dropped CNN (d-CNN) is firstly developed to extract deep features from a probabilistic speed insensitive color image. This image expresses both spatial distributions and temporal evolutions of skeleton joints meanwhile avoids the effect of speed variations. To enhance the temporal discriminative power, we extend d-CNN to a hierarchical structure (h-CNN), where multiple scales of temporal information are encoded. Extensive experiments on benchmark MSRC-12 dataset and the largest NTU RGB+D dataset verify the effectiveness and robustness of the proposed method. Fanyang Meng, Hong Liu 0008, Yongsheng Liang 0001, Mengyuan Liu 0001, Wei Liu 0065 |
ICME | 4 |
| 2018 | 3D Action Recognition Using Multiscale Energy-Based Global Ternary ImageabstractThis paper presents an effective multiscale energy-based global ternary image (E-GTI) representation for action recognition from depth sequences. The unique property of our representation is that it takes the spatiotemporal discrimination and action speed variations into account, intending to solve the problems of distinguishing similar actions and identifying the actions with different speeds in one goal. The entire method is carried out in two stages. In the first stage, consecutive depth frames are used to generate global ternary image (GTI) features, which implicitly capture both inter-frame motion regions and motion directions. Specifically, each pixel in the GTI represents one of three possible states, namely, positive, negative, and neutral, which indicate the increased, decreased, and same depth values, respectively. To cope with speed variations in actions, energy-based sampling method is utilized, leading to multiscale E-GTI features, where the multiscale scheme can efficiently capture the temporal relationships among frames. In the second stage, all the E-GTI features are transformed by Radon transform (RT) as robust descriptors, which are aggregated by the bag-of-visual-words model as a compact representation. Extensive experiments on benchmark data sets show that our representation outperforms state-of-the-art approaches, since it captures discriminating spatiotemporal information of actions. Due to the merits of energy-based sampling and RT methods, our representation shows robustness to speed variations, depth noise, and partial occlusions. Mengyuan Liu 0001, Hong Liu 0008, Chen Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Robust 3D Action Recognition Through Sampling Local Appearances and Global DistributionsabstractThree-dimensional (3-D) action recognition has broad applications in human-computer interaction and intelligent surveillance. However, recognizing similar actions remains challenging since previous literature fails to capture motion and shape cues effectively from noisy depth data. In this paper, we propose a novel two-layer Bag-of-Visual-Words (BoVW) model, which suppresses the noise disturbances and jointly encodes both motion and shape cues. First, background clutter is removed by a background modeling method that is designed for depth data. Then, motion and shape cues are jointly used to generate robust and distinctive spatial-temporal interest points (STIPs): motion-based STIPs and shape-based STIPs. In the first layer of our model, a multiscale 3-D local steering kernel descriptor is proposed to describe local appearances of cuboids around motion-based STIPs. In the second layer, a spatial-temporal vector descriptor is proposed to describe the spatial-temporal distributions of shape-based STIPs. Using the BoVW model, motion and shape cues are combined to form a fused action representation. Our model performs favorably compared with common STIP detection and description methods. Thorough experiments verify that our model is effective in distinguishing similar actions and robust to background clutter, partial occlusions and pepper noise. Mengyuan Liu 0001, Hong Liu 0008, Chen Chen 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Depth Context: a new descriptor for human activity recognition by using sole depth sequences
Mengyuan Liu 0001, Hong Liu 0008 |
Neurocomputing | 1 |
| 2015 | SDM-BSM: A fusing depth scheme for human action recognitionabstractDepth map has shown promising capability in human action recognition, however it always be auxiliary of RGB features in previous work. As to sufficiently exploring depth map, we propose an innovative descriptor for human action recognition using solo depth data. First, Salient Depth Map (SDM) is calculated between two consecutive depth frames, which is superior for action description as it is located on salient moving objects. Moreover, Binary Shape Map (BSM) is proposed to depict the silhouettes induced by the lateral component of the scene action parallel to the image plane. Then, for implementation, a new framework as Bag-of-Map-Words is employed after concatenating SDM and BSM feature vectors. Experiments on NHA database demonstrate the superiority and high efficiency of the proposed method. We also give detailed comparisons with other features and analysis for parameters as a guidance of further applications. Hong Liu 0008, Mengyuan Liu 0001, Hao Tang 0005 |
ICIP | 3 |