Pichao Wang

dblp:151/6390 · DBLP profile ↗
← Back
86ranked-venue papers
13as first author
51since 2021 · last 2026
0000-0002-1430-0237ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 10 first-author · 32 since 2021Artificial intelligence and machine learning · 52 · 8 first-author · 38 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 H2OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
abstract
Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper, we present a hierarchical plug-and-play pruning-and-recovering framework, calledHierarchicalHourglassTokenizer (H2OT), for efficient transformer-based 3D human pose estimation from videos. H2OT begins with progressively pruning pose tokens of redundant frames and ends with recovering full-length sequences, resulting in a few pose tokens in the intermediate transformer blocks and thus improving the model efficiency. It works with two key modules, namely, a Token Pruning Module (TPM) and a Token Recovering Module (TRM). TPM dynamically selects a few representative tokens to eliminate the redundancy of video frames, while TRM restores the detailed spatio-temporal information based on the selected tokens, thereby expanding the network output to the original full-length temporal resolution for fast inference. Our method is general-purpose: it can be easily incorporated into common VPT models on bothseq2seqandseq2framepipelines while effectively accommodating different token pruning and recovery strategies. In addition, our H2OT reveals that maintaining the full pose sequence is unnecessary, and a few pose tokens of representative frames can achieve both high efficiency and estimation accuracy. Extensive experiments on multiple benchmark datasets demonstrate both the effectiveness and efficiency of the proposed method. Code and models are available athttps://github.com/NationalGAILab/HoT.
Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Pichao Wang, Shijian Lu, Nicu Sebe
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Style-Aware Contrastive Test-Time Adaptation: A Dual-Cache Model for Robust Vision-Language Alignment
abstract
Test-time adaptation (TTA) has emerged as a key strategy to enhance vision-language models (VLMs) under real-world distribution changes. However, existing methods always face two problems: 1) The fundamental trade-off dilemma: parameter-free TTA retains inference efficiency but fails to correct modality misalignment, while prompt tuning adapts to shifts, incurs high computational costs, and lacks knowledge retention. 2) Discriminative collapse also exists in TTA when faced with fine-grained downstream tasks. To alleviate these two bottlenecks, we introduce Style-aware Contrastive Test-Time Adaptation (SCTTA), a novel framework that jointly addresses modality misalignment and discriminative collapse. Firstly, we introduce Style-aware Embedding Adaptation (SEA), which dynamically refines text embeddings by incorporating domain-specific style attributes, improving alignment between visual and textual modalities. Secondly, we propose Fine-grained Contrastive Adaptation (FCA), which enhances feature separation by enforcing contrastive learning with adaptive prototypes, reducing inter-class feature overlap in fine-grained tasks. In addition, we introduce Dual-Cache Model (DCM), which extends prior unimodal cache model to a multimodal cache for the first time. Eventually, it accumulates adaptation knowledge through a visual-cache (capturing evolving domain styles) and a textual-cache (retaining discriminative semantics), enabling long-term adaptation without additional overhead. Extensive experiments on 15 datasets demonstrate that our approach achieves state-of-the-art performance for both fine-grained and out-of-distribution dataset benchmarks. Furthermore, SCTTA continuously improves as more test samples accumulate, validating its sustainable adaptation capacity. Our code is available at https://github.com/alusi123/SCTTA.
Shanshan Wang 0008, ALuSi, Xun Yang 0001, Pichao Wang, Ke Xu 0011, Xingyi Zhang 0001
IEEE Trans. Image Process.4
2026 TypiCD: Cognitive Diagnosis via Problem-Type-Guided Bias Correction
Shanshan Wang 0008, Yali Ye, Xun Yang 0001, Pichao Wang, Mengzhu Wang, Xingyi Zhang 0001
IEEE Trans. Knowl. Data Eng.4
2025 Beyond Speaker Identity: Text Guided Target Speech Extraction
abstract
Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker’s identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided TSE model StyleTSE that uses natural language descriptions of speaking style in addition to the audio clue to extract the desired speech from a given mixture. Our model integrates a speech separation network adapted from SepFormer with a bi-modality clue network that flexibly processes both audio and text clues. To train and evaluate our model, we introduce a new dataset TextrolMix with speech mixtures and natural language descriptions. Experimental results demonstrate that our method effectively separates speech based not only on who is speaking, but also on how they are speaking, enhancing TSE when traditional audio clues are absent. Demos are at: https://mingyue66.github.io/TextrolMix/demo/
Mingyue Huo, Cong Phuoc Huynh, Fanjie Kong, Pichao Wang, Vimal Bhat
ICASSP5
2025 Learning Rich Speech Representations with Acoustic-Semantic Factorization
abstract
Self-supervised pretraining has transformed speech representation learning, enabling models to generalize across various downstream tasks. However, empirical studies have highlighted two notable gaps. First, different speech tasks require varying levels of acoustic and semantic information, which are encoded at different layers within the model. This adds the extra complexity of layer selection on downstream tasks to reach optimal performance. Second, the entanglement of acoustic and semantic information can undermine model robustness, particularly in varied acoustic environments. To address these issues, we propose a two-branch multitask finetuning strategy that integrates Automatic Speech Recognition and transcript-aligned audio reconstruction, designed to preserve and disentangle semantic and acoustic information in a final layer of a pretrained model. Experiments with the pretrained Wav2Vec 2.0 model demonstrate that our approach surpasses ASR-only finetuning across multiple downstream tasks, and it significantly improves ASR robustness in acoustically varied (emotional) speech.
Minxue Niu, Najmeh Sadoughi, Abhishek Yanamandra, Pichao Wang, Vimal Bhat, S. Elizabeth Norred
ICASSP4
2025 Training-Free Text-Guided Image Editing with Visual Autoregressive Model
abstract
Text-guided image editing is an essential task that enables users to modify images through natural language descriptions. Recent advances in diffusion models and rectified flows have significantly improved editing quality, primarily relying on inversion techniques to extract structured noise from input images. However, inaccuracies in inversion can propagate errors, leading to unintended modifications and compromising fidelity. Moreover, even with perfect inversion, the entanglement between textual prompts and image features often results in global changes when only local edits are intended. To address these challenges, we propose a novel text-guided image editing framework based on VAR (Visual AutoRegressive modeling), which eliminates the need for explicit inversion while ensuring precise and controlled modifications. Our method introduces a caching mechanism that stores token indices and probability distributions from the original image, capturing the relationship between the source prompt and the image. Using this cache, we design an adaptive fine-grained masking strategy that dynamically identifies and constrains modifications to relevant regions, preventing unintended changes. A token reassembling approach further refines the editing process, enhancing diversity, fidelity, and control. Our framework operates in a training-free manner and achieves high-fidelity editing with faster inference speeds, processing a 1K resolution image in as fast as 1.2 seconds. Extensive experiments demonstrate that our method achieves performance comparable to, or even surpassing, existing diffusion- and rectified flow-based approaches in both quantitative metrics and visual quality. The code will be released.
Yufei Wang 0006, Lanqing Guo, Jiaxing Huang 0001, Pichao Wang, Bihan Wen
ICCV5
2025 Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
abstract
As online video content rapidly grows, the task of text-video retrieval (TVR) becomes increasingly important. A key challenge in TVR is the information asymmetry between video and text: videos are inherently richer in information, while their textual descriptions often capture only fragments of this complexity. This paper introduces a novel, data-centric framework to bridge this gap by enriching textual representations to better match the richness of video content. During training, videos are segmented into event-level clips and captioned to ensure comprehensive coverage. During retrieval, a large language model (LLM) generates semantically diverse queries to capture a broader range of possible matches. To enhance retrieval efficiency, we propose a query selection mechanism that identifies the most relevant and diverse queries, reducing computational cost while improving accuracy. Our method achieves state-of-the-art results across multiple benchmarks, demonstrating the power of data-centric approaches in addressing information asymmetry in TVR. This work paves the way for new research focused on leveraging data to improve cross-modal retrieval.
Zechen Bai, Tianjun Xiao, Tong He 0002, Pichao Wang, Zheng Zhang 0001, Thomas Brox, Zheng Shou 0001
ICLR4
2025 Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
Ju-Seung Byun, Micha Elsner, Pichao Wang, Andrew Perrault
INTERSPEECH4
2025 SparseDiT: Token Sparsification for Efficient Diffusion Transformer
abstract
Diffusion Transformers (DiT) are renowned for their impressive generative performance; however, they are significantly constrained by considerable computational costs due to the quadratic complexity in self-attention and the extensive sampling steps required. While advancements have been made in expediting the sampling process, the underlying architectural inefficiencies within DiT remain underexplored. We introduce SparseDiT, a novel framework that implements token sparsification across spatial and temporal dimensions to enhance computational efficiency while preserving generative quality. Spatially, SparseDiT employs a tri-segment architecture that allocates token density based on feature requirements at each layer: Poolingformer in the bottom layers for efficient global feature extraction, Sparse-Dense Token Modules (SDTM) in the middle layers to balance global context with local detail, and dense tokens in the top layers to refine high-frequency details. Temporally, SparseDiT dynamically modulates token density across denoising stages, progressively increasing token count as finer details emerge in later timesteps. This synergy between SparseDiT’s spatially adaptive architecture and its temporal pruning strategy enables a unified framework that balances efficiency and fidelity throughout the generation process. Our experiments demonstrate SparseDiT’s effectiveness, achieving a 55\% reduction in FLOPs and a 175\% improvement in inference speed on DiT-XL with similar FID score on 512$\times$512 ImageNet, a 56\% reduction in FLOPs across video generation datasets, and a 69\% improvement in inference speed on PixArt-$\alpha$ on text-to-image generation task with a 0.24 FID score decrease. SparseDiT provides a scalable solution for high-quality diffusion-based generation compatible with sampling optimization techniques. Code is available at https://github.com/changsn/SparseDiT.
Shuning Chang, Pichao Wang, Jiasheng Tang, Fan Wang 0019, Yi Yang 0001
NeurIPS2
2025 EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose Estimation
abstract
Locating 3D objects from a single RGB image via Perspective-n-Point (PnP) is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest interpreting PnP as a differentiable layer, allowing for partial learning of 2D-3D point correspondences by backpropagating the gradients of pose loss. Yet, learning the entire correspondences from scratch is highly challenging, particularly for ambiguous pose solutions, where the globally optimal pose is theoretically non-differentiable w.r.t. the points. In this paper, we propose the EPro-PnP, a probabilistic PnP layer for general end-to-end pose estimation, which outputs a distribution of pose with differentiable probability density on the SE(3) manifold. The 2D-3D coordinates and corresponding weights are treated as intermediate variables learned by minimizing the KL divergence between the predicted and target pose distribution. The underlying principle generalizes previous approaches, and resembles the attention mechanism. EPro-PnP can enhance existing correspondence networks, closing the gap between PnP-based method and the task-specific leaders on the LineMOD 6DoF pose estimation benchmark. Furthermore, EPro-PnP helps to explore new possibilities of network design, as we demonstrate a novel deformable correspondence network with the state-of-the-art pose accuracy on the nuScenes 3D object detection benchmark.
Hansheng Chen 0001, Wei Tian 0001, Pichao Wang, Fan Wang 0019, Lu Xiong 0001, Hao Li 0030
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation
abstract
Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper, we present a plug-and-play pruning-and-recovering framework, called Hourglass Tokenizer (HoT), for efficient transformer-based 3D human pose estimation from videos. Our HoT begins with pruning pose tokens of re-dundant frames and ends with recovering full-length tokens, resulting in a few pose tokens in the intermediate transformer blocks and thus improving the model efficiency. To effectively achieve this, we propose a token pruning cluster (TPC) that dynamically selects a few representative tokens with high semantic diversity while eliminating the redundancy of video frames. In addition, we develop a token recovering attention (TRA) to restore the detailed spatio-temporal information based on the selected tokens, thereby expanding the network output to the original full-length temporal resolution for fast inference. Extensive experiments on two benchmark datasets (i.e., Human3.6M and MPI-INF-3DHP) demonstrate that our method can achieve both high efficiency and estimation accuracy compared to the original VPT models. For instance, applying to MotionBERT and MixSTE on Hu-man3.6M, our HoT can save nearly 50% FLOPs without sacrificing accuracy and nearly 40% FLOPs with only 0.2% accuracy drop, respectively. Code and models are available at https://github.com/NationalGAILab/HoT.
Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Pichao Wang, Jialun Cai, Nicu Sebe
CVPR4
2024 Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
abstract
The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and re-silient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the de-termination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% ~ 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five bench-mark datasets, including MSRVTT, LSMDC, DiDeMo, VA-TEX, and Charades. Code and models are available here.
Jiamian Wang, Pichao Wang, Dongfang Liu, Sohail A. Dianat, Raghuveer M. Rao, Majid Rabbani, Zhiqiang Tao
CVPR2
2024 Adaptive Query Selection for Camouflaged Instance Segmentation
abstract
Camouflaged instance segmentation is a challenging task due to the various aspects such as color, structure, lighting, etc., of object instances embedded in complex backgrounds. Although the current DETR-based scheme simplifies the pipeline, it suffers from a large number of object queries, leading to many false positive instances. To address this issue, we propose an adaptive query selection mechanism. Our research reveals that a large number of redundant queries scatter the extracted features of the camouflaged instances. To remove these redundant queries with weak correlation, we evaluate the importance of the object query from the perspectives of information entropy and volatility. Moreover, we observed that occlusion and overlapping instances significantly impact the accuracy of the selection mechanism. Therefore, we design a boundary location embedding mechanism that incorporates fake instance boundaries to obtain better location information for more accurate query instance matching. We conducted extensive experiments on two challenging camouflaged instance segmentation datasets, namely COD10K and NC4K, and demonstrated the effectiveness of our proposed model. Compared with the OSFormer, our model significantly improves the performance by 3.8% AP and 5.6% AP with less computational cost, achieving the state-of-the-art of 44.8 AP and 48.1 AP with ResNet-50 on the COD10K and NC4K test-dev sets, respectively.
Pichao Wang, Hao Luo 0004, Fan Wang 0019
ACM Multimedia2
2024 Align2Concept: Language Guided Interpretable Image Recognition by Visual Prototype and Textual Concept Alignment
abstract
Most works of interpretable neural networks strive for learning the semantics concepts merely from single modal information such as images. However, humans usually learn semantic concepts from multiple modalities and the semantics is encoded by the brain from fused multi-modal information. Inspired by cognitive science and vision-language learning, we propose a Prototype-Concept Alignment Network (ProCoNet) for learning visual prototypes under the guidance of textual concepts. In the ProCoNet, we have designed a visual encoder to decompose the input image into regional features of prototypes, while also developing a prompt generation strategy that incorporates in-context learning to prompt large language models to generate textual concepts. To align visual prototypes with textual concepts, we leverage the multimodal space provided by the pre-trained CLIP as a bridge. Specifically, the regional features from the vision space and the cropped regions of prototypes encoded by CLIP reside on different but semantically highly correlated manifolds, i.e. follow a multi-manifold distribution. We transform the multi-manifold distribution alignment problem into optimizing the projection matrix by Cayley transform on the Stiefel manifold. Through the learned projection matrix, visual prototypes can be projected into the multimodal space to align with semantically similar textual concept features encoded by CLIP. We conducted two case studies on the CUB-200-2011 and Oxford Flower dataset. Our experiments show that the ProCoNet provides higher accuracy and better interpretability compared to the single-modality interpretable model. Furthermore, ProCoNet offers a level of interpretability not previously available in other interpretable methods.
Jiaqi Wang 0006, Pichao Wang, Huafeng Liu 0001, Chang Gao 0007, Liping Jing
ACM Multimedia2
2024 One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
abstract
We introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoning capabilities and world knowledge of large language models, and augmented by the Segment Anything Model, VideoLISA generates temporally consistent segmentation masks in videos based on language instructions. Existing image-based methods, such as LISA, struggle with video tasks due to the additional temporal dimension, which requires temporal dynamic understanding and consistent segmentation across frames. VideoLISA addresses these challenges by integrating a Sparse Dense Sampling strategy into the video-LLM, which balances temporal context and spatial detail within computational constraints. Additionally, we propose a One-Token-Seg-All approach using a specially designed <TRK> token, enabling the model to segment and track objects across multiple frames. Extensive evaluations on diverse benchmarks, including our newly introduced ReasonVOS benchmark, demonstrate VideoLISA's superior performance in video object segmentation tasks involving complex reasoning, temporal understanding, and object tracking. While optimized for videos, VideoLISA also shows promising generalization to image segmentation, revealing its potential as a unified foundation model for language-instructed object segmentation. Code and model will be available at: https://github.com/showlab/VideoLISA.
Zechen Bai, Tong He 0002, Haiyang Mei, Pichao Wang, Ziteng Gao, Joya Chen, Zheng Zhang 0001, Zheng Shou 0001
NeurIPS4
2024 Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning
abstract
Despite advancements in Text-to-Video (T2V) generation, producing videos with realistic motion remains challenging. Current models often yield static or minimally dynamic outputs, failing to capture complex motions described by text. This issue stems from the internal biases in text encoding which overlooks motions, and inadequate conditioning mechanisms in T2V generation models. To address this, we propose a novel framework called DEcomposed MOtion (DEMO), which enhances motion synthesis in T2V generation by decomposing both text encoding and conditioning into content and motion components. Our method includes a content encoder for static elements and a motion encoder for temporal dynamics, alongside separate content and motion conditioning mechanisms. Crucially, we introduce text-motion and video-motion supervision to improve the model's understanding and generation of motion. Evaluations on benchmarks such as MSR-VTT, UCF-101, WebVid-10M, EvalCrafter, and VBench demonstrate DEMO's superior ability to produce videos with enhanced motion dynamics while maintaining high visual quality. Our approach significantly advances T2V generation by integrating comprehensive motion understanding directly from textual descriptions. Project page: https://PR-Ryan.github.io/DEMO-project/
Penghui Ruan, Pichao Wang, Divya Saxena, Jiannong Cao 0001, Yuhui Shi 0001
NeurIPS2
2024 Diffusion-Inspired Truncated Sampler for Text-Video Retrieval
abstract
Prevalent text-to-video retrieval methods represent multimodal text-video data in a joint embedding space, aiming at bridging the relevant text-video pairs and pulling away irrelevant ones. One main challenge in state-of-the-art retrieval methods lies in the modality gap, which stems from the substantial disparities between text and video and can persist in the joint space. In this work, we leverage the potential of Diffusion models to address the text-video modality gap by progressively aligning text and video embeddings in a unified space. However, we identify two key limitations of existing Diffusion models in retrieval tasks: The L2 loss does not fit the ranking problem inherent in text-video retrieval, and the generation quality heavily depends on the varied initial point drawn from the isotropic Gaussian, causing inaccurate retrieval. To this end, we introduce a new Diffusion-Inspired Truncated Sampler (DITS) that jointly performs progressive alignment and modality gap modeling in the joint embedding space. The key innovation of DITS is to leverage the inherent proximity of text and video embeddings, defining a truncated diffusion flow from the fixed text embedding to the video embedding, enhancing controllability compared to adopting the isotropic Gaussian. Moreover, DITS adopts the contrastive loss to jointly consider the relevant and irrelevant pairs, not only facilitating alignment but also yielding a discriminatively structured embedding. Experiments on five benchmark datasets suggest the state-of-the-art performance of DITS. We empirically find that DITS can also improve the structure of the CLIP embedding space. Code is available at https://github.com/Jiamian- Wang/DITS-text-video-retrieval
Jiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao
NeurIPS2
2024 DFN: A deep fusion network for flexible single and multi-modal action recognition
abstract
Multi-modal action recognition methods can be generally classified into two categories: (1) fusing multi-modal features with simple concatenation or fusing the classification scores of individual modalities without considering the interaction among the multi-modalities; (2) using one of the modalities as privileged information in training to boost the recognition on the other modalities in inference. The former approach usually is not able to deal with the cases where one of the modalities is missing. In the latter, the trained classifier does not work on the privileged modality. To address these shortcomings, this paper presents a novel end-to-end trainable deep fusion network (DFN) that is able to improve the performance not only in the cases where all modalities are available and also in the cases where there is a missing modality. The DFN is simple yet effective with the capability of retrieving an estimation of one modality by using another modality through a Multilayer Perceptron (MLP). In order to better preserve structure information, the DFN first maps the individual modality features to a high dimensional Kronecker-product space and subsequently learns a low-dimensional discriminative space for classification. The effectiveness of the proposed DFN has been verified on three benchmark datasets: the large NTU RGB+D, UTD-MHAD, and SYSU-3D datasets and it has achieved state-of-the-art results.
Chuankun Li, Yonghong Hou, Wanqing Li 0001, Zewei Ding, Pichao Wang
Expert Syst. Appl.5
2024 SCT: A Simple Baseline for Parameter-Efficient Fine-Tuning via Salient Channels
Hengyuan Zhao, Pichao Wang, Hao Luo 0004, Fan Wang 0019, Zheng Shou 0001
Int. J. Comput. Vis.2
2023 Head-Free Lightweight Semantic Segmentation with Linear Transformer
abstract
Existing semantic segmentation works have been mainly focused on designing effective decoders; however, the computational load introduced by the overall structure has long been ignored, which hinders their applications on resource-constrained hardwares. In this paper, we propose a head-free lightweight architecture specifically for semantic segmentation, named Adaptive Frequency Transformer (AFFormer). AFFormer adopts a parallel architecture to leverage prototype representations as specific learnable local descriptions which replaces the decoder and preserves the rich image semantics on high-resolution features. Although removing the decoder compresses most of the computation, the accuracy of the parallel structure is still limited by low computational resources. Therefore, we employ heterogeneous operators (CNN and vision Transformer) for pixel embedding and prototype representations to further save computational costs. Moreover, it is very difficult to linearize the complexity of the vision Transformer from the perspective of spatial domain. Due to the fact that semantic segmentation is very sensitive to frequency information, we construct a lightweight prototype learning block with adaptive frequency filter of complexity O(n) to replace standard self attention with O(n^2). Extensive experiments on widely adopted datasets demonstrate that AFFormer achieves superior accuracy while retaining only 3M parameters. On the ADE20K dataset, AFFormer achieves 41.8 mIoU and 4.6 GFLOPs, which is 4.4 mIoU higher than Segformer, with 45% less GFLOPs. On the Cityscapes dataset, AFFormer achieves 78.7 mIoU and 34.4 GFLOPs, which is 2.5 mIoU higher than Segformer with 72.5% less GFLOPs. Code is available at https://github.com/dongbo811/AFFormer.
Pichao Wang, Fan Wang 0019
AAAI2
2023 Frequency Domain Disentanglement for Arbitrary Neural Style Transfer
abstract
Arbitrary neural style transfer has been a popular research topic due to its rich application scenarios. Effective disentanglement of content and style is the critical factor for synthesizing an image with arbitrary style. The existing methods focus on disentangling feature representations of content and style in the spatial domain where the content and style components are innately entangled and difficult to be disentangled clearly. Therefore, these methods always suffer from low-quality results because of the sub-optimal disentanglement. To address such a challenge, this paper proposes the frequency mixer (FreMixer) module that disentangles and re-entangles the frequency spectrum of content and style components in the frequency domain. Since content and style components have different frequency-domain characteristics (frequency bands and frequency patterns), the FreMixer could well disentangle these two components. Based on the FreMixer module, we design a novel Frequency Domain Disentanglement (FDD) framework for arbitrary neural style transfer. Qualitative and quantitative experiments verify that the proposed method can render better stylized results compared to the state-of-the-art methods.
Hao Luo 0004, Pichao Wang, Zhibin Wang 0004, Shang Liu 0002, Fan Wang 0019
AAAI3
2023 Making Vision Transformers Efficient from A Token Sparsification View
abstract
The quadratic computational complexity to the number of tokens limits the practical applications of Vision Transformers (ViTs). Several works propose to prune redundant tokens to achieve efficient ViTs. However, these methods generally suffer from (i) dramatic accuracy drops, (ii) application difficulty in the local vision transformer, and (iii) non-general-purpose networks for downstream tasks. In this work, we propose a novel Semantic Token ViT (STViT), for efficient global and local vision transformers, which can also be revised to serve as backbone for downstream tasks. The semantic tokens represent cluster centers, and they are initialized by pooling image tokens in space and recovered by attention, which can adaptively represent global or local semantic information. Due to the cluster properties, a few semantic tokens can attain the same effect as vast image tokens, for both global and local vision transformers. For instance, only 16 semantic tokens on DeiT-(Tiny,Small,Base) can achieve the same accuracy with more than 100% inference speed improvement and nearly 60% FLOPs reduction; on Swin-(Tiny,Small,Base), we can employ 16 semantic tokens in each window to further speed it up by around 20% with slight accuracy increase. Besides great success in image classification, we also extend our method to video recognition. In addition, we design a STViT-R(ecovery) network to restore the detailed spatial information based on the STViT, making it work for downstream tasks, which is powerless for previous token sparsification methods. Experiments demonstrate that our method can achieve competitive results compared to the original networks in object detection and instance segmentation, with over 30% FLOPs reduction for backbone.
Shuning Chang, Pichao Wang, Ming Lin 0002, Fan Wang 0019, Junhao Zhang 0001, Rong Jin 0001, Zheng Shou 0001
CVPR2
2023 Selective Structured State-Spaces for Long-Form Video Understanding
abstract
Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence ($S4$) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all imagetokens equally as done by$S4$model can adversely affect its efficiency and accuracy. To address this limitation, we present a novel Selective$S4$(i.e.,$S5)$model that employs a lightweight mask generator to adaptively select informative image tokens resulting in more efficient and accurate modeling of long-term spatiotemporal dependencies in videos. Unlike previous mask-based token reduction methods used in transformers, our$S5$model avoids the dense self-attention calculation by making use of the guidance of the momentum-updated$S4$model. This enables our model to efficiently discard less informative tokens and adapt to various long-form video understanding tasks more effectively. However, as is the case for most token reduction methods, the informative image tokens could be dropped incorrectly. To improve the robustness and the temporal horizon of our model, we propose a novel long-short masked contrastive learning (LSMCL) approach that enables our model to predict longer temporal context using shorter input videos. We present extensive comparative results using three challenging long-form video understanding datasets (LVU, COIN and Breakfast), demonstrating that our approach consistently outperforms the previous state-of-the-art S4 model by up to 9.6% accuracy while reducing its memory footprint by 23%.
Pichao Wang, Linda Liu, Mohamed Omar, Raffay Hamid
CVPR3
2023 PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation
abstract
Recently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relations of human joints in each video frame and human dynamics across frames with cascaded transformer layers and has achieved impressive performance. However, in real scenarios, the performance of PoseFormer and its follow-ups is limited by two factors: (a) The length of the input joint sequence; (b) The quality of 2D joint detection. Existing methods typically apply self-attention to all frames of the input sequence, causing a huge computational burden when the frame number is increased to obtain advanced estimation accuracy, and they are not robust to noise naturally brought by the limited capability of 2D joint detectors. In this paper, we propose PoseFormerV2, which exploits a compact representation of lengthy skeleton sequences in the frequency domain to efficiently scale up the receptive field and boost robustness to noisy 2D joint detection. With minimum modifications to PoseFormer, the proposed method effectively fuses features both in the time domain and frequency domain, enjoying a better speed-accuracy trade-off than its precursor. Extensive experiments on two benchmark datasets (i.e., Human3.6M and MPI-INF-3DHP) demonstrate that the proposed approach significantly outperforms the original PoseFormer and other transformer-based variants. Code is released at https://github.com/ QitaoZhao/PoseFormerV2.
Qitao Zhao, Mengyuan Liu 0001, Pichao Wang, Chen Chen 0001
CVPR4
2023 Revisiting Vision Transformer from the View of Path Ensemble
abstract
Vision Transformers (ViTs) are normally regarded as a stack of transformer layers. In this work, we propose a novel view of ViTs showing that they can be seen as ensemble networks containing multiple parallel paths with different lengths. Specifically, we equivalently transform the traditional cascade of multi-head self-attention (MSA) and feed-forward network (FFN) into three parallel paths in each transformer layer. Then, we utilize the identity connection in our new transformer form and further transform the ViT into an explicit multi-path ensemble network. From the new perspective, these paths perform two functions: the first is to provide the feature for the classifier directly, and the second is to provide the lower-level feature representation for subsequent longer paths. We investigate the influence of each path for the final prediction and discover that some paths even pull down the performance. Therefore, we propose the path pruning and EnsembleScale skills for improvement, which cut out the underperforming paths and reweight the ensemble components, respectively, to optimize the path combination and make the short paths focus on providing high-quality representation for subsequent paths. We also demonstrate that our path combination strategies can help ViTs go deeper and act as high-pass filters to filter out partial low-frequency signals. To further enhance the representation of paths served for subsequent paths, self-distillation is applied to transfer knowledge from the long paths to the short paths. This work calls for more future research to explain and design ViTs from new perspectives.
Shuning Chang, Pichao Wang, Hao Luo 0004, Fan Wang 0019, Zheng Shou 0001
ICCV2
2023 Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment
abstract
Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality while disregarding the audio signal for this task. Nevertheless, a recent advancement by ECLIPSE has improved long-range text-to-video retrieval by developing an audiovisual video representation. Nonetheless, the objective of the text-to-video retrieval task is to capture the complementary audio and video information that is pertinent to the text query rather than simply achieving better audio and video alignment. To address this issue, we introduce TEFAL, a TExt-conditioned Feature ALignment method that produces both audio and video representations conditioned on the text query. Instead of using only an audiovisual attention block, which could suppress the audio information relevant to the text query, our approach employs two independent cross-modal attention blocks that enable the text to attend to the audio and video representations separately. Our proposed method’s efficacy is demonstrated on four benchmark datasets that include audio: MSR-VTT, LSMDC, VATEX, and Charades, and achieves better than state-of-the-art performance consistently across the four datasets. This is attributed to the additional text-query-conditioned audio representation and the complementary information it adds to the text-query-conditioned video representation.
Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg, Ashutosh Sanan, Mohamed Omar
ICCV3
2023 Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture Recognition
abstract
RGB-D action and gesture recognition remain an interesting topic in human-centered scene understanding, primarily due to the multiple granularities and large variation in human motion. Although many RGB-D based action and gesture recognition approaches have demonstrated remarkable results by utilizing highly integrated spatio-temporal representations across multiple modalities (i.e., RGB and depth data), they still encounter several challenges. Firstly, vanilla 3D convolution makes it hard to capture fine-grained motion differences between local clips under different modalities. Secondly, the intricate nature of highly integrated spatio-temporal modeling can lead to optimization difficulties. Thirdly, duplicate and unnecessary information can add complexity and complicate entangled spatio-temporal modeling. To address the above issues, we propose an innovative heuristic architecture called Multi-stage Factorized Spatio-Temporal (MFST) for RGB-D action and gesture recognition. The proposed MFST model comprises a 3D Central Difference Convolution Stem (CDC-Stem) module and multiple factorized spatio-temporal stages. The CDC-Stem enriches fine-grained temporal perception, and the multiple hierarchical spatio-temporal stages construct dimension-independent higher-order semantic primitives. Specifically, the CDC-Stem module captures bottom-level spatio-temporal features and passes them successively to the following spatio-temporal factored stages to capture the hierarchical spatial and temporal features through the Multi-Scale Convolution and Transformer (MSC-Trans) hybrid block and Weight-shared Multi-Scale Transformer (WMS-Trans) block. The seamless integration of these innovative designs results in a robust spatio-temporal representation that outperforms state-of-the-art approaches on RGB-D action and gesture recognition datasets.
Yujun Ma, Benjia Zhou, Ruili Wang 0001, Pichao Wang
ACM Multimedia4
2023 What Limits the Performance of Local Self-attention?
Jingkai Zhou, Pichao Wang, Jiasheng Tang, Fan Wang 0019, Hao Li 0030, Rong Jin 0001
Int. J. Comput. Vis.2
2023 FT-HID: a large-scale RGB-D dataset for first- and third-person human interaction analysis
Zihui Guo, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001
Neural Comput. Appl.3
2023 A Unified Multimodal De- and Re-Coupling Framework for RGB-D Motion Recognition
abstract
Motion recognition is a promising direction in computer vision, but the training of video classification models is much harder than images due to insufficient data and considerable parameters. To get around this, some works strive to explore multimodal cues from RGB-D data. Although improving motion recognition to some extent, these methods still face sub-optimal situations in the following aspects: (i) Data augmentation, i.e., the scale of the RGB-D datasets is still limited, and few efforts have been made to explore novel data augmentation strategies for videos; (ii) Optimization mechanism, i.e., the tightly space-time-entangled network structure brings more challenges to spatiotemporal information modeling; And (iii) cross-modal knowledge fusion, i.e., the high similarity between multimodal representations leads to insufficient late fusion. To alleviate these drawbacks, we propose to improve RGB-D-based motion recognition both from data and algorithm perspectives in this article. In more detail, firstly, we introduce a novel video data augmentation method dubbed ShuffleMix, which acts as a supplement to MixUp, to provide additional temporal regularization for motion recognition. Secondly, a Unified Multimodal De-coupling and multi-stage Re-coupling framework, termed UMDR, is proposed for video representation learning. Finally, a novel cross-modal Complement Feature Catcher (CFCer) is explored to mine potential commonalities features in multimodal information as the auxiliary fusion stream, to improve the late fusion results. The seamless combination of these novel designs forms a robust spatiotemporal representation and achieves better performance than state-of-the-art methods on four public motion datasets. Specifically, UMDR achieves unprecedented improvements of ↑ 4.5% on the Chalearn IsoGD dataset.
Benjia Zhou, Pichao Wang, Jun Wan 0001, Yanyan Liang 0001, Fan Wang 0019
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Multi-hypothesis representation learning for transformer-based 3D human pose estimation
abstract
Despite significant progress, estimating 3D human poses from monocular videos remains a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasible solutions ( i.e. , hypotheses) exist. To relieve this limitation, we propose a Multi-Hypothesis Transformer that learns spatio-temporal representations of multiple plausible pose hypotheses. In order to effectively model multi-hypothesis dependencies and build strong relationships across hypothesis features, we introduce a one-to-many-to-one three-stage framework: (i) Generate multiple initial hypothesis representations; (ii) Model self-hypothesis communication, merge multiple hypotheses into a single converged representation and then partition it into several diverged hypotheses; (iii) Learn cross-hypothesis communication and aggregate the multi-hypothesis features to synthesize the final 3D pose. Through the above processes, the final representation is enhanced and the synthesized pose is much more accurate. Extensive experiments show that the proposed method achieves state-of-the-art results on two challenging datasets: Human3.6M and MPI-INF-3DHP. The code and models are available at https://github.com/Vegetebird/MHFormer .
Wenhao Li 0002, Hong Liu 0008, Hao Tang 0005, Pichao Wang
Pattern Recognit.4
2023 BP-triplet net for unsupervised domain adaptation: A Bayesian perspective
Shanshan Wang 0008, Lei Zhang 0038, Pichao Wang, Mengzhu Wang, Xingyi Zhang 0001
Pattern Recognit.3
2023 Exploiting Temporal Contexts With Strided Transformer for 3D Human Pose Estimation
abstract
Despite the great progress in 3D human pose estimation from videos, it is still an open problem to take full advantage of a redundant 2D pose sequence to learn representative representations for generating one 3D pose. To this end, we propose an improved Transformer-based architecture, called Strided Transformer, which simply and effectively lifts a long sequence of 2D joint locations to a single 3D pose. Specifically, a Vanilla Transformer Encoder (VTE) is adopted to model long-range dependencies of 2D pose sequences. To reduce the redundancy of the sequence, fully-connected layers in the feed-forward network of VTE are replaced with strided convolutions to progressively shrink the sequence length and aggregate information from local contexts. The modified VTE is termed as Strided Transformer Encoder (STE), which is built upon the outputs of VTE. STE not only effectively aggregates long-range information to a single-vector representation in a hierarchical global and local fashion, but also significantly reduces the computation cost. Furthermore, a full-to-single supervision scheme is designed at both full sequence and single target frame scales applied to the outputs of VTE and STE, respectively. This scheme imposes extra temporal smoothness constraints in conjunction with the single target frame supervision and hence helps produce smoother and more accurate 3D poses. The proposed Strided Transformer is evaluated on two challenging benchmark datasets, Human3.6 M and HumanEva-I, and achieves state-of-the-art results with fewer parameters. Code and models are available athttps://github.com/Vegetebird/StridedTransformer-Pose3D.
Wenhao Li 0002, Hong Liu 0008, Runwei Ding, Mengyuan Liu 0001, Pichao Wang, Wenming Yang
IEEE Trans. Multim.5
2022 Scaled ReLU Matters for Training Vision Transformers
abstract
Vision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs). However, the training of ViTs is much harder than CNNs, as it is sensitive to the training parameters, such as learning rate, optimizer and warmup epoch. The reasons for training difficulty are empirically analysed in the paper Early Convolutions Help Transformers See Better, and the authors conjecture that the issue lies with the patchify-stem of ViT models. In this paper, we further investigate this problem and extend the above conclusion: only early convolutions do not help for stable training, but the scaled ReLU operation in the convolutional stem (conv-stem) matters. We verify, both theoretically and empirically, that scaled ReLU in conv-stem not only improves training stabilization, but also increases the diversity of patch tokens, thus boosting peak performance with a large margin via adding few parameters and flops. In addition, extensive experiments are conducted to demonstrate that previous ViTs are far from being well trained, further showing that ViTs have great potential to be a better substitute of CNNs.
Pichao Wang, Xue Wang 0010, Hao Luo 0004, Jingkai Zhou, Fan Wang 0019, Hao Li 0030, Rong Jin 0001
AAAI1
2022 Focal and Global Spatial-Temporal Transformer for Skeleton-Based Action Recognition
Zhimin Gao, Peitao Wang, Pei Lv, Xiaoheng Jiang, Qidong Liu 0001, Pichao Wang, Mingliang Xu 0001, Wanqing Li 0001
ACCV (4)6
2022 EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose Estimation
abstract
Locating 3D objects from a single RGB image via Perspective-n-Points (PnP) is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest interpreting PnP as a differentiable layer, so that 2D-3D point correspondences can be partly learned by backpropagating the gradient w.r.t. object pose. Yet, learning the entire set of unrestricted 2D-3D points from scratch fails to converge with existing approaches, since the deterministic pose is inherently non-differentiable. In this paper, we propose the EPro-PnP a probabilistic PnP layer for general end-to-end pose estimation, which outputs a distribution of pose on the SE(3) manifold, essentially bringing categorical Softmax to the continuous domain. The 2D-3D coordinates and corresponding weights are treated as intermediate variables learned by minimizing the KL divergence between the predicted and target pose distribution. The underlying principle unifies the existing approaches and resembles the attention mechanism. EPro-PnP significantly outperforms competitive baselines, closing the gap between PnP-based method and the task-specific leaders on the LineMOD 6DoF pose estimation and nuScenes 3D object detection benchmarks.3
Hansheng Chen 0001, Pichao Wang, Fan Wang 0019, Wei Tian 0001, Lu Xiong 0001, Hao Li 0030
CVPR2
2022 MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
abstract
Estimating 3D human poses from monocular videos is a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasible solutions (i.e., hypotheses) exist. To relieve this limitation, we propose a Multi-Hypothesis Transformer (MHFormer) that learns spatio-temporal representations of multiple plausible pose hypotheses. In order to effectively model multi-hypothesis dependencies and build strong relationships across hypothesis features, the task is decomposed into three stages: (i) Generate multiple initial hypothesis representations; (ii) Model self-hypothesis communication, merge multiple hypotheses into a single converged representation and then partition it into several diverged hypotheses; (iii) Learn cross-hypothesis communication and aggregate the multi-hypothesis features to synthesize the final 3D pose. Through the above processes, the final representation is enhanced and the synthesized pose is much more accurate. Extensive experiments show that MHFormer achieves state-of-the-art results on two challenging datasets: Human3.6M and MPI-INF-3DHP. Without bells and whistles, its performance surpasses the previous best result by a large margin of 3% on Human3.6M. Code and models are available at https://github.com/Vegetebird/MHFormer.
Wenhao Li 0002, Hong Liu 0008, Hao Tang 0005, Pichao Wang, Luc Van Gool
CVPR4
2022 Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion Recognition
abstract
Decoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising performance through the tightly coupled multi-modal spatiotemporal representation, they still suffer from (i) optimization difficulty under small data setting due to the tightly spatiotemporal-entangled modeling; (ii) information redundancy as it usually contains lots of marginal information that is weakly relevant to classification; and (iii) low interaction between multi-modal spatiotemporal information caused by insufficient late fusion. To alleviate these drawbacks, we propose to decouple and recouple spatiotemporal representation for RGB-D-based motion recognition. Specifically, we disentangle the task of learning spatiotemporal representation into 3 sub-tasks: (1) Learning high-quality and dimension independent features through a decoupled spatial and temporal modeling network. (2) Recoupling the decoupled representation to establish stronger space-time dependency. (3) Introducing a Cross-modal Adaptive Posterior Fusion (CAPF) mechanism to capture cross-modal spatiotemporal information from RGB-D data. Seamless combination of these novel designs forms a robust spatiotemporal representation and achieves better performance than state-of-the-art methods on four public motion datasets. Our code is available at https://github.com/damo-cv/MotionRGBD.
Benjia Zhou, Pichao Wang, Jun Wan 0001, Yanyan Liang 0001, Fan Wang 0019, Zhen Lei 0001, Hao Li 0030, Rong Jin 0001
CVPR2
2022 KVT: k-NN Attention for Boosting Vision Transformers
Pichao Wang, Xue Wang 0010, Fan Wang 0019, Ming Lin 0002, Shuning Chang, Hao Li 0030, Rong Jin 0001
ECCV (24)1
2022 TransFGU: A Top-Down Approach to Fine-Grained Unsupervised Semantic Segmentation
Zhaoyuan Yin, Pichao Wang, Fan Wang 0019, Hanling Zhang, Hao Li 0030, Rong Jin 0001
ECCV (29)2
2022 Image-to-Video Re-Identification via Mutual Discriminative Knowledge Transfer
abstract
The gap in representations between image and video makes Image-to-Video Re-identification (I2V Re-ID) challenging, and recent works formulate this problem as a knowledge distillation (KD) process. In this paper, we propose a mutual discriminative knowledge distillation framework to transfer a video-based richer representation to an image based representation more effectively. Specifically, we propose the triplet contrast loss (TCL), a novel loss designed for KD. During the KD process, the TCL loss transfers the local structure, exploits the higher order information, and mitigates the misalignment of the heterogeneous output of teacher and student networks. Compared with other losses for KD, the proposed TCL loss selectively transfers the local discriminative features from teacher to student, making it effective in the ReID. Besides the TCL loss, we adopt mutual learning to regularize both the teacher and student networks training. Extensive experiments demonstrate the effectiveness of our method on the MARS, DukeMTMC-VideoReID and VeRi-776 benchmarks.
Pichao Wang, Fan Wang 0019, Hao Li 0030
ICASSP1
2022 CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation
Tongkun Xu, Pichao Wang, Fan Wang 0019, Hao Li 0030, Rong Jin 0001
ICLR3
2022 VTC-LFC: Vision Transformer Compression with Low-Frequency Components
abstract
Although Vision transformers (ViTs) have recently dominated many vision tasks, deploying ViT models on resource-limited devices remains a challenging problem. To address such a challenge, several methods have been proposed to compress ViTs. Most of them borrow experience in convolutional neural networks (CNNs) and mainly focus on the spatial domain. However, the compression only in the spatial domain suffers from a dramatic performance drop without fine-tuning and is not robust to noise, as the noise in the spatial domain can easily confuse the pruning criteria, leading to some parameters/channels being pruned incorrectly. Inspired by recent findings that self-attention is a low-pass filter and low-frequency signals/components are more informative to ViTs, this paper proposes compressing ViTs with low-frequency components. Two metrics named low-frequency sensitivity (LFS) and low-frequency energy (LFE) are proposed for better channel pruning and token pruning. Additionally, a bottom-up cascade pruning scheme is applied to compress different dimensions jointly. Extensive experiments demonstrate that the proposed method could save 40% ~ 60% of the FLOPs in ViTs, thus significantly increasing the throughput on practical devices with less than 1% performance drop on ImageNet-1K.
Zhenyu Wang 0008, Hao Luo 0004, Pichao Wang, Fan Wang 0019, Hao Li 0030
NeurIPS3
2022 Class-Aware Feature Aggregation Network for Video Object Detection
abstract
Recent progress in video object detection (VOD) has shown that aggregating features from other frames to capture long-range contextual information is very important to deal with the challenges in VOD, such as partial occlusion, motion blur, etc. To exploit more effective feature aggregation, we propose several improvements over previous works in this paper: (1) a class-aware pixel-level feature aggregation module, which characterizes a pixel by exploiting the context information lying in the instances from both the current frame and other frames. Different from the previous non-local operation, the proposed class-aware pixel-level feature aggregation filters out most of the noisy information from the large scope of background and objects in different classes, and only enhances representation of a foreground pixel with the same class instances with limited ambiguous information; (2) a class-aware instance-level feature aggregation module, which aggregates features for object proposals by learning two kinds of relations: the temporal dependencies among the same class object proposals from support frames sampled in a long time range or even the whole sequence, and spatial topology relation among proposals of different objects in the target frame. The homogeneity constraint in instance-level feature aggregation filters out many defective proposals, making the feature aggregation more accurate; and (3) a correlation-based feature alignment module embedded in the instance-level feature aggregation, which aligns the feature maps of the support and target proposals. Without bells or whistles, the proposed method achieves state-of-the-art performance on the ImageNet VID dataset without any post-processing methods. This project is publicly availablehttps://github.com/LiangHann/Class-aware-Feature-Aggregation-Network-for-Video-Object-Detection.
Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030
IEEE Trans. Circuits Syst. Video Technol.2
2021 TransReID: Transformer-based Object Re-Identification
abstract
Extracting robust feature representation is one of the key challenges in object re-identification (ReID). Although convolution neural network (CNN)-based methods have achieved great success, they only process one local neighborhood at a time and suffer from information loss on details caused by convolution and downsampling operators (e.g. pooling and strided convolution). To overcome these limitations, we propose a pure transformer-based object ReID framework named TransReID. Specifically, we first encode an image as a sequence of patches and build a transformer-based strong baseline with a few critical improvements, which achieves competitive results on several ReID benchmarks with CNN-based methods. To further enhance the robust feature learning in the context of transformers, two novel modules are carefully designed. (i) The jigsaw patch module (JPM) is proposed to rearrange the patch embeddings via shift and patch shuffle operations which generates robust features with improved discrimination ability and more diversified coverage. (ii) The side information embeddings (SIE) is introduced to mitigate feature bias towards camera/view variations by plugging in learnable embeddings to incorporate these non-visual clues. To the best of our knowledge, this is the first work to adopt a pure transformer for ReID research. Experimental results of TransReID are superior promising, which achieve state-of-the-art performance on both person and vehicle ReID benchmarks. Code is available at https://github.com/heshuting555/TransReID.
Shuting He, Hao Luo 0004, Pichao Wang, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009
ICCV3
2021 Zen-NAS: A Zero-Shot NAS for High-Performance Image Recognition
abstract
Accuracy predictor is a key component in Neural Architecture Search (NAS) for ranking architectures. Building a high-quality accuracy predictor usually costs enormous computation. To address this issue, instead of using an accuracy predictor, we propose a novel zero-shot index dubbed Zen-Score to rank the architectures. The Zen-Score represents the network expressivity and positively correlates with the model accuracy. The calculation of Zen-Score only takes a few forward inferences through a randomly initialized network, without training network parameters. Built upon the Zen-Score, we further propose a new NAS algorithm, termed as Zen-NAS, by maximizing the Zen-Score of the target network under given inference budgets. Within less than half GPU day, Zen-NAS is able to directly search high performance architectures in a data-free style. Comparing with previous NAS methods, the proposed Zen-NAS is magnitude times faster on multiple server-side and mobile-side GPU platforms with state-of-the-art accuracy on ImageNet. Searching and training code as well as pre-trained models are available from https://github.com/idstcv/ZenNAS.
Ming Lin 0002, Pichao Wang, Zhenhong Sun, Hesen Chen, Xiuyu Sun, Qi Qian 0001, Hao Li 0030, Rong Jin 0001
ICCV2
2021 Context and Structure Mining Network for Video Object Detection
Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030
Int. J. Comput. Vis.2
2021 Transformer guided geometry model for flow-based unsupervised visual odometry
Xiangyu Li 0009, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu 0001, Wanqing Li 0001
Neural Comput. Appl.3
2021 TransRPPG: Remote Photoplethysmography Transformer for 3D Mask Face Presentation Attack Detection
abstract
3D mask face presentation attack detection (PAD) plays a vital role in securing face recognition systems from emergent 3D mask attacks. Recently, remote photoplethysmography (rPPG) has been developed as an intrinsic liveness clue for 3D mask PAD without relying on the mask appearance. However, the rPPG features for 3D mask PAD are still needed expert knowledge to design manually, which limits its further progress in the deep learning and big data era. In this letter, we propose a pure rPPG transformer (TransRPPG) framework for learning intrinsic liveness representation efficiently. At first, rPPG-based multi-scale spatial-temporal maps (MSTmap) are constructed from facial skin and background regions. Then the transformer fully mines the global relationship within MSTmaps for liveness representation, and gives a binary prediction for 3D mask detection. Comprehensive experiments are conducted on two benchmark datasets to demonstrate the efficacy of the TransRPPG on both intra- and cross-dataset testings. Our TransRPPG is lightweight and efficient (with only 547 K parameters and 763 M FLOPs), which is promising for mobile-level applications.
Zitong Yu, Pichao Wang, Guoying Zhao 0001
IEEE Signal Process. Lett.3
2021 Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture Recognition
abstract
Gesture recognition has attracted considerable attention owing to its great potential in applications. Although the great progress has been made recently in multi-modal learning methods, existing methods still lack effective integration to fully explore synergies among spatio-temporal modalities effectively for gesture recognition. The problems are partially due to the fact that the existing manually designed network architectures have low efficiency in the joint learning of multi-modalities. In this paper, we propose the first neural architecture search (NAS)-based method for RGB-D gesture recognition. The proposed method includes two key components: 1) enhanced temporal representation via the proposed 3D Central Difference Convolution (3D-CDC) family, which is able to capture rich temporal context via aggregating temporal difference information; and 2) optimized backbones for multi-sampling-rate branches and lateral connections among varied modalities. The resultant multi-modal multi-rate network provides a new perspective to understand the relationship between RGB and depth modalities and their temporal dynamics. Comprehensive experiments are performed on three benchmark datasets (IsoGD, NvGesture, and EgoGesture), demonstrating the state-of-the-art performance in both single- and multi-modality settings. The code is available at https://github.com/ZitongYu/3DCDC-NAS.
Zitong Yu, Benjia Zhou, Jun Wan 0001, Pichao Wang, Haoyu Chen 0001, Xin Liu 0012, Stan Z. Li, Guoying Zhao 0001
IEEE Trans. Image Process.4
2021 BR$^2$Net: Defocus Blur Detection Via a Bidirectional Channel Attention Residual Refining Network
abstract
Due to the remarkable potential applications, defocus blur detection, which aims to separate blurry regions from an image, has attracted much attention. Although significant progress has been made by many methods, there are still various challenges that hinder the results, e.g., confusing background areas, sensitivity to the scale and missing the boundary details of the defocus blur regions. To solve these issues, in this paper, we propose a deep convolutional neural network (CNN) for defocus blur detection via a Bi-directional Residual Refining network (BR2Net). Specifically, a residual learning and refining module (RLRM) is designed to correct the prediction errors in the intermediate defocus blur map. Then, we develop a bidirectional residual feature refining network with two branches by embedding multiple RLRMs into it for recurrently combining and refining the residual features. One branch of the network refines the residual features from the shallow layers to the deep layers, and the other branch refines the residual features from the deep layers to the shallow layers. In such a manner, both the low-level spatial details and highlevel semantic information can be encoded step by step in two directions to suppress background clutter and enhance the detected region details. The outputs of the two branches are fused to generate the final results. In addition, with the observation that different feature channels have different extents of discrimination for detecting blurred regions, we add a channel attention module to each feature extraction layer to select more discriminative features for residual learning. To promote further research on defocus blur detection, we create a new dataset with various challenging images and manually annotate their corresponding pixelwise ground truths. The proposed network is validated on two commonly used defocus blur detection datasets and our newly collected dataset by comparing it with 10 other state-of-the-art methods. Extensive experiments with ablation studies demonstrate that BR2Net consistently and significantly outperforms the competitors in terms of both the efficiency and accuracy.
Chang Tang, Xinwang Liu 0002, Shan An, Pichao Wang
IEEE Trans. Multim.4
2020 R²MRF: Defocus Blur Detection via Recurrently Refining Multi-Scale Residual Features
abstract
Defocus blur detection aims to separate the in-focus and out-of-focus regions in an image. Although attracting more and more attention due to its remarkable potential applications, there are still several challenges for accurate defocus blur detection, such as the interference of background clutter, sensitivity to scales and missing boundary details of defocus blur regions. In order to address these issues, we propose a deep neural network which Recurrently Refines Multi-scale Residual Features (R2MRF) for defocus blur detection. We firstly extract multi-scale deep features by utilizing a fully convolutional network. For each layer, we design a novel recurrent residual refinement branch embedded with multiple residual refinement modules (RRMs) to more accurately detect blur regions from the input image. Considering that the features from bottom layers are able to capture rich low-level features for details preservation while the features from top layers are capable of characterizing the semantic information for locating blur regions, we aggregate the deep features from different layers to learn the residual between the intermediate prediction and the ground truth for each recurrent step in each residual refinement branch. Since the defocus degree is sensitive to image scales, we finally fuse the side output of each branch to obtain the final blur detection map. We evaluate the proposed network on two commonly used defocus blur detection benchmark datasets by comparing it with other 11 state-of-the-art methods. Extensive experimental results with ablation studies demonstrate that R2MRF consistently and significantly outperforms the competitors in terms of both efficiency and accuracy.
Chang Tang, Xinwang Liu 0002, Xinzhong Zhu, En Zhu, Kun Sun 0002, Pichao Wang, Lizhe Wang 0001, Albert Y. Zomaya
AAAI6
2020 Exploiting Better Feature Aggregation for Video Object Detection
abstract
Video object detection (VOD) has been a rising topic in recent years due to the challenges such as occlusion, motion blur, etc. To deal with these challenges, feature aggregation from local or global support frames is verified effective. To exploit better feature aggregation, in this paper, we propose two improvements over previous works: a class-constrained spatial-temporal relation network and a correlation-based feature alignment module. For the class constrained spatial-temporal relation network, it operates on object region proposals, and learns two kinds of relations: (1) the dependencies among region proposals of the same object class from support frames sampled in a long time range or even the whole sequence, and (2) spatial relations among proposals of different objects in the target frame. The homogeneity constraint in spatial-temporal relation network not only filters out many defective proposals but also implicitly embeds the traditional post-processing strategies (e.g., Seq-NMS), leading to a unified end-to-end training networks. In the feature alignment module, we propose a correlation based feature alignment method to align the support and target frames for feature aggregation in the temporal domain. Our experiments show that the proposed method improves the accuracy of single-frame detectors significantly, and outperforms previous temporal or spatial relation networks. Without bells or whistles, the proposed method achieves state-of-the-art performance on the ImageNet VID dataset (84.80% with ResNet-101) without any post-processing methods.
Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030
ACM Multimedia2
2020 SAR-NAS: Skeleton-based action recognition via neural architecture searching
Yonghong Hou, Pichao Wang, Zihui Guo, Wanqing Li 0001
J. Vis. Commun. Image Represent.3
2020 A Review of Dynamic Maps for 3D Human Motion Recognition Using ConvNets and Its Improvement
Zhimin Gao, Pichao Wang, Huogen Wang, Mingliang Xu 0001, Wanqing Li 0001
Neural Process. Lett.2
2019 Salient Object Detection via Recurrently Aggregating Spatial Attention Weighted Cross-Level Deep Features
abstract
This paper proposes a novel deep saliency detection network by recurrently aggregating and refining features in a cross-level and spatial attention-aware manner. In this way, the features integrated from multiple layers can be used to refine layer-wise features step by step and the complementary information in different layers can be fully captured for detecting salient objects with different scales, i.e., the features integrated from low-level layers can serve to refine the details of detected salient objects while the features integrated from high-level layers with semantic information can benefit the locating of salient objects. In addition, by considering that only partial regions of an image are salient, we embed a spatial attention-aware module to suppress the non-salient regions and highlight salient objects. Finally, different saliency detection results from different layers are fused to generate the final saliency map. Experimental results on five benchmark datasets demonstrate that our proposed method outperforms other 14 state-of-the-art competitors.
Chang Tang, Xinzhong Zhu, Xinwang Liu 0002, Pichao Wang
ICME4
2019 Self-Attention Guided Deep Features for Action Recognition
abstract
Skeleton based human action recognition is an important task in computer vision. However, it is very challenging due to the complex spatio-temporal variations of skeleton joints. In this work, we propose an end-to-end trainable network consisting of a Deep Convolutional Model (DCM) and a Self-Attention Model (SAM) for human action recognition from skeleton data. Specifically, skeleton sequences are encoded into color images and fed into DCM to extract deep features. In the SAM, handcrafted features representing the motion degree of joints are extracted and the attention weights are learned by a simple yet effective linear mapping. The effectiveness of proposed method has been verified on NTU RGB+D, SYSU-3D and UTD-MHAD datasets and achieved state-of-the-art results.
Renyi Xiao, Yonghong Hou, Zihui Guo, Chuankun Li, Pichao Wang, Wanqing Li 0001
ICME5
2019 Light Weight Stereo Matching via Deep Extraction and Integration of Low and High Level Information
abstract
Deep convolutional neural networks (CNN) have demonstrated remarkable progress in stereo matching recently. However, disparity estimation in the ill-posed regions is still difficult. In addition, CNN based stereo matching methods often have impractical computational complexity and memory consumption. To address these problems we propose an end-to-end light weight CNN architecture to effectively learn and integrate low and high level information. To achieve this, a novel enhancement block built upon group convolution and dilated-convolution is proposed. Compared with state-of-the-art methods, the proposed method achieved competitive performance with the least number of network parameters on the Flyingthings3d and KITTI datasets.
Yonghong Hou, Pichao Wang, Zhongyu Jiang, Wanqing Li 0001
ICME3
2019 DVONet: Unsupervised Monocular Depth Estimation and Visual Odometry
abstract
This paper proposes an unsupervised learning framework for monocular depth estimation and visual odometry (VO), referred to as DVONet. The framework is trained using stereo image sequences and is able to estimate absolute-scale scene depth and camera poses from monocular images. To mitigate the effect of stereo occlusions in training and improve the depth estimation, left-right occlusion mask is introduced. In addition, a novel VO network is proposed where the feature extraction network is shared between pose estimation and optical flow estimation. The proposed DVONet achieves state-of-the-art results for both depth estimation and VO tasks on the KITTI driving dataset, outperforming the existing unsupervised methods and being comparable to the traditional ones.
Xiangyu Li 0009, Yonghong Hou, Pichao Wang, Wanqing Li 0001
VCIP4
2019 Learning attentive dynamic maps (ADMs) for Understanding Human Actions
Chuankun Li, Yonghong Hou, Wanqing Li 0001, Pichao Wang
J. Vis. Commun. Image Represent.4
2019 Unsupervised feature selection via latent representation learning and manifold regularization
Chang Tang, Meiru Bian, Xinwang Liu 0002, Miaomiao Li 0001, Pichao Wang, Hailin Yin
Neural Networks6
2019 Multiview-Based 3-D Action Recognition Using Deep Networks
abstract
In multiview learning, views may be obtained from multiple sources or extracted from a single source as different features. In this paper, effective multiple views from skeleton sequences are proposed to learn the discriminative features using multiple networks for three-dimensional human action recognition. Specifically, three views are constructed in the spatial domain and fed to a stack of long short-term memory networks to exploit temporal information and three views are constructed using the improved joint trajectory maps and fed to three convolutional neural networks to exploit spatial information. Multiply fusion is used to combine the recognition scores of all views. The proposed method has been verified and achieved the state-of-the-art results on the widely used UTD-MHAD, MSRC-12 Kinect Gesture, and NTU red, green, blue (RGB)+D datasets.
Chuankun Li, Yonghong Hou, Pichao Wang, Wanqing Li 0001
IEEE Trans. Hum. Mach. Syst.3
2019 Adaptive Hypergraph Embedded Semi-Supervised Multi-Label Image Annotation
abstract
Multilabel image annotation attracts a lot of research interest due to its practicability in multimedia and computer vision fields, while the need for a large amount of labeled training data to achieve promising performance makes it a challenging task. Fortunately, unlabeled and relevant data are widely available and these data can be used to serve the annotation task. To this end, we propose a novel adaptive hypergraph learning (AHL) method for multilabel image annotation in a semisupervised way, in which both the limited labeled data and abundant unlabeled data are utilized to facilitate the annotation performance. In detail, we seek a multilabel propagation scheme by learning a hypergraph which is used to preserve the local geometric structures of data in a high-order manner. Meanwhile, a feature projection is integrated into AHL to obtain a latent feature space where unlabeled instances can be effectively and robustly assigned with multiple labels. Experiments on six widely used image datasets are conducted to evaluate our model and the results demonstrate that the proposed AHL outperforms other state-of-the-art semisupervised methods.
Chang Tang, Xinwang Liu 0002, Pichao Wang, Changqing Zhang 0002, Miaomiao Li 0001, Lizhe Wang 0001
IEEE Trans. Multim.3
2019 Learning a Joint Affinity Graph for Multiview Subspace Clustering
abstract
With the ability to exploit the internal structure of data, graph-based models have received a lot of attention and have achieved great success in multiview subspace clustering for multimedia data. Most of the existing methods individually construct an affinity graph for each single view and fuse the result obtained from each single graph. However, the common representation shared by different views and the complementary diversity across these views are not efficiently exploited. In addition, noise and outliers are often mixed in original data, which adversely degenerate the clustering performance of many existing methods. In this paper, we propose addressing these issues by learning a joint affinity graph for multiview subspace clustering based on a low-rank representation with diversity regularization and a rank constraint. Specifically, a low-rank representation model is employed to learn a shared sample representation coefficient matrix to generate the affinity graph. At the same time, we use diversity regularization to learn the optimal weights for each view, which can suppress the redundancy and enhance the diversity among different feature views. In addition, the cluster number is used to promote affinity graph learning by using a rank constraint. The final clustering result is obtained by using normalized cuts on the learned affinity graph. An efficient algorithm based on an augmented Lagrangian multiplier with alternating direction minimization is carefully designed to solve the resulting optimization problem. Extensive experiments on various real-world datasets are conducted, and the results demonstrate well the effectiveness of the proposed algorithm.
Chang Tang, Xinzhong Zhu, Xinwang Liu 0002, Miaomiao Li 0001, Pichao Wang, Changqing Zhang 0002, Lizhe Wang 0001
IEEE Trans. Multim.5
2018 Cooperative Training of Deep Aggregation Networks for RGB-D Action Recognition
abstract
A novel deep neural network training paradigm that exploits the conjoint information in multiple heterogeneous sources is proposed. Specifically, in a RGB-D based action recognition task, it cooperatively trains a single convolutional neural network (named c-ConvNet) on both RGB visual features and depth features, and deeply aggregates the two kinds of features for action recognition. Differently from the conventional ConvNet that learns the deep separable features for homogeneous modality-based classification with only one softmax loss function, the c-ConvNet enhances the discriminative power of the deeply learned features and weakens the undesired modality discrepancy by jointly optimizing a ranking loss and a softmax loss for both homogeneous and heterogeneous modalities. The ranking loss consists of intra-modality and cross-modality triplet losses, and it reduces both the intra-modality and cross-modality feature variations. Furthermore, the correlations between RGB and depth data are embedded in the c-ConvNet, and can be retrieved by either of the modalities and contribute to the recognition in the case even only one of the modalities is available. The proposed method was extensively evaluated on two large RGB-D action recognition datasets, ChaLearn LAP IsoGD and NTU RGB+D datasets, and one small dataset, SYSU 3D HOI, and achieved state-of-the-art results.
Pichao Wang, Wanqing Li 0001, Jun Wan 0001, Philip Ogunbona, Xinwang Liu 0002
AAAI1
2018 RGB-D-based human motion recognition with deep learning: A survey
Pichao Wang, Wanqing Li 0001, Philip Ogunbona, Jun Wan 0001, Sergio Escalera
Comput. Vis. Image Underst.1
2018 Robust graph regularized unsupervised feature selection
abstract
Recent research indicates the critical importance of preserving local geometric structure of data in unsupervised feature selection (UFS), and the well studied graph Laplacian is usually deployed to capture this property. By using a squared l 2 -norm, we observe that conventional graph Laplacian is sensitive to noisy data, leading to unsatisfying data processing performance. To address this issue, we propose a unified UFS framework via feature self-representation and robust graph regularization , with the aim at reducing the sensitivity to outliers from the following two aspects: i) an l 2, 1 -norm is used to characterize the feature representation residual matrix; and ii) an l 1 -norm based graph Laplacian regularization term is adopted to preserve the local geometric structure of data. By this way, the proposed framework is able to reduce the effect of noisy data on feature selection. Furthermore, the proposed l 1 -norm based graph Laplacian is readily extendible, which can be easily integrated into other UFS methods and machine learning tasks with local geometrical structure of data being preserved. As demonstrated on ten challenging benchmark data sets, our algorithm significantly and consistently outperforms state-of-the-art UFS methods in the literature, suggesting the effectiveness of the proposed UFS framework.
Chang Tang, Xinzhong Zhu, Jiajia Chen 0010, Pichao Wang, Xinwang Liu 0002, Jie Tian 0001
Expert Syst. Appl.4
2018 Saliency detection via affinity graph learning and weighted manifold ranking
Xinzhong Zhu, Chang Tang, Pichao Wang, Minhui Wang, Jiajia Chen 0010, Jie Tian 0001
Neurocomputing3
2018 Online human action recognition based on incremental learning of weighted covariance descriptors
Chang Tang, Wanqing Li 0001, Pichao Wang, Lizhe Wang 0001
Inf. Sci.3
2018 Consensus learning guided multi-view unsupervised feature selection
Chang Tang, Jiajia Chen 0010, Xinwang Liu 0002, Miaomiao Li 0001, Pichao Wang, Minhui Wang
Knowl. Based Syst.5
2018 Robust unsupervised feature selection via dual self-representation and manifold regularization
Chang Tang, Xinwang Liu 0002, Miaomiao Li 0001, Pichao Wang, Jiajia Chen 0010, Lizhe Wang 0001, Wanqing Li 0001
Knowl. Based Syst.4
2018 Action recognition based on joint trajectory maps with convolutional neural networks
Pichao Wang, Wanqing Li 0001, Chuankun Li, Yonghong Hou
Knowl. Based Syst.1
2018 Skeleton Optical Spectra-Based Action Recognition Using Convolutional Neural Networks
abstract
This letter presents an effective method to encode the spatiotemporal information of a skeleton sequence into color texture images, referred to as skeleton optical spectra, and employs convolutional neural networks (ConvNets) to learn the discriminative features for action recognition. Such spectrum representation makes it possible to use a standard ConvNet architecture to learn suitable “dynamic” features from skeleton sequences without training millions of parameters afresh and it is especially valuable when there is insufficient annotated training video data. Specifically, the encoding consists of four steps: mapping of joint distribution, spectrum coding of joint trajectories, spectrum coding of body parts, and joint velocity weighted saturation and brightness. Experimental results on three widely used datasets have demonstrated the efficacy of the proposed method.
Yonghong Hou, Pichao Wang, Wanqing Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 Depth Pooling Based Large-Scale 3-D Action Recognition With Convolutional Neural Networks
abstract
This paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as dynamic depth images (DDI), dynamic depth normal images (DDNI), and dynamic depth motion normal images (DDMNI), for both isolated and continuous action recognition. These dynamic images are constructed from a segmented sequence of depth maps using hierarchical bidirectional rank pooling to effectively capture the spatial-temporal information. Specifically, DDI exploits the dynamics of postures over time, and DDNI and DDMNI exploit the 3-D structural information captured by depth maps. Upon the proposed representations, a convolutional neural network (ConvNet)-based method is developed for action recognition. The image-based representations enable us to fine-tune the existing ConvNet models trained on image data without training a large number of parameters from scratch. The proposed method achieved the state-of-art results on three large datasets, namely, the large-scale continuous gesture recognition dataset (means the Jaccard index 0.4109), the large-scale isolated gesture recognition dataset (59.21%), and the NTU RGB+D dataset (87.08% cross-subject and 84.22% cross-view) even though only the depth modality was used.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona
IEEE Trans. Multim.1
2017 Scene Flow to Action Map: A New Representation for RGB-D Based Action Recognition with Convolutional Neural Networks
abstract
Scene flow describes the motion of 3D objects in real world and potentially could be the basis of a good feature for 3D action recognition. However, its use for action recognition, especially in the context of convolutional neural networks (ConvNets), has not been previously studied. In this paper, we propose the extraction and use of scene flow for action recognition from RGB-D data. Previous works have considered the depth and RGB modalities as separate channels and extract features for later fusion. We take a different approach and consider the modalities as one entity, thus allowing feature extraction for action recognition at the beginning. Two key questions about the use of scene flow for action recognition are addressed: how to organize the scene flow vectors and how to represent the long term dynamics of videos based on scene flow. In order to calculate the scene flow correctly on the available datasets, we propose an effective self-calibration method to align the RGB and depth data spatially without knowledge of the camera parameters. Based on the scene flow vectors, we propose a new representation, namely, Scene Flow to Action Map (SFAM), that describes several long term spatio-temporal dynamics for action recognition. We adopt a channel transform kernel to transform the scene flow vectors to an optimal color space analogous to RGB. This transformation takes better advantage of the trained ConvNets models over ImageNet. Experimental results indicate that this new representation can surpass the performance of state-of-the-art methods on two large public datasets.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona
CVPR1
2017 Weakly structured information aggregation for upper-body posture assessment using ConvNets
abstract
Posture assessment aims to determine the risk associated with poor posture and thus avoid injury in subjects. Upper-body posture assessment from images offers an attractive alternative to manual methods by directly extracting relevant features for classification. A deep convolutional neural network is proposed to extract structured features from different body parts and learn shared features that are used to determine the appropriate assessment. The structured features are learned with triplet-based rank constraints based on head and torso separately. The shared feature and assessment function are learned with soft-max constraints based on posture risk measurements. Experimental evaluation on a self-collected upper-body posture dataset has verified the efficacy of the proposed method and network architecture.
Zewei Ding, Wanqing Li 0001, Pichao Wang, Philip Ogunbona
ICME3
2017 Joint Distance Maps Based Action Recognition With Convolutional Neural Networks
abstract
Motivated by the promising performance achieved by deep learning, an effective yet simple method is proposed to encode the spatio-temporal information of skeleton sequences into color texture images, referred to as joint distance maps (JDMs), and convolutional neural networks are employed to exploit the discriminative features from the JDMs for human action and interaction recognition. The pair-wise distances between joints over a sequence of single or multiple person skeletons are encoded into color variations to capture temporal information. The efficacy of the proposed method has been verified by the state-of-the-art results on the large RGB+D Dataset and small UTD-MHAD Dataset in both single-view and cross-view settings.
Chuankun Li, Yonghong Hou, Pichao Wang, Wanqing Li 0001
IEEE Signal Process. Lett.3
2017 Salient Object Detection via Weighted Low Rank Matrix Recovery
abstract
Image-based salient object detection is a useful and important technique, which can promote the efficiency of several applications such as object detection, image classification/retrieval, object co-segmentation, and content-based image editing. In this letter, we present a novel weighted low-rank matrix recovery (WLRR) model for salient object detection. In order to facilitate efficient salient objects-background separation, a high-level background prior map is estimated by employing the property of the color, location, and boundary connectivity, and then this prior map is ensembled into a weighting matrix which indicates the likelihood that each image region belongs to the background. The final salient object detection task is formulated as the WLRR model with the weighting matrix. Both quantitative and qualitative experimental results on three challenging datasets show competitive results as compared with 24 state-of-the-art methods.
Chang Tang, Pichao Wang, Changqing Zhang 0002, Wanqing Li 0001
IEEE Signal Process. Lett.2
2016 Large-scale Isolated Gesture Recognition using Convolutional Neural Networks
abstract
This paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as Dynamic Depth Images (DDI), Dynamic Depth Normal Images (DDNI) and Dynamic Depth Motion Normal Images (DDMNI). These dynamic images are constructed from a sequence of depth maps using bidirectional rank pooling to effectively capture the spatial-temporal information. Such image-based representations enable us to fine-tune the existing ConvNets models trained on image data for classification of depth sequences, without introducing large parameters to learn. Upon the proposed representations, a convolutional Neural networks (ConvNets) based method is developed for gesture recognition and evaluated on the Large-scale Isolated Gesture Recognition at the ChaLearn Looking at People (LAP) challenge 2016. The method achieved 55.57% classification accuracy and ranked 2ndplace in this challenge but was very close to the best performance even though we only used depth data.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona
ICPR1
2016 Large-scale Continuous Gesture Recognition Using Convolutional Neural Networks
abstract
This paper addresses the problem of continuous gesture recognition from sequences of depth maps using Convolutional Neural networks (ConvNets). The proposed method first segments individual gestures from a depth sequence based on quantity of movement (QOM). For each segmented gesture, an Improved Depth Motion Map (IDMM), which converts the depth sequence into one image, is constructed and fed to a ConvNet for recognition. The IDMM effectively encodes both spatial and temporal information and allows the fine-tuning with existing ConvNet models for classification without introducing millions of parameters to learn. The proposed method is evaluated on the Large-scale Continuous Gesture Recognition of the ChaLearn Looking at People (LAP) challenge 2016. It achieved the performance of 0.2655 (Mean Jaccard Index) and ranked 3rdplace in this challenge.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Philip Ogunbona
ICPR1
2016 Action Recognition Based on Joint Trajectory Maps Using Convolutional Neural Networks
abstract
Recently, Convolutional Neural Networks (ConvNets) have shown promising performances in many computer vision tasks, especially image-based recognition. How to effectively use ConvNets for video-based recognition is still an open problem. In this paper, we propose a compact, effective yet simple method to encode spatio-temporal information carried in 3D skeleton sequences into multiple 2D images, referred to as Joint Trajectory Maps (JTM), and ConvNets are adopted to exploit the discriminative features for real-time human action recognition. The proposed method has been evaluated on three public benchmarks, i.e., MSRC-12 Kinect gesture dataset (MSRC-12), G3D dataset and UTD multimodal human action dataset (UTD-MHAD) and achieved the state-of-the-art results.
Pichao Wang, Yonghong Hou, Wanqing Li 0001
ACM Multimedia1
2016 Salient object detection using color spatial distribution and minimum spanning tree weight
Chang Tang, Chunping Hou, Pichao Wang, Zhanjie Song
Multim. Tools Appl.3
2016 RGB-D-based action recognition datasets: A survey
Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona, Pichao Wang, Chang Tang
Pattern Recognit.4
2016 A Spectral and Spatial Approach of Coarse-to-Fine Blurred Image Region Detection
abstract
Blur exists in many digital images, it can be mainly categorized into two classes: defocus blur which is caused by optical imaging systems and motion blur which is caused by the relative motion between camera and scene objects. In this letter, we propose a simple yet effective automatic blurred image region detection method. Based on the observation that blur attenuates high-frequency components of an image, we present a blur metric based on the log averaged spectrum residual to get a coarse blur map. Then, a novel iterative updating mechanism is proposed to refine the blur map from coarse to fine by exploiting the intrinsic relevance of similar neighbor image regions. The proposed iterative updating mechanism can partially resolve the problem of differentiating an in-focus smooth region and a blurred smooth region. In addition, our iterative updating mechanism can be integrated into other image blurred region detection algorithms to refine the final results. Both quantitative and qualitative experimental results demonstrate that our proposed method is more reliable and efficient compared to various state-of-the-art methods.
Chang Tang, Yonghong Hou, Pichao Wang, Wanqing Li 0001
IEEE Signal Process. Lett.4
2016 Action Recognition From Depth Maps Using Deep Convolutional Neural Networks
abstract
This paper proposes a new method, i.e., weighted hierarchical depth motion maps (WHDMM) + three-channel deep convolutional neural networks (3ConvNets), for human action recognition from depth maps on small training datasets. Three strategies are developed to leverage the capability of ConvNets in mining discriminative features for recognition. First, different viewpoints are mimicked by rotating the 3-D points of the captured depth maps. This not only synthesizes more data, but also makes the trained ConvNets view-tolerant. Second, WHDMMs at several temporal scales are constructed to encode the spatiotemporal motion patterns of actions into 2-D spatial structures. The 2-D spatial structures are further enhanced for recognition by converting the WHDMMs into pseudocolor images. Finally, the three ConvNets are initialized with the models obtained from ImageNet and fine-tuned independently on the color-coded WHDMMs constructed in three orthogonal planes. The proposed algorithm was evaluated on the MSRAction3D, MSRAction3DExt, UTKinect-Action, and MSRDailyActivity3D datasets using cross-subject protocols. In addition, the method was evaluated on the large dataset constructed from the above datasets. The proposed method achieved 2-9% better results on most of the individual datasets. Furthermore, the proposed method maintained its performance on the large dataset, whereas the performance of existing methods decreased with the increased number of actions.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Jing Zhang 0017, Chang Tang, Philip Ogunbona
IEEE Trans. Hum. Mach. Syst.1
2015 ConvNets-Based Action Recognition from Depth Maps through Virtual Cameras and Pseudocoloring
abstract
In this paper, we propose to adopt ConvNets to recognize human actions from depth maps on relatively small datasets based on Depth Motion Maps (DMMs). In particular, three strategies are developed to effectively leverage the capability of ConvNets in mining discriminative features for recognition. Firstly, different viewpoints are mimicked by rotating virtual cameras around subject represented by the 3D points of the captured depth maps. This not only synthesizes more data from the captured ones, but also makes the trained ConvNets view-tolerant. Secondly, DMMs are constructed and further enhanced for recognition by encoding them into Pseudo-RGB images, turning the spatial-temporal motion patterns into textures and edges. Lastly, through transferring learning the models originally trained over ImageNet for image classification, the three ConvNets are trained independently on the color-coded DMMs constructed in three orthogonal planes. The proposed algorithm was extensively evaluated on MSRAction3D, MSRAction3DExt and UTKinect-Action datasets and achieved the state-of-the-art results on these datasets.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Jing Zhang 0017, Philip Ogunbona
ACM Multimedia1