Xiaojiang Peng

dblp:133/6556 · DBLP profile ↗
← Back
86ranked-venue papers
9as first author
50since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 8 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 7 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Noisy correspondence decomposition for robust composed image retrieval
Zhaopan Xu, Chenrui Zhou, Xi Chen 0110, Wangbo Zhao, Jianning Zhang, Xiaojiang Peng, Hongxun Yao, Kaipeng Zhang
Expert Syst. Appl.8
2026 3D landmark detection on human point clouds: A benchmark and a dual cascade point transformer framework
Fan Zhang 0111, Shuyi Mao, Xiaojiang Peng
Expert Syst. Appl.4
2026 A survey of robotic manipulation: From bottom-up approaches to end-to-end paradigms with LLMs
Qing Li 0001, Zhijian He, Bowen Zhang 0005, Xianghua Fu, Zhi-Qi Cheng, Yan Yan 0001, Xiaojiang Peng
Neurocomputing10
2026 Generative models for noise-robust training in unsupervised domain adaptation
Zhongying Deng, Da Li 0001, Junjun He, Xiaojiang Peng, Yi-Zhe Song, Tao Xiang 0002
Pattern Recognit.4
2026 BORT2: Bi-level optimization for robust target training in multi-source domain adaptation
Zhongying Deng, Da Li 0001, Xiaojiang Peng, Yi-Zhe Song, Tao Xiang 0002
Pattern Recognit.3
2026 MEERA: Multimodal experts for evidence routing and sparse LLM adaptation in sentiment analysis
Fuqiang Niu, Xiaojiang Peng, Hu Huang 0009, Bowen Zhang 0005
Pattern Recognit.3
2026 Learning With Dual Noisy Labels for Text-to-Image Person Re-Identification
abstract
Text-to-image person re-identification (TIReID) aims to identify a target person from a given textual description. Although recent work has made significant progress, most of it implicitly assumes that the sample annotations are correct and that the cross-modal correspondence in each image-text pair is well aligned. However, such an assumption requires elaborately annotated datasets, which are expensive and even impossible to obtain in practice. To alleviate this issue, in this letter, we explore a new TIReID setting, termed learning with dual noisy labels, in which the model learns from data with both noisy identity labels and noisy correspondence. We propose a general framework called TDTD (Two stage framework forDual noise ofTIReID) to achieve this. In the first stage, a Noise-Aware Preliminary Learning (NAPL) strategy selects “easy” triplets to train a noise-tolerant initial model. In the second stage, the model leverages reliable representations from NAPL to automatically correct both identity and correspondence errors via soft-label estimation and is then fine-tuned on the entire dataset using a dual noise-robust triplet loss. Extensive experiments on three public benchmarks, CUHK-PEDES, ICFG-PEDES, and RSTPReID, demonstrate the performance and robustness of TDTD, achieving state-of-the-art results under dual noise conditions.
Zhaopan Xu, Wangbo Zhao, Xiaojiang Peng, Hongxun Yao
IEEE Signal Process. Lett.5
2025 DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal Retrieval
abstract
With the burst of big data, 2D-3D cross-modal retrieval has received increasing attention, which aims to retrieve relevant data from one modality given the query from the other modality. In this paper, we study an underexplored yet practical problem of semi-supervised 2D-3D cross-modal retrieval, which could suffer from serious label scarcity in real-world applications. Moreover, the huge heterogeneous gap could deteriorate the process of learning from unlabeled data. In this work, we propose a novel approach named Decoupled Discriminative Learning with Bigraph-aware Alignment (DREAM) for semi-supervised 2D-3D cross-modal retrieval. The core of our DREAM is to decouple the label prediction and reliability measurement processes to reduce overconfident samples in discriminative learning. In particular, we enhance a label prediction module with label propagation from labeled samples and additionally introduce a reliability measurement module to learn the scores of predicted labels. To reduce class-related bias, we compare reliability scores with class-specific adaptive thresholds to identify samples for additional learning. In addition, negative labels are estimated for unselected samples, which guides soft semantic learning to make the best use of all the information. To further minimize the heterogeneous gap, we build a bigraph graph that connects cross-modal similar examples and then conduct learning to cluster with most edges kept for alignment. Extensive experiments on several benchmark datasets validate the superiority of the proposed DREAM.
Fan Zhang 0111, Changhu Wang, Zebang Cheng, Xiaojiang Peng, Dongjie Wang 0001, Yijia Xiao, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001
AAAI4
2025 Mamba-Enhanced Text-Audio-Video Alignment Network for Emotion Recognition in Conversations
Xiaomao Fan, Qingyang Wu, Xiaojiang Peng, Ye Li 0002
ADMA (3)4
2025 A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training
abstract
Training diffusion models is always a computation-intensive task. In this paper, we introduce a novel speed-up method for diffusion model training, called SpeeD, which is based on a closer look at time steps. Our key findings are: i) Time steps can be empirically divided into acceleration, deceleration, and convergence areas based on the process increment. ii) These time steps are imbalanced, with many concentrated in the convergence area. iii) The concentrated steps provide limited benefits for diffusion training. To address this, we design an asymmetric sampling strategy that reduces the frequency of steps from the convergence area while increasing the sampling probability in other areas. Additionally, we propose a weighting strategy to emphasize the importance of time steps with rapid-change process increments. As a plug-and-play and architecture-agnostic approach, SpeeD consistently achieves 3 × acceleration across various diffusion architectures, datasets, and tasks. Notably, due to its simple design, our approach significantly reduces the cost of diffusion model training with minimal overhead. Our research enables more researchers to train diffusion models at a lower cost.
Kai Wang 0036, Mingjia Shi, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Hanwang Zhang, Yang You 0001
CVPR7
2025 Enhancing 6D Pose Estimation with Cross-modal Fusion Network and Density-peak Keypoint Localization
abstract
Current dual-fusion models for 6D pose estimation often lead to increased computational complexity and risk of overfitting with the addition of more networks. To address this, we propose a Cross-modal Fusion Network (CFN), which extracts robust dual-modal features while reducing computation energy and overfitting risks. The CFN consists of multiple Cross-modal Fusion Modules (CFM), featuring two key components: 1) the Spiking-based Cross-Attention Block (SCA), which utilizes only mask and addition operations, significantly lowering computational energy compared to traditional self-attention; and 2) the Specificity Preserving Block (SPB), designed to mitigate overfitting from multiple CFM layers. Additionally, we introduce a density-peak keypoint localization (DKL) method that resists noise and sparse data, eliminating the need for iterative processes. Extensive experiments on multiple 6D pose estimation benchmarks demonstrate that our CFN method significantly outperforms existing state-of-the-art approaches.
Liming Zhang 0007, Chuan Yan, Xiaojiang Peng
ICASSP5
2025 UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
abstract
Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional labels or single-modality input. In this paper, we introduce the Unified Multimodal Prompt-Induced Emotional Text-to-Speech System (UMETTS), a novel framework that leverages emotional cues from multiple modalities to generate highly expressive and emotionally resonant speech. The core of UMETTS consists of two key components: the Emotion Prompt Alignment Module (EP-Align) and the Emotion Embedding-Induced TTS Module(EMI-TTS). (1) EP-Align employs contrastive learning to align emotional features across text, audio, and visual modalities, ensuring a coherent fusion of multimodal information. (2) Subsequently, EMI-TTS integrates the aligned emotional embeddings with state-of-the-art TTS models to synthesize speech that accurately reflects the intended emotions. Extensive evaluations show that UMETTS achieves significant improvements in emotion accuracy and speech naturalness, outperforming traditional E-TTS methods on both objective and subjective metrics. To facilitate reproducibility and further research, we have made our code publicly available at https://github.com/KTTRCDL/UMETTS.
Zhi-Qi Cheng, Jun-Yan He, Junyao Chen, Xiaomao Fan, Xiaojiang Peng, Alex Hauptmann 0001
ICASSP6
2025 Open-Vocabulary Visual Emotion Adaptation via Prompt Learning
abstract
Visual Emotion Recognition (VER) aims to identify emotions from visual content and has garnered significant attention in recent years due to its wide-ranging applications. Although deep learning-based methods have shown success in VER, they require extensive labeled data, which is costly. Unsupervised Domain Adaptation (UDA) methods can reduce reliance on annotated data by transferring models trained on labeled datasets to unlabeled data. However, these methods assume that the source and target domains share the same label space. In practice, this assumption is often violated due to the inherent ambiguity and subjectivity in emotion labeling. To address this limitation, we propose a novel prompt learning paradigm for open-vocabulary visual emotion UDA, termed Domain-specific Ensemble Prompting (DSEP). DSEP leverages psychological emotion models to unify emotion labels into a common space in an ensemble manner, enhancing the open-vocabulary capabilities of UDA. It then combines ensemble label prompts with domain-specific content prompts to achieve open-vocabulary UDA. To our knowledge, we are the first to explore open-vocabulary adaptation for VER. Extensive experiments demonstrate that DSEP consistently outperforms state-of-the-art methods across four public benchmarks.
Zhaopan Xu, Sicheng Zhao, Xiaojiang Peng, Hongxun Yao
ICASSP3
2025 DeformAvatar: Point-Based Human Avatar Re-targeting and Rendering
abstract
In this paper, we present the DeformAvatar, a novel architecture for human avatar re-targetting and rendering based on point clouds. Given the multiple views of a person, we first build a point-model-paired human representation containing a raw point cloud and an optimal parametric model. Then, we repurpose several advanced neural point-based rendering and Gaussian Splatting techniques for 3D avatar modeling. Finally, to enhance photorealistic re-targeting of body shapes and poses, we propose a dual point-adaptive (DPA) regularization based on traditional linear blend skinning. Extensive experiments demonstrate that our DeformAvatar framework can synthesize highly realistic novel views in new shape and pose parameters. We also find that 3DGS-based avatar modeling is superior to others in 3D avatar re-targeting.
Renyi Zhan, Zhi-Qi Cheng, Junyao Chen, Xiaojiang Peng
ICASSP4
2025 EA-Vit: Efficient Adaptation for Elastic Vision Transformer
abstract
Vision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming and energy-intensive. To address this issue, we propose an efficient ViT adaptation framework that enables a single adaptation process to generate multiple models of varying sizes for deployment on platforms with various resource constraints. Our approach comprises two stages. In the first stage, we enhance a pre-trained ViT with a nested elastic architecture that enables structural flexibility across MLP expansion ratio, number of attention heads, embedding dimension, and network depth. To preserve pre-trained knowledge and ensure stable adaptation, we adopt a curriculum-based training strategy that progressively increases elasticity. In the second stage, we design a lightweight router to select submodels according to computational budgets and downstream task demands. Initialized with Pareto-optimal configurations derived via a customized NSGA-II algorithm, the router is then jointly optimized with the backbone. Extensive experiments on multiple benchmarks demonstrate the effectiveness and versatility of EA-ViT. The code is available at https://github.com/zcxcf/EA-ViT.
Wangbo Zhao, Yuhao Zhou 0004, Weidong Tang, Shuo Wang 0001, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Kai Wang 0036
ICCV9
2025 Let Your Features Tell The Differences: Understanding Graph Convolution By Feature Splitting
abstract
Graph Neural Networks (GNNs) have demonstrated strong capabilities in processing structured data. While traditional GNNs typically treat each feature dimension equally important during graph convolution, we raise an important question: **Is the graph convolution operation equally beneficial for each feature?** If not, the convolution operation on certain feature dimensions can possibly lead to harmful effects, even worse than convolution-free models. Therefore, it is required to distinguish convolution-favored and convolution-disfavored features. Traditional feature selection methods mainly focus on identifying informative features or reducing redundancy, but they are not suitable for structured data as they overlook graph structures. In graph community, some studies have investigated the performance of GNN with respect to node features using feature homophily metrics, which assess feature consistency across graph topology. Unfortunately, these metrics do not effectively align with GNN performance and cannot be reliably used for feature selection in GNNs. To address these limitations, we introduce a novel metric, Topological Feature Informativeness (TFI), to distinguish GNN-favored and GNN-disfavored features, where its effectiveness is validated through both theoretical analysis and empirical observations. Based on TFI, we propose a simple yet effective Graph Feature Selection (GFS) method, which processes GNN-favored and GNN-disfavored features with GNNs and non-GNN models separately. Compared to original GNNs, GFS significantly improves the extraction of useful topological information from each feature with comparable computational costs. Extensive experiments show that after applying GFS to $\textbf{8}$ baseline and state-of-the-art (SOTA) GNN architectures across $\textbf{10}$ datasets, $\textbf{90\%}$ of the GFS-augmented cases show significant performance boosts. Furthermore, our proposed TFI metric outperforms other feature selection methods for GFS. These results verify the effectiveness of both GFS and TFI. Additionally, we demonstrate that GFS's improvements are robust to hyperparameter tuning, highlighting its potential as a universally valid method for enhancing various GNN architectures. To facilitate reproducibility and further research, we have made our code publicly available at https://github.com/KTTRCDL/graph-feature-selection.
Yilun Zheng, Sitao Luan, Xiaojiang Peng
ICLR4
2025 DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing
abstract
Fashion image editing is a crucial tool for designers to convey their creative ideas by visualizing design concepts interactively. However, current fashion image editing techniques often struggle to accurately identify editing regions and preserve the desired garment texture detail. To address these challenges, we present Detail-Preserved Diffusion Models (DPDEdit), a new multimodal fashion image editing architecture based on latent diffusion models. To precisely locate the editing region, we introduce Grounded-SAM to predict the editing region. To transfer the detail of the given garment texture into the target image, we propose a texture injection and refinement mechanism. This mechanism employs a decoupled cross-attention layer to integrate textual descriptions and texture images, and incorporates an auxiliary U-Net to preserve the high-frequency details of generated garment texture. Additionally, we extend the VITON-HD dataset using a multimodal large language model to generate paired samples with texture images and textual descriptions. Extensive experiments show that our DPDEdit outperforms state-of-the-art methods in terms of image fidelity and coherence with the given multimodal input.
Zhi-Qi Cheng, Huizi Xue, Xiaojiang Peng
ICME5
2025 AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
abstract
The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.
Zheng Lian 0004, Haoyu Chen 0001, Lan Chen 0005, Haiyang Sun 0004, Licai Sun, Yong Ren 0006, Zebang Cheng, Bin Liu 0041, Rui Liu 0008, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao 0001
ICML10
2025 MER 2025: When Affective Computing Meets Large Language Models
abstract
MER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality).
Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001
ACM Multimedia12
2025 Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths
abstract
Mamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's chain-like scanning mechanism, which we hypothesize not only induces cascading losses in token connectivity but also limits the diversity of spatial receptive fields. In this paper, we propose Asymmetric Multi-scale Vision Mamba (AMVim), a novel architecture designed to enhance pruning robustness. AMVim employs a dual-path structure, integrating a window-aware scanning mechanism into one path while retaining sequential scanning in the other. This asymmetry design promotes token connection diversity and enables multi-scale information flow, reinforcing spatial awareness. Empirical results demonstrate that AMVim achieves state-of-the-art pruning robustness. During token reduction, AMVim-T achieves a substantial 34\% improvement in training-free accuracy with identical model sizes and FLOPs. Meanwhile, AMVim-S exhibits only a 1.5\% accuracy drop, performing comparably to ViT. Notably, AMVim also delivers superior performance during pruning-free settings, further validating its architectural advantages.
Jindi Lv, Yuhao Zhou 0004, Mingjia Shi, Zhiyuan Liang, Xiaojiang Peng, Wangbo Zhao, Jiancheng Lv 0001, Kai Wang 0036
NeurIPS6
2025 TSFNet: A Temporal-Spectral Fusion Network for advanced speech emotion recognition in medical applications
abstract
Speech emotion recognition (SER) is a critical component in enhancing communication systems and human-machine interaction, with significant potential for applications in the medical field. Although existing SER methods that combine temporal and spectral features have achieved notable advancements, they still encounter a big challenge in capturing emotional nuances, which are vital in medical diagnostics and patient care. In this study, we introduce a straightforward yet highly efficient network called TSFNet, which is the Temporal-Spectral Fusion Network via a Large-scale Pre-trained Model. This network is specifically designed to effectively process intricate emotional nuances by seamlessly integrating temporal and spectral information present in speech signals. By leveraging the capabilities of a large-scale pre-trained model, which serves as a powerful plug-and-play component for extracting and learning the temporal characteristics of speech, TSFNet enables a more accurate capture of complex emotional details crucial for medical applications. Extensive experiments are conducted on publicly available datasets, to evaluate the performance of TSFNet. Extensive experiments conducted on six public datasets demonstrate that TSFNet significantly outperforms existing baselines, achieving unweighted accuracies of 95.57% for Savee, 92.67% for Crema-D, 85.71% for IEMOCAP, 100.00% for Tess, 95.86% for Emovo, and 80.43% for Meld. It means that TSFNet has the potential in advancing medical diagnostic tools and patient monitoring systems.
Peilin Huang, Xiaojiang Peng, Feng Sha, Xiaomao Fan, Ye Li 0002
Artif. Intell. Medicine3
2025 LEAF: Unveiling two sides of the same coin in semi-supervised facial expression recognition
Fan Zhang 0111, Zhi-Qi Cheng, Jian Zhao 0006, Xiaojiang Peng, Xuelong Li 0001
Comput. Vis. Image Underst.4
2025 Self-ensembling for 3D point cloud domain adaptation
Qing Li 0058, Xiaojiang Peng, Chuan Yan
Image Vis. Comput.2
2025 Large Language Model Enhanced Logic Tensor Network for Stance Detection
Genan Dai, Jiayu Liao, Sicheng Zhao, Xianghua Fu, Xiaojiang Peng, Hu Huang 0009, Bowen Zhang 0005
Neural Networks5
2025 miMamba: EEG-Based Emotion Recognition With Multi-Scale Inverted Mamba Models
abstract
EEG-based emotion recognition holds significant potential in the field of brain-computer interfaces. A key challenge is extracting discriminative spatiotemporal features from electroencephalogram (EEG) signals. Existing studies often rely on domain-specific time-frequency features and analyze temporal dependencies and spatial characteristics separately, neglecting the local-global relationships and the interaction in spatiotemporal dynamics. To address this, we propose a novel network called Multi-scale Inverted Mamba (miMamba), which consists of Multi-Scale Temporal Blocks (MSTB) and Temporal-Spatial Fusion Blocks (TSFB). Specifically, MSTBs are designed to capture both local details and global temporal dependencies across different scale subsequences. The TSFBs, implemented with an inverted Mamba structure, focus on the interaction between dynamic temporal dependencies and spatial characteristics. The primary advantage of miMamba lies in its ability to leverage transformed multi-scale EEG sequences, exploiting the interaction between temporal and spatial features without the need for domain-specific time-frequency feature extraction. Experiments show that using only four EEG channels, miMamba achieves remarkable average recognition accuracies for Valence and Arousal classification: 94.86% on the DEAP dataset, 94.94% on the DREAMER dataset, and 91.36% on the SEED dataset. These results underscore the model's superior performance in multidimensional emotion recognition tasks and its potential for practical applications in resource-constrained affective computing scenarios.
Dawei Huang, Xiaojiang Peng
IEEE Trans. Affect. Comput.3
2025 Tucker Decomposition-Enhanced Dynamic Graph Convolutional Networks for Crowd Flows Prediction
abstract
Crowd flows prediction is an important problem for traffic management and public safety. Graph Convolutional Network (GCN), known for its ability to effectively capture and utilize topological information, has demonstrated significant advancements in addressing this problem. However, GCN-based models were often based on predefined crowd-flow graphs via historical movement behaviors of human beings and traffic vehicles, which ignored the abnormal changes in crowd flows. In this study, we propose a multi-scale fusion GCN-based framework with Tucker decomposition named mTDNet to enhance dynamic GCN for crowd flows prediction. Following the paradigm of extant methods, we also employ the predefined crowd-flow graphs as a part of mTDNet to effectively capture the historical movement behaviors of crowd flows. To capture the abnormal changes, we propose a Tucker decomposition-based network with the product of the adjacency matrix of historical movement pattern graphs and an Adaptive Learning Tensor ( ALT ) by reconstructing the crowd flows. Particularly, we utilize the Tucker decomposition scheme to decompose ALT , which enhances the dynamic learning of graph structures, allowing for effective capturing of the dynamic changes in crowd flow, including abnormal changes. Furthermore, a multi-scale 3DGCN is utilized to mine and fuse the multi-scale spatio-temporal information from crowd flows, to further boost the mTDNet prediction performance. Experiments conducted on two real-world datasets showed that the proposed mTDNet surpasses other crowd flow prediction methods.
Genan Dai, Weiyang Kong, Bowen Zhang 0005, Xiaojiang Peng, Xiaomao Fan, Hu Huang 0009
ACM Trans. Intell. Syst. Technol.5
2025 cVAN: A Novel Sleep Staging Method via Cross-View Alignment Network
abstract
Sleep staging is imperative for evaluating sleep quality and diagnosing sleep disorders. Extant sleep staging methods with fusing multiple data-views of physiological signals have achieved promising results. However, they remain neglectful of the relationship among different data-views at different feature scales with view position-alignment. To address this, we propose a novel cross-view alignment network, termed cVAN, utilising scale-aware attention for sleep stages classification. Specifically, cVAN principally incorporates two sub-networks of a residual-like network which learn spectral information from time-frequency images and a transformer-like network which learns corresponding temporal information. The prime advantage of cVAN is to adaptively align the learned feature scales among the different data-views of physiological signals with a scale-aware attention by reorganizing feature maps. Extensive experiments on three public sleep datasets demonstrate that cVAN can achieve a new state-of-the-art result, which is superior to existing counterparts.
Zhanjiang Yang, Meiyu Qiu, Xiaomao Fan, Genan Dai, Wenjun Ma, Xiaojiang Peng, Xianghua Fu, Ye Li 0002
IEEE J. Biomed. Health Informatics6
2025 Facial Action Units as a Joint Dataset Training Bridge for Facial Expression Recognition
abstract
Label biases in facial expression recognition (FER) datasets, caused by annotators' subjectivity, pose challenges in improving the performance of target datasets when auxiliary labeled data are used. Moreover, training with multiple datasets can lead to visible degradations in the target dataset. To address these issues, we propose a novel framework called the AU-aware Vision Transformer (AU-ViT), which leverages unified action unit (AU) information and discards expression annotations of auxiliary data. AU-ViT integrates an elaborately designed AU branch in the middle part of a master ViT to enhance representation learning during training. Through qualitative and quantitative analyses, we demonstrate that AU-ViT effectively captures expression regions and is robust to real-world occlusions. Additionally, we observe that AU-ViT also yields performance improvements on the target dataset, even without auxiliary data, by utilizing pseudo AU labels. Our AU-ViT achieves performances superior to, or comparable to, that of the state-of-the-art methods on FERPlus, RAFDB, AffectNet, LSD and the other three occlusion test datasets.
Shuyi Mao, Xinpeng Li 0004, Fan Zhang 0111, Xiaojiang Peng, Yang Yang 0002
IEEE Trans. Multim.4
2024 A Challenge Dataset and Effective Models for Conversational Stance Detection
abstract
Previous stance detection studies typically concentrate on evaluating stances within individual instances, thereby exhibiting limitations in effectively modeling multi-party discussions concerning the same specific topic, as naturally transpire in authentic social media interactions. This constraint arises primarily due to the scarcity of datasets that authentically replicate real social media contexts, hindering the research progress of conversational stance detection. In this paper, we introduce a new multi-turn conversation stance detection dataset (called MT-CSD), which encompasses multiple targets for conversational stance detection. To derive stances from this challenging dataset, we propose a global-local attention network (GLAN) to address both long and short-range dependencies inherent in conversational data. Notably, even state-of-the-art stance detection methods, exemplified by GLAN, exhibit an accuracy of only 50.47%, highlighting the persistent challenges in conversational stance detection. Furthermore, our MT-CSD dataset serves as a valuable resource to catalyze advancements in cross-domain stance detection, where a classifier is adapted from a different yet related target. We believe that MT-CSD will contribute to advancing real-world applications of stance detection research. Our source code, data, and models are available at https://github.com/nfq729/MT-CSD.
Fuqiang Niu, Min Yang 0007, Ang Li 0047, Baoquan Zhang, Xiaojiang Peng, Bowen Zhang 0005
LREC/COLING5
2024 Dataset Growth
Ziheng Qin, Zhaopan Xu, Zangwei Zheng, Zebang Cheng, Hao Tang 0005, Baigui Sun, Xiaojiang Peng, Radu Timofte, Hongxun Yao, Kai Wang 0036, Yang You 0001
ECCV (9)9
2024 DSMix: Distortion-Induced Sensitivity Map Based Pre-training for No-Reference Image Quality Assessment
Jinsong Shi 0002, Pan Gao 0001, Xiaojiang Peng, Jie Qin 0004
ECCV (70)3
2024 NuwaDynamics: Discovering and Updating in Causal Spatio-Temporal Modeling
abstract
Spatio-temporal (ST) prediction plays a pivotal role in earth sciences, such as meteorological prediction, urban computing. Adequate high-quality data, coupled with deep models capable of inference, are both indispensable and prerequisite for achieving meaningful results. However, the sparsity of data and the high costs associated with deploying sensors lead to significant data imbalances. Models that are overly tailored and lack causal relationships further compromise the generalizabilities of inference methods. Towards this end, we first establish a causal concept for ST predictions, named NuwaDynamics, which targets to identify causal regions in data and endow model with causal reasoning ability in a two-stage process. Concretely, we initially leverage upstream self-supervision to discern causal important patches, imbuing the model with generalized information and conducting informed interventions on complementary trivial patches to extrapolate potential test distributions. This phase is referred to as the discovery step. Advancing beyond discovery step, we transfer the data to downstream tasks for targeted ST objectives, aiding the model in recognizing a broader potential distribution and fostering its causal perceptual capabilities (refer as Update step). Our concept aligns seamlessly with the contemporary backdoor adjustment mechanism in causality theory. Extensive experiments on six real-world ST benchmarks showcase that models can gain outcomes upon the integration of the NuwaDynamics concept. NuwaDynamics also can significantly benefit a wide range of changeable ST tasks like extreme weather and long temporal step super-resolution predictions.
Kun Wang 0056, Hao Wu 0083, Yifan Duan, Guibin Zhang, Kai Wang 0036, Xiaojiang Peng, Yu Zheng 0004, Yuxuan Liang 0002, Yang Wang 0015
ICLR6
2024 The Snowflake Hypothesis: Training and Powering GNN with One Node One Receptive Field
abstract
Despite Graph Neural Networks (GNNs) demonstrating considerable promise in graph representation learning tasks, GNNs predominantly face significant issues with overfitting and over-smoothing as they go deeper as models of computer vision (CV) realm.The success of artificial intelligence in computer vision and natural language processing largely stems from its ability to train deep models effectively.We have thus conducted a systematic study on deep GNN models.Our findings indicate that the current success of deep GNNs primarily stems from (I) the adoption of innovations from CNNs, such as residual/skip connections, or (II) the tailor-made aggregation algorithms like DropEdge.However, these algorithms often lack intrinsic interpretability and indiscriminately treat all nodes within a given layer in a similar manner, thereby failing to capture the nuanced differences among various nodes.In this paper, we introduce the Snowflake Hypothesis -a novel paradigm underpinning the concept of "one node, one receptive field".The hypothesis draws inspiration from the unique and individualistic patterns of * Contribute equally to this research.
Kun Wang 0056, Guohao Li 0001, Shilong Wang 0002, Guibin Zhang, Kai Wang 0036, Yang You 0001, Junfeng Fang, Xiaojiang Peng, Yuxuan Liang 0002, Yang Wang 0015
KDD8
2024 Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context Transformer
abstract
Emotion recognition aims to discern the emotional state of subjects within an image, relying on subject-centric and contextual visual cues. Current approaches typically follow a two-stage pipeline: first localize subjects by off-the-shelf detectors, then perform emotion classification through the late fusion of subject and context features. However, the complicated paradigm suffers from disjoint training stages and limited fine-grained interaction between subject-context elements. To address the challenge, we present a single-stage emotion recognition approach, employing a Decoupled Subject-Context Transformer (DSCT), for simultaneous subject localization and emotion classification. Rather than compartmentalizing training stages, we jointly leverage box and emotion signals as supervision to enrich subject-centric feature learning. Furthermore, we introduce DSCT to facilitate interactions between fine-grained subject-context cues in a ''decouple-then-fuse'' manner. The decoupled query tokens-subject queries and context queries-gradually intertwine across layers within DSCT, during which spatial and semantic relations are exploited and aggregated. We evaluate our single-stage framework on two widely used context-aware emotion recognition datasets, CAER-S and EMOTIC. Our approach surpasses two-stage alternatives with fewer parameter numbers, achieving a 3.39% accuracy improvement and a 6.46% average precision gain on CAER-S and EMOTIC datasets, respectively. Code and models are available at: https://github.com/Sampson-Lee/DSCT.
Xinpeng Li 0004, Teng Wang 0007, Jian Zhao 0006, Shuyi Mao, Jinbao Wang 0001, Feng Zheng 0001, Xiaojiang Peng, Xuelong Li 0001
ACM Multimedia7
2024 Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
abstract
Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the proliferation of diverse multimodal social media content including text, and images multimodal stance detection (MSD) has become a crucial research area. However, existing MSD studies have focused on modeling stance within individual text-image pairs, overlooking the multi-party conversational contexts that naturally occur on social media. This limitation stems from a lack of datasets that authentically capture such conversational scenarios, hindering progress in conversational MSD. To address this, we introduce a new multimodal multi-turn conversational stance detection dataset (called MmMtCSD). To derive stances from this challenging dataset, we propose a novel multimodal large language model stance detection framework (MLLM-SD), that learns joint stance representations from textual and visual modalities. Experiments on MmMtCSD show state-of-the-art performance of our proposed MLLM-SD approach for multimodal stance detection. We believe that MmMtCSD will contribute to advancing real-world applications of stance detection research.
Fuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng, Genan Dai, Hu Huang 0009, Bowen Zhang 0005
ACM Multimedia4
2024 Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
abstract
Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset.
Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 0036, Zheng Lian 0004, Xiaojiang Peng, Alex Hauptmann 0001
NeurIPS7
2024 Semi-supervised Knowledge Transfer Across Multi-omic Single-cell Data
abstract
Knowledge transfer between multi-omic single-cell data aims to effectively transfer cell types from scRNA-seq data to unannotated scATAC-seq data. Several approaches aim to reduce the heterogeneity of multi-omic data while maintaining the discriminability of cell types with extensive annotated data. However, in reality, the cost of collecting both a large amount of labeled scRNA-seq data and scATAC-seq data is expensive. Therefore, this paper explores a practical yet underexplored problem of knowledge transfer across multi-omic single-cell data under cell type scarcity. To address this problem, we propose a semi-supervised knowledge transfer framework named Dual label scArcity elimiNation with Cross-omic multi-samplE Mixup (DANCE). To overcome the label scarcity in scRNA-seq data, we generate pseudo-labels based on optimal transport and merge them into the labeled scRNA-seq data. Moreover, we adopt a divide-and-conquer strategy which divides the scATAC-seq data into source-like and target-specific data. For source-like samples, we employ consistency regularization with random perturbations while for target-specific samples, we select a few candidate labels and progressively eliminate incorrect cell types from the label set for additional supervision. Next, we generate virtual scRNA-seq samples with multi-sample Mixup based on the class-wise similarity to reduce cell heterogeneity. Extensive experiments on many benchmark datasets suggest the superiority of our DANCE over a series of state-of-the-art methods.
Fan Zhang 0111, Tianyu Liu 0005, Xiaojiang Peng, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001, Hongyu Zhao 0003
NeurIPS4
2024 A Comprehensive Exploration on Detecting Fake Images Generated by Stable Diffusion
Zhijian He, Xiaojiang Peng
PRCV (1)4
2024 Invisible gas detection: An RGB-thermal cross attention network and a new benchmark
Shuaibao Chen, Xiaojiang Peng
Comput. Vis. Image Underst.7
2024 Graph Attentive Dual Ensemble learning for Unsupervised Domain Adaptation on point clouds
Qing Li 0058, Chuan Yan, Xiaojiang Peng
Pattern Recognit.4
2023 Semi-Supervised Multimodal Emotion Recognition with Expression MAE
abstract
The Multimodal Emotion Recognition (MER 2023) challenge aims to recognize emotion with audio, language, and visual signals, facilitating innovative technologies of affective computing. This paper presents our submission approach on the Semi-Supervised Learning Sub-Challenge (MER-SEMI). First, with large-scale unlabeled emotional videos, we train both image-based and video-based Masked Autoencoders to extract visual features, which termed as expression MAE (expMAE) for simplicity. The expMAE features are found to be largely complementary with other official baseline features. Second, since there is only a few labeled data, we use a classifier to generate pseudo labels for unlabeled videos which have high confidence for a certain category. In addition, we also explore several advanced large models for cross-feature extraction like CLIP, and apply factorized bilinear pooling (FBP) for multimodal feature fusion. Our methods finally achieved 88.55% in F1 score on MER-SEMI, ranking second place among all participating teams.
Zebang Cheng, Zhaoru Chen, Xiang Li 0130, Shuyi Mao, Fan Zhang 0111, Daijun Ding, Bowen Zhang 0005, Xiaojiang Peng
ACM Multimedia9
2022 An Efficient Training Approach for Very Large Scale Face Recognition
abstract
Face recognition has achieved significant progress in deep learning era due to the ultra-large-scale and well- labeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is in- dispensable. The heavy computational and memory costs mainly result from the million-level dimensionality of the fully connected (FC) layer. To this end, we propose a novel training approach, termed Faster Face Classification (F2C), to alleviate time and cost without sacrificing the performance. This method adopts Dynamic Class Pool (DCP) for storing and updating the identities' features dy-namically, which could be regarded as a substitute for the FC layer. DCP is efficiently time-saving and cost-saving, as its smaller size with the independence from the whole face identities together. We further validate the proposed F2C method across several face benchmarks and private datasets, and display comparable results, meanwhile the speed is faster than state-of-the-art FC-based methods in terms of recognition accuracy and hardware costs. More-over, our method is further improved by a well-designed dual data loader including indentity-based and instance- based loaders, which makes it more efficient for updating DCP parameters.
Kai Wang 0036, Shuo Wang 0001, Xiaojiang Peng, Baigui Sun, Hao Li 0030, Yang You 0001
CVPR7
2022 Video Frame Interpolation Based on Deformable Kernel Region
abstract
Video frame interpolation task has recently become more and more prevalent in the computer vision field. At present, a number of researches based on deep learning have achieved great success. Most of them are either based on optical flow information, or interpolation kernel, or a combination of these two methods. However, these methods have ignored that there are grid restrictions on the position of kernel region during synthesizing each target pixel. These limitations result in that they cannot well adapt to the irregularity of object shape and uncertainty of motion, which may lead to irrelevant reference pixels used for interpolation. In order to solve this problem, we revisit the deformable convolution for video interpolation, which can break the fixed grid restrictions on the kernel region, making the distribution of reference points more suitable for the shape of the object, and thus warp a more accurate interpolation frame. Experiments are conducted on four datasets to demonstrate the superior performance of the proposed model in comparison to the state-of-the-art alternatives.
Haoyue Tian, Pan Gao 0001, Xiaojiang Peng
IJCAI3
2022 Rail Detection: An Efficient Row-based Network and a New Benchmark
abstract
Rail detection, essential for railroad anomaly detection, aims to identify the railroad region in video frames. Although various studies on rail detection exist, neither an open benchmark nor a high-speed network is available in the community, making al- gorithm comparison and development difficult. Inspired by the growth of lane detection, we propose a rail database and a row- based rail detection method. In detail, we make several contribu- tions: (i) We present a real-world railway dataset, Rail-DB, with 7432 pairs of images and annotations. The images are collected from different situations in lighting, road structures, and views. The rails are labeled with polylines, and the images are catego- rized into nine scenes. The Rail-DB is expected to facilitate the improvement of rail detection algorithms. (ii) We present an ef- ficient row-based rail detection method, Rail-Net, containing a lightweight convolutional backbone and an anchor classifier. Specif- ically, we formulate the process of rail detection as a row-based selecting problem. This strategy reduces the computational cost compared to alternative segmentation methods. (iii) We evaluate the Rail-Net on Rail-DB with extensive experiments, including cross-scene settings and network backbones ranging from ResNet to Vision Transformers. Our method achieves promising perfor- mance in terms of both speed and accuracy. Notably, a lightweight version could achieve 92.77% accuracy and 312 frames per second. The Rail-Net outperforms the traditional method by 50.65% and the segmentation one by 5.86%. The database and code are available at: https://github.com/Sampson-Lee/Rail-Detection.
Xinpeng Li 0004, Xiaojiang Peng
ACM Multimedia2
2022 Joint 3D facial shape reconstruction and texture completion from a single image
abstract
Recent years have witnessed significant progress in image-based 3D face reconstruction using deep convolutional neural networks. However, current reconstruction methods often perform improperly in self-occluded regions and can lead to inaccurate correspondences between a 2D input image and a 3D face template, hindering use in real applications. To address these problems, we propose a deep shape reconstruction and texture completion network, SRTC-Net, which jointly reconstructs 3D facial geometry and completes texture with correspondences from a single input face image. In SRTC-Net, we leverage the geometric cues from completed 3D texture to reconstruct detailed structures of 3D shapes. The SRTC-Net pipeline has three stages. The first introduces a correspondence network to identify pixel-wise correspondence between the input 2D image and a 3D template model, and transfers the input 2D image to a U - V texture map. Then we complete the invisible and occluded areas in the U - V texture map using an inpainting network. To get the 3D facial geometries, we predict coarse shape ( U - V position maps) from the segmented face from the correspondence network using a shape network, and then refine the 3D coarse shape by regressing the U - V displacement map from the completed U - V texture map in a pixel-to-pixel way. We examine our methods on 3D reconstruction tasks as well as face frontalization and pose invariant face recognition tasks, using both in-the-lab datasets (MICC, MultiPIE) and in-the-wild datasets (CFP). The qualitative and quantitative results demonstrate the effectiveness of our methods on inferring 3D facial geometry and complete texture; they outperform or are comparable to the state-of-the-art.
Xiaoxing Zeng, Zhelun Wu, Xiaojiang Peng, Yu Qiao 0001
Comput. Vis. Media3
2022 Unsupervised person re-identification with multi-label learning guided self-paced clustering
Qing Li 0058, Xiaojiang Peng, Yu Qiao 0001
Pattern Recognit.2
2021 Affordance Transfer Learning for Human-Object Interaction Detection
abstract
Reasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspired by this, we introduce an affordance transfer learning approach to jointly detect HOIs with novel object and recognize affordances. Specifically, HOI representations can be decoupled into a combination of affordance and object representations, making it possible to compose novel interactions by combining affordance representations and novel object representations from additional images, i.e. transferring the affordance to novel objects. With the proposed affordance transfer learning, the model is also capable of inferring the affordances of novel objects from known affordance representations. The proposed method can thus be used to 1) improve the performance of HOI detection, especially for the HOIs with unseen objects; and 2) infer the affordances of novel objects. Experimental results on two datasets, HICO-DET and HOI-COCO (from V-COCO), demonstrate significant improvements over recent state-of-the-art methods for HOI detection and object affordance detection. Code is available at https://github.com/zhihou7/HOI-CL.
Zhi Hou, Baosheng Yu, Yu Qiao 0001, Xiaojiang Peng, Dacheng Tao
CVPR4
2021 Detecting Human-Object Interaction via Fabricated Compositional Learning
abstract
Human-Object Interaction (HOI) detection, inferring the relationships between human and objects from images/videos, is a fundamental task for high-level scene understanding. However, HOI detection usually suffers from the open long-tailed nature of interactions with objects, while human has extremely powerful compositional perception ability to cognize rare or unseen HOI samples. Inspired by this, we devise a novel HOI compositional learning framework, termed as Fabricated Compositional Learning (FCL), to address the problem of open long-tailed HOI detection. Specifically, we introduce an object fabricator to generate effective object representations, and then combine verbs and fabricated objects to compose new HOI samples. With the proposed object fabricator, we are able to generate large-scale HOI samples for rare and unseen categories to alleviate the open long-tailed issues in HOI detection. Extensive experiments on the most popular HOI detection dataset, HICO-DET, demonstrate the effectiveness of the proposed method for imbalanced HOI detection and significantly improve the state-of-the-art performance on rare and unseen HOI categories. Code is available at https://github.com/zhihou7/HOI-CL.
Zhi Hou, Baosheng Yu, Yu Qiao 0001, Xiaojiang Peng, Dacheng Tao
CVPR4
2021 Sequential Interactive Biased Network for Context-Aware Emotion Recognition
abstract
Emotion context information is crucial yet complicated for emotion recognition. How to process it is a challenging problem. Existing works mainly extract context representations of the face, body and scene independently. These strategies may be limited in the understanding of emotional context relation. To address this problem, we propose Sequential Interactive Biased Network (SIB-Net), which is motivated by the studies that the context contains sequential, interactive and biased relation. Specifically, SIB-Net captures and utilizes the context relation by three modules: i) a Sequential Context Module captures consecutive relation with a GRU-like architecture, ii) an Interactive Context Module acquires cooperative context with global correlated linear fusion, and iii) a Biased Context Module benefits from the biased relation with distribution labels and the L1 loss. Extensive experiments on EMOTIC and CAER datasets show that our SIB-Net improves baseline significantly and achieves comparable results to the state-of-the-art methods.
Xinpeng Li 0004, Xiaojiang Peng, Changxing Ding
IJCB2
2021 TTPP: Temporal Transformer with Progressive Prediction for efficient action anticipation
Wen Wang 0012, Xiaojiang Peng, Yanzhou Su, Yu Qiao 0001, Jian Cheng 0003
Neurocomputing2
2020 Suppressing Uncertainties for Large-Scale Facial Expression Recognition
abstract
Annotating a qualitative large-scale facial expression dataset is extremely difficult due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. These uncertainties suspend the progress of large-scale Facial Expression Recognition (FER) in data-driven deep learning era. To address this problelm, this paper proposes to suppress the uncertainties by a simple yet efficient Self-Cure Network (SCN). Specifically, SCN suppresses the uncertainty from two different aspects: 1) a self-attention mechanism over FER dataset to weight each sample in training with a ranking regularization, and 2) a careful relabeling mechanism to modify the labels of these samples in the lowest-ranked group. Experiments on synthetic FER datasets and our collected WebEmotion dataset validate the effectiveness of our method. Results on public benchmarks demonstrate that our SCN outperforms current state-of-the-art methods with \textbf{88.14}\% on RAF-DB, \textbf{60.23}\% on AffectNet, and \textbf{89.35}\% on FERPlus.
Kai Wang 0036, Xiaojiang Peng, Jianfei Yang 0001, Shijian Lu, Yu Qiao 0001
CVPR2
2020 Visual Compositional Learning for Human-Object Interaction Detection
Zhi Hou, Xiaojiang Peng, Yu Qiao 0001, Dacheng Tao
ECCV (15)2
2020 Suppressing Mislabeled Data via Grouping and Self-attention
Xiaojiang Peng, Kai Wang 0036, Zhaoyang Zeng, Qing Li 0058, Jianfei Yang 0001, Yu Qiao 0001
ECCV (16)1
2020 Attention-Driven Dynamic Graph Convolutional Network for Multi-label Image Recognition
Jin Ye 0002, Junjun He, Xiaojiang Peng, Yu Qiao 0001
ECCV (21)3
2020 Learning Discriminative Representation For Facial Expression Recognition From Uncertainties
abstract
Recent progresses on Facial Expression Recognition (FER) heavily rely on deep learning models trained with large scale datasets. However, large-scale facial expression datasets always suffer from annotation uncertainties caused by ambiguous expressions, low-quality facial images, and the subjectiveness of annotators, which limits FER performance. To address this challenge, this paper introduces novel Rayleigh and weighted-softmax loss from two aspects. First, we propose Rayleigh loss to extract discriminative representation, which aims at minimizing within-class distances and maximizing inter-class distances simultaneously. Moreover, Rayleigh loss has a Euclidean form which make it easily be optimized with SGD and be combined with other forms. Second, we introduce a weight to measure the uncertainty of a given sample, by considering its distance to class center. Extensive experiments on RAF-DB, FERPlus and AffectNet show the effectiveness of our method with SOTA performance.
Xingyu Fan, Zhongying Deng, Kai Wang 0036, Xiaojiang Peng, Yu Qiao 0001
ICIP4
2020 Product image recognition with guidance learning and noisy supervision
Qing Li 0058, Xiaojiang Peng, Liangliang Cao, Wenbin Du, Yu Qiao 0001, Qiang Peng
Comput. Vis. Image Underst.2
2020 Cascade multi-head attention networks for action recognition
Xiaojiang Peng, Yu Qiao 0001
Comput. Vis. Image Underst.2
2020 Finding hard faces with better proposals and classifier
Xiaoxing Zeng, Xiaojiang Peng, Yali Wang 0001, Yu Qiao 0001
Mach. Vis. Appl.2
2020 Learning label correlations for multi-label image recognition with graph networks
Qing Li 0058, Xiaojiang Peng, Yu Qiao 0001, Qiang Peng
Pattern Recognit. Lett.2
2020 Region Attention Networks for Pose and Occlusion Robust Facial Expression Recognition
abstract
Occlusion and pose variations, which can change facial appearance significantly, are two major obstacles for automatic Facial Expression Recognition (FER). Though automatic FER has made substantial progresses in the past few decades, occlusion-robust and pose-invariant issues of FER have received relatively less attention, especially in real-world scenarios. This paper addresses the real-world pose and occlusion robust FER problem in the following aspects. First, to stimulate the research of FER under real-world occlusions and variant poses, we annotate several in-the-wild FER datasets with pose and occlusion attributes for the community. Second, we propose a novel Region Attention Network (RAN), to adaptively capture the importance of facial regions for occlusion and pose variant FER. The RAN aggregates and embeds varied number of region features produced by a backbone convolutional neural network into a compact fixed-length representation. Last, inspired by the fact that facial expressions are mainly defined by facial action units, we propose a region biased loss to encourage high attention weights for the most important regions. We validate our RAN and region biased loss on both our built test datasets and four popular datasets: FERPlus, AffectNet, RAF-DB, and SFEW. Extensive experiments show that our RAN and region biased loss largely improve the performance of FER with occlusion and variant pose. Our method also achieves state-of-the-art results on FERPlus, AffectNet, RAF-DB, and SFEW. Code and the collected test data will be publicly available.
Kai Wang 0036, Xiaojiang Peng, Jianfei Yang 0001, Debin Meng, Yu Qiao 0001
IEEE Trans. Image Process.2
2019 Residual Compensation Networks for Heterogeneous Face Recognition
abstract
Heterogeneous Face Recognition (HFR) is a challenging task due to large modality discrepancy as well as insufficient training images in certain modalities. In this paper, we propose a new two-branch network architecture, termed as Residual Compensation Networks (RCN), to learn separated features for different modalities in HFR. The RCN incorporates a residual compensation (RC) module and a modality discrepancy loss (MD loss) into traditional convolutional neural networks. The RC module reduces modal discrepancy by adding compensation to one of the modalities so that its representation can be close to the other modality. The MD loss alleviates modal discrepancy by minimizing the cosine distance between different modalities. In addition, we explore different architectures and positions for the RC module, and evaluate different transfer learning strategies for HFR. Extensive experiments on IIIT-D Viewed Sketch, Forensic Sketch, CASIA NIR-VIS 2.0 and CUHK NIR-VIS show that our RCN outperforms other state-of-the-art methods significantly.
Zhongying Deng, Xiaojiang Peng, Yu Qiao 0001
AAAI2
2019 DF2Net: A Dense-Fine-Finer Network for Detailed 3D Face Reconstruction
abstract
Reconstructing the detailed geometric structure from a single face image is a challenging problem due to its ill-posed nature and the fine 3D structures to be recovered. This paper proposes a deep Dense-Fine-Finer Network (DF2Net) to address this challenging problem. DF2Net decomposes the reconstruction process into three stages, each of which is processed by an elaborately-designed network, namely D-Net, F-Net, and Fr-Net. D-Net exploits a U-net architecture to map the input image to a dense depth image. F-Net refines the output of D-Net by integrating features from depth and RGB domains, whose output is further enhanced by Fr-Net with a novel multi-resolution hypercolumn architecture. In addition, we introduce three types of data to train these networks, including 3D model synthetic data, 2D image reconstructed data, and fine facial images. We elaborately exploit different datasets (or combination) together with well-designed losses to train different networks. Qualitative evaluation indicates that our DF2Net can effectively reconstruct subtle facial details such as small crow's feet and wrinkles. Our DF2Net achieves performance superior or comparable to state-of-the-art algorithms in qualitative and quantitative analyses on real-world images and the BU-3DFE dataset. Code and the collected 70K image-depth data will be publicly available.
Xiaoxing Zeng, Xiaojiang Peng, Yu Qiao 0001
ICCV2
2019 Frame Attention Networks for Facial Expression Recognition in Videos
abstract
The video-based facial expression recognition aims to classify a given video into several basic emotions. How to integrate facial features of individual frames is crucial for this task. In this paper, we propose the Frame Attention Networks (FAN)1, to automatically highlight some discriminative frames in an end-to-end framework. The network takes a video with a variable number of face images as its input and produces a fixed-dimension representation. The whole network is composed of two modules. The feature embedding module is a deep Convolutional Neural Network (CNN) which embeds face images into feature vectors. The frame attention module learns multiple attention weights which are used to adaptively aggregate the feature vectors to form a single discriminative video representation. We conduct extensive experiments on CK+ and AFEW8.0 datasets. Our proposed FAN shows superior performance compared to other CNN based methods and achieves state-of-the-art performance on CK+.
Debin Meng, Xiaojiang Peng, Kai Wang 0036, Yu Qiao 0001
ICIP2
2019 Visual-Textual Sentiment Analysis in Product Reviews
abstract
Sentiment analysis has attracted increasing attention recently due to its potential wide applications in opinion analysis, recommendation system, etc. Visual-textual sentiment analysis aims to improve the performance of sentiment analysis by leveraging both visual and textual signals. In this paper, we address the visual-textual sentiment analysis in product reviews. Our main contributions are two-fold. First, instead of crawling data from Flickr or Twitter with positive and negative labels in existing works, we introduce a new dataset for visual-textual sentiment analysis, termed as Product Reviews150K (PR-150K), which is collected from the product reviews of online shopping websites. Second, we propose a deep Tucker fusion method for visual-textual sentiment analysis, which efficiently combines visual and textual deep representations based on the Tucker decomposition and a bilinear pooling operation. Extensive experiments on our PR-150K, MVSO, and VSO datasets show that our method outperforms several state-of-the-art methods.
Jin Ye 0002, Xiaojiang Peng, Yu Qiao 0001, Rongrong Ji
ICIP2
2019 Exploring Regularizations with Face, Body and Image Cues for Group Cohesion Prediction
abstract
This paper presents our approach for the group cohesion prediction sub-challenge in the EmotiW 2019. The task is to predict group cohesiveness in images. We mainly explore several regularizations with three types of visual cues, namely face, body ,and global image. Our main contribution is two-fold. First, we jointly train the group cohesion prediction task and group emotion recognition task using multi-task learning strategy with all visual cues. Second, we elaborately design two regularizations, namely a rank loss and a hourglass loss, where the former aims to give a margin between the distance of distant categories and near categories and the later to avoid centralization predictions with only MSE loss. With careful evaluations, we finally achieve the second place in this sub-challenge with MSE of 0.43821 on the testing set. https://github.com/DaleAG/Group_Cohesion_Prediction
Da Guo, Kai Wang 0036, Jianfei Yang 0001, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001
ICMI5
2019 Bootstrap Model Ensemble and Rank Loss for Engagement Intensity Regression
abstract
This paper presents our approach for the engagement intensity regression task of EmotiW 2019. The task is to predict the engagement intensity value of a student when he or she is watching an online MOOCs video in various conditions. Based on our winner solution last year, we mainly explore head features and body features with a bootstrap strategy and two novel loss functions in this paper. We maintain the framework of multi-instance learning with long short-term memory (LSTM) network, and make three contributions. First, besides of the gaze and head pose features, we explore facial landmark features in our framework. Second, inspired by the fact that engagement intensity can be ranked in values, we design a rank loss as a regularization which enforces a distance margin between the features of distant category pairs and adjacent category pairs. Third, we use the classical bootstrap aggregation method to perform model ensemble which randomly samples a certain training data by several times and then averages the model predictions. We evaluate the performance of our method and discuss the influence of each part on the validation dataset. Our methods finally win 3rd place with MSE of 0.0626 on the testing set. https://github.com/kaiwang960112/EmotiW_2019_ engagement_regression
Kai Wang 0036, Jianfei Yang 0001, Da Guo, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001
ICMI5
2019 Exploring Emotion Features and Fusion Strategies for Audio-Video Emotion Recognition
abstract
The audio-video based emotion recognition aims to classify a given video into basic emotions. In this paper, we describe our approaches in EmotiW 2019, which mainly explores emotion features and feature fusion strategies for audio and visual modality. For emotion features, we explore audio feature with both speech-spectrogram and Log Mel-spectrogram and evaluate several facial features with different CNN models and different emotion pretrained strategies. For fusion strategies, we explore intra-modal and cross-modal fusion methods, such as designing attention mechanisms to highlights important emotion feature, exploring feature concatenation and factorized bilinear pooling (FBP) for cross-modal feature fusion. With careful evaluation, we obtain 65.5% on the AFEW validation set and 62.48% on the test set and rank third in the challenge.
Hengshun Zhou, Debin Meng, Xiaojiang Peng, Jun Du 0002, Kai Wang 0036, Yu Qiao 0001
ICMI4
2019 AnoPCN: Video Anomaly Detection via Deep Predictive Coding Network
abstract
Video anomaly detection is a challenging problem due to the ambiguity and complexity of how anomalies are defined. Recent approaches for this task mainly utilize deep reconstruction methods and deep prediction ones, but their performances suffer when they cannot guarantee either higher reconstruction errors for abnormal events or lower prediction errors for normal events. Inspired by the predictive coding mechanism explaining how brains detect events violating regularities, we address the Anomaly detection problem with a novel deep Predictive Coding Network, termed as AnoPCN, which consists of a Predictive Coding Module (PCM) and an Error Refinement Module (ERM). Specifically, PCM is designed as a convolutional recurrent neural network with feedback connections carrying frame predictions and feedforward connections carrying prediction errors. By using motion information explicitly, PCM yields better prediction results. To further solve the problem of narrow regularity score gaps in deep reconstruction methods, we decompose reconstruction into prediction and refinement, introducing ERM to reconstruct current prediction error and refine the coarse prediction. AnoPCN unifies reconstruction and prediction methods in an end-to-end framework, and it achieves state-of-the-art performance with better prediction results and larger regularity score gaps on three benchmark datasets including ShanghaiTech Campus, CUHK Avenue, and UCSD Ped2.
Muchao Ye, Xiaojiang Peng, Weihao Gan, Wei Wu 0021, Yu Qiao 0001
ACM Multimedia2
2019 Mutual Component Convolutional Neural Networks for Heterogeneous Face Recognition
abstract
HHeterogeneous face recognition (HFR) aims to identify a person from different facial modalities such as visible and near-infrared images. The main challenges of HFR lie in the large modality discrepancy and insufficient training samples. In this paper, we propose the Mutual Component Convolutional Neural Network (MC-CNN), a modal-invariant deep learning framework, to tackle these two issues simultaneously. Our MCCNN incorporates a generative module, i.e. the Mutual Component Analysis (MCA) [1], into modern deep convolutional neural networks by viewing MCA as a special fully-connected (FC) layer. Based on deep features, this FC layer is designed to extract modal-independent hidden factors, and is updated according to maximum likelihood analytic formulation instead of back propagation which prevents over-fitting from limited data naturally. In addition, we develop an MCA loss to update the network for modal-invariant feature learning. Extensive experiments show that our MC-CNN outperforms several finetuned baseline models significantly. Our methods achieve the state-of-the-art performance on CASIA NIR-VIS 2.0, CUHK NIR-VIS and IIIT-D Sketch dataset.
Zhongying Deng, Xiaojiang Peng, Zhifeng Li 0001, Yu Qiao 0001
IEEE Trans. Image Process.2
2018 Cascade Attention Networks For Group Emotion Recognition with Face, Body and Image Cues
abstract
This paper presents our approach for group-level emotion recognition sub-challenge in the EmotiW 2018. The task is to classify an image into one of the group emotions such as positive, negative, and neutral. Our approach mainly explores three cues, namely face, body and global image with recent deep networks. Our main contribution is two-fold. First, we introduce body based Convolutional Neural Networks (CNNs) into this task based on our previous winner method [18]. For body based CNNs, we crop all bodies in an image with the state-of-the-art human pose estimation method and train CNNs with the image-level label to capture. The body cue captures a full view of an individual. Second, we propose a cascade attention network for the face cue in images. This network exploits the importance of each face in an image to generates a global representation based on all faces. The cascade attention network is not only complementary with other models but also improves the naive average pooling method by about 2%. We finally achieve the second place in this sub-challenge with classification accuracies of 86.9% and 67.48% on the validation set and testing set, respectively.
Kai Wang 0036, Xiaoxing Zeng, Jianfei Yang 0001, Debin Meng, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001
ICMI6
2018 Deep Recurrent Multi-instance Learning with Spatio-temporal Features for Engagement Intensity Prediction
abstract
This paper elaborates the winner approach for engagement intensity prediction in the EmotiW Challenge 2018. The task is to predict the engagement level of a subject when he or she is watching an educational video in diverse conditions and different environments. Our approach formulates the prediction task as a multi-instance regression problem. We divide an input video sequence into segments and calculate the temporal and spatial features of each segment for regressing the intensity. Subject engagement, that is intuitively related with body and face changes in time domain, can be characterized by long short-term memory (LSTM) network. Hence, we build a multi-modal regression model based on multi-instance mechanism as well as LSTM. To make full use of training and validation data, we train different models for different data split and conduct model ensemble finally. Experimental results show that our method achieves mean squared error (MSE) of 0.0717 in the validation set, which improves the baseline results by 28%. Our methods finally win the challenge with MSE of 0.0626 on the testing set.
Jianfei Yang 0001, Kai Wang 0036, Xiaojiang Peng, Yu Qiao 0001
ICMI3
2018 Frankenstein: Learning Deep Face Representations Using Small Data
abstract
Deep convolutional neural networks have recently proven extremely effective for difficult face recognition problems in uncontrolled settings. To train such networks, very large training sets are needed with millions of labeled images. For some applications, such as near-infrared (NIR) face recognition, such large training data sets are not publicly available and difficult to collect. In this paper, we propose a method to generate very large training data sets of synthetic images by compositing real face images in a given data set. We show that this method enables to learn models from as few as 10 000 training images, which perform on par with models trained from 500 000 images. Using our approach, we also obtain state-of-the-art results on the CASIA NIR-VIS2.0 heterogeneous face recognition data set.
Guosheng Hu, Xiaojiang Peng, Yongxin Yang, Timothy M. Hospedales, Jakob Verbeek
IEEE Trans. Image Process.2
2017 Group emotion recognition with individual facial emotion CNNs and global image based CNNs
abstract
This paper presents our approach for group-level emotion recognition in the Emotion Recognition in the Wild Challenge 2017. The task is to classify an image into one of the group emotion such as positive, neutral or negative. Our approach is based on two types of Convolutional Neural Networks (CNNs), namely individual facial emotion CNNs and global image based CNNs. For the individual facial emotion CNNs, we first extract all the faces in an image, and assign the image label to all faces for training. In particular, we utilize a large-margin softmax loss for discriminative learning and we train two CNNs on both aligned and non-aligned faces. For the global image based CNNs, we compare several recent state-of-the-art network structures and data augmentation strategies to boost performance. For a test image, we average the scores from all faces and the image to predict the final group emotion category. We win the challenge with accuracies 83.9% and 80.9% on the validation set and testing set respectively, which improve the baseline results by about 30%.
Lianzhi Tan, Kaipeng Zhang, Kai Wang 0036, Xiaoxing Zeng, Xiaojiang Peng, Yu Qiao 0001
ICMI5
2016 Multi-region Two-Stream R-CNN for Action Detection
Xiaojiang Peng, Cordelia Schmid
ECCV (4)1
2016 Bag of visual words and fusion methods for action recognition: Comprehensive study and good practice
Xiaojiang Peng, Limin Wang 0002, Yu Qiao 0001
Comput. Vis. Image Underst.1
2016 An example-based approach to 3D man-made object reconstruction from line drawings
Changqing Zou, Tianfan Xue, Xiaojiang Peng, Honghua Li, Baochang Zhang 0001, Jianzhuang Liu
Pattern Recognit.3
2015 Multi-view descriptor mining via codeword net for action recognition
abstract
Action recognition is an important yet challenging task in computer vision. A successful and widely used framework in this field is the Bag of Visual Words (BoVW), wherein the first step is to extract local features. One critical property of local features is that they are often multi-view, e.g., dense trajectory feature includes both appearance and motion properties. Different types of features are aligned together in coding and pooling thus leading the process to be heavily entangled. Our motivation is to disentangle each sub-descriptor and let them contribute to the maximum extent. To achieve this, a codeword net is constructed via exploiting the relation between features and codewords. Based on the codeword net, features from the same viewpoint are pooled together. Experiments on two large scale action recognition datasets, UCF50 and HMDB51, demonstrate that our approach can enhance the state-of-the-art algorithms.
Jingyu Liu 0004, Yongzhen Huang, Xiaojiang Peng, Liang Wang 0001
ICIP3
2015 Sketch-based 3-D modeling for piecewise planar objects in single images
Changqing Zou, Xiaojiang Peng, Shifeng Chen, Hongbo Fu 0001, Jianzhuang Liu
Comput. Graph.2
2015 Gradient-based compressive sensing for noise image and video reconstruction
abstract
In this study, a fast gradient‐based compressive sensing (FGB‐CS) for noise image and video is proposed. Given a noise image or video, the authors first make it sparse by orthogonal transformation, and then reconstruct it by solving a convex optimisation problem with a novel gradient‐based method. The main contribution is twofold. Firstly, they deal with the noise signal reconstruction as a convex minimisation problem, and propose a new compressive sensing based on gradient‐based method for noise image and video. Secondly, to improve the computational efficiency of gradient‐based compressive sensing, they formulate the convex optimisation of noise signal reconstruction under Lipschitz gradient and replace the iteration parameter by the Lipschitz constant. With this strategy, the convergence of our FGB‐CS is reduced from O (1/ k ) to O (1/ k 2 ). Experimental results indicate that their FGB‐CS method is able to achieve better performance than several classical algorithms.
Yaonan Wang 0001, Xiaojiang Peng
IET Commun.3
2014 Multi-view Super Vector for Action Recognition
abstract
Images and videos are often characterized by multiple types of local descriptors such as SIFT, HOG and HOF, each of which describes certain aspects of object feature. Recognition systems benefit from fusing multiple types of these descriptors. Two widely applied fusion pipelines are descriptor concatenation and kernel average. The first one is effective when different descriptors are strongly correlated, while the second one is probably better when descriptors are relatively independent. In practice, however, different descriptors are neither fully independent nor fully correlated, and previous fusion methods may not be satisfying. In this paper, we propose a new global representation, Multi-View Super Vector (MVSV), which is composed of relatively independent components derived from a pair of descriptors. Kernel average is then applied on these components to produce recognition result. To obtain MVSV, we develop a generative mixture model of probabilistic canonical correlation analyzers (M-PCCA), and utilize the hidden factors and gradient vectors of M-PCCA to construct MVSV for video representation. Experiments on video based action recognition tasks show that MVSV achieves promising results, and outperforms FV and VLAD with descriptor concatenation or kernel average fusion strategy.
Zhuowei Cai, Limin Wang 0002, Xiaojiang Peng, Yu Qiao 0001
CVPR3
2014 Boosting VLAD with Supervised Dictionary Learning and High-Order Statistics
Xiaojiang Peng, Limin Wang 0002, Yu Qiao 0001, Qiang Peng
ECCV (3)1
2014 Action Recognition with Stacked Fisher Vectors
Xiaojiang Peng, Changqing Zou, Yu Qiao 0001, Qiang Peng
ECCV (5)1
2014 A Joint Evaluation of Dictionary Learning and Feature Encoding for Action Recognition
abstract
Many mid-level representations have been developed to replace traditional bag-of-words model (VQ+k-means) such as sparse coding, OMP-k with k-SVD, and fisher vector with GMM in image domain. These approaches can be split into a dictionary learning phase and a feature encoding phase which are often closely related. In this paper, we jointly evaluate the effect of these two phases for video-based action recognition. Specially, we compare several dictionary learning methods and feature encoding schemes through extensive experiments on the KTH and HMDB51 datasets. Experimental results indicate that fisher vector performs consistently better than the other encoding methods, and sparse coding is robust to different dictionaries even random weights. In addition, we observe that the advantages of sophisticated mid-level representations do not come from their specific dictionaries but the encoding mechanisms, and we can just use randomly selected exemplars as dictionaries for most of encoding methods. Finally, we achieve the state-of-the-art results on the HMDB51 and UCF101 by combining our configurations with improved dense trajectory features.
Xiaojiang Peng, Limin Wang 0002, Yu Qiao 0001, Qiang Peng
ICPR1
2014 Motion boundary based sampling and 3D co-occurrence descriptors for action recognition
Xiaojiang Peng, Yu Qiao 0001, Qiang Peng
Image Vis. Comput.1
2014 Large Margin Dimensionality Reduction for Action Similarity Labeling
abstract
Action recognition in videos is receiving extensive research interest due to its wide applications. This task needs to assign a specific action class for each video. In this paper, we study the problem of action similarity labeling (ASLAN) that is to verify whether two action videos present the same type of action or not. We show that both Fisher vector (FV) and vector of locally aggregated descriptors (VLAD) with dense trajectory features can achieve state-of-the-art performance on the ASLAN benchmark. Our main contribution is to develop a large margin dimensionality reduction (LMDR) method to compress high-dimensional FV and VLAD. Specially, we leverage the hinge loss objective function and stochastic gradient descent to optimize the discriminative projection matrix of these vectors. Extensive experiments on the ASLAN dataset indicate that our LMDR method not only reduces the dimension significantly but also improves the verification performance.
Xiaojiang Peng, Yu Qiao 0001, Qiang Peng, Qiong-Hua Wang
IEEE Signal Process. Lett.1
2013 Exploring Motion Boundary based Sampling and Spatial-Temporal Context Descriptors for Action Recognition
abstract
Feature representation is important for human action recognition.Recently, Wang et al. [25] proposed dense trajectory (DT) based features for action video representation and achieved state-of-the-art performance on several action datasets.In this paper, we improve the DT method in two folds.Firstly, we introduce a motion boundary based dense sampling strategy, which greatly reduces the number of valid trajectories while preserves the discriminative power.Secondly, we develop a set of new descriptors which describe the spatial-temporal context of motion trajectories.To evaluate the performance of the proposed methods, we conduct extensive experiments on three benchmarks including K-TH, YouTube and HMDB51.The results show that our sampling strategy significantly reduces the computational cost of point tracking without degrading performance.Meanwhile, we achieve superior performance than the state-of-the-art methods by utilizing our spatial-temporal context descriptors.
Xiaojiang Peng, Yu Qiao 0001, Qiang Peng, Xianbiao Qi
BMVC1