Xinxiao Wu

dblp:67/3509 · DBLP profile ↗
← Back
97ranked-venue papers
15as first author
47since 2021 · last 2026
0000-0002-2056-6947ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 67 · 8 first-author · 31 since 2021Artificial intelligence and machine learning · 58 · 12 first-author · 29 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 What to Trust? A Trust-aware Knowledge-guided Method for Zero-shot Object State Understanding in Videos
abstract
Object state understanding aims at recognizing the co-occurrence and transitions of multiple object states in videos. While learning from videos handles seen object states well, it struggles with novel ones. We address this task in a zero-shot setting by extracting state-specific knowledge from pre-trained models and using Vision-Language Models (VLMs) to verify whether such knowledge is visually grounded in videos. However, the extracted knowledge varies in its ability to distinguish states, and VLM observations are not always trustworthy. To address this issue, we propose a trust-aware knowledge-guided method to model knowledge trustworthiness and emphasize highly discriminative knowledge that VLMs can reliably observe. Specifically, we collect spatial knowledge for each object state from retrieved images and cues generated from a Large Language Model, then use VLMs to vote on each knowledge element by scoring its visual consistency with the video. In addition to a single scene, temporal dependencies of object states across scenes are also captured using a generative VLM. Under spatial and temporal constraints, we propose an adaptive knowledge refinement module that iteratively updates knowledge reliability weights to achieve a global consensus in object state inference across the video. Finally, object states are inferred by combining the refined weights with VLM voting results. Experiments on two datasets demonstrate the effectiveness of our method.
Yayun Qi, Xinxiao Wu
AAAI2
2026 TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
abstract
Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. However, a major challenge in developing generalized GUI agents is the lack of sufficient trajectory data across various operating systems and applications, mainly due to the high cost of manual annotations. In this paper, we propose the TongUI framework that transforms millions of multimodal web tutorials into GUI trajectories for generalized GUI agents. Concretely, we crawl GUI videos and articles from the Internet and process them into GUI agent trajectory data. Based on this, we construct the GUI-Net-1M dataset, which contains 1 million trajectories across five operating systems and over 280 applications. To the best of our knowledge, this is the largest open-source GUI trajectory dataset. We develop the TongUI agent by fine-tuning Qwen2.5-VL-3B/7B/32B models on GUI-Net-1M, which shows consistent performance improvements on commonly used grounding and navigation benchmarks, outperforming baseline agents by 10\% on multiple benchmarks, showing the effectiveness of the GUI-Net-1M dataset and underscoring the significance of our TongUI framework.
Bofei Zhang, Zirui Shang, Zhi Gao 0002, Xiaojian Ma 0001, Xinxiao Wu, Song-Chun Zhu, Qing Li 0003
AAAI8
2026 Open-vocabulary multi-label action recognition in movies via LLM-enhanced prompt tuning
Rongjiang Zhu, Xinxiao Wu, Yuheng Shi
Comput. Vis. Image Underst.2
2026 Challenging and enhancing the reasoning capacity of multimodal LLMs in context-violating images
Hongxi Li, Qiyuan Cheng, Xinxiao Wu
Pattern Recognit.4
2026 Advancing zero-shot humorous video understanding with test-time humor knowledge augmentation
Yayun Qi, Xinxiao Wu
Pattern Recognit.2
2026 Domain adaptive video summarization using generalized transformer
Yubo Zhu, Xinxiao Wu
Pattern Recognit.3
2026 Image-free multi-label image recognition via LLM-powered hierarchical prompt tuning
Shuo Yang 0002, Zirui Shang, Derong Deng, Xinxiao Wu, Qiyuan Cheng
Pattern Recognit.6
2026 A Comprehensive Survey on Video Summarization: Challenges and Advances
abstract
Video data is growing exponentially daily due to the popularity of video-sharing platforms and the proliferation of video capture devices. The video summarization task has been proposed to remove redundancy while maintaining as many critical parts of the video as possible so that users can browse and process videos more effectively, which has received increasing attention from researchers. The existing research addresses the challenges faced by video summarization methods from various perspectives, such as temporal dependency, data scarcity, user preference, and high precision. This paper reviews representative and state-of-the-art methods, analyzes recent research advances, datasets, and performance evaluations, and discusses future directions. We hope this survey can help future research explore the potential directions of video summarization methods.
Hongxi Li, Yubo Zhu, Zirui Shang, Xinxiao Wu
IEEE Trans. Circuits Syst. Video Technol.5
2026 How Vision-Language Tasks Benefit From Large Pre-Trained Models: A Survey
abstract
The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelligence and continuously attracts the attention of the research community. Despite improvements in overall performance, classic challenges still exist in vision-language tasks and hinder the development of this area. In recent years, the rise of pre-trained models is driving the research on vision-language tasks. Because of the massive scale of training data and model parameters, pre-trained models have exhibited excellent performance in numerous downstream tasks. Inspired by the powerful capabilities of pre-trained models, new paradigms have emerged to solve the classic challenges. Such methods have become mainstream in current research with increasing attention and rapid advances. In this paper, we present a comprehensive overview of how vision-language tasks benefit from pre-trained models. First, we review several main challenges in vision-language tasks and discuss the limitations of previous solutions before the era of pre-training. Next, we summarize the recent advances in incorporating pre-trained models to address the challenges in vision-language tasks. Finally, we analyze the potential risks associated with the inherent limitations of pre-trained models, discuss possible solutions, and attempt to provide future research directions.
Yayun Qi, Hongxi Li, Yiqi Song, Xinxiao Wu, Jiebo Luo 0001
IEEE Trans. Multim.4
2026 Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
abstract
Domain generalization (DG) aims to learn a model from source domains and apply it to unseen target domains with out-of-distribution data. Owing to CLIP's strong ability to encode semantic concepts, it has attracted increasing interest in domain generalization. However, CLIP often struggles to focus on task-relevant regions across domains,i.e.,domain-invariant regions, resulting in suboptimal performance on unseen target domains. To address this challenge, we propose an attention-refocusing scheme, calledSimulate, Refocus and Ensemble (SRE), which learns to reduce the domain shift by aligning the attention maps in CLIP via attention refocusing. SRE first simulates domain shifts by performing augmentation on the source data to generate simulated target domains. SRE then learns to reduce the domain shifts by refocusing the attention in CLIP between the source and simulated target domains. Finally, SRE utilizes ensemble learning to enhance the ability to capture domain-invariant attention maps between the source data and the simulated target data. Extensive experimental results on several datasets demonstrate that SRE generally achieves better results than state-of-the-art methods.
Zhi Gao 0002, Jin Chen 0009, Qingjie Zhao, Xinxiao Wu, Jiebo Luo 0001
IEEE Trans. Multim.5
2025 Video Summarization Using Denoising Diffusion Probabilistic Model
abstract
Video summarization aims to eliminate visual redundancy while retaining key parts of video to construct concise and comprehensive synopses. Most existing methods use discriminative models to predict the importance scores of video frames. However, these methods are susceptible to annotation inconsistency caused by the inherent subjectivity of different annotators when annotating the same video. In this paper, we introduce a generative framework for video summarization that learns how to generate summaries from a probability distribution perspective, effectively reducing the interference of subjective annotation noise. Specifically, we propose a novel diffusion summarization method based on the Denoising Diffusion Probabilistic Model (DDPM), which learns the probability distribution of training data through noise prediction, and generates summaries by iterative denoising. Our method is more resistant to subjective annotation noise, and is less prone to overfitting the training data than discriminative methods, with strong generalization ability. Moreover, to facilitate training DDPM with limited data, we employ an unsupervised video summarization model to implement the earlier denoising process. Extensive experiments on various datasets (TVSum, SumMe, and FPVSum) demonstrate the effectiveness of our method.
Zirui Shang, Yubo Zhu, Hongxi Li, Shuo Yang 0002, Xinxiao Wu
AAAI5
2025 LLM-Enhanced Action-Aware Multi-Modal Prompt Tuning for Image-Text Matching
Mengxiao Tian, Xinxiao Wu, Shuo Yang 0002
ICCV2
2025 METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
abstract
Open-vocabulary video visual relationship detection aims to detect objects and their relationships in videos without being restricted by predefined object or relationship categories. Existing methods leverage the rich semantic knowledge of pre-trained vision-language models such as CLIP to identify novel categories. They typically adopt a cascaded pipeline to first detect objects and then classify relationships based on the detected objects, which may lead to error propagation and thus suboptimal performance. In this paper, we propose Mutual EnhancemenT of Objects and Relationships (METOR), a query-based unified framework to jointly model and mutually enhance object detection and relationship classification in open-vocabulary scenarios. Under this framework, we first design a CLIP-based contextual refinement encoding module that extracts visual contexts of objects and relationships to refine the encoding of text features and object queries, thus improving the generalization of encoding to novel categories. Then we propose an iterative enhancement module to alternatively enhance the representations of objects and relationships by fully exploiting their interdependence to improve recognition performance. Extensive experiments on two public datasets, VidVRD and VidOR, demonstrate that our framework achieves state-of-the-art performance. Codes are at https://github.com/wangyongqi558/METOR.
Xinxiao Wu
IJCAI2
2025 Storyboard-guided Alignment for Fine-grained Video Action Recognition
abstract
Fine-grained video action recognition can be formulated as a video–text matching problem. Previous approaches primarily rely on global video semantics to consolidate video embeddings, often leading to misaligned video–text pairs due to inaccurate atomic-level action understanding. This inaccuracy arises due to i) videos with distinct global semantics may share similar atomic actions or visual appearances, and ii) atomic actions can be momentary, gradual, or not directly aligned with overarching video semantics. Inspired by storyboarding, where a script is segmented into individual shots, we propose a multi-granularity framework, SFAR. SFAR generates fine-grained descriptions of common atomic actions for each global semantic using a large language model. Unlike existing works that refine global semantics with auxiliary video frames, SFAR introduces a filtering metric to ensure correspondence between the descriptions and the global semantics, eliminating the need for direct video involvement and thereby enabling more nuanced recognition of subtle actions. By leveraging both global semantics and fine-grained descriptions, our SFAR effectively identifies prominent frames within videos, thereby improving the accuracy of embedding aggregation. Extensive experiments on various video action recognition datasets demonstrate the competitive performance of our SFAR in supervised, few-shot, and zero-shot settings.
Enqi Liu, Liyuan Pan, Yan Yang 0011, Yiran Zhong, Zhijing Wu 0001, Xinxiao Wu, Liu Liu 0009
NeurIPS6
2025 End-to-End Open-Vocabulary Video Visual Relationship Detection Using Multi-Modal Prompting
abstract
Open-vocabulary video visual relationship detection aims to expand video visual relationship detection beyond annotated categories by detecting unseen relationships between both seen and unseen objects in videos. Existing methods usually use trajectory detectors trained on closed datasets to detect object trajectories, and then feed these trajectories into large-scale pre-trained vision-language models to achieve open-vocabulary classification. Such heavy dependence on the pre-trained trajectory detectors limits their ability to generalize to novel object categories, leading to performance degradation. To address this challenge, we propose to unify object trajectory detection and relationship classification into an end-to-end open-vocabulary framework. Under this framework, we propose a relationship-aware open-vocabulary trajectory detector. It primarily consists of a query-based Transformer decoder, where the visual encoder of CLIP is distilled for frame-wise open-vocabulary object detection, and a trajectory associator. To exploit relationship context during trajectory detection, a relationship query is embedded into the Transformer decoder, and accordingly, an auxiliary relationship loss is designed to enable the decoder to perceive the relationships between objects explicitly. Moreover, we propose an open-vocabulary relationship classifier that leverages the rich semantic knowledge of CLIP to discover novel relationships. To adapt CLIP well to relationship classification, we design a multi-modal prompting method that employs spatio-temporal visual prompting for visual representation and vision-guided language prompting for language input. Extensive experiments on two public datasets, VidVRD and VidOR, demonstrate the effectiveness of our framework. Our framework is also applied to a more difficult cross-dataset scenario to further demonstrate its generalization ability.
Xinxiao Wu, Shuo Yang 0002, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Relational Distant Supervision for Image Captioning without Image-Text Pairs
abstract
Unsupervised image captioning aims to generate descriptions of images without relying on any image-sentence pairs for training. Most existing works use detected visual objects or concepts as bridge to connect images and texts. Considering that the relationship between objects carries more information, we use the object relationship as a more accurate connection between images and texts. In this paper, we adapt the idea of distant supervision that extracts the knowledge about object relationships from an external corpus and imparts them to images to facilitate inferring visual object relationships, without introducing any extra pre-trained relationship detectors. Based on these learned informative relationships, we construct pseudo image-sentence pairs for captioning model training. Specifically, our method consists of three modules: (1) a relationship learning module that learns to infer relationships from images under the distant supervision; (2) a relationship-to-sentence module that transforms the inferred relationships into sentences to generate pseudo image-sentence pairs; (3) an image captioning module that is trained by using the generated image-sentence pairs. Promising results on three datasets show that our method outperforms the state-of-the-art methods of unsupervised image captioning.
Yayun Qi, Wentian Zhao, Xinxiao Wu
AAAI3
2024 Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship Detection
abstract
Open-vocabulary video visual relationship detection aims to extend video visual relationship detection beyond annotated categories by detecting unseen relationships between objects in videos. Recent progresses in open-vocabulary perception, primarily driven by large-scale image-text pre-trained models like CLIP, have shown remarkable success in recognizing novel objects and semantic categories. However, directly applying CLIP-like models to video visual relationship detection encounters significant challenges due to the substantial gap between images and video object relationships. To address this challenge, we propose a multi-modal prompting method that adapts CLIP well to open-vocabulary video visual relationship detection by prompt-tuning on both visual representation and language input. Specifically, we enhance the image encoder of CLIP by using spatio-temporal visual prompting to capture spatio-temporal contexts, thereby making it suitable for object-level relationship representation in videos. Furthermore, we propose visual-guided language prompting to leverage CLIP's comprehensive semantic knowledge for discovering unseen relationship categories, thus facilitating recognizing novel video relationships. Extensive experiments on two public datasets, VidVRD and VidOR, demonstrate the effectiveness of our method, especially achieving a significant gain of nearly 10% in mAP on novel relationship categories on the VidVRD dataset.
Shuo Yang 0002, Xiaofeng Ji, Xinxiao Wu
AAAI4
2024 Event-based Few-shot Fine-grained Human Action Recognition
abstract
Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as high-dynamic or low lighting conditions. Event cameras can independently and sparsely capture brightness changes in a scene at microsecond resolution and high dynamic range, which offer a promising solution. However, the modality differences between events and RGB frames, and the lack of paired fine-grained data hinder the development of event-based FGH action recognition. Therefore, in this paper, we introduce the first Event Camera Fine-grained Human Action (E-FAction) dataset. This dataset comprises 3304 paired ‘event stream and RGB sequence’, covering 15 coarse action classes and 128 fine-grained actions. Then, we develop a versatile event feature extractor. Considering the spatial sparsity of event stream, we design two modules to mine the temporal motion and semantic features under the guidance of paired RGB frames, facilitating robust weight initialization for the feature extractor in few-shot FGH action recognition. We conduct extensive experiments on both published and our built synthetic and real datasets, and consistently achieve state-of-the-art performance compared to existing baselines. Code and dataset will be available at link.
Zonglin Yang 0002, Yan Yang 0011, Yuheng Shi, Hao Yang 0040, Ruikun Zhang, Liu Liu 0009, Xinxiao Wu, Liyuan Pan
IROS7
2024 Efficient Language-Driven Action Localization by Feature Aggregation and Prediction Adjustment
Zirui Shang, Shuo Yang 0002, Xinxiao Wu
PRCV (5)3
2024 DiffCLIP: Leveraging Stable Diffusion for Language Grounded 3D Classification
abstract
Large pre-trained models have revolutionized the field of computer vision by facilitating multi-modal learning. Notably, the CLIP model has exhibited remarkable proficiency in tasks such as image classification, object detection, and semantic segmentation. Nevertheless, its efficacy in processing 3D point clouds is restricted by the domain gap between the depth maps derived from 3D projection and the training images of CLIP.This paper introduces DiffCLIP, a novel pre-training framework that seamlessly integrates stable diffusion with ControlNet. The primary objective of DiffCLIP is to bridge the domain gap inherent in the visual branch. Furthermore, to address few-shot tasks in the textual branch, we incorporate a style-prompt generation module.Extensive experiments on the ModelNet10, ModelNet40, and ScanObjectNN datasets show that DiffCLIP has strong abilities for 3D understanding. By using stable diffusion and style-prompt generation, DiffCLIP achieves an accuracy of 43.2% for zero-shot classification on OBJ_BG of ScanObjectNN, which is state-of-the-art performance, and an accuracy of 82.4% for zero-shot classification on Model-Net10, which is also state-of-the-art performance.
Sitian Shen, Zilin Zhu, Linqian Fan, Harry Zhang, Xinxiao Wu
WACV5
2024 Hide and track: Towards blind video watermarking network in frequency domain
Guisong Wang, Xinxiao Wu
Neurocomputing4
2024 Source-Free Image-Text Matching via Uncertainty-Aware Learning
abstract
When applying a trained image-text matching model to a new scenario, the performance may largely degrade due to domain shift, which makes it impractical in real-world applications. In this paper, we make the first attempt on adapting the image-text matching model well-trained on a labeled source domain to an unlabeled target domain in the absence of source data, namely, source-free image-text matching. This task is challenging since it has no direct access to the source data when learning to reduce the doma in shift. To address this challenge, we propose a simple yet effective method that introduces uncertainty-aware learning to generate high-quality pseudo-pairs of image and text for target adaptation. Specifically, starting with using the pre-trained source model to retrieve several top-ranked image-text pairs from the target domain as pseudo-pairs, we then model uncertainty of each pseudo-pair by calculating the variance of retrieved texts (resp. images) given the paired image (resp. text) as query, and finally incorporate the uncertainty into an objective function to down-weight noisy pseudo-pairs for better training, thereby enhancing adaptation. This uncertainty-aware training approach can be generally applied on all existing models. Extensive experiments on the COCO and Flickr30K datasets demonstrate the effectiveness of the proposed method.
Mengxiao Tian, Shuo Yang 0002, Xinxiao Wu, Yunde Jia
IEEE Signal Process. Lett.3
2024 Commonsense Knowledge Prompting for Few-Shot Action Recognition in Videos
abstract
Few-shot action recognition in videos is challenging as the lack of supervision makes it extremely difficult to generalize well to unseen actions. To address this challenge, we propose a simple yet effective method, called knowledge prompting, which leverages commonsense knowledge of actions from external resources to prompt-tune a powerful pre-trained vision-language model for few-shot classification. To that end, we first collect a large-scale corpus of language descriptions of actions, defined as text proposals, to build an action knowledge base. The collection of text proposals is done by filling in a handcraft sentence template with an external action-related corpus or by extracting action-related phrases from captions of Web instruction videos. Next, we feed these text proposals to a pre-trained vision-language model along with video frames to generate matching scores of the proposals for each frame, and the scores can be treated as action semantics with strong generalization. Finally, we design a lightweight temporal modeling network to capture the temporal evolution of action semantics for classification. Extensive experiments on six benchmark datasets demonstrate that our method generally achieves state-of-the-art performance while reducing the training computational cost to 0.1% of the existing methods. Code is available athttps://github.com/OldStone0124/Knowledge-Prompting-for-FSAR.
Yuheng Shi, Xinxiao Wu, Hanxi Lin, Jiebo Luo 0001
IEEE Trans. Multim.2
2024 Dynamic Pathway for Query-Aware Feature Learning in Language-Driven Action Localization
abstract
Language-driven action localization aims to search a video segment in an untrimmed video, which is semantically relevant to an input language query. This task is challenging since language queries describe diverse actions with different motion characteristics and semantic granularities. Some actions, such as“the person takes off their shoes, and goes to the door”, are characterized by complex motion relationships, while others, such as“a person is standing holding a mirror in one hand”, are distinguished by salient body postures. In this paper, we propose a dynamic pathway between an exploitation module and an exploration module for query-aware feature learning to handle the diversity of actions. The exploitation module works in a coarse-to-fine manner, first learns the feature of general motion relationships to search the coarse segment of the target action and then learns the feature of subtle motion changes to predict the refined action boundaries. The exploration module functions in a point-to-area diffusion fashion, first learns the feature of sub-action pattern to search the salient postures of the target action and then learns the feature of temporal dependency to expand the posture frames to the action segment. The exploitation module and the exploration module are dynamically and adaptively selected to learn comprehensive representations of diverse actions to improve the action localization accuracy. Extensive experiments on the Charades-STA and TACoS datasets demonstrate that our method performs better than existing methods.
Shuo Yang 0002, Xinxiao Wu, Zirui Shang, Jiebo Luo 0001
IEEE Trans. Multim.2
2024 Boosting Entity-Aware Image Captioning With Multi-Modal Knowledge Graph
abstract
Entity-aware image captioning aims to describe named entities and events related to the image by utilizing the background knowledge in the associated article. This task remains challenging as it is difficult to learn the association between named entities and visual cues due to the long-tail distribution of named entities. Furthermore, the complexity of the article brings difficulty in extracting fine-grained relationships between entities to generate informative event descriptions about the image. To tackle these challenges, we propose a novel approach that constructs a multi-modal knowledge graph (MMKG) to associate the visual objects with named entities and capture the relationship between entities simultaneously with the help of external knowledge collected from the web. Specifically, we build a text sub-graph by extracting named entities and their relationships from the article, and build an image sub-graph by detecting the objects in the image. To connect these two sub-graphs, we propose a cross-modal entity matching module trained using a knowledge base that contains Wikipedia entries and the corresponding images. Finally, the MMKG is integrated into the captioning model via a graph attention mechanism. Extensive experiments on both GoodNews and NYTimes800 k datasets demonstrate the effectiveness of our method.
Wentian Zhao, Xinxiao Wu
IEEE Trans. Multim.2
2023 Meta-Causal Learning for Single Domain Generalization
abstract
Single domain generalization aims to learn a model from a single training domain (source domain) and apply it to multiple unseen test domains (target domains). Existing methods focus on expanding the distribution of the training domain to cover the target domains, but without estimating the domain shift between the source and target domains. In this paper, we propose a new learning paradigm, namely simulate-analyze-reduce, which first simulates the domain shift by building an auxiliary domain as the target domain, then learns to analyze the causes of domain shift, and finally learns to reduce the domain shift for model adaptation. Under this paradigm, we propose a meta-causal learning method to learn meta-knowledge, that is, how to infer the causes of domain shift between the auxiliary and source domains during training. We use the meta-knowledge to analyze the shift between the target and source domains during testing. Specifically, we perform multiple transformations on source data to generate the auxiliary domain, perform counterfactual inference to learn to discover the causal factors of the shift between the auxiliary and source domains, and incorporate the inferred causality into factor-aware domain alignments. Extensive experiments on several benchmarks of image classification show the effectiveness of our method.
Jin Chen 0009, Zhi Gao 0002, Xinxiao Wu, Jiebo Luo 0001
CVPR3
2023 Counterfactual Inference for Visual Relationship Detection in Videos
abstract
Visual relationship detection in videos is a challenging task since it requires not only to detect static relationships but also to infer dynamic relationships. Recent progress has been made through enriching visual representations by appearance and motion fusion or spatial and temporal reasoning, but without exploring the intrinsic causality between representations and predictions. In this paper, we propose a novel counterfactual inference method for video relationship detection, which infers the causal effects of appearance, motion and language features on the predictions of static and dynamic relationships. Specifically, starting with building a causal graph to represent the causality between features and relationship categories, we then construct counterfactual scenes by intervening the features to infer their effects on prediction, and finally incorporate the inferred effects into the relationship categorization by adaptively learning the weights of appearance, motion and language. Extensive experiments on two benchmark datasets demonstrate the effectiveness of our method.
Xiaofeng Ji, Xinxiao Wu
ICME3
2023 Teaching What You Should Teach: A Data-Based Distillation Method
abstract
In real teaching scenarios, an excellent teacher always teaches what he (or she) is good at but the student is not. This gives the student the best assistance in making up for his (or her) weaknesses and becoming a good one overall. Enlightened by this, we introduce the "Teaching what you Should Teach" strategy into a knowledge distillation framework, and propose a data-based distillation method named "TST" that searches for desirable augmented samples to assist in distilling more efficiently and rationally. To be specific, we design a neural network-based data augmentation module with priori bias to find out what meets the teacher's strengths but the student's weaknesses, by learning magnitudes and probabilities to generate suitable data samples. By training the data augmentation module and the generalized distillation paradigm alternately, a student model is learned with excellent generalization ability. To verify the effectiveness of our method, we conducted extensive comparative experiments on object recognition, detection, and segmentation tasks. The results on the CIFAR-100, ImageNet-1k, MS-COCO, and Cityscapes datasets demonstrate that our method achieves state-of-the-art performance on almost all teacher-student pairs. Furthermore, we conduct visualization studies to explore what magnitudes and probabilities are needed for the distillation process.
Shitong Shao, Huanran Chen, Zhen Huang 0007, Linrui Gong, Shuai Wang 0048, Xinxiao Wu
IJCAI6
2023 Probability Distribution Based Frame-supervised Language-driven Action Localization
abstract
Frame-supervised language-driven action localization aims to localize action boundaries in untrimmed videos corresponding to the input natural language query, with only a single frame annotation within the target action in training. This task is challenging due to the absence of complete and accurate annotation of action boundaries, hindering visual-language alignment and action boundary prediction. To address this challenge, we propose a novel method that introduces distribution functions to model both the probability of action frame and that of boundary frame. Specifically, we assign each video frame the probability of being the action frame based on the estimated shape parameters of the distribution function, serving as a foreground pseudo-label that guides cross-modal feature learning. Moreover, we model the probabilities of start frame and end frame of the target action using different distribution functions, and then estimate the probability of each action candidate being a positive candidate based on its start and end boundaries, which facilitates predicting action boundaries by exploring more positive terms in training. Experiments on two benchmark datasets demonstrate that our method outperforms existing methods, achieving a gain of more than 10% of R1@ 0.5 on the challenging TACoS dataset. These results emphasize the significance of generating pseudo labels with appropriate probabilities via distribution functions to address the challenge of frame-supervised language-driven action localization.
Shuo Yang 0002, Zirui Shang, Xinxiao Wu
ACM Multimedia3
2023 Sentimental Visual Captioning using Multimodal Transformer
Xinxiao Wu
Int. J. Comput. Vis.1
2023 Topic-aware video summarization using multimodal transformer
Yubo Zhu, Wentian Zhao, Xinxiao Wu
Pattern Recognit.4
2023 Adaptive Latent Graph Representation Learning for Image-Text Matching
abstract
Image-text matching is a challenging task due to the modality gap. Many recent methods focus on modeling entity relationships to learn a common embedding space of image and text. However, these methods suffer from distractions of entity relationships such as irrelevant visual regions in an image and noisy textual words in a text. In this paper, we propose an adaptive latent graph representation learning method to reduce the distractions of entity relationships for image-text matching. Specifically, we use an improved graph variational autoencoder to separate the distracting factors and latent factor of relationships and jointly learn latent textual graph representations, latent visual graph representations, and a visual-textual graph embedding space. We also introduce an adaptive cross-attention mechanism to perform feature attending on the latent graph representations across images and texts, thus further narrowing the modality gap to boost the matching performance. Extensive experiments on two public datasets, Flickr30K and COCO, show the effectiveness of our method.
Mengxiao Tian, Xinxiao Wu, Yunde Jia
IEEE Trans. Image Process.2
2022 Adaptive Image-to-Video Scene Graph Generation via Knowledge Reasoning and Adversarial Learning
abstract
Scene graph in a video conveys a wealth of information about objects and their relationships in the scene, thus benefiting many downstream tasks such as video captioning and visual question answering. Existing methods of scene graph generation require large-scale training videos annotated with objects and relationships in each frame to learn a powerful model. However, such comprehensive annotation is time-consuming and labor-intensive. On the other hand, it is much easier and less cost to annotate images with scene graphs, so we investigate leveraging annotated images to facilitate training a scene graph generation model for unannotated videos, namely image-to-video scene graph generation. This task presents two challenges: 1) infer unseen dynamic relationships in videos from static relationships in images due to the absence of motion information in images; 2) adapt objects and static relationships from images to video frames due to the domain shift between them. To address the first challenge, we exploit external commonsense knowledge to infer the unseen dynamic relationship from the temporal evolution of static relationships. We tackle the second challenge by hierarchical adversarial learning to reduce the data distribution discrepancy between images and video frames. Extensive experiment results on two benchmark video datasets demonstrate the effectiveness of our method.
Jin Chen 0009, Xiaofeng Ji, Xinxiao Wu
AAAI3
2022 Adaptive Recursive Circle Framework for Fine-Grained Action Recognition
abstract
Intuitively, distinguishing fine-grained actions in videos requires recursively capturing subtle visual cues and learning abstract features. However, existing deep neural network based methods are counter-intuitive in that their network layers do not explicitly model the recursive feature abstraction. Therefore, we are motivated to propose an Adaptive Recursive Circle (ARC) framework that equips common neural network layers with recursive attention and recursive fusion. ARC layer inherits the same operators and parameters as the original layer, but, most critically, it treats the layer input as an evolving state, thus explicitly achieving recursive feature abstraction by alternating the state update and the feature generation. Specifically, at each recursive step, the input state is firstly updated via both recursive attention and recursive fusion from the previously generated features, and then the feature abstraction is performed with the newly updated input state. Significant improvements are observed on multiple datasets. For example, an ARC-equipped TSM-ResNet-18 outperforms TSM-ResNet-50 on the Something-Something V1 and Diving48 datasets with only half over-heads. Code will be available at: https://github.com/0HaNC/ARC-ActionRecog.
Hanxi Lin, Wentian Zhao, Xinxiao Wu
ICME3
2022 Entity-aware and Motion-aware Transformers for Language-driven Action Localization
abstract
Language-driven action localization in videos is a challenging task that involves not only visual-linguistic matching but also action boundary prediction. Recent progress has been achieved through aligning language queries to video segments, but estimating precise boundaries is still under-explored. In this paper, we propose entity-aware and motion-aware Transformers that progressively localize actions in videos by first coarsely locating clips with entity queries and then finely predicting exact boundaries in a shrunken temporal region with motion queries. The entity-aware Transformer incorporates the textual entities into visual representation learning via cross-modal and cross-frame attentions to facilitate attending action-related video clips. The motion-aware Transformer captures fine-grained motion changes at multiple temporal scales via integrating long short-term memory into the self-attention module to further improve the precision of action boundary prediction. Extensive experiments on the Charades-STA and TACoS datasets demonstrate that our method achieves better performance than existing methods.
Shuo Yang 0002, Xinxiao Wu
IJCAI2
2022 Learning Cooperative Neural Modules for Stylized Image Captioning
Xinxiao Wu, Wentian Zhao, Jiebo Luo 0001
Int. J. Comput. Vis.1
2022 Exploiting Informative Video Segments for Temporal Action Localization
abstract
We propose a novel method of exploiting informative video segments by learning segment weights for temporal action localization in untrimmed videos. Informative video segments represent the intrinsic motion and appearance of an action, and thus contribute crucially to action localization. The learned segment weights represent the informativeness of video segments to recognize actions and help infer the boundaries required to temporally localize actions. We build a supervised temporal attention network (STAN) that includes a supervised segment-level attention module to dynamically learn the weights of video segments, and a feature-level attention module to effectively fuse multiple features of segments. Through the cascade of the attention modules, STAN exploits informative video segments and generates descriptive and discriminative video representations. We use a proposal generator and a classifier to estimate the boundaries of actions and classify the classes of actions. Extensive experiments are conducted on two public benchmarks, i.e., THUMOS2014 and ActivityNet1.3. The results demonstrate that our proposed method achieves competitive performance compared with existing state-of-the-art methods. Moreover, compared with the baseline method that treats video segments equally, STAN achieves significant improvements with an increase of the mean average precision from 30.4% to 39.8% on the THUMOS2014 dataset, and from 31.4% to 35.9% on the ActivityNet1.3 dataset, demonstrating the effectiveness of learning informative video segments for temporal action localization.
Che Sun, Hao Song 0002, Xinxiao Wu, Yunde Jia, Jiebo Luo 0001
IEEE Trans. Multim.3
2022 Domain Adversarial Reinforcement Learning for Partial Domain Adaptation
abstract
Partial domain adaptation aims to transfer knowledge from a label-rich source domain to a label-scarce target domain (i.e., the target categories are a subset of the source ones), which relaxes the common assumption in traditional domain adaptation that the label space is fully shared across different domains. In this more general and practical scenario on partial domain adaptation, a major challenge is how to select source instances from the shared categories to ensure positive transfer for the target domain. To address this problem, we propose a domain adversarial reinforcement learning (DARL) framework to progressively select source instances to learn transferable features between domains by reducing the domain shift. Specifically, we employ a deep Q-learning to learn policies for an agent to make selection decisions by approximating the action-value function. Moreover, domain adversarial learning is introduced to learn a common feature subspace for the selected source instances and the target instances, and also to contribute to the reward calculation for the agent that is based on the relevance of the selected source instances with respect to the target domain. Extensive experiments on several benchmark data sets clearly demonstrate the superior performance of our proposed DARL over existing state-of-the-art methods for partial domain adaptation.
Jin Chen 0009, Xinxiao Wu, Lixin Duan, Shenghua Gao
IEEE Trans. Neural Networks Learn. Syst.2
2021 Spatial-temporal Causal Inference for Partial Image-to-video Adaptation
abstract
Image-to-video adaptation leverages off-the-shelf learned models in labeled images to help classification in unlabeled videos, thus alleviating the high computation overhead of training a video classifier from scratch. This task is very challenging since there exist two types of domain shifts between images and videos: 1) spatial domain shift caused by static appearance variance between images and video frames, and 2) temporal domain shift caused by the absence of dynamic motion in images. Moreover, for different video classes, these two domain shifts have different effects on the domain gap and should not be treated equally during adaptation. In this paper, we propose a spatial-temporal causal inference framework for image-to-video adaptation. We first construct a spatial-temporal causal graph to infer the effects of the spatial and temporal domain shifts by performing counterfactual causality. We then learn causality-guided bidirectional heterogeneous mappings between images and videos to adaptively reduce the two domain shifts. Moreover, to relax the assumption that the label spaces of the image and video domains are the same by the existing methods, we incorporate class-wise alignment into the learning of image-video mappings to perform partial image-to-video adaptation where the image label space subsumes the video label space. Extensive experiments on several video datasets have validated the effectiveness of our proposed method.
Jin Chen 0009, Xinxiao Wu, Yao Hu 0002, Jiebo Luo 0001
AAAI2
2021 Anticipating Future Relations via Graph Growing for Action Prediction
abstract
Predicting actions from partially observed videos is challenging as the partial videos containing incomplete action executions have insufficient discriminative information for classification. Recent progress has been made through enriching the features of the observed video part or generating the features for the unobserved video part, but without explicitly modeling the fine-grained evolution of visual object relations over both space and time. In this paper, we investigate how the interaction and correlation between visual objects evolve and propose a graph growing method to anticipate future object relations from limited video observations for reliable action prediction. There are two tasks in our method. First, we work with spatial-temporal graph neural networks to reason object relations in the observed video part. Then, we synthesize the spatial-temporal relation representation for the unobserved video part via graph node generation and aggregation. These two tasks are jointly learned to enable the anticipated future relation representation informative to action prediction. Experimental results on two action video datasets demonstrate the effectiveness of our method.
Xinxiao Wu
AAAI1
2021 Image Captioning with Inherent Sentiment
abstract
We propose a new task called sentimental image captioning which aims to generate captions with the inherent sentiment reflected by the image. Compared with the stylized image captioning task that requires a predefined style independent of the image, our new task can automatically analyze the inherent sentiment tendency within the image. With this in mind, we propose an Inherent Sentiment Image Captioning (InSenti-Cap) method that first extracts the content and sentiment information from the image, and then fuses these information into the sentimental sentence generation via an attention mechanism. To effectively train the proposed model using the pairs of image and factual caption in existing captioning dataset and the extra sentiment corpus, we propose a two-stage training strategy that involves a sentimental regularization and a sentimental reward to enable the model to generate fluent and relevant sentences with inherent sentimental styles. Experiments demonstrate the effectiveness of our method.
Yunhui Hu, Xinxiao Wu
ICME3
2021 Multi-modal Dependency Tree for Video Captioning
abstract
Generating fluent and relevant language to describe visual content is critical for the video captioning task. Many existing methods generate captions using sequence models that predict words in a left-to-right order. In this paper, we investigate a graph-structured model for caption generation by explicitly modeling the hierarchical structure in the sentences to further improve the fluency and relevance of sentences. To this end, we propose a novel video captioning method that generates a sentence by first constructing a multi-modal dependency tree and then traversing the constructed tree, where the syntactic structure and semantic relationship in the sentence are represented by the tree topology. To take full advantage of the information from both vision and language, both the visual and textual representation features are encoded into each tree node. Different from existing dependency parsing methods that generate uni-modal dependency trees for language understanding, our method construct s multi-modal dependency trees for language generation of images and videos. We also propose a tree-structured reinforcement learning algorithm to effectively optimize the captioning model where a novel reward is designed by evaluating the semantic consistency between the generated sub-tree and the ground-truth tree. Extensive experiments on several video captioning datasets demonstrate the effectiveness of the proposed method.
Wentian Zhao, Xinxiao Wu, Jiebo Luo 0001
NeurIPS2
2021 Spatial-Temporal Relation Reasoning for Action Prediction in Videos
Xinxiao Wu, Jingyi Hou, Hanxi Lin, Jiebo Luo 0001
Int. J. Comput. Vis.1
2021 Boundary discrimination and proposal evaluation for temporal action proposal generation
Bing Bing, Xinxiao Wu
Multim. Tools Appl.3
2021 Joint Learning of Multiple Latent Domains and Deep Representations for Domain Adaptation
abstract
In domain adaptation, the automatic discovery of multiple latent source domains has succeeded by capturing the intrinsic structure underlying the source data. Different from previous works that mainly rely on shallow models for domain discovery, we propose a novel unified framework based on deep neural networks to jointly address latent domain prediction from source data and deep representation learning from both source and target data. Within this framework, an iterative algorithm is proposed to alternate between 1) utilizing a new probabilistic hierarchical clustering method to separate the source domain into latent clusters and 2) training deep neural networks by using the domain membership as the supervision to learn deep representations. The key idea behind this joint learning framework is that good representations can help to improve the prediction accuracy of latent domains and, in turn, domain prediction results can provide useful supervisory information for feature learning. During the training of the deep model, a domain prediction loss, a domain confusion loss, and a task-specific classification loss are effectively integrated to enable the learned feature to distinguish between different latent source domains, transfer between source and target domains, and become semantically meaningful among different classes. Trained in an end-to-end fashion, our framework outperforms the state-of-the-art methods for latent domain discovery, as validated by extensive experiments on both object classification and human action-recognition tasks.
Xinxiao Wu, Jin Chen 0009, Feiwu Yu, Mingyu Yao, Jiebo Luo 0001
IEEE Trans. Cybern.1
2021 Sequential Instance Refinement for Cross-Domain Object Detection in Images
abstract
Cross-domain object detection in images has attracted increasing attention in the past few years, which aims at adapting the detection model learned from existing labeled images (source domain) to newly collected unlabeled ones (target domain). Existing methods usually deal with the cross-domain object detection problem through direct feature alignment between the source and target domains at the image level, the instance level (i.e., region proposals) or both. However, we have observed that directly aligning features of all object instances from the two domains often results in the problem of negative transfer, due to the existence of (1) outlier target instances that contain confusing objects not belonging to any category of the source domain and thus are hard to be captured by detectors and (2) low-relevance source instances that are considerably statistically different from target instances although their contained objects are from the same category. With this in mind, we propose a reinforcement learning based method, coined as sequential instance refinement, where two agents are learned to progressively refine both source and target instances by taking sequential actions to remove both outlier target instances and low-relevance source instances step by step. Extensive experiments on several benchmark datasets demonstrate the superior performance of our method over existing state-of-the-art baselines for cross-domain object detection.
Jin Chen 0009, Xinxiao Wu, Lixin Duan, Lin Chen 0021
IEEE Trans. Image Process.2
2021 Cross-Domain Image Captioning via Cross-Modal Retrieval and Model Adaptation
abstract
In recent years, large scale datasets of paired images and sentences have enabled the remarkable success in automatically generating descriptions for images, namely image captioning. However, it is labour-intensive and time-consuming to collect a sufficient number of paired images and sentences in each domain. It may be beneficial to transfer the image captioning model trained in an existing domain with pairs of images and sentences (i.e., source domain) to a new domain with only unpaired data (i.e., target domain). In this paper, we propose a cross-modal retrieval aided approach to cross-domain image captioning that leverages a cross-modal retrieval model to generate pseudo pairs of images and sentences in the target domain to facilitate the adaptation of the captioning model. To learn the correlation between images and sentences in the target domain, we propose an iterative cross-modal retrieval process where a cross-modal retrieval model is first pre-trained using the source domain data and then applied to the target domain data to acquire an initial set of pseudo image-sentence pairs. The pseudo image-sentence pairs are further refined by iteratively fine-tuning the retrieval model with the pseudo image-sentence pairs and updating the pseudo image-sentence pairs using the retrieval model. To make the linguistic patterns of the sentences learned in the source domain adapt well to the target domain, we propose an adaptive image captioning model with a self-attention mechanism fine-tuned using the refined pseudo image-sentence pairs. Experimental results on several settings where MSCOCO is used as the source domain and five different datasets (Flickr30k, TGIF, CUB-200, Oxford-102 and Conceptual) are used as the target domains demonstrate that our method achieves mostly better or comparable performance against the state-of-the-art methods. We also extend our method to cross-domain video captioning where MSR-VTT is used as the source domain and two other datasets (MSVD and Charades Captions) are used as the target domains to further demonstrate the effectiveness of our method.
Wentian Zhao, Xinxiao Wu, Jiebo Luo 0001
IEEE Trans. Image Process.2
2020 Joint Commonsense and Relation Reasoning for Image and Video Captioning
abstract
Exploiting relationships between objects for image and video captioning has received increasing attention. Most existing methods depend heavily on pre-trained detectors of objects and their relationships, and thus may not work well when facing detection challenges such as heavy occlusion, tiny-size objects, and long-tail classes. In this paper, we propose a joint commonsense and relation reasoning method that exploits prior knowledge for image and video captioning without relying on any detectors. The prior knowledge provides semantic correlations and constraints between objects, serving as guidance to build semantic graphs that summarize object relationships, some of which cannot be directly perceived from images or videos. Particularly, our method is implemented by an iterative learning algorithm that alternates between 1) commonsense reasoning for embedding visual regions into the semantic space to build a semantic graph and 2) relation reasoning for encoding semantic graphs to generate sentences. Experiments on several benchmark datasets validate the effectiveness of our prior knowledge-based approach.
Jingyi Hou, Xinxiao Wu, Xiaoxun Zhang, Yayun Qi, Yunde Jia, Jiebo Luo 0001
AAAI2
2020 MemCap: Memorizing Style Knowledge for Image Captioning
abstract
Generating stylized captions for images is a challenging task since it requires not only describing the content of the image accurately but also expressing the desired linguistic style appropriately. In this paper, we propose MemCap, a novel stylized image captioning method that explicitly encodes the knowledge about linguistic styles with memory mechanism. Rather than relying heavily on a language model to capture style factors in existing methods, our method resorts to memorizing stylized elements learned from training corpus. Particularly, we design a memory module that comprises a set of embedding vectors for encoding style-related phrases in training corpus. To acquire the style-related phrases, we develop a sentence decomposing algorithm that splits a stylized sentence into a style-related part that reflects the linguistic style and a content-related part that contains the visual content. When generating captions, our MemCap first extracts content-relevant style knowledge from the memory module via an attention mechanism and then incorporates the extracted knowledge into a language model. Extensive experiments on two stylized image captioning datasets (SentiCap and FlickrStyle10K) demonstrate the effectiveness of our method.
Wentian Zhao, Xinxiao Wu, Xiaoxun Zhang
AAAI2
2020 Preserving Global and Local Temporal Consistency for Arbitrary Video Style Transfer
abstract
Video style transfer is a challenging task that requires not only stylizing video frames but also preserving temporal consistency among them. Many existing methods resort to optical flow for maintaining the temporal consistency in stylized videos. However, optical flow is sensitive to occlusions and rapid motions, and its training processing speed is quite slow, which makes it less practical in real-world applications. In this paper, we propose a novel fast method that explores both global and local temporal consistency for video style transfer without estimating optical flow. To preserve the temporal consistency of the entire video (i.e., global consistency), we use structural similarity index instead of flow optical and propose a self-similarity loss to ensure the temporal structure similarity between the stylized video and the source video. Furthermore, to enhance the coherence between adjacent frames (i.e., local consistency), a self-attention mechanism is designed to attend the previous stylized frame for synthesizing the current frame. Extensive experiments demonstrate that our method generally achieves better visual results and runs faster than the state-of-the-art methods, which validates the superiority of simultaneously preserving global and local temporal consistency for video style transfer
Xinxiao Wu
ACM Multimedia1
2020 Hierarchical Matching and Reasoning for Action Localization via Language Query
Xinxiao Wu
PRCV (3)2
2020 Incremental transfer learning for video annotation via grouped heterogeneous sources
abstract
Here, the authors focus on incrementally acquiring heterogeneous knowledge from both internet and publicly available datasets to reduce the tedious and expensive labelling efforts required in video annotation. An incremental transfer learning framework is presented to integrate heterogeneous source knowledge and update the annotation model incrementally during the transfer learning process. Under this framework, web images and existing action videos form the source domain to provide labelled static and motion information of the target domain videos, respectively. Moreover, according to the semantic of the source domain data, all the source domain data are partitioned into several groups. Different from traditional methods, which compare the entire target domain videos with each source group from the source domain, the authors treat the group weights as sample‐specific variables and optimise them along with new adding data. Two regularisers are used to prevent the incremental learning process from negative transfer. Experimental results on the two large‐scale consumer video datasets (i.e. multimedia event detection (MED) and Columbia consumer video (CCV)) show the effectiveness of the proposed method.
Hao Song 0002, Xinxiao Wu, Yunde Jia
IET Comput. Vis.3
2020 Confidence-Guided Self Refinement for Action Prediction in Untrimmed Videos
abstract
Many existing methods formulate the action prediction task as recognizing early parts of actions in trimmed videos. In this paper, we focus on predicting actions from ongoing untrimmed videos where actions might not happen at the very beginning of videos. It is extremely challenging to predict actions in such untrimmed videos due to ambiguous or even no information of actions in the early parts of videos. To address this problem, we propose a prediction confidence that assesses the decision quality of a prediction model. Guided by the confidence, the model continuously refines the prediction results by itself with the increasing observed video frames. Specifically, we build a Self Prediction Refining Network (SPR-Net) which incrementally learns the confidence for action prediction. SPR-Net consists of three modules: a temporal hybrid network, an incremental confidence learner, and a self-refining Gumbel softmax sampler. The temporal hybrid network generates the action category distributions by integrating static scene and dynamic motion information. The incremental confidence learner calculates the confidence in an incremental manner, judging the extent to which the temporal hybrid network should believe its prediction result. The self-refining Gumbel softmax sampler models the mutual relationship between the prediction confidence and the category distribution, which enables them to be jointly learned in an end-to-end fashion. We also present a sparse self-attention mechanism to encode local spatio-temporal features into the frame-level motion representation to further improve the prediction performance. Extensive experiments on five datasets (i.e., UT-Interaction, BIT-Interaction, UCF101, THUMOS14, and ActivityNet) validate the effectiveness of the proposed method.
Jingyi Hou, Xinxiao Wu, Jiebo Luo 0001, Yunde Jia
IEEE Trans. Image Process.2
2020 Learning Normal Patterns via Adversarial Attention-Based Autoencoder for Abnormal Event Detection in Videos
abstract
Automatically detecting anomalies in videos is a challenging problem due to non-deterministic definitions of abnormal events and lack of sufficient training data. To address these issues, we propose an autoencoder coupled with attention model to discover normal patterns in videos via adversarial learning. Abnormal events are detected by diverging them from the normal patterns with the reconstruction error produced by the autoencoder. To this end, we build an end-to-end trainable adversarial attention-based autoencoder network, called Ada-Net, to make the reconstructed frames indistinguishable from original frames. The Ada-Net combines an autoencoder network and a GAN model that is used to benefit enhancing the reconstruction ability of the autoencoder. To further improve the reconstruction performance, we integrate an attention model into the decoder to dynamically select informative parts of encoding features for decoding. The attenion mechanism is helpful to preserving important information for learning intrinsic normal patterns. Evaluations on four challenging datasets, including the Subway, the UCSD Pedestrian, the CUHK Avenue, and the ShanghaiTech datasets, demonstrate the effectiveness of the proposed method.
Hao Song 0002, Che Sun, Xinxiao Wu, Yunde Jia
IEEE Trans. Multim.3
2019 Joint Syntax Representation Learning and Visual Cue Translation for Video Captioning
abstract
Video captioning is a challenging task that involves not only visual perception but also syntax representation learning. Recent progress in video captioning has been achieved through visual perception, but syntax representation learning is still under-explored. We propose a novel video captioning approach that takes into account both visual perception and syntax representation learning to generate accurate descriptions of videos. Specifically, we use sentence templates composed of Part-of-Speech (POS) tags to represent the syntax structure of captions, and accordingly, syntax representation learning is performed by directly inferring POS tags from videos. The visual perception is implemented by a mixture model which translates visual cues into lexical words that are conditional on the learned syntactic structure of sentences. Thus, a video captioning task consists of two sub-tasks: video POS tagging and visual cue translation, which are jointly modeled and trained in an end-to-end fashion. Evaluations on three public benchmark datasets demonstrate that our proposed method achieves substantially better performance than the state-of-the-art methods, which validates the superiority of joint modeling of syntax representation learning and visual perception for video captioning.
Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo 0001, Yunde Jia
ICCV2
2019 Learning Weighted Video Segments for Temporal Action Localization
Che Sun, Hao Song 0002, Xinxiao Wu, Yunde Jia
PRCV (1)3
2019 Exploiting Human Pose for Weakly-Supervised Temporal Action Localization
Xinxiao Wu
PRCV (3)3
2019 Combining multiple deep cues for action recognition
Xinxiao Wu
Multim. Tools Appl.2
2019 Exploiting Images for Video Recognition: Heterogeneous Feature Augmentation via Symmetric Adversarial Learning
abstract
Training deep models of video recognition usually requires sufficient labeled videos in order to achieve good performance without over-fitting. However, it is quite labor-intensive and time-consuming to collect and annotate a large amount of videos. Moreover, training deep neural networks on large-scale video datasets always demands huge computational resources which further hold back many researchers and practitioners. To resolve that, collecting and training on annotated images are much easier. However, thoughtlessly applying images to help recognize videos may result in noticeable performance degeneration due to the well-known domain shift and feature heterogeneity. This proposes a novel symmetric adversarial learning approach for heterogeneous image-to-video adaptation, which augments deep image and video features by learning domain-invariant representations of source images and target videos. Primarily focusing on an unsupervised scenario where the labeled source images are accompanied by unlabeled target videos in the training phrase, we present a data-driven approach to respectively learn the augmented features of images and videos with superior transformability and distinguishability. Starting with learning a common feature space (called image-frame feature space) between images and video frames, we then build new symmetric generative adversarial networks (Sym-GANs) where one GAN maps image-frame features to video features and the other maps video features to image-frame features. Using the Sym-GANs, the source image feature is augmented with the generated video-specific representation to capture the motion dynamics while the target video feature is augmented with the image-specific representation to take the static appearance information. Finally, the augmented features from the source domain are fed into a network with fully connected layers for classification. Thanks to an end-to-end training procedure of the Sym-GANs and the classification network, our approach achieves better results than other state-of-the-arts, which is clearly validated by experiments on two video datasets, i.e., the UCF101 and HMDB51 datasets.
Feiwu Yu, Xinxiao Wu, Lixin Duan
IEEE Trans. Image Process.2
2019 Temporal Action Localization in Untrimmed Videos Using Action Pattern Trees
abstract
In this paper, we present a novel framework of automatically localizing action instances based on action pattern trees (AP-Trees) in a long untrimmed video. For localizing action instances in videos with varied temporal lengths, we first split videos into sequential segments and then use the AP-Trees to produce precise temporal boundaries of action instances. The AP-Trees can exploit the temporal information between segments of videos based on the label vectors of segments, by learning the occurrence frequency and order of segments. In AP-Trees, nodes stand for action class labels of segments and edges represent the temporal relationships between two consecutive segments. Thus, we can discover the occurrence frequencies of segments by searching paths of AP-Trees. In order to obtain accurate labels of video segments, we introduce deep neural networks to annotate the segments by simultaneously leveraging the spatio-temporal information and the high-level semantic feature of segments. In the networks, informative action maps are generated by a global average pooling layer to retain the spatio-temporal information of segments. An overlap loss function is employed to further improve the precision of label vectors of segments by considering the temporal overlap between segments and the ground truth. The experiments on THUMOS2014, MSR ActionII, and MPII Cooking datasets demonstrate the effectiveness of the method.
Hao Song 0002, Xinxiao Wu, Yuwei Wu 0001, Yunde Jia
IEEE Trans. Multim.2
2018 Unsupervised Deep Learning of Mid-Level Video Representation for Action Recognition
abstract
Current deep learning methods for action recognition rely heavily on large scale labeled video datasets. Manually annotating video datasets is laborious and may introduce unexpected bias to train complex deep models for learning video representation. In this paper, we propose an unsupervised deep learning method which employs unlabeled local spatial-temporal volumes extracted from action videos to learn midlevel video representation for action recognition. Specifically, our method simultaneously discovers mid-level semantic concepts by discriminative clustering and optimizes local spatial-temporal features by two relatively small and simple deep neural networks. The clustering generates semantic visual concepts that guide the training of the deep networks, and the networks in turn guarantee the robustness of the semantic concepts. Experiments on the HMDB51 and the UCF101 datasets demonstrate the superiority of the proposed method, even over several supervised learning methods.
Jingyi Hou, Xinxiao Wu, Jin Chen 0009, Jiebo Luo 0001, Yunde Jia
AAAI2
2018 Exploiting Images for Video Recognition with Hierarchical Generative Adversarial Networks
abstract
Existing deep learning methods of video recognition usually require a large number of labeled videos for training. But for a new task, videos are often unlabeled and it is also time-consuming and labor-intensive to annotate them. Instead of human annotation, we try to make use of existing fully labeled images to help recognize those videos. However, due to the problem of domain shifts and heterogeneous feature representations, the performance of classifiers trained on images may be dramatically degraded for video recognition tasks. In this paper, we propose a novel method, called Hierarchical Generative Adversarial Networks (HiGAN), to enhance recognition in videos (i.e., target domain) by transferring knowledge from images (i.e., source domain). The HiGAN model consists of a \emph{low-level} conditional GAN and a \emph{high-level} conditional GAN. By taking advantage of these two-level adversarial learning, our method is capable of learning a domain-invariant feature representation of source images and target videos. Comprehensive experiments on two challenging video recognition datasets (i.e. UCF101 and HMDB51) demonstrate the effectiveness of the proposed method when compared with the existing state-of-the-art domain adaptation methods.
Feiwu Yu, Xinxiao Wu, Yuchao Sun, Lixin Duan
IJCAI2
2018 Action recognition with motion map 3D network
Yuchao Sun, Xinxiao Wu, Wennan Yu, Feiwu Yu
Neurocomputing2
2018 A discriminative structural model for joint segmentation and recognition of human actions
Cuiwei Liu, Jingyi Hou, Xinxiao Wu, Yunde Jia
Multim. Tools Appl.3
2018 Content-Attention Representation by Factorized Action-Scene Network for Action Recognition
abstract
During action recognition in videos, irrelevant motions in the background can greatly degrade the performance of recognizing specific actions with which we actually concern ourself here. In this paper, a novel deep neural network, called factorized action-scene network (FASNet), is proposed to encode and fuse the most relevant and informative semantic cues for action recognition. Specifically, we decompose the FASNet into two components. One is a newly designed encoding network, named content attention network (CANet), which encodes local spatial-temporal features to learn the action representations with good robustness to the noise of irrelevant motions. The other is a fusion network, which integrates the pretrained CANet to fuse the encoded spatial-temporal features with contextual scene feature extracted from the same video, for learning more descriptive and discriminative action representations. Moreover, different from the existing deep learning based tasks for generic action recognition, which applies softmax loss function as the training guidance, we formulate two loss functions for guiding the proposed model to accomplish more specific action recognition tasks, i.e., the multilabel correlation loss for multilabel action recognition and the triplet loss for complex event detection. Extensive experiments on the Hollywood2 dataset and the TRECVID MEDTest 14 dataset show that our method achieves superior performance compared with the state-of-the-art methods.
Jingyi Hou, Xinxiao Wu, Yuchao Sun, Yunde Jia
IEEE Trans. Multim.2
2018 Extracting Key Segments of Videos for Event Detection by Learning From Web Sources
abstract
In this paper, we present a novel approach of extracting the key segments for event detection in unconstrained videos. The key segments are automatically extracted by transferring the knowledge learned from Web images and Web videos to consumer videos. We propose an adaptive latent structural support vector machine model, where the locations of key segments in videos are regarded as latent variables due to the unavailability of the ground truth of key-segment locations in training data. In order to alleviate the time-consuming and labor-expensive manual annotation of huge amounts of training videos, a large number of loosely labeled Web images as well as videos are collected from the Web sources. Additionally, a limited number of labeled consumer videos are utilized to guarantee the precision of the model. Considering the semantic diversity of key segments, we learn a set of concepts as the semantic description of key segments and explore the temporal information of concepts to capture the sequential relations between the segments. The concepts are automatically discovered by using Web images and videos with their associated tags and description sentences. Comprehensive experiments on the Columbia's consumer video and the TRECVID 2014 Multimedia Event Detection datasets demonstrate that our method outperforms the state-of-the-art methods.
Hao Song 0002, Xinxiao Wu, Wennan Yu, Yunde Jia
IEEE Trans. Multim.2
2017 Heterogeneous Multi-group Adaptation for Event Recognition in Consumer Videos
Mingyu Yao, Xinxiao Wu, Yunde Jia
ICIG (1)2
2017 Heterogeneous domain adaptation method for video annotation
abstract
In this study, the authors study the video annotation problem over heterogeneous domains, in which data from the image source domain and the video target domain is represented by heterogeneous features with different dimensions and physical meanings. A novel feature learning method, called heterogeneous discriminative analysis of canonical correlation (HDCC), is proposed to discover a common feature subspace in which heterogeneous features can be compared. The HDCC utilises discriminative information from the source domain as well as topology information from the target domain to learn two different projection matrices. By using these two matrices, heterogeneous data can be projected onto a common subspace and different features can be compared. They additionally design a group weighting learning framework for multi‐domain adaptation to effectively leverage knowledge learned from the source domain. Under this framework, source domain images are organised in groups according to their semantic meanings, and different weights are assigned to these groups according to their relevancies to the target domain videos. Extensive experiments on the Columbia Consumer Video and Kodak datasets demonstrate the effectiveness of their HDCC and group weighting methods.
Xinxiao Wu, Yunde Jia
IET Comput. Vis.2
2017 Recognizing key segments of videos for video annotation by learning from web image sets
Hao Song 0002, Xinxiao Wu, Wei Liang 0008, Yunde Jia
Multim. Tools Appl.2
2016 Multimedia event detection via deep spatial-temporal neural networks
abstract
This paper proposes a novel method using deep spatial-temporal neural networks based on deep Convolutional Neural Network (CNN) for multimedia event detection. To sufficiently take advantage of the motion and appearance information of events from videos, our networks contain two branches: a temporal neural network and a spatial neural network. The temporal neural network captures motion information by Recurrent Neural Networks with the mutation of gated recurrent unit. The spatial neural network catches object information by using the deep CNN, to encode the CNN features as a bag of semantics with more discriminative representations. Both the temporal and spatial features are beneficial for event detection in a fully coupled way. Finally, we employ the generalized multiple kernel learning method to effectively fuse these two types of heterogeneous and complementary features for action recognition. Experiments on TRECVID MEDTest 14 dataset show that our method achieves better performance than the state of the art.
Jingyi Hou, Xinxiao Wu, Feiwu Yu, Yunde Jia
ICME2
2016 Multi-group-multi-class domain adaptation for event recognition
abstract
In this study, the authors propose a multi‐group–multi‐class domain adaptation framework to recognise events in consumer videos by leveraging a large number of web videos. The authors’ framework is extended from multi‐class support vector machine by adding a novel data‐dependent regulariser, which can force the event classifier to become consistent in consumer videos. To obtain web videos, they search them using several event‐related keywords and refer the videos returned by one keyword search as a group. They also leverage a video representation which is the average of convolutional neural networks features of the video frames for better performance. Comprehensive experiments on the two real‐world consumer video datasets demonstrate the effectiveness of their method for event recognition in consumer videos.
Xinxiao Wu, Yunde Jia
IET Comput. Vis.2
2016 A Hierarchical Video Description for Complex Activity Understanding
Cuiwei Liu, Xinxiao Wu, Yunde Jia
Int. J. Comput. Vis.2
2016 Heterogeneous discriminant analysis for cross-view action recognition
Wanchen Sui, Xinxiao Wu, Yunde Jia
Neurocomputing2
2016 Transfer Latent SVM for Joint Recognition and Localization of Actions in Videos
abstract
In this paper, we develop a novel transfer latent support vector machine for joint recognition and localization of actions by using Web images and weakly annotated training videos. The model takes training videos which are only annotated with action labels as input for alleviating the laborious and time-consuming manual annotations of action locations. Since the ground-truth of action locations in videos are not available, the locations are modeled as latent variables in our method and are inferred during both training and testing phrases. For the purpose of improving the localization accuracy with some prior information of action locations, we collect a number of Web images which are annotated with both action labels and action locations to learn a discriminative model by enforcing the local similarities between videos and Web images. A structural transformation based on randomized clustering forest is used to map the Web images to videos for handling the heterogeneous features of Web images and videos. Experiments on two public action datasets demonstrate the effectiveness of the proposed model for both action localization and action recognition.
Cuiwei Liu, Xinxiao Wu, Yunde Jia
IEEE Trans. Cybern.2
2015 A Multiple Image Group Adaptation Approach for Event Recognition in Consumer Videos
Dengfeng Zhang, Wei Liang 0008, Hao Song 0002, Zhen Dong 0002, Xinxiao Wu
ICIG (1)5
2015 Heterogeneous Discriminant Analysis for Cross-View Action Recognition
Wanchen Sui, Xinxiao Wu, Wei Liang 0008, Yunde Jia
ICONIP (4)2
2015 Cross-domain structural model for video event annotation via web images
Xiabi Liu, Xinxiao Wu, Yunde Jia
Multim. Tools Appl.3
2015 Cross-View Action Recognition Over Heterogeneous Feature Spaces
abstract
In cross-view action recognition, what you saw in one view is different from what you recognize in another view, since the data distribution even the feature space can change from one view to another. In this paper, we address the problem of transferring action models learned in one view (source view) to another different view (target view), where action instances from these two views are represented by heterogeneous features. A novel learning method, called heterogeneous transfer discriminant-analysis of canonical correlations (HTDCC), is proposed to discover a discriminative common feature space for linking source view and target view to transfer knowledge between them. Two projection matrices are learned to, respectively, map data from the source view and the target view into a common feature space via simultaneously minimizing the canonical correlations of interclass training data, maximizing the canonical correlations of intraclass training data, and reducing the data distribution mismatch between the source and target views in the common feature space. In our method, the source view and the target view neither share any common features nor have any corresponding action instances. Moreover, our HTDCC method is capable of handling only a few or even no labeled samples available in the target view, and can also be easily extended to the situation of multiple source views. We additionally propose a weighting learning framework for multiple source views adaptation to effectively leverage action knowledge learned from multiple source views for the recognition task in the target view. Under this framework, different source views are assigned different weights according to their different relevances to the target view. Each weight represents how contributive the corresponding source view is to the target view. Extensive experiments on the IXMAS data set demonstrate the effectiveness of HTDCC on learning the common feature space for heterogeneous cross-view action recognition. In addition, the weighting learning framework can achieve promising results on automatically adapting multiple transferred source-view knowledge to the target view.
Xinxiao Wu, Han Wang 0001, Cuiwei Liu, Yunde Jia
IEEE Trans. Image Process.1
2014 Weakly Supervised Action Recognition and Localization Using Web Images
Cuiwei Liu, Xinxiao Wu, Yunde Jia
ACCV (5)2
2014 Video Annotation by Incremental Learning from Grouped Heterogeneous Sources
Hao Song 0002, Xinxiao Wu, Yunde Jia
ACCV (5)3
2014 Multi-group Adaptation for Event Recognition from Videos
abstract
Recognizing events in consumer videos is becoming increasingly important because of the enormous growth of consumer videos in recent years. Current researches mainly focus on learning from numerous labeled videos, which is time consuming and labor expensive due to labeling the consumer videos. To alleviate the labeling process, we utilize a large number of loosely labeled Web videos (e.g., from YouTube) for visual event recognition in consumer videos. Web videos are noisy and diverse, so brute force transfer of Web videos to consumer videos may hurt the performance. To address such a negative transfer problem, we propose a novel Multi-Group Adaptation (MGA) framework to divide the training Web videos into several semantic groups and seek the optimal weight of each group. Each weight represents how relative the corresponding group is to the consumer domain. The final classifier for event recognition is learned using the weighted combination of classifiers learned from Web videos and enforced to be smooth on the consumer domain. Comprehensive experiments on three real-world consumer video datasets demonstrate the effectiveness of MGA for event recognition in consumer videos.
Xinxiao Wu, Han Wang 0001
ICPR2
2014 Modeling the Relationship of Action, Object, and Scene
abstract
In the task of action recognition, object and scene can provide rich source of contextual information for analyzing human actions, as human actions often occur under particular scene settings with certain related objects. Therefore, we try to utilize the contextual object and scene for improving the performance of action recognition. Specifically, a latent structural SVM is introduced to build the co-occurrence relationship among action, object and scene, in which the object class label and scene class label are treated as latent variables. Using this framework, we can simultaneously predict action class labels, object class labels as well as scene class labels. Moreover, we use a mid-level discriminative feature to separately describe the information of action, object and scene. The feature is actually a set of decision values from the pre-learned classifiers of each class, measuring the likelihood that the input video belongs to the corresponding class. In this paper, we use SVM as action and scene pre-learned classifiers, and use deformable part-based object detector as the object pre-learned classifier, so that object location can be obtained as a by-product. Experimental results on UCF Sports, YouTube and UCF50 datasets demonstrate the effectiveness of the proposed approach.
Xinxiao Wu
ICPR2
2014 Learning a discriminative mid-level feature for action recognition
Cuiwei Liu, Mingtao Pei, Xinxiao Wu, Yu Kong 0001, Yunde Jia
Sci. China Inf. Sci.3
2014 Video Annotation via Image Groups from the Web
abstract
Searching desirable events in uncontrolled videos is a challenging task. Current researches mainly focus on obtaining concepts from numerous labeled videos. But it is time consuming and labor expensive to collect a large amount of required labeled videos for training event models under various circumstances. To alleviate this problem, we propose to leverage abundant Web images for videos since Web images contain a rich source of information with many events roughly annotated and taken under various conditions. However, knowledge from the Web is noisy and diverse, brute force knowledge transfer of images may hurt the video annotation performance. Therefore, we propose a novel Group-based Domain Adaptation (GDA) learning framework to leverage different groups of knowledge (source domain) queried from the Web image search engine to consumer videos (target domain). Different from traditional methods using multiple source domains of images, our method organizes the Web images according to their intrinsic semantic relationships instead of their sources. Specifically, two different types of groups (i.e., event-specific groups and concept-specific groups) are exploited to respectively describe the event-level and concept-level semantic meanings of target-domain videos. Under this framework, we assign different weights to different image groups according to the relevances between the source groups and the target domain, and each group weight represents how contributive the corresponding source image group is to the knowledge transferred to the target video. In order to make the group weights and group classifiers mutually beneficial and reciprocal, a joint optimization algorithm is presented for simultaneously learning the weights and classifiers, using two novel data-dependent regularizers. Experimental results on three challenging video datasets (i.e., CCV, Kodak, and YouTube) demonstrate the effectiveness of leveraging grouped knowledge gained from Web images for video annotation.
Xinxiao Wu, Yunde Jia
IEEE Trans. Multim.2
2013 Cross-View Action Recognition over Heterogeneous Feature Spaces
abstract
In cross-view action recognition, "what you saw" in one view is different from "what you recognize" in another view. The data distribution even the feature space can change from one view to another due to the appearance and motion of actions drastically vary across different views. In this paper, we address the problem of transferring action models learned in one view (source view) to another different view (target view), where action instances from these two views are represented by heterogeneous features. A novel learning method, called Heterogeneous Transfer Discriminantanalysis of Canonical Correlations (HTDCC), is proposed to learn a discriminative common feature space for linking source and target views to transfer knowledge between them. Two projection matrices that respectively map data from source and target views into the common space are optimized via simultaneously minimizing the canonical correlations of inter-class samples and maximizing the intraclass canonical correlations. Our model is neither restricted to corresponding action instances in the two views nor restricted to the same type of feature, and can handle only a few or even no labeled samples available in the target view. To reduce the data distribution mismatch between the source and target views in the common feature space, a nonparametric criterion is included in the objective function. We additionally propose a joint weight learning method to fuse multiple source-view action classifiers for recognition in the target view. Different combination weights are assigned to different source views, with each weight presenting how contributive the corresponding source view is to the target view. The proposed method is evaluated on the IXMAS multi-view dataset and achieves promising results.
Xinxiao Wu, Han Wang 0001, Cuiwei Liu, Yunde Jia
ICCV1
2013 Scene image retrieval via re-ranking semantic and packed dense interestpoints
Wei Liang 0008, Xinxiao Wu, Peng Teng
Neurocomputing3
2013 Action Recognition Using Multilevel Features and Latent Structural SVM
abstract
We first propose a new low-level visual feature, called spatio-temporal context distribution feature of interest points, to describe human actions. Each action video is expressed as a set of relative XYT coordinates between pairwise interest points in a local region. We learn a global Gaussian mixture model (GMM) (referred to as a universal background model) using the relative coordinate features from all the training videos, and then we represent each video as the normalized parameters of a video-specific GMM adapted from the global GMM. In order to capture the spatio-temporal relationships at different levels, multiple GMMs are utilized to describe the context distributions of interest points over multiscale local regions. Motivated by the observation that some actions share similar motion patterns, we additionally propose a novel mid-level class correlation feature to capture the semantic correlations between different action classes. Each input action video is represented by a set of decision values obtained from the pre-learned classifiers of all the action classes, with each decision value measuring the likelihood that the input video belongs to the corresponding action class. Moreover, human actions are often associated with some specific natural environments and also exhibit high correlation with particular scene classes. It is therefore beneficial to utilize the contextual scene information for action recognition. In this paper, we build the high-level co-occurrence relationship between action classes and scene classes to discover the mutual contextual constraints between action and scene. By treating the scene class label as a latent variable, we propose to use the latent structural SVM (LSSVM) model to jointly capture the compatibility between multilevel action features (e.g., low-level visual context distribution feature and the corresponding mid-level class correlation feature) and action classes, the compatibility between multilevel scene features (i.e., SIFT feature and the corresponding class correlation feature) and scene classes, and the contextual relationship between action classes and scene classes. Extensive experiments on UCF Sports, YouTube and UCF50 datasets demonstrate the effectiveness of the proposed multilevel features and action-scene interaction based LSSVM model for human action recognition. Moreover, our method generally achieves higher recognition accuracy than other state-of-the-art methods on these datasets.
Xinxiao Wu, Dong Xu 0001, Lixin Duan, Jiebo Luo 0001, Yunde Jia
IEEE Trans. Circuits Syst. Video Technol.1
2012 View-Invariant Action Recognition Using Latent Kernelized Structural SVM
Xinxiao Wu, Yunde Jia
ECCV (5)1
2012 Action recognition with discriminative mid-level features
Cuiwei Liu, Yu Kong 0001, Xinxiao Wu, Yunde Jia
ICPR3
2012 Annotating videos from the web images
Xinxiao Wu, Yunde Jia
ICPR2
2011 Action recognition using context and appearance distribution features
abstract
We first propose a new spatio-temporal context distribution feature of interest points for human action recognition. Each action video is expressed as a set of relative XYT coordinates between pairwise interest points in a local region. We learn a global GMM (referred to as Universal Background Model, UBM) using the relative coordinate features from all the training videos, and then represent each video as the normalized parameters of a video-specific GMM adapted from the global GMM. In order to capture the spatio-temporal relationships at different levels, multiple GMMs are utilized to describe the context distributions of interest points over multi-scale local regions. To describe the appearance information of an action video, we also propose to use GMM to characterize the distribution of local appearance features from the cuboids centered around the interest points. Accordingly, an action video can be represented by two types of distribution features: 1) multiple GMM distributions of spatio-temporal context; 2) GMM distribution of local video appearance. To effectively fuse these two types of heterogeneous and complementary distribution features, we additionally propose a new learning algorithm, called Multiple Kernel Learning with Augmented Features (AFMKL), to learn an adapted classifier based on multiple kernels and the pre-learned classifiers of other action classes. Extensive experiments on KTH, multi-view IXMAS and complex UCF sports datasets demonstrate that our method generally achieves higher recognition accuracy than other state-of-the-art methods.
Xinxiao Wu, Dong Xu 0001, Lixin Duan, Jiebo Luo 0001
CVPR1
2010 Discriminative human action recognition in the learned hierarchical manifold space
Xinxiao Wu, Wei Liang 0008, Guangming Hou, Yunde Jia
Image Vis. Comput.2
2010 Incremental discriminant-analysis of canonical correlations for action recognition
Xinxiao Wu, Yunde Jia, Wei Liang 0008
Pattern Recognit.1
2009 Incremental discriminative-analysis of canonical correlations for action recognition
abstract
Human action recognition is a challenging problem due to the large changes of human appearance in the cases of partial occlusions, non-rigid deformations and high irregularities. It is difficult to collect a large set of training samples with the hope of covering all possible variations of an action. In this paper, we propose an online recognition method, namely Incremental Discriminant-Analysis of Canonical Correlations (IDCC), whose discriminative model is incrementally updated to capture the changes of human appearance and thereby facilitates the recognition task in changing environments. As the training sets are acquired sequentially instead of being given completely in advance, our method is able to compute a new discriminant matrix by updating the existing one using the eigenspace merging algorithm. Experimental results on both Weizmann and KTH action data sets show that our method performs better than state-of-the-art methods on both accuracy and efficiency. Moreover, the robustness of our method is demonstrated on the irregular action recognition.
Xinxiao Wu, Wei Liang 0008, Yunde Jia
ICCV1
2009 Tracking articulated objects by learning intrinsic structure of motion
Xinxiao Wu, Wei Liang 0008, Yunde Jia
Pattern Recognit. Lett.1
2009 Action recognition feedback-based framework for human pose reconstruction from monocular images
Xinxiao Wu, Wei Liang 0008, Yunde Jia
Pattern Recognit. Lett.1
2008 Human action recognition using discriminative models in the learned hierarchical manifold space
abstract
A hierarchical learning based approach for human action recognition is proposed in this paper. It consists of hierarchical nonlinear dimensionality reduction based feature extraction and cascade discriminative model based action modeling. Human actions are inferred from human body joint motions and human bodies are decomposed into several physiological body parts according to inherent hierarchy (e.g. right arm, left arm and head all belong to upper body). We explore the underlying hierarchical structures of high-dimensional human pose space using hierarchical Gaussian process latent variable model (HGPLVM) and learn a representative motion pattern set for each body part. In the hierarchical manifold space, the bottom-up cascade conditional random fields (CRFs) are used to predict the corresponding motion pattern in each manifold subspace, and then the final action label is estimated for each observation by a discriminative classifier on the current motion pattern set.
Wei Liang 0008, Xinxiao Wu, Yunde Jia
FG3