Xiaodong Yang 0001

dblp:19/1551-1 · DBLP profile ↗
← Back
46ranked-venue papers
12as first author
4since 2021 · last 2024
0009-0003-4638-8039ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 9 first-author · 1 since 2021Artificial intelligence and machine learning · 29 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
25 papers
Video understanding and tracking · 22% Face, body and person analysis · 17% Generative modeling · 17%
Computer graphics and multimedia
4 papers
Image and video processing · 71% Audio and music processing · 29%
Human-computer interaction and pervasive computing
4 papers
Interaction techniques and input · 72% Accessibility and assistive technology · 28%

Topics — the 30 heaviest of 63, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
person re-identification
1.232020
Joint Disentangling and Adaptation for Cross-Domain Person Re-Identification · ECCV (2) 2020
PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data · ICCV 2019
Joint Discriminative and Generative Learning for Person Re-Identification · CVPR 2019
Machine learning › Generative modeling
synthetic data generation
1.222024
Attribute Descent: Simulating Object-Centric Datasets on the Content Level and Beyond · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Simulating Content Consistent Vehicle Datasets with Attribute Descent · ECCV (6) 2020
Computer vision › Image recognition and object detection
object detection
1.032020
UFO2: A Unified Framework Towards Omni-supervised Object Detection · ECCV (19) 2020
Instance-Aware, Context-Focused, and Memory-Efficient Weakly Supervised Object Detection · CVPR 2020
CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification · CVPR 2019
Computer vision › 3D vision › motion estimation › optical flow
CNN-based optical flow
0.822020
Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2020
PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume · CVPR 2018
Computer vision › 3D vision › motion estimation
optical flow
0.822020
Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2020
PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume · CVPR 2018
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.812024
Attribute Descent: Simulating Object-Centric Datasets on the Content Level and Beyond · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › Face, body and person analysis › person re-identification
vehicle re-identification
0.822019
PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data · ICCV 2019
CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification · CVPR 2019
Computer vision › Video understanding and tracking
action recognition
0.742018
Super Normal Vector for Human Activity Recognition with Depth Cameras · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Action Recognition Using Super Sparse Coding Vector with Spatio-temporal Awareness · ECCV (2) 2014
Recognizing actions using depth motion maps-based histograms of oriented gradients · ACM Multimedia 2012
Computer vision › Image recognition and object detection
image classification
0.712023
Partial Convolution for Padding, Inpainting, and Image Synthesis · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Image and video processing › image restoration
image inpainting
0.712023
Partial Convolution for Padding, Inpainting, and Image Synthesis · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Video understanding and tracking › action recognition › 3d action recognition
depth-based action recognition
0.632017
Super Normal Vector for Human Activity Recognition with Depth Cameras · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Super Normal Vector for Activity Recognition Using Depth Sequences · CVPR 2014
Recognizing actions using depth motion maps-based histograms of oriented gradients · ACM Multimedia 2012
Machine learning › Generative modeling › diffusion model
3d shape generation
0.412020
Simulating Content Consistent Vehicle Datasets with Attribute Descent · ECCV (6) 2020
Computer vision › Face, body and person analysis › person re-identification
cross-domain person re-identification
0.412020
Joint Disentangling and Adaptation for Cross-Domain Person Re-Identification · ECCV (2) 2020
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.412020
Joint Disentangling and Adaptation for Cross-Domain Person Re-Identification · ECCV (2) 2020
Machine learning › Learning paradigms › semi-supervised learning
omni-supervised learning
0.412020
UFO2: A Unified Framework Towards Omni-supervised Object Detection · ECCV (19) 2020
Computer vision › Vision and language › visual grounding
phrase grounding
0.412020
Contrastive Learning for Weakly Supervised Phrase Grounding · ECCV (3) 2020
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection
0.412020
Instance-Aware, Context-Focused, and Memory-Efficient Weakly Supervised Object Detection · CVPR 2020
Computer vision › Vision and language › visual grounding › phrase grounding
weakly supervised phrase grounding
0.412020
Contrastive Learning for Weakly Supervised Phrase Grounding · ECCV (3) 2020
Computer vision › Face, body and person analysis
face alignment
0.422018
Dynamic Facial Analysis: From Bayesian Filtering to Recurrent Neural Network · CVPR 2017
Making Convolutional Networks Recurrent for Visual Sequence Learning · CVPR 2018
Machine learning › Generative modeling
image generation
0.412019
Joint Discriminative and Generative Learning for Person Re-Identification · CVPR 2019
Machine learning › Generative modeling
multimodal generation
0.412019
Dancing to Music · NeurIPS 2019
Computer vision › Video understanding and tracking › multi-camera tracking
multi-target multi-camera tracking
0.412019
CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification · CVPR 2019
Machine learning › Generative modeling › motion generation
music-to-dance generation
0.412019
Dancing to Music · NeurIPS 2019
Computer vision › Video understanding and tracking › action detection
spatio-temporal action localization
0.412019
STEP: Spatio-Temporal Progressive Learning for Video Action Detection · CVPR 2019
Audio and music processing
music analysis
0.412019
Dancing to Music · NeurIPS 2019
Machine learning › Generative modeling
generative adversarial network
0.312018
MoCoGAN: Decomposing Motion and Content for Video Generation · CVPR 2018
Computer vision › Video understanding and tracking
motion analysis
0.312018
MoCoGAN: Decomposing Motion and Content for Video Generation · CVPR 2018
Computer vision › Video understanding and tracking › motion analysis › motion modeling
motion and content decomposition
0.312018
MoCoGAN: Decomposing Motion and Content for Video Generation · CVPR 2018
Computer vision › Video understanding and tracking › video representation learning
motion-view disentanglement
0.312018
MoCoGAN: Decomposing Motion and Content for Video Generation · CVPR 2018
Machine learning › Deep learning architectures and training
recurrent neural network
0.312018
Making Convolutional Networks Recurrent for Visual Sequence Learning · CVPR 2018

Methods — techniques the papers use, named apart from their topics

renormalization · 1.3binary masking · 1.3recurrent neural network · 1.0warping · 0.8cost volume · 0.8graphic engine simulation · 0.8attribute optimization · 0.8convolutional neural network · 0.6spatio-temporal pyramid · 0.5fisher kernel · 0.5synthesis-by-analysis · 0.4pose generation · 0.4benchmark dataset construction · 0.4wavelet subband analysis · 0.2structural feature · 0.2connectionist temporal classification · 0.2confidence margin feature combination · 0.23d convolutional neural network · 0.2
YearPublicationVenuePosition
2024 Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds
abstract
3D perception in LiDAR point clouds is crucial for a self-driving vehicle to properly act in 3D environment. However, manually labeling point clouds is hard and costly. There has been a growing interest in self-supervised pre-training of 3D perception models. Following the success of contrastive learning in images, current methods mostly conduct contrastive pre-training on point clouds only. Yet an autonomous driving vehicle is typically supplied with multiple sensors including cameras and LiDAR. In this context, we systematically study single modality, cross-modality, and multi-modality for contrastive learning of point clouds, and show that cross-modality wins over other alternatives. In addition, considering the huge difference between the training sources in 2D images and 3D point clouds, it remains unclear how to design more effective contrastive units for LiDAR. We therefore propose the instance-aware and similarity-balanced contrastive units that are tailored for self-driving point clouds. Extensive experiments reveal that our approach achieves remarkable performance gains over various point cloud models across the downstream perception tasks of LiDAR based 3D object detection and 3D semantic segmentation on the four popular benchmarks including Waymo Open Dataset, nuScenes, SemanticKITTI and ONCE.
Mu Cai, Chenxu Luo, Yong Jae Lee, Xiaodong Yang 0001
IROS4
2024 Attribute Descent: Simulating Object-Centric Datasets on the Content Level and Beyond
abstract
This article aims to use graphic engines to simulate a large number of training data that have free annotations and possibly strongly resemble to real-world data. Between synthetic and real, a two-level domain gap exists, involving content level and appearance level. While the latter is concerned with appearance style, the former problem arises from a different mechanism, i.e., content mismatch in attributes such as camera viewpoint, object placement and lighting conditions. In contrast to the widely-studied appearance-level gap, the content-level discrepancy has not been broadly studied. To address the content-level misalignment, we propose an attribute descent approach that automatically optimizes engine attributes to enable synthetic data to approximate real-world data. We verify our method on object-centric tasks, wherein an object takes up a major portion of an image. In these tasks, the search space is relatively small, and the optimization of each attribute yields sufficiently obvious supervision signals. We collect a new synthetic asset VehicleX, and reformat and reuse existing the synthetic assets ObjectX and PersonX. Extensive experiments on image classification and object re-identification confirm that adapted synthetic data can be effectively used in three scenarios: training with synthetic data only, training data augmentation and numerically understanding dataset content.
Yue Yao 0001, Liang Zheng 0001, Xiaodong Yang 0001, Milind Napthade, Tom Gedeon
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Partial Convolution for Padding, Inpainting, and Image Synthesis
abstract
Partial convolution weights convolutions with binary masks and renormalizes on valid pixels. It was originally proposed for image inpainting task because a corrupted image processed by a standard convolutional often leads to artifacts. Therefore, binary masks are constructed that define the valid and corrupted pixels, so that partial convolution results are only calculated based on valid pixels. It has been also used for conditional image synthesis task, so that when a scene is generated, convolution results of an instance depend only on the feature values that belong to the same instance. One of the unexplored applications for partial convolution is padding which is a critical component of modern convolutional networks. Common padding schemes make strong assumptions about how the padded data should be extrapolated. We show that these padding schemes impair model accuracy, whereas partial convolution based padding provides consistent improvements across a range of tasks. In this article, we review partial convolution applications under one framework. We conduct a comprehensive study of the partial convolution based padding on a variety of computer vision tasks, including image classification, 3D-convolution-based action recognition, and semantic segmentation. Our results suggest that partial convolution-based padding shows promising improvements over strong baselines.
Guilin Liu, Aysegul Dundar, Kevin J. Shih, Ting-Chun Wang, Fitsum A. Reda, Karan Sapra, Zhiding Yu, Xiaodong Yang 0001, Andrew Tao, Bryan Catanzaro
IEEE Trans. Pattern Anal. Mach. Intell.8
2021 Hierarchical Contrastive Motion Learning for Video Action Recognition
Xitong Yang, Xiaodong Yang 0001, Sifei Liu, Deqing Sun, Larry Davis 0001, Jan Kautz
BMVC2
2020 Instance-Aware, Context-Focused, and Memory-Efficient Weakly Supervised Object Detection
abstract
Weakly supervised learning has emerged as a compelling tool for object detection by reducing the need for strong supervision during training. However, major challenges remain: (1) differentiation of object instances can be ambiguous; (2) detectors tend to focus on discriminative parts rather than entire objects; (3) without ground truth, object proposals have to be redundant for high recalls, causing significant memory consumption. Addressing these challenges is difficult, as it often requires to eliminate uncertainties and trivial solutions. To target these issues we develop an instance-aware and context-focused unified framework. It employs an instance-aware self-training algorithm and a learnable Concrete DropBlock while devising a memory-efficient sequential batch back-propagation. Our proposed method achieves state-of-the-art results on COCO (12.1% AP, 24.8% AP50), VOC 2007 (54.9% AP), and VOC 2012 (52.1% AP), improving baselines by great margins. In addition, the proposed method is the first to benchmark ResNet based models and weakly supervised video object detection. Refer to our project page for code, models, and more details: https://github.com/NVlabs/wetectron.
Zhongzheng Ren, Zhiding Yu, Xiaodong Yang 0001, Ming-Yu Liu 0001, Yong Jae Lee, Alexander G. Schwing, Jan Kautz
CVPR3
2020 Contrastive Learning for Weakly Supervised Phrase Grounding
Tanmay Gupta, Arash Vahdat, Gal Chechik, Xiaodong Yang 0001, Jan Kautz, Derek Hoiem
ECCV (3)4
2020 UFO2: A Unified Framework Towards Omni-supervised Object Detection
Zhongzheng Ren, Zhiding Yu, Xiaodong Yang 0001, Ming-Yu Liu 0001, Alexander G. Schwing, Jan Kautz
ECCV (19)3
2020 Simulating Content Consistent Vehicle Datasets with Attribute Descent
Yue Yao 0001, Liang Zheng 0001, Xiaodong Yang 0001, Milind Naphade, Tom Gedeon
ECCV (6)3
2020 Joint Disentangling and Adaptation for Cross-Domain Person Re-Identification
Yang Zou 0003, Xiaodong Yang 0001, Zhiding Yu, B. V. K. Vijaya Kumar, Jan Kautz
ECCV (2)2
2020 Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation
abstract
We investigate two crucial and closely-related aspects of CNNs for optical flow estimation: models and training. First, we design a compact but effective CNN model, called PWC-Net, according to simple and well-established principles: pyramidal processing, warping, and cost volume processing. PWC-Net is 17 times smaller in size, 2 times faster in inference, and 11 percent more accurate on Sintel final than the recent FlowNet2 model. It is the winning entry in the optical flow competition of the robust vision challenge. Next, we experimentally analyze the sources of our performance gains. In particular, we use the same training procedure for PWC-Net to retrain FlowNetC, a sub-network of FlowNet2. The retrained FlowNetC is 56 percent more accurate on Sintel final than the previously trained one and even 5 percent more accurate than the FlowNet2 model. We further improve the training procedure and increase the accuracy of PWC-Net on Sintel by 10 percent and on KITTI 2012 and 2015 by 20 percent. Our newly trained model parameters and training protocols are available on https://github.com/NVlabs/PWC-Net.
Deqing Sun, Xiaodong Yang 0001, Ming-Yu Liu 0001, Jan Kautz
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification
abstract
Urban traffic optimization using traffic cameras as sensors is driving the need to advance state-of-the-art multi-target multi-camera (MTMC) tracking. This work introduces CityFlow, a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 10 intersections, with the longest distance between two simultaneous cameras being 2.5 km. To the best of our knowledge, CityFlow is the largest-scale dataset in terms of spatial coverage and the number of cameras/videos in an urban environment. The dataset contains more than 200K annotated bounding boxes covering a wide range of scenes, viewing angles, vehicle models, and urban traffic flow conditions. Camera geometry and calibration information are provided to aid spatio-temporal analysis. In addition, a subset of the benchmark is made available for the task of image-based vehicle re-identification (ReID). We conducted an extensive experimental evaluation of baselines/state-of-the-art approaches in MTMC tracking, multi-target single-camera (MTSC) tracking, object detection, and image-based ReID on this dataset, analyzing the impact of different network architectures, loss functions, spatio-temporal models and their combinations on task effectiveness. An evaluation server is launched with the release of our benchmark at the 2019 AI City Challenge (https://www.aicitychallenge.org/) that allows researchers to compare the performance of their newest techniques. We expect this dataset to catalyze research in this field, propel the state-of-the-art forward, and lead to deployed traffic optimization(s) in the real world.
Milind Naphade, Ming-Yu Liu 0001, Xiaodong Yang 0001, Stanley T. Birchfield, Ratnesh Kumar 0004, David C. Anastasiu, Jenq-Neng Hwang
CVPR4
2019 STEP: Spatio-Temporal Progressive Learning for Video Action Detection
abstract
In this paper, we propose Spatio-TEmporal Progressive (STEP) action detector-a progressive learning framework for spatio-temporal action detection in videos. Starting from a handful of coarse-scale proposal cuboids, our approach progressively refines the proposals towards actions over a few steps. In this way, high-quality proposals (i.e., adhere to action movements) can be gradually obtained at later steps by leveraging the regression outputs from previous steps. At each step, we adaptively extend the proposals in time to incorporate more related temporal context. Compared to the prior work that performs action detection in one run, our progressive learning framework is able to naturally handle the spatial displacement within action tubes and therefore provides a more effective way for spatio-temporal modeling. We extensively evaluate our approach on UCF101 and AVA, and demonstrate superior detection results. Remarkably, we achieve mAP of 75.0% and 18.6% on the two datasets with 3 progressive steps and using respectively only 11 and 34 initial proposals.
Xitong Yang, Xiaodong Yang 0001, Ming-Yu Liu 0001, Fanyi Xiao, Larry Davis 0001, Jan Kautz
CVPR2
2019 Joint Discriminative and Generative Learning for Person Re-Identification
abstract
Person re-identification (re-id) remains challenging due to significant intra-class variations across different cameras. Recently, there has been a growing interest in using generative models to augment training data and enhance the invariance to input changes. The generative pipelines in existing methods, however, stay relatively separate from the discriminative re-id learning stages. Accordingly, re-id models are often trained in a straightforward manner on the generated data. In this paper, we seek to improve learned re-id embeddings by better leveraging the generated data. To this end, we propose a joint learning framework that couples re-id learning and data generation end-to-end. Our model involves a generative module that separately encodes each person into an appearance code and a structure code, and a discriminative module that shares the appearance encoder with the generative module. By switching the appearance or structure codes, the generative module is able to generate high-quality cross-id composed images, which are online fed back to the appearance encoder and used to improve the discriminative module. The proposed joint learning framework renders significant improvement over the baseline without using generated data, leading to the state-of-the-art performance on several benchmark datasets.
Zhedong Zheng, Xiaodong Yang 0001, Zhiding Yu, Liang Zheng 0001, Yi Yang 0001, Jan Kautz
CVPR2
2019 PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data
abstract
In comparison with person re-identification (ReID), which has been widely studied in the research community, vehicle ReID has received less attention. Vehicle ReID is challenging due to 1) high intra-class variability (caused by the dependency of shape and appearance on viewpoint), and 2) small inter-class variability (caused by the similarity in shape and appearance between vehicles produced by different manufacturers). To address these challenges, we propose a Pose-Aware Multi-Task Re-Identification (PAMTRI) framework. This approach includes two innovations compared with previous methods. First, it overcomes viewpoint-dependency by explicitly reasoning about vehicle pose and shape via keypoints, heatmaps and segments from pose estimation. Second, it jointly classifies semantic vehicle attributes (colors and types) while performing ReID, through multi-task learning with the embedded pose representations. Since manually labeling images with detailed pose and attribute information is prohibitive, we create a large-scale highly randomized synthetic dataset with automatically annotated vehicle attributes for training. Extensive experiments validate the effectiveness of each proposed component, showing that PAMTRI achieves significant improvement over state-of-the-art on two mainstream vehicle ReID benchmarks: VeRi and CityFlow-ReID.
Milind Naphade, Stanley T. Birchfield, Jonathan Tremblay, William Hodge, Ratnesh Kumar 0004, Xiaodong Yang 0001
ICCV8
2019 Dancing to Music
abstract
Dancing to music is an instinctive move by humans. Learning to model the music-to-dance generation process is, however, a challenging problem. It requires significant efforts to measure the correlation between music and dance as one needs to simultaneously consider multiple aspects, such as style and beat of both music and dance. Additionally, dance is inherently multimodal and various following movements of a pose at any moment are equally likely. In this paper, we propose a synthesis-by-analysis learning framework to generate dance from music. In the top-down analysis phase, we decompose a dance into a series of basic dance units, through which the model learns how to move. In the bottom-up synthesis phase, the model learns how to compose a dance by combining multiple basic dancing movements seamlessly according to input music. Experimental qualitative and quantitative results demonstrate that the proposed method can synthesize realistic, diverse, style-consistent, and beat-matching dances from music.
Hsin-Ying Lee 0001, Xiaodong Yang 0001, Ming-Yu Liu 0001, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang 0001, Jan Kautz
NeurIPS2
2019 Discovering spatio-temporal action tubes
Yuancheng Ye, Xiaodong Yang 0001, Yingli Tian
J. Vis. Commun. Image Represent.2
2018 Budget-Aware Activity Detection with A Recurrent Policy Network
Behrooz Mahasseni, Xiaodong Yang 0001, Pavlo Molchanov 0001, Jan Kautz
BMVC2
2018 PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume
abstract
We present a compact but effective CNN model for optical flow, called PWC-Net. PWC-Net has been designed according to simple and well-established principles: pyramidal processing, warping, and the use of a cost volume. Cast in a learnable feature pyramid, PWC-Net uses the current optical flow estimate to warp the CNN features of the second image. It then uses the warped features and features of the first image to construct a cost volume, which is processed by a CNN to estimate the optical flow. PWC-Net is 17 times smaller in size and easier to train than the recent FlowNet2 model. Moreover, it outperforms all published optical flow methods on the MPI Sintel final pass and KITTI 2015 benchmarks, running at about 35 fps on Sintel resolution (1024 Ã × 436) images. Our models are available on our project website.
Deqing Sun, Xiaodong Yang 0001, Ming-Yu Liu 0001, Jan Kautz
CVPR2
2018 MoCoGAN: Decomposing Motion and Content for Video Generation
abstract
Visual signals in a video can be divided into content and motion. While content specifies which objects are in the video, motion describes their dynamics. Based on this prior, we propose the Motion and Content decomposed Generative Adversarial Network (MoCoGAN) framework for video generation. The proposed framework generates a video by mapping a sequence of random vectors to a sequence of video frames. Each random vector consists of a content part and a motion part. While the content part is kept fixed, the motion part is realized as a stochastic process. To learn motion and content decomposition in an unsupervised manner, we introduce a novel adversarial learning scheme utilizing both image and video discriminators. Extensive experimental results on several challenging datasets with qualitative and quantitative comparison to the state-of-the-art approaches, verify effectiveness of the proposed framework. In addition, we show that MoCoGAN allows one to generate videos with same content but different motion as well as videos with different content and same motion. Our code is available at https://github.com/sergeytulyakov/mocogan.
Sergey Tulyakov, Ming-Yu Liu 0001, Xiaodong Yang 0001, Jan Kautz
CVPR3
2018 Making Convolutional Networks Recurrent for Visual Sequence Learning
abstract
Recurrent neural networks (RNNs) have emerged as a powerful model for a broad range of machine learning problems that involve sequential data. While an abundance of work exists to understand and improve RNNs in the context of language and audio signals such as language modeling and speech recognition, relatively little attention has been paid to analyze or modify RNNs for visual sequences, which by nature have distinct properties. In this paper, we aim to bridge this gap and present the first large-scale exploration of RNNs for visual sequence learning. In particular, with the intention of leveraging the strong generalization capacity of pre-trained convolutional neural networks (CNNs), we propose a novel and effective approach, PreRNN, to make pre-trained CNNs recurrent by transforming convolutional layers or fully connected layers into recurrent layers. We conduct extensive evaluations on three representative visual sequence learning tasks: sequential face alignment, dynamic hand gesture recognition, and action recognition. Our experiments reveal that PreRNN consistently outperforms the traditional RNNs and achieves state-of-the-art results on the three applications, suggesting that PreRNN is more suitable for visual sequence learning.
Xiaodong Yang 0001, Pavlo Molchanov 0001, Jan Kautz
CVPR1
2018 Video you only look once: Overall temporal convolutions for action recognition
Longlong Jing, Xiaodong Yang 0001, Yingli Tian
J. Vis. Commun. Image Represent.2
2017 Dynamic Facial Analysis: From Bayesian Filtering to Recurrent Neural Network
abstract
Facial analysis in videos, including head pose estimation and facial landmark localization, is key for many applications such as facial animation capture, human activity recognition, and human-computer interaction. In this paper, we propose to use a recurrent neural network (RNN) for joint estimation and tracking of facial features in videos. We are inspired by the fact that the computation performed in an RNN bears resemblance to Bayesian filters, which have been used for tracking in many previous methods for facial analysis from videos. Bayesian filters used in these methods, however, require complicated, problem-specific design and tuning. In contrast, our proposed RNN-based method avoids such tracker-engineering by learning from training data, similar to how a convolutional neural network (CNN) avoids feature-engineering for image classification. As an end-to-end network, the proposed RNN-based method provides a generic and holistic solution for joint estimation and tracking of various types of facial features from consecutive video frames. Extensive experimental results on head pose estimation and facial landmark localization from videos demonstrate that the proposed RNN-based method outperforms frame-wise models and Bayesian filtering. In addition, we create a large-scale synthetic dataset for head pose estimation, with which we achieve state-of-the-art performance on a benchmark dataset.
Jinwei Gu, Xiaodong Yang 0001, Shalini De Mello, Jan Kautz
CVPR2
2017 3D convolutional neural network with multi-model framework for action recognition
abstract
In this paper, we propose an efficient and effective action recognition framework by combining multiple feature models from dynamic image, optical flow and raw frame, with 3D convolutional neural network (CNN). Dynamic image preserves the long-term temporal information, while optical flow captures short-term temporal information, and raw frame represents the appearance information. Experiments demonstrate that dynamic image provides complementary information to raw frame feature and optical flow feature. Furthermore, with the approximate rank pooling, the computation of dynamic images is about 360 times faster than optical flow, and the dynamic image requires far less memory than optical flow and raw frame.
Longlong Jing, Yuancheng Ye, Xiaodong Yang 0001, Yingli Tian
ICIP3
2017 Super Normal Vector for Human Activity Recognition with Depth Cameras
abstract
The advent of cost-effectiveness and easy-operation depth cameras has facilitated a variety of visual recognition tasks including human activity recognition. This paper presents a novel framework for recognizing human activities from video sequences captured by depth cameras. We extend the surface normal to polynormal by assembling local neighboring hypersurface normals from a depth sequence to jointly characterize local motion and shape information. We then propose a general scheme of super normal vector (SNV) to aggregate the low-level polynormals into a discriminative representation, which can be viewed as a simplified version of the Fisher kernel representation. In order to globally capture the spatial layout and temporal order, an adaptive spatio-temporal pyramid is introduced to subdivide a depth video into a set of space-time cells. In the extensive experiments, the proposed approach achieves superior performance to the state-of-the-art methods on the four public benchmark datasets, i.e., MSRAction3D, MSRDailyActivity3D, MSRGesture3D, and MSRActionPairs3D.
Xiaodong Yang 0001, Yingli Tian
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Evaluation of Low-Level Features for Real-World Surveillance Event Detection
abstract
Event detection targets at recognizing and localizing specified spatio-temporal patterns in videos. Most research of human activity recognition in the past decades experimented on relatively clean scenes with limited actors performing explicit actions. Recently, more efforts have been paid to the real-world surveillance videos in which the human activity recognition is more challenging due to large variations caused by factors, such as scaling, resolution, viewpoint, cluttered background, and crowdedness. In this paper, we systematically evaluate seven different types of low-level spatio-temporal features in the context of surveillance event detection (SED) using a uniform experimental setup. Fisher vector is employed to aggregate low-level features as the representation of each video clip. A set of random forests is then learned as the classification models. To bridge the research efforts and real-world applications, we utilize the NIST TRECVID SED as our testbed in which seven events are predefined involving different levels of human activity analysis. Strengths and limitations for each low-level feature type are analyzed and discussed.
Yang Xian, Xuejian Rong, Xiaodong Yang 0001, Yingli Tian
IEEE Trans. Circuits Syst. Video Technol.3
2016 Online Detection and Classification of Dynamic Hand Gestures with Recurrent 3D Convolutional Neural Networks
abstract
Automatic detection and classification of dynamic hand gestures in real-world systems intended for human computer interaction is challenging as: 1) there is a large diversity in how people perform gestures, making detection and classification difficult, 2) the system must work online in order to avoid noticeable lag between performing a gesture and its classification, in fact, a negative lag (classification before the gesture is finished) is desirable, as feedback to the user can then be truly instantaneous. In this paper, we address these challenges with a recurrent three-dimensional convolutional neural network that performs simultaneous detection and classification of dynamic hand gestures from multi-modal data. We employ connectionist temporal classification to train the network to predict class labels from inprogress gestures in unsegmented input streams. In order to validate our method, we introduce a new challenging multimodal dynamic hand gesture dataset captured with depth, color and stereo-IR sensors. On this challenging dataset, our gesture recognition system achieves an accuracy of 83:8%, outperforms competing state-of-the-art algorithms, and approaches human accuracy of 88:4%. Moreover, our method achieves state-of-the-art performance on SKIG and ChaLearn2014 benchmarks.
Pavlo Molchanov 0001, Xiaodong Yang 0001, Shalini Gupta, Stephen Tyree, Jan Kautz
CVPR2
2016 Towards selecting robust hand gestures for automotive interfaces
abstract
Driver distraction is a serious threat to automotive safety. The visual-manual interfaces in cars are a source of distraction for drivers. Automotive touch-less hand gesture-based user interfaces can help to reduce driver distraction and enhance safety and comfort. The choice of hand gestures in automotive interfaces is central to their success and widespread adoption. In this work we evaluate the recognition accuracy of 25 different gestures for state-of-the-art computer vision-based gesture recognition algorithms and for human observers. We show that some gestures are consistently recognized more accurately than others by both vision-based algorithms and humans. We further identify similarities in the hand gesture recognition abilities of vision-based systems and humans. Lastly, by merging pairs of gestures with high miss-classification rates, we propose ten robust hand gestures for automotive interfaces, which are classified with high and equal accuracy by vision-based algorithms.
Shalini Gupta, Pavlo Molchanov 0001, Xiaodong Yang 0001, Stephen Tyree, Jan Kautz
Intelligent Vehicles Symposium3
2016 Region Trajectories for Video Semantic Concept Detection
abstract
Recently, with the advent of the convolutional neural network (CNN), many CNN-based object detection algorithms have been proposed and achieved encouraging results. In this paper, we introduce an algorithm based on region trajectories to establish the connections between object localizations in individual frames and video sequences. To detect object regions in the individual frames of a video, we enhance the region-based convolutional neural network (R-CNN), by incorporating EdgeBox with the Selective Search to generate candidate region proposals and combining the GoogLeNet with the AlexNet to improve the discriminability of the feature representations. The DeepMatching algorithm is employed in our proposed region trajectory method to track the points in the detected object regions. The experiments are conducted on the validation split of the TRECVID 2015 Localization dataset. As demonstrated by the experimental results, our proposed approach improves the object detection accuracy in both temporal and spatial measurements.
Yuancheng Ye, Xuejian Rong, Xiaodong Yang 0001, Yingli Tian
ICMR3
2016 Multilayer and Multimodal Fusion of Deep Neural Networks for Video Classification
abstract
This paper presents a novel framework to combine multiple layers and modalities of deep neural networks for video classification. We first propose a multilayer strategy to simultaneously capture a variety of levels of abstraction and invariance in a network, where the convolutional and fully connected layers are effectively represented by our proposed feature aggregation methods. We further introduce a multimodal scheme that includes four highly complementary modalities to extract diverse static and dynamic cues at multiple temporal scales. In particular, for modeling the long-term temporal information, we propose a new structure, FC-RNN, to effectively transform pre-trained fully connected layers into recurrent layers. A robust boosting model is then introduced to optimize the fusion of multiple layers and modalities in a unified way. In the extensive experiments, we achieve state-of-the-art results on two public benchmark datasets: UCF101 and HMDB51.
Xiaodong Yang 0001, Pavlo Molchanov 0001, Jan Kautz
ACM Multimedia1
2015 Exploring Pooling Strategies based on Idiosyncrasies of Spatio-Temporal Interest Points
abstract
Recent studies have demonstrated that the implementation of local space-time interest points has good competence and robustness in the area of human action recognition, which has become one of the challenging problems in multimedia analysis. While most research focuses on the techniques of detecting feature points or capturing spatial and temporal information around those points, there has been very limited research on delving into the pooling strategies which are also important components of action recognition algorithms. In this paper, we propose a novel pooling framework by categorizing the interest points with respect to their idiosyncrasies. Specifically, we discuss three pooling strategies based on the optical flow orientation, foreground weight and spatio-temporal locations respectively and further investigate the fusion of different pooling strategies. For the encoding process, instead of the popular bag-of-visual words (BoV) method, we adopt the improved Fisher Vector (FV) approach. Our proposed methods are evaluated on a benchmark dataset with controlled settings (KTH), and two more challenging datasets with realistic background (HMDB51 and UCF101). The experimental results demonstrate that pooling strategies based on the appropriate idiosyncrasies of individual interest points can improve the performance of action classification.
Yuancheng Ye, Xiaodong Yang 0001, Yingli Tian
ICMR2
2015 Discriminative Hierarchical K-Means Tree for Large-Scale Image Classification
abstract
A key challenge in large-scale image classification is how to achieve efficiency in terms of both computation and memory without compromising classification accuracy. The learning-based classifiers achieve the state-of-the-art accuracies, but have been criticized for the computational complexity that grows linearly with the number of classes. The nonparametric nearest neighbor (NN)-based classifiers naturally handle large numbers of categories, but incur prohibitively expensive computation and memory costs. In this brief, we present a novel classification scheme, i.e., discriminative hierarchical K-means tree (D-HKTree), which combines the advantages of both learning-based and NN-based classifiers. The complexity of the D-HKTree only grows sublinearly with the number of categories, which is much better than the recent hierarchical support vector machines-based methods. The memory requirement is the order of magnitude less than the recent Naïve Bayesian NN-based approaches. The proposed D-HKTree classification scheme is evaluated on several challenging benchmark databases and achieves the state-of-the-art accuracies, while with significantly lower computation cost and memory requirement.
Shizhi Chen, Xiaodong Yang 0001, Yingli Tian
IEEE Trans. Neural Networks Learn. Syst.2
2014 Super Normal Vector for Activity Recognition Using Depth Sequences
abstract
This paper presents a new framework for human activity recognition from video sequences captured by a depth camera. We cluster hypersurface normals in a depth sequence to form the polynormal which is used to jointly characterize the local motion and shape information. In order to globally capture the spatial and temporal orders, an adaptive spatio-temporal pyramid is introduced to subdivide a depth video into a set of space-time grids. We then propose a novel scheme of aggregating the low-level polynormals into the super normal vector (SNV) which can be seen as a simplified version of the Fisher kernel representation. In the extensive experiments, we achieve classification results superior to all previous published results on the four public benchmark datasets, i.e., MSRAction3D, MSRDailyActivity3D, MSRGesture3D, and MSRActionPairs3D.
Xiaodong Yang 0001, Yingli Tian
CVPR1
2014 Action Recognition Using Super Sparse Coding Vector with Spatio-temporal Awareness
Xiaodong Yang 0001, Yingli Tian
ECCV (2)1
2014 Scene text recognition in multiple frames based on text tracking
abstract
Text signage as visual indicators in natural scene plays an important role in navigation and notification in our daily life. Most previous methods of scene text extraction are developed from a single scene image. In this paper, we propose a multi-frame based scene text recognition method by tracking text regions in a video captured by a moving camera. The main contributions of this paper are as follows. First, we present a framework of scene text recognition in multiple frames based on feature representation of scene text character (STC) for character prediction and conditional random field (CRF) model for word configuration. Second, a feature representation of STC is employed from dense sampled SIFT descriptors and Fisher Vector. Third, we collect a dataset for text information extraction from natural scene videos. Our proposed multi-frame scene text recognition is more compatible with image/video-based mobile applications. The experimental results demonstrate that STC prediction and word configuration in multiple frames based on text tracking significantly improves the performance of scene text recognition.
Xuejian Rong, Chucai Yi, Xiaodong Yang 0001, Yingli Tian
ICME3
2014 Effective 3D action recognition using EigenJoints
Xiaodong Yang 0001, Yingli Tian
J. Vis. Commun. Image Represent.1
2014 Assistive Clothing Pattern Recognition for Visually Impaired People
abstract
Choosing clothes with complex patterns and colors is a challenging task for visually impaired people. Automatic clothing pattern recognition is also a challenging research problem due to rotation, scaling, illumination, and especially large intraclass pattern variations. We have developed a camera-based prototype system that recognizes clothing patterns in four categories (plaid, striped, patternless, and irregular) and identifies 11 clothing colors. The system integrates a camera, a microphone, a computer, and a Bluetooth earpiece for audio description of clothing patterns and colors. A camera mounted upon a pair of sunglasses is used to capture clothing images. The clothing patterns and colors are described to blind users verbally. This system can be controlled by speech input through microphone. To recognize clothing patterns, we propose a novel Radon Signature descriptor and a schema to extract statistical properties from wavelet subbands to capture global features of clothing patterns. They are combined with local features to recognize complex clothing patterns. To evaluate the effectiveness of the proposed approach, we used the CCNY Clothing Pattern dataset. Our approach achieves 92.55% recognition accuracy which significantly outperforms the state-of-the-art texture analysis methods on clothing pattern recognition. The prototype was also used by ten visually impaired participants. Most thought such a system would support more independence in their daily life but they also made suggestions for improvements.
Xiaodong Yang 0001, Yingli Tian
IEEE Trans. Hum. Mach. Syst.1
2013 Visual speech learning from an e-tutor via dynamic lip movement-based video segmentation and comparison
abstract
This paper is motivated by the difficulties that deaf students encounter when learning speechreading and speaking; the skills that enable them to effectively communicate with hearing people. In this paper, we propose a speech learning prototype system based on the analysis and comparison of lip movements of an E-Tutor and those of a deaf student in a video. The main framework of our proposed system can be divided into two stages: lip movement segmentation and speech comparison. Lip movement segmentation fragments the frames of each word from a visual speech video sequence by analyzing the movement and shape of lips. Comparison determines whether a student is producing a correct word utterance or not, this is accomplished by comparing the lip shape and movements according to that of an e-tutor. To model lip movement, we compute two dynamic-based features by using a lip tracking method, which employs landmark points to define lip shapes. We utilize these dynamic features along with Space-Time Interest Points (STIP) to capture lip movements. In order to evaluate the effectiveness of our proposed methods, we collect a visual speech learning dataset consisting of 220 videos and 1100 word utterances. The proposed system achieves promising performances in both visual speech segmentation and visual speech comparison on this dataset.
Carol Mazuera, Xiaodong Yang 0001, Yingli Tian
BIBM2
2013 Feature Representations for Scene Text Character Recognition: A Comparative Study
abstract
Recognizing text character from natural scene images is a challenging problem due to background interferences and multiple character patterns. Scene Text Character (STC) recognition, which generally includes feature representation to model character structure and multi-class classification to predict label and score of character class, mostly plays a significant role in word-level text recognition. The contribution of this paper is a complete performance evaluation of image-based STC recognition, by comparing different sampling methods, feature descriptors, dictionary sizes, coding and pooling schemes, and SVM kernels. We systematically analyze the impact of each option in the feature representation and classification. The evaluation results on two datasets CHARS74K and ICDAR2003 demonstrate that Histogram of Oriented Gradient (HOG) descriptor, soft-assignment coding, max pooling, and Chi-Square Support Vector Machines (SVM) obtain the best performance among local sampling based feature representations. To improve STC recognition, we apply global sampling feature representation. We generate Global HOG (GHOG) by computing HOG descriptor from global sampling. GHOG enables better character structure modeling and obtains better performance than local sampling based feature representations. The GHOG also outperforms existing methods in the two benchmark datasets.
Chucai Yi, Xiaodong Yang 0001, Yingli Tian
ICDAR2
2013 Toward a computer vision-based wayfinding aid for blind persons to access unfamiliar indoor environments
Yingli Tian, Xiaodong Yang 0001, Chucai Yi, Aries Arditi
Mach. Vis. Appl.2
2013 Texture representations using subspace embeddings
Xiaodong Yang 0001, Yingli Tian
Pattern Recognit. Lett.1
2012 Recognizing actions using depth motion maps-based histograms of oriented gradients
abstract
In this paper, we propose an effective method to recognize human actions from sequences of depth maps, which provide additional body shape and motion information for action recognition. In our approach, we project depth maps onto three orthogonal planes and accumulate global activities through entire video sequences to generate the Depth Motion Maps (DMM). Histograms of Oriented Gradients (HOG) are then computed from DMM as the representation of an action video. The recognition results on Microsoft Research (MSR) Action3D dataset show that our approach significantly outperforms the state-of-the-art methods, although our representation is much more compact. In addition, we investigate how many frames are required in our framework to recognize actions on the MSR Action3D dataset. We observe that a short sub-sequence of 30-35 frames is sufficient to achieve comparable results to that operating on entire video sequences.
Xiaodong Yang 0001, Chenyang Zhang 0001, Yingli Tian
ACM Multimedia1
2012 Robust and Effective Component-Based Banknote Recognition for the Blind
abstract
We develop a novel camera-based computer vision technology to automatically recognize banknotes for assisting visually impaired people. Our banknote recognition system is robust and effective with the following features: 1) high accuracy: high true recognition rate and low false recognition rate, 2) robustness: handles a variety of currency designs and bills in various conditions, 3) high efficiency: recognizes banknotes quickly, and 4) ease of use: helps blind users to aim the target for image capture. To make the system robust to a variety of conditions including occlusion, rotation, scaling, cluttered background, illumination change, viewpoint variation, and worn or wrinkled bills, we propose a component-based framework by using Speeded Up Robust Features (SURF). Furthermore, we employ the spatial relationship of matched SURF features to detect if there is a bill in the camera view. This process largely alleviates false recognition and can guide the user to correctly aim at the bill to be recognized. The robustness and generalizability of the proposed system is evaluated on a dataset including both positive images (with U.S. banknotes) and negative images (no U.S. banknotes) collected under a variety of conditions. The proposed algorithm, achieves 100% true recognition rate and 0% false recognition rate. Our banknote recognition system is also tested by blind users.
Faiz M. Hasanuzzaman, Xiaodong Yang 0001, Yingli Tian
IEEE Trans. Syst. Man Cybern. Part C2
2011 Recognizing clothes patterns for blind people by confidence margin based feature combination
abstract
Clothes pattern recognition is a challenging task for blind or visually impaired people. Automatic clothes pattern recognition is also a challenging problem in computer vision due to the large pattern variations. In this paper, we present a new method to classify clothes patterns into 4 categories: stripe, lattice, special, and patternless. While existing texture analysis methods mainly focused on textures varying with distinctive pattern changes, they cannot achieve the same level of accuracy for clothes pattern recognition because of the large intra-class variations in each clothes pattern category. To solve this problem, we extract both structural feature and statistical feature from image wavelet subbands. Furthermore, we develop a new feature combination scheme based on the confidence margin of a classifier to combine the two types of features to form a novel local image descriptor in a compact and discriminative format. The recognition experiment is conducted on a database with 627 clothes images of 4 categories of patterns. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art texture analysis methods in the context of clothes pattern recognition.
Xiaodong Yang 0001
ACM Multimedia1
2011 Recognizing clothes patterns for blind people by confidence margin based feature combination
abstract
Clothes pattern recognition is a challenging task for blind or visually impaired people. Automatic clothes pattern recognition is also a challenging problem in computer vision due to the large pattern variations. In this paper, we present a new method to classify clothes patterns into 4 categories: stripe, lattice, special, and patternless. While existing texture analysis methods mainly focused on textures varying with distinctive pattern changes, they cannot achieve the same level of accuracy for clothes pattern recognition because of the large intra-class variations in each clothes pattern category. To solve this problem, we extract both structural feature and statistical feature from image wavelet subbands. Furthermore, we develop a new feature combination scheme based on the confidence margin of a classifier to combine the two types of features to form a novel local image descriptor in a compact and discriminative format. The recognition experiment is conducted on a database with 627 clothes images of 4 categories of patterns. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art texture analysis methods in the context of clothes pattern recognition.
Xiaodong Yang 0001, Yingli Tian
ACM Multimedia1
2010 Computer Vision-Based Door Detection for Accessibility of Unfamiliar Environments to Blind Persons
Yingli Tian, Xiaodong Yang 0001, Aries Arditi
ICCHP (2)2
2010 Context-based indoor object detection as an aid to blind persons accessing unfamiliar environments
abstract
Independent travel is a well known challenge for blind or visually impaired persons. In this paper, we propose a computer vision-based indoor wayfinding system for assisting blind people to independently access unfamiliar buildings. In order to find different rooms (i.e. an office, a lab, or a bathroom) and other building amenities (i.e. an exit or an elevator), we incorporate door detection with text recognition. First we develop a robust and efficient algorithm to detect doors and elevators based on general geometric shape, by combining edges and corners. The algorithm is generic enough to handle large intra-class variations of the object model among different indoor environments, as well as small inter-class differences between different objects such as doors and elevators. Next, to distinguish an office door from a bathroom door, we extract and recognize the text information associated with the detected objects. We first extract text regions from indoor signs with multiple colors. Then text character localization and layout analysis of text strings are applied to filter out background interference. The extracted text is recognized by using off-the-shelf optical character recognition (OCR) software products. The object type, orientation, and location can be displayed as speech for blind travelers.
Xiaodong Yang 0001, Yingli Tian, Chucai Yi, Aries Arditi
ACM Multimedia1