Gang Wang 0012

dblp:71/4292-12 · DBLP profile ↗
← Back
126ranked-venue papers
12as first author
10since 2021 · last 2026
0000-0002-1816-1457ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 85 · 6 first-authorArtificial intelligence and machine learning · 74 · 8 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Security and privacy · 3Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
87 papers
Video understanding and tracking · 27% Segmentation and scene understanding · 17% Face, body and person analysis · 12%
Computer networks
2 papers
Internet of things and sensor networks · 43% Edge and fog computing · 29% Vehicular, aerial and satellite networks · 29%
Computer graphics and multimedia
11 papers
Multimedia analysis and retrieval · 51% Image and video processing · 38% Computational photography and imaging · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 56% Emerging computing paradigms · 44%

Topics — the 30 heaviest of 163, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
3.4122020
Semantic Segmentation With Context Encoding and Multi-Path Decoding · IEEE Trans. Image Process. 2020
Toward Achieving Robust Low-Level and High-Level Scene Parsing · IEEE Trans. Image Process. 2019
Boundary-Aware Feature Propagation for Scene Segmentation · ICCV 2019
Computer vision › Video understanding and tracking
action recognition
2.172020
Skeleton-Based Online Action Prediction Using Scale Selection Network · IEEE Trans. Pattern Anal. Mach. Intell. 2020
NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Deep Multimodal Feature Analysis for Action Recognition in RGB+D Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
2.162020
Skeleton-Based Online Action Prediction Using Scale Selection Network · IEEE Trans. Pattern Anal. Mach. Intell. 2020
NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Skeleton-Based Human Action Recognition With Global Context-Aware Attention LSTM Networks · IEEE Trans. Image Process. 2018
Vehicular, aerial and satellite networks
UAV-assisted communication
1.722025
CSMAAC: Multi-Agent Reinforcement Learning Based Flight Control in Partially Observable Multi-UAV Assisted Crowd Sensing Systems · IEEE Trans. Mob. Comput. 2025
Transfer Learning for Joint Trajectory Control and Task Offloading in Large-Scale Partially Observable UAV-Assisted MEC · IEEE Trans. Mob. Comput. 2025
Computer vision › Face, body and person analysis
person re-identification
1.452018
Person Re-Identification With Cascaded Pairwise Convolutions · CVPR 2018
Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification · CVPR 2018
Learning Invariant Color Features for Person Reidentification · IEEE Trans. Image Process. 2016
Computer vision › Vision and language
image captioning
1.342019
Unpaired Image Captioning via Scene Graph Alignments · ICCV 2019
Unpaired Image Captioning by Language Pivoting · ECCV (1) 2018
Stack-Captioning: Coarse-to-Fine Learning for Image Captioning · AAAI 2018
Computer vision › Face, body and person analysis
face recognition
1.262017
Simultaneous Feature and Dictionary Learning for Image Set Based Face Recognition · IEEE Trans. Image Process. 2017
Reconstruction-Based Metric Learning for Unconstrained Face Verification · IEEE Trans. Inf. Forensics Secur. 2015
Joint Feature Learning for Face Recognition · IEEE Trans. Inf. Forensics Secur. 2015
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection
0.932018
Progressive Attention Guided Recurrent Network for Salient Object Detection · CVPR 2018
A Bi-Directional Message Passing Model for Salient Object Detection · CVPR 2018
Deep Level Sets for Salient Object Detection · CVPR 2017
Edge and fog computing › mobile edge computing
computation offloading
0.912025
Transfer Learning for Joint Trajectory Control and Task Offloading in Large-Scale Partially Observable UAV-Assisted MEC · IEEE Trans. Mob. Comput. 2025
Internet of things and sensor networks › wireless sensor network
data collection
0.912025
CSMAAC: Multi-Agent Reinforcement Learning Based Flight Control in Partially Observable Multi-UAV Assisted Crowd Sensing Systems · IEEE Trans. Mob. Comput. 2025
Internet of things and sensor networks
mobile crowdsensing
0.912025
CSMAAC: Multi-Agent Reinforcement Learning Based Flight Control in Partially Observable Multi-UAV Assisted Crowd Sensing Systems · IEEE Trans. Mob. Comput. 2025
Edge and fog computing
mobile edge computing
0.912025
Transfer Learning for Joint Trajectory Control and Task Offloading in Large-Scale Partially Observable UAV-Assisted MEC · IEEE Trans. Mob. Comput. 2025
Internet of things and sensor networks
trajectory control
0.912025
Transfer Learning for Joint Trajectory Control and Task Offloading in Large-Scale Partially Observable UAV-Assisted MEC · IEEE Trans. Mob. Comput. 2025
Distributed systems
consensus
0.912025
Practical Iterative Quantum Consensus Protocol With Sharding Construction · IEEE J. Sel. Areas Commun. 2025
Emerging computing paradigms
quantum computing
0.912025
Practical Iterative Quantum Consensus Protocol With Sharding Construction · IEEE J. Sel. Areas Commun. 2025
Machine learning › Deep learning architectures and training
recurrent neural network
0.852019
Learning Contextual Dependence With Convolutional Hierarchical Recurrent Neural Networks · IEEE Trans. Image Process. 2016
DAG-Recurrent Neural Networks for Scene Labeling · CVPR 2016
Early Action Prediction by Soft Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Computer vision › Image recognition and object detection
pedestrian detection
0.822020
Graininess-Aware Deep Feature Learning for Robust Pedestrian Detection · IEEE Trans. Image Process. 2020
Graininess-Aware Deep Feature Learning for Pedestrian Detection · ECCV (9) 2018
Computer vision › Video understanding and tracking
video object segmentation
0.822020
Motion-Guided Cascaded Refinement Network for Video Object Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Motion-Guided Cascaded Refinement Network for Video Object Segmentation · CVPR 2018
Machine learning › Deep learning architectures and training
attention mechanism
0.732020
Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification · CVPR 2018
Episodic CAMN: Contextual Attention-Based Memory Networks with Iterative Feedback for Scene Labeling · CVPR 2017
Graininess-Aware Deep Feature Learning for Robust Pedestrian Detection · IEEE Trans. Image Process. 2020
Computer vision › Segmentation and scene understanding › multimodal segmentation
RGB-D segmentation
0.732018
Multimodal Recurrent Neural Networks With Information Transfer Layers for Indoor Scene Labeling · IEEE Trans. Multim. 2018
Unsupervised Joint Feature Learning and Encoding for RGB-D Scene Labeling · IEEE Trans. Image Process. 2015
Multi-modal Unsupervised Feature Learning for RGB-D Scene Labeling · ECCV (5) 2014
Computer vision › Vision and language › image captioning
unpaired image captioning
0.722019
Unpaired Image Captioning via Scene Graph Alignments · ICCV 2019
Unpaired Image Captioning by Language Pivoting · ECCV (1) 2018
Computer vision › Video understanding and tracking
object tracking
0.732016
Object Instance Search in Videos via Spatio-Temporal Trajectory Discovery · IEEE Trans. Multim. 2016
Video Tracking Using Learned Hierarchical Features · IEEE Trans. Image Process. 2015
Real-time part-based visual tracking via adaptive correlation filters · CVPR 2015
Computer vision › Video understanding and tracking › action recognition › multimodal action recognition
RGB-D action recognition
0.722020
NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020
NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis · CVPR 2016
Computer vision › Video understanding and tracking › action recognition
human action recognition
0.712023
Human Action Recognition From Various Data Modalities: A Review · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Segmentation and scene understanding › image segmentation
scene segmentation
0.722018
Scene Segmentation with DAG-Recurrent Neural Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Context Contrasted Feature and Gated Multi-Scale Aggregation for Scene Segmentation · CVPR 2018
Computer vision › Segmentation and scene understanding
scene parsing
0.622019
Toward Achieving Robust Low-Level and High-Level Scene Parsing · IEEE Trans. Image Process. 2019
Scene Parsing With Integration of Parametric and Non-Parametric Models · IEEE Trans. Image Process. 2016
Machine learning › Deep learning architectures and training › recurrent neural network › attentive recurrent network
attention LSTM
0.622018
Skeleton-Based Human Action Recognition With Global Context-Aware Attention LSTM Networks · IEEE Trans. Image Process. 2018
Global Context-Aware Attention LSTM Networks for 3D Action Recognition · CVPR 2017
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.632015
Reconstruction-Based Metric Learning for Unconstrained Face Verification · IEEE Trans. Inf. Forensics Secur. 2015
Multi-manifold deep metric learning for image set classification · CVPR 2015
Neighborhood repulsed metric learning for kinship verification · CVPR 2012
Computer vision › Video understanding and tracking › action recognition
3d action recognition
0.522017
Global Context-Aware Attention LSTM Networks for 3D Action Recognition · CVPR 2017
Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition · ECCV (3) 2016
Computer vision › Video understanding and tracking
activity recognition
0.522019
Early Action Prediction by Soft Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Learning to Recognize Unsuccessful Activities Using a Two-Layer Latent Structural Model · ECCV (3) 2012

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 2.8deep learning · 2.0recurrent neural network · 1.8transfer learning · 1.7policy transfer · 1.7multi-agent actor-critic · 1.7mean field multi-agent actor-critic · 1.7causal inference · 1.7optical flow · 1.6attention mechanism · 0.9sharding · 0.9safety layer · 0.9greenberger-horne-zeilinger states · 0.9aharonov states · 0.9minimum spanning tree · 0.6variational method · 0.3spatial pyramid · 0.3spanning tree · 0.3
YearPublicationVenuePosition
2026 PLRF: A Personalized Learning Recommendation Framework Based on Federated Knowledge Graphs
Yiping Teng, Tiantian Yu, Gang Wang 0012, Zhen Song 0004, Bin Hu 0001
DASFAA (1)3
2026 Ponzitracker: A General Detection Framework for Ponzi Scheme in Blockchains
Gang Wang 0012, Yiping Teng, Zhen Song 0004, Leyang Li, Qinnan Zhang, Yanfeng Zhang 0001, Ge Yu 0001
DASFAA (6)1
2025 MMFed: A Multimodal Federated Learning Framework for Heterogeneous Devices
abstract
Existing federated learning frameworks are primarily designed for single-modal data. However, real-world scenarios require processing multi-modal data on heterogeneous devices. The gap between existing methods and real-world scenarios presents challenges in processing multimodal data on heterogeneous devices, significantly impacting model training efficiency. To address these issues, we propose a multimodal federated learning framework, which integrates multimodal algorithms with a semi-synchronous training method. The multimodal algorithm trains local autoencoders on different data modalities. By leveraging the similarity of encodings across different modalities with the same data labels, we further train and aggregate these local autoencoders into a global autoencoder, which is then deployed on the blockchain to perform downstream classification tasks. In the semi-synchronous training method, each device updates its parameters independently during a round. At the end of each round, a global aggregation combines the updates from devices. We conduct an empirical evaluation of our framework on various multimodal datasets, including Opportunity (Opp) Challenge, mHealth, and UR Fall Detection datasets. Experimental results demonstrate that our federated learning framework, outperforms the state-of-the-art multimodal frameworks on three multimodal datasets, achieving an average accuracy improvement of 9.07%. Furthermore, in terms of training speed, MMFed is obviously superior to synchronization strategies when it is extended to a large number of clients.
Gang Wang 0012, Yanfeng Zhang 0001, Chenhao Ying 0001, Qinnan Zhang, Zehui Xiong, Jiakang Wang, Ge Yu 0001
IEEE Internet Things J.1
2025 Practical Iterative Quantum Consensus Protocol With Sharding Construction
abstract
With the development of quantum blockchain, the quantum consensus protocols have garnered increasing attention, which play a crucial role in driving the implementation of quantum blockchains. However, existing protocols, derived from the classical consensus algorithms, face practical application challenges due to current quantum technology limitations. The first challenge is the bottleneck in generating large-scale entangled quantum states. The second challenge arises from the generation of malicious quantum states. The final challenge involves privacy concerns. To address these challenges, we propose a practical iterative QUantum consensus protocol with sharding construction, namely, Q-Union. In fact, Q-Union employs an iterative consensus algorithm where participating nodes are divided into multiple smaller shards, with the consensus process occurring within the current shard, and new shards are involved only if consensus is not achieved. Leveraging Greenberger-Horne- Zeilinge states and Aharonov states, Q-Union harnesses the advantages of quantum mechanics to achieve anonymous consensus, protecting the private information of participating nodes. Additionally, by integrating state verification, Q-Union ensures the correctness of the consensus procedure in the presence of malicious nodes generating adversarial quantum states. Finally, it is proven that Q-Union can also defend against Byzantine attacks from adversarial nodes, maintaining the same security level as traditional non-sharded consensus protocols. Specifically, it consistently outputs the correct consensus when the fraction of adversaries among participating nodes is less than 1/2 with synchronous communication. Both the theoretical analysis and performance illustration demonstrate the superior performance of the proposed Q-Union compared to state-of-the-art protocols.
Chenhao Ying 0001, Weiting Zhang, Xikun Jiang, Gang Wang 0012, Haiming Jin, Jie Li 0002, Yuan Luo 0003, Dacheng Tao
IEEE J. Sel. Areas Commun.5
2025 Transfer Learning for Joint Trajectory Control and Task Offloading in Large-Scale Partially Observable UAV-Assisted MEC
abstract
Existing joint trajectory control and task offloading (JTCTO) algorithms offer ultra-low latency services for smart devices (SDs) in unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC). However, these JTCTO algorithms typically require large training datasets to learn the optimal policies, leading to low learning efficiency. Additionally, most existing JTCTO algorithms are difficult to scale to environments with more than a few UAVs, as their complexity increases exponentially with the number of UAVs. In this paper, we propose a decentralized JTCTO algorithm based on the Policy Transfer and Mean Field-based Multi-Agent Actor-Critic (PTMF-MAAC). First, a novel policy transfer algorithm is proposed to determine which UAV's JTCTO strategy is helpful for each UAV and when to terminate the strategy to accelerate the learning efficiency of the UAV. Second, we propose a partially observable mean field algorithm that significantly reduces the model space by replacing the influence of all other UAVs on a particular UAV with an average value, thereby adapting to large-scale UAV scenarios. Experiments have shown that compared to the baseline, PTMF-MAAC reduces the system cost by 18.44%$\sim$28.57% and improves the model learning efficiency and adaptability to partially observable large-scale UAV-assisted MEC.
Gang Wang 0012, Lei Yang 0016, Yu Dai 0001
IEEE Trans. Mob. Comput.2
2025 CSMAAC: Multi-Agent Reinforcement Learning Based Flight Control in Partially Observable Multi-UAV Assisted Crowd Sensing Systems
abstract
In mobile crowd sensing systems, existing flight control methods enable unmanned aerial vehicles (UAVs) to provide high-quality data collection services for various applications. However, due to limited communication range, UAVs typically collect data under partial observability, hindering optimal performance without global environmental information. Additionally, many methods fail to enforce critical safety constraints. This paper proposes a communication-assisted safe multi-agent actor-critic-based UAV flight control method (CSMAAC). First, we propose an independent prediction communication partner model to address the partial observability problem. Based on the UAV's local observation, causal inference is used to obtain prior communication information between UAVs through a feed-forward neural network to help UAVs determine potential communication partners. Second, we utilize a critic-network to predict and quantify inter-UAV influence and determine the necessity of communication. By exchanging necessary information inter-UAV, UAVs can perceive global information, thereby solving the UAV's partial observability problem and reducing communication overhead. Moreover, we propose a similarity enhancement mechanism to improve the learning efficiency of the model by enhancing the connection between UAV observations and the policies of other UAVs. Finally, we introduce a safety layer to Actor-Network to ensure safe UAV flight. The simulation results show that the proposed method outperforms the baselines.
Gang Wang 0012, Lei Yang 0016, Chenhao Ying 0001
IEEE Trans. Mob. Comput.2
2024 Hammer: A General Blockchain Evaluation Framework
abstract
With the rising proliferation of blockchain systems and applications, choosing the appropriate blockchains to deploy applications is critical to achieving optimal performance. Evaluation frameworks provide a systematic approach to assessing and comparing different blockchain systems, guiding application developers to choose the most suitable one. However, existing evaluation frameworks still have limitations that affect their accuracy. First, most frameworks utilize workloads initially de-signed for traditional databases, which fail to capture the unique characteristics and requirements of blockchain systems. Second, these frameworks fail to generate correct results under heavy workloads due to their imbalanced task processing algorithms. Third, existing frameworks are tailored only for non-sharding blockchain architectures, limiting their ability to evaluate diverse blockchains. This paper introduces Hammer, a general blockchain evaluation framework that addresses the above limitations. It consists of two key components: workload prediction and asynchronous task processing. Workload prediction accurately predicts real-world workload trends by expanding the scope of temporal control sequences, providing a more realistic evaluation of blockchain performance. Asynchronous task processing handles heavy-load situations, enabling accurate evaluation of blockchain performance. Extensive experiments on various blockchains under Smallbank workload empower application developers to make informed decisions about blockchain selection and optimization.
Gang Wang 0012, Yanfeng Zhang 0001, Chenhao Ying 0001, Xiaohua Li 0004, Ge Yu 0001
ICDCS1
2023 Human Action Recognition From Various Data Modalities: A Review
abstract
Human Action Recognition (HAR) aims to understand human behavior and assign a label to each action. It has a wide range of applications, and therefore has been attracting increasing attention in the field of computer vision. Human actions can be represented using various data modalities, such as RGB, skeleton, depth, infrared, point cloud, event stream, audio, acceleration, radar, and WiFi signal, which encode different sources of useful yet distinct information and have various advantages depending on the application scenarios. Consequently, lots of existing works have attempted to investigate different types of approaches for HAR using various modalities. In this article, we present a comprehensive survey of recent progress in deep learning methods for HAR based on the type of input data modality. Specifically, we review the current mainstream deep learning methods for single data modalities and multiple data modalities, including the fusion-based and the co-learning-based frameworks. We also present comparative results on several benchmark datasets for HAR, together with insightful observations and inspiring future research directions.
Zehua Sun, Qiuhong Ke, Hossein Rahmani 0001, Mohammed Bennamoun, Gang Wang 0012, Jun Liu 0036
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 What have we learned from OpenReview?
Gang Wang 0012, Qi Peng 0004, Yanfeng Zhang 0001, Mingyang Zhang 0009
World Wide Web (WWW)1
2022 An Improved Empirical Mode Decomposition of Electroencephalogram Signals for Depression Detection
abstract
Depression is a mental disorder characterized by persistent low mood that affects a person’s thoughts, behavior, feelings, and sense of well-being. According to the World Health Organization (WHO), depression will become the second major life-threatening illness in 2020. Electroencephalogram (EEG) signals, which reflect the working status of human brain, are regarded as the best physiological tool for depression detection. Previous studies used the Empirical Mode Decomposition (EMD) method, which can deal with the highly complex, nonlinear and non-stationary nature of EEG, to extract features from EEG signals. However, for some special data, the neighboring components extracted through EMD could certainly have sections of data carrying the same frequency at different time durations. Thus, the Intrinsic Mode Functions (IMFs) of the data could be linearly dependent and the features coefficients of expansion based on IMFs could not be extracted, which can make the pre-proposed EMD-based feature extraction method impractical. In order to solve this problem, an improved EMD applying Singular Value Decomposition (SVD)-based feature extraction method was proposed in this study, which can extract the features coefficients of expansion based on all IMFs as accurately as possible, ignoring potentially linear dependence of IMFs. Experiments were conducted on four EEG databases for detecting depression. The improved EMD-based feature extraction method can extract feature from all three channels (Fp1, Fpz, and Fp2) on the four EEG databases. The average classification results of the proposed method on the four EEG databases including depressed patients and healthy subjects reached 83.27, 85.19, 81.98 and 88.07 percent, respectively, which were comparable with the pre-proposed EMD-based feature extraction method.
Jian Shen 0004, Xiaowei Zhang 0001, Gang Wang 0012, Zhijie Ding, Bin Hu 0001
IEEE Trans. Affect. Comput.3
2020 Object Tracking Via ImageNet Classification Scores
abstract
Object tracking is a challenging task in computer vision. The correlation filter based trackers are widely used for visual tracking due to their efficiencies. However, they cannot handle occlusion very well. In this paper, an effective method is proposed for occlusion detection based on high-level classification scores from the Convolutional Neural Network (CNN) trained on the ImageNet dataset. Also, we propose a novel tracking method by holistically considering multiple tracking models trained previously. In each frame, multiple correlation filters are first trained using hierarchical convolutional features, and then progressively selected according to the so-called tracking quality (status). Finally, a linear motion model is adopted to effectively re-detect the lost target. Experimental results have demonstrated that our method achieved good performance for handling occlusion.
Li Wang 0057, Ting Liu 0009, Bing Wang 0003, Jie Lin 0001, Xulei Yang, Gang Wang 0012
ICIP6
2020 MRP-Net: A Light Multiple Region Perception Neural Network for Multi-label AU Detection
abstract
Facial Action Units (AUs) are of great significance in communication. Automatic AU detection can improve the understanding of psychological condition and emotional status. Recently, a number of deep learning methods have been proposed to take charge with problems in automatic AU detection. Several challenges, like unbalanced labels and ignorance of local information, remain to be addressed. In this paper, we propose a fast and light neural network called MRP-Net, which is an end-to-end trainable method for facial AU detection to solve these problems. First, we design a Multiple Region Perception (MRP) module aimed at capturing different locations and sizes of features in the deeper level of the network without facial landmark points. Then, in order to balance the positive and negative samples in the large dataset, a batch balanced method adjusting the weight of every sample in one batch in our loss function is suggested. Experimental results on two popular AU datasets, BP4D and DISFA prove that MRP-Net outperforms state-of-the-art methods. Compared with the best method, not only does MRP-Net have an average F1 score improvement of 2.95% on BP4D and 5.43% on DISFA, and it also decreases the number of network parameters by 54.62% and the number of network FLOPs by 19.6%.
Honggang Zhang 0002, Gang Wang 0012
ICPR4
2020 DIPNet: Dynamic Identity Propagation Network for Video Object Segmentation
abstract
Many recent methods for semi-supervised Video Object Segmentation (VOS) have achieved good performance by exploiting the annotated first frame via one-shot fine-tuning or mask propagation. However, heavily relying on the first frame may weaken the robustness for VOS, since video objects can show large variations through time. In this work, we propose a Dynamic Identity Propagation Network (DIPNet) that adaptively propagates and accurately segments the video objects over time. To achieve this, DIPNet factors the VOS task at each time step into a dynamic propagation phase and a spatial segmentation phase. The former utilizes a novel identity representation to adaptively propagate objects’ reference information over time, which enhances the robustness to videos’ temporal variations. The segmentation phase uses the propagated information to tackle the object segmentation as an easier static image problem that can be optimized via light-weight fine-tuning on the first frame, thus reducing the computational cost. As a result, by optimizing these two components to complement each other, we can achieve a robust system for VOS. Evaluations on four benchmark datasets show that DIPNet provides state-of-the-art performance with time efficiency.
Ping Hu 0001, Jun Liu 0036, Gang Wang 0012, Vitaly Ablavsky, Kate Saenko, Stan Sclaroff
WACV3
2020 Motion-Guided Cascaded Refinement Network for Video Object Segmentation
abstract
In this work, we propose a motion-guided cascaded refinement network for video object segmentation. By assuming the foreground objects show different motion patterns from the background, for each video frame we apply an active contour model on optical flow to coarsely segment the foreground. The proposed Cascaded Refinement Network (CRN) then takes as guidance the coarse segmentation to generate an accurate segmentation in full resolution. In this way, the motion information and the deep CNNs can complement each other well to accurately segment the foreground objects from video frames. To deal with multi-instance cases, we extend our method with a spatial-temporal instance embedding model that further segments the foreground regions into instances and propagates instance labels. We further introduce a single-channel residual attention module in CRN to incorporate the coarse segmentation map as attention, which makes the network effective and efficient in both training and testing. We perform experiments on popular benchmarks and the results show that our method achieves state-of-the-art performance with high time efficiency.
Ping Hu 0001, Gang Wang 0012, Xiangfei Kong, Jason Kuen, Yap-Peng Tan
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Feature Boosting Network For 3D Pose Estimation
abstract
In this paper, a feature boosting network is proposed for estimating 3D hand pose and 3D body pose from a single RGB image. In this method, the features learned by the convolutional layers are boosted with a new long short-term dependence-aware (LSTD) module, which enables the intermediate convolutional feature maps to perceive the graphical long short-term dependency among different hand (or body) parts using the designed Graphical ConvLSTM. Learning a set of features that are reliable and discriminatively representative of the pose of a hand (or body) part is difficult due to the ambiguities, texture and illumination variation, and self-occlusion in the real application of 3D pose estimation. To improve the reliability of the features for representing each body part and enhance the LSTD module, we further introduce a context consistency gate (CCG) in this paper, with which the convolutional feature maps are modulated according to their consistency with the context representations. We evaluate the proposed method on challenging benchmark datasets for 3D hand pose estimation and 3D full body pose estimation. Experimental results show the effectiveness of our method that achieves state-of-the-art performance on both of the tasks.
Jun Liu 0036, Henghui Ding, Amir Shahroudy, Ling-Yu Duan, Xudong Jiang 0001, Gang Wang 0012, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
abstract
Research on depth-based human activity analysis achieved outstanding performance and demonstrated the effectiveness of 3D representation for action recognition. The existing depth-based and RGB+D-based action recognition benchmarks have a number of limitations, including the lack of large-scale training samples, realistic number of distinct class categories, diversity in camera views, varied environmental conditions, and variety of human subjects. In this work, we introduce a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames. This dataset contains 120 different action classes including daily, mutual, and health-related activities. We evaluate the performance of a series of existing 3D activity analysis methods on this dataset, and show the advantage of applying deep learning methods for 3D-based human action recognition. Furthermore, we investigate a novel one-shot 3D activity recognition problem on our dataset, and a simple yet effective Action-Part Semantic Relevance-aware (APSR) framework is proposed for this task, which yields promising results for recognition of the novel action classes. We believe the introduction of this large-scale dataset will enable the community to apply, adapt, and develop various data-hungry learning techniques for depth-based and RGB+D-based human activity understanding.
Jun Liu 0036, Amir Shahroudy, Mauricio Perez, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Skeleton-Based Online Action Prediction Using Scale Selection Network
abstract
Action prediction is to recognize the class label of an ongoing activity when only a part of it is observed. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is introduced to model the motion dynamics in temporal dimension via a sliding window over the temporal axis. Since there are significant temporal scale variations in the observed part of the ongoing action at different time steps, a novel window scale selection method is proposed to make our network focus on the performed part of the ongoing action and try to suppress the possible incoming interference from the previous actions at each step. An activation sharing scheme is also proposed to handle the overlapping computations among the adjacent time steps, which enables our framework to run more efficiently. Moreover, to enhance the performance of our framework for action prediction with the skeletal input data, a hierarchy of dilated tree convolutions are also designed to learn the multi-level structured semantic representations over the skeleton joints at each frame. Our proposed approach is evaluated on four challenging datasets. The extensive experiments demonstrate the effectiveness of our method for skeleton-based online action prediction.
Jun Liu 0036, Amir Shahroudy, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Context-Aware Deep Spatiotemporal Network for Hand Pose Estimation From Depth Images
abstract
As a fundamental and challenging problem in computer vision, hand pose estimation aims to estimate the hand joint locations from depth images. Typically, the problems are modeled as learning a mapping function from images to hand joint coordinates in a data-driven manner. In this paper, we propose a context-aware deep spatiotemporal network, a novel method to jointly model the spatiotemporal properties for hand pose estimation. Our proposed network is able to learn the representations of the spatial information and the temporal structure from the image sequences. Moreover, by adopting the adaptive fusion method, the model is capable of dynamically weighting different predictions to lay emphasis on sufficient context. Our method is examined on two common benchmarks, the experimental results demonstrate that our proposed approach achieves the best or the second-best performance with the state-of-the-art methods and runs in 60 fps.
Yiming Wu 0005, Wei Ji 0008, Xi Li 0001, Gang Wang 0012, Jianwei Yin, Fei Wu 0001
IEEE Trans. Cybern.4
2020 Semantic Segmentation With Context Encoding and Multi-Path Decoding
abstract
Semantic image segmentation aims to classify every pixel of a scene image to one of many classes. It implicitly involves object recognition, localization, and boundary delineation. In this paper, we propose a segmentation network called CGBNet to enhance the paring results by context encoding and multi-path decoding. We first propose a context encoding module that generates context contrasted local feature to make use of the informative context and the discriminative local information. This context encoding module greatly improves the segmentation performance, especially for inconspicuous objects. Furthermore, we propose a scale-selection scheme to selectively fuse the parsing results from different-scales of features at every spatial position. It adaptively selects appropriate score maps from rich scales of features. To improve the parsing results of boundary, we further propose a boundary delineation module that encourages the location-specific very-low-level feature near the boundaries to take part in the final prediction and suppresses them far from the boundaries. Without bells and whistles, the proposed segmentation network achieves very competitive performance in terms of all three different evaluation metrics consistently on the four popular scene segmentation datasets, Pascal Context, SUN-RGBD, Sift Flow, and COCO Stuff.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
IEEE Trans. Image Process.5
2020 Graininess-Aware Deep Feature Learning for Robust Pedestrian Detection
abstract
In this paper, we propose a graininess-aware deep feature learning method for pedestrian detection. Unlike most existing methods which utilize the convolutional features without explicit distinction, we appropriately exploit multiple convolutional layers and dynamically select most informative features. Specifically, we train a multi-scale pedestrian attention via pixel-wise segmentation supervision to efficiently identify the pedestrian of particular scales. We encodes the fine-grained attention map into the feature maps of the detection layers to guide them to highlight the pedestrians of specific scale and avoid the background interference. The graininess-aware feature maps generated with our attention mechanism are more focused on pedestrians, and in particular on the small-scale and occluded targets. We further introduce a zoom-in-zoom-out module to enhances the features by incorporating local details and context information. Extensive experimental results on five challenging pedestrian detection benchmarks show that our method achieves very competitive or even better performance with the state-of-the-arts and is faster than most existing approaches.
Chunze Lin, Jiwen Lu, Gang Wang 0012, Jie Zhou 0001
IEEE Trans. Image Process.3
2019 Depression Detection from Electroencephalogram Signals Induced by Affective Auditory Stimuli
abstract
Depression is a mental disorder characterized by emotional and cognitive dysfunction, which appears a state of low mood and aversion to activity. Depression can affect a person's thoughts, behavior, feelings, and sense of well-being. Depression is projected to be the second major life-threatening illness in 2020 by World Health Organization (WHO). Thus, it is urgent to detect and treat depression. Electroencephalogram (EEG) signals, which objectively reflect the working status of the human brain, are considered as promising physiological tools for depression detection. Negatively biased processing of affective stimuli in depression has been proven. In order to detect depression more effectively, we proposed an affective auditory stimuli induced depression detection method from EEG signals. In this method, we applied negative, positive and neutral affective auditory stimuli with several frequency selected from the International Affective Digitized Sounds (IADS-2) to induce negative affective bias in patients with depression. We synchronously collected EEG signals with three electrodes located on the prefrontal lobe (Fpl, Fpz, and Fp2), then extracted efficacious features by Empirical Mode Decomposition (EMD) based feature extraction method to detect depression effectively. The results of the proposed method showed that high-frequency affective auditory stimuli were more effective in depression detection and the frequency of affective auditory stimuli was a crucial property, which can influence the effectiveness of affective auditory stimuli in depression detection.
Jian Shen 0004, Xiaowei Zhang 0001, Junlei Li, Yuanxi Li 0001, Lei Feng 0005, Changqing Hu, Zhijie Ding, Gang Wang 0012, Bin Hu 0001
ACII8
2019 Semantic Correlation Promoted Shape-Variant Context for Segmentation
abstract
Context is essential for semantic segmentation. Due to the diverse shapes of objects and their complex layout in various scene images, the spatial scales and shapes of contexts for different objects have very large variation. It is thus ineffective or inefficient to aggregate various context information from a predefined fixed region. In this work, we propose to generate a scale- and shape-variant semantic mask for each pixel to confine its contextual region. To this end, we first propose a novel paired convolution to infer the semantic correlation of the pair and based on that to generate a shape mask. Using the inferred spatial scope of the contextual region, we propose a shape-variant convolution, of which the receptive field is controlled by the shape mask that varies with the appearance of input. In this way, the proposed network aggregates the context information of a pixel from its semantic-correlated region instead of a predefined fixed region. Furthermore, this work also proposes a labeling denoising model to reduce wrong predictions caused by the noisy low-level features. Without bells and whistles, the proposed segmentation network achieves new state-of-the-arts consistently on the six public segmentation datasets.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
CVPR5
2019 Boundary-Aware Feature Propagation for Scene Segmentation
abstract
In this work, we address the challenging issue of scene segmentation. To increase the feature similarity of the same object while keeping the feature discrimination of different objects, we explore to propagate information throughout the image under the control of objects' boundaries. To this end, we first propose to learn the boundary as an additional semantic class to enable the network to be aware of the boundary layout. Then, we propose unidirectional acyclic graphs (UAGs) to model the function of undirected cyclic graphs (UCGs), which structurize the image via building graphic pixel-by-pixel connections, in an efficient and effective way. Furthermore, we propose a boundary-aware feature propagation (BFP) module to harvest and propagate the local features within their regions isolated by the learned boundaries in the UAG-structured image. The proposed BFP is capable of splitting the feature propagation into a set of semantic groups via building strong connections among the same segment region but weak connections between different segment regions. Without bells and whistles, our approach achieves new state-of-the-art segmentation performance on three challenging semantic segmentation datasets, i.e., PASCAL-Context, CamVid, and Cityscapes.
Henghui Ding, Xudong Jiang 0001, Ai Qun Liu, Nadia Magnenat-Thalmann, Gang Wang 0012
ICCV5
2019 Unpaired Image Captioning via Scene Graph Alignments
abstract
Most of current image captioning models heavily rely on paired image-caption datasets. However, getting large scale image-caption paired data is labor-intensive and time-consuming. In this paper, we present a scene graph-based approach for unpaired image captioning. Our framework comprises an image scene graph generator, a sentence scene graph generator, a scene graph encoder, and a sentence decoder. Specifically, we first train the scene graph encoder and the sentence decoder on the text modality. To align the scene graphs between images and sentences, we propose an unsupervised feature alignment method that maps the scene graph features from the image to the sentence modality. Experimental results show that our proposed model can generate quite promising results without using any image-caption training pairs, outperforming existing methods by a wide margin.
Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai 0001, Handong Zhao, Xu Yang 0021, Gang Wang 0012
ICCV6
2019 Learning Hierarchical Features for Visual Object Tracking With Recursive Neural Networks
abstract
Recently, deep learning has achieved very promising results in visual object tracking. Deep neural networks in existing tracking methods require a lot of training data to learn a large number of parameters. However, training data is not sufficient for visual object tracking as annotations of a target object are only available in the first frame of a test sequence. In this paper, we propose to learn hierarchical features for visual object tracking by using tree structure based Recursive Neural Networks (RNN), which have a relatively small number of parameters compared to other deep neural networks (e.g. Convolutional Neural Networks (CNN)) due to all basic modules in RNN share only one set of parameters. Experimental results demonstrate that our feature learning algorithm can significantly improve tracking performance on benchmark datasets.
Li Wang 0057, Ting Liu 0009, Bing Wang 0003, Jie Lin 0001, Xulei Yang, Gang Wang 0012
ICIP6
2019 2D LiDAR Map Prediction via Estimating Motion Flow with GRU
abstract
It is a significant problem to predict the 2D LiDAR map at next moment for robotics navigation and path-planning. To tackle this problem, we resort to the motion flow between adjacent maps, as motion flow is a powerful tool to process and analyze the dynamic data, which is named optical flow in video processing. However, unlike video, which contains abundant visual features in each frame, a 2D LiDAR map lacks distinctive local features. To alleviate this challenge, we propose to estimate the motion flow based on deep neural networks inspired by its powerful representation learning ability in estimating the optical flow of the video. To this end, we design a recurrent neural network based on gated recurrent unit, which is named LiDAR-FlowNet. As a recurrent neural network can encode the temporal dynamic information, our LiDAR-FlowNet can estimate motion flow between the current map and the unknown next map only from the current frame and previous frames. A self-supervised strategy is further designed to train the LiDAR-FlowNet model effectively, while no training data need to be manually annotated. With the estimated motion flow, it is straightforward to predict the 2D LiDAR map at the next moment. Experimental results verify the effectiveness of our LiDAR-FlowNet as well as the proposed training strategy. The results of the predicted LiDAR map also show the advantages of our motion flow based method.
Yafei Song 0002, Yonghong Tian 0001, Gang Wang 0012, Mingyang Li 0001
ICRA3
2019 Early Action Prediction by Soft Regression
abstract
We propose a novel approach for predicting on-going action with the assistance of a low-cost depth camera. Our approach introduces a soft regression-based early prediction framework. In this framework, we estimate soft labels for the subsequences at different progress levels, jointly learned with an action predictor. Our formulation of soft regression framework 1) overcomes a usual assumption in existing early action prediction systems that the progress level of on-going sequence is given in the testing stage; and 2) presents a theoretical framework to better resolve the ambiguity and uncertainty of subsequences at early performing stage. The proposed soft regression framework is further enhanced in order to take the relationships among subsequences and the discrepancy of soft labels over different classes into consideration, so that a Multiple Soft labels Recurrent Neural Network (MSRNN) is finally developed. For real-time performance, we also introduce a new RGB-D feature called "local accumulative frame feature (LAFF)", which can be computed efficiently by constructing an integral feature map. Our experiments on three RGB-D benchmark datasets and an unconstrained RGB action set demonstrate that the proposed regression-based early action prediction model outperforms existing models significantly and also show that the early action prediction on RGB-D sequence is more accurate than that on RGB channel.
Jianfang Hu, Wei-Shi Zheng 0001, Lianyang Ma, Gang Wang 0012, Jian-Huang Lai, Jianguo Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Toward Achieving Robust Low-Level and High-Level Scene Parsing
abstract
In this paper, we address the challenging task of scene segmentation. We first discuss and compare two widely used approaches to retain detailed spatial information from pretrained CNN - "dilation" and "skip". Then, we demonstrate that the parsing performance of "skip" network can be noticeably improved by modifying the parameterization of skip layers. Furthermore, we introduce a "dense skip" architecture to retain a rich set of low-level information from pre-trained CNN, which is essential to improve the low-level parsing performance. Meanwhile, we propose a convolutional context network (CCN) and place it on top of pre-trained CNNs, which is used to aggregate contexts for high-level feature maps so that robust high-level parsing can be achieved. We name our segmentation network enhanced fully convolutional network (EFCN) based on its significantly enhanced structure over FCN. Extensive experimental studies justify each contribution separately. Without bells and whistles, EFCN achieves state-of-the-arts on segmentation datasets of ADE20K, Pascal Context, SUN-RGBD and Pascal VOC 2012.
Bing Shuai, Henghui Ding, Ting Liu 0009, Gang Wang 0012, Xudong Jiang 0001
IEEE Trans. Image Process.4
2019 Multimodal Depression Detection: Fusion of Electroencephalography and Paralinguistic Behaviors Using a Novel Strategy for Classifier Ensemble
abstract
Currently, depression has become a common mental disorder and one of the main causes of disability worldwide. Due to the difference in depressive symptoms evoked by individual differences, how to design comprehensive and effective depression detection methods has become an urgent demand. This study explored from physiological and behavioral perspectives simultaneously and fused pervasive electroencephalography (EEG) and vocal signals to make the detection of depression more objective, effective and convenient. After extraction of several effective features for these two types of signals, we trained six representational classifiers on each modality, then denoted diversity and correlation of decisions from different classifiers using co-decision tensor and combined these decisions into the ultimate classification result with multi-agent strategy. Experimental results on 170 (81 depressed patients and 89 normal controls) subjects showed that the proposed multi-modal depression detection strategy is superior to the single-modal classifiers or other typical late fusion strategies in accuracy, f1-score and sensitivity. This work indicates that late fusion of pervasive physiological and behavioral signals is promising for depression detection and the multi-agent strategy can take advantage of diversity and correlation of different classifiers effectively to gain a better final decision.
Xiaowei Zhang 0001, Jian Shen 0004, Zia Ud Din, Jinyong Liu, Gang Wang 0012, Bin Hu 0001
IEEE J. Biomed. Health Informatics5
2018 Stack-Captioning: Coarse-to-Fine Learning for Image Captioning
abstract
The existing image captioning approaches typically train a one-stage sentence decoder, which is difficult to generate rich fine-grained descriptions. On the other hand, multi-stage image caption model is hard to train due to the vanishing gradient problem. In this paper, we propose a coarse-to-fine multi-stage prediction framework for image captioning, composed of multiple decoders each of which operates on the output of the previous stage, producing increasingly refined image descriptions. Our proposed learning approach addresses the difficulty of vanishing gradients during training by providing a learning objective function that enforces intermediate supervisions. Particularly, we optimize our model with a reinforcement learning approach which utilizes the output of each intermediate decoder's test-time inference algorithm as well as the output of its preceding decoder to normalize the rewards, which simultaneously solves the well-known exposure bias problem and the loss-evaluation mismatch problem. We extensively evaluate the proposed approach on MSCOCO and show that our approach can achieve the state-of-the-art performance.
Jiuxiang Gu, Jianfei Cai 0001, Gang Wang 0012, Tsuhan Chen
AAAI3
2018 SSNet: Scale Selection Network for Online 3D Action Prediction
abstract
In action prediction (early action recognition), the goal is to predict the class label of an ongoing action using its observed part so far. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is introduced to model the motion dynamics in temporal dimension via a sliding window over the time axis. As there are significant temporal scale variations of the observed part of the ongoing action at different progress levels, we propose a novel window scale selection scheme to make our network focus on the performed part of the ongoing action and try to suppress the noise from the previous actions at each time step. Furthermore, an activation sharing scheme is proposed to deal with the overlapping computations among the adjacent steps, which allows our model to run more efficiently. The extensive experiments on two challenging datasets show the effectiveness of the proposed action prediction framework.
Jun Liu 0036, Amir Shahroudy, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot
CVPR3
2018 Context Contrasted Feature and Gated Multi-Scale Aggregation for Scene Segmentation
abstract
Scene segmentation is a challenging task as it need label every pixel in the image. It is crucial to exploit discriminative context and aggregate multi-scale features to achieve better segmentation. In this paper, we first propose a novel context contrasted local feature that not only leverages the informative context but also spotlights the local information in contrast to the context. The proposed context contrasted local feature greatly improves the parsing performance, especially for inconspicuous objects and background stuff. Furthermore, we propose a scheme of gated sum to selectively aggregate multi-scale features for each spatial position. The gates in this scheme control the information flow of different scale features. Their values are generated from the testing image by the proposed network learnt from the training data so that they are adaptive not only to the training data, but also to the specific testing image. Without bells and whistles, the proposed approach achieves the state-of-the-arts consistently on the three popular scene segmentation datasets, Pascal Context, SUN-RGBD and COCO Stuff.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
CVPR5
2018 Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval With Generative Models
abstract
Textual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval performance. Unlike existing image-text retrieval approaches that embed image-text pairs as single feature vectors in a common representational space, we propose to incorporate generative processes into the cross-modal feature embedding, through which we are able to learn not only the global abstract features but also the local grounded features. Extensive experiments show that our framework can well match images and sentences with complex content, and achieve the state-of-the-art cross-modal retrieval results on MSCOCO dataset.
Jiuxiang Gu, Jianfei Cai 0001, Shafiq R. Joty, Li Niu 0002, Gang Wang 0012
CVPR5
2018 Motion-Guided Cascaded Refinement Network for Video Object Segmentation
abstract
Deep CNNs have achieved superior performance in many tasks of computer vision and image understanding. However, it is still difficult to effectively apply deep CNNs to video object segmentation(VOS) since treating video frames as separate and static will lose the information hidden in motion. To tackle this problem, we propose a Motion-guided Cascaded Refinement Network for VOS. By assuming the object motion is normally different from the background motion, for a video frame we first apply an active contour model on optical flow to coarsely segment objects of interest. Then, the proposed Cascaded Refinement Network(CRN) takes the coarse segmentation as guidance to generate an accurate segmentation of full resolution. In this way, the motion information and the deep CNNs can well complement each other to accurately segment objects from video frames. Furthermore, in CRN we introduce a Single-channel Residual Attention Module to incorporate the coarse segmentation map as attention, making our network effective and efficient in both training and testing. We perform experiments on the popular benchmarks and the results show that our method achieves state-of-the-art performance at a much faster speed.
Ping Hu 0001, Gang Wang 0012, Xiangfei Kong, Jason Kuen, Yap-Peng Tan
CVPR2
2018 Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
abstract
It is desirable to train convolutional networks (CNNs) to run more efficiently during inference. In many cases however, the computational budget that the system has for inference cannot be known beforehand during training, or the inference budget is dependent on the changing real-time resource availability. Thus, it is inadequate to train just inference-efficient CNNs, whose inference costs are not adjustable and cannot adapt to varied inference budgets. We propose a novel approach for cost-adjustable inference in CNNs - Stochastic Downsampling Point (SDPoint). During training, SDPoint applies feature map downsampling to a random point in the layer hierarchy, with a random downsampling ratio. The different stochastic downsampling configurations known as SDPoint instances (of the same model) have computational costs different from each other, while being trained to minimize the same prediction loss. Sharing network parameters across different instances provides significant regularization boost. During inference, one may handpick a SDPoint instance that best fits the inference budget. The effectiveness of SDPoint, as both a cost-adjustable inference approach and a regularizer, is validated through extensive experiments on image classification.
Jason Kuen, Xiangfei Kong, Zhe Lin 0001, Gang Wang 0012, Jianxiong Yin, Simon See, Yap-Peng Tan
CVPR4
2018 Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification
abstract
Typical person re-identification (ReID) methods usually describe each pedestrian with a single feature vector and match them in a task-specific metric space. However, the methods based on a single feature vector are not sufficient enough to overcome visual ambiguity, which frequently occurs in real scenario. In this paper, we propose a novel end-to-end trainable framework, called Dual ATtention Matching network (DuATM), to learn context-aware feature sequences and perform attentive sequence comparison simultaneously. The core component of our DuATM framework is a dual attention mechanism, in which both intrasequence and inter-sequence attention strategies are used for feature refinement and feature-pair alignment, respectively. Thus, detailed visual cues contained in the intermediate feature sequences can be automatically exploited and properly compared. We train the proposed DuATM network as a siamese network via a triplet loss assisted with a decorrelation loss and a cross-entropy loss. We conduct extensive experiments on both image and video based ReID benchmark datasets. Experimental results demonstrate the significant advantages of our approach compared to the state-of-the-art methods.
Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jason Kuen, Xiangfei Kong, Alex Chichung Kot, Gang Wang 0012
CVPR7
2018 Person Re-Identification With Cascaded Pairwise Convolutions
abstract
In this paper, a novel deep architecture named BraidNet is proposed for person re-identification. BraidNet has a specially designed WConv layer, and the cascaded WConv structure learns to extract the comparison features of two images, which are robust to misalignments and color differences across cameras. Furthermore, a Channel Scaling layer is designed to optimize the scaling factor of each input channel, which helps mitigate the zero gradient problem in the training phase. To solve the problem of imbalanced volume of negative and positive training samples, a Sample Rate Learning strategy is proposed to adaptively update the ratio between positive and negative samples in each batch. Experiments conducted on CUHK03-Detected, CUHK03-Labeled, CUHK01, Market-1501 and DukeMTMC-reID datasets demonstrate that our method achieves competitive performance when compared to state-of-the-art methods.
Gang Wang 0012
CVPR4
2018 A Bi-Directional Message Passing Model for Salient Object Detection
abstract
Recent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detection. In this paper, we propose a novel bi-directional message passing model to integrate multi-level features for salient object detection. At first, we adopt a Multi-scale Context-aware Feature Extraction Module (MCFEM) for multi-level feature maps to capture rich context information. Then a bi-directional structure is designed to pass messages between multi-level features, and a gate function is exploited to control the message passing rate. We use the features after message passing, which simultaneously encode semantic information and spatial details, to predict saliency maps. Finally, the predicted results are efficiently combined to generate the final saliency map. Quantitative and qualitative experiments on five benchmark datasets demonstrate that our proposed model performs favorably against the state-of-the-art methods under different evaluation metrics.
Lu Zhang 0053, Ju Dai, Huchuan Lu, You He 0002, Gang Wang 0012
CVPR5
2018 Progressive Attention Guided Recurrent Network for Salient Object Detection
abstract
Effective convolutional features play an important role in saliency estimation but how to learn powerful features for saliency is still a challenging task. FCN-based methods directly apply multi-level convolutional features without distinction, which leads to sub-optimal results due to the distraction from redundant details. In this paper, we propose a novel attention guided network which selectively integrates multi-level contextual information in a progressive manner. Attentive features generated by our network can alleviate distraction of background thus achieve better performance. On the other hand, it is observed that most of existing algorithms conduct salient object detection by exploiting side-output features of the backbone feature extraction network. However, shallower layers of backbone network lack the ability to obtain global semantic information, which limits the effective feature learning. To address the problem, we introduce multi-path recurrent feedback to enhance our proposed progressive attention driven framework. Through multi-path recurrent connections, global semantic information from the top convolutional layer is transferred to shallower layers, which intrinsically refines the entire network. Experimental results on six benchmark datasets demonstrate that our algorithm performs favorably against the state-of-the-art approaches.
Tiantian Wang 0002, Jinqing Qi, Huchuan Lu, Gang Wang 0012
CVPR5
2018 Unpaired Image Captioning by Language Pivoting
Jiuxiang Gu, Shafiq R. Joty, Jianfei Cai 0001, Gang Wang 0012
ECCV (1)4
2018 Graininess-Aware Deep Feature Learning for Pedestrian Detection
Chunze Lin, Jiwen Lu, Gang Wang 0012, Jie Zhou 0001
ECCV (9)3
2018 QL-Net: Quantized-by-LookUp CNN
abstract
Convolutional Neural Networks (CNNs) have achieved a state-of-the-art performance in the different computer vision tasks. However, CNN algorithms are computationally and power intensive, which makes them difficult to run on wearable and embedded systems. One way to address this constraint is to reduce the number of computational operations performed. Recently, several approaches addressed the problem of the computational complexity in the CNNs. Most of these methods, however, require a dedicated hardware. We propose a new method for the computation reduction in CNNs that substitutes Multiply and Accumulate (MAC) operations with a codebook lookup and can be executed on the generic hardware. The proposed method called QL-Net combines several concepts: (i) a codebook construction, (ii) a layer-wise retraining strategy, and (iii) a substitution of the MAC operations with the lookup of the convolution responses at inference time. The proposed QL-Net achieves a 98.6% accuracy on the MNIST dataset with a 5.8x reduction in runtime, when compared to MAC-based CNN model that achieved a 99.2% accuracy.
Kamila Abdiyeva, Kim-Hui Yap, Gang Wang 0012, Narendra Ahuja, Martin Lukac
ICARCV3
2018 Learning fine-grained features via a CNN Tree for Large-scale Classification
Zhenhua Wang 0002, Gang Wang 0012
Neurocomputing3
2018 Spatiotemporal GMM for Background Subtraction with Superpixel Hierarchy
abstract
We propose a background subtraction algorithm using hierarchical superpixel segmentation, spanning trees and optical flow. First, we generate superpixel segmentation trees using a number of Gaussian Mixture Models (GMMs) by treating each GMM as one vertex to construct spanning trees. Next, we use the -smoother to enhance the spatial consistency on the spanning trees and estimate optical flow to extend the -smoother to the temporal domain. Experimental results on synthetic and real-world benchmark datasets show that the proposed algorithm performs favorably for background subtraction in videos against the state-of-the-art methods in spite of frequent and sudden changes of pixel values.
Xing Wei 0001, Qingxiong Yang, Qing Li 0001, Gang Wang 0012, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Skeleton-Based Action Recognition Using Spatio-Temporal LSTM Network with Trust Gates
abstract
Skeleton-based human action recognition has attracted a lot of research attention during the past few years. Recent works attempted to utilize recurrent neural networks to model the temporal dependencies between the 3D positional configurations of human body joints for better analysis of human activities in the skeletal data. The proposed work extends this idea to spatial domain as well as temporal domain to better analyze the hidden sources of action-related information within the human skeleton sequences in both of these domains simultaneously. Based on the pictorial structure of Kinect's skeletal data, an effective tree-structure based traversal framework is also proposed. In order to deal with the noise in the skeletal data, a new gating mechanism within LSTM module is introduced, with which the network can learn the reliability of the sequential data and accordingly adjust the effect of the input data on the updating procedure of the long-term context representation stored in the unit's memory cell. Moreover, we introduce a novel multi-modal feature fusion strategy within the LSTM unit in this paper. The comprehensive experimental results on seven challenging benchmark datasets for human action recognition demonstrate the effectiveness of the proposed method.
Jun Liu 0036, Amir Shahroudy, Dong Xu 0001, Alex Chichung Kot, Gang Wang 0012
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Deep Multimodal Feature Analysis for Action Recognition in RGB+D Videos
abstract
Single modality action recognition on RGB or depth sequences has been extensively explored recently. It is generally accepted that each of these two modalities has different strengths and limitations for the task of action recognition. Therefore, analysis of the RGB+D videos can help us to better study the complementary properties of these two types of modalities and achieve higher levels of performance. In this paper, we propose a new deep autoencoder based shared-specific feature factorization network to separate input multimodal signals into a hierarchy of components. Further, based on the structure of the features, a structured sparsity learning machine is proposed which utilizes mixed norms to apply regularization within components and group selection between them for better classification performance. Our experimental results show the effectiveness of our cross-modality feature analysis framework by achieving state-of-the-art accuracy for action classification on five challenging benchmark datasets.
Amir Shahroudy, Tian-Tsong Ng, Yihong Gong, Gang Wang 0012
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Scene Segmentation with DAG-Recurrent Neural Networks
abstract
In this paper, we address the challenging task of scene segmentation. In order to capture the rich contextual dependencies over image regions, we propose Directed Acyclic Graph-Recurrent Neural Networks (DAG-RNN) to perform context aggregation over locally connected feature maps. More specifically, DAG-RNN is placed on top of pre-trained CNN (feature extractor) to embed context into local features so that their representative capability can be enhanced. In comparison with plain CNN (as in Fully Convolutional Networks-FCN), DAG-RNN is empirically found to be significantly more effective at aggregating context. Therefore, DAG-RNN demonstrates noticeably performance superiority over FCNs on scene segmentation. Besides, DAG-RNN entails dramatically less parameters as well as demands fewer computation operations, which makes DAG-RNN more favorable to be potentially applied on resource-constrained embedded devices. Meanwhile, the class occurrence frequencies are extremely imbalanced in scene segmentation, so we propose a novel class-weighted loss to train the segmentation network. The loss distributes reasonably higher attention weights to infrequent classes during network training, which is essential to boost their parsing performance. We evaluate our segmentation network on three challenging public scene segmentation benchmarks: Sift Flow, Pascal Context and COCO Stuff. On top of them, we achieve very impressive segmentation performance.
Bing Shuai, Zhen Zuo, Bing Wang 0003, Gang Wang 0012
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Recent advances in convolutional neural networks
Jiuxiang Gu, Zhenhua Wang 0002, Jason Kuen, Lianyang Ma, Amir Shahroudy, Bing Shuai, Ting Liu 0009, Gang Wang 0012, Jianfei Cai 0001, Tsuhan Chen
Pattern Recognit.9
2018 Hierarchical Spatial Sum-Product Networks for Action Recognition in Still Images
abstract
Recognizing actions from still images has been popularly studied recently. In this paper, we model an action class as a flexible number of spatial configurations of body parts by proposing a new spatial sum–product network (SPN). First, we discover a set of parts in image collections via unsupervised learning. Then, our new spatial SPN is applied to model the spatial relationship and also the high-order correlations of parts. To learn robust networks, we further develop a hierarchical spatial SPN method, which models pairwise spatial relationship between parts inside subimages and models the correlation of subimages via extra layers of SPN. Our method is shown to be effective on two benchmark data sets.
Gang Wang 0012
IEEE Trans. Circuits Syst. Video Technol.2
2018 Skeleton-Based Human Action Recognition With Global Context-Aware Attention LSTM Networks
abstract
Human action recognition in 3D skeleton sequences has attracted a lot of research attention. Recently, long short-term memory (LSTM) networks have shown promising performance in this task due to their strengths in modeling the dependencies and dynamics in sequential data. As not all skeletal joints are informative for action recognition, and the irrelevant joints often bring noise which can degrade the performance, we need to pay more attention to the informative ones. However, the original LSTM network does not have explicit attention ability. In this paper, we propose a new class of LSTM network, global context-aware attention LSTM, for skeleton-based action recognition, which is capable of selectively focusing on the informative joints in each frame by using a global context memory cell. To further improve the attention capability, we also introduce a recurrent attention mechanism, with which the attention performance of our network can be enhanced progressively. Besides, a two-stream framework, which leverages coarse-grained attention and fine-grained attention, is also introduced. The proposed method achieves state-of-the-art performance on five challenging datasets for skeleton-based action recognition.
Jun Liu 0036, Gang Wang 0012, Ling-Yu Duan, Kamila Abdiyeva, Alex Chichung Kot
IEEE Trans. Image Process.2
2018 Deep Context-Sensitive Facial Landmark Detection With Tree-Structured Modeling
abstract
Facial landmark detection is typically cast as a point-wise regression problem that focuses on how to build an effective image-to-point mapping function. In this paper, we propose an end-to-end deep learning approach for contextually discriminative feature construction together with effective facial structure modeling. The proposed learning approach is able to predict more contextually discriminative facial landmarks by capturing their associated contextual information. Moreover, we present a tree model to characterize human face structure and a structural loss function to measure the deformation cost between the ground-truth and predicted tree model, which are further incorporated into the proposed learning approach and jointly optimized within a unified framework. The presented tree model is able to well characterize the spatial layout patterns of facial landmarks for capturing the facial structure information. Experimental results demonstrate the effectiveness of the proposed approach against the state-of-the-art over the MTFL and AFLW-full data sets.
Jiajian Zeng, Xi Li 0001, Debbah Abderrahmane Mahdi, Fei Wu 0001, Gang Wang 0012
IEEE Trans. Image Process.6
2018 Multimodal Recurrent Neural Networks With Information Transfer Layers for Indoor Scene Labeling
abstract
This paper proposes a new method called multimodal recurrent neural networks (RNNs) for RGB-D scene semantic segmentation. It is optimized to classify image pixels given two input sources: RGB color channels and depth maps. It simultaneously performs training of two RNNs that are crossly connected through information transfer layers, which are learnt to adaptively extract relevant cross-modality features. Each RNN model learns its representations from its own previous hidden states and transferred patterns from the other RNNs previous hidden states; thus, both model-specific and cross-modality features are retained. We exploit the structure of quad-directional 2D-RNNs to model the short- and long-range contextual information in the 2D input image. We carefully designed various baselines to efficiently examine our proposed model structure. We test our multimodal RNNs method on popular RGB-D benchmarks and show how it outperforms previous methods significantly and achieves competitive results with other state-of-the-art works.
Abrar H. Abdulnabi, Bing Shuai, Zhen Zuo, Lap-Pui Chau, Gang Wang 0012
IEEE Trans. Multim.5
2018 Recurrent Spatial Pyramid CNN for Optical Flow Estimation
abstract
Optical flow estimation plays an important role in many multimedia and computer vision tasks. Although great progress has been made in applying convolutional neural networks (CNNs) to estimate optical flow in recent works, it is still difficult for CNNs to generate optical flow with the desired effectiveness and efficiency. Compared to CNN-based methods, conventional variational methods normally perform to optimize an energy function and produce optical flow with more precise details. Inspired by the effectiveness of variational methods and deep CNNs, we propose a recurrent spatial pyramid (RecSPy) network for optical flow estimation. To deal with large displacements and to decrease the number of parameters, we formulate the spatial pyramid as a recurrent process, and adopt a CNN to refine optical flow at each spatial scale. Furthermore, to improve the results with more precise details, we propose an energy function that encodes structure and constancy constraints to help refine the optical flow at each spatial scale. The combination of the proposed RecSPy network and the proposed energy-based refinement enables our system to estimate optical flow effectively and efficiently. Experimental results on the benchmarks validate the effectiveness and efficiency of the proposed method.
Ping Hu 0001, Gang Wang 0012, Yap-Peng Tan
IEEE Trans. Multim.2
2017 Ensemble-based depression detection in speech
abstract
Depression detection using speech signal is becoming an attractive topic because it is fast, convenient and non-invasive. Many researches aimed at improving depression classification performance. This study investigated application of ensemble learners in depression detection and compared three speaking styles (interview, reading and picture description) in ensembles. A speech dataset collecting from 184 subjects (92 depressed patients and 92 healthy controls) was used for these goals. The results showed that ensemble learners perform better than individual learners apparently. Interview is a more effective speaking style than reading and picture description for speech acquisition. These findings suggest us ensemble model based multi-utterance in interview is the best way to detect depression.
Zhenyu Liu 0006, Changcong Li, Gang Wang 0012
BIBM4
2017 Eye movement pattern and mental retardation in depression
abstract
In order to explore the mental retardation of patients with depression, the visual search paradigm was used in this study. The emotional expression (happy, sad) and neutral expression were used as interdependent or search targets in this paradigm. Subjects were asked to search the target face form a matrix containing 16 emotional faces after watching the target face. The measurement indices of the search process - scanpath duration (SPD), scanpath length (SPL), convex hull area (CHA)-were collected and analyzed. The three indexes of patients were all larger than those in control group and there were significant differences between the groups. It can be seen that two kinds of emotional faces reduce the search efficiency of depression patients and there is a slow response and mental retardation phenomenon for depression in the search process.
Xingwang Liu, Shengfu Lu, Dachao Liu, Lei Feng 0005, Bingbing Fu, Gang Wang 0012, Ning Zhong 0001
BIBM8
2017 Episodic CAMN: Contextual Attention-Based Memory Networks with Iterative Feedback for Scene Labeling
abstract
Scene labeling can be seen as a sequence-sequence prediction task (pixels-labels), and it is quite important to leverage relevant context to enhance the performance of pixel classification. In this paper, we introduce an episodic attention-based memory network to achieve the goal. We present a unified framework that mainly consists of a Convolutional Neural Network (CNN), specifically, Fully Convolutional Network (FCN) and an attention-based memory module with feedback connections to perform context selection and refinement. The full model produces context-aware representation for each target patch by aggregating the activated context and its original local representation produced by the convolution layers. We evaluate our model on PASCAL Context, SIFT Flow and PASCAL VOC 2011 datasets and achieve competitive results to other state-of-the-art methods in scene labeling.
Abrar H. Abdulnabi, Bing Shuai, Gang Wang 0012
CVPR4
2017 Deep Level Sets for Salient Object Detection
abstract
Deep learning has been applied to saliency detection in recent years. The superior performance has proved that deep networks can model the semantic properties of salient objects. Yet it is difficult for a deep network to discriminate pixels belonging to similar receptive fields around the object boundaries, thus deep networks may output maps with blurred saliency and inaccurate boundaries. To tackle such an issue, in this work, we propose a deep Level Set network to produce compact and uniform saliency maps. Our method drives the network to learn a Level Set function for salient objects so it can output more accurate boundaries and compact saliency. Besides, to propagate saliency information among pixels and recover full resolution saliency map, we extend a superpixel-based guided filter to be a layer in the network. The proposed network has a simple structure and is trained end-to-end. During testing, the network can produce saliency maps by efficiently feedforwarding testing images at a speed over 12FPS on GPUs. Evaluations on benchmark datasets show that the proposed method achieves state-of-the-art performance.
Ping Hu 0001, Bing Shuai, Jun Liu 0036, Gang Wang 0012
CVPR4
2017 Global Context-Aware Attention LSTM Networks for 3D Action Recognition
abstract
Long Short-Term Memory (LSTM) networks have shown superior performance in 3D human action recognition due to their power in modeling the dynamics and dependencies in sequential data. Since not all joints are informative for action analysis and the irrelevant joints often bring a lot of noise, we need to pay more attention to the informative ones. However, original LSTM does not have strong attention capability. Hence we propose a new class of LSTM network, Global Context-Aware Attention LSTM (GCA-LSTM), for 3D action recognition, which is able to selectively focus on the informative joints in the action sequence with the assistance of global contextual information. In order to achieve a reliable attention representation for the action sequence, we further propose a recurrent attention mechanism for our GCA-LSTM network, in which the attention performance is improved iteratively. Experiments show that our end-to-end network can reliably focus on the most informative joints in each frame of the skeleton sequence. Moreover, our network yields state-of-the-art performance on three challenging datasets for 3D action recognition.
Jun Liu 0036, Gang Wang 0012, Ping Hu 0001, Ling-Yu Duan, Alex Chichung Kot
CVPR2
2017 An Empirical Study of Language CNN for Image Captioning
abstract
Language models based on recurrent neural networks have dominated recent image caption generation tasks. In this paper, we introduce a language CNN model which is suitable for statistical language modeling tasks and shows competitive performance in image captioning. In contrast to previous models which predict next word based on one previous word and hidden state, our language CNN is fed with all the previous words and can model the long-range dependencies in history words, which are critical for image captioning. The effectiveness of our approach is validated on two datasets: Flickr30K and MS COCO. Our extensive experimental results show that our method outperforms the vanilla recurrent neural network based language models and is competitive with the state-of-the-art methods.
Jiuxiang Gu, Gang Wang 0012, Jianfei Cai 0001, Tsuhan Chen
ICCV2
2017 Tracklet Association by Online Target-Specific Metric Learning and Coherent Dynamics Estimation
abstract
In this paper, we present a novel method based on online target-specific metric learning and coherent dynamics estimation for tracklet (track fragment) association by network flow optimization in long-term multi-person tracking. Our proposed framework aims to exploit appearance and motion cues to prevent identity switches during tracking and to recover missed detections. Furthermore, target-specific metrics (appearance cue) and motion dynamics (motion cue) are proposed to be learned and estimated online, i.e., during the tracking process. Our approach is effective even when such cues fail to identify or follow the target due to occlusions or object-to-object interactions. We also propose to learn the weights of these two tracking cues to handle the difficult situations, such as severe occlusions and object-to-object interactions effectively. Our method has been validated on several public datasets and the experimental results show that it outperforms several state-of-the-art tracking methods.
Bing Wang 0003, Gang Wang 0012, Kap Luk Chan, Li Wang 0057
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Simultaneous Feature and Dictionary Learning for Image Set Based Face Recognition
abstract
In this paper, we propose a simultaneous feature and dictionary learning (SFDL) method for image set-based face recognition, where each training and testing example contains a set of face images, which were captured from different variations of pose, illumination, expression, resolution, and motion. While a variety of feature learning and dictionary learning methods have been proposed in recent years and some of them have been successfully applied to image set-based face recognition, most of them learn features and dictionaries for facial image sets individually, which may not be powerful enough because some discriminative information for dictionary learning may be compromised in the feature learning stage if they are applied sequentially, and vice versa. To address this, we propose a SFDL method to learn discriminative features and dictionaries simultaneously from raw face pixels so that discriminative information from facial image sets can be jointly exploited by a one-stage learning procedure. To better exploit the nonlinearity of face samples from different image sets, we propose a deep SFDL (D-SFDL) method by jointly learning hierarchical non-linear transformations and class-specific dictionaries to further improve the recognition performance. Extensive experimental results on five widely used face data sets clearly shows that our SFDL and D-SFDL achieve very competitive or even better performance with the state-of-the-arts.
Jiwen Lu, Gang Wang 0012, Jie Zhou 0001
IEEE Trans. Image Process.2
2016 Recurrent Attentional Networks for Saliency Detection
abstract
Convolutional-deconvolution networks can be adopted to perform end-to-end saliency detection. But, they do not work well with objects of multiple scales. To overcome such a limitation, in this work, we propose a recurrent attentional convolutional-deconvolution network (RACDNN). Using spatial transformer and recurrent network units, RACDNN is able to iteratively attend to selected image sub-regions to perform saliency refinement progressively. Besides tackling the scale problem, RACDNN can also learn context-aware features from past iterations to enhance saliency refinement in future iterations. Experiments on several challenging saliency detection datasets validate the effectiveness of RACDNN, and show that RACDNN outperforms state-of-the-art saliency detection methods.
Jason Kuen, Zhenhua Wang 0002, Gang Wang 0012
CVPR3
2016 NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis
abstract
Recent approaches in depth-based human activity analysis achieved outstanding performance and proved the effectiveness of 3D representation for classification of action classes. Currently available depth-based and RGB+Dbased action recognition benchmarks have a number of limitations, including the lack of training samples, distinct class labels, camera views and variety of subjects. In this paper we introduce a large-scale dataset for RGB+D human action recognition with more than 56 thousand video samples and 4 million frames, collected from 40 distinct subjects. Our dataset contains 60 different action classes including daily, mutual, and health-related actions. In addition, we propose a new recurrent neural network structure to model the long-term temporal correlation of the features for each body part, and utilize them for better action classification. Experimental results show the advantages of applying deep learning methods over state-of-the-art handcrafted features on the suggested cross-subject and cross-view evaluation criteria for our dataset. The introduction of this large scale dataset will enable the community to apply, develop and adapt various data-hungry learning techniques for the task of depth-based and RGB+D-based human activity analysis.
Amir Shahroudy, Jun Liu 0036, Tian-Tsong Ng, Gang Wang 0012
CVPR4
2016 DAG-Recurrent Neural Networks for Scene Labeling
abstract
In image labeling, local representations for image units are usually generated from their surrounding image patches, thus long-range contextual information is not effectively encoded. In this paper, we introduce recurrent neural networks (RNNs) to address this issue. Specifically, directed acyclic graph RNNs (DAG-RNNs) are proposed to process DAG-structured images, which enables the network to model long-range semantic dependencies among image units. Our DAG-RNNs are capable of tremendously enhancing the discriminative power of local representations, which significantly benefits the local classification. Meanwhile, we propose a novel class weighting function that attends to rare classes, which phenomenally boosts the recognition accuracy for non-frequent classes. Integrating with convolution and deconvolution layers, our DAG-RNNs achieve new state-of-the-art results on the challenging SiftFlow, CamVid and Barcelona benchmarks.
Bing Shuai, Zhen Zuo, Bing Wang 0003, Gang Wang 0012
CVPR4
2016 Real-Time RGB-D Activity Prediction by Soft Regression
Jianfang Hu, Wei-Shi Zheng 0001, Lianyang Ma, Gang Wang 0012, Jian-Huang Lai
ECCV (1)4
2016 Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition
Jun Liu 0036, Amir Shahroudy, Dong Xu 0001, Gang Wang 0012
ECCV (3)4
2016 Gated Siamese Convolutional Neural Network Architecture for Human Re-identification
Rahul Rama Varior, Mrinal Haloi, Gang Wang 0012
ECCV (8)3
2016 A Siamese Long Short-Term Memory Architecture for Human Re-identification
Rahul Rama Varior, Bing Shuai, Jiwen Lu, Dong Xu 0001, Gang Wang 0012
ECCV (7)5
2016 Learning Common and Specific Features for RGB-D Semantic Segmentation with Deconvolutional Networks
Zhenhua Wang 0002, Dacheng Tao, Simon See, Gang Wang 0012
ECCV (5)5
2016 ROML: A Robust Feature Correspondence Approach for Matching Objects in A Set of Images
Kui Jia, Tsung-Han Chan, Shenghua Gao, Gang Wang 0012, Tianzhu Zhang 0001, Yi Ma 0001
Int. J. Comput. Vis.5
2016 Multimodal Multipart Learning for Action Recognition in Depth Videos
abstract
The articulated and complex nature of human actions makes the task of action recognition difficult. One approach to handle this complexity is dividing it to the kinetics of body parts and analyzing the actions based on these partial descriptors. We propose a joint sparse regression based learning method which utilizes the structured sparsity to model each action as a combination of multimodal features from a sparse set of body parts. To represent dynamics and appearance of parts, we employ a heterogeneous set of depth and skeleton based features. The proper structure of multimodal multipart features are formulated into the learning framework via the proposed hierarchical mixed norm, to regularize the structured features of each part and to apply sparsity between them, in favor of a group feature selection. Our experimental results expose the effectiveness of the proposed learning method in which it outperforms other methods in all three tested datasets while saturating one of them by achieving perfect accuracy.
Amir Shahroudy, Tian-Tsong Ng, Qingxiong Yang, Gang Wang 0012
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Exploring Local and Overall Ordinal Information for Robust Feature Description
abstract
This paper aims to build robust feature descriptors by exploring intensity order information in a patch. To this end, the local intensity order pattern (LIOP) and the overall intensity order pattern (OIOP) are proposed to effectively encode intensity order information of each pixel in different aspects. Specifically, LIOP captures the local ordinal information by using the intensity relationships among all the neighbouring sampling points around a pixel, while OIOP exploits the coarsely quantized overall intensity order of these sampling points. These two kinds of patterns are then separately aggregated into different ordinal bins, leading to two kinds of feature descriptors. Furthermore, as these two kinds of descriptors could encode complementary ordinal information, they are combined together to obtain a discriminative and compact mixed intensity order pattern descriptor. All these descriptors are constructed on the basis of relative relationships of intensities in a rotationally invariant way, making them be inherently invariant to image rotation and any monotonic intensity changes. Experimental results on image matching and object recognition are encouraging, demonstrating the superiorities of our descriptors over the state of the art.
Zhenhua Wang 0002, Bin Fan 0001, Gang Wang 0012, Fuchao Wu
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Localized Multifeature Metric Learning for Image-Set-Based Face Recognition
abstract
This paper presents a new approach to image-set-based face recognition, where each training and testing example is a set of face images captured from varying poses, illuminations, expressions, and resolutions. While a number of image set based face recognition methods have been proposed in recent years, most of them model each face image set as a single linear subspace or as the union of linear subspaces, which may lose some discriminative information for face image set representation. To address this shortcoming, we propose exploiting statistics information as feature representations for face image sets and develop a localized multikernel metric learning algorithm to effectively combine different statistics for recognition. Moreover, we propose a localized multikernel multimetric learning method to jointly learn multiple feature-specific distance metrics in the kernel spaces, one for each statistic feature, to better exploit complementary information for recognition. Our methods achieve state-of-the-art performance on four widely used video face datasets including the Honda, MoBo, YouTube Celebrities, and YouTube Face datasets.
Jiwen Lu, Gang Wang 0012, Pierre Moulin
IEEE Trans. Circuits Syst. Video Technol.2
2016 Scene Parsing With Integration of Parametric and Non-Parametric Models
abstract
We adopt convolutional neural networks (CNNs) to be our parametric model to learn discriminative features and classifiers for local patch classification. Based on the occurrence frequency distribution of classes, an ensemble of CNNs (CNN-Ensemble) are learned, in which each CNN component focuses on learning different and complementary visual patterns. The local beliefs of pixels are output by CNN-Ensemble. Considering that visually similar pixels are indistinguishable under local context, we leverage the global scene semantics to alleviate the local ambiguity. The global scene constraint is mathematically achieved by adding a global energy term to the labeling energy function, and it is practically estimated in a non-parametric framework. A large margin-based CNN metric learning method is also proposed for better global belief estimation. In the end, the integration of local and global beliefs gives rise to the class likelihood of pixels, based on which maximum marginal inference is performed to generate the label prediction maps. Even without any post-processing, we achieve the state-of-the-art results on the challenging SiftFlow and Barcelona benchmarks.
Bing Shuai, Zhen Zuo, Gang Wang 0012, Bing Wang 0003
IEEE Trans. Image Process.3
2016 Learning Invariant Color Features for Person Reidentification
abstract
Matching people across multiple camera views known as person reidentification is a challenging problem due to the change in visual appearance caused by varying lighting conditions. The perceived color of the subject appears to be different under different illuminations. Previous works use color as it is or address these challenges by designing color spaces focusing on a specific cue. In this paper, we propose an approach for learning color patterns from pixels sampled from images across two camera views. The intuition behind this work is that, even though varying lighting conditions across views affect the pixel values of the same color, the final representation of a particular color should be stable and invariant to these variations, i.e., they should be encoded with the same values. We model color feature generation as a learning problem by jointly learning a linear transformation and a dictionary to encode pixel values. We also analyze different photometric invariant color spaces as well as popular color constancy algorithm for person reidentification. Using color as the only cue, we compare our approach with all the photometric invariant color spaces and show superior performance over all of them. Combining with other learned low-level and high-level features, we obtain promising results in VIPeR, Person Re-ID 2011, and CAVIAR4REID data sets.
Rahul Rama Varior, Gang Wang 0012, Jiwen Lu, Ting Liu 0009
IEEE Trans. Image Process.2
2016 Learning Contextual Dependence With Convolutional Hierarchical Recurrent Neural Networks
abstract
Deep convolutional neural networks (CNNs) have shown their great success on image classification. CNNs mainly consist of convolutional and pooling layers, both of which are performed on local image areas without considering the dependence among different image regions. However, such dependence is very important for generating explicit image representation. In contrast, recurrent neural networks (RNNs) are well known for their ability of encoding contextual information in sequential data, and they only require a limited number of network parameters. Thus, we proposed the hierarchical RNNs (HRNNs) to encode the contextual dependence in image representation. In HRNNs, each RNN layer focuses on modeling spatial dependence among image regions from the same scale but different locations. While the cross RNN scale connections target on modeling scale dependencies among regions from the same location but different scales. Specifically, we propose two RNN models: 1) hierarchical simple recurrent network (HSRN), which is fast and has low computational cost and 2) hierarchical long-short term memory recurrent network, which performs better than HSRN with the price of higher computational cost. In this paper, we integrate CNNs with HRNNs, and develop end-to-end convolutional hierarchical RNNs (C-HRNNs) for image classification. C-HRNNs not only utilize the discriminative representation power of CNNs, but also utilize the contextual dependence learning ability of our HRNNs. On four of the most challenging object/scene image classification benchmarks, our C-HRNNs achieve the state-of-the-art results on Places 205, SUN 397, and MIT indoor, and the competitive results on ILSVRC 2012.
Zhen Zuo, Bing Shuai, Gang Wang 0012, Bing Wang 0003, Yushi Chen 0002
IEEE Trans. Image Process.3
2016 Object Instance Search in Videos via Spatio-Temporal Trajectory Discovery
abstract
Given a specific object as query, object instance search aims to not only retrieve the images or frames that contain the query, but also locate all its occurrences. In this work, we explore the use of spatio-temporal cues to improve the quality of object instance search from videos. To this end, we formulate this problem as the spatio-temporal trajectory search problem, where a trajectory is a sequence of bounding boxes that locate the object instance in each frame. The goal is to find the top- K trajectories that are likely to contain the target object. Despite the large number of trajectory candidates, we build on a recent spatio- temporal search algorithm for event detection to efficiently find the optimal spatio- temporal trajectories in large video volumes , with complexity linear to the video volume size. We solve the key bottleneck in applying this approach to object instance search by leveraging a randomized approach to enable fast scoring of any bounding boxes in the video volume. In addition , we present a new dataset for video object instance search. Experimental results on a 73-hour video dataset demonstrate that our approach improves the performance of video object instance search and localization over the state-of-the-art search and tracking methods.
Jingjing Meng, Junsong Yuan 0001, Gang Wang 0012, Yap-Peng Tan
IEEE Trans. Multim.4
2015 Deep hashing for compact binary codes learning
abstract
In this paper, we propose a new deep hashing (DH) approach to learn compact binary codes for large scale visual search. Unlike most existing binary codes learning methods which seek a single linear projection to map each sample into a binary vector, we develop a deep neural network to seek multiple hierarchical non-linear transformations to learn these binary codes, so that the nonlinear relationship of samples can be well exploited. Our model is learned under three constraints at the top layer of the deep network: 1) the loss between the original real-valued feature descriptor and the learned binary vector is minimized, 2) the binary codes distribute evenly on each bit, and 3) different bits are as independent as possible. To further improve the discriminative power of the learned binary codes, we extend DH into supervised DH (SDH) by including one discriminative term into the objective function of DH which simultaneously maximizes the inter-class variations and minimizes the intra-class variations of the learned binary codes. Experimental results show the superiority of the proposed approach over the state-of-the-arts.
Venice Erin Liong, Jiwen Lu, Gang Wang 0012, Pierre Moulin, Jie Zhou 0001
CVPR3
2015 Real-time part-based visual tracking via adaptive correlation filters
abstract
Robust object tracking is a challenging task in computer vision. To better solve the partial occlusion issue, part-based methods are widely used in visual object trackers. However, due to the complicated online training and updating process, most of these part-based trackers cannot run in real-time. Correlation filters have been used in tracking tasks recently because of the high efficiency. However, the conventional correlation filter based trackers cannot deal with occlusion. Furthermore, most correlation filter based trackers fix the scale and rotation of the target which makes the trackers unreliable in long-term tracking tasks. In this paper, we propose a novel tracking method which track objects based on parts with multiple correlation filters. Our method can run in real-time. Additionally, the Bayesian inference framework and a structural constraint mask are adopted to enable our tracker to be robust to various appearance changes. Extensive experiments have been done to prove the effectiveness of our method.
Ting Liu 0009, Gang Wang 0012, Qingxiong Yang
CVPR2
2015 Multi-manifold deep metric learning for image set classification
abstract
In this paper, we propose a multi-manifold deep metric learning (MMDML) method for image set classification, which aims to recognize an object of interest from a set of image instances captured from varying viewpoints or under varying illuminations. Motivated by the fact that manifold can be effectively used to model the nonlinearity of samples in each image set and deep learning has demonstrated superb capability to model the nonlinearity of samples, we propose a MMDML method to learn multiple sets of nonlinear transformations, one set for each object class, to nonlinearly map multiple sets of image instances into a shared feature subspace, under which the manifold margin of different class is maximized, so that both discriminative and class-specific information can be exploited, simultaneously. Our method achieves the state-of-the-art performance on five widely used datasets.
Jiwen Lu, Gang Wang 0012, Weihong Deng, Pierre Moulin, Jie Zhou 0001
CVPR2
2015 Integrating parametric and non-parametric models for scene labeling
abstract
We adopt Convolutional Neural Networks (CNN) as our parametric model to learn discriminative features and classifiers for local patch classification. As visually similar pixels are indistinguishable from local context, we alleviate such ambiguity by introducing a global scene constraint. We estimate the global potential in a non-parametric framework. Furthermore, a large margin based CNN metric learning method is proposed for better global potential estimation. The final pixel class prediction is performed by integrating local and global beliefs. Even without any post-processing, we achieve state-of-the-art performance on SiftFlow and competitive results on Stanford Background benchmark.
Bing Shuai, Gang Wang 0012, Zhen Zuo, Bing Wang 0003, Lifan Zhao
CVPR2
2015 Visual tracking using learned color features
abstract
Robust object tracking is a challenging task in computer vision. Color features have been popularly used in visual tracking. However, most conventional color-based trackers either rely on luminance information or use simple color representations for image description. During the tracking sequences, the perceived color of the target may change because of the varying lighting conditions. In this paper, we learn the color patterns offline from pixels sampled from images across different camera views. In the new color feature space, the proposed tracking method performs robustly in various environment. The new color feature space is learned by learning a linear transformation and a dictionary to encode pixel values. To speedup the feature extraction, we use the marginal regression to calculate the sparse feature codes. Experimental results demonstrate that significant improvement can be achieved by using our learned color features, especially on the video sequences with complicated lighting conditions.
Ting Liu 0009, Rahul Rama Varior, Gang Wang 0012
ICASSP3
2015 A data-driven color feature learning scheme for image retrieval
abstract
This paper addresses content based image retrieval based on color features. Several previous works have addressed color based image retrieval based on hand-crafted features. In this paper, a data-driven learning framework is proposed for generating color based signatures. To obtain the features, a linear transformation is learned from the pixel values based on its reconstruction error. Using this linear transformation, the original pixel values are transformed into a higher dimensional space. In the higher dimensional space, a dictionary is learned to obtain the sparse codes of the pixels. A max pooling strategy is used to obtain the dominant color features of a region and the final feature vector for an image is obtained by concatenating the pooled features. We evaluate our approach following the standard evaluation criteria for the INRIA Holidays and University of Kentucky Benchmark datasets. The approach is compared with several baselines such as histograms in RGB, HSV, YUV and Lab color spaces and several other color based features proposed for addressing this problem. Our approach shows competitive results on these datasets and outperforms all the baselines.
Rahul Rama Varior, Gang Wang 0012
ICASSP2
2015 Fast object instance search in videos from one example
abstract
We present an efficient approach to search for and locate all occurrences of a specific object in large video volumes, given a single query example. Locations of object occurrences are returned as spatio-temporal trajectories in the 3D video volume. Despite much work on object instance search in image datasets, these methods locate the object independently in each image, therefore do not preserve the spatio-temporal consistency in consecutive video frames. This results in sub-optimal performance if directly applied to videos, as will be shown in our experiments. We propose to locate the object jointly across video frames using spatio-temporal search. The efficiency and effectiveness of the proposed approach is demonstrated on a consumer video dataset consisting of crawled YouTube videos and mobile captured consumer clips. Our method significantly improves the localized search accuracy over the baseline, which treats each frame independently. Moreover, it is able to find the top 100 object trajectories in the 5.5-hour dataset within 30 seconds.
Jingjing Meng, Junsong Yuan 0001, Yap-Peng Tan, Gang Wang 0012
ICIP4
2015 Exemplar based Deep Discriminative and Shareable Feature Learning for scene image classification
Zhen Zuo, Gang Wang 0012, Bing Shuai, Lifan Zhao, Qingxiong Yang
Pattern Recognit.2
2015 Visual Tracking via Temporally Smooth Sparse Coding
abstract
Sparse representation has been popular in visual tracking recently for its robustness and accuracy. However, for most conventional sparse coding based trackers, the target candidates are considered independently between consecutive frames. This paper shows that the temporal correlation of these frames can be exploited to improve the performance of tracking and makes the tracker more robust to noise. Furthermore, to improve the tracking speed, we revisit a more efficient method for ℓ1norm problem, marginal regression, which can solve the sparse coding problem more efficiently. Consequently we can realize real-time tracking based on the temporal smooth sparse representation. Extensive experiments have been done to demonstrate the effectiveness and efficiency of our method.
Ting Liu 0009, Gang Wang 0012, Li Wang 0057, Kap Luk Chan
IEEE Signal Process. Lett.2
2015 Quaddirectional 2D-Recurrent Neural Networks For Image Labeling
abstract
We adopt Convolutional Neural Networks (CNN) to learn discriminative features for local patch classification. We further introduce quaddirectional 2D Recurrent Neural Networks to model the long range dependencies among pixels. Our quaddirectional 2D-RNN is able to embed the global image context into the compact local representation, which significantly enhance their discriminative power. Our experiments demonstrate that the integration of CNN and quaddirectional 2D-RNN achieves very promising results which are comparable to state-of-the-art on real-world image labeling benchmarks.
Bing Shuai, Zhen Zuo, Gang Wang 0012
IEEE Signal Process. Lett.3
2015 Joint Feature Learning for Face Recognition
abstract
This paper presents a new joint feature learning (JFL) approach to automatically learn feature representation from raw pixels for face recognition. Unlike many existing face recognition systems, where conventional feature descriptors, such as local binary patterns and Gabor features, are used for face representation, we propose an unsupervised feature learning method to learn hierarchical feature representation. Since different face regions have different physical characteristics, we propose to use different feature dictionaries to represent them, and to learn multiple yet related feature projection matrices for these regions simultaneously. Hence position-specific discriminative information can be exploited for face representation. Having learned these feature projections for different face regions, we perform spatial pooling for face patches within each region to enhance the representative power of the learned features. Moreover, we stack our JFL model into a deep architecture to exploit hierarchical information for feature representation and further improve the recognition performance. Experimental results on five widely used face data sets show the effectiveness of our proposed approach.
Jiwen Lu, Venice Erin Liong, Gang Wang 0012, Pierre Moulin
IEEE Trans. Inf. Forensics Secur.3
2015 Reconstruction-Based Metric Learning for Unconstrained Face Verification
abstract
In this paper, we propose a reconstruction-based metric learning method to learn a discriminative distance metric for unconstrained face verification. Unlike conventional metric learning methods, which only consider the label information of training samples and ignore the reconstruction residual information in the learning procedure, we apply a reconstruction criterion to learn a discriminative distance metric. For each training example, the distance metric is learned by enforcing a margin between the interclass sparse reconstruction residual and interclass sparse reconstruction residual, so that the reconstruction residual of training samples can be effectively exploited to compute the between-class and within-class variations. To better use multiple features for distance metric learning, we propose a reconstruction-based multimetric learning method to collaboratively learn multiple distance metrics, one for each feature descriptor, to remove uncorrelated information for recognition. We evaluate our proposed methods on the Labelled Faces in the Wild (LFW) and YouTube face data sets and our experimental results clearly show the superiority of our methods over both previous metric learning methods and several state-of-the-art unconstrained face verification methods.
Jiwen Lu, Gang Wang 0012, Weihong Deng, Kui Jia
IEEE Trans. Inf. Forensics Secur.2
2015 Unsupervised Joint Feature Learning and Encoding for RGB-D Scene Labeling
abstract
Most existing approaches for RGB-D indoor scene labeling employ hand-crafted features for each modality independently and combine them in a heuristic manner. There has been some attempt on directly learning features from raw RGB-D data, but the performance is not satisfactory. In this paper, we propose an unsupervised joint feature learning and encoding (JFLE) framework for RGB-D scene labeling. The main novelty of our learning framework lies in the joint optimization of feature learning and feature encoding in a coherent way, which significantly boosts the performance. By stacking basic learning structure, higher level features are derived and combined with lower level features for better representing RGB-D data. Moreover, to explore the nonlinear intrinsic characteristic of data, we further propose a more general joint deep feature learning and encoding (JDFLE) framework that introduces the nonlinear mapping into JFLE. The experimental results on the benchmark NYU depth dataset show that our approaches achieve competitive performance, compared with the state-of-the-art methods, while our methods do not need complex feature handcrafting and feature combination and can be easily applied to other data sets.
Anran Wang 0001, Jiwen Lu, Jianfei Cai 0001, Gang Wang 0012, Tat-Jen Cham
IEEE Trans. Image Process.4
2015 Video Tracking Using Learned Hierarchical Features
abstract
In this paper, we propose an approach to learn hierarchical features for visual object tracking. First, we offline learn features robust to diverse motion patterns from auxiliary video sequences. The hierarchical features are learned via a two-layer convolutional neural network. Embedding the temporal slowness constraint in the stacked architecture makes the learned features robust to complicated motion transformations, which is important for visual object tracking. Then, given a target video sequence, we propose a domain adaptation module to online adapt the pre-learned features according to the specific target object. The adaptation is conducted in both layers of the deep feature learning module so as to include appearance information of the specific target object. As a result, the learned hierarchical features can be robust to both complicated motion transformations and appearance changes of target objects. We integrate our feature learning algorithm into three tracking methods. Experimental results demonstrate that significant improvement can be achieved using our learned hierarchical features, especially on video sequences with complicated motion transformations.
Li Wang 0057, Ting Liu 0009, Gang Wang 0012, Kap Luk Chan, Qingxiong Yang
IEEE Trans. Image Process.3
2015 Large-Margin Multi-Modal Deep Learning for RGB-D Object Recognition
abstract
Most existing feature learning-based methods for RGB-D object recognition either combine RGB and depth data in an undifferentiated manner from the outset, or learn features from color and depth separately, which do not adequately exploit different characteristics of the two modalities or utilize the shared relationship between the modalities. In this paper, we propose a general CNN-based multi-modal learning framework for RGB-D object recognition. We first construct deep CNN layers for color and depth separately, which are then connected with a carefully designed multi-modal layer. This layer is designed to not only discover the most discriminative features for each modality, but is also able to harness the complementary relationship between the two modalities. The results of the multi-modal layer are back-propagated to update parameters of the CNN layers, and the multi-modal feature learning and the back-propagation are iteratively performed until convergence. Experimental results on two widely used RGB-D object datasets show that our method for general multi-modal learning achieves comparable performance to state-of-the-art methods specifically designed for RGB-D data.
Anran Wang 0001, Jiwen Lu, Jianfei Cai 0001, Tat-Jen Cham, Gang Wang 0012
IEEE Trans. Multim.5
2015 Multi-Task CNN Model for Attribute Prediction
abstract
This paper proposes a joint multi-task learning algorithm to better predict attributes in images using deep convolutional neural networks (CNN). We consider learning binary semantic attributes through a multi-task CNN model, where each CNN will predict one binary attribute. The multi-task learning allows CNN models to simultaneously share visual knowledge among different attribute categories. Each CNN will generate attribute-specific feature representations, and then we apply multi-task learning on the features to predict their attributes. In our multi-task framework, we propose a method to decompose the overall model's parameters into a latent task matrix and combination matrix. Furthermore, under-sampled classifiers can leverage shared statistics from other classifiers to improve their performance. Natural grouping of attributes is applied such that attributes in the same group are encouraged to share more knowledge. Meanwhile, attributes in different groups will generally compete with each other, and consequently share less knowledge. We show the effectiveness of our method on two popular attribute datasets.
Abrar H. Abdulnabi, Gang Wang 0012, Jiwen Lu, Kui Jia
IEEE Trans. Multim.2
2014 DL-SFA: Deeply-Learned Slow Feature Analysis for Action Recognition
abstract
Most of the previous work on video action recognition use complex hand-designed local features, such as SIFT, HOG and SURF, but these approaches are implemented sophisticatedly and difficult to be extended to other sensor modalities. Recent studies discover that there are no universally best hand-engineered features for all datasets, and learning features directly from the data may be more advantageous. One such endeavor is Slow Feature Analysis (SFA) proposed by Wiskott and Sejnowski [33]. SFA can learn the invariant and slowly varying features from input signals and has been proved to be valuable in human action recognition [34]. It is also observed that the multi-layer feature representation has succeeded remarkably in widespread machine learning applications. In this paper, we propose to combine SFA with deep learning techniques to learn hierarchical representations from the video data itself. Specifically, we use a two-layered SFA learning structure with 3D convolution and max pooling operations to scale up the method to large inputs and capture abstract and structural features from the video. Thus, the proposed method is suitable for action recognition. At the same time, sharing the same merits of deep learning, the proposed method is generic and fully automated. Our classification results on Hollywood2, KTH and UCF Sports are competitive with previously published results. To highlight some, on the KTH dataset, our recognition rate shows approximately 1% improvement in comparison to state-of-the-art methods even without supervision or dense sampling.
Lin Sun 0004, Kui Jia, Tsung-Han Chan, Yuqiang Fang, Gang Wang 0012, Shuicheng Yan
CVPR5
2014 Tracklet Association with Online Target-Specific Metric Learning
abstract
This paper presents a novel introduction of online target-specific metric learning in track fragment (tracklet) association by network flow optimization for long-term multi-person tracking. Different from other network flow formulation, each node in our network represents a tracklet, and each edge represents the likelihood of neighboring tracklets belonging to the same trajectory as measured by our proposed affinity score. In our method, target-specific similarity metrics are learned, which give rise to the appearance-based models used in the tracklet affinity estimation. Trajectory-based tracklets are refined by using the learned metrics to account for appearance consistency and to identify reliable tracklets. The metrics are then re-learned using reliable tracklets for computing tracklet affinity scores. Long-term trajectories are then obtained through network flow optimization. Occlusions and missed detections are handled by a trajectory completion step. Our method is effective for long-term tracking even when the targets are spatially close or completely occluded by others. We validate our proposed framework on several public datasets and show that it outperforms several state of art methods.
Bing Wang 0003, Gang Wang 0012, Kap Luk Chan, Li Wang 0057
CVPR2
2014 Spatiotemporal Background Subtraction Using Minimum Spanning Tree and Optical Flow
Qingxiong Yang, Qing Li 0001, Gang Wang 0012, Ming-Hsuan Yang 0001
ECCV (7)4
2014 Simultaneous Feature and Dictionary Learning for Image Set Based Face Recognition
Jiwen Lu, Gang Wang 0012, Weihong Deng, Pierre Moulin
ECCV (1)2
2014 Multi-modal Unsupervised Feature Learning for RGB-D Scene Labeling
Anran Wang 0001, Jiwen Lu, Gang Wang 0012, Jianfei Cai 0001, Tat-Jen Cham
ECCV (5)3
2014 Learning Discriminative and Shareable Features for Scene Classification
Zhen Zuo, Gang Wang 0012, Bing Shuai, Lifan Zhao, Qingxiong Yang, Xudong Jiang 0001
ECCV (1)2
2014 Pedestrian detection in highly crowded scenes using "online" dictionary learning for occlusion handling
abstract
Pedestrian detection is one of the most important task for video analytics of an intelligent surveillance system. In this paper, we propose a framework to improve the detection performance of a generic pedestrian detector for highly crowded scenes. The generic offline-trained pedestrian detectors usually cannot handle the problem of detecting pedestrians in highly crowded scenes due to the severe mutual occlusions of the pedestrians. In our approach, we firstly enhance the head detection and suppress the detections of other body parts in the deformable part-based model because the heads of pedestrians less likely to be occluded in highly crowded scenes. Then we propose to utilize multiple-instance dictionary learning to refine the previous detection responses. Compared to other related work, our approach builds a data-adaptive dictionary (codebook) for the heads of pedestrians, hence it can better handle the problem of detecting pedestrians in highly crowded scenes. The experiments on three datasets containing video clips of crowded scenes demonstrated the effectiveness of our proposed approach, significantly improving the state-of-the-art detector.
Bing Wang 0003, Kap Luk Chan, Gang Wang 0012
ICIP3
2014 Learning deep features for multiple object tracking by using a multi-task learning strategy
abstract
Model-free object tracking is still challenging because of the limited prior knowledge and the unexpected variation of the target object. In this paper, we propose a feature learning algorithm for model-free multiple object tracking. First, we pre-learn generic features invariant to diverse motion transformations from auxiliary video data by using a deep network of anto-encoder. Then, we adapt the pre-learned features according to multiple target objects respectively in a multi-task learning manner. We treat the feature adaptation for each target object as one single task. We simultaneously learn the common feature shared by all target objects and the individual feature of each object. Experimental results demonstrate that our feature learning algorithm can significantly improve multiple object tracking performance.
Li Wang 0057, Nam Trung Pham, Tian-Tsong Ng, Gang Wang 0012, Kap Luk Chan, Karianto Leman
ICIP4
2014 Estimating spatial layout of rooms from RGB-D videos
abstract
Spatial layout estimation of indoor rooms plays an important role in many visual analysis applications such as robotics and human-computer interaction. While many methods have been proposed for recovering spatial layout of rooms in recent years, their performance is still far from satisfactory due to high occlusion caused by the presence of objects that clutter the scene. In this paper, we propose a new approach to estimate the spatial layout of rooms from RGB-D videos. Unlike most existing methods which estimate the layout from still images, RGB-D videos provide more spatial-temporal and depth information, which are helpful to improve the estimation performance because more contextual information can be exploited in RGB-D videos. Given a RGB-D video, we first estimate the spatial layout of the scene in each single frame and compute the camera trajectory using the simultaneous localization and mapping (SLAM) algorithm. Then, the estimated spatial layouts of different frames are integrated to infer temporally consistent layouts of the room throughout the whole video. Our method is evaluated on the NYU RGB-D dataset, and the experimental results show the efficacy of the proposed approach.
Anran Wang 0001, Jiwen Lu, Jianfei Cai 0001, Gang Wang 0012, Tat-Jen Cham
MMSP4
2014 Optimizing LBP Structure For Visual Recognition Using Binary Quadratic Programming
abstract
Local binary pattern (LBP) and its variants have shown promising results in visual recognition applications. However, most existing approaches rely on a pre-defined structure to extract LBP features. We argue that the optimal LBP structure should be task-dependent and propose a new method to learn discriminative LBP structures. We formulate it as a point selection problem: Given a set of point candidates, the goal is to select an optimal subset to compose the LBP structure. In view of the problems of current feature selection algorithms, we propose a novel Maximal Joint Mutual Information criterion. Then, the point selection is converted into a binary quadratic programming problem and solved efficiently via the branch and bound algorithm. The proposed LBP structures demonstrate superior performance to the state-of-the-art approaches on classifying both spatial patterns in scene recognition and spatial-temporal patterns in dynamic texture recognition.
Jianfeng Ren, Xudong Jiang 0001, Junsong Yuan 0001, Gang Wang 0012
IEEE Signal Process. Lett.4
2014 Human Identity and Gender Recognition From Gait Sequences With Arbitrary Walking Directions
abstract
We investigate the problem of human identity and gender recognition from gait sequences with arbitrary walking directions. Most current approaches make the unrealistic assumption that persons walk along a fixed direction or a pre-defined path. Given a gait sequence collected from arbitrary walking directions, we first obtain human silhouettes by background subtraction and cluster them into several clusters. For each cluster, we compute the cluster-based averaged gait image as features. Then, we propose a sparse reconstruction based metric learning method to learn a distance metric to minimize the intra-class sparse reconstruction errors and maximize the inter-class sparse reconstruction errors simultaneously, so that discriminative information can be exploited for recognition. The experimental results show the efficacy of our approach.
Jiwen Lu, Gang Wang 0012, Pierre Moulin
IEEE Trans. Inf. Forensics Secur.2
2014 Tree Filtering: Efficient Structure-Preserving Smoothing With a Minimum Spanning Tree
abstract
We present a new efficient edge-preserving filter-"tree filter"-to achieve strong image smoothing. The proposed filter can smooth out high-contrast details while preserving major edges, which is not achievable for bilateral-filter-like techniques. Tree filter is a weighted-average filter, whose kernel is derived by viewing pixel affinity in a probabilistic framework simultaneously considering pixel spatial distance, color/intensity difference, as well as connectedness. Pixel connectedness is acquired by treating pixels as nodes in a minimum spanning tree (MST) extracted from the image. The fact that an MST makes all image pixels connected through the tree endues the filter with the power to smooth out high-contrast, fine-scale details while preserving major image structures, since pixels in small isolated region will be closely connected to surrounding majority pixels through the tree, while pixels inside large homogeneous region will be automatically dragged away from pixels outside the region. The tree filter can be separated into two other filters, both of which turn out to have fast algorithms. We also propose an efficient linear time MST extraction algorithm to further improve the whole filtering speed. The algorithms give tree filter a great advantage in low computational complexity (linear to number of image pixels) and fast speed: it can process a 1-megapixel 8-bit image at ~ 0.25 s on an Intel 3.4 GHz Core i7 CPU (including the construction of MST). The proposed tree filter is demonstrated on a variety of applications.
Linchao Bao, Yibing Song, Qingxiong Yang, Gang Wang 0012
IEEE Trans. Image Process.5
2014 Nonnegative Tensor Cofactorization and Its Unified Solution
abstract
In this paper, we present a new joint factorization algorithm, called Nonnegative Tensor Co-Factorization (NTCoF). The key idea is to simultaneously factorize multiple visual features of the same data into nonnegative dimensionality-reduced representations, and meanwhile, to maximize the correlations of the low-dimensional representations. The data is generally encoded as tensors of arbitrary order, rather than vectors, to preserve the original data structures. NTCoF provides a simple and efficient way to fuse multiple complementary features for enhancing the discriminative power of the desired rank-reduced representations under the nonnegative constraints. We formulate the related objectives with a block-wise quadratic nonnegative function. To optimize, a unified convergence provable solution is developed. This solution is applicable for any nonnegative optimization problems with block-wise quadratic objective functions, and thus offer an unified platform based on which specific solution can be directly derived by skipping over tedious proof about algorithmic convergence. We apply the proposed algorithm and solution on three image tasks, face recognition, multi-class image categorization and multi-label image annotation. Results with comparisons on public challenging datasets show that the proposed algorithm can outperform both the traditional nonnegative methods and the popular feature combination methods.
Xiaobai Liu, Shuicheng Yan, Gang Wang 0012, Hai Jin 0001, Seong-Whan Lee
IEEE Trans. Image Process.4
2013 Image Set Classification Using Holistic Multiple Order Statistics Features and Localized Multi-kernel Metric Learning
abstract
This paper presents a new approach for image set classification, where each training and testing example contains a set of image instances of an object captured from varying viewpoints or under varying illuminations. While a number of image set classification methods have been proposed in recent years, most of them model each image set as a single linear subspace or mixture of linear subspaces, which may lose some discriminative information for classification. To address this, we propose exploring multiple order statistics as features of image sets, and develop a localized multi-kernel metric learning (LMKML) algorithm to effectively combine different order statistics information for classification. Our method achieves the state-of-the-art performance on four widely used databases including the Honda/UCSD, CMU Mobo, and Youtube face datasets, and the ETH-80 object dataset.
Jiwen Lu, Gang Wang 0012, Pierre Moulin
ICCV2
2013 Learning to Share Latent Tasks for Action Recognition
abstract
Sharing knowledge for multiple related machine learning tasks is an effective strategy to improve the generalization performance. In this paper, we investigate knowledge sharing across categories for action recognition in videos. The motivation is that many action categories are related, where common motion pattern are shared among them (e.g. diving and high jump share the jump motion). We propose a new multi-task learning method to learn latent tasks shared across categories, and reconstruct a classifier for each category from these latent tasks. Compared to previous methods, our approach has two advantages: (1) The learned latent tasks correspond to basic motion patterns instead of full actions, thus enhancing discrimination power of the classifiers. (2) Categories are selected to share information with a sparsity regularizer, avoiding falsely forcing all categories to share knowledge. Experimental results on multiple public data sets show that the proposed approach can effectively transfer knowledge between different action categories to improve the performance of conventional single task learning methods.
Gang Wang 0012, Kui Jia
ICCV2
2013 Discriminative Multimanifold Analysis for Face Recognition from a Single Training Sample per Person
abstract
Conventional appearance-based face recognition methods usually assume that there are multiple samples per person (MSPP) available for discriminative feature extraction during the training phase. In many practical face recognition applications such as law enhancement, e-passport, and ID card identification, this assumption, however, may not hold as there is only a single sample per person (SSPP) enrolled or recorded in these systems. Many popular face recognition methods fail to work well in this scenario because there are not enough samples for discriminant learning. To address this problem, we propose in this paper a novel discriminative multimanifold analysis (DMMA) method by learning discriminative features from image patches. First, we partition each enrolled face image into several nonoverlapping patches to form an image set for each sample per person. Then, we formulate the SSPP face recognition as a manifold-manifold matching problem and learn multiple DMMA feature spaces to maximize the manifold margins of different persons. Finally, we present a reconstruction-based manifold-manifold distance to identify the unlabeled subjects. Experimental results on three widely used face databases are presented to demonstrate the efficacy of the proposed approach.
Jiwen Lu, Yap-Peng Tan, Gang Wang 0012
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Improved Object Categorization and Detection Using Comparative Object Similarity
abstract
Due to the intrinsic long-tailed distribution of objects in the real world, we are unlikely to be able to train an object recognizer/detector with many visual examples for each category. We have to share visual knowledge between object categories to enable learning with few or no training examples. In this paper, we show that local object similarity information--statements that pairs of categories are similar or dissimilar--is a very useful cue to tie different categories to each other for effective knowledge transfer. The key insight: Given a set of object categories which are similar and a set of categories which are dissimilar, a good object model should respond more strongly to examples from similar categories than to examples from dissimilar categories. To exploit this category-dependent similarity regularization, we develop a regularized kernel machine algorithm to train kernel classifiers for categories with few or no training examples. We also adapt the state-of-the-art object detector to encode object similarity constraints. Our experiments on hundreds of categories from the Labelme dataset show that our regularized kernel classifiers can make significant improvement on object categorization. We also evaluate the improved object detector on the PASCAL VOC 2007 benchmark dataset.
Gang Wang 0012, David A. Forsyth, Derek Hoiem
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Image-to-Set Face Recognition Using Locality Repulsion Projections and Sparse Reconstruction-Based Similarity Measure
abstract
For many practical face recognition systems such as law enforcement, e-passport, and ID card identification, there is usually only a single sample per person (SSPP) enrolled in these systems, and many existing face recognition methods may fail to work well because there are not enough samples for discriminative feature extraction in this scenario. However, the probe samples of these face recognition systems are usually captured on the spot, and it is possible to collect multiple face images per person for on-location probing, which is potentially useful to improve the recognition performance. In this paper, we propose a method based on locality repulsion projections (LRP) and a sparse reconstruction-based similarity measure (SRSM) to address the problem of SSPP face recognition using multiple probe images. The LRP method is motivated by our observation that similar face images from different people may lie in a locality in the feature space and cause misclassifications. We design the method with the aim of separating the samples of different classes within a neighborhood through subspace projections for easier classification. To better characterize the similarity between each gallery face and the probe image set, we propose a SRSM method for assigning a label to each probe image set. Experimental results on five widely used face datasets are presented to demonstrate the effectiveness of the proposed approach.
Jiwen Lu, Yap-Peng Tan, Gang Wang 0012, Gao Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2013 Fusion of Median and Bilateral Filtering for Range Image Upsampling
abstract
We present a new upsampling method to enhance the spatial resolution of depth images. Given a low-resolution depth image from an active depth sensor and a potentially high-resolution color image from a passive RGB camera, we formulate it as an adaptive cost aggregation problem and solve it using the bilateral filter. The formulation synergistically combines the median and bilateral filters thus it better preserves the depth edges and is more robust to noise. Numerical and visual evaluations on a total of 37 Middlebury data sets demonstrate the effectiveness of our method. A real-time high-resolution depth capturing system is also developed using commercial active depth sensor based on the proposed upsampling method.
Qingxiong Yang, Narendra Ahuja, Ruigang Yang, Kar-Han Tan, James Davis 0001, W. Bruce Culbertson, John G. Apostolopoulos, Gang Wang 0012
IEEE Trans. Image Process.8
2013 Reinforced Similarity Integration in Image-Rich Information Networks
abstract
Social multimedia sharing and hosting websites, such as Flickr and Facebook, contain billions of user-submitted images. Popular Internet commerce websites such as Amazon.com are also furnished with tremendous amounts of product-related images. In addition, images in such social networks are also accompanied by annotations, comments, and other information, thus forming heterogeneous image-rich information networks. In this paper, we introduce the concept of (heterogeneous) image-rich information network and the problem of how to perform information retrieval and recommendation in such networks. We propose a fast algorithm heterogeneous minimum order k-SimRank (HMok-SimRank) to compute link-based similarity in weighted heterogeneous information networks. Then, we propose an algorithm Integrated Weighted Similarity Learning (IWSL) to account for both link-based and content-based similarities by considering the network structure and mutually reinforcing link similarity and feature weight learning. Both local and global feature learning methods are designed. Experimental results on Flickr and Amazon data sets show that our approach is significantly better than traditional methods in terms of both relevance and speed. A new product search and recommendation system for e-commerce has been implemented based on our algorithm.
Xin Jin 0001, Jiebo Luo 0001, Jie Yu 0001, Gang Wang 0012, Dhiraj Joshi, Jiawei Han 0001
IEEE Trans. Knowl. Data Eng.4
2012 Neighborhood repulsed metric learning for kinship verification
abstract
Kinship verification from facial images is a challenging problem in computer vision, and there is a very few attempts on tackling this problem in the literature. In this paper, we propose a new neighborhood repulsed metric learning (NRML) method for kinship verification. Motivated by the fact that interclass samples (without kinship relations) with higher similarity usually lie in a neighborhood and are more easily misclassified than those with lower similarity, we aim to learn a distance metric under which the intraclass samples (with kinship relations) are pushed as close as possible and interclass samples lying in a neighborhood are repulsed and pulled as far as possible, simultaneously, such that more discriminative information can be exploited for verification. Moreover, we propose a multiview NRM-L (MNRML) method to seek a common distance metric to make better use of multiple feature descriptors to further improve the verification performance. Experimental results are presented to demonstrate the efficacy of the proposed methods.
Jiwen Lu, Junlin Hu 0001, Xiuzhuang Zhou, Yap-Peng Tan, Gang Wang 0012
CVPR6
2012 Learning to Recognize Unsuccessful Activities Using a Two-Layer Latent Structural Model
Gang Wang 0012
ECCV (3)2
2012 Gait-based gender classification in unconstrained environments
Jiwen Lu, Gang Wang 0012, Thomas S. Huang
ICPR2
2012 Learning Image Similarity from Flickr Groups Using Fast Kernel Machines
abstract
Measuring image similarity is a central topic in computer vision. In this paper, we propose to measure image similarity by learning from the online Flickr image groups. We do so by: Choosing 103 Flickr groups, building a one-versus-all multiclass classifier to classify test images into a group, taking the set of responses of the classifiers as features, calculating the distance between feature vectors to measure image similarity. Experimental results on the Corel dataset and the PASCAL VOC 2007 dataset show that our approach performs better on image matching, retrieval, and classification than using conventional visual features. To build our similarity measure, we need one-versus-all classifiers that are accurate and can be trained quickly on very large quantities of data. We adopt an SVM classifier with a histogram intersection kernel. We describe a novel fast training algorithm for this classifier: the Stochastic Intersection Kernel MAchine (SIKMA) training algorithm. This method can produce a kernel classifier that is more accurate than a linear classifier on tens of thousands of examples in minutes.
Gang Wang 0012, Derek Hoiem, David A. Forsyth
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Discriminative multi-manifold analysis for face recognition from a single training sample per person
abstract
Conventional appearance-based face recognition methods usually assume there are multiple samples per person (MSPP) available during the training phase for discriminative feature extraction. In many practical face recognition applications such as law enhancement, e-passport and ID card identification, this assumption, however, may not hold as there is only a single sample per person (SSPP) enrolled or recorded in these systems. Many popular face recognition methods fail to work well in this scenario because there are not enough samples for discriminant learning. To address this problem, we propose in this paper a novel discriminative multi-manifold analysis (DMMA) method by learning discriminative features from image patches. First, we partition each enrolled image into several non-overlapping patches to form an image set for each sample per person. Then, we formulate the SSPP face recognition as a manifold-manifold matching problem and learn multiple DMMA feature spaces to maximize the manifold margins of different persons. Lastly, we propose a reconstruction-based manifold-manifold distance to identify the unlabeled subjects. Experimental results on three widely used face databases are presented to demonstrate the efficacy of the proposed approach.
Jiwen Lu, Yap-Peng Tan, Gang Wang 0012
ICCV3
2010 Comparative object similarity for improved recognition with few or no examples
abstract
Learning models for recognizing objects with few or no training examples is important, due to the intrinsic long-tailed distribution of objects in the real world. In this paper, we propose an approach to use comparative object similarity. The key insight is that: given a set of object categories which are similar and a set of categories which are dissimilar, a good object model should respond more strongly to examples from similar categories than to examples from dissimilar categories. We develop a regularized kernel machine algorithm to use this category dependent similarity regularization. Our experiments on hundreds of categories show that our method can make significant improvement, especially for categories with no examples.
Gang Wang 0012, David A. Forsyth, Derek Hoiem
CVPR1
2010 Seeing People in Social Context: Recognizing People and Social Relationships
Gang Wang 0012, Andrew C. Gallagher, Jiebo Luo 0001, David A. Forsyth
ECCV (5)1
2010 iRIN: image retrieval in image-rich information networks
abstract
In this demo, we present a system called iRIN designed for performing image retrieval in image-rich information networks. We first introduce MoK-SimRank to significantly improve the speed of SimRank, one of the most popular algorithms for computing node similarity in information networks. Next, we propose an algorithm called SimLearn to (1) extend MoK-SimRank to heterogeneous image-rich information network, and (2) account for both link-based and content-based similarities by seamlessly integrating reinforcement learning with feature learning.
Xin Jin 0001, Jiebo Luo 0001, Jie Yu 0001, Gang Wang 0012, Dhiraj Joshi, Jiawei Han 0001
WWW4
2010 It's All About the Data
abstract
Modern computer vision research consumes labelled data in quantity, and building datasets has become an important activity. The Internet has become a tremendous resource for computer vision researchers. By seeing the Internet as a vast, slightly disorganized collection of visual data, we can build datasets. The key point is that visual data are surrounded by contextual information like text and HTML tags, which is a strong, if noisy, cue to what the visual data means. In a series of case studies, we illustrate how useful this contextual information is. It can be used to build a large and challenging labelled face dataset with no manual intervention. With very small amounts of manual labor, contextual data can be used together with image data to identify pictures of animals. In fact, these contextual data are sufficiently reliable that a very large pool of noisily tagged images can be used as a resource to build image features, which reliably improve on conventional visual features. By seeing the Internet as a marketplace that can connect sellers of annotation services to researchers, we can obtain accurately annotated datasets quickly and cheaply. We describe methods to prepare data, check quality, and set prices for work for this annotation process. The problems posed by attempting to collect very big research datasets are fertile for researchers because collecting datasets requires us to focus on two important questions: What makes a good picture? What is the meaning of a picture?
Tamara L. Berg, Alexander Sorokin, Gang Wang 0012, David A. Forsyth, Derek Hoiem, Ian Endres, Ali Farhadi
Proc. IEEE3
2009 Building text features for object image classification
abstract
We introduce a text-based image feature and demonstrate that it consistently improves performance on hard object classification problems. The feature is built using an auxiliary dataset of images annotated with tags, downloaded from the Internet. We do not inspect or correct the tags and expect that they are noisy. We obtain the text feature of an unannotated image from the tags of its k-nearest neighbors in this auxiliary collection. A visual classifier presented with an object viewed under novel circumstances (say, a new viewing direction) must rely on its visual examples. Our text feature may not change, because the auxiliary dataset likely contains a similar picture. While the tags associated with images are noisy, they are more stable when appearance changes. We test the performance of this feature using PASCAL VOC 2006 and 2007 datasets. Our feature performs well, consistently improves the performance of visual object classifiers, and is particularly effective when the training dataset is small.
Gang Wang 0012, Derek Hoiem, David A. Forsyth
CVPR1
2009 Joint learning of visual attributes, object classes and visual saliency
abstract
We present a method to learn visual attributes (eg.“red”, “metal”, “spotted”) and object classes (eg. “car”, “dress”, “umbrella”) together. We assume images are labeled with category, but not location, of an instance. We estimate models with an iterative procedure: the current model is used to produce a saliency score, which, together with a homogeneity cue, identifies likely locations for the object (resp. attribute); then those locations are used to produce better models with multiple instance learning. Crucially, the object and attribute models must agree on the potential locations of an object. This means that the more accurate of the two models can guide the improvement of the less accurate model. Our method is evaluated on two data sets of images of real scenes, one in which the attribute is color and the other in which it is material. We show that our joint learning produces improved detectors. We demonstrate generalization by detecting attribute-object pairs which do not appear in our training data. The iteration gives significant improvement in performance.
Gang Wang 0012, David A. Forsyth
ICCV1
2009 Learning image similarity from Flickr groups using Stochastic Intersection Kernel MAchines
abstract
Measuring image similarity is a central topic in computer vision. In this paper, we learn similarity from Flickr groups and use it to organize photos. Two images are similar if they are likely to belong to the same Flickr groups. Our approach is enabled by a fast Stochastic Intersection Kernel MAchine (SIKMA) training algorithm, which we propose. This proposed training method will be useful for many vision problems, as it can produce a classifier that is more accurate than a linear classifier, trained on tens of thousands of examples in two minutes. The experimental results show our approach performs better on image matching, retrieval, and classification than using conventional visual features.
Gang Wang 0012, Derek Hoiem, David A. Forsyth
ICCV1
2008 Object image retrieval by exploiting online knowledge resources
abstract
We describe a method to retrieve images found on web pages with specified object class labels, using an analysis of text around the image and of image appearance. Our method determines whether an object is both described in text and appears in a image using a discriminative image model and a generative text model. Our models are learnt by exploiting established online knowledge resources (Wikipedia pages for text; Flickr and Caltech data sets for image). These resources provide rich text and object appearance information. We describe results on two data sets. The first is Berg’s collection of ten animal categories; on this data set, we outperform previous approaches [7, 33]. We have also collected five more categories. Experimental results show the effectiveness of our approach on this new data set.
Gang Wang 0012, David A. Forsyth
CVPR1