Wei Wang 0115

dblp:35/7092-0115 · DBLP profile ↗
← Back
64ranked-venue papers
4as first author
20since 2021 · last 2025
0000-0002-5750-6980ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 38 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Bayesian Active Learning for Bivariate Causal Discovery
abstract
Determining the direction of relationships between variables is fundamental for understanding complex systems across scientific domains. While observational data can uncover relationships between variables, it cannot distinguish between cause and effect without experimental interventions. To effectively uncover causality, previous works have proposed intervention strategies that sequentially optimize the intervention values. However, most of these approaches primarily maximized information-theoretic gains that may not effectively measure the reliability of direction determination. In this paper, we formulate the causal direction identification as a hypothesis-testing problem, and propose a Bayes factor-based intervention strategy, which can quantify the evidence strength of one hypothesis (e.g., causal) over the other (e.g., non-causal). To balance the immediate and future gains of testing strength, we propose a sequential intervention objective over intervention values in multiple steps. By analyzing the objective function, we develop a dynamic programming algorithm that reduces the complexity from non-polynomial to polynomial. Experimental results on bivariate systems, tree-structured graphs, and an embodied AI environment demonstrate the effectiveness of our framework in direction determination and its extensibility to both multivariate settings and real-world applications.
Yuxuan Wang 0005, Mingzhou Liu 0001, Xinwei Sun 0001, Wei Wang 0115, Yizhou Wang 0001
ICML4
2025 Facial Expression Generation from Text with FaceCLIP
Wenwen Fu, Wenjuan Gong, Chen-Yang Yu, Wei Wang 0115, Jordi Gonzàlez 0001
J. Comput. Sci. Technol.4
2025 Pedestrian Attribute Recognition via CLIP-Based Prompt Vision-Language Fusion
abstract
Existing pedestrian attribute recognition (PAR) algorithms adopt pre-trained CNN (e.g., ResNet) as their backbone network for visual feature learning, which might obtain sub-optimal results due to the insufficient employment of the relations between pedestrian images and attribute labels. In this paper, we formulate PAR as a vision-language fusion problem and fully exploit the relations between pedestrian images and attribute labels. Specifically, the attribute phrases are first expanded into sentences, and then the pre-trained vision-language model CLIP is adopted as our backbone for feature embedding of visual images and attribute descriptions. The contrastive learning objective connects the vision and language modalities well in the CLIP-based feature space, and the Transformer layers used in CLIP can capture the long-range relations between pixels. Then, a multi-modal Transformer is adopted to fuse the dual features effectively and feed-forward network is used to predict attributes. To optimize our network efficiently, we propose the region-aware prompt tuning technique to adjust very few parameters (i.e., only the prompt vectors and classification heads) and fix both the pre-trained VL model and multi-modal Transformer. Our proposed PAR algorithm only adjusts 0.75% learnable parameters compared with the fine-tuning strategy. It also achieves new state-of-the-art performance on both standard and zero-shot settings for PAR, including RAPv1, RAPv2, WIDER, PA100K, and PETA-ZS, RAP-ZS datasets. The source code and pre-trained models will be released onhttps://github.com/Event-AHU/OpenPAR.
Xiao Wang 0014, Jiandong Jin, Chenglong Li 0002, Jin Tang 0001, Cheng Zhang 0010, Wei Wang 0115
IEEE Trans. Circuits Syst. Video Technol.6
2024 Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real World
abstract
We introduce Bongard-OpenWorld, a new benchmark for evaluating real-world few-shot reasoning for machine vision. It originates from the classical Bongard Problems (BPs): Given two sets of images (positive and negative), the model needs to identify the set that query images belong to by inducing the visual concepts, which is exclusively depicted by images from the positive set. Our benchmark inherits the few-shot concept induction of the original BPs while adding the two novel layers of challenge: 1) open-world free-form concepts, as the visual concepts in Bongard-OpenWorld are unique compositions of terms from an open vocabulary, ranging from object categories to abstract visual attributes and commonsense factual knowledge; 2) real-world images, as opposed to the synthetic diagrams used by many counterparts. In our exploration, Bongard-OpenWorld already imposes a significant challenge to current few-shot reasoning algorithms. We further investigate to which extent the recently introduced Large Language Models (LLMs) and Vision-Language Models (VLMs) can solve our task, by directly probing VLMs, and combining VLMs and LLMs in an interactive reasoning scheme. We even conceived a neuro-symbolic reasoning approach that reconciles LLMs & VLMs with logical reasoning to emulate the human problem-solving process for Bongard Problems. However, none of these approaches manage to close the human-machine gap, as the best learner achieves 64% accuracy while human participants easily reach 91%. We hope Bongard-OpenWorld can help us better understand the limitations of current visual intelligence and facilitate future research on visual agents with stronger few-shot visual reasoning capabilities.
Rujie Wu, Xiaojian Ma 0001, Zhenliang Zhang 0002, Wei Wang 0115, Qing Li 0003, Song-Chun Zhu, Yizhou Wang 0001
ICLR4
2024 Learning Energy-Based Models for 3D Human Pose Estimation
abstract
Recently, 3D human pose estimation has attracted more attention due to its promising applications. In general, existing methods usually directly predict a target 3D pose for a given input using a Deep Neural Network (DNN), and train the DNN by minimizing the mean squared error (MSE) loss. Despite the impressive performance of these methods, they create a fixed-variance Gaussian model of the conditional target density (the distribution for the target 3D pose given the input) from a probabilistic perspective, which significantly restricts the expressive capabilities of the learned conditional target density. Thus, this hinders the complete utilization of the predictive potential embedded within the DNN. We tackle this problem by delving into the latest developments in conditional energy-based models (EBMs) for probabilistic regression. In this work, we design a simple yet effective network to learn an energy function from 2D and 3D joints pairs. Then a gradient-based refinement procedure is adopted to minimize the energy function to find the corresponding target 3D pose. In this way, we can apply the energy-based model to refine the initial 3D joints estimated by the state-of-the-art 3D human pose estimator. Extensive experiments are conducted on two popular benchmarks on human pose estimation and the results demonstrate the superiority of our method over existing state-of-the-art approaches.
Xianglu Zhu, Zhang Zhang 0001, Wei Wang 0115, Zilei Wang, Liang Wang 0001
IJCNN3
2024 Learning Concept-Based Causal Transition and Symbolic Reasoning for Visual Planning
abstract
Visual planning simulates how humans make decisions to achieve desired goals in the form of searching for visual causal transitions between an initial visual state and a final visual goal state. It has become increasingly important in egocentric vision with its advantages in guiding agents to perform daily tasks in complex environments. In this paper, we propose an interpretable and generalizable visual planning framework consisting of i) a novel Substitution-based Concept Learner (SCL) that abstracts visual inputs into disentangled concept representations, ii) symbol abstraction and reasoning that performs task planning via the learned symbols, and iii) a Visual Causal Transition model (ViCT) that grounds visual causal transitions to semantically similar real-world actions. Given an initial state, we perform goal-conditioned visual planning with a symbolic reasoning method fueled by the learned representations and causal transitions to reach the goal state. To verify the effectiveness of the proposed model, we collect a large-scale visual planning dataset based on AI2-THOR, dubbed as CCTP. Extensive experiments on this challenging dataset demonstrate the superior performance of our method in visual planning. Empirically, we show that our framework can generalize to unseen task trajectories, unseen object categories, and real-world data. Further details of this work are provided at https://fqyqc.github.io/ConTranPlan/.
Yilue Qian, Peiyu Yu, Ying Nian Wu, Yao Su 0001, Wei Wang 0115, Lifeng Fan
IROS5
2024 On the Emergence of Symmetrical Reality
abstract
Artificial intelligence (AI) has revolutionized human cognitive abilities and facilitated the development of new AI entities capable of interacting with humans in both physical and virtual environments. Despite the existence of virtual reality, mixed reality, and augmented reality for many years, integrating these technical fields remains a formidable challenge due to their disparate application directions. The advent of AI agents, capable of autonomous perception and action, further compounds this issue by exposing the limitations of traditional human-centered research approaches. It is imperative to establish a comprehensive framework that accommodates the dual perceptual centers of humans and AI agents in both physical and virtual worlds. In this paper, we introduce the symmetrical reality framework, which offers a unified representation encompassing various forms of physical-virtual amalgamations. This framework enables researchers to better comprehend how AI agents can collaborate with humans and how distinct technical pathways of physical-virtual integration can be consolidated from a broader perspective. We then delve into the coexistence of humans and AI, demonstrating a prototype system that exemplifies the operation of symmetrical reality systems for specific tasks, such as pouring water. Finally, we propose an instance of an AI-driven active assistance service that illustrates the potential applications of symmetrical reality. This paper aims to offer beneficial perspectives and guidance for researchers and practitioners in different fields, thus contributing to the ongoing research about human-AI coexistence in both physical and virtual environments.
Zhenliang Zhang 0002, Zeyu Zhang 0001, Ziyuan Jiao, Yao Su 0001, Hangxin Liu, Wei Wang 0115, Song-Chun Zhu
VR6
2024 Text-to-Image Vehicle Re-Identification: Multi-Scale Multi-View Cross-Modal Alignment Network and a Unified Benchmark
abstract
Vehicle Re-IDentification (Re-ID) aims to retrieve the most similar images with a given query vehicle image from a set of images captured by non-overlapping cameras, and plays a crucial role in intelligent transportation systems and has made impressive advancements in recent years. In real-world scenarios, we can often acquire the text descriptions of target vehicle through witness accounts, and then manually search the image queries for vehicle Re-ID, which is time-consuming and labor-intensive. To solve this problem, this paper introduces a new fine-grained cross-modal retrieval task called text-to-image vehicle re-identification, which seeks to retrieve target vehicle images based on the given text descriptions. To bridge the significant gap between language and visual modalities, we propose a novel Multi-scale multi-view Cross-modal Alignment Network (MCANet). In particular, we incorporate view masks and multi-scale features to align image and text features in a progressive way. In addition, we design the Masked Bidirectional InfoNCE (MB-InfoNCE) loss to enhance the training stability and make the best use of negative samples. To provide an evaluation platform for text-to-image vehicle re-identification, we create a Text-to-Image Vehicle Re-Identification dataset (T2I VeRi), which contains 2465 image-text pairs from 776 vehicles with an average sentence length of 26.8 words. Extensive experiments conducted on T2I VeRi demonstrate MCANet outperforms the current state-of-art (SOTA) method by 2.2% in rank-1 accuracy.
Leqi Ding, Lei Liu 0049, Yan Huang 0008, Chenglong Li 0002, Cheng Zhang 0010, Wei Wang 0115, Liang Wang 0001
IEEE Trans. Intell. Transp. Syst.6
2024 Collaborative License Plate Recognition via Association Enhancement Network With Auxiliary Learning and a Unified Benchmark
abstract
Since the standard license plate of large vehicle is easily affected by occlusion and stain, the traffic management department introduces the enlarged license plate at the rear of the large vehicle to assist license plate recognition. However, current researches regards standard license plate recognition and enlarged license plate recognition as independent tasks, and do not take advantage of the complementary benefits from the two types of license plates. In this work, we propose a new computer vision task called collaborative license plate recognition, aiming to leverage the complementary advantages of standard and enlarged license plates for achieving more accurate license plate recognition. To achieve this goal, we propose an Association Enhancement Network (AENet), which achieves robust collaborative licence plate recognition by capturing the correlations between characters within a single licence plate and enhancing the associations between two license plates. In particular, we design an association enhancement branch, which supervises the fusion of two licence plate information using the complete licence plate number to mine the association between them. To enhance the representation ability of each type of licence plates, we design an auxiliary learning branch in the training stage, which supervises the learning of individual license plates in the association enhancement between two license plates. In addition, we contribute a comprehensive benchmark dataset called CLPR, which consists of a total of 19,782 standard and enlarged licence plates from 24 provinces in China and covers most of the challenges in real scenarios, for collaborative license plate recognition. Extensive experiments on the proposed CLPR dataset demonstrate the effectiveness of the proposed AENet against several state-of-the-art methods.
Yifei Deng, Guohao Wang, Chenglong Li 0002, Wei Wang 0115, Cheng Zhang 0010, Jin Tang 0001
IEEE Trans. Multim.4
2024 Meta-MMFNet: Meta-learning-based Multi-model Fusion Network for Micro-expression Recognition
abstract
Despite its wide applications in criminal investigations and clinical communications with patients suffering from autism, automatic micro-expression recognition remains a challenging problem because of the lack of training data and imbalanced classes problems. In this study, we proposed a meta-learning-based multi-model fusion network (Meta-MMFNet) to solve the existing problems. The proposed method is based on the metric-based meta-learning pipeline, which is specifically designed for few-shot learning and is suitable for model-level fusion. The frame difference and optical flow features were fused, deep features were extracted from the fused feature, and finally in the meta-learning-based framework, weighted sum model fusion method was applied for micro-expression classification. Meta-MMFNet achieved better results than state-of-the-art methods on four datasets. The code is available at https://github.com/wenjgong/meta-fusion-based-method .
Wenjuan Gong, Yue Zhang 0087, Wei Wang 0115, Peng Cheng 0008, Jordi Gonzàlez 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Transbuilding: An End-to-End Polygonal Building Extraction with Transformers
abstract
In this paper, we propose a simple yet powerful network, called TransBuilding, for high-quality polygonal building extraction from remote sensing images. Unlike many previous methods that vectorize building masks through mask refinement and fitting or vertex prediction and assembling, our approach predicts the building vertex sequence with a vertex transformer (termed as VertexFormer) branch without any additional processing. The VertexFormer branch represents a polygon as a Bi-directional Ring without start or end vertex hypothesis, which leads to a simple and elegant representation of polygons avoiding ambiguous of defining the start vertex in polygons. Furthermore, three self-attention modules in row-wise, column-wise, and vertex-wise are integrated in parallel together to better capture geometric structures of building polygons. We graft the VertexFormer module onto the standard Faster RCNN detector and train the model end-to-endly using the novel Bi-Ring loss developed by the new perspective of Bi-directional Ring. Extensive experiments on the benchmark CrowdAI dataset demonstrate that our method outperforms state-of-the-art methods by considerable margins.
Weiming Zhang 0001, Qingjie Liu 0001, Wei Wang 0115, Yunhong Wang 0001
ICIP3
2023 TERNformer: Topology-Enhanced Road Network Extraction by Exploring Local Connectivity
abstract
Remote-sensing images provide us with rich information for extracting road networks. However, there are still great challenges ahead, such as occlusions caused by trees and shadows, and complex topology. In this work, we focus on the topology of road networks. Inspired by the observation that road networks are composed of road fragments in a bottom-up way and the breaks between fragments tend to be connected within a local area, we propose a Topology-Enhanced Road Network extraction (termed TERNformer) method by exploring local connectivity. First, a transformer-based network is built for road feature extraction to capture long-range context. Furthermore, we propose parallel depth-wise separable dilated convolution blocks (DSDB) to extract local information within different ranges. Thereafter, a minimum spanning tree-based local structure exploring block (LSEB) is built to enhance the topology of the road network. Finally, a simple but effective shortest-path-based method is used to refine the road network connectivity within a local threshold. Experiments conducted on two datasets demonstrate the superiority of TERNformer. TERNformer outperforms the state-of-the-art methods on CityScale dataset with the best topology performance. The result on DeepGlobe dataset improves 4.83% APLS to state-of-the-art methods.
Qingjie Liu 0001, Wei Wang 0115, Yunhong Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Pose-Appearance Relational Modeling for Video Action Recognition
abstract
Recent studies of video action recognition can be classified into two categories: the appearance-based methods and the pose-based methods. The appearance-based methods generally cannot model temporal dynamics of large motion well by virtue of optical flow estimation, while the pose-based methods ignore the visual context information such as typical scenes and objects, which are also important cues for action understanding. In this paper, we tackle these problems by proposing a Pose-Appearance Relational Network (PARNet), which models the correlation between human pose and image appearance, and combines the benefits of these two modalities to improve the robustness towards unconstrained real-world videos. There are three network streams in our model, namely pose stream, appearance stream and relation stream. For the pose stream, a Temporal Multi-Pose RNN module is constructed to obtain the dynamic representations through temporal modeling of 2D poses. For the appearance stream, a Spatial Appearance CNN module is employed to extract the global appearance representation of the video sequence. For the relation stream, a Pose-Aware RNN module is built to connect pose and appearance streams by modeling action-sensitive visual context information. Through jointly optimizing the three modules, PARNet achieves superior performances compared with the state-of-the-arts on both the pose-complete datasets (KTH, Penn-Action, UCF11) and the challenging pose-incomplete datasets (UCF101, HMDB51, JHMDB), demonstrating its robustness towards complex environments and noisy skeletons. Its effectiveness on NTU-RGBD dataset is also validated even compared with 3D skeleton-based methods. Furthermore, an appearance-enhanced PARNet equipped with a RGB-based I3D stream is proposed, which outperforms the Kinetics pre-trained competitors on UCF101 and HMDB51. The better experimental results verify the potentials of our framework by integrating various modules.
Mengmeng Cui, Wei Wang 0115, Kunbo Zhang, Zhenan Sun, Liang Wang 0001
IEEE Trans. Image Process.2
2023 Multi-Level Adversarial Spatio-Temporal Learning for Footstep Pressure Based FoG Detection
abstract
Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease, which is a neurodegenerative disorder of the central nervous system impacting millions of people around the world. To address the pressing need to improve the quality of treatment for FoG, devising a computer-aided detection and quantification tool for FoG has been increasingly important. As a non-invasive technique for collecting motion patterns, the footstep pressure sequences obtained from pressure sensitive gait mats provide a great opportunity for evaluating FoG in the clinic and potentially in the home environment. In this study, FoG detection is formulated as a sequential modelling task and a novel deep learning architecture, namely Adversarial Spatio-temporal Network (ASTN), is proposed to learn FoG patterns across multiple levels. ASTN introduces a novel adversarial training scheme with a multi-level subject discriminator to obtain subject-independent FoG representations, which helps to reduce the over-fitting risk due to the high inter-subject variance. As a result, robust FoG detection can be achieved for unseen subjects. The proposed scheme also sheds light on improving subject-level clinical studies from other scenarios as it can be integrated with many existing deep architectures. To the best of our knowledge, this is one of the first studies of footstep pressure-based FoG detection and the approach of utilizing ASTN is the first deep neural network architecture in pursuit of subject-independent representations. In our experiments on 393 trials collected from 21 subjects, the proposed ASTN achieved an AUC 0.85, clearly outperforming conventional learning methods.
Kun Hu 0008, Shaohui Mei, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Simon J. G. Lewis, David Dagan Feng, Zhiyong Wang 0001
IEEE J. Biomed. Health Informatics3
2022 Cross-Domain Cross-Set Few-Shot Learning via Learning Compact and Aligned Representations
Zhang Zhang 0001, Wei Wang 0115, Liang Wang 0001, Zilei Wang, Tieniu Tan
ECCV (20)3
2021 Locate Then Segment: A Strong Pipeline for Referring Image Segmentation
abstract
Referring image segmentation aims to segment the objects referred by a natural language expression. Previous methods usually focus on designing an implicit and recurrent feature interaction mechanism to fuse the visual-linguistic features to directly generate the final segmentation mask without explicitly modeling the localization information of the referent instances. To tackle these problems, we view this task from another perspective by decoupling it into a "Locate-Then-Segment" (LTS) scheme. Given a language expression, people generally first perform attention to the corresponding target image regions, then generate a fine segmentation mask about the object based on its context. The LTS first extracts and fuses both visual and textual features to get a cross-modal representation, then applies a cross-model interaction on the visual-textual features to locate the referred object with position prior, and finally generates the segmentation result with a light-weight segmentation network. Our LTS is simple but surprisingly effective. On three popular benchmark datasets, the LTS outperforms all the previous state-of-the-arts methods by a large margin (e.g., +3.2% on RefCOCO+ and +3.4% on RefCOCOg). In addition, our model is more interpretable with explicitly locating the object, which is also proved by visualization experiments. We believe this framework is promising to serve as a strong baseline for referring image segmentation.
Ya Jing, Tao Kong, Wei Wang 0115, Liang Wang 0001, Lei Li 0005, Tieniu Tan
CVPR3
2021 Representation and Correlation Enhanced Encoder-Decoder Framework for Scene Text Recognition
Mengmeng Cui, Wei Wang 0115, Liang Wang 0001
ICDAR (4)2
2021 Few-Shot Learning with Part Discovery and Augmentation from Unlabeled Images
abstract
Few-shot learning is a challenging task since only few instances are given for recognizing an unseen class. One way to alleviate this problem is to acquire a strong inductive bias via meta-learning on similar tasks. In this paper, we show that such inductive bias can be learned from a flat collection of unlabeled images, and instantiated as transferable representations among seen and unseen classes. Specifically, we propose a novel part-based self-supervised representation learning scheme to learn transferable representations by maximizing the similarity of an image to its discriminative part. To mitigate the overfitting in few-shot classification caused by data scarcity, we further propose a part augmentation strategy by retrieving extra images from a base dataset. We conduct systematic studies on miniImageNet and tieredImageNet benchmarks. Remarkably, our method yields impressive results, outperforming the previous best unsupervised methods by 7.74% and 9.24% under 5-way 1-shot and 5-way 5-shot settings, which are comparable with state-of-the-art supervised methods.
Chenyang Si, Wei Wang 0115, Liang Wang 0001, Zilei Wang, Tieniu Tan
IJCAI3
2021 Joint Learning Appearance and Motion Models for Visual Tracking
Wenmei Xu, Hongyuan Yu, Wei Wang 0115, Chenglong Li 0002, Liang Wang 0001
PRCV (1)3
2021 Learning Aligned Image-Text Representations Using Graph Attentive Relational Network
abstract
Image-text matching aims to measure the similarities between images and textual descriptions, which has made great progress recently. The key to this cross-modal matching task is to build the latent semantic alignment between visual objects and words. Due to the widespread variations of sentence structures, it is very difficult to learn the latent semantic alignment using only global cross-modal features. Many previous methods attempt to learn the aligned image-text representations by the attention mechanism but generally ignore the relationships within textual descriptions which determine whether the words belong to the same visual object. In this paper, we propose a graph attentive relational network (GARN) to learn the aligned image-text representations by modeling the relationships between noun phrases in a text for the identity-aware image-text matching. In the GARN, we first decompose images and texts into regions and noun phrases, respectively. Then a skip graph neural network (skip-GNN) is proposed to learn effective textual representations which are a mixture of textual features and relational features. Finally, a graph attention network is further proposed to obtain the probabilities that the noun phrases belong to the image regions by modeling the relationships between noun phrases. We perform extensive experiments on the CUHK Person Description dataset (CUHK-PEDES), Caltech-UCSD Birds dataset (CUB), Oxford-102 Flowers dataset and Flickr30K dataset to verify the effectiveness of each component in our model. Experimental results show that our approach achieves the state-of-the-art results on these four benchmark datasets.
Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
IEEE Trans. Image Process.2
2020 Pose-Guided Multi-Granularity Attention Network for Text-Based Person Search
abstract
Text-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting visual contents corresponding to the human description is the key to this cross-modal matching problem. Moreover, correlated images and descriptions involve different granularities of semantic relevance, which is usually ignored in previous methods. To exploit the multilevel corresponding visual contents, we propose a pose-guided multi-granularity attention network (PMA). Firstly, we propose a coarse alignment network (CA) to select the related image regions to the global description by a similarity-based attention. To further capture the phrase-related visual body part, a fine-grained alignment network (FA) is proposed, which employs pose information to learn latent semantic alignment between visual body part and textual noun phrase. To verify the effectiveness of our model, we perform extensive experiments on the CUHK Person Description Dataset (CUHK-PEDES) which is currently the only available dataset for text-based person search. Experimental results show that our approach outperforms the state-of-the-art methods by 15 % in terms of the top-1 metric.
Ya Jing, Chenyang Si, Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
AAAI4
2020 Cross-Modal Cross-Domain Moment Alignment Network for Person Search
abstract
Text-based person search has drawn increasing attention due to its wide applications in video surveillance. However, most of the existing models depend heavily on paired image-text data, which is very expensive to acquire. Moreover, they always face huge performance drop when directly exploiting them to new domains. To overcome this problem, we make the first attempt to adapt the model to new target domains in the absence of pairwise labels, which combines the challenges from both cross-modal (text-based) person search and cross-domain person search. Specially, we propose a moment alignment network (MAN) to solve the cross-modal cross-domain person search task in this paper. The idea is to learn three effective moment alignments including domain alignment (DA), cross-modal alignment (CA) and exemplar alignment (EA), which together can learn domain-invariant and semantic aligned cross-modal representations to improve model generalization. Extensive experiments are conducted on CUHK Person Description dataset (CUHK-PEDES) and Richly Annotated Pedestrian dataset (RAP). Experimental results show that our proposed model achieves the state-of-the-art performances on five transfer tasks.
Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
CVPR2
2020 Adversarial Self-supervised Learning for Semi-supervised 3D Action Recognition
Chenyang Si, Xuecheng Nie, Wei Wang 0115, Liang Wang 0001, Tieniu Tan, Jiashi Feng
ECCV (7)3
2020 Image and Sentence Matching via Semantic Concepts and Order Learning
abstract
Image and sentence matching has made great progress recently, but it remains challenging due to the existing large visual-semantic discrepancy. This mainly arises from two aspects: 1) images consist of unstructured content which is not semantically abstract as the words in the sentences, so they are not directly comparable, and 2) arranging semantic concepts in different semantic order could lead to quite diverse meanings. The words in the sentences are sequentially arranged in a grammatical manner, while the semantic concepts in the images are usually unorganized. In this work, we propose a semantic concepts and order learning framework for image and sentence matching, which can improve the image representation by first predicting semantic concepts and then organizing them in a correct semantic order. Given an image, we first use a multi-regional multi-label CNN to predict its included semantic concepts in terms of object, property and action. These word-level semantic concepts are directly comparable with the words of noun, adjective and verb in the matched sentence. Then, to organize these concepts and make them express similar meanings as the matched sentence, we use a context-modulated attentional LSTM to learn the semantic order. It regards the predicted semantic concepts and image global scene as context at each timestep, and selectively attends to concept-related image regions by referring to the context in a sequential order. To further enhance the semantic order, we perform additional sentence generation on the image representation, by using the groundtruth order in the matched sentence as supervision. After obtaining the improved image representation, we learn the sentence representation with a conventional LSTM, and then jointly perform image and sentence matching and sentence generation for model learning. Extensive experiments demonstrate the effectiveness of our learned semantic concepts and order, by achieving the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Qi Wu 0001, Wei Wang 0115, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Relational graph neural network for situation recognition
Ya Jing, Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
Pattern Recognit.3
2020 Skeleton-based action recognition with hierarchical spatial reasoning and temporal stack learning network
Chenyang Si, Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
Pattern Recognit.3
2020 Learning visual relationship and context-aware attention for image captioning
Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Zhiyong Wang 0001, David Dagan Feng, Tieniu Tan
Pattern Recognit.2
2020 Graph Sequence Recurrent Neural Network for Vision-Based Freezing of Gait Detection
abstract
Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease (PD), a neurodegenerative disorder which impacts millions of people around the world. Accurate assessment of FoG is critical for the management of PD and to evaluate the efficacy of treatments. Currently, the assessment of FoG requires well-trained experts to perform time-consuming annotations via vision-based observations. Thus, automatic FoG detection algorithms are needed. In this study, we formulate vision-based FoG detection, as a fine-grained graph sequence modelling task, by representing the anatomic joints in each temporal segment with a directed graph, since FoG events can be observed through the motion patterns of joints. A novel deep learning method is proposed, namely graph sequence recurrent neural network (GS-RNN), to characterize the FoG patterns by devising graph recurrent cells, which take graph sequences of dynamic structures as inputs. For the cases of which prior edge annotations are not available, a data-driven based adjacency estimation method is further proposed. To the best of our knowledge, this is one of the first studies on vision-based FoG detection using deep neural networks designed for graph sequences of dynamic structures. Experimental results on more than 150 videos collected from 45 patients demonstrated promising performance of the proposed GS-RNN for FoG detection with an AUC value of 0.90.
Kun Hu 0008, Zhiyong Wang 0001, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Tieniu Tan, Simon J. G. Lewis, David Dagan Feng
IEEE Trans. Image Process.3
2020 Attribute-Guided Attention for Referring Expression Generation and Comprehension
abstract
Referring expression is a special kind of verbal expression. The goal of referring expression is to refer to a particular object in some scenarios. Referring expression generation and comprehension are two inverse tasks within the field. Considering the critical role that visual attributes play in distinguishing the referred object from other objects, we propose an attribute-guided attention model to address the two tasks. In our proposed framework, attributes collected from referring expressions are used as explicit supervision signals on the generation and comprehension modules. The online predicted attributes of the visual object can benefit both tasks in two aspects: First, attributes can be directly embedded into the generation and comprehension modules, distinguishing the referred object as additional visual representations. Second, since attributes have their correspondence in both visual and textual space, an attribute-guided attention module is proposed as a bridging part to link the counterparts in visual representation and textual expression. Attention weights learned on both visual feature and word embeddings validate our motivation. We experiment on three standard datasets of RefCOCO, RefCOCO+ and RefCOCOg commonly used in this field. Both quantitative and qualitative results demonstrate the effectiveness of our proposed framework. The experimental results show significant improvements over baseline methods, and are favorably comparable to the state-of-the-art results. Further ablation study and analysis clearly demonstrate the contribution of each module, which could provide useful inspirations to the community.
Jingyu Liu 0004, Wei Wang 0115, Liang Wang 0001, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.2
2019 An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition is an important task that requires the adequate understanding of movement characteristics of a human action from the given skeleton sequence. Recent studies have shown that exploring spatial and temporal features of the skeleton sequence is vital for this task. Nevertheless, how to effectively extract discriminative spatial and temporal features is still a challenging problem. In this paper, we propose a novel Attention Enhanced Graph Convolutional LSTM Network (AGC-LSTM) for human action recognition from skeleton data. The proposed AGC-LSTM can not only capture discriminative features in spatial configuration and temporal dynamics but also explore the co-occurrence relationship between spatial and temporal domains. We also present a temporal hierarchical architecture to increase temporal receptive fields of the top AGC-LSTM layer, which boosts the ability to learn the high-level semantic representation and significantly reduces the computation cost. Furthermore, to select discriminative spatial information, the attention mechanism is employed to enhance information of key joints in each AGC-LSTM layer. Experimental results on two datasets are provided: NTU RGB+D dataset and Northwestern-UCLA dataset. The comparison results demonstrate the effectiveness of our approach and show that our approach outperforms the state-of-the-art methods on both datasets.
Chenyang Si, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
CVPR3
2019 Stacked Memory Network for Video Summarization
abstract
In recent years, supervised video summarization has achieved promising progress with various recurrent neural networks (RNNs) based methods, which treats video summarization as a sequence-to-sequence learning problem to exploit temporal dependency among video frames across variable ranges. However, RNN has limitations in modelling the long-term temporal dependency for summarizing videos with thousands of frames due to the restricted memory storage unit. Therefore, in this paper we propose a stacked memory network called SMN to explicitly model the long dependency among video frames so that redundancy could be minimized in the video summaries produced. Our proposed SMN consists of two key components: Long Short-Term Memory (LSTM) layer and memory layer, where each LSTM layer is augmented with an external memory layer. In particular, we stack multiple LSTM layers and memory layers hierarchically to integrate the learned representation from prior layers. By combining the hidden states of the LSTM layers and the read representations of the memory layers, our SMN is able to derive more accurate video summaries for individual video frames. Compared with the existing RNN based methods, our SMN is particularly good at capturing long temporal dependency among frames with few additional training parameters. Experimental results on two widely used public benchmark datasets: SumMe and TVsum, demonstrate that our proposed model is able to clearly outperform a number of state-of-the-art ones under various settings.
Junbo Wang 0003, Wei Wang 0115, Zhiyong Wang 0001, Liang Wang 0001, David Dagan Feng, Tieniu Tan
ACM Multimedia2
2018 Multistage Adversarial Losses for Pose-Based Human Image Synthesis
abstract
Human image synthesis has extensive practical applications e.g. person re-identification and data augmentation for human pose estimation. However, it is much more challenging than rigid object synthesis, e.g. cars and chairs, due to the variability of human posture. In this paper, we propose a pose-based human image synthesis method which can keep the human posture unchanged in novel viewpoints. Furthermore, we adopt multistage adversarial losses separately for the foreground and background generation, which fully exploits the multi-modal characteristics of generative loss to generate more realistic looking images. We perform extensive experiments on the Human3.6M dataset and verify the effectiveness of each stage of our method. The generated human images not only keep the same pose as the input image, but also have clear detailed foreground and background. The quantitative comparison results illustrate that our approach achieves much better results than several state-of-the-art methods.
Chenyang Si, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
CVPR2
2018 M3: Multimodal Memory Modelling for Video Captioning
abstract
Video captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, video captioning has made great progress. However, learning an effective mapping from the visual sequence space to the language space is still a challenging problem due to the long-term multimodal dependency modelling and semantic misalignment. Inspired by the facts that memory modelling poses potential advantages to long-term sequential problems [35] and working memory is the key factor of visual attention [33], we propose a Multimodal Memory Model (M3) to describe videos, which builds a visual and textual shared memory to model the long-term visual-textual dependency and further guide visual attention on described visual targets to solve visual-textual alignments. Specifically, similar to [10], the proposed M3 attaches an external memory to store and retrieve both visual and textual contents by interacting with video and sentence with multiple read and write operations. To evaluate the proposed model, we perform experiments on two public datasets: MSVD and MSR-VTT. The experimental results demonstrate that our method outperforms most of the state-of-the-art methods in terms of BLEU and METEOR.
Junbo Wang 0003, Wei Wang 0115, Yan Huang 0008, Liang Wang 0001, Tieniu Tan
CVPR2
2018 Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning
Chenyang Si, Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ECCV (1)3
2018 Hough Transform Guided Deep Feature Extraction for Dense Building Detection in Remote Sensing Images
abstract
Detecting dense buildings without elevation information is an important and challenging task in remote sensing applications. In this paper, we present a novel cascaded deep neural network architecture, incorporating multi -stage region proposal detection and Hough transform to obtain better mid-level semantic information for man-made objects. This proposed network can be trained end-to-end by multi-loss jointly. We train and test it on a large building dataset collected from Google Earth, including buildings from urban, suburban and rural areas. Experiments demonstrate great robustness and superiority of our method to various buildings over other convolutional neural network (CNN) based detection methods.
Qingpeng Li, Yunhong Wang 0001, Qingjie Liu 0001, Wei Wang 0115
ICASSP4
2018 RotateConv: Making Asymmetric Convolutional Kernels Rotatable
abstract
In deep Convolutional Neural Networks(CNN), the design of kernel shapes influences a lot on the model size and performance. In this work, our proposed method, RotateConv, applies a novel kernel shape to massively reduce the number of parameters while maintaining considerable performance. The new shape is extremely simple as a line segment one, and we equip it with the rotatable ability which aims to learn diverse features with respect to different angles. The kernel weights and angles are learned simultaneously during end-to-end training via the standard back-propagation algorithm. There are two variants of RotateConv that only have 2 and 4 parameters respectively depending on whether using weight sharing, which are much compressed than the normal 3×3 kernel with 9 parameters. In experiments, we validate our RotateConv with two classical models, ResNet and DenseNet, on four image classification benchmark datasets, namely MNIST, CIFAR10, CIFAR100 and SVHN.
Jiabin Ma, Weiyu Guo, Wei Wang 0115, Liang Wang 0001
ICPR3
2018 Hierarchical Memory Modelling for Video Captioning
abstract
Translating videos into natural language sentences has drawn much attention recently. The framework of combining visual attention with Long Short-Term Memory (LSTM) based text decoder has achieved much progress. However, the vision-language translation still remains unsolved due to the semantic gap and misalignment between video content and described semantic concept. In this paper, we propose a Hierarchical Memory Model (HMM) - a novel deep video captioning architecture which unifies a textual memory, a visual memory and an attribute memory in a hierarchical way. These memories can guide attention for efficient video representation extraction and semantic attribute selection in addition to modelling the long-term dependency for video sequence and sentences, respectively. Compared with traditional vision-based text decoder, the proposed attribute-based text decoder can largely reduce the semantic discrepancy between video and sentence. To prove the effectiveness of the proposed model, we perform extensive experiments on two public benchmark datasets: MSVD and MSR-VTT. Experiments show that our model not only can discover appropriate video representation and semantic attributes but also can achieve comparable or superior performances than state-of-the-art methods on these datasets.
Junbo Wang 0003, Wei Wang 0115, Yan Huang 0008, Liang Wang 0001, Tieniu Tan
ACM Multimedia2
2018 Video Super-Resolution via Bidirectional Recurrent Convolutional Networks
abstract
Super resolving a low-resolution video, namely video super-resolution (SR), is usually handled by either single-image SR or multi-frame SR. Single-Image SR deals with each video frame independently, and ignores intrinsic temporal dependency of video frames which actually plays a very important role in video SR. Multi-Frame SR generally extracts motion information, e.g., optical flow, to model the temporal dependency, but often shows high computational cost. Considering that recurrent neural networks (RNNs) can model long-term temporal dependency of video sequences well, we propose a fully convolutional RNN named bidirectional recurrent convolutional network for efficient multi-frame SR. Different from vanilla RNNs, 1) the commonly-used full feedforward and recurrent connections are replaced with weight-sharing convolutional connections. So they can greatly reduce the large number of network parameters and well model the temporal dependency in a finer level, i.e., patch-based rather than frame-based, and 2) connections from input layers at previous timesteps to the current hidden layer are added by 3D feedforward convolutions, which aim to capture discriminate spatio-temporal patterns for short-term fast-varying motions in local adjacent frames. Due to the cheap convolutional operations, our model has a low computational complexity and runs orders of magnitude faster than other multi-frame SR methods. With the powerful temporal dependency modeling, our model can super resolve videos with complex motions and achieve well performance.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Instance-Aware Image and Sentence Matching with Selective Multimodal LSTM
abstract
Effective image and sentence matching depends on how to well measure their global visual-semantic similarity. Based on the observation that such a global similarity arises from a complex aggregation of multiple local similarities between pairwise instances of image (objects) and sentence (words), we propose a selective multimodal Long Short-Term Memory network (sm-LSTM) for instance-aware image and sentence matching. The sm-LSTM includes a multimodal context-modulated attention scheme at each timestep that can selectively attend to a pair of instances of image and sentence, by predicting pairwise instance-aware saliency maps for image and sentence. For selected pairwise instances, their representations are obtained based on the predicted saliency maps, and then compared to measure their local similarity. By similarly measuring multiple local similarities within a few timesteps, the sm-LSTM sequentially aggregates them with hidden states to obtain a final matching score as the desired global similarity. Extensive experiments show that our model can well match image and sentence with complex content, and achieve the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
CVPR2
2017 See the Forest for the Trees: Joint Spatial and Temporal Recurrent Neural Networks for Video-Based Person Re-identification
abstract
Surveillance cameras have been widely used in different scenes. Accordingly, a demanding need is to recognize a person under different cameras, which is called person re-identification. This topic has gained increasing interests in computer vision recently. However, less attention has been paid to video-based approaches, compared with image-based ones. Two steps are usually involved in previous approaches, namely feature learning and metric learning. But most of the existing approaches only focus on either feature learning or metric learning. Meanwhile, many of them do not take full use of the temporal and spatial information. In this paper, we concentrate on video-based person re-identification and build an end-to-end deep neural network architecture to jointly learn features and metrics. The proposed method can automatically pick out the most discriminative frames in a given video by a temporal attention model. Moreover, it integrates the surrounding information at each location by a spatial recurrent model when measuring the similarity with another pedestrian video. That is, our method handles spatial and temporal information simultaneously in a unified manner. The carefully designed experiments on three public datasets show the effectiveness of each component of the proposed deep network, performing better in comparison with the state-of-the-art methods.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
CVPR3
2017 Ground substrate classification for adaptive quadruped locomotion
abstract
In order to realize adaptive quadruped locomotion on terrains with different properties (such as surface friction or elasticity modulus), we plan to collect the foot-ground contact force and gyroscope information during locomotion on different ground substrates, then classify the ground substrates with the feature vector extracted from the collected data using Support Vector Machine (SVM) algorithm. However, the quadruped walk gait generated by Central Pattern Generators (CPGs) does not perform well on certain ground substrates, e.g., robots may be stuck in the soft ground substrates with small elasticity modulus. Therefore, for one thing, we present a Center of Gravity (COG) adjustment method to eliminate the offset between the control signal generated by CPGs and the actual phase of the quadruped limb, so the limb in theoretical swing phase is able to lift off the ground. For another, we combine CPGs with a foot path planning method to make the lift height controllable. Using these methods, the quadruped robot Biodog realizes the sensor data collection on five different ground substrates. Then we train and classify the sensor data with the SVM. About 99.33% of the five ground substrates can be classified correctly.
Xiaoqi Li 0006, Wei Wang 0115, Jianqiang Yi
ICRA2
2017 Conditional High-Order Boltzmann Machines for Supervised Relation Learning
abstract
Relation learning is a fundamental problem in many vision tasks. Recently, high-order Boltzmann machine and its variants have shown their great potentials in learning various types of data relation in a range of tasks. But most of these models are learned in an unsupervised way, i.e., without using relation class labels, which are not very discriminative for some challenging tasks, e.g., face verification. In this paper, with the goal to perform supervised relation learning, we introduce relation class labels into conventional high-order multiplicative interactions with pairwise input samples, and propose a conditional high-order Boltzmann Machine (CHBM), which can learn to classify the data relation in a binary classification way. To be able to deal with more complex data relation, we develop two improved variants of CHBM: 1) latent CHBM, which jointly performs relation feature learning and classification, by using a set of latent variables to block the pathway from pairwise input samples to output relation labels and 2) gated CHBM, which untangles factors of variation in data relation, by exploiting a set of latent variables to multiplicatively gate the classification of CHBM. To reduce the large number of model parameters generated by the multiplicative interactions, we approximately factorize high-order parameter tensors into multiple matrices. Then, we develop efficient supervised learning algorithms, by first pretraining the models using joint likelihood to provide good parameter initialization, and then finetuning them using conditional likelihood to enhance the discriminant ability. We apply the proposed models to a series of tasks including invariant recognition, face verification, and action similarity labeling. Experimental results demonstrate that by exploiting supervised relation labels, our models can greatly improve the performance.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
IEEE Trans. Image Process.2
2016 How scenes imply actions in realistic videos?
abstract
People drive on the road and eat in the kitchen. Can the road imply driving or the kitchen imply eating? This paper addresses such a problem by studying the relations between actions and scenes. To get effective scene representation, we use a deep convolutional neural networks (CNN) model trained from a scene-centric database to predict scene responses for videos. We employ two encoding schemes based on frame features to represent the scene and its changes, respectively. We conduct experiments on two challenging datasets, HMDB51 and Hollywood2, and compare action recognition results of different encodings based on different scene features. Our results demonstrate that scene features, when combined with motion features, improve the state-of-the-art results for action recognition. Finally, we explore the relationship between actions and scenes by analyzing scene preferences to a particular action qualitatively and quantitatively.
Hongsong Wang 0001, Wei Wang 0115, Liang Wang 0001
ICIP2
2016 An attention model based on spatial transformers for scene recognition
abstract
Scene recognition is an important and challenging task in computer vision. We propose an end-to-end pipeline by combing convolutional neural networks (CNNs) with explicit attention model to determine several meaningful regions of original images for scene recognition. In the proposed pipeline, the spatial transformer network is leveraged as the attention module, which can automatically learn the scales and movements of centers of attention windows. As for feature extraction, the basic CNN architecture is utilized. Furthermore, the stronger descriptors of scenes are constructed by feature fusion. The highlight of our proposed network is that it is capable to localize discriminative regions from an image in a data-driven manner without any additional supervision. We conduct experiments on a subset of the Places205 database to evaluate the performance of the proposed basic network and the involved parameters. Our model achieves state-of-the-art top-1 accuracy 82.10% on the evaluation dataset comparing with fine-tuned PlacesCNN (80.98%). We find that our model is able to learn informative attention regions for discriminating scene categories.
Shuxuan Guo, Li Liu 0002, Wei Wang 0115, Songyang Lao, Liang Wang 0001
ICPR3
2016 CNN based suburban building detection using monocular high resolution Google Earth images
abstract
This paper proposes a deep convolutional neural networks (CNNs) based method to automatically detect suburban buildings from high resolution Google Earth imagery. Traditional methods based on low-level hand-engineered features or mid-level bag of features have great limitations in complex environment, especially in suburban areas. Inspired by the astounding achievement of CNNs in object recognition and detection, we develop a novel method to detect buildings in cluttered images which consists of three main steps. Firstly, a multi-scale saliency computation is employed to extract built-up areas and a sliding windows approach is applied to generate candidate regions. Then, a CNN is applied to classify the regions. Finally, an improved non maximum suppression is used to remove false buildings. We test our method on a collection of very challenging Google Earth images and achieve 89% precision, which shows robustness and efficiency of our method.
Qinchuan Zhang, Yunhong Wang 0001, Qingjie Liu 0001, Wei Wang 0115
IGARSS5
2016 Joint Feature Selection and Subspace Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval has recently drawn much attention due to the widespread existence of multimodal data. It takes one type of data as the query to retrieve relevant data objects of another type, and generally involves two basic problems: the measure of relevance and coupled feature selection. Most previous methods just focus on solving the first problem. In this paper, we aim to deal with both problems in a novel joint learning framework. To address the first problem, we learn projection matrices to map multimodal data into a common subspace, in which the similarity between different modalities of data can be measured. In the learning procedure, the l21-norm penalties are imposed on the projection matrices separately to solve the second problem, which selects relevant and discriminative features from different feature spaces simultaneously. A multimodal graph regularization term is further imposed on the projected data,which preserves the inter-modality and intra-modality similarity relationships.An iterative algorithm is presented to solve the proposed joint learning problem, along with its convergence analysis. Experimental results on cross-modal retrieval tasks demonstrate that the proposed method outperforms the state-of-the-art subspace approaches.
Kaiye Wang, Ran He 0001, Liang Wang 0001, Wei Wang 0115, Tieniu Tan
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Hierarchical recurrent neural network for skeleton based action recognition
abstract
Human actions can be represented by the trajectories of skeleton joints. Traditional methods generally model the spatial structure and temporal dynamics of human skeleton with hand-crafted features and recognize human actions by well-designed classifiers. In this paper, considering that recurrent neural network (RNN) can model the long-term contextual information of temporal sequences well, we propose an end-to-end hierarchical RNN for skeleton based action recognition. Instead of taking the whole skeleton as the input, we divide the human skeleton into five parts according to human physical structure, and then separately feed them to five subnets. As the number of layers increases, the representations extracted by the subnets are hierarchically fused to be the inputs of higher layers. The final representations of the skeleton sequences are fed into a single-layer perceptron, and the temporally accumulated output of the perceptron is the final decision. We compare with five other deep RNN architectures derived from our model to verify the effectiveness of the proposed network, and also compare with several other methods on three publicly available datasets. Experimental results demonstrate that our model achieves the state-of-the-art performance with high computational efficiency.
Wei Wang 0115, Liang Wang 0001
CVPR2
2015 Conditional High-Order Boltzmann Machine: A Supervised Learning Model for Relation Learning
abstract
Relation learning is a fundamental operation in many computer vision tasks. Recently, high-order Boltzmann machine and its variants have exhibited the great power of modelling various data relation. However, most of them are unsupervised learning models which are not very discriminative and thus cannot server as a standalone solution to relation learning tasks. In this paper, we explore supervised learning algorithms and propose a new model named Conditional High-order Boltzmann Machine (CHBM), which can be directly used as a bilinear classifier to assign similarity scores for pairwise images. Then, to better deal with complex data relation, we propose a gated version of CHBM which untangles factors of variation by exploiting a set of latent variables to gate classification. We perform four-order tensor factorization for parameter reduction, and present two efficient supervised learning algorithms from the perspectives of being generative and discriminative, respectively. The experimental results of image transformation visualization, binary-way classification and face verification demonstrate that, by performing supervised learning, our models can greatly improve the performance.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
ICCV2
2015 Learning unified sparse representations for multi-modal data
abstract
Cross-modal retrieval has become one of interesting and important research problem recently, where users can take one modality of data (e.g., text, image or video) as the query to retrieve relevant data of another modality. In this paper, we present a Multi-modal Unified Representation Learning (MURL) algorithm for cross-modal retrieval, which learns unified sparse representations for multi-modal data representing the same semantics via joint dictionary learning. The ℓ1-norm is imposed on the unified representations to explicitly encourage sparsity, which makes our algorithm more robust. Furthermore, a constraint regularization term is imposed to force the representations to be similar if their corresponding multi-modal data have must-links or to be far apart if their corresponding multi-modal data have cannot-links. An iterative algorithm is also proposed to solve the objective function. The effectiveness of the proposed method is verified by extensive results on two real-world datasets.
Kaiye Wang, Wei Wang 0115, Liang Wang 0001
ICIP2
2015 Scene text recognition with deeper convolutional neural networks
abstract
Scene text recognition plays an important role in many applications such as video indexing and house number localization in maps. Recently, some feature learning methods have been proposed to handle this problem, which often exploit deep architectures with no more than 5 layers and relatively large receptive fields. Meanwhile, to avoid model overfitting, they generally take advantage of large amount of additional data. Inspired by the great success of GoogleLeNet with a deeper network and VGG networks with smaller receptive fields in the ImageNet competition, in this paper, we adopt a much deeper network with up to 15 layers and smaller receptive fields (3×3) to learn better features for scene text recognition. Particularly, even without additional training data, our model can achieve better performance. Experiments on scene text datasets (ICDAR 2003, SVT, Chars74K) demonstrate that our method achieves the state-of-the-art performance on character classification and competitive performance on cropped word recognition.
Yuqi Zhang 0001, Wei Wang 0115, Liang Wang 0001, Liuan Wang
ICIP2
2015 A Two-step Approach to Cross-modal Hashing
abstract
With the rapid growth of multimedia data, it is very desirable to effectively and efficiently search objects of interest across different modalities from large scale databases. Cross-modal hashing provides a very promising way to address such problem. In this paper, we propose a two-step cross-modal hashing approach to obtain compact hash codes and learn hash functions from multimodal data. Our approach decomposes the cross-modal hashing problem into two steps: generating hash code and learning hash function. In the first step, we obtain the hash codes for all modalities of data via a joint multi-modal graph, which takes into consideration both the intra-modality and inter-modality similarity. In the second step, learning hashing function is formulated as a binary classification problem. We train binary classifiers to predict the hash code for any data object unseen before. Experimental results on two cross-modal datasets show the effectiveness of our proposed approach.
Kaiye Wang, Wei Wang 0115, Liang Wang 0001, Ran He 0001
ICMR2
2015 Bidirectional Recurrent Convolutional Networks for Multi-Frame Super-Resolution
abstract
Super resolving a low-resolution video is usually handled by either single-image super-resolution (SR) or multi-frame SR. Single-Image SR deals with each video frame independently, and ignores intrinsic temporal dependency of video frames which actually plays a very important role in video super-resolution. Multi-Frame SR generally extracts motion information, e.g. optical flow, to model the temporal dependency, which often shows high computational cost. Considering that recurrent neural network (RNN) can model long-term contextual information of temporal sequences well, we propose a bidirectional recurrent convolutional network for efficient multi-frame SR.Different from vanilla RNN, 1) the commonly-used recurrent full connections are replaced with weight-sharing convolutional connections and 2) conditional convolutional connections from previous input layers to current hidden layer are added for enhancing visual-temporal dependency modelling. With the powerful temporal dependency modelling, our model can super resolve videos with complex motions and achieve state-of-the-art performance. Due to the cheap convolution operations, our model has a low computational complexity and runs orders of magnitude faster than other multi-frame methods.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
NIPS2
2015 Unconstrained Multimodal Multi-Label Learning
abstract
Multimodal learning has been mostly studied by assuming that multiple label assignments are independent of each other and all the modalities are available. In this paper, we consider a more general problem where the labels contain dependency relationships and some modalities are likely to be missing. To this end, we propose a multi-label conditional restricted Boltzmann machine (ML-CRBM), which handles modality completion , fusion, and multi-label prediction in a unified framework. The proposed model is able to generate missing modalities based on observed ones, by explicitly modelling and sampling their conditional distributions. After that, it can discriminatively fuse multiple modalities to obtain shared representations under the supervision of class labels. To consider the co-occurrence of the labels, the proposed model formulates the multi-label prediction as a max-margin-based multi-task learning problem. Model parameters can be jointly learned by seeking a balance between being generative for modality generation and being discriminative for label prediction. We perform a series of experiments in terms of classification, visualization, and retrieval, and the experimental results clearly demonstrate the effectiveness of our method.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
IEEE Trans. Multim.2
2014 Deep Embedding Network for Clustering
abstract
Clustering is a fundamental technique widely used for exploring the inherent data structure in pattern recognition and machine learning. Most of the existing methods focus on modeling the similarity/dissimilarity relationship among instances, such as k-means and spectral clustering, and ignore to extract more effective representation for clustering. In this paper, we propose a deep embedding network for representation learning, which is more beneficial for clustering by considering two constraints on learned representations. We first utilize a deep auto encoder to learn the reduced representations from the raw data. To make the learned representations suitable for clustering, we first impose a locality-persevering constraint on the learned representations, which aims to embed original data into its underlying manifold space. Then, different from spectral clustering which extracts representations from the block diagonal similarity matrix, we apply a group sparsity constraint for the learned representations, and aim to learn block diagonal representations in which the nonzero groups correspond to its cluster. After obtaining the learned representations, we use k-means to cluster them. To evaluate the proposed deep embedding network, we compare its performance with k-means and spectral clustering on three commonly-used datasets. The experiments demonstrate that the proposed method achieves promising performance.
Peihao Huang, Yan Huang 0008, Wei Wang 0115, Liang Wang 0001
ICPR3
2014 A General Nonlinear Embedding Framework Based on Deep Neural Network
abstract
Recently there has been increasing interest in deep neural network due to its powerful represent ability in several successful applications such as speech recognition and image classification. In this paper, we propose a general nonlinear embedding framework based on deep neural network which can be utilized to implement a family of dimensionality reduction algorithms. The objective function of our framework consists of two terms: 1) an embedding term which transforms the input to a low-dimensional representation with a multilayer network, and 2) a regularization term which computes the reconstruction error of the original input by unrolling the multilayer network to a deep auto encoder. We adopt a layer-by-layer pretraining procedure to obtain good initial weights for the network, and then minimize the objective function by back propagating derivatives of the two terms. To evaluate the proposed framework, we perform face recognition and digit classification experiments. The experiments demonstrate that the proposed framework achieves better results than the state-of-the-art algorithms. The success of our framework further verifies deep neural network's advantages in representation learning.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ICPR2
2013 Learning Coupled Feature Spaces for Cross-Modal Matching
abstract
Cross-modal matching has recently drawn much attention due to the widespread existence of multimodal data. It aims to match data from different modalities, and generally involves two basic problems: the measure of relevance and coupled feature selection. Most previous works mainly focus on solving the first problem. In this paper, we propose a novel coupled linear regression framework to deal with both problems. Our method learns two projection matrices to map multimodal data into a common feature space, in which cross-modal data matching can be performed. And in the learning procedure, the ell_21-norm penalties are imposed on the two projection matrices separately, which leads to select relevant and discriminative features from coupled feature spaces simultaneously. A trace norm is further imposed on the projected data as a low-rank constraint, which enhances the relevance of different modal data with connections. We also present an iterative algorithm based on half-quadratic minimization to solve the proposed regularized linear regression problem. The experimental results on two challenging cross-modal datasets demonstrate that the proposed method outperforms the state-of-the-art approaches.
Kaiye Wang, Ran He 0001, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ICCV3
2013 Multi-task deep neural network for multi-label learning
abstract
This paper proposes a multi-task deep neural network (MT-DNN) architecture to handle the multi-label learning problem, in which each label learning is defined as a binary classification task, i.e., a positive class for “an instance owns this label” and a negative class for “an instance does not own this label”. Multi-label learning is accordingly transformed to multiple binary-class classification tasks. Considering that a deep neural nets (DNN) architecture can learn good intermediate representations shared across tasks, we generalize one classification task of traditional DNN into multiple binary classification tasks through defining the output layer with a negative class node and a positive class node for each label. After a similar pretraining process to deep belief nets, we redefine the label assignment error of MT-DNN and perform the back-propagation algorithm to fine-tune the network. To evaluate the proposed model, we carry out image annotation experiments on two public image datasets, with 2000 images and 30,000 images respectively. The experiments demonstrate that the proposed model achieves the state-of-the-art performance.
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ICIP2
2012 Baseline Results for Violence Detection in Still Images
abstract
Recognizing objectionable content draws more and more attention nowadays given the rapid proliferation of images and videos on the Internet. Although there are some investigations about violence video detection and pornographic information filtering, very few existing methods touch on the problem of violence detection in still images. However, given its potential use in violence webpage filtering, online public opinion monitoring and some other aspects, recognizing violence in still images is worth being deeply investigated. To this end, we first establish a new database containing 500 violence images and 1500 non-violence images. And we use the Bag-of-Words (BoW) model which is frequently adopted in image classification domain to discriminate violence images and non-violence images. The effectiveness of four different feature representations are tested within the BoW framework. Finally the baseline results for violence image detection on our newly built database are reported.
Dong Wang 0004, Zhang Zhang 0001, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
AVSS3
2012 An effective regional saliency model based on extended site entropy rate
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan
ICPR2
2011 Simulating human saccadic scanpaths on natural images
abstract
Human saccade is a dynamic process of information pursuit. Based on the principle of information maximization, we propose a computational model to simulate human saccadic scanpaths on natural images. The model integrates three related factors as driven forces to guide eye movements sequentially - reference sensory responses, fovea-periphery resolution discrepancy, and visual working memory. For each eye movement, we compute three multi-band filter response maps as a coherent representation for the three factors. The three filter response maps are combined into multi-band residual filter response maps, on which we compute residual perceptual information (RPI) at each location. The RPI map is a dynamic saliency map varying along with eye movements. The next fixation is selected as the location with the maximal RPI value. On a natural image dataset, we compare the saccadic scanpaths generated by the proposed model and several other visual saliency-based models against human eye movement data. Experimental results demonstrate that the proposed model achieves the best prediction accuracy on both static fixation locations and dynamic scanpaths.
Wei Wang 0115, Cheng Chen 0004, Yizhou Wang 0001, Tingting Jiang 0001, Fang Fang 0003, Yuan Yao 0011
CVPR1
2010 Measuring visual saliency by Site Entropy Rate
abstract
In this paper, we propose a new computational model for visual saliency derived from the information maximization principle. The model is inspired by a few well acknowledged biological facts. To compute the saliency spots of an image, the model first extracts a number of sub-band feature maps using learned sparse codes. It adopts a fully-connected graph representation for each feature map, and runs random walks on the graphs to simulate the signal/information transmission among the interconnected neurons. We propose a new visual saliency measure called Site Entropy Rate (SER) to compute the average information transmitted from a node (neuron) to all the others during the random walk on the graphs/network. This saliency definition also explains the center-surround mechanism from computation aspect. We further extend our model to spatial-temporal domain so as to detect salient spots in videos. To evaluate the proposed model, we do extensive experiments on psychological stimuli, two well known image data sets, as well as a public video dataset. The experiments demonstrate encouraging results that the proposed model achieves the state-of-the-art performance of saliency detection in both still images and videos.
Wei Wang 0115, Yizhou Wang 0001, Qingming Huang, Wen Gao 0001
CVPR1
2008 Symmetric segment-based stereo matching of motion blurred images with illumination variations
abstract
Most existing methods of stereo matching focus on dealing with clear image pairs. Consequently, there is a lack of approaches capable of handling degraded images captured under challenging real situations, e.g. motion blur is present and an image pair is in different illumination conditions. In this paper we propose a novel approach to handling these challenging situations by formulating the problem into a Maximum a Posteriori (MAP) estimation framework, and adopt a segment-based symmetric stereo matching method to infer a mask of disparity map which indicates whether a disparity is affected by motion blur and estimate the disparity value. The experimental results show that our stereo matching method is able to compute more accurate disparity maps of this type of degraded images.
Wei Wang 0115, Yizhou Wang 0001, Longshe Huo, Qingming Huang, Wen Gao 0001
ICPR1
2005 Double layer sliding mode control for second-order underactuated mechanical systems
abstract
A new stable sliding mode control method for a class of underactuated mechanical systems is proposed in this paper. The controller has the double-layer structure. Firstly, the system states are divided into several different subsystems. For each of these subsystems, a first-layer sliding plane is constructed. From these first-layer sliding planes, then we further construct a second-layer sliding plane. By analyzing the features of the mathematical model of the underactuated mechanical systems, we derive the sliding-mode control law and indicate the ranges of the controller parameters. Using Lyapunov law, the paper proves the stability of all the sliding planes theoretically. The simulation results show the validity of this method for this class of underactuated mechanical systems.
Wei Wang 0115, Jianqiang Yi, Dongbin Zhao
IROS1
2005 Cascade sliding-mode controller for large-scale underactuated systems
abstract
On the basis of sliding mode control, a new cascade sliding-mode controller (CSMC) for a class of large-scale underactuated systems is proposed. The large-scale underactuated systems include several subsystems. Firstly, two states are chosen to construct the first-layer sliding surface. Secondly, the first-layer sliding surface and one of the left states are used to construct the second-layer sliding surface. This process continues till the last-layer sliding surface is obtained. By theoretical analysis, the cascade sliding-mode controller is proved to be globally stable in the sense that all signals involved are bounded. The simulation results show the validity of this method.
Jianqiang Yi, Wei Wang 0115, Dongbin Zhao
IROS2