Bing Shuai

dblp:118/3676 · DBLP profile ↗
← Back
35ranked-venue papers
11as first author
13since 2021 · last 2024
0000-0001-5811-9950ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 10 first-author · 12 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 11 since 2021
YearPublicationVenuePosition
2024 Self-Supervised Multi-Object Tracking with Path Consistency
abstract
In this paper, we propose a novel concept of path consis-tency to learn robust object matching without using manual object identity supervision. Our key idea is that, to track a object through frames, we can obtain multiple different as-sociation results from a model by varying the frames it can observe, i.e., skipping frames in observation. As the differ-ences in observations do not alter the identities of objects, the obtained association results should be consistent. Based on this rationale, we generate multiple observation paths, each specifying a different set of frames to be skipped, and formulate the Path Consistency Loss that enforces the as-sociation results are consistent across different observation paths. We use the proposed loss to train our object matching model with only self-supervision. By extensive experiments on three tracking datasets (MOT17, PersonPath22, KITTI), we demonstrate that our method outperforms existing unsu-pervised methods with consistent margins on various eval-uation metrics, and even achieves performance close to su-pervised methods.
Zijia Lu, Bing Shuai, Yanbei Chen, Zhenlin Xu, Davide Modolo
CVPR2
2024 Robustness Preserving Fine-Tuning Using Neuron Importance
Guangrui Li 0005, Rahul Duggal, Aaditya Singh, Kaustav Kundu, Bing Shuai, Jon Wu
ECCV (68)5
2023 SkeleTR: Towards Skeleton-based Action Recognition in the Wild
abstract
We present SkeleTR, a new framework for skeleton-based action recognition. In contrast to prior work, which focuses mainly on controlled environments, we target more general scenarios that typically involve a variable number of people and various forms of interaction between people. SkeleTR works with a two-stage paradigm. It first models the intra-person skeleton dynamics for each skeleton sequence with graph convolutions, and then uses stacked Transformer encoders to capture person interactions that are important for action recognition in general scenarios. To mitigate the negative impact of inaccurate skeleton associations, SkeleTR takes relative short skeleton sequences as input and increases the number of sequences. As a unified solution, SkeleTR can be directly applied to multiple skeleton-based action tasks, including video-level action classification, instance-level action detection, and group-level activity recognition. It also enables transfer learning and joint training across different action tasks and datasets, which result in performance improvement. When evaluated on various skeleton-based action recognition benchmarks, SkeleTR achieves the state-of-the-art performance.
Haodong Duan, Bing Shuai, Davide Modolo, Zhuowen Tu, Joseph Tighe, Alessandro Bergamo
ICCV3
2023 MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation
abstract
Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently processing long-form videos (> 60min). In this paper, we introduce Multimodal alignmEnt aGgregation and distillAtion (MEGA) for cinematic long-video segmentation. MEGA tackles the challenge by leveraging multiple media modalities. The method coarsely aligns inputs of variable lengths and different modalities with alignment positional encoding. To maintain temporal synchronization while reducing computation, we further introduce an enhanced bottleneck fusion layer which uses temporal alignment. Additionally, MEGA employs a novel contrastive loss to synchronize and transfer labels across modalities, enabling act segmentation from labeled synopsis sentences on video shots. Our experimental results show that MEGA outperforms state-of-the-art methods on MovieNet dataset for scene segmentation (with an Average Precision improvement of +1.19%) and on TRIPOD dataset for act segmentation (with a Total Agreement improvement of +5.51%).
Najmeh Sadoughi, Xinyu Li 0003, Avijit Vajpayee, David Fan 0001, Bing Shuai, Hector J. Santos-Villalobos, Vimal Bhat, Rohith MV
ICCV5
2023 Object-Centric Multiple Object Tracking
abstract
Unsupervised object-centric learning methods allow the partitioning of scenes into entities without additional localization information and are excellent candidates for reducing the annotation burden of multiple-object tracking (MOT) pipelines. Unfortunately, they lack two key properties: objects are often split into parts and are not consistently tracked over time. In fact, state-of-the-art models achieve pixel-level accuracy and temporal consistency by relying on supervised object detection with additional ID labels for the association through time. This paper proposes a video object-centric model for MOT. It consists of an index-merge module that adapts the object-centric slots into detection outputs and an object memory module that builds complete object prototypes to handle occlusions. Benefited from object-centric learning, we only require sparse detection labels (0%-6.25%) for object localization and feature binding. Relying on our self-supervised Expectation-Maximization-inspired loss for object association, our approach requires no ID labels. Our experiments significantly narrow the gap between the existing object-centric model and the fully supervised state-of-the-art and outperform several unsupervised trackers. Code is available at https://github.com/amazon-science/object-centric-multiple-object-tracking.
Max Horn, Yizhuo Ding, Tong He 0002, Zechen Bai, Dominik Zietlow, Carl-Johann Simon-Gabriel, Bing Shuai, Zhuowen Tu, Thomas Brox, Bernt Schiele, Yanwei Fu 0001, Francesco Locatello, Zheng Zhang 0001, Tianjun Xiao
ICCV9
2022 Id-Free Person Similarity Learning
abstract
Learning a unified person detection and re-identification model is a key component of modern trackers. However, training such models usually relies on the availability of training images / videos that are manually labeled with both person boxes and their identities. In this work, we explore training such a model by only using person box annotations, thus removing the necessity of manually labeling a training dataset with additional person identity annotation as these are expensive to collect. To this end, we present a contrastive learning framework to learn person similarity without using manually labeled identity annotations. First, we apply image-level augmentation to images on public person detection datasets, based on which we learn a strong model for general person detection as well as for short-term person re-identification. To learn a model capable of longerterm re-identification, we leverage the natural appearance evolution of each person in videos to serve as instance-level appearance augmentation in our contrastive loss formulation. Without access to the target dataset or person identity annotation, our model achieves competitive results compared to existing fully-supervised state-of-the-art methods on both person search and person tracking tasks. Our model also shows promising results for saving the annotation cost that is needed to achieve a certain level of performance on the person search task.
Bing Shuai, Xinyu Li 0003, Kaustav Kundu, Joseph Tighe
CVPR1
2022 TubeR: Tubelet Transformer for Video Action Detection
abstract
We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an offline actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in a video by simultaneously performing action localization and recognition from a single representation. TubeR learns a set of tubelet-queries and utilizes a tubelet-attention module to model the dynamic spatio-temporal nature of a video clip, which effectively reinforces the model capacity compared to using actor-positional hypotheses in the spatio-temporal space. For videos containing transitional states or scene changes, we propose a context aware classification head to utilize short-term and long-term context to strengthen action classification, and an action switch regression head for detecting the precise temporal action extent. TubeR directly produces action tubelets with variable lengths and even maintains good results for long video clips. TubeR outperforms the previous state-of-the-art on commonly used action detection datasets AVA, UCF101-24 and JHMDB51-21. Code will be available on GluonCV(https://cv.gluon.ai/).
Jiaojiao Zhao, Yanyi Zhang, Xinyu Li 0003, Hao Chen 0024, Bing Shuai, Chunhui Liu 0002, Kaustav Kundu, Yuanjun Xiong, Davide Modolo, Ivan Marsic, Cees Snoek, Joseph Tighe
CVPR5
2022 Large Scale Real-World Multi-person Tracking
Bing Shuai, Alessandro Bergamo, Uta Büchler, Andrew G. Berneshawi, Alyssa Boden, Joseph Tighe
ECCV (8)1
2022 An In-depth Study of Stochastic Backpropagation
abstract
In this paper, we provide an in-depth study of Stochastic Backpropagation (SBP) when training deep neural networks for standard image classification and object detection tasks. During backward propagation, SBP calculates gradients by using only a subset of feature maps to save GPU memory and computational cost. We interpret SBP as an efficient way to implement stochastic gradient decent by performing backpropagation dropout, which leads to significant memory saving and training run-time reduction, with a minimal impact on the overall model accuracy. We offer best practices to apply SBP for training image recognition models, which can be adopted in learning a wide range of deep neural networks. Experiments on image classification and object detection show that SBP can save up to 40% of GPU memory with less than 1% accuracy degradation. Code is available at: https://github.com/amazon-research/stochastic-backpropagation
Hao Chen 0024, Bing Shuai, Zhuowen Tu, Joseph Tighe
NeurIPS4
2022 NUTA: Non-uniform Temporal Aggregation for Action Recognition
abstract
In the world of action recognition research, one primary focus has been on how to construct and train networks to model the spatial-temporal volume of an input video. These methods typically uniformly sample a segment of an input clip (along the temporal dimension). However, not all parts of a video are equally important to determine the action in the clip. In this work, we focus instead on learning where to extract features, so as to focus on the most informative parts of the video. We propose a method called the non-uniform temporal aggregation (NUTA), which aggregates features only from informative temporal segments. We also introduce a synchronization method that allows our NUTA features to be temporally aligned with traditional uniformly sampled video features, so that both local and clip-level features can be combined. Our model has achieved state-of-the-art performance on four widely used large-scale action-recognition datasets (Kinetics400, Kinetics700, Something-something V2 and Charades). In addition, we have created a visualization to illustrate how the proposed NUTA method selects only the most relevant parts of a video clip.
Xinyu Li 0003, Chunhui Liu 0002, Bing Shuai, Yi Zhu 0001, Hao Chen 0024, Joseph Tighe
WACV3
2021 SiamMOT: Siamese Multi-Object Tracking
abstract
In this paper, we focus on improving online multi-object tracking (MOT). In particular, we introduce a region-based Siamese Multi-Object Tracking network, which we name SiamMOT. SiamMOT includes a motion model that estimates the instance’s movement between two frames such that detected instances are associated. To explore how the motion modelling affects its tracking capability, we present two variants of Siamese tracker, one that implicitly models motion and one that models it explicitly. We carry out extensive quantitative experiments on three different MOT datasets: MOT17, TAO-person and Caltech Roadside Pedestrians, showing the importance of motion modelling for MOT and the ability of SiamMOT to substantially outperform the state-of-the-art. Finally, SiamMOT also outperforms the winners of ACM MM’20 HiEve Grand Challenge on HiEve dataset. Moreover, SiamMOT is efficient, and it runs at 17 FPS for 720P videos on a single modern GPU.
Bing Shuai, Andrew G. Berneshawi, Xinyu Li 0003, Davide Modolo, Joseph Tighe
CVPR1
2021 VidTr: Video Transformer Without Convolutions
abstract
We introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatiotemporal information via stacked attentions and provide better performance with higher efficiency. We first introduce the vanilla video transformer and show that transformer module is able to perform spatio-temporal modeling from raw pixels, but with heavy memory usage. We then present VidTr which reduces the memory cost by 3.3× while keeping the same performance. To further optimize the model, we propose the standard deviation based topK pooling for attention (pooltopK_std), which reduces the computation by dropping non-informative features along temporal dimension. VidTr achieves state-of-the-art performance on five commonly used datasets with lower computational requirement, showing both the efficiency and effectiveness of our design. Finally, error analysis and visualization show that VidTr is especially good at predicting actions that require long-term temporal reasoning.
Yanyi Zhang, Xinyu Li 0003, Chunhui Liu 0002, Bing Shuai, Yi Zhu 0001, Biagio Brattoli, Hao Chen 0024, Ivan Marsic, Joseph Tighe
ICCV4
2021 Understanding the impact of mistakes on background regions in crowd counting
abstract
In crowd counting we often observe wrong predictions on image regions not containing any person. But how often do these mistakes happen and how much do they affect the overall performance? In this paper we analyze this problem in depth and present an extensive analysis on five of the most important crowd counting datasets. We present this analysis in two parts. First, we quantify the number of mistakes made. Our results show that (i) mistakes on back-ground are substantial and they are responsible for 18-49% of the total error, (ii) models do not generalize well to different kinds of backgrounds and perform poorly on completely background images, and (iii) models make many more mistakes than those captured by the standard Mean Absolute Error (MAE) metric, as counting on background compensates substantially for misses on foreground. And second, we quantify the performance change gained by helping the model better deal with this problem. We enrich a popular crowd counting network with a segmentation branch trained to suppress background predictions. This simple addition (i) reduces background error by 10-83%, (ii) reduces fore-ground error by up to 26% and (iii) improves overall crowd counting performance up to 20%. When compared against the literature, this simple technique achieves very competitive results on all datasets, showing the importance of tack-ling the background problem.
Davide Modolo, Bing Shuai, Rahul Rama Varior, Joseph Tighe
WACV2
2020 Directional Temporal Modeling for Action Recognition
Xinyu Li 0003, Bing Shuai, Joseph Tighe
ECCV (6)2
2020 Application of Multi-Object Tracking with Siamese Track-RCNN to the Human in Events Dataset
abstract
Multi-object tracking systems often consist of a combination of a detector, a short term linker, a re-identification feature extractor and a solver that takes the output from these separate components and makes a final prediction. Differently, this work aims to unify all these in a single tracking system. Towards this, we propose Siamese Track-RCNN, a two stage detect-and-track framework which consists of three functional branches: (1) the detection branch localizes object instances; (2) the Siamese-based track branch estimates the object motion and (3) the object re-identification branch re-activates the previously terminated tracks when they re-emerge. We used this design and apply it to the Human in Events dataset.
Bing Shuai, Andrew G. Berneshawi, Manchen Wang, Chunhui Liu 0002, Davide Modolo, Xinyu Li 0003, Joseph Tighe
ACM Multimedia1
2020 Semantic Segmentation With Context Encoding and Multi-Path Decoding
abstract
Semantic image segmentation aims to classify every pixel of a scene image to one of many classes. It implicitly involves object recognition, localization, and boundary delineation. In this paper, we propose a segmentation network called CGBNet to enhance the paring results by context encoding and multi-path decoding. We first propose a context encoding module that generates context contrasted local feature to make use of the informative context and the discriminative local information. This context encoding module greatly improves the segmentation performance, especially for inconspicuous objects. Furthermore, we propose a scale-selection scheme to selectively fuse the parsing results from different-scales of features at every spatial position. It adaptively selects appropriate score maps from rich scales of features. To improve the parsing results of boundary, we further propose a boundary delineation module that encourages the location-specific very-low-level feature near the boundaries to take part in the final prediction and suppresses them far from the boundaries. Without bells and whistles, the proposed segmentation network achieves very competitive performance in terms of all three different evaluation metrics consistently on the four popular scene segmentation datasets, Pascal Context, SUN-RGBD, Sift Flow, and COCO Stuff.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
IEEE Trans. Image Process.3
2019 Semantic Correlation Promoted Shape-Variant Context for Segmentation
abstract
Context is essential for semantic segmentation. Due to the diverse shapes of objects and their complex layout in various scene images, the spatial scales and shapes of contexts for different objects have very large variation. It is thus ineffective or inefficient to aggregate various context information from a predefined fixed region. In this work, we propose to generate a scale- and shape-variant semantic mask for each pixel to confine its contextual region. To this end, we first propose a novel paired convolution to infer the semantic correlation of the pair and based on that to generate a shape mask. Using the inferred spatial scope of the contextual region, we propose a shape-variant convolution, of which the receptive field is controlled by the shape mask that varies with the appearance of input. In this way, the proposed network aggregates the context information of a pixel from its semantic-correlated region instead of a predefined fixed region. Furthermore, this work also proposes a labeling denoising model to reduce wrong predictions caused by the noisy low-level features. Without bells and whistles, the proposed segmentation network achieves new state-of-the-arts consistently on the six public segmentation datasets.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
CVPR3
2019 Toward Achieving Robust Low-Level and High-Level Scene Parsing
abstract
In this paper, we address the challenging task of scene segmentation. We first discuss and compare two widely used approaches to retain detailed spatial information from pretrained CNN - "dilation" and "skip". Then, we demonstrate that the parsing performance of "skip" network can be noticeably improved by modifying the parameterization of skip layers. Furthermore, we introduce a "dense skip" architecture to retain a rich set of low-level information from pre-trained CNN, which is essential to improve the low-level parsing performance. Meanwhile, we propose a convolutional context network (CCN) and place it on top of pre-trained CNNs, which is used to aggregate contexts for high-level feature maps so that robust high-level parsing can be achieved. We name our segmentation network enhanced fully convolutional network (EFCN) based on its significantly enhanced structure over FCN. Extensive experimental studies justify each contribution separately. Without bells and whistles, EFCN achieves state-of-the-arts on segmentation datasets of ADE20K, Pascal Context, SUN-RGBD and Pascal VOC 2012.
Bing Shuai, Henghui Ding, Ting Liu 0009, Gang Wang 0012, Xudong Jiang 0001
IEEE Trans. Image Process.1
2018 Improving Fast Segmentation With Teacher-Student Learning
Jiafeng Xie, Bing Shuai, Jianfang Hu, Wei-Shi Zheng 0001
BMVC2
2018 Context Contrasted Feature and Gated Multi-Scale Aggregation for Scene Segmentation
abstract
Scene segmentation is a challenging task as it need label every pixel in the image. It is crucial to exploit discriminative context and aggregate multi-scale features to achieve better segmentation. In this paper, we first propose a novel context contrasted local feature that not only leverages the informative context but also spotlights the local information in contrast to the context. The proposed context contrasted local feature greatly improves the parsing performance, especially for inconspicuous objects and background stuff. Furthermore, we propose a scheme of gated sum to selectively aggregate multi-scale features for each spatial position. The gates in this scheme control the information flow of different scale features. Their values are generated from the testing image by the proposed network learnt from the training data so that they are adaptive not only to the training data, but also to the specific testing image. Without bells and whistles, the proposed approach achieves the state-of-the-arts consistently on the three popular scene segmentation datasets, Pascal Context, SUN-RGBD and COCO Stuff.
Henghui Ding, Xudong Jiang 0001, Bing Shuai, Ai Qun Liu, Gang Wang 0012
CVPR3
2018 Scene Segmentation with DAG-Recurrent Neural Networks
abstract
In this paper, we address the challenging task of scene segmentation. In order to capture the rich contextual dependencies over image regions, we propose Directed Acyclic Graph-Recurrent Neural Networks (DAG-RNN) to perform context aggregation over locally connected feature maps. More specifically, DAG-RNN is placed on top of pre-trained CNN (feature extractor) to embed context into local features so that their representative capability can be enhanced. In comparison with plain CNN (as in Fully Convolutional Networks-FCN), DAG-RNN is empirically found to be significantly more effective at aggregating context. Therefore, DAG-RNN demonstrates noticeably performance superiority over FCNs on scene segmentation. Besides, DAG-RNN entails dramatically less parameters as well as demands fewer computation operations, which makes DAG-RNN more favorable to be potentially applied on resource-constrained embedded devices. Meanwhile, the class occurrence frequencies are extremely imbalanced in scene segmentation, so we propose a novel class-weighted loss to train the segmentation network. The loss distributes reasonably higher attention weights to infrequent classes during network training, which is essential to boost their parsing performance. We evaluate our segmentation network on three challenging public scene segmentation benchmarks: Sift Flow, Pascal Context and COCO Stuff. On top of them, we achieve very impressive segmentation performance.
Bing Shuai, Zhen Zuo, Bing Wang 0003, Gang Wang 0012
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Recent advances in convolutional neural networks
Jiuxiang Gu, Zhenhua Wang 0002, Jason Kuen, Lianyang Ma, Amir Shahroudy, Bing Shuai, Ting Liu 0009, Gang Wang 0012, Jianfei Cai 0001, Tsuhan Chen
Pattern Recognit.6
2018 Multimodal Recurrent Neural Networks With Information Transfer Layers for Indoor Scene Labeling
abstract
This paper proposes a new method called multimodal recurrent neural networks (RNNs) for RGB-D scene semantic segmentation. It is optimized to classify image pixels given two input sources: RGB color channels and depth maps. It simultaneously performs training of two RNNs that are crossly connected through information transfer layers, which are learnt to adaptively extract relevant cross-modality features. Each RNN model learns its representations from its own previous hidden states and transferred patterns from the other RNNs previous hidden states; thus, both model-specific and cross-modality features are retained. We exploit the structure of quad-directional 2D-RNNs to model the short- and long-range contextual information in the 2D input image. We carefully designed various baselines to efficiently examine our proposed model structure. We test our multimodal RNNs method on popular RGB-D benchmarks and show how it outperforms previous methods significantly and achieves competitive results with other state-of-the-art works.
Abrar H. Abdulnabi, Bing Shuai, Zhen Zuo, Lap-Pui Chau, Gang Wang 0012
IEEE Trans. Multim.2
2017 Episodic CAMN: Contextual Attention-Based Memory Networks with Iterative Feedback for Scene Labeling
abstract
Scene labeling can be seen as a sequence-sequence prediction task (pixels-labels), and it is quite important to leverage relevant context to enhance the performance of pixel classification. In this paper, we introduce an episodic attention-based memory network to achieve the goal. We present a unified framework that mainly consists of a Convolutional Neural Network (CNN), specifically, Fully Convolutional Network (FCN) and an attention-based memory module with feedback connections to perform context selection and refinement. The full model produces context-aware representation for each target patch by aggregating the activated context and its original local representation produced by the convolution layers. We evaluate our model on PASCAL Context, SIFT Flow and PASCAL VOC 2011 datasets and achieve competitive results to other state-of-the-art methods in scene labeling.
Abrar H. Abdulnabi, Bing Shuai, Gang Wang 0012
CVPR2
2017 Deep Level Sets for Salient Object Detection
abstract
Deep learning has been applied to saliency detection in recent years. The superior performance has proved that deep networks can model the semantic properties of salient objects. Yet it is difficult for a deep network to discriminate pixels belonging to similar receptive fields around the object boundaries, thus deep networks may output maps with blurred saliency and inaccurate boundaries. To tackle such an issue, in this work, we propose a deep Level Set network to produce compact and uniform saliency maps. Our method drives the network to learn a Level Set function for salient objects so it can output more accurate boundaries and compact saliency. Besides, to propagate saliency information among pixels and recover full resolution saliency map, we extend a superpixel-based guided filter to be a layer in the network. The proposed network has a simple structure and is trained end-to-end. During testing, the network can produce saliency maps by efficiently feedforwarding testing images at a speed over 12FPS on GPUs. Evaluations on benchmark datasets show that the proposed method achieves state-of-the-art performance.
Ping Hu 0001, Bing Shuai, Jun Liu 0036, Gang Wang 0012
CVPR2
2016 DAG-Recurrent Neural Networks for Scene Labeling
abstract
In image labeling, local representations for image units are usually generated from their surrounding image patches, thus long-range contextual information is not effectively encoded. In this paper, we introduce recurrent neural networks (RNNs) to address this issue. Specifically, directed acyclic graph RNNs (DAG-RNNs) are proposed to process DAG-structured images, which enables the network to model long-range semantic dependencies among image units. Our DAG-RNNs are capable of tremendously enhancing the discriminative power of local representations, which significantly benefits the local classification. Meanwhile, we propose a novel class weighting function that attends to rare classes, which phenomenally boosts the recognition accuracy for non-frequent classes. Integrating with convolution and deconvolution layers, our DAG-RNNs achieve new state-of-the-art results on the challenging SiftFlow, CamVid and Barcelona benchmarks.
Bing Shuai, Zhen Zuo, Bing Wang 0003, Gang Wang 0012
CVPR1
2016 A Siamese Long Short-Term Memory Architecture for Human Re-identification
Rahul Rama Varior, Bing Shuai, Jiwen Lu, Dong Xu 0001, Gang Wang 0012
ECCV (7)2
2016 Decomposed human localization from social photo album
Shaozi Li, Songzhi Su, Bing Shuai, Rongrong Ji
Multim. Syst.4
2016 Scene Parsing With Integration of Parametric and Non-Parametric Models
abstract
We adopt convolutional neural networks (CNNs) to be our parametric model to learn discriminative features and classifiers for local patch classification. Based on the occurrence frequency distribution of classes, an ensemble of CNNs (CNN-Ensemble) are learned, in which each CNN component focuses on learning different and complementary visual patterns. The local beliefs of pixels are output by CNN-Ensemble. Considering that visually similar pixels are indistinguishable under local context, we leverage the global scene semantics to alleviate the local ambiguity. The global scene constraint is mathematically achieved by adding a global energy term to the labeling energy function, and it is practically estimated in a non-parametric framework. A large margin-based CNN metric learning method is also proposed for better global belief estimation. In the end, the integration of local and global beliefs gives rise to the class likelihood of pixels, based on which maximum marginal inference is performed to generate the label prediction maps. Even without any post-processing, we achieve the state-of-the-art results on the challenging SiftFlow and Barcelona benchmarks.
Bing Shuai, Zhen Zuo, Gang Wang 0012, Bing Wang 0003
IEEE Trans. Image Process.1
2016 Learning Contextual Dependence With Convolutional Hierarchical Recurrent Neural Networks
abstract
Deep convolutional neural networks (CNNs) have shown their great success on image classification. CNNs mainly consist of convolutional and pooling layers, both of which are performed on local image areas without considering the dependence among different image regions. However, such dependence is very important for generating explicit image representation. In contrast, recurrent neural networks (RNNs) are well known for their ability of encoding contextual information in sequential data, and they only require a limited number of network parameters. Thus, we proposed the hierarchical RNNs (HRNNs) to encode the contextual dependence in image representation. In HRNNs, each RNN layer focuses on modeling spatial dependence among image regions from the same scale but different locations. While the cross RNN scale connections target on modeling scale dependencies among regions from the same location but different scales. Specifically, we propose two RNN models: 1) hierarchical simple recurrent network (HSRN), which is fast and has low computational cost and 2) hierarchical long-short term memory recurrent network, which performs better than HSRN with the price of higher computational cost. In this paper, we integrate CNNs with HRNNs, and develop end-to-end convolutional hierarchical RNNs (C-HRNNs) for image classification. C-HRNNs not only utilize the discriminative representation power of CNNs, but also utilize the contextual dependence learning ability of our HRNNs. On four of the most challenging object/scene image classification benchmarks, our C-HRNNs achieve the state-of-the-art results on Places 205, SUN 397, and MIT indoor, and the competitive results on ILSVRC 2012.
Zhen Zuo, Bing Shuai, Gang Wang 0012, Bing Wang 0003, Yushi Chen 0002
IEEE Trans. Image Process.2
2015 Integrating parametric and non-parametric models for scene labeling
abstract
We adopt Convolutional Neural Networks (CNN) as our parametric model to learn discriminative features and classifiers for local patch classification. As visually similar pixels are indistinguishable from local context, we alleviate such ambiguity by introducing a global scene constraint. We estimate the global potential in a non-parametric framework. Furthermore, a large margin based CNN metric learning method is proposed for better global potential estimation. The final pixel class prediction is performed by integrating local and global beliefs. Even without any post-processing, we achieve state-of-the-art performance on SiftFlow and competitive results on Stanford Background benchmark.
Bing Shuai, Gang Wang 0012, Zhen Zuo, Bing Wang 0003, Lifan Zhao
CVPR1
2015 Exemplar based Deep Discriminative and Shareable Feature Learning for scene image classification
Zhen Zuo, Gang Wang 0012, Bing Shuai, Lifan Zhao, Qingxiong Yang
Pattern Recognit.3
2015 Quaddirectional 2D-Recurrent Neural Networks For Image Labeling
abstract
We adopt Convolutional Neural Networks (CNN) to learn discriminative features for local patch classification. We further introduce quaddirectional 2D Recurrent Neural Networks to model the long range dependencies among pixels. Our quaddirectional 2D-RNN is able to embed the global image context into the compact local representation, which significantly enhance their discriminative power. Our experiments demonstrate that the integration of CNN and quaddirectional 2D-RNN achieves very promising results which are comparable to state-of-the-art on real-world image labeling benchmarks.
Bing Shuai, Zhen Zuo, Gang Wang 0012
IEEE Signal Process. Lett.1
2014 Learning Discriminative and Shareable Features for Scene Classification
Zhen Zuo, Gang Wang 0012, Bing Shuai, Lifan Zhao, Qingxiong Yang, Xudong Jiang 0001
ECCV (1)3
2013 Decomposed human localization in personal photo albums
abstract
Recent years have seen tremendous progress in human detection, whereas only upright poses are usually considered. In this paper, we relax this constraint to localizing highly deformable persons, as commonly exhibited in personal photo albums. Human localization based on arbitrary pose is extremely challenging, due to the large pose variances, disabling the traditional part based template detectors. To tackle this issue, we propose a decomposition-based human localization model dealing with this issue in three-step: a stable upper-body is firstly detected, then a set of bigger bounding boxes are extended, from which the most appropriate instance is distinguished by a discriminative Whole Person Model. The experiment results demonstrated that our decomposition-based model worked very well at localizing deformable persons, which boosted the average precision by 10% compared to state-of-the-art person detectors. On the other hand, Similar Pose Feature(SPF) provides the feasibility of projecting persons with similar poses into same clusters, facilitating a novel pose-based photo album browsing functionality.
Bing Shuai, Songzhi Su, Shaozi Li, Rongrong Ji
VCIP1