Xiaobai Liu

dblp:86/3967 · DBLP profile ↗
← Back
63ranked-venue papers
30as first author
8since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 25 first-author · 6 since 2021Artificial intelligence and machine learning · 34 · 13 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2024 Policy-driven Auto-Augmentation with Distillment Rewards for Scene Text Recognition
Yibiao Zhao, Xiaobai Liu
MMAsia3
2024 Focal Diffusion Process for Object-Aware 3D LiDAR Generation
abstract
LiDAR data generation has seen significant advancements and success in recent years, particularly in applications such as autonomous driving and robotics.However, existing generative models often struggle with poor foreground object quality due to two main reasons: (1) foreground objects typically occupy a small portion of the 3D point cloud in a frame; and (2) objects of the same class often exhibit large intra-class variances.To address these challenges, we propose a variant of the Denoising Diffusion Probabilistic Model (DDPM) and introduce an object focal loss to increase emphasis on foreground objects during training, particularly those with less frequent appearances.The proposed focal diffusion process is able to enhance the distinctiveness of foreground objects in the generated lidar frames.Experiments on the public KITTI-360 dataset showed that the proposed focal diffusion process outperformed existing generation approaches, establishing a new state-of-the-art in 3D LiDAR generation.Ablation studies further validated the superiority of the proposed object focal loss.
Xiaobai Liu
MMAsia2
2024 An Iterative Semi-supervised Approach with Pixel-wise Contrastive Loss for Road Extraction in Aerial Images
abstract
Extracting roads in aerial images has numerous applications in artificial intelligence and multimedia computing, including traffic pattern analysis and parking space planning. Learning deep neural networks, though very successful, demand vast amounts of high-quality annotations, of which acquisition is time-consuming and expensive. In this work, we propose a semi-supervised approach for image-based road extraction in which only a small set of labeled images are available for training to address this challenge. We design a pixel-wise contrastive loss to self-supervise the network training to utilize the large corpus of unlabeled images. The key idea is to identify pairs of overlapping image regions (positive) or non-overlapping image regions (negative) and encourage the network to make similar outputs for positive pairs or dissimilar outputs for negative pairs. We also develop a negative sampling strategy to filter false-negative samples during the process. An iterative procedure is introduced to apply the network over raw images to generate pseudo-labels, filter and select high-quality labels with the proposed contrastive loss, and retrain the network with the enlarged training dataset. We repeat these iterative steps until convergence. We validate the effectiveness of the proposed methods by performing extensive experiments on the public SpaceNet3 and DeepGlobe Road datasets. Results show that our proposed method achieves state-of-the-art results on public image segmentation benchmarks and significantly outperforms other semi-supervised methods.
Xiaobai Liu, Xianfeng Terry Yang
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Siamese-Based Twin Attention Network for Visual Tracking
abstract
Recently, object tracking have achieved remarkable progress in terms of both efficiency and accuracy. However, exiting methods still cannot satisfy challenging tasks under complicated scenarios, such as occlusion, scale variations, and etc. To this end, we propose a novel Siamese-based Twin Attention Network for visual tracking. First, a multi-branch fusion module is presented. By leveaging the fusion scheme, we merge the low-level features with the high-level features extracted from different convolution layers. Then, the representation ability of the target can be enhanced effectively. Second, to fully capture the contextual information in the tracking process, we introduced a global context module into the search branch. Third, to attain robust performance, a saliency mine scheme is employed in the proposed network. Specifically, the self-attention operation is utilized to capture the contextual information from the spatial and channel domain, while the cross-attention operation is to enrich the contextual information relevance by fusing the features between the template and search region. By utilizing these schemes, our tracker can cope well with different challenging scenes. Extensive experiments were conducted on several popular benchmarks, including VOT2016, VOT2018, VOT2019, VOT2021, OTB2013, OTB2015, GOT10k, LaSOT, and NFS. The results demonstrate that the proposed method is effective and achieves competitive results.
Hua Bao, Ping Shu, Xiaobai Liu
IEEE Trans. Circuits Syst. Video Technol.4
2023 Learning Stage-Wise GANs for Whistle Extraction in Time-Frequency Spectrograms
abstract
Whistle contour extraction aims to derive animal whistles from time-frequency spectrograms as polylines. For toothed whales, whistle extraction results can serve as the basis for analyzing animal abundance, species identity, and social activities. During the last few decades, as long-term recording systems have become affordable, automated whistle extraction algorithms were proposed to process large volumes of recording data. Recently, a deep learning-based method demonstrated superior performance in extracting whistles under varying noise conditions. However, training such networks requires a large amount of labor-intensive annotation, which is not available for many species. To overcome this limitation, we present a framework of stage-wise generative adversarial networks (GANs), which compile new whistle data suitable for deep model training via three stages: generation of background noise in the spectrogram, generation of whistle contours, and generation of whistle signals. By separating the generation of different components in the samples, our framework composes visually promising whistle data and labels even when few expert annotated data are available. Regardless of the amount of human-annotated data, the proposed data augmentation framework leads to a consistent improvement in performance of the whistle extraction model, with a maximum increase of 1.69 in the whistle extraction mean F1-score. Our stage-wise GAN also surpasses one single GAN in improving whistle extraction models with augmented data.
Marie A. Roch, Holger Klinck, Erica Fleishman, Douglas Gillespie, Eva-Marie Nosal, Yu Shiu, Xiaobai Liu
IEEE Trans. Multim.8
2022 A Hamiltonian Monte Carlo Method for Probabilistic Adversarial Attack and Learning
abstract
Although deep convolutional neural networks (CNNs) have demonstrated remarkable performance on multiple computer vision tasks, researches on adversarial learning have shown that deep models are vulnerable to adversarial examples, which are crafted by adding visually imperceptible perturbations to the input images. Most of the existing adversarial attack methods only create a single adversarial example for the input, which just gives a glimpse of the underlying data manifold of adversarial examples. An attractive solution is to explore the solution space of the adversarial examples and generate a diverse bunch of them, which could potentially improve the robustness of real-world systems and help prevent severe security threats and vulnerabilities. In this paper, we present an effective method, called Hamiltonian Monte Carlo with Accumulated Momentum (HMCAM), aiming to generate a sequence of adversarial examples. To improve the efficiency of HMC, we propose a new regime to automatically control the length of trajectories, which allows the algorithm to move with adaptive step sizes along the search direction at different positions. Moreover, we revisit the reason for high computational cost of adversarial training under the view of MCMC and design a new generative method called Contrastive Adversarial Training (CAT), which approaches equilibrium distribution of adversarial examples with only few iterations by building from small modifications of the standard Contrastive Divergence (CD) and achieve a trade-off between efficiency and accuracy. Both quantitative and qualitative analysis on several natural image datasets and practical systems have confirmed the superiority of the proposed algorithm.
Hongjun Wang 0005, Guanbin Li, Xiaobai Liu, Liang Lin 0004
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Monocular 3D Pose Estimation via Pose Grammar and Data Augmentation
abstract
In this paper, we propose a pose grammar to tackle the problem of 3D human pose estimation from a monocular RGB image. Our model takes estimated 2D pose as the input and learns a generalized 2D-3D mapping function to leverage into 3D pose. The proposed model consists of a base network which efficiently captures pose-aligned features and a hierarchy of Bi-directional RNNs (BRNNs) on the top to explicitly incorporate a set of knowledge regarding human body configuration (i.e., kinematics, symmetry, motor coordination). The proposed model thus enforces high-level constraints over human poses. In learning, we develop a data augmentation algorithm to further improve model robustness against appearance variations and cross-view generalization ability. We validate our method on public 3D human pose benchmarks and propose a new evaluation protocol working on cross-view setting to verify the generalization capability of different methods. We empirically observe that most state-of-the-art methods encounter difficulty under such setting while our method can well handle such challenges.
Yuanlu Xu, Wenguan Wang, Tengyu Liu, Xiaobai Liu, Jianwen Xie, Song-Chun Zhu
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Learning Sample-Specific Policies for Sequential Image Augmentation
abstract
This paper presents a policy-driven sequential image augmentation approach for image-related tasks. Our approach applies a sequence of image transformations (e.g., translation, rotation) over a training image, one transformation at a time, with the augmented image from the previous time step treated as the input for the next transformation. This sequential data augmentation substantially improves sample diversity, leading to improved test performance, especially for data-hungry models (e.g., deep neural networks). However, the search for the optimal transformation of each image at each time step of the sequence has high complexity due to its combination nature. To address this challenge, we formulate the search task as a sequential decision process and introduce a deep policy network that learns to produce transformations based on image content. We also develop an iterative algorithm to jointly train a classifier and the policy network in the reinforcement learning setting. The immediate reward of a potential transformation is defined to encourage transformations producing hard samples for the current classifier. At each iteration, we employ the policy network to augment the training dataset, train a classifier with the augmented data, and train the policy net with the aid of the classifier. We apply the above approach to both public image classification benchmarks and a newly collected image dataset for material recognition. Comparisons to alternative augmentation approaches show that our policy-driven approach achieves comparable or improved classification performance while using significantly fewer augmented images. The code is available at https://github.com/Paul-LiPu/rl_autoaug.
Xiaobai Liu, Xiaohui Xie
ACM Multimedia2
2020 Learning Knowledge-Rich Sequential Model for Planar Homography Estimation in Aerial Video
abstract
This paper presents an unsupervised approach that leverages raw aerial videos to learn to estimate planar homographic transformation between consecutive video frames. Previous learning-based estimators work on pairs of images to estimate their planar homographic transformations but suffer from severe over-fitting issues, especially when applying over aerial videos. To address this concern, we develop a sequential estimator that directly processes a sequence of video frames and estimates their pairwise planar homographic transformations in batches. We also incorporate a set of spatial-temporal knowledge to regularize the learning of such a sequence-to-sequence model. We collect a set of challenging aerial videos and compare the proposed method to the alternative algorithms. Empirical studies suggest that our sequential model achieves significant improvement over alternative image-based methods and the knowledge-rich regularization further boosts our system performance. Our codes and dataset could be found at https://github.com/Paul-LiPu/DeepVideoHomography.
Xiaobai Liu
ICPR2
2020 Learning Deep Models from Synthetic Data for Extracting Dolphin Whistle Contours
abstract
We present a learning-based method for extracting whistles of toothed whales (Odontoceti) in hydrophone recordings. Our method represents audio signals as time-frequency spectrograms and decomposes each spectrogram into a set of time-frequency patches. A deep neural network learns archetypical patterns (e.g., crossings, frequency modulated sweeps) from the spectrogram patches and predicts time-frequency peaks that are associated with whistles. We also developed a comprehensive method to synthesize training samples from background environments and train the network with minimal human annotation effort. We applied the proposed learn-from-synthesis method to a subset of the public Detection, Classification, Localization, and Density Estimation (DCLDE) 2011 workshop data to extract whistle confidence maps, which we then processed with an existing contour extractor to produce whistle annotations. The F1-score of our best synthesis method was 0.158 greater than our baseline whistle extraction algorithm (~25% improvement) when applied to common dolphin (Delphinus spp.) and bottlenose dolphin (Tursiops truncatus) whistles.
Xiaobai Liu, K. J. Palmer, Erica Fleishman, Douglas Gillespie, Eva-Marie Nosal, Yu Shiu, Holger Klinck, Danielle Cholewiak, Tyler A. Helble, Marie A. Roch
IJCNN2
2020 Discrete Optimal Graph Clustering
abstract
Graph-based clustering is one of the major clustering methods. Most of it works in three separate steps: 1) similarity graph construction; 2) clustering label relaxing; and 3) label discretization with k -means (KM). Such common practice has three disadvantages: 1) the predefined similarity graph is often fixed and may not be optimal for the subsequent clustering; 2) the relaxing process of cluster labels may cause significant information loss; and 3) label discretization may deviate from the real clustering result since KM is sensitive to the initialization of cluster centroids. To tackle these problems, in this paper, we propose an effective discrete optimal graph clustering framework. A structured similarity graph that is theoretically optimal for clustering performance is adaptively learned with a guidance of reasonable rank constraints. Besides, to avoid the information loss, we explicitly enforce a discrete transformation on the intermediate continuous label, which derives a tractable optimization problem with a discrete solution. Furthermore, to compensate for the unreliability of the learned labels and enhance the clustering accuracy, we design an adaptive robust module that learns the prediction function for the unseen data based on the learned discrete cluster labels. Finally, an iterative optimization strategy guaranteed with convergence is developed to directly solve the clustering results. Extensive experiments conducted on both real and synthetic datasets demonstrate the superiority of our proposed methods compared with several state-of-the-art clustering approaches.
Lei Zhu 0002, Zhiyong Cheng 0001, Jingjing Li 0001, Xiaobai Liu
IEEE Trans. Cybern.5
2020 Learning Semisupervised Multilabel Fully Convolutional Network for Hierarchical Object Parsing
abstract
This article presents a semisupervised multilabel fully convolutional network (FCN) for hierarchical object parsing of images. We consider each object part (e.g., eye and head) as a class label and learn to assign every image pixel to multiple coherent part labels. Different from previous methods that consider part labels as independent classes, our method explicitly models the internal relationships between object parts, e.g., that a pixel highly scored for eyes should be highly scored for heads as well. Such relationships directly reflect the structure of the semantic space and thus should be respected while learning the deep representation. We achieve this objective by introducing a multilabel softmax loss function over both labeled and unlabeled images and regularizing it with two pairwise ranking constraints. The first constraint is based on a manifold assumption that image pixels being visually and spatially close to each other should be collaboratively classified as the same part label. The other constraint is used to enforce that no pixel receives significant scores from more than one label that are semantically conflicting with each other. The proposed loss function is differentiable with respect to network parameters and hence can be optimized by standard stochastic gradient methods. We evaluate the proposed method on two public image data sets for hierarchical object parsing and compare it with the alternative parsing methods. Extensive comparisons showed that our method can achieve state-of-the-art performance while using 50% less labeled training samples than the alternatives.
Xiaobai Liu, Grayson Adkins, Eric Medwedeff, Liang Lin 0004, Shuicheng Yan
IEEE Trans. Neural Networks Learn. Syst.1
2019 Revisiting Jump-Diffusion Process for Visual Tracking: A Reinforcement Learning Approach
abstract
In this paper, we revisit the classical stochastic jump-diffusion process and develop an effective variant for estimating visibility statuses of objects while tracking them in videos. Dealing with partial or full occlusions is a long standing problem in computer vision but largely remains unsolved. In this paper, we cast the above problem as a Markov decision process and develop a policy-based jump-diffusion method to jointly track object locations in videos and estimate their visibility statuses. Our method employs a set of jump dynamics to change visibility statuses of objects and a set of diffusion dynamics to track objects in videos. Different from the traditional jump-diffusion process that stochastically generates dynamics, we utilize deep policy functions to determine the best dynamic for the present state and learn the optimal policies using reinforcement learning methods. Our method is capable of tracking objects with full or partial occlusions in crowded scenes. We evaluate the proposed method over challenging video sequences and compare it to alternative tracking methods. Significant improvements are made particularly for videos with frequent interactions or occlusions.
Xiaobai Liu, Thuan Chau, Yadong Mu, Lei Zhu 0002, Shuicheng Yan
IEEE Trans. Circuits Syst. Video Technol.1
2019 Neural Task Planning With AND-OR Graph Representations
abstract
This paper focuses on semantic task planning, that is, predicting a sequence of actions toward accomplishing a specific task under a certain scene, which is a new problem in computer vision research. The primary challenges are how to model the task-specific knowledge and how to integrate this knowledge into the learning procedure. In this paper, we propose training a recurrent long short-term memory (LSTM) network to address this problem, that is, taking a scene image (including prelocated objects) and the specified task as input and recurrently predicting action sequences. However, training such a network generally requires large numbers of annotated samples to cover the semantic space (e.g., diverse action decomposition and ordering). To overcome this issue, we introduce a knowledge and-or graph (AOG) for task description, which hierarchically represents a task as atomic actions. With this AOG representation, we can produce many valid samples (i.e., action sequences according to common sense) by training another auxiliary LSTM network with a small set of annotated samples. Furthermore, these generated samples (i.e., task-oriented action sequences) effectively facilitate training of the model for semantic task planning. In our experiments, we create a new dataset that contains diverse daily tasks and extensively evaluates the effectiveness of our approach.
Tianshui Chen, Riquan Chen, Lin Nie, Xiaobai Liu, Liang Lin 0004
IEEE Trans. Multim.5
2018 Learning Pose Grammar to Encode Human Body Configuration for 3D Pose Estimation
abstract
In this paper, we propose a pose grammar to tackle the problem of 3D human pose estimation. Our model directly takes 2D pose as input and learns a generalized 2D-3D mapping function. The proposed model consists of a base network which efficiently captures pose-aligned features and a hierarchy of Bi-directional RNNs (BRNN) on the top to explicitly incorporate a set of knowledge regarding human body configuration (i.e., kinematics, symmetry, motor coordination). The proposed model thus enforces high-level constraints over human poses. In learning, we develop a pose sample simulator to augment training samples in virtual camera views, which further improves our model generalizability. We validate our method on public 3D human pose benchmarks and propose a new evaluation protocol working on cross-view setting to verify the generalization capability of different methods. We empirically observe that most state-of-the-art methods encounter difficulty under such setting while our method can well handle such challenges.
Haoshu Fang, Yuanlu Xu, Wenguan Wang, Xiaobai Liu, Song-Chun Zhu
AAAI4
2018 A Causal And-Or Graph Model for Visibility Fluent Reasoning in Tracking Interacting Objects
abstract
Tracking humans that are interacting with the other subjects or environment remains unsolved in visual tracking, because the visibility of the human of interests in videos is unknown and might vary over time. In particular, it is still difficult for state-of-the-art human trackers to recover complete human trajectories in crowded scenes with frequent human interactions. In this work, we consider the visibility status of a subject as a fluent variable, whose change is mostly attributed to the subject's interaction with the surrounding, e.g., crossing behind another object, entering a building, or getting into a vehicle, etc. We introduce a Causal And-Or Graph (C-AOG) to represent the causal-effect relations between an object's visibility fluent and its activities, and develop a probabilistic graph model to jointly reason the visibility fluent change (e.g., from visible to invisible) and track humans in videos. We formulate this joint task as an iterative search of a feasible causal graph structure that enables fast search algorithm, e.g., dynamic programming method. We apply the proposed method on challenging video sequences to evaluate its capabilities of estimating visibility fluent changes of subjects and tracking subjects of interests over time. Results with comparisons demonstrate that our method outperforms the alternative trackers and can recover complete trajectories of humans in complicated scenarios with frequent human interactions.
Yuanlu Xu, Xiaobai Liu, Jianwen Xie, Song-Chun Zhu
CVPR3
2018 Unsupervised Learning based Jump-Diffusion Process for Object Tracking in Video Surveillance
abstract
This paper presents a principled way for dealing with occlusions in visual tracking which is a long-standing issue in computer vision but largely remains unsolved. As the major innovation, we develop a learning-based jump-diffusion process to jointly track object locations and estimate their visibility statuses over time. Our method employs in particular a set of jump dynamics to change object's visibility statuses and a set of diffusion dynamics to track objects in videos. Different from the traditional jump-diffusion process that stochastically generates dynamics, we utilize deep policy functions to determine the best dynamic at the present step and learn the optimal policies from raw videos using reinforcement learning methods.Our method is capable of tracking objects with severe occlusions in crowded scenes and thus recovers the complete trajectories of objects that undergo multiple interactions with others. We evaluate the proposed method on challenging video sequences and compare it to alternative methods. Significant improvements are obtained particularly for the videos including frequent interactions or occlusions.
Xiaobai Liu, Donovan Lo, Chau Thuan
IJCAI1
2018 FIPIP: A novel fine-grained parallel partition based intra-frame prediction on heterogeneous many-core systems
Wenbin Jiang 0001, Laurence T. Yang, Xiaobai Liu, Hai Jin 0001, Alan L. Yuille, Ye Chi
Future Gener. Comput. Syst.4
2018 Single-View 3D Scene Reconstruction and Parsing by Attribute Grammar
abstract
In this paper, we present an attribute grammar for solving two coupled tasks: i) parsing a 2D image into semantic regions; and ii) recovering the 3D scene structures of all regions. The proposed grammar consists of a set of production rules, each describing a kind of spatial relation between planar surfaces in 3D scenes. These production rules are used to decompose an input image into a hierarchical parse graph representation where each graph node indicates a planar surface or a composite surface. Different from other stochastic image grammars, the proposed grammar augments each graph node with a set of attribute variables to depict scene-level global geometry, e.g., camera focal length, or local geometry, e.g., surface normal, contact lines between surfaces. These geometric attributes impose constraints between a node and its off-springs in the parse graph. Under a probabilistic framework, we develop a Markov Chain Monte Carlo method to construct a parse graph that optimizes the 2D image recognition and 3D scene reconstruction purposes simultaneously. We evaluated our method on both public benchmarks and newly collected datasets. Experiments demonstrate that the proposed method is capable of achieving state-of-the-art scene reconstruction of a single image.
Xiaobai Liu, Yibiao Zhao, Song-Chun Zhu
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Iterative image segmentation with feature driven heuristic four-color labeling
Kunqian Li, Wenbing Tao, Xiaobai Liu, Liman Liu
Pattern Recognit.3
2018 A Stochastic Attribute Grammar for Robust Cross-View Human Tracking
abstract
In computer vision, tracking humans across camera views remain challenging, especially for complex scenarios with frequent occlusions, significant lighting changes, and other difficulties. Under such conditions, most existing appearance and geometric cues are not reliable enough to distinguish humans across camera views. To address these challenges, this paper presents a stochastic attribute grammar model for leveraging complementary and discriminative human attributes for enhancing cross-view tracking. The key idea of our method is to introduce a hierarchical representation, parse graph, to describe a subject and its movement trajectory in both space and time domains. These results in a hierarchical compositional representation, comprising trajectory entities of varying level, including human boxes, 3D human boxes, tracklets, and trajectories. We use a set of grammar rules to decompose a graph node (e.g., tracklet) into a set of children nodes (e.g., 3D human boxes), and augment each node with a set of attributes, including geometry (e.g., moving speed and direction), accessories (e.g., bags), and/or activities (e.g., walking and running). These attributes serve as valuable cues, in addition to appearance features (e.g., colors), in determining the associations of human detection boxes across cameras. In particular, the attributes of a parent node are inherited by its children nodes, resulting in consistency constraints over the feasible parse graph. Thus, we cast cross-view human tracking as finding the most discriminative parse graph for each subject in videos. We develop a learning method to train this attribute grammar model from weakly supervised training data. To infer the optimal parse graph and its attributes, we develop an alternative parsing method that employs both top-down and bottom-up computations to search the optimal solution. We also explicitly reason the occlusion status of each entity in order to deal with significant changes of camera viewpoints. We evaluate the proposed method over public video benchmarks, and demonstrate with extensive experiments that our method clearly outperforms the state-of-the-art tracking methods.
Xiaobai Liu, Yuanlu Xu, Lei Zhu 0002, Yadong Mu
IEEE Trans. Circuits Syst. Video Technol.1
2018 High-Precision Camera Localization in Scenes with Repetitive Patterns
abstract
This article presents a high-precision multi-modal approach for localizing moving cameras with monocular videos, which has wide potentials in many intelligent applications, including robotics, autonomous vehicles, and so on. Existing visual odometry methods often suffer from symmetric or repetitive scene patterns, e.g., windows on buildings or parking stalls. To address this issue, we introduce a robust camera localization method that contributes in two aspects. First, we formulate feature tracking, the critical step of visual odometry, as a hierarchical min-cost network flow optimization task, and we regularize the formula with flow constraints, cross-scale consistencies, and motion heuristics. The proposed regularized formula is capable of adaptively selecting distinctive features or feature combinations, which is more effective than traditional methods that detect and group repetitive patterns in a separate step. Second, we develop a joint formula for integrating dense visual odometry and sparse GPS readings in a common reference coordinate. The fusion process is guided with high-order statistics knowledge to suppress the impacts of noises, clusters, and model drifting. We evaluate the proposed camera localization method on both public video datasets and a newly created dataset that includes scenes full of repetitive patterns. Results with comparisons show that our method can achieve comparable performance to state-of-the-art methods and is particularly effective for addressing repetitive pattern issues.
Xiaobai Liu, Yadong Mu, Jiadi Yang, Liang Lin 0004, Shuicheng Yan
ACM Trans. Intell. Syst. Technol.1
2018 Learning Multi-Instance Deep Ranking and Regression Network for Visual House Appraisal
abstract
This paper presents a weakly supervised regression model for the visual house appraisal problem, which aims to predict the value of a house from its photos and textual descriptions (e.g., number of bedrooms). The key idea of our approach is a multi-layer neural network, called multi-instance Deep Ranking and Regression (MiDRR) net, which jointly solves two coupled tasks: ranking and regression, in the multiple instance setting. The network is trained using weakly supervised data, which do not require intensive human annotations. We also design a set of human heuristics to promote deep features through imposing constraints over the solution space, e.g., a house with three bedrooms often has a higher value than that with only two bedrooms. While these constraints are specific to the studied problem, the developed formula can be easily generalized to the other regression applications. For test and evaluation purposes, we collect a comprehensive house image benchmark that includes 900,000 photos from 30,000 houses recently traded in the USA, and apply the proposed MiDRR net to predict house values. Extensive evaluations with comparisons demonstrate that additional usage of imagery data as well as human heuristics can significantly boost system performance and that the proposed MiDRR net clearly outperforms the alternative methods.
Xiaobai Liu, Jingjie Yang, Jacob Thalman, Shuicheng Yan, Jiebo Luo 0001
IEEE Trans. Knowl. Data Eng.1
2017 Cross-View People Tracking by Scene-Centered Spatio-Temporal Parsing
abstract
In this paper, we propose a Spatio-temporal Attributed Parse Graph (ST-APG) to integrate semantic attributes with trajectories for cross-view people tracking. Given videos from multiple cameras with overlapping field of view (FOV), our goal is to parse the videos and organize the trajectories of all targets into a scene-centered representation. We leverage rich semantic attributes of human, e.g., facing directions, postures and actions, to enhance cross-view tracklet associations, besides frequently used appearance and geometry features in the literature.In particular, the facing direction of a human in 3D, once detected, often coincides with his/her moving direction or trajectory. Similarly, the actions of humans, once recognized, provide strong cues for distinguishing one subject from the others. The inference is solved by iteratively grouping tracklets with cluster sampling and estimating people semantic attributes by dynamic programming.In experiments, we validate our method on one public dataset and create another new dataset that records people's daily life in public, e.g., food court, office reception and plaza, each of which includes 3-4 cameras. We evaluate the proposed method on these challenging videos and achieve promising multi-view tracking results.
Yuanlu Xu, Xiaobai Liu, Song-Chun Zhu
AAAI2
2017 Single-Image 3D Scene Parsing Using Geometric Commonsense
abstract
This paper presents a unified grammatical framework capable of reconstructing a variety of scene types (e.g., urban, campus, county etc.) from a single input image. The key idea of our approach is to study a novel commonsense reasoning framework that mainly exploits two types of prior knowledges: (i) prior distributions over a single dimension of objects, e.g., that the length of a sedan is about 4.5 meters; (ii) pair-wise relationships between the dimensions of scene entities, e.g., that the length of a sedan is shorter than a bus. These unary or relative geometric knowledge, once extracted, are fairly stable across different types of natural scenes, and are informative for enhancing the understanding of various scenes in both 2D images and 3D world. Methodologically, we propose to construct a hierarchical graph representation as a unified representation of the input image and related geometric knowledge. We formulate these objectives with a unified probabilistic formula and develop a data-driven Monte Carlo method to infer the optimal solution with both bottom-to-up and top-down computations. Results with comparisons on public datasets showed that our method clearly outperforms the alternative methods.
Chengcheng Yu, Xiaobai Liu, Song-Chun Zhu
IJCAI2
2017 Place-centric Visual Urban Perception with Deep Multi-instance Regression
abstract
This paper presents a unified framework to learn to quantify perceptual attributes (e.g., safety, attractiveness) of physical urban environments using crowd-sourced street-view photos without human annotations. The efforts of this work include two folds. First, we collect a large-scale urban image dataset in multiple major cities in U.S.A., which consists of multiple street-view photos for every place. Instead of using subjective annotations as in previous works, which are neither accurate nor consistent, we collect for every place the safety score from government's crime event records as objective safety indicators. Second, we observe that the place-centric perception task is by nature a multi-instance regression problem since the labels are only available for places (bags), rather than images or image regions (instances). We thus introduce a deep convolutional neural network (CNN) to parameterize the instance-level scoring function, and develop an EM algorithm to alternatively estimate the primary instances (images or image regions) which affect the safety scores and train the proposed network. Our method is capable of localizing interesting images and image regions for each place. We evaluate the proposed method on a newly created dataset and a public dataset. Results with comparisons showed that our method can clearly outperform the alternative perception methods and more importantly, is capable of generating region-level safety scores to facilitate interpretations of the perception process.
Xiaobai Liu, Lei Zhu 0002, Yuanlu Xu, Liang Lin 0004
ACM Multimedia1
2017 VSCC'2017: Visual Analysis for Smart and Connected Communities
abstract
This paper presents a brief summary of the first workshop on Visual Analysis for Smart and Connected Communities (VSCC'2017), which is held in conjunction with the ACM Conference on Multimedia, 2017. VSCC'2017 is the first workshop on visual analysis for smart and connected communities, and is aimed at the creation of a multi-discipline community with the concentration on visual data analysis. The topics addressed in VSCC'2017 cover all aspects of smart communities, including safety, security, retrieval, transportation, information technologies, Internet of Things, etc. The focus is on the development of fundamental theories, algorithms, models towards the parsing of visual data generated by various digital devices (e.g., smart phones, vehicles) in smart communities.
Xiaobai Liu, Yadong Mu, Yu-Gang Jiang 0001, Jiebo Luo 0001
ACM Multimedia1
2017 Stochastic Gradient Made Stable: A Manifold Propagation Approach for Large-Scale Optimization
abstract
Stochastic gradient descent (SGD) holds as a classical method to build large scale machine learning models over big data. A stochastic gradient is typically calculated from a limited number of samples (known as mini-batch), which potentially incurs a high variance and causes the estimated parameters to bounce around the optimal solution. To improve the stability of stochastic gradient, recent years have witnessed the proposal of several semi-stochastic gradient descent algorithms, which distinguish themselves from standard SGD by incorporating global information into gradient computation. In this paper, we contribute a novel stratified semi-stochastic gradient descent (S3GD) algorithm to this nascent research area, accelerating the optimization of a large family of composite convex functions. Though theoretically converging faster, prior semi-stochastic algorithms are found to suffer from high iteration complexity, which makes them even slower than SGD in practice on many datasets. In our proposed S3GD, the semi-stochastic gradient is calculated based on efficient manifold propagation, which can be numerically accomplished by sparse matrix multiplications. This way S3GD is able to generate a highly-accurate estimate of the exact gradient from each mini-batch with largely-reduced computational complexity. Theoretic analysis reveals that the proposed S3GD elegantly balances the geometric algorithmic convergence rate against the space and time complexities during the optimization. The efficacy of S3GD is also experimentally corroborated on several large-scale benchmark datasets.
Yadong Mu, Wei Liu 0005, Xiaobai Liu, Wei Fan 0001
IEEE Trans. Knowl. Data Eng.3
2017 Discrete Multimodal Hashing With Canonical Views for Robust Mobile Landmark Search
abstract
Mobile landmark search (MLS) recently receives increasing attention for its great practical values. However, it still remains unsolved due to two important challenges. One is high bandwidth consumption of query transmission, and the other is the huge visual variations of query images sent from mobile devices. In this paper, we propose a novel hashing scheme, named as canonical view based discrete multimodal hashing (CV-DMH), to handle these problems. First, a submodular function is designed to measure visual representativeness and redundancy of a view set. With it, canonical views, which capture key visual appearances of landmark with limited redundancy, are efficiently discovered with an iterative mining strategy. Second, multimodal sparse coding is applied to transform visual features from multiple modalities into an intermediate representation. It can robustly and adaptively characterize visual contents of varied landmark images with certain canonical views. Finally, compact binary codes are learned on intermediate representation within a tailored discrete binary embedding model which preserves visual relations of images measured with canonical views and removes the involved noises. In this part, we develop a new augmented Lagrangian multiplier (ALM) based optimization method to directly solve the discrete binary codes. We can not only explicitly deal with the discrete constraint, but also consider the bit-uncorrelated constraint and balance constraint together. The proposed solution can desirably avoid accumulated quantization errors in conventional optimization method which simply adopts two-step "relaxing+rounding'' framework. Experiments on real world landmark datasets demonstrate the superior performance of CV-DMH over several state-of-the-art methods.
Lei Zhu 0002, Zi Huang, Xiaobai Liu, Xiangnan He 0001, Jiande Sun 0001, Xiaofang Zhou 0001
IEEE Trans. Multim.3
2016 Multi-View 3D Human Tracking in Crowded Scenes
abstract
This paper presents a robust multi-view method for tracking people in 3D scene. Our method distinguishes itself from previous works in two aspects. Firstly, we define a set of binary spatial relationships for individual subjects or pairs of subjects that appear at the same time, e.g. being left or right, being closer or further to the camera, etc. These binary relationships directly reflect relative positions of subjects in 3D scene and thus should be persisted during inference. Secondly, we introduce an unified probabilistic framework to exploit binary spatial constraints for simultaneous 3D localization and cross-view human tracking. We develop a cluster Markov Chain Monte Carlo method to search the optimal solution. We evaluate our method on both public video benchmarks and newly built multi-view video dataset. Results with comparisons showed that our method could achieve state-of-the-art tracking results and meter-level 3D localization on challenging videos.
Xiaobai Liu
AAAI1
2016 Multi-view People Tracking via Hierarchical Trajectory Composition
abstract
This paper presents a hierarchical composition approach for multi-view object tracking. The key idea is to adaptively exploit multiple cues in both 2D and 3D, e.g., ground occupancy consistency, appearance similarity, motion coherence etc., which are mutually complementary while tracking the humans of interests over time. While feature online selection has been extensively studied in the past literature, it remains unclear how to effectively schedule these cues for the tracking purpose especially when encountering various challenges, e.g. occlusions, conjunctions, and appearance variations. To do so, we propose a hierarchical composition model and re-formulate multi-view multi-object tracking as a problem of compositional structure optimization. We setup a set of composition criteria, each of which corresponds to one particular cue. The hierarchical composition process is pursued by exploiting different criteria, which impose constraints between a graph node and its offsprings in the hierarchy. We learn the composition criteria using MLE on annotated data and efficiently construct the hierarchical graph by an iterative greedy pursuit algorithm. In the experiments, we demonstrate superior performance of our approach on three public datasets, one of which is newly created by us to test various challenges in multi-view multi-object tracking.
Yuanlu Xu, Xiaobai Liu, Yang Liu 0266, Song-Chun Zhu
CVPR2
2016 A Stochastic Image Grammar for Fine-Grained 3D Scene Reconstruction
Xiaobai Liu, Yadong Mu, Liang Lin 0004
IJCAI1
2016 Geometric Scene Parsing with Hierarchical LSTM
Zhanglin Peng, Ruimao Zhang, Xiaodan Liang, Xiaobai Liu, Liang Lin 0004
IJCAI4
2016 Learning Compact Visual Representation with Canonical Views for Robust Mobile Landmark Search
Lei Zhu 0002, Jialie Shen 0001, Xiaobai Liu, Liang Xie 0001, Liqiang Nie
IJCAI3
2016 V3I-STAL: Visual Vehicle-to-Vehicle Interaction via Simultaneous Tracking and Localization
abstract
This paper investigates a visual interaction system for vehicle-to-vehicle (V2V) platform, called V3I. Our system employs common visual cameras that are mounted on connected vehicles to perceive the existence of isolated vehicles in the same roadway, and provides human drivers with imagery situational awareness. This allows effective interactions between vehicles even with a low permeation rate of V2V devices. The underlying research problem for V3I includes two aspects: i) tracking isolated vehicles of interest over time through local cameras; ii) at each time-step fusing the results of local visual perceptions to obtain a global location map that involves both isolated and connected vehicles. In this paper, we introduce a unified probabilistic approach to solve the above two problems, i.e., tracking and localization, in a joint fashion. Our approach will explore both the visual features of individual vehicles in images and the pair-wise spatial relationships between vehicles. We develop a fast Markov Chain Monte Carlo (MCMC) algorithm to search the joint solution space efficiently, which enables real-time application. To evaluate the performance of the proposed approach, we collect and annotate a set of video sequences captured with a group of vehicle-resident cameras. Extensive experiments with comparisons clearly demonstrate that the proposed V3I approach can precisely recover the dynamic location map of the surrounding and thus enable direct visual interactions between vehicles .
Xiaobai Liu
ACM Multimedia1
2016 Learning Collaborative Sparse Representation for Grayscale-Thermal Tracking
abstract
Integrating multiple different yet complementary feature representations has been proved to be an effective way for boosting tracking performance. This paper investigates how to perform robust object tracking in challenging scenarios by adaptively incorporating information from grayscale and thermal videos, and proposes a novel collaborative algorithm for online tracking. In particular, an adaptive fusion scheme is proposed based on collaborative sparse representation in Bayesian filtering framework. We jointly optimize sparse codes and the reliable weights of different modalities in an online way. In addition, this paper contributes a comprehensive video benchmark, which includes 50 grayscale-thermal sequences and their ground truth annotations for tracking purpose. The videos are with high diversity and the annotations were finished by one single person to guarantee consistency. Extensive experiments against other state-of-the-art trackers with both grayscale and grayscale-thermal inputs demonstrate the effectiveness of the proposed tracking approach. Through analyzing quantitative results, we also provide basic insights and potential future research directions in grayscale-thermal tracking.
Chenglong Li 0002, Shiyi Hu, Xiaobai Liu, Jin Tang 0001, Liang Lin 0004
IEEE Trans. Image Process.4
2016 Learning Compositional Shape Models of Multiple Distance Metrics by Information Projection
abstract
This paper presents a novel compositional contour-based shape model by incorporating multiple distance metrics to account for varying shape distortions or deformations. Our approach contains two key steps: 1) contour feature generation and 2) generative model pursuit. For each category, we first densely sample an ensemble of local prototype contour segments from a few positive shape examples and describe each segment using three different types of distance metrics. These metrics are diverse and complementary with each other to capture various shape deformations. We regard the parameterized contour segment plus an additive residual ϵ as a basic subspace, namely, ϵ -ball, in the sense that it represents local shape variance under the certain distance metric. Using these ϵ -balls as features, we then propose a generative learning algorithm to pursue the compositional shape model, which greedily selects the most representative features under the information projection principle. In experiments, we evaluate our model on several public challenging data sets, and demonstrate that the integration of multiple shape distance metrics is capable of dealing various shape deformations, articulations, and background clutter, hence boosting system performance.
Ping Luo 0002, Liang Lin 0004, Xiaobai Liu
IEEE Trans. Neural Networks Learn. Syst.3
2014 Detect What You Can: Detecting and Representing Objects Using Holistic Models and Body Parts
abstract
Detecting objects becomes difficult when we need to deal with large shape deformation, occlusion and low resolution. We propose a novel approach to i) handle large deformations and partial occlusions in animals (as examples of highly deformable objects), ii) describe them in terms of body parts, and iii) detect them when their body parts are hard to detect (e.g., animals depicted at low resolution). We represent the holistic object and body parts separately and use a fully connected model to arrange templates for the holistic object and body parts. Our model automatically decouples the holistic object or body parts from the model when they are hard to detect. This enables us to represent a large number of holistic object and body part combinations to better deal with different "detectability" patterns caused by deformations, occlusion and/or low resolution. We apply our method to the six animal categories in the PASCAL VOC dataset and show that our method significantly improves state-of-the-art (by 4.1% AP) and provides a richer representation for objects. During training we use annotations for body parts (e.g., head, torso, etc.), making use of a new dataset of fully annotated object parts for PASCAL VOC 2010, which provides a mask for each part.
Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, Alan L. Yuille
CVPR3
2014 Single-View 3D Scene Parsing by Attributed Grammar
abstract
In this paper, we present an attributed grammar for parsing man-made outdoor scenes into semantic surfaces, and recovering its 3D model simultaneously. The grammar takes superpixels as its terminal nodes and use five production rules to generate the scene into a hierarchical parse graph. Each graph node actually correlates with a surface or a composite of surfaces in the 3D world or the 2D image. They are described by attributes for the global scene model, e.g. focal length, vanishing points, or the surface properties, e.g. surface normal, contact line with other surfaces, and relative spatial location etc. Each production rule is associated with some equations that constraint the attributes of the parent nodes and those of their children nodes. Given an input image, our goal is to construct a hierarchical parse graph by recursively applying the five grammar rules while preserving the attributes constraints. We develop an effective top-down/bottom-up cluster sampling procedure which can explore this constrained space efficiently. We evaluate our method on both public benchmarks and newly built datasets, and achieve state-of-the-art performances in terms of layout estimation and region segmentation. We also demonstrate that our method is able to recover detailed 3D model with relaxed Manhattan structures which clearly advances the state-of-the-arts of single-view 3D reconstruction.
Xiaobai Liu, Yibiao Zhao, Song-Chun Zhu
CVPR1
2014 The Role of Context for Object Detection and Semantic Segmentation in the Wild
abstract
In this paper we study the role of context in existing state-of-the-art detection and segmentation approaches. Towards this goal, we label every pixel of PASCAL VOC 2010 detection challenge with a semantic category. We believe this data will provide plenty of challenges to the community, as it contains 520 additional classes for semantic segmentation and object detection. Our analysis shows that nearest neighbor based approaches perform poorly on semantic segmentation of contextual classes, showing the variability of PASCAL imagery. Furthermore, improvements of existing contextual models for detection is rather modest. In order to push forward the performance in this difficult scenario, we propose a novel deformable part-based model, which exploits both local context around each candidate detection as well as global context at the level of the scene. We show that this contextual reasoning significantly helps in detecting objects at all scales.
Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, Alan L. Yuille
CVPR3
2014 MsLRR: A Unified Multiscale Low-Rank Representation for Image Segmentation
abstract
In this paper, we present an efficient multiscale low-rank representation for image segmentation. Our method begins with partitioning the input images into a set of superpixels, followed by seeking the optimal superpixel-pair affinity matrix, both of which are performed at multiple scales of the input images. Since low-level superpixel features are usually corrupted by image noise, we propose to infer the low-rank refined affinity matrix. The inference is guided by two observations on natural images. First, looking into a single image, local small-size image patterns tend to recur frequently within the same semantic region, but may not appear in semantically different regions. The internal image statistics are referred to as replication prior, and we quantitatively justified it on real image databases. Second, the affinity matrices at different scales should be consistently solved, which leads to the cross-scale consistency constraint. We formulate these two purposes with one unified formulation and develop an efficient optimization procedure. The proposed representation can be used for both unsupervised or supervised image segmentation tasks. Our experiments on public data sets demonstrate the presented method can substantially improve segmentation accuracy.
Xiaobai Liu, Jiayi Ma 0001, Hai Jin 0001, Yanduo Zhang
IEEE Trans. Image Process.1
2014 Nonnegative Tensor Cofactorization and Its Unified Solution
abstract
In this paper, we present a new joint factorization algorithm, called Nonnegative Tensor Co-Factorization (NTCoF). The key idea is to simultaneously factorize multiple visual features of the same data into nonnegative dimensionality-reduced representations, and meanwhile, to maximize the correlations of the low-dimensional representations. The data is generally encoded as tensors of arbitrary order, rather than vectors, to preserve the original data structures. NTCoF provides a simple and efficient way to fuse multiple complementary features for enhancing the discriminative power of the desired rank-reduced representations under the nonnegative constraints. We formulate the related objectives with a block-wise quadratic nonnegative function. To optimize, a unified convergence provable solution is developed. This solution is applicable for any nonnegative optimization problems with block-wise quadratic objective functions, and thus offer an unified platform based on which specific solution can be directly derived by skipping over tedious proof about algorithmic convergence. We apply the proposed algorithm and solution on three image tasks, face recognition, multi-class image categorization and multi-label image annotation. Results with comparisons on public challenging datasets show that the proposed algorithm can outperform both the traditional nonnegative methods and the popular feature combination methods.
Xiaobai Liu, Shuicheng Yan, Gang Wang 0012, Hai Jin 0001, Seong-Whan Lee
IEEE Trans. Image Process.1
2013 Robust Region Grouping via Internal Patch Statistics
abstract
In this work, we present an efficient multi-scale low-rank representation for image segmentation. Our method begins with partitioning the input images into a set of super pixels, followed by seeking the optimal super pixel-pair affinity matrix, both of which are performed at multiple scales of the input images. Since low-level super pixel features are usually corrupted by image noises, we propose to infer the low-rank refined affinity matrix. The inference is guided by two observations on natural images. First, looking into a single image, local small-size image patterns tend to recur frequently within the same semantic region, but may not appear in semantically different regions. We call this internal image statistics as replication prior, and quantitatively justify it on real image databases. Second, the affinity matrices at different scales should be consistently solved, which leads to the cross-scale consistency constraint. We formulate these two purposes with one unified formulation and develop an efficient optimization procedure. Our experiments demonstrate the presented method can substantially improve segmentation accuracy.
Xiaobai Liu, Liang Lin 0004, Alan L. Yuille
CVPR1
2013 Human Re-identification by Matching Compositional Template with Cluster Sampling
abstract
This paper aims at a newly raising task in visual surveillance: re-identifying people at a distance by matching body information, given several reference examples. Most of existing works solve this task by matching a reference template with the target individual, but often suffer from large human appearance variability (e.g. different poses/views, illumination) and high false positives in matching caused by conjunctions, occlusions or surrounding clutters. Addressing these problems, we construct a simple yet expressive template from a few reference images of a certain individual, which represents the body as an articulated assembly of compositional and alternative parts, and propose an effective matching algorithm with cluster sampling. This algorithm is designed within a candidacy graph whose vertices are matching candidates (i.e. a pair of source and target body parts), and iterates in two steps for convergence. (i) It generates possible partial matches based on compatible and competitive relations among body parts. (ii) It confirms the partial matches to generate a new matching solution, which is accepted by the Markov Chain Monte Carlo (MCMC) mechanism. In the experiments, we demonstrate the superior performance of our approach on three public databases compared to existing methods.
Yuanlu Xu, Liang Lin 0004, Wei-Shi Zheng 0001, Xiaobai Liu
ICCV4
2012 Object categorization with sketch representation and generalized samples
Liang Lin 0004, Xiaobai Liu, Shaowu Peng, Hongyang Chao, Yongtian Wang, Bo Jiang 0002
Pattern Recognit.2
2012 Visual Classification With Multitask Joint Sparse Representation
abstract
We address the problem of visual classification with multiple features and/or multiple instances. Motivated by the recent success of multitask joint covariate selection, we formulate this problem as a multitask joint sparse representation model to combine the strength of multiple features and/or instances for recognition. A joint sparsity-inducing norm is utilized to enforce class-level joint sparsity patterns among the multiple representation vectors. The proposed model can be efficiently optimized by a proximal gradient method. Furthermore, we extend our method to the setup where features are described in kernel matrices. We then investigate into two applications of our method to visual classification: 1) fusing multiple kernel features for object categorization and 2) robust face recognition in video with an ensemble of query images. Extensive experiments on challenging real-world data sets demonstrate that the proposed method is competitive to the state-of-the-art methods in respective applications.
Xiao-Tong Yuan, Xiaobai Liu, Shuicheng Yan
IEEE Trans. Image Process.2
2012 Image label completion by pursuing contextual decomposability
abstract
This article investigates how to automatically complete the missing labels for the partially annotated images, without image segmentation. The label completion procedure is formulated as a nonnegative data factorization problem, to decompose the global image representations that are used for describing the entire images, for instance, various image feature descriptors, into their corresponding label representations, that are used for describing the local semantic regions within images. The solution provided in this work is motivated by following observations. First, label representations of the regions with the same label often share certain commonness, yet may be essentially different due to the large intraclass variations. Thus, each label or concept should be represented by using a subspace spanned by an ensemble of basis, instead of a single one, to characterize the intralabel diversities. Second, the subspaces for different labels are different from each other. Third, while two images are similar with each other, the corresponding label representations should be similar. We formulate this cross-image context as well as the given partial label annotations in the framework of nonnegative data factorization and then propose an efficient multiplicative nonnegative update rules to alternately optimize the subspaces and the reconstruction coefficients. We also provide the theoretic proof of algorithmic convergence and correctness. Extensive experiments over several challenging image datasets clearly demonstrate the effectiveness of our proposed solution in boosting the quality of image label completion and image annotation accuracy. Based on the same formulation, we further develop a label ranking algorithms, to refine the noised image labels without any manual supervision. We compare the proposed label ranking algorithm with the state-of-the-arts over the popular evaluation databases and achieve encouragingly improvements.
Xiaobai Liu, Shuicheng Yan, Tat-Seng Chua, Hai Jin 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2012 Label-to-region with continuity-biased bi-layer sparsity priors
abstract
In this work, we investigate how to reassign the fully annotated labels at image level to those contextually derived semantic regions, namely Label-to-Region (L2R), in a collective manner. Given a set of input images with label annotations, the basic idea of our approach to L2R is to first discover the patch correspondence across images, and then propagate the common labels shared in image pairs to these correlated patches. Specially, our approach consists of following aspects. First, each of the input images is encoded as a Bag-of-Hierarchical-Patch (BOP) for capturing the rich cues at variant scales, and the individual patches are expressed by patch-level feature descriptors. Second, we present a sparse representation formulation for discovering how well an image or a semantic region can be robustly reconstructed by all the other image patches from the input image set. The underlying philosophy of our formulation is that an image region can be sparsely reconstructed with the image patches belonging to the other images with common labels, while the robustness in label propagation across images requires that these selected patches come from very few images. This preference of being sparse at both patch and image level is namedbi-layer sparsity prior. Meanwhile, we enforce the preference of choosing larger-size patches in reconstruction, referred to ascontinuity-biased priorin this work, which may further enhance the reliability of L2R assignment. Finally, we harness the reconstruction coefficients to propagate the image labels to the matched patches, and fuse the propagation results over all patches to finalize the L2R task. As a by-product, the proposed continuity-biased bi-layer sparse representation formulation can be naturally applied to perform image annotation on new testing images. Extensive experiments on three public image datasets clearly demonstrate the effectiveness of our proposed framework in both L2R assignment and image annotation.
Xiaobai Liu, Shuicheng Yan, Bin Cheng 0001, Jinhui Tang 0001, Tat-Seng Chua, Hai Jin 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2011 Segment an image by looking into an image corpus
abstract
This paper investigates how to segment an image into semantic regions by harnessing an unlabeled image corpus. First, the image segmentation task is recast as a small-size patch grouping problem. Then, we discover two novel patch-pair priors, namely the first-order patch-pair density prior and the second-order patch-pair co-occurrence prior, founded on two statistical observations from the natural image corpus. The underlying rationalities are: 1) a patch-pair falling within the same object region generally has higher density than a patch-pair falling on different objects, and 2) two patch-pairs with high co-occurrence frequency are likely to bear similar semantic consistence confidences (SCCs), i.e. the confidence of the consisted two patches belonging to the same semantic concept. These two discriminative priors are further integrated into a unified objective function in order to augment the intrinsic patch-pair similarities, originally calculated using patch-level visual features, into the semantic consistence confidences. Nonnegative constraint is also imposed over the output variables and an efficient iterative procedure is provided to seek the optimal solution. The ultimate patch grouping is conducted by first building a similarity graph, which takes the atomic patches as vertices and the augmented patch-pair SCCs as edge weights, and then employing the popular Normalized Cut approach to group patches into semantic clusters. Extensive image segmentation experiments on two public databases clearly demonstrate the superiority of the proposed approach over various state-of-the-arts unsupervised image segmentation algorithms.
Xiaobai Liu, Jiashi Feng, Shuicheng Yan, Liang Lin 0004, Hai Jin 0001
CVPR1
2011 Multi-class semi-supervised SVMs with Positiveness Exclusive Regularization
abstract
In this work, we address the problem of multi-class classification problem in semi-supervised setting. A regularized multi-task learning approach is presented to train multiple binary-class Semi-Supervised Support Vector Machines (S3VMs) using the one-vs-rest strategy within a joint framework. A novel type of regularization, namely Positiveness Exclusive Regularization (PER), is introduced to induce the following prior: if an unlabeled sample receives significant positive response from one of the classifiers, it is less likely for this sample to receive positive responses from the other classifiers. That is, we expect an exclusive relationship among different S3VMs for evaluating the same unlabeled sample. We propose to use an ℓ1,2-norm regularizer as an implementation of PER. The objective of our approach is to minimize an empirical risk regularized by a PER term and a manifold regularization term. An efficient Nesterov-type smoothing approximation based method is developed for optimization. Evaluations with comparisons are conducted on several benchmarks for visual classification to demonstrate the advantages of the proposed method.
Xiaobai Liu, Xiao-Tong Yuan, Shuicheng Yan, Hai Jin 0001
ICCV1
2011 Adaptive Object Tracking by Learning Hybrid Template Online
abstract
This paper presents an adaptive tracking algorithm by learning hybrid object templates online in video. The templates consist of multiple types of features, each of which describes one specific appearance structure, such as flatness, texture, or edge/corner. Our proposed solution consists of three aspects. First, in order to make the features of different types comparable with each other, a unified statistical measure is defined to select the most informative features to construct the hybrid template. Second, we propose a simple yet powerful generative model for representing objects. This model is characterized by its simplicity since it could be efficiently learnt from the currently observed frames. Last, we present an iterative procedure to learn the object template from the currently observed frames, and to locate every feature of the object template within the observed frames. The former step is referred to as feature pursuit, and the latter step is referred to as feature alignment, both of which are performed over a batch of observations. We fuse the results of feature alignment to locate objects within frames. The proposed solution to object tracking is in essence robust against various challenges, including background clutters, low-resolution, scale changes, and severe occlusions. Extensive experiments are conducted over several publicly available databases and the results with comparisons show that our tracking algorithm clearly outperforms the state-of-the-art methods.
Xiaobai Liu, Liang Lin 0004, Shuicheng Yan, Hai Jin 0001, Wenbin Jiang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2011 Integrating Spatio-Temporal Context With Multiview Representation for Object Recognition in Visual Surveillance
abstract
We present in this paper an integrated solution to rapidly recognizing dynamic objects in surveillance videos by exploring various contextual information. This solution consists of three components. The first one is a multi-view object representation. It contains a set of deformable object templates, each of which comprises an ensemble of active features for an object category in a specific view/pose. The template can be efficiently learned via a small set of roughly aligned positive samples without negative samples. The second component is a unified spatio-temporal context model, which integrates two types of contextual information in a Bayesian way. One is the spatial context, including main surface property (constraints on object type and density) and camera geometric parameters (constraints on object size at a specific location). The other is the temporal context, containing the pixel-level and instance-level consistency models, used to generate the foreground probability map and local object trajectory prediction. We also combine the above spatial and temporal contextual information to estimate the object pose in scene and use it as a strong prior for inference. The third component is a robust sampling-based inference procedure. Taking the spatio-temporal contextual knowledge as the prior model and deformable template matching as the likelihood model, we formulate the problem of object category recognition as a maximum-a-posteriori problem. The probabilistic inference can be achieved by a simple Markov chain Mento Carlo sampler, owing to the informative spatio-temporal context model which is able to greatly reduce the computation complexity and the category ambiguities. The system performance and benefit gain from the spatio-temporal contextual information are quantitatively evaluated on several challenging datasets and the comparison results clearly demonstrate that our proposed algorithm outperforms other state-of-the-art algorithms.
Xiaobai Liu, Liang Lin 0004, Shuicheng Yan, Hai Jin 0001, Wenbing Tao
IEEE Trans. Circuits Syst. Video Technol.1
2010 Nonparametric Label-to-Region by search
abstract
In this work, we investigate how to propagate annotated labels for a given single image from the image-level to their corresponding semantic regions, namely Label-to-Region (L2R), by utilizing the auxiliary knowledge from Internet image search with the annotated image labels as queries. A nonparametric solution is proposed to perform L2R for single image with complete labels. First, each label of the image is used as query for online image search engines to obtain a set of semantically related and visually similar images, which along with the input image are encoded as Bags-of-Hierarchical-Patches. Then, an efficient two-stage feature mining procedure is presented to discover those input-image specific, salient and descriptive features for each label from the proposed Interpolation SIFT (iSIFT) feature pool. These features consequently constitute a patch-level representation, and the continuity-biased sparse coding is proposed to select few patches from the online images with preference to larger patches to reconstruct a candidate region, which randomly merges the spatially connected patches of the input image. Such candidate regions are further ranked according to the reconstruction errors, and the top regions are used to derive the label confidence vector for each patch of the input image. Finally, a patch clustering procedure is performed as postprocessing to finalize L2R for the input image. Extensive experiments on three public databases demonstrate the encouraging performance of the proposed nonparametric L2R solution.
Xiaobai Liu, Shuicheng Yan, Jiebo Luo 0001, Jinhui Tang 0001, ZhongYang Huang, Hai Jin 0001
CVPR1
2010 Multi-View Object Detection by Classifier Interpolation
Xiaobai Liu, Haifeng Gong, Shuicheng Yan, Hai Jin 0001
ICASSP1
2010 Image segmentation with patch-pair density priors
abstract
In this paper, we investigate how an unlabeled image corpus can facilitate the segmentation of any given image. A simple yet efficient multi-task joint sparse representation model is presented to augment the patch-pair similarities by harnessing the newly discovered patch-pair density priors. First, each image in over-segmented as a set of patches, and the adjacent patch-pair density priors, statistically calculated from the unlabeled image corpus, bring an intuitively explainable and informative observation that kindred patch-pairs generally have higher densities that inhomogeneous patch-pairs. Then for each adjacent patch-pair within the given image, high-density biased multi-task joint sparse reconstruction is pursued such that 1) both individual patches and patch-pair can be reconstructed with few patch-pairs from the unlabeled image corpus, and 2) the patch-pairs selected for reconstruction are high-density biased, namely, preferring patch-pairs belonging to the same semantic region. In this way, the overall reconstruction residue well conveys the discriminative information on whether these two patches belong to the same semantic region, and consequently the patch affinity matrix is augmented by reconstruction residues for all adjacent patch-pairs within the given image. The ultimate image segmentation is derived by employing the popular normalized cut approach over the augmented patch affinity matrix. Extensive image segmentation experiments over two public databases clearly demonstrate the superiority of the proposed solution over several state-of-the-art algorithms. Furthermore, the algorithmic practicality is well validated with comparison experiments on content-based image retrieval and multi-label image annotation performed over image segmentation outputs.
Xiaobai Liu, Jiashi Feng, Shuicheng Yan, Hai Jin 0001
ACM Multimedia1
2010 Layered Graph Matching with Composite Cluster Sampling
abstract
This paper presents a framework of layered graph matching for integrating graph partition and matching. The objective is to find an unknown number of corresponding graph structures in two images. We extract discriminative local primitives from both images and construct a candidacy graph whose vertices are matching candidates (i.e., a pair of primitives) and whose edges are either negative for mutual exclusion or positive for mutual consistence. Then we pose layered graph matching as a multicoloring problem on the candidacy graph and solve it using a composite cluster sampling algorithm. This algorithm assigns some vertices into a number of colors, each being a matched layer, and turns off all the remaining candidates. The algorithm iterates two steps: 1) Sampling the positive and negative edges probabilistically to form a composite cluster, which consists of a few mutually conflicting connected components (CCPs) in different colors and 2) assigning new colors to these CCPs with consistence and exclusion relations maintained, and the assignments are accepted by the Markov Chain Monte Carlo (MCMC) mechanism to preserve detailed balance. This framework demonstrates state-of-the-art performance on several applications, such as multi-object matching with large motion, shape matching and retrieval, and object localization in cluttered background.
Liang Lin 0004, Xiaobai Liu, Song-Chun Zhu
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Projective Nonnegative Graph Embedding
abstract
We present in this paper a general formulation for nonnegative data factorization, called projective nonnegative graph embedding (PNGE), which 1) explicitly decomposes the data into two nonnegative components favoring the characteristics encoded by the so-called intrinsic and penalty graphs , respectively, and 2) explicitly describes how to transform each new testing sample into its low-dimensional nonnegative representation. In the past, such a nonnegative decomposition was often obtained for the training samples only, e.g., nonnegative matrix factorization (NMF) and its variants, nonnegative graph embedding (NGE) and its refined version multiplicative nonnegative graph embedding (MNGE). Those conventional approaches for out-of-sample extension either suffer from the high computational cost or violate the basic nonnegative assumption. In this work, PNGE offers a unified solution to out-of-sample extension problem, and the nonnegative coefficient vector of each datum is assumed to be projected from its original feature representation with a universal nonnegative transformation matrix. A convergency provable multiplicative nonnegative updating rule is then derived to learn the basis matrix and transformation matrix. Extensive experiments compared with the state-of-the-art algorithms on nonnegative data factorization demonstrate the algorithmic properties in convergency, sparsity, and classification power.
Xiaobai Liu, Shuicheng Yan, Hai Jin 0001
IEEE Trans. Image Process.1
2009 Layered graph matching by composite cluster sampling with collaborative and competitive interactions
abstract
This paper studies a framework for matching an unknown number of corresponding structures in two images (shapes), motivated by detecting objects in cluttered background and learning parts from articulated motion. Due to the large distortion between shapes and ambiguity caused by symmetric or cluttered structures, many inference algorithms often get stuck in local minimums and converge slowly. We propose a composite cluster sampling algorithm with a “candidacy graph” representation, where each vertex (candidate) is a possible match for a pair of source and target primitives (local structure or small curves), and the layered matching is then formulated as a multiple coloring problem. Each two vertices can be linked by either a competitive edge or a collaborative edge. These edges indicate the connected vertices should/shouldn't be assigned the same color. With this representation, the stochastic sampling contains two steps: (i) Sampling the competitive and collaborative edges to form a composite cluster, in which a few mutual-conflicting connected components are in different colors; (ii) Sampling the new colors to this cluster remaining consistency with Markov Chain Monte Carlo (MCMC) mechanism. The algorithm is applied to many applications on many public datasets and outperform the state of the art approaches.
Liang Lin 0004, Xiaobai Liu, Song-Chun Zhu
CVPR3
2009 Trajectory parsing by cluster sampling in spatio-temporal graph
abstract
The objective of this paper is to parse object trajectories in surveillance video against occlusion, interruption, and background clutter. We present a spatio-temporal graph (ST-Graph) representation and a cluster sampling algorithm via deferred inference. An object trajectory in the ST-Graph is represented by a bundle of “motion primitives”, each of which consists of a small number of matched features (interesting patches) generated by adaptive feature pursuit and a tracking process. Each motion primitive is a graph vertex and has six bonds connecting to neighboring vertices. Based on the ST-Graph, we jointly solve three tasks: 1)spatial segmentation; 2)temporal correspondence and 3)object recognition, by flipping the labels of the motion primitives. We also adapt the scene geometric and statistical information as strong prior. Then the inference computation is formulated in a Markov Chain and solved by an efficient cluster sampling. We apply the proposed approach to various challenging videos from a number of public datasets and show it outperform other state of the art methods.
Xiaobai Liu, Liang Lin 0004, Song-Chun Zhu, Hai Jin 0001
CVPR1
2009 Unified Solution to Nonnegative Data Factorization Problems
abstract
In this paper, we restudy the non-convex data factorization problems (regularized or not, unsupervised or supervised), where the optimization is confined in the nonnegative orthant, and provide a unified convergency provable solution based on multiplicative nonnegative update rules. This solution is general for optimization problems with block-wisely quadratic objective functions, and thus direct update rules can be derived by skipping over the tedious specific procedure deduction process and algorithmic convergence proof. By taking this unified solution as a general template, we i) re-explain several existing nonnegative data factorization algorithms, ii) develop a variant of nonnegative matrix factorization formulation for handling out-of-sample data, and Hi) propose a new nonnegative data factorization algorithm, called correlated co-decomposition (CCD), to simultaneously factorize two feature spaces by exploring the inter-correlated information. Experiments on both face recognition and multi-label image annotation tasks demonstrate the wide applicability of the unified solution as well as the effectiveness of two proposed new algorithms.
Xiaobai Liu, Shuicheng Yan, Jun Yan 0001, Hai Jin 0001
ICDM1
2009 Label to region by bi-layer sparsity priors
abstract
In this work, we investigate how to automatically reassign the manually annotated labels at the image-level to those contextually derived semantic regions. First, we propose a bi-layer sparse coding formulation for uncovering how an image or semantic region can be robustly reconstructed from the over-segmented image patches of an image set. We then harness it for the automatic label to region assignment of the entire image set. The solution to bi-layer sparse coding is achieved by convex l1-norm minimization. The underlying philosophy of bi-layer sparse coding is that an image or semantic region can be sparsely reconstructed via the atomic image patches belonging to the images with common labels, while the robustness in label propagation requires that these selected atomic patches come from very few images. Each layer of sparse coding produces the image label assignment to those selected atomic patches and merged candidate regions based on the shared image labels. The results from all bi-layer sparse codings over all candidate regions are then fused to obtain the entire label to region assignments. Besides, the presenting bi-layer sparse coding framework can be naturally applied to perform image annotation on new test images. Extensive experiments on three public image datasets clearly demonstrate the effectiveness of our proposed framework in both label to region assignment and image annotation tasks.
Xiaobai Liu, Bin Cheng 0001, Shuicheng Yan, Jinhui Tang 0001, Tat-Seng Chua, Hai Jin 0001
ACM Multimedia1
2008 Object-of-interest extraction by integrating stochastic inference with learnt active shape sketch
abstract
This article presents a novel integrated approach to object of interest extraction, including learning to define target pattern and extracting by combining detection and segmentation. The learning stage captures both shape sketch and appearance information of target pattern as prior knowledge. The extraction stage utilizes a stochastic Markov Chain Monte Carlo (MCMC) algorithm under the Bayesian framework. By employing a proposed measurement for the similarity between continuous region boundary and discrete learnt sketch, the shape prior knowledge is embedded into the inference process, playing an important role in segmentation. The experiment shows that our method can perform well for both small and large size objects, even in the occluded case, and outperform the comparable methods.
Xiaobai Liu, Lanfang Dong
ICPR4
2008 Layered shape matching and registration: Stochastic sampling with hierarchical graph representation
abstract
To automatically register foreground target in cluttered images, we present a novel hierarchical graph representation and a stochastic computing strategy in Bayesian framework. The graph representation, which contains point-(image primitives), seedgraph-, and subgraph- three levels, are built up following the primal sketch theory to capture geometric, topological, and spatial information both in local and global scale. We use two types of bottom-up algorithms for searching matching candidates to generate the point-level and seedgraph-level representations respectively. Then, the Swendsen-Wang Cuts and Gibbs sampling methods are performed for global optimal solution to generate the final subgraph-level representation, where a mixture bending function and a set of topological operators are defined for matching measurement. Experiments with comparison are demonstrated on standard dataset with outperforming results. Results show that our method can work well even with clutter noise and complex background.
Xiaobai Liu, Liang Lin 0004, Hai Jin 0001, Wenbing Tao
ICPR1