Bo Liu 0043

dblp:58/2670-43 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0002-4188-9147ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 5 since 2021Computer networks · 1
YearPublicationVenuePosition
2024 UGG: Unified Generative Grasping
Jiaxin Lu 0001, Hao Kang, Bo Liu 0043, Yiding Yang, Qixing Huang, Gang Hua 0001
ECCV (67)4
2023 Flexible Visual Recognition by Evidential Modeling of Confusion and Ignorance
abstract
In real-world scenarios, typical visual recognition systems could fail under two major causes, i.e., the misclassification between known classes and the excusable misbehavior on unknown-class images. To tackle these deficiencies, flexible visual recognition should dynamically predict multiple classes when they are unconfident between choices and reject making predictions when the input is entirely out of the training distribution. Two challenges emerge along with this novel task. First, prediction uncertainty should be separately quantified as confusion depicting inter-class uncertainties and ignorance identifying out-of-distribution samples. Second, both confusion and ignorance should be comparable between samples to enable effective decision-making. In this paper, we propose to model these two sources of uncertainty explicitly with the theory of Subjective Logic. Regarding recognition as an evidence-collecting process, confusion is then defined as conflicting evidence, while ignorance is the absence of evidence. By predicting Dirichlet concentration parameters for singletons, comprehensive subjective opinions, including confusion and ignorance, could be achieved via further evidence combinations. Through a series of experiments on synthetic data analysis, visual recognition, and open-set detection, we demonstrate the effectiveness of our methods in quantifying two sources of uncertainties and dealing with flexible recognition.
Lei Fan 0005, Bo Liu 0043, Ying Wu 0001, Gang Hua 0001
ICCV2
2022 Breadcrumbs: Adversarial Class-Balanced Sampling for Long-Tailed Recognition
Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos
ECCV (24)1
2021 BEV-Net: Assessing Social Distancing Compliance by Joint People Localization and Geometric Reasoning
abstract
Social distancing, an essential public health measure to limit the spread of contagious diseases, has gained significant attention since the outbreak of the COVID-19 pandemic. In this work, the problem of visual social distancing compliance assessment in busy public areas, with wide field-of-view cameras, is considered. A dataset of crowd scenes with people annotations under a bird’s eye view (BEV) and ground truth for metric distances is introduced, and several measures for the evaluation of social distance detection systems are proposed. A multi-branch network, BEV-Net, is proposed to localize individuals in world coordinates and identify high-risk regions where social distancing is violated. BEV-Net combines detection of head and feet locations, camera pose estimation, a differentiable homography module to map image into BEV coordinates, and geometric reasoning to produce a BEV map of the people locations in the scene. Experiments on complex crowded scenes demonstrate the power of the approach and show superior performance over baselines derived from methods in the literature. Applications of interest for public health decision makers are finally discussed. Datasets, code and pretrained models are publicly available at GitHub1.
Zhirui Dai, Yuepeng Jiang, Yi Li 0051, Bo Liu 0043, Antoni B. Chan, Nuno Vasconcelos
ICCV4
2021 GistNet: a Geometric Structure Transfer Network for Long-Tailed Recognition
abstract
The problem of long-tailed recognition, where the number of examples per class is highly unbalanced, is considered. It is hypothesized that the well known tendency of standard classifier training to overfit to popular classes can be exploited for effective transfer learning. Rather than eliminating this overfitting, e.g. by adopting popular class-balanced sampling methods, the learning algorithm should instead leverage this overfitting to transfer geometric information from popular to low-shot classes. A new classifier architecture, GistNet, is proposed to support this goal, using constellations of classifier parameters to encode the class geometry. A new learning algorithm is then proposed for GeometrIc Structure Transfer (GIST), with resort to a combination of loss functions that combine class-balanced and random sampling to guarantee that, while overfitting to the popular classes is restricted to geometric parameters, it is leveraged to transfer class geometry from popular to few-shot classes. This enables better generalization for few-shot classes without the need for the manual specification of class weights, or even the explicit grouping of classes into different types. Experiments on two popular long-tailed recognition datasets show that GistNet outperforms existing solutions to this problem.
Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos
ICCV1
2020 Exploit Clues From Views: Self-Supervised and Regularized Learning for Multiview Object Recognition
abstract
Multiview recognition has been well studied in the literature and achieves decent performance in object recognition and retrieval task. However, most previous works rely on supervised learning and some impractical underlying assumptions, such as the availability of all views in training and inference time. In this work, the problem of multiview self-supervised learning (MV-SSL) is investigated, where only image to object association is given. Given this setup, a novel surrogate task for self-supervised learning is proposed by pursuing "object invariant" representation. This is solved by randomly selecting an image feature of an object as object prototype, accompanied with multiview consistency regularization, which results in view invariant stochastic prototype embedding (VISPE). Experiments shows that the recognition and retrieval results using VISPE outperform that of other self-supervised learning methods on seen and unseen data. VISPE can also be applied to semi-supervised scenario and demonstrates robust performance with limited data available. Code is available at https://github.com/chihhuiho/VISPE.
Chih-Hui Ho, Bo Liu 0043, Tz-Ying Wu, Nuno Vasconcelos
CVPR2
2020 Few-Shot Open-Set Recognition Using Meta-Learning
abstract
The problem of open-set recognition is considered. While previous approaches only consider this problem in the context of large-scale classifier training, we seek a unified solution for this and the low-shot classification setting. It is argued that the classic softmax classifier is a poor solution for open-set recognition, since it tends to overfit on the training classes. Randomization is then proposed as a solution to this problem. This suggests the use of meta-learning techniques, commonly used for few-shot classification, for the solution of open-set recognition. A new oPen sEt mEta LEaRning (PEELER) algorithm is then introduced. This combines the random selection of a set of novel classes per episode, a loss that maximizes the posterior entropy for examples of those classes, and a new metric learning formulation based on the Mahalanobis distance. Experimental results show that PEELER achieves state of the art open set recognition performance for both few-shot and large-scale recognition. On CIFAR and miniImageNet, it achieves substantial gains in seen/unseen class detection AUROC for a given seen-class classification accuracy.
Bo Liu 0043, Hao Kang, Gang Hua 0001, Nuno Vasconcelos
CVPR1
2020 SPOT: Selective Point Cloud Voting for Better Proposal in Point Cloud Object Detection
Hongyuan Du, Linjun Li, Bo Liu 0043, Nuno Vasconcelos
ECCV (11)3
2018 Feature Space Transfer for Data Augmentation
abstract
The problem of data augmentation in feature space is considered. A new architecture, denoted the FeATure TransfEr Network (FATTEN), is proposed for the modeling of feature trajectories induced by variations of object pose. This architecture exploits a parametrization of the pose manifold in terms of pose and appearance. This leads to a deep encoder/decoder network architecture, where the encoder factors into an appearance and a pose predictor. Unlike previous attempts at trajectory transfer, FATTEN can be efficiently trained end-to-end, with no need to train separate feature transfer functions. This is realized by supplying the decoder with information about a target pose and the use of a multi-task loss that penalizes category- and pose-mismatches. In result, FATTEN discourages discontinuous or non-smooth trajectories that fail to capture the structure of the pose manifold, and generalizes well on object recognition tasks involving large pose variation. Experimental results on the artificial ModelNet database show that it can successfully learn to map source features to target features of a desired pose, while preserving class identity. Most notably, by using feature space transfer for data augmentation (w.r.t. pose and depth) on SUN-RGBD objects, we demonstrate considerable performance improvements on one/few-shot object recognition in a transfer learning setup, compared to current state-of-the-art methods.
Bo Liu 0043, Xudong Wang 0007, Mandar Dixit, Roland Kwitt, Nuno Vasconcelos
CVPR1
2015 Bayesian Model Adaptation for Crowd Counts
abstract
The problem of transfer learning is considered in the domain of crowd counting. A solution based on Bayesian model adaptation of Gaussian processes is proposed. This is shown to produce intuitive model updates, which are tractable, and lead to an adapted model (predictive distribution) that accounts for all information in both training and adaptation data. The new adaptation procedure achieves significant gains over previous approaches, based on multi-task learning, while requiring much less computation to deploy. This makes it particularly suited for the problem of expanding the capacity of crowd counting camera networks. A large video dataset for the evaluation of adaptation approaches to crowd counting is also introduced. This contains a number of adaptation tasks, involving information transfer across video collected by 1) a single camera under different scene conditions (different times of the day) and 2) video collected from different cameras. Evaluation of the proposed model adaptation procedure in this dataset shows good performance in realistic operating conditions.
Bo Liu 0043, Nuno Vasconcelos
ICCV1
2012 iCEnergy: augmented reality display for intuitive energy monitoring
abstract
Energy saving is the main goal in most building energy monitoring applications. These systems, however, are operated by people. For this reason, an intuitive user interface is an essential element that will affect users' data understandability, thereby determining the system's usability. Traditional energy monitoring generally focuses on getting energy information by utilizing graphs or static text interfaces. However, these approaches are not related to the physical space. With the increase of information in energy monitoring systems, intuitive and efficient ways of displaying information are needed. In this demo, we present iCEnergy, a vision-based mobile information interface that provides power monitoring using augmented reality. Using existing system data, the system overlays an interactive "energy cloud" over corresponding devices in order to illustrate information about the physical environment. This approach aims to provide users with a comfortable interaction experience through its intuitive information display.
Shijia Pan, Bo Liu 0043, Lin Zhang 0001, Pei Zhang 0001
SenSys2