Gabriel L. Oliveira

dblp:117/2073 · also Gabriel Leivas Oliveira · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0003-0099-9873ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 3 since 2021Systems, architecture and hardware · 7 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Efficient and distributed learning · 18% Probabilistic and Bayesian machine learning · 13% Deep learning architectures and training · 11%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
active learning
0.912025
Uncertainty Herding: One Active Learning Method for All Label Budgets · ICLR 2025
Machine learning › Efficient and distributed learning › active learning
uncertainty sampling
0.912025
Uncertainty Herding: One Active Learning Method for All Label Budgets · ICLR 2025
Machine learning › Learning theory
generalization
0.812024
Forget Sharpness: Perturbed Forgetting of Model Biases Within SAM Dynamics · ICML 2024
Machine learning › Trustworthy machine learning › fairness
model bias
0.812024
Forget Sharpness: Perturbed Forgetting of Model Biases Within SAM Dynamics · ICML 2024
Machine learning › Optimization for machine learning › gradient-based optimization
sharpness-aware minimization
0.812024
Forget Sharpness: Perturbed Forgetting of Model Biases Within SAM Dynamics · ICML 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
Forget Sharpness: Perturbed Forgetting of Model Biases Within SAM Dynamics · ICML 2024
Machine learning › Transfer learning and domain adaptation
meta-learning
0.712023
Meta Temporal Point Processes · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.712023
Meta Temporal Point Processes · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
temporal point process
0.712023
Meta Temporal Point Processes · ICLR 2023
Computer vision › Image recognition and object detection
object recognition
0.322014
Sparse Spatial Coding: A Novel Approach to Visual Recognition · IEEE Trans. Image Process. 2014
Sparse Spatial Coding: A novel approach for efficient and accurate object recognition · ICRA 2012
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network architecture
0.312018
DPDB-Net: Exploiting Dense Connections for Convolutional Encoders · ICRA 2018
Computer vision › Video understanding and tracking
action detection
0.312017
Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection · ICCV 2017
Computer vision › Video understanding and tracking
action recognition
0.312017
Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection · ICCV 2017
Robotics › Robot navigation and mapping › localization
long-term localization
0.312017
Semantics-aware visual localization under challenging perceptual conditions · ICRA 2017
Computer vision › Video understanding and tracking › action detection
spatio-temporal action localization
0.312017
Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection · ICCV 2017
Robotics › Robot navigation and mapping › place recognition
visual place recognition
0.312017
Semantics-aware visual localization under challenging perceptual conditions · ICRA 2017
Computer vision › Segmentation and scene understanding › part segmentation
body part segmentation
0.212016
Deep learning for human part discovery in images · ICRA 2016
Computer vision › Image recognition and object detection
scene recognition
0.212014
Sparse Spatial Coding: A Novel Approach to Visual Recognition · IEEE Trans. Image Process. 2014
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.112012
Sparse Spatial Coding: A novel approach for efficient and accurate object recognition · ICRA 2012
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112018
DPDB-Net: Exploiting Dense Connections for Convolutional Encoders · ICRA 2018
Computer vision › Face, body and person analysis
human pose estimation
0.112017
Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection · ICCV 2017
Computer vision › Vision and language
scene description
0.112017
Semantics-aware visual localization under challenging perceptual conditions · ICRA 2017

Methods — techniques the papers use, named apart from their topics

uncertainty coverage · 0.9greedy optimization · 0.9perturbation · 0.8information bottleneck · 0.8meta-learning · 0.7sparse coding · 0.3residual connections · 0.3dense connections · 0.3multi-stream network · 0.3markov chain · 0.3
YearPublicationVenuePosition
2025 Uncertainty Herding: One Active Learning Method for All Label Budgets
abstract
Most active learning research has focused on methods which perform well when many labels are available, but can be dramatically worse than random selection when label budgets are small. Other methods have focused on the low-budget regime, but do poorly as label budgets increase. As the line between "low" and "high" budgets varies by problem, this is a serious issue in practice. We propose *uncertainty coverage*, an objective which generalizes a variety of low- and high-budget objectives, as well as natural, hyperparameter-light methods to smoothly interpolate between low- and high-budget regimes. We call greedy optimization of the estimate Uncertainty Herding; this simple method is computationally fast, and we prove that it nearly optimizes the distribution-level coverage. In experimental validation across a variety of active learning tasks, our proposal matches or beats state-of-the-art performance in essentially all cases; it is the only method of which we are aware that reliably works well in both low- and high-budget settings.
Wonho Bae, Danica J. Sutherland, Gabriel L. Oliveira
ICLR3
2024 Forget Sharpness: Perturbed Forgetting of Model Biases Within SAM Dynamics
abstract
Despite attaining high empirical generalization, the sharpness of models trained with sharpness-aware minimization (SAM) do not always correlate with generalization error. Instead of viewing SAM as minimizing sharpness to improve generalization, our paper considers a new perspective based on SAM’s training dynamics. We propose that perturbations in SAM perform perturbed forgetting, where they discard undesirable model biases to exhibit learning signals that generalize better. We relate our notion of forgetting to the information bottleneck principle, use it to explain observations like the better generalization of smaller perturbation batches, and show that perturbed forgetting can exhibit a stronger correlation with generalization than flatness. While standard SAM targets model biases exposed by the steepest ascent directions, we propose a new perturbation that targets biases exposed through the model’s outputs. Our output bias forgetting perturbations outperform standard SAM, GSAM, and ASAM on ImageNet, robustness benchmarks, and transfer to CIFAR-10,100, while sometimes converging to sharper regions. Our results suggest that the benefits of SAM can be explained by alternative mechanistic principles that do not require flatness of the loss surface.
Ankit Vani, Frederick Tung, Gabriel L. Oliveira, Hossein Sharifi-Noghabi
ICML3
2023 Meta Temporal Point Processes
Wonho Bae, Mohamed Osama Ahmed, Frederick Tung, Gabriel L. Oliveira
ICLR4
2022 Heterogeneous Multi-Task Learning With Expert Diversity
abstract
Predicting multiple heterogeneous biological and medical targets is a challenge for traditional deep learning models. In contrast to single-task learning, in which a separate model is trained for each target, multi-task learning (MTL) optimizes a single model to predict multiple related targets simultaneously. To address this challenge, we propose the Multi-gate Mixture-of-Experts with Exclusivity (MMoEEx). Our work aims to tackle the heterogeneous MTL setting, in which the same model optimizes multiple tasks with different characteristics. Such a scenario can overwhelm current MTL approaches due to the challenges in balancing shared and task-specific representations and the need to optimize tasks with competing optimization paths. Our method makes two key contributions: first, we introduce an approach to induce more diversity among experts, thus creating representations more suitable for highly imbalanced and heterogenous MTL learning; second, we adopt a two-step optimization (Finn et al., 2017 and Lee et al., 2020) approach to balancing the tasks at the gradient level. We validate our method on three MTL benchmark datasets, including UCI-Census-income dataset, Medical Information Mart for Intensive Care (MIMIC-III) and PubChem BioAssay (PCBA).
Raquel Y. S. Aoki, Frederick Tung, Gabriel L. Oliveira
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 DPDB-Net: Exploiting Dense Connections for Convolutional Encoders
abstract
Densely connected networks for classification enable feature exploration and result in state-of-the-art performance on multiple classification tasks. The alternative to dense networks is the residual network which enables feature re-usage. In this work, we combine these orthogonal concepts for encoder-decoder architectures, which we call Dual-Path Dense-Block Network (DPDB-Net). We introduce a dense block which incorporates feature re-usage and new feature exploration in the encoder. Moreover, we discuss that feature re-usage by the residual network architecture leads to a feature map explosion in the decoder and, thus, is not advantageous in this part of the network. We evaluated our proposed architecture in multiple segmentation tasks and report state-of-the-art performance on the Freiburg Forest dataset and competitive results on the Cam Vid dataset.
Gabriel L. Oliveira, Wolfram Burgard, Thomas Brox
ICRA1
2017 Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection
abstract
General human action recognition requires understanding of various visual cues. In this paper, we propose a network architecture that computes and integrates the most important visual cues for action recognition: pose, motion, and the raw images. For the integration, we introduce a Markov chain model which adds cues successively. The resulting approach is efficient and applicable to action classification as well as to spatial and temporal action localization. The two contributions clearly improve the performance over respective baselines. The overall approach achieves state-of-the-art action classification performance on HMDB51, J-HMDB and NTU RGB+D datasets. Moreover, it yields state-of-the-art spatio-temporal action localization results on UCF101 and J-HMDB.
Mohammadreza Zolfaghari, Gabriel L. Oliveira, Nima Sedaghat, Thomas Brox
ICCV2
2017 Semantics-aware visual localization under challenging perceptual conditions
abstract
Visual place recognition under difficult perceptual conditions remains a challenging problem due to changing weather conditions, illumination and seasons. Long-term visual navigation approaches for robot localization should be robust to these dynamics of the environment. Existing methods typically leverage feature descriptions of whole images or image regions from Deep Convolutional Neural Networks. Some approaches also exploit sequential information to alleviate the problem of spatially inconsistent and non-perfect image matches. In this paper, we propose a novel approach for learning a discriminative holistic image representation which exploits the image content to create a dense and salient scene description. These salient descriptions are learnt over a variety of datasets under large perceptual changes. Such an approach enables us to precisely segment the regions of an image which are geometrically stable over large time lags. We combine features from these salient regions and an off-the-shelf holistic representation to form a more robust scene descriptor. We also introduce a semantically labeled dataset which captures extreme perceptual and structural scene dynamics over the course of 3 years. We evaluated our approach with extensive experiments on data collected over several kilometers in Freiburg and show that our learnt image representation outperforms off-the-shelf features from the deep networks and hand-crafted features.
Tayyab Naseer, Gabriel L. Oliveira, Thomas Brox, Wolfram Burgard
ICRA2
2017 Deep semantic classification for 3D LiDAR data
abstract
Robots are expected to operate autonomously in dynamic environments. Understanding the underlying dynamic characteristics of objects is a key enabler for achieving this goal. In this paper, we propose a method for pointwise semantic classification of 3D LiDAR data into three classes: non-movable, movable and dynamic. We concentrate on understanding these specific semantics because they characterize important information required for an autonomous system. To learn the distinction between movable and non-movable points in the environment, we introduce an approach based on deep neural network and for detecting the dynamic points, we estimate pointwise motion. We propose a Bayes filter framework for combining the learned semantic cues with the motion cues to infer the required semantic classification. In extensive experiments, we compare our approach with other methods on a standard benchmark dataset and report competitive results in comparison to the existing state-of-the-art. Furthermore, we show an improvement in the classification of points by combining the semantic cues retrieved from the neural network with the motion cues.
Ayush Dewan, Gabriel L. Oliveira, Wolfram Burgard
IROS2
2017 Perspectives on Deep Multimodel Robot Learning
Wolfram Burgard, Abhinav Valada, Noha Radwan, Tayyab Naseer, Jingwei Zhang 0001, Johan Vertens, Oier Mees, Andreas Eitel, Gabriel L. Oliveira
ISRR9
2017 Topometric Localization with Deep Learning
Gabriel L. Oliveira, Noha Radwan, Wolfram Burgard, Thomas Brox
ISRR1
2016 Deep learning for human part discovery in images
abstract
This paper addresses the problem of human body part segmentation in conventional RGB images, which has several applications in robotics, such as learning from demonstration and human-robot handovers. The proposed solution is based on Convolutional Neural Networks (CNNs). We present a network architecture that assigns each pixel to one of a predefined set of human body part classes, such as head, torso, arms, legs. After initializing weights with a very deep convolutional network for image classification, the network can be trained end-to-end and yields precise class predictions at the original input resolution. Our architecture particularly improves on over-fitting issues in the up-convolutional part of the network. Relying only on RGB rather than RGB-D images also allows us to apply the approach outdoors. The network achieves state-of-the-art performance on the PASCAL Parts dataset. Moreover, we introduce two new part segmentation datasets, the Freiburg sitting people dataset and the Freiburg people in disaster dataset. We also present results obtained with a ground robot and an unmanned aerial vehicle.
Gabriel L. Oliveira, Abhinav Valada, Claas Bollen, Wolfram Burgard, Thomas Brox
ICRA1
2016 Efficient deep models for monocular road segmentation
abstract
This paper addresses the problem of road scene segmentation in conventional RGB images by exploiting recent advances in semantic segmentation via convolutional neural networks (CNNs). Segmentation networks are very large and do not currently run at interactive frame rates. To make this technique applicable to robotics we propose several architecture refinements that provide the best trade-off between segmentation quality and runtime. This is achieved by a new mapping between classes and filters at the expansion side of the network. The network is trained end-to-end and yields precise road/lane predictions at the original input resolution in roughly 50ms. Compared to the state of the art, the network achieves top accuracies on the KITTI dataset for road and lane segmentation while providing a 20× speed-up. We demonstrate that the improved efficiency is not due to the road segmentation task. Also on segmentation datasets with larger scene complexity, the accuracy does not suffer from the large speed-up.
Gabriel L. Oliveira, Wolfram Burgard, Thomas Brox
IROS1
2014 On the improvement of human action recognition from depth map sequences using Space-Time Occupancy Patterns
Antônio Wilson Vieira, Erickson R. Nascimento, Gabriel L. Oliveira, Zicheng Liu 0001, Mario Fernando Montenegro Campos
Pattern Recognit. Lett.3
2014 Sparse Spatial Coding: A Novel Approach to Visual Recognition
abstract
Successful image-based object recognition techniques have been constructed founded on powerful techniques such as sparse representation, in lieu of the popular vector quantization approach. However, one serious drawback of sparse space-based methods is that local features that are quite similar can be quantized into quite distinct visual words. We address this problem with a novel approach for object recognition, called sparse spatial coding, which efficiently combines a sparse coding dictionary learning and spatial constraint coding stage. We performed experimental evaluation using the Caltech 101, Caltech 256, Corel 5000, and Corel 10000 data sets, which were specifically designed for object recognition evaluation. Our results show that our approach achieves high accuracy comparable with the best single feature method previously published on those databases. Our method outperformed, for the same bases, several multiple feature methods, and provided equivalent, and in few cases, slightly less accurate results than other techniques specifically designed to that end. Finally, we report state-of-the-art results for scene recognition on COsy Localization Dataset (COLD) and high performance results on the MIT-67 indoor scene recognition, thus demonstrating the generalization of our approach for such tasks.
Gabriel L. Oliveira, Erickson R. Nascimento, Antônio Wilson Vieira, Mario Fernando Montenegro Campos
IEEE Trans. Image Process.1
2013 On the development of a robust, fast and lightweight keypoint descriptor
Erickson R. Nascimento, Gabriel L. Oliveira, Antônio Wilson Vieira, Mario Fernando Montenegro Campos
Neurocomputing2
2012 STOP: Space-Time Occupancy Patterns for 3D Action Recognition from Depth Map Sequences
Antônio Wilson Vieira, Erickson R. Nascimento, Gabriel L. Oliveira, Zicheng Liu 0001, Mario Fernando Montenegro Campos
CIARP3
2012 Sparse Spatial Coding: A novel approach for efficient and accurate object recognition
abstract
Successful state-of-the-art object recognition techniques from images have been based on powerful methods, such as sparse representation, in order to replace the also popular vector quantization (VQ) approach. Recently, sparse coding, which is characterized by representing a signal in a sparse space, has raised the bar on several object recognition benchmarks. However, one serious drawback of sparse space based methods is that similar local features can be quantized into different visual words. We present in this paper a new method, called Sparse Spatial Coding (SSC), which combines a sparse coding dictionary learning, a spatial constraint coding stage and an online classification method to improve object recognition. An efficient new off-line classification algorithm is also presented. We overcome the problem of techniques which make use of sparse representation alone by generating the final representation with SSC and max pooling, presented for an online learning classifier. Experimental results obtained on the Caltech 101, Caltech 256, Corel 5000 and Corel 10000 databases, show that, to the best of our knowledge, our approach supersedes in accuracy the best published results to date on the same databases. As an extension, we also show high performance results on the MIT-67 indoor scene recognition dataset.
Gabriel L. Oliveira, Erickson R. Nascimento, Antônio Wilson Vieira, Mario Fernando Montenegro Campos
ICRA1
2012 BRAND: A robust appearance and depth descriptor for RGB-D images
abstract
This work introduces a novel descriptor called Binary Robust Appearance and Normals Descriptor (BRAND), that efficiently combines appearance and geometric shape information from RGB-D images, and is largely invariant to rotation and scale transform. The proposed approach encodes point information as a binary string providing a descriptor that is suitable for applications that demand speed performance and low memory consumption. Results of several experiments demonstrate that as far as precision and robustness are concerned, BRAND achieves improved results when compared to state of the art descriptors based on texture, geometry and combination of both information. We also demonstrate that our descriptor is robust and provides reliable results in a registration task even when a sparsely textured and poorly illuminated scene is used.
Erickson R. Nascimento, Gabriel L. Oliveira, Mario Fernando Montenegro Campos, Antônio Wilson Vieira, William Robson Schwartz
IROS2