Ricardo Silveira Cabral

dblp:65/10697 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
2since 2021 · last 2025
0000-0002-4919-8711ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
3D vision · 21% Language models and text generation · 20% Video understanding and tracking · 12%
Computer graphics and multimedia
1 paper
Rendering · 100%
Theoretical computer science
2 papers
Mathematical optimization · 65% Algorithms and data structures · 35%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
0.912025
Logic.py: Bridging the Gap between LLMs and Constraint Solvers · NeurIPS 2025
Computer vision › 3D vision
implicit neural representation
0.612022
PINs: Progressive Implicit Networks for Multi-Scale Neural Representations · ICML 2022
Rendering › neural rendering
neural scene representation
0.612022
PINs: Progressive Implicit Networks for Multi-Scale Neural Representations · ICML 2022
Machine learning › Learning paradigms
multi-label classification
0.322015
Matrix Completion for Weakly-Supervised Multi-Label Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Matrix Completion for Multi-label Image Classification · NIPS 2011
Machine learning › Learning theory › statistical estimation › robust statistics
robust regression
0.322016
Robust Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Robust Regression · ECCV (4) 2012
Computer vision › 3D vision
3d object detection
0.212016
Motion from Structure (MfS): Searching for 3D Objects in Cluttered Point Trajectories · CVPR 2016
Computer vision › Video understanding and tracking
action recognition
0.212016
Feature and Region Selection for Visual Learning · IEEE Trans. Image Process. 2016
Machine learning › Representation and self-supervised learning › visual representation › image representation
bag of visual words
0.212016
Feature and Region Selection for Visual Learning · IEEE Trans. Image Process. 2016
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.212016
Feature and Region Selection for Visual Learning · IEEE Trans. Image Process. 2016
Computer vision › Video understanding and tracking
motion segmentation
0.212016
Motion from Structure (MfS): Searching for 3D Objects in Cluttered Point Trajectories · CVPR 2016
Computer vision › Image recognition and object detection
region selection
0.212016
Feature and Region Selection for Visual Learning · IEEE Trans. Image Process. 2016
Machine learning › Trustworthy machine learning
robustness
0.212016
Robust Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification
0.212015
Matrix Completion for Weakly-Supervised Multi-Label Image Classification · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Algorithms and data structures › numerical linear algebra › matrix factorization
low-rank matrix factorization
0.212013
Unifying Nuclear Norm and Bilinear Factorization Approaches for Low-Rank Matrix Decomposition · ICCV 2013
Mathematical optimization › continuous optimization › convex optimization › norm optimization
nuclear norm minimization
0.212013
Unifying Nuclear Norm and Bilinear Factorization Approaches for Low-Rank Matrix Decomposition · ICCV 2013
Mathematical optimization › optimization under uncertainty
robust optimization
0.112012
Robust Regression · ECCV (4) 2012
Machine learning › Learning theory
matrix completion
0.112011
Matrix Completion for Multi-label Image Classification · NIPS 2011
Machine learning and data management
matrix completion
0.112011
Matrix Completion for Multi-label Image Classification · NIPS 2011
Computer vision › 3D vision
photometric stereo
0.012013
Unifying Nuclear Norm and Bilinear Factorization Approaches for Low-Rank Matrix Decomposition · ICCV 2013
Computer vision › 3D vision
structure from motion
0.012013
Unifying Nuclear Norm and Bilinear Factorization Approaches for Low-Rank Matrix Decomposition · ICCV 2013
Machine learning › Learning theory
statistical learning theory
0.012012
Robust Regression · ECCV (4) 2012

Methods — techniques the papers use, named apart from their topics

positional encoding · 1.1prompting · 0.9domain-specific language · 0.9constraint solver · 0.9multilayer perceptron · 0.6multi-layer perceptron · 0.6rank minimization · 0.5convex optimization · 0.5point trajectory analysis · 0.2kernel regression · 0.23d shape models · 0.2robust PCA · 0.2rank continuation · 0.2robust statistics · 0.1
YearPublicationVenuePosition
2025 Logic.py: Bridging the Gap between LLMs and Constraint Solvers
abstract
We present a novel approach to formalise and solve search-based problems using large language models, which significantly improves upon previous state-of-the-art results. We demonstrate the efficacy of this approach on benchmarks like the logic puzzles tasks in ZebraLogicBench. Instead of letting the LLM attempt to directly solve the puzzles, our method prompts the model to formalise the problem in a logic-focused, human-readable domain-specific language (DSL) called Logic.py. This formalised representation is then solved using a constraint solver, leveraging the strengths of both the language model and the solver. Our approach achieves a remarkable 65% absolute improvement over the baseline performance of Llama 3.1 70B on ZebraLogicBench, setting a new state-of-the-art with an accuracy of over 90%. This significant advancement demonstrates the potential of combining language models with domain-specific languages and auxiliary tools on traditionally challenging tasks for LLMs.
Pascal Kesseli, Peter W. O'Hearn, Ricardo Silveira Cabral
NeurIPS3
2022 PINs: Progressive Implicit Networks for Multi-Scale Neural Representations
abstract
Multi-layer perceptrons (MLP) have proven to be effective scene encoders when combined with higher-dimensional projections of the input, commonly referred to as positional encoding. However, scenes with a wide frequency spectrum remain a challenge: choosing high frequencies for positional encoding introduces noise in low structure areas, while low frequencies results in poor fitting of detailed regions. To address this, we propose a progressive positional encoding, exposing a hierarchical MLP structure to incremental sets of frequency encodings. Our model accurately reconstructs scenes with wide frequency bands and learns a scene representation at progressive level of detail without explicit per-level supervision. The architecture is modular: each level encodes a continuous implicit representation that can be leveraged separately for its respective resolution, meaning a smaller network for coarser reconstructions. Experiments on several 2D and 3D datasets shows improvements in reconstruction accuracy, representational capacity and training speed compared to baselines.
Zoe Landgraf, Alexander Sorkine-Hornung, Ricardo Silveira Cabral
ICML3
2016 Motion from Structure (MfS): Searching for 3D Objects in Cluttered Point Trajectories
abstract
Object detection has been a long standing problem in computer vision, and state-of-the-art approaches rely on the use of sophisticated features and/or classifiers. However, these learning-based approaches heavily depend on the quality and quantity of labeled data, and do not generalize well to extreme poses or textureless objects. In this work, we explore the use of 3D shape models to detect objects in videos in an unsupervised manner. We call this problem Motion from Structure (MfS): given a set of point trajectories and a 3D model of the object of interest, find a subset of trajectories that correspond to the 3D model and estimate its alignment (i.e., compute the motion matrix). MfS is related to Structure from Motion (SfM) and motion segmentation problems: unlike SfM, the structure of the object is known but the correspondence between the trajectories and the object is unknown, unlike motion segmentation, the MfS problem incorporates 3D structure, providing robustness to tracking mismatches and outliers. Experiments illustrate how our MfS algorithm outperforms alternative approaches in both synthetic data and real videos extracted from YouTube.
Jayakorn Vongkulbhisal, Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira
CVPR2
2016 Robust Regression
abstract
Discriminative methods (e.g., kernel regression, SVM) have been extensively used to solve problems such as object recognition, image alignment and pose estimation from images. These methods typically map image features ( X) to continuous (e.g., pose) or discrete (e.g., object category) values. A major drawback of existing discriminative methods is that samples are directly projected onto a subspace and hence fail to account for outliers common in realistic training sets due to occlusion, specular reflections or noise. It is important to notice that existing discriminative approaches assume the input variables X to be noise free. Thus, discriminative methods experience significant performance degradation when gross outliers are present. Despite its obvious importance, the problem of robust discriminative learning has been relatively unexplored in computer vision. This paper develops the theory of robust regression (RR) and presents an effective convex approach that uses recent advances on rank minimization. The framework applies to a variety of problems in computer vision including robust linear discriminant analysis, regression with missing data, and multi-label classification. Several synthetic and real examples with applications to head pose estimation from images, image and video classification and facial attribute classification with missing data are used to illustrate the benefits of RR.
Dong Huang 0007, Ricardo Silveira Cabral, Fernando De la Torre
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Feature and Region Selection for Visual Learning
abstract
Visual learning problems, such as object classification and action recognition, are typically approached using extensions of the popular bag-of-words (BoWs) model. Despite its great success, it is unclear what visual features the BoW model is learning. Which regions in the image or video are used to discriminate among classes? Which are the most discriminative visual words? Answering these questions is fundamental for understanding existing BoW models and inspiring better models for visual recognition. To answer these questions, this paper presents a method for feature selection and region selection in the visual BoW model. This allows for an intermediate visualization of the features and regions that are important for visual learning. The main idea is to assign latent weights to the features or regions, and jointly optimize these latent variables with the parameters of a classifier (e.g., support vector machine). There are four main benefits of our approach: 1) our approach accommodates non-linear additive kernels, such as the popular χ(2) and intersection kernel; 2) our approach is able to handle both regions in images and spatio-temporal regions in videos in a unified way; 3) the feature selection problem is convex, and both problems can be solved using a scalable reduced gradient method; and 4) we point out strong connections with multiple kernel learning and multiple instance learning approaches. Experimental results in the PASCAL VOC 2007, MSR Action Dataset II and YouTube illustrate the benefits of our approach.
Ji Zhao 0001, Liantao Wang, Ricardo Silveira Cabral, Fernando De la Torre
IEEE Trans. Image Process.3
2015 Matrix Completion for Weakly-Supervised Multi-Label Image Classification
abstract
In the last few years, image classification has become an incredibly active research topic, with widespread applications. Most methods for visual recognition are fully supervised, as they make use of bounding boxes or pixelwise segmentations to locate objects of interest. However, this type of manual labeling is time consuming, error prone and it has been shown that manual segmentations are not necessarily the optimal spatial enclosure for object classifiers. This paper proposes a weakly-supervised system for multi-label image classification. In this setting, training images are annotated with a set of keywords describing their contents, but the visual concepts are not explicitly segmented in the images. We formulate the weakly-supervised image classification as a low-rank matrix completion problem. Compared to previous work, our proposed framework has three advantages: (1) Unlike existing solutions based on multiple-instance learning methods, our model is convex. We propose two alternative algorithms for matrix completion specifically tailored to visual data, and prove their convergence. (2) Unlike existing discriminative methods, our algorithm is robust to labeling errors, background noise and partial occlusions. (3) Our method can potentially be used for semantic segmentation. Experimental validation on several data sets shows that our method outperforms state-of-the-art classification algorithms, while effectively capturing each class appearance.
Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Multi-label Discriminative Weakly-Supervised Human Activity Recognition and Localization
Ehsan Adeli-Mosabbeb, Ricardo Silveira Cabral, Fernando De la Torre, Mahmood Fathy
ACCV (5)2
2014 Optimal no-intersection multi-label binary localization for time series using totally unimodular linear programming
abstract
We propose a new model for simultaneously localizing different classes in the same media, casting it as an integer optimization problem. Our model subsumes into a single formulation previous single and multi-class localization methods, as well as allows us to exploit optimal relaxations to the linear domain. We apply our model to the problem of multi-label multiple instance learning for tagging video collections. Given weakly labeled training samples, where tags for actions in video and objects in images are known but not their locations, our aim is to train classifiers for both detection and localization of said classes on new data. Experimental results demonstrate our approach obtains similar performances when compared to fully supervised methods.
Ricardo Silveira Cabral, João Paulo Costeira, Alexandre Bernardino, Fernando De la Torre
ICIP1
2013 Unifying Nuclear Norm and Bilinear Factorization Approaches for Low-Rank Matrix Decomposition
abstract
Low rank models have been widely used for the representation of shape, appearance or motion in computer vision problems. Traditional approaches to fit low rank models make use of an explicit bilinear factorization. These approaches benefit from fast numerical methods for optimization and easy kernelization. However, they suffer from serious local minima problems depending on the loss function and the amount/type of missing data. Recently, these low-rank models have alternatively been formulated as convex problems using the nuclear norm regularizer, unlike factorization methods, their numerical solvers are slow and it is unclear how to kernelize them or to impose a rank a priori. This paper proposes a unified approach to bilinear factorization and nuclear norm regularization, that inherits the benefits of both. We analyze the conditions under which these approaches are equivalent. Moreover, based on this analysis, we propose a new optimization algorithm and a "rank continuation'' strategy that outperform state-of-the-art approaches for Robust PCA, Structure from Motion and Photometric Stereo with outliers and missing data.
Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino
ICCV1
2012 Robust Regression
Dong Huang 0007, Ricardo Silveira Cabral, Fernando De la Torre
ECCV (4)2
2011 Fast incremental method for matrix completion: An application to trajectory correction
abstract
We address the problem of incrementally recovering a matrix of tracked image points, based on partial observations of their trajectories. Besides partial observability, we assume the existence of gross, but sparse, noise on the known entries. This problem has obvious applications in real-time tracking and structure from motion, where observations are plagued by self-occlusion and outliers. Recently, work in the optimization community has spun optimal methods for matrix completion when this matrix is known to be low rank by minimizing the nuclear norm, the sum of its singular values. Despite exhibiting several optimality properties, no available algorithms perform this minimization incrementally. In this paper, we build upon the Nuclear Norm Robust PCA method and SPectrally Optimal Completion to propose a fast and incremental algorithm which is able to cope with outliers. We present experiments showing the competitive speed of our method while maintaining performance comparable to the state-of-the-art.
Ricardo Silveira Cabral, João Paulo Costeira, Fernando De la Torre, Alexandre Bernardino
ICIP1
2011 Matrix Completion for Multi-label Image Classification
abstract
Recently, image categorization has been an active research topic due to the urgent need to retrieve and browse digital images via semantic keywords. This paper formulates image categorization as a multi-label classification problem using recent advances in matrix completion. Under this setting, classification of testing data is posed as a problem of completing unknown label entries on a data matrix that concatenates training and testing features with training labels. We propose two convex algorithms for matrix completion based on a Rank Minimization criterion specifically tailored to visual data, and prove its convergence properties. A major advantage of our approach w.r.t. standard discriminative classification methods for image categorization is its robustness to outliers, background noise and partial occlusions both in the feature and label space. Experimental validation on several datasets shows how our method outperforms state-of-the-art algorithms, while effectively capturing semantic concepts of classes.
Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino
NIPS1