Cristina Nader Vasconcelos

dblp:37/135 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0003-2112-4806ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts
Ibtihel Amara, Ahmed Imtiaz Humayun, Ivana Kajic, Zarana Parekh, Natalie Harris, Sarah Young, Chirag Nagpal, Najoung Kim, Junfeng He, Cristina Nader Vasconcelos, Deepak Ramachandran, Golnoosh Farnadi, Katherine A. Heller, Mohammad Havaei, Negar Rostamzadeh
ICCV10
2025 What Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models
abstract
Deep Generative Models are frequently used to learn continuous representations of complex data distributions by training on a finite number of samples. For any generative model, including pre-trained foundation models with Diffusion or Transformer architectures, generation performance can significantly vary across the learned data manifold. In this paper, we study the local geometry of the learned manifold and its relationship to generation outcomes for a wide range of generative models, including DDPM, Diffusion Transformer (DiT), and Stable Diffusion 1.4. Building on the theory of continuous piecewise-linear (CPWL) generators, we characterize the local geometry in terms of three geometric descriptors - scaling ($\psi$), rank ($\nu$), and complexity/un-smoothness ($\delta$). We provide quantitative and qualitative evidence showing that for a given latent vector, the local descriptors are indicative of post-generation aesthetics, generation diversity, and memorization by the generative model. Finally, we demonstrate that by training a reward model on the 'local scaling' for Stable Diffusion, we can self-improve both generation aesthetics and diversity using geometry sensitive guidance during denoising. Website: https://imtiazhumayun.github.io/generative_geometry.
Ahmed Imtiaz Humayun, Ibtihel Amara, Cristina Nader Vasconcelos, Deepak Ramachandran, Candice Schumann, Junfeng He, Katherine A. Heller, Golnoosh Farnadi, Negar Rostamzadeh, Mohammad Havaei
ICLR3
2024 PLANTED: A Dataset for Planted Forest Identification from Multi-Satellite Time Series
abstract
Protecting and restoring forest ecosystems is critical for bio-diversity conservation and climate change mitigation. Forest monitoring on a global scale is essential for prioritizing and assessing conservation efforts. Satellite-based remote sensing is the only viable solution for providing global coverage, but to date, large-scale forest monitoring is limited mostly to single modalities or single time points. In this paper, we present a dataset consisting of data from five public satellites for recognizing forest plantations and planted tree species across the globe. Each satellite modality consists of a multi-year time series. The dataset, named "Planted" includes over 2M examples of 64 tree label classes (46 genera and 40 distinctive species), distributed among 41 countries. This dataset is released to foster research in forest monitoring using multimodal, multi-scale, multi-temporal data sources. Additionally, we present initial baseline results and evaluate modality fusion and data augmentation approaches for this dataset.
Luis Miguel Pazos-Outón, Cristina Nader Vasconcelos, Anton Raichuk, Anurag Arnab, Dan Morris 0001, Maxim Neumann
IGARSS2
2023 CUF: Continuous Upsampling Filters
abstract
Neural fields have rapidly been adopted for representing 3D signals, but their application to more classical 2D image-processing has been relatively limited. In this paper, we consider one of the most important operations in image processing: upsampling. In deep learning, learnable upsampling layers have extensively been used for single image super-resolution. We propose to parameterize upsampling kernels as neural fields. This parameterization leads to a compact architecture that obtains a 40-fold reduction in the number of parameters when compared with competing arbitrary-scale super-resolution architectures. When upsampling images of size 256×256 we show that our architecture is 2x-10x more efficient than competing arbitrary-scale super-resolution architectures, and more efficient than sub-pixel convolutions when instantiated to a single-scale model. In the general setting, these gains grow polynomially with the square of the target scale. We validate our method on standard benchmarks showing such efficiency gains can be achieved without sacrifices in super-resolution performance. https://cuf-paper.github.io
Cristina Nader Vasconcelos, A. Cengiz Öztireli, Mark J. Matthews, Milad Hashemi, Kevin Swersky, Andrea Tagliasacchi
CVPR1
2023 Scaling Vision Transformers to 22 Billion Parameters
abstract
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there.
Mostafa Dehghani 0001, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner 0001, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang 0038, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu 0001, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Collier, Alexey A. Gritsenko, Vighnesh Birodkar, Cristina Nader Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov 0003, Filip Pavetic, Dustin Tran, Thomas Kipf, Mario Lucic, Xiaohua Zhai, Daniel Keysers, Jeremiah J. Harmsen, Neil Houlsby
ICML31
2023 An analysis of ConformalLayers' robustness to corruptions in natural images
Eduardo Vera Sousa, Cristina Nader Vasconcelos, Leandro A. F. Fernandes
Pattern Recognit. Lett.2
2022 Proper Reuse of Image Classification Features Improves Object Detection
abstract
A common practice in transfer learning is to initialize the downstream model weights by pre-training on a data-abundant upstream task. In object detection specifically, the feature backbone is typically initialized with ImageNet classifier weights and fine-tuned on the object detection task. Recent works show this is not strictly necessary under longer training regimes and provide recipes for training the backbone from scratch. We investigate the opposite direction of this end-to-end training trend: we show that an extreme form of knowledge preservation-freezing the classifier-initialized backbone— consistently improves many different detection models, and leads to considerable resource savings. We hypothesize and corroborate experimentally that the remaining detector components capacity and structure is a crucial factor in leveraging the frozen backbone. Immediate applications of our findings include performance improvements on hard cases like detection of long-tail object classes and computational and memory resource savings that contribute to making the field more accessible to researchers with access to fewer computational resources.
Cristina Nader Vasconcelos, Vighnesh Birodkar, Vincent Dumoulin
CVPR1
2022 Spatially and color consistent environment lighting estimation using deep neural networks for mixed reality
Bruno A. D. Marques, Esteban Walter Gonzalez Clua, Anselmo Antunes Montenegro, Cristina Nader Vasconcelos
Comput. Graph.4
2021 Impact of Aliasing on Generalization in Deep Convolutional Networks
abstract
We investigate the impact of aliasing on generalization in Deep Convolutional Networks and show that data augmentation schemes alone are unable to prevent it due to structural limitations in widely used architectures. Drawing insights from frequency analysis theory, we take a closer look at ResNet and EfficientNet architectures and review the trade-off between aliasing and information loss in each of their major components. We show how to mitigate aliasing by inserting non-trainable low-pass filters at key locations, particularly where networks lack the capacity to learn them. These simple architectural changes lead to substantial improvements in generalization on i.i.d. and even more on out-of-distribution conditions, such as image classification under natural corruptions on ImageNet-C [11] and few-shot learning on Meta-Dataset [26]. State-of-the art results are achieved on both datasets without introducing additional trainable parameters and using the default hyper-parameters of open source codebases.
Cristina Nader Vasconcelos, Hugo Larochelle, Vincent Dumoulin, Rob Romijnders, Nicolas Le Roux, Ross Goroshin
ICCV1
2020 Experiments using deep learning for dermoscopy image analysis
Cristina Nader Vasconcelos, Bárbara Nader Vasconcelos
Pattern Recognit. Lett.1
2019 A Model Based on LSTM Neural Networks to Identify Five Different Types of Malware
abstract
Identifying malware has always been a great challenge. Much money and time has been invested by companies and governments to mitigate the impact of these threats. Nowadays, with the increasing amount of data available, it is possible to use more precise classification techniques. However, most large datasets that include malicious and non-malicious softwares are not public, which hinders the quest for solutions based in technologies that rely on the availability of large amounts of data, such as deep learning. To overcome this limitation, this article introduces a new large dataset for malware classification, which was made publicly available. We then propose a model to train a multiclass classification recurrent neural network (RNN), more specifically a long short-term memory neural network (LSTM) on our dataset. This model for analyzing unstructured malware data is then tested on unseen programs and the accuracy obtained reaches 67.60%, including six classes with five different types of malware.
Eduardo de Oliveira Andrade, José Viterbo, Cristina Nader Vasconcelos, Joris Guérin, Flavia Bernardini
KES3
2018 Deep spherical harmonics light probe estimator for mixed reality games
Bruno A. D. Marques, Esteban Walter Gonzalez Clua, Cristina Nader Vasconcelos
Comput. Graph.3
2016 Towards Deep Learning Invariant Pedestrian Detection by Data Enrichment
abstract
Deep learning models have recently achieved the state-of-the-art results on a well-known pedestrian detection dataset. However, such images were obtained from open scenarios with fixed imaging geometry parameters, which may produce a network not suitable for detecting a person in more general settings, such as the ones found in surveillance systems. As gathering and annotating data is a highly expensive manual task, we propose a methodology for artificially augmenting the positive training set with automatically generated local image affine and perspective transforms. Furthermore, to enrich the variability of background images, we include to the negative training set images that resemble human figures automatically obtained by the proposed methodology over images from commonly found surveillance scenarios. Extensive results show that by providing the enriched data as the input to a Convolutional Neural Network it is possible to precisely detect pedestrians in a number of public datasets. The data enrichment proposed here may also be used in other detectors based on supervised learning architectures, as the process is independent from the learning algorithm employed.
Cristina Nader Vasconcelos, Aline Paes, Anselmo Antunes Montenegro
ICMLA1
2015 Revisiting ant colony algorithms to seismic faults detection
Walther Maciel, Cristina Nader Vasconcelos, Pedro Mario Silva, Marcelo Gattass
ESANN2
2015 Unsupervised cosegmentation based on global clustering and saliency
abstract
This paper introduces a new method for unsupervised cosegmentation. Our method combines saliency information with a Global Clustering step, which reveals parts of the objects by detecting similar subregions across image collections, based on a low dimensional descriptor that includes color, texture and positional features. The saliency information is used to yield a classification of the global clusters into foreground and background and also classify regions not detected as global clusters into potential background or foreground. These four types of regions are the input seeds for a Graph Cuts procedure that computes the final cosegmentation. The Graph Cuts result can also be used to compute a refined version of the saliency information which enables us to define an iterative cosegmentation pipeline. Our framework produces remarkable results in comparison with state-of-the-art works, even in challenging datasets with illumination variance, occluded objects and identical background.
Lucas Lattari, Anselmo Antunes Montenegro, Cristina Nader Vasconcelos
ICIP3
2008 Real-Time Video Processing for Multi-Object Chromatic Tracking
abstract
This paper presents MOCT, a multi-object chromatic tracking technique for real-time natural video processing. Its main step is the MOCT localization algorithm, that performs local data evaluations in order to apply a multiple output parallel reduction operator to the image. The reduction operator is used to localize the positions of the object centroids, to compute the number of pixels occupied by an object and its bounding boxes, and to update object trajectories in image space. The operator is analyzed using three different computation layouts and tested over several reduction factors. 1
Cristina Nader Vasconcelos, Asla Medeiros Sá, Lucas Teixeira, Paulo C. P. Carvalho, Marcelo Gattass
BMVC1
2006 Shot segmentation based on the encoder signature
abstract
This paper presents a new approach for finding camera shot transitions in the compressed stream of MPEG-1 and MPEG-2 videos. False transitions patterns on the compressed data are considered to be related to the choices made by the encoder, which has the freedom of making several decisions without violating the MPEG-1 and MPEG-2 standards. These choices contaminate the similarity metrics with patterns that are called encoder signatures in this paper. The present work proposes a video segmentation method that filters the encoder signatures, through a step-wise refinement process using multiple similarity metrics. This process is general and can be easily adapted to MPEG-4 videos
Cristina Nader Vasconcelos, Bruno Feijó, Dilza Szwarcman, Mónica Costa
MMM1