Victor G. T. da Costa

dblp:203/9327 · also Victor G. Turrisi da Costa, Victor Guilherme Turrisi da Costa · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0003-2597-4998ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Multimodal Autoregressive Pre-training of Large Vision Encoders
abstract
We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.
Enrico Fini, Mustafa Shukor, Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Victor G. T. da Costa, Louis Béthune, Zhe Gan, Alexander Toshev, Marcin Eichner, Moin Nabi, Yinfei Yang, Joshua M. Susskind, Alaaeldin El-Nouby
CVPR8
2025 Scaling Laws for Native Multimodal Models
abstract
Building general-purpose models that can effectively perceive the world through multimodal signals has been a long-standing goal. Current approaches involve integrating separately pre-trained components, such as connecting vision encoders to LLMs and continuing multimodal training. While such approaches exhibit remarkable sample efficiency, it remains an open question whether such late-fusion architectures are inherently superior. In this work, we revisit the architectural design of native multimodal models (NMMs)-those trained from the ground up on all modalities-and conduct an extensive scaling laws study, spanning 457 trained models with different architectures and training mixtures. Our investigation reveals no inherent advantage to late-fusion architectures over early-fusion ones, which do not rely on image encoders or tokenizers. On the contrary, early-fusion exhibits stronger performance at lower parameter counts, is more efficient to train, and is easier to deploy. Motivated by the strong performance of the early-fusion architectures, we show that incorporating Mixture of Experts (MoEs) allows models to learn modality-specific weights, significantly benefiting performance.
Mustafa Shukor, Enrico Fini, Victor G. T. da Costa, Matthieu Cord, Joshua M. Susskind, Alaaeldin El-Nouby
ICCV3
2024 Enhancing the Domain Robustness of Self-Supervised pre-Training with Synthetic Images
abstract
We present a novel method for improving the adaptability of self-supervised (SSL) pre-trained models across different domains. Our approach uses synthetic images that are generated using an auxiliary diffusion model, namely InstructPix2Pix. More specifically, starting from a real image, we prompt the diffusion model to generate synthetic versions of that image in the style of the target domains. This allows us to generate a diverse set of multi-domain images that share the same semantics as real images. Integrating these synthetic images into the training dataset enhances the model’s capacity to generalize to other domains. We pre-trained different SSL methods on Imagenet-100 with and without the synthetic images and evaluated their performance on three multi-domain datasets, DomainNet, PACS, and Office-Home. Our results show significant improvements in all datasets and methods, encouraging new research in the direction of leveraging synthetic data to improve the robustness of pre-trained models. Code is available at https://github.com/has97/Diffusion_pre-training.
Mohamad Hassan N C, Avigyan Bhattacharya, Victor G. T. da Costa, Biplab Banerjee, Elisa Ricci 0001
ICASSP3
2024 Simplifying open-set video domain adaptation with contrastive learning
Giacomo Zara, Victor G. T. da Costa, Subhankar Roy, Paolo Rota, Elisa Ricci 0001
Comput. Vis. Image Underst.2
2023 Bayesian Prompt Learning for Image-Language Model Generalization
abstract
Foundational image-language models have generated considerable interest due to their efficient adaptation to downstream tasks by prompt learning. Prompt learning treats part of the language model input as trainable while freezing the rest, and optimizes an Empirical Risk Mini-mization objective. However, Empirical Risk Minimization is known to suffer from distributional shifts which hurt gen-eralizability to prompts unseen during training. By leveraging the regularization ability of Bayesian methods, we frame prompt learning from the Bayesian perspective and formulate it as a variational inference problem. Our approach regularizes the prompt space, reduces overfitting to the seen prompts and improves the prompt generalization on unseen prompts. Our framework is implemented by modeling the input prompt space in a probabilistic manner, as an a priori distribution which makes our proposal compatible with prompt learning approaches that are unconditional or conditional on the image. We demonstrate empirically on 15 benchmarks that Bayesian prompt learning provides an appropriate coverage of the prompt space, prevents learning spurious features, and exploits transferable invariant features. This results in better generalization of unseen prompts, even across different datasets and domains.Code available at: https://github.com/saic-fi/Bayesian-Prompt-Learning
Mohammad Mahdi Derakhshani, Enrique Sanchez, Adrian Bulat, Victor G. T. da Costa, Cees Snoek, Georgios Tzimiropoulos, Brais Martínez
ICCV4
2022 Self-Supervised Models are Continual Learners
abstract
Self-supervised models have been shown to produce comparable or better visual representations than their su-pervised counterparts when trained offline on unlabeled data at scale. However, their efficacy is catastrophically reduced in a Continual Learning (CL) scenario where data is presented to the model sequentially. In this paper, we show that self-supervised loss functions can be seamlessly converted into distillation mechanisms for CL by adding a predictor network that maps the current state of the repre-sentations to their past state. This enables us to devise a framework for Continual self-supervised visual representation Learning that (i) significantly improves the quality of the learned representations, (ii) is compatible with several state-of-the-art self-supervised objectives, and (iii) needs little to no hyperparameter tuning. We demonstrate the ef-fectiveness of our approach empirically by training six pop-ular self-supervised models in various CL settings. Code: github.com/DonkeyShot21/cassle.
Enrico Fini, Victor G. T. da Costa, Xavier Alameda-Pineda, Elisa Ricci 0001, Karteek Alahari, Julien Mairal
CVPR2
2022 Unsupervised Domain Adaptation for Video Transformers in Action Recognition
abstract
Over the last few years, Unsupervised Domain Adaptation (UDA) techniques have acquired remarkable importance and popularity in computer vision. However, when compared to the extensive literature available for images, the field of videos is still relatively unexplored. On the other hand, the performance of a model in action recognition is heavily affected by domain shift. In this paper, we propose a simple and novel UDA approach for video action recognition. Our approach leverages recent advances on spatio-temporal transformers to build a robust source model that better generalises to the target domain. Furthermore, our architecture learns domain invariant features thanks to the introduction of a novel alignment loss term derived from the Information Bottleneck principle. We report results on two video action recognition benchmarks for UDA, showing state-of-the-art performance on HMDB ↔ UCF, as well as on Kinetics→NEC-Drone, which is more challenging. This demonstrates the effectiveness of our method in handling different levels of domain shift. The source code is available at https://github.com/vturrisi/UDAVT.
Victor G. T. da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, Elisa Ricci 0001
ICPR1
2022 Dual-Head Contrastive Domain Adaptation for Video Action Recognition
abstract
Unsupervised domain adaptation (UDA) methods have become very popular in computer vision. However, while several techniques have been proposed for images, much less attention has been devoted to videos. This paper introduces a novel UDA approach for action recognition from videos, inspired by recent literature on contrastive learning. In particular, we propose a novel two-headed deep architecture that simultaneously adopts cross-entropy and contrastive losses from different network branches to robustly learn a target classifier. Moreover, this work introduces a novel large-scale UDA dataset, Mixamo→Kinetics, which, to the best of our knowledge, is the first dataset that considers the domain shift arising when transferring knowledge from synthetic to real video sequences. Our extensive experimental evaluation conducted on three publicly available benchmarks and on our new Mixamo→Kinetics dataset demonstrate the effectiveness of our approach, which outperforms the current state-of-the-art methods. Code is available at https://github.com/vturrisi/CO2A.
Victor G. T. da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, Elisa Ricci 0001
WACV1
2022 solo-learn: A Library of Self-supervised Methods for Visual Representation Learning
abstract
This paper presents solo-learn, a library of self-supervised methods for visual representation learning. Implemented in Python, using Pytorch and Pytorch lightning, the library fits both research and industry needs by featuring distributed training pipelines with mixed-precision, faster data loading via Nvidia DALI, online linear evaluation for better prototyping, and many additional training tricks. Our goal is to provide an easy-to-use library comprising a large amount of Self-supervised Learning (SSL) methods, that can be easily extended and fine-tuned by the community. solo-learn opens up avenues for exploiting large-budget SSL solutions on inexpensive smaller infrastructures and seeks to democratize SSL by making it accessible to all. The source code is available at https://github.com/vturrisi/solo-learn.
Victor G. T. da Costa, Enrico Fini, Moin Nabi, Nicu Sebe, Elisa Ricci 0001
J. Mach. Learn. Res.1
2022 Deep computer vision system for cocoa classification
abstract
Abstract Cocoa hybridisation generates new varieties which are resistant to several plant diseases, but has individual chemical characteristics that affect chocolate production. Image analysis is a useful method for visual discrimination of cocoa beans, while deep learning (DL) has emerged as thede factotechnique for image processing . However, these algorithms require a large amount of data and careful tuning of hyperparameters. Since it is necessary to acquire a large number of images to encompass the wide range of agricultural products, in this paper, we compare a Deep Computer Vision System (DCVS) and a traditional Computer Vision System (CVS) to classify cocoa beans into different varieties. For DCVS, we used a Resnet18 and Resnet50 as backbone, while for CVS, we experimented traditional machine learning algorithms, Support Vector Machine (SVM), and Random Forest (RF). All the algorithms were selected since they provide good classification performance and their potential application for food classification A dataset with 1,239 samples was used to evaluate both systems. The best accuracy was 96.82% for DCVS (ResNet 18), compared to 85.71% obtained by the CVS using SVM. The essential handcrafted features were reported and discussed regarding their influence on cocoa bean classification. Class Activation Maps was applied to DCVS’s predictions, providing a meaningful visualisation of the most important regions of the images in the model.
Jessica Fernandes Lopes, Victor G. T. da Costa, Douglas Fernandes Barbin, Luis Jam Pier Cruz-Tirado, Baeten Vincent, Sylvio Barbon Junior
Multim. Tools Appl.2
2020 Evaluating the Four-Way Performance Trade-Off for Data Stream Classification in Edge Computing
abstract
Edge computing (EC) is a promising technology capable of bridging the gap between Cloud computing services and the demands of emerging technologies such as the Internet of Things (IoT). Most EC-based solutions, from wearable devices to smart cities architectures, benefit from Machine Learning (ML) methods to perform various tasks, such as classification. In these cases, ML solutions need to deal efficiently with a huge amount of data, while balancing predictive performance, memory and time costs, and energy consumption. The fact that these data usually come in the form of a continuous and evolving data stream makes the scenario even more challenging. Many algorithms have been proposed to cope with data stream classification, e.g., Very Fast Decision Tree (VFDT) and Strict VFDT (SVFDT). Recently, Online Local Boosting (OLBoost) has also been introduced to improve predictive performance without modifying the underlying structure of the decision tree produced by these algorithms. In this work, we compared the four-way relationship among time efficiency, energy consumption, predictive performance, and memory costs, tuning the hyperparameters of VFDT and the two versions of SVFDT with and without OLBoost. Experiments over 6 benchmark datasets using an EC device revealed that VFDT and SVFDT-I were the most energy-friendly algorithms, with SVFDT-I also significantly reducing memory consumption. OLBoost, as expected, improved the predictive performance, but caused a deterioration in memory and energy consumption.
Jessica Fernandes Lopes, Everton Jose Santana, Victor G. T. da Costa, Bruno Bogaz Zarpelão, Sylvio Barbon Junior
IEEE Trans. Netw. Serv. Manag.3
2019 Evaluating the Four-Way Performance Trade-Off for Stream Classification
Victor G. T. da Costa, Everton Jose Santana, Jessica Fernandes Lopes, Sylvio Barbon Junior
GPC1
2018 U-Healthcare System for Pre-Diagnosis of Parkinson's Disease from Voice Signal
abstract
With the ageing and growth of the population, some chronic diseases, such as Parkinson's disease (PD), urge the society to a health-conscious looking for better health system designs. Some recent research endeavour has been supported by solutions grounded in ubiquitous healthcare (u-Health) coupling telemedicine, context awareness and decision support capabilities. In this work, we propose a u-healthcare system to pre-diagnose PD based on the speech signal of people under voice call. The speech stream is sampled as well as processed to support the pre-diagnose using machine learning (ML). Experiments were conducted over a PD voice dataset composed of 40 individuals by using five different ML algorithms. Based on a linear Support Vector Machine (SVM) model, a false negative rate of 10% was obtained when classifying the locution of number "three".
Sylvio Barbon Junior, Victor G. T. da Costa, Shi-Huang Chen, Rodrigo Capobianco Guido
ISM2
2018 Strict Very Fast Decision Tree: A memory conservative algorithm for data stream mining
Victor G. T. da Costa, André C. P. L. F. de Carvalho, Sylvio Barbon Junior
Pattern Recognit. Lett.1
2017 Detecting mobile botnets through machine learning and system calls analysis
abstract
Botnets have been a serious threat to the Internet security. With the constant sophistication and the resilience of them, a new trend has emerged, shifting botnets from the traditional desktop to the mobile environment. As in the desktop domain, detecting mobile botnets is essential to minimize the threat that they impose. Along the diverse set of strategies applied to detect these botnets, the ones that show the best and most generalized results involve discovering patterns in their anomalous behavior. In the mobile botnet field, one way to detect these patterns is by analyzing the operation parameters of this kind of applications. In this paper, we present an anomaly-based and host-based approach to detect mobile botnets. The proposed approach uses machine learning algorithms to identify anomalous behaviors in statistical features extracted from system calls. Using a self-generated dataset containing 13 families of mobile botnets and legitimate applications, we were able to test the performance of our approach in a close-to-reality scenario. The proposed approach achieved great results, including low false positive rates and high true detection rates.
Victor G. T. da Costa, Sylvio Barbon Junior, Rodrigo Sanches Miani, Joel J. P. C. Rodrigues, Bruno Bogaz Zarpelão
ICC1