VLDB 2026 Research / reviewers in the wild / expert
Moacir Ponti
dblp:69/5836 · also Moacir A. Ponti, Moacir Antonelli Ponti, Moacir P. Ponti Jr.
· DBLP profile ↗
29ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0003-2059-9463ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Role of Cyclopean-Eye in Stereo Vision
Sherlon Almeida da Silva, Davi Geiger, Luiz Velho 0001, Moacir Ponti |
CIARP | 4 |
| 2023 | ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion
Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti |
INTERSPEECH | 7 |
| 2023 | Scene designer: compositional sketch-based image retrieval with contrastive learning and an auxiliary synthesis task
Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir Ponti |
Multim. Tools Appl. | 4 |
| 2022 | YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneabstractYourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual training. We achieved state-of-the-art (SOTA) results in zero-shot multi-speaker TTS and results comparable to SOTA in zero-shot voice conversion on the VCTK dataset. Additionally, our approach achieves promising results in a target language with a single-speaker dataset, opening possibilities for zero-shot multi-speaker TTS and zero-shot voice conversion systems in low-resource languages. Finally, it is possible to fine-tune the YourTTS model with less than 1 minute of speech and achieve state-of-the-art results in voice similarity and with reasonable quality. This is important to allow synthesis for speakers with a very different voice or recording characteristics from those seen during training. Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Cândido Jr., Eren Gölge, Moacir Ponti |
ICML | 6 |
| 2022 | A classification and quantification approach to generate features in soundscape ecology using neural networks
Fábio Felix Dias, Moacir Ponti, Rosane Minghim |
Neural Comput. Appl. | 2 |
| 2022 | Robust image features for classification and zero-shot tasks by merging visual and semantic attributes
Damares Oliveira de Resende, Moacir Ponti |
Neural Comput. Appl. | 2 |
| 2021 | Transfer Learning and Data Augmentation Techniques to the COVID-19 Identification Tasks in ComParE 2021abstractIn this work, we propose several techniques to address data scarceness in ComParE 2021 COVID-19 identification tasks for the application of deep models such as Convolutional Neural Networks.Data is initially preprocessed into spectrogram or MFCC-gram formats.After preprocessing, we combine three different data augmentation techniques to be applied in model training.Then we employ transfer learning techniques from pretrained audio neural networks.Those techniques are applied to several distinct neural architectures.For COVID-19 identification in speech segments, we obtained competitive results.On the other hand, in the identification task based on cough data, we succeeded in producing a noticeable improvement on existing baselines, reaching 75.9% unweighted average recall (UAR). Edresson Casanova, Arnaldo Cândido Jr., Ricardo Corso Fernandes Junior, Marcelo Finger, Lucas Gris, Moacir Ponti, Daniel Peixoto Pinto da Silva |
Interspeech | 6 |
| 2021 | SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech ModelabstractIn this paper, we propose SC-GlowTTS: an efficient zero-shot multi-speaker text-to-speech model that improves similarity for speakers unseen during training. We propose a speaker-conditional architecture that explores a flow-based decoder that works in a zero-shot scenario. As text encoders, we explore a dilated residual convolutional-based encoder, gated convolutional-based encoder, and transformer-based encoder. Additionally, we have shown that adjusting a GAN-based vocoder for the spectrograms predicted by the TTS model on the training dataset can significantly improve the similarity and speech quality for new speakers. Our model converges using only 11 speakers, reaching state-of-the-art results for similarity with new speakers, as well as high speech quality. Edresson Casanova, Christopher Shulby, Eren Gölge, Nicolas M. Müller, Frederico Santos de Oliveira, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti |
Interspeech | 9 |
| 2020 | Sketchformer: Transformer-Based Representation for Sketched StructureabstractSketchformer is a novel transformer-based representation for encoding free-hand sketches input in a vector form, i.e. as a sequence of strokes. Sketchformer effectively addresses multiple tasks: sketch classification, sketch based image retrieval (SBIR), and the reconstruction and interpolation of sketches. We report several variants exploring continuous and tokenized input representations, and contrast their performance. Our learned embedding, driven by a dictionary learning tokenization scheme, yields state of the art performance in classification and image retrieval tasks, when compared against baseline representations driven by LSTM sequence to sequence architectures: SketchRNN and derivatives. We show that sketch reconstruction and interpolation are improved significantly by the Sketchformer embedding for complex sketches with longer stroke sequences. Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir Ponti |
CVPR | 4 |
| 2020 | Learning image features with fewer labels using a semi-supervised deep convolutional networkabstractLearning feature embeddings for pattern recognition is a relevant task for many applications. Deep learning methods such as convolutional neural networks can be employed for this assignment with different training strategies: leveraging pre-trained models as baselines; training from scratch with the target dataset; or fine-tuning from the pre-trained model. Although there are separate systems used for learning features from labelled and unlabelled data, there are few models combining all available information. Therefore, in this paper, we present a novel semi-supervised deep network training strategy that comprises a convolutional network and an autoencoder using a joint classification and reconstruction loss function. We show our network improves the learned feature embedding when including the unlabelled data in the training process. The results using the feature embedding obtained by our network achieve better classification accuracy when compared with competing methods, as well as offering good generalisation in the context of transfer learning. Furthermore, the proposed network ensemble and loss function is highly extensible and applicable in many recognition tasks. Fernando Pereira dos Santos, Cemre Zor, Josef Kittler, Moacir Ponti |
Neural Networks | 4 |
| 2019 | Homogeneity Index as Stopping Criterion for Anisotropic Diffusion Filter
Fernando Pereira dos Santos, Moacir Ponti |
CAIP (2) | 2 |
| 2019 | Combining clustering and active learning for the detection and learning of new image classes
Luiz F. S. Coletta, Moacir Ponti, Eduardo R. Hruschka, Ayan Acharya, Joydeep Ghosh |
Neurocomputing | 2 |
| 2019 | Generalization of feature embeddings transferred from different video anomaly detection domainsabstractDetecting anomalous activity in video surveillance often suffers from limited availability of training data. Transfer learning may close this gap, allowing to use existing annotated data from some source domain. However, analyzing the source feature space in terms of its potential for transfer of learning to another context is still to be investigated. This paper reports a study on video anomaly detection, focusing on the analysis of feature embeddings of pre-trained CNNs with the use of novel cross-domain generalization measures that allow to study how source features generalize for different target video domains. This generalization analysis represents not only a theoretical approach, can be useful in practice as a path to understand which datasets allow better transfer of knowledge. Our results confirm this, achieving better anomaly detectors for video frames and allowing analysis of transfer learning's positive and negative aspects. Fernando Pereira dos Santos, Leo Sampaio Ferraz Ribeiro, Moacir Ponti |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Deep Manifold Alignment for Mid-Grain Sketch Based Image Retrieval
Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse |
ACCV (3) | 3 |
| 2018 | Sketching out the details: Sketch-based image retrieval using convolutional neural networks with multi-stage regressionabstractWe propose and evaluate several deep network architectures for measuring the similarity between sketches and photographs, within the context of the sketch based image retrieval (SBIR) task.We study the ability of our networks to generalize across diverse object categories from limited training data, and explore in detail strategies for weight sharing, pre-processing, data augmentation and dimensionality reduction.In addition to a detailed comparative study of network configurations, we contribute by describing a hybrid multi-stage training network that exploits both contrastive and triplet networks to exceed state of the art performance on several SBIR benchmarks by a significant margin.Datasets and models are available at www.cvssp.org. Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse |
Comput. Graph. | 3 |
| 2017 | Deep Convolutional Neural Networks and Noisy Images
Tiago S. Nazaré, Gabriel B. Paranhos da Costa, Welinton A. Contato, Moacir Ponti |
CIARP | 4 |
| 2017 | Optical-flow features empirical mode decomposition for motion anomaly detectionabstractIn video data analysis of dynamic scenes, temporal characteristics of moving objects play an important role in decision-making. However, the temporal consistency of typical features used for video interpretation is low due to the overlap of the spectra of informative video signal component and the stochastic variations perturbing it. We propose a novel method for object motion anomaly detection in video designed to overcome this problem. It is based on empirical mode decomposition. We show in experiments on a benchmarking dataset that the deterministic component of an optical flow feature obtained using the proposed method is able to isolate the periodic behaviour of the motion from the stochastic values, facilitating much simpler analysis of the motion patterns and achieving impressive anomaly detection performance. Moacir Ponti, Tiago S. Nazaré, Josef Kittler |
ICASSP | 1 |
| 2017 | Compact descriptors for sketch-based image retrieval using a triplet loss convolutional neural networkabstractWe present an efficient representation for sketch based image retrieval (SBIR) derived from a triplet loss convolutional neural network (CNN). We treat SBIR as a cross-domain modelling problem, in which a depiction invariant embedding of sketch and photo data is learned by regression over a siamese CNN architecture with half-shared weights and modified triplet loss function. Uniquely, we demonstrate the ability of our learned image descriptor to generalise beyond the categories of object present in our training data, forming a basis for general cross-category SBIR. We explore appropriate strategies for training, and for deriving a compact image descriptor from the learned representation suitable for indexing data on resource constrained e. g. mobile devices. We show the learned descriptors to outperform state of the art SBIR on the defacto standard Flickr15k dataset using a significantly more compact (56 bits per image, i. e. ≈ 105KB total) search index than previous methods. Datasets and models are available from the CVSSP datasets server at www.cvssp.org. Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse |
Comput. Vis. Image Underst. | 3 |
| 2017 | An incremental linear-time learning algorithm for the Optimum-Path Forest classifierabstractWe present a classification method with incremental capabilities based on the Optimum-Path Forest classifier (OPF). The OPF considers instances as nodes of a fully-connected training graph, arc weights represent distances between two feature vectors. Our algorithm includes new instances in an OPF in linear-time, while keeping similar accuracies when compared with the original quadratic-time model. Moacir Ponti, Mateus Riva |
Inf. Process. Lett. | 1 |
| 2017 | A decision cognizant Kullback-Leibler divergenceabstractIn decision making systems involving multiple classifiers there is the need to assess classifier (in)congruence, that is to gauge the degree of agreement between their outputs. A commonly used measure for this purpose is the Kullback–Leibler (KL) divergence. We propose a variant of the KL divergence, named decision cognizant Kullback–Leibler divergence (DC-KL), to reduce the contribution of the minority classes, which obscure the true degree of classifier incongruence. We investigate the properties of the novel divergence measure analytically and by simulation studies. The proposed measure is demonstrated to be more robust to minority class clutter. Its sensitivity to estimation noise is also shown to be considerably lower than that of the classical KL divergence. These properties render the DC-KL divergence a much better statistic for discriminating between classifier congruence and incongruence in pattern recognition systems. Moacir Ponti, Josef Kittler, Mateus Riva, Teófilo Emídio de Campos, Cemre Zor |
Pattern Recognit. | 1 |
| 2016 | One-Class to Multi-Class Model Update Using the Class-Incremental Optimum-Path Forest ClassifierabstractIncremental learning capabilities of classifiers is a relevant topic, specially when dealing with scenarios such as data stream mining, concept drift and active learning. We investigate the capabilities of an incremental version of the Optimum-Path Forest classifier (OPF-CI) in the context of learning new classes and compare its behavior against Support Vector Machines and k Nearest Neighbours classifiers. The OPF-CI classifier is a parameter-free, graph-based approach to incremental training that runs in linear time with respect to the number of instances. Our results show OPF-CI keeps the running time low while producing an accuracy behavior similar to the other classifiers for increments of instances. Also, it is robust to variations on the order of the learned classes, demonstrating the applicability of the method. Mateus Riva, Moacir Ponti, Teófilo Emídio de Campos |
ECAI | 2 |
| 2016 | Image quantization as a dimensionality reduction procedure in color and texture feature extractionabstractThe image-based visual recognition pipeline includes a step that converts color images into images with a single channel, obtaining a color-quantized image that can be processed by feature extraction methods. In this paper we explore this step in order to produce compact features that can be used in retrieval and classification systems. We show that different quantization methods produce very different results in terms of accuracy. While compared with more complex methods, this procedure allows the feature extraction in order to achieve a significant dimensionality reduction, while preserving or improving system accuracy. The results indicate that quantization simplify images before feature extraction and dimensionality reduction, producing more compact vectors and reducing system complexity. Moacir Ponti, Tiago S. Nazaré, Gabriela S. Thumé |
Neurocomputing | 1 |
| 2015 | Color description of low resolution images using fast bitwise quantization and border-interior classificationabstractImage classification often require preprocessing and feature extraction steps that are directly related to the accuracy and speed of the whole task. In this paper we investigate color features extracted from low resolution images, assessing the influence of the resolution settings on the final classification accuracy. We propose a border-interior classification extractor with a logarithmic distance function in order to maintain the discrimination capability in different resolutions. Our study shows that the overall computational effort can be reduced in 98%. Besides, a fast bitwise quantization is performed for its efficiency on converting RGB images to one channel images. The contributions can benefit many applications, when dealing with a large number of images or in scenarios with limited network bandwidth and concerns with power consumption. Moacir Ponti, Camila T. Picon |
ICASSP | 1 |
| 2013 | Green Coverage Detection on Sub-orbital Plantation Images Using Anomaly Detection
Gabriel B. Paranhos da Costa, Moacir Ponti |
CIARP (2) | 2 |
| 2013 | Hand-Raising Gesture Detection with Lienhart-Maydt Method in Videoconference and Distance Learning
Tiago S. Nazaré, Moacir Ponti |
CIARP (2) | 2 |
| 2013 | Segmentation of Low-Cost Remote Sensing Images Combining Vegetation Indices and Mean ShiftabstractThe development of low-cost remote sensing systems is important in small agriculture business, particularly in developing countries, to allow feasible use of images to gather information. However, images obtained through such systems with uncalibrated cameras have often illumination variations, shadows, and other elements that can hinder the analysis by image processing techniques. This letter investigates the combination of vegetation indices (color index of vegetation extraction, visual vegetation index, and excess green) and the mean-shift algorithm, based on the local density estimation in the color space on images acquired by a low-cost system. The objective is to detect green coverage, gaps, and degraded areas. The results showed that combining local density estimation and vegetation indices improves the segmentation accuracy when compared with the competing methods. It deals well with images in different conditions and with regions of imbalanced sizes, confirming the practical application of the low-cost system. Moacir Ponti |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2012 | Improving restoration of microscopy images using iterative prototypes and a sequence of support constraintsabstractImages obtained by wide-field microscopy are specially corrupted by an out-of-focus blur on the axial direction. Constrained restoration algorithms are able to restore some of the frequencies lost on the acquisition process by using prior knowledge. We report an improved restoration algorithm using prototype images produced by the iterative Richardson-Lucy algorithm and a sequence of finite support constraint sets. The finite support shrinks as the nonlinear restoration is performed. Results showed an improved restoration on images with dense background, achieving a larger passband and restoration on axial direction. It encourages the development of new restoration algorithms based on the projection onto sequence of sets. Moacir Ponti |
ICIP | 1 |
| 2011 | A Markov Random Field Model for Combining Optimum-Path Forest Classifiers Using Decision Graphs and Game Strategy Approach
Moacir Ponti, João Paulo Papa, Alexandre L. M. Levada |
CIARP | 1 |
| 2011 | Feature selection through gravitational search algorithmabstractIn this paper we deal with the problem of feature selection by introducing a new approach based on Gravitational Search Algorithm (GSA). The proposed algorithm combines the optimization behavior of GSA together with the speed of Optimum-Path Forest (OPF) classifier in order to provide a fast and accurate framework for feature selection. Experiments on datasets obtained from a wide range of applications, such as vowel recognition, image classification and fraud detection in power distribution systems are conducted in order to asses the robustness of the proposed technique against Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA) and a Particle Swarm Optimization (PSO)-based algorithm for feature selection. João Paulo Papa, Andre Pagnin, Silvana Artioli Schellini, André Augusto Spadotto, Rodrigo Capobianco Guido, Moacir Ponti, Giovani Chiachia, Alexandre X. Falcão |
ICASSP | 6 |