Moacir Ponti

dblp:69/5836 · also Moacir A. Ponti, Moacir Antonelli Ponti, Moacir P. Ponti Jr. · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0003-2059-9463ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 The Role of Cyclopean-Eye in Stereo Vision
Sherlon Almeida da Silva, Davi Geiger, Luiz Velho 0001, Moacir Ponti
CIARP4
2023 ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion
Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti
INTERSPEECH7
2023 Scene designer: compositional sketch-based image retrieval with contrastive learning and an auxiliary synthesis task
Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir Ponti
Multim. Tools Appl.4
2022 YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for Everyone
abstract
YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual training. We achieved state-of-the-art (SOTA) results in zero-shot multi-speaker TTS and results comparable to SOTA in zero-shot voice conversion on the VCTK dataset. Additionally, our approach achieves promising results in a target language with a single-speaker dataset, opening possibilities for zero-shot multi-speaker TTS and zero-shot voice conversion systems in low-resource languages. Finally, it is possible to fine-tune the YourTTS model with less than 1 minute of speech and achieve state-of-the-art results in voice similarity and with reasonable quality. This is important to allow synthesis for speakers with a very different voice or recording characteristics from those seen during training.
Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Cândido Jr., Eren Gölge, Moacir Ponti
ICML6
2022 A classification and quantification approach to generate features in soundscape ecology using neural networks
Fábio Felix Dias, Moacir Ponti, Rosane Minghim
Neural Comput. Appl.2
2022 Robust image features for classification and zero-shot tasks by merging visual and semantic attributes
Damares Oliveira de Resende, Moacir Ponti
Neural Comput. Appl.2
2021 Transfer Learning and Data Augmentation Techniques to the COVID-19 Identification Tasks in ComParE 2021
abstract
In this work, we propose several techniques to address data scarceness in ComParE 2021 COVID-19 identification tasks for the application of deep models such as Convolutional Neural Networks.Data is initially preprocessed into spectrogram or MFCC-gram formats.After preprocessing, we combine three different data augmentation techniques to be applied in model training.Then we employ transfer learning techniques from pretrained audio neural networks.Those techniques are applied to several distinct neural architectures.For COVID-19 identification in speech segments, we obtained competitive results.On the other hand, in the identification task based on cough data, we succeeded in producing a noticeable improvement on existing baselines, reaching 75.9% unweighted average recall (UAR).
Edresson Casanova, Arnaldo Cândido Jr., Ricardo Corso Fernandes Junior, Marcelo Finger, Lucas Gris, Moacir Ponti, Daniel Peixoto Pinto da Silva
Interspeech6
2021 SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech Model
abstract
In this paper, we propose SC-GlowTTS: an efficient zero-shot multi-speaker text-to-speech model that improves similarity for speakers unseen during training. We propose a speaker-conditional architecture that explores a flow-based decoder that works in a zero-shot scenario. As text encoders, we explore a dilated residual convolutional-based encoder, gated convolutional-based encoder, and transformer-based encoder. Additionally, we have shown that adjusting a GAN-based vocoder for the spectrograms predicted by the TTS model on the training dataset can significantly improve the similarity and speech quality for new speakers. Our model converges using only 11 speakers, reaching state-of-the-art results for similarity with new speakers, as well as high speech quality.
Edresson Casanova, Christopher Shulby, Eren Gölge, Nicolas M. Müller, Frederico Santos de Oliveira, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti
Interspeech9
2020 Sketchformer: Transformer-Based Representation for Sketched Structure
abstract
Sketchformer is a novel transformer-based representation for encoding free-hand sketches input in a vector form, i.e. as a sequence of strokes. Sketchformer effectively addresses multiple tasks: sketch classification, sketch based image retrieval (SBIR), and the reconstruction and interpolation of sketches. We report several variants exploring continuous and tokenized input representations, and contrast their performance. Our learned embedding, driven by a dictionary learning tokenization scheme, yields state of the art performance in classification and image retrieval tasks, when compared against baseline representations driven by LSTM sequence to sequence architectures: SketchRNN and derivatives. We show that sketch reconstruction and interpolation are improved significantly by the Sketchformer embedding for complex sketches with longer stroke sequences.
Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir Ponti
CVPR4
2020 Learning image features with fewer labels using a semi-supervised deep convolutional network
abstract
Learning feature embeddings for pattern recognition is a relevant task for many applications. Deep learning methods such as convolutional neural networks can be employed for this assignment with different training strategies: leveraging pre-trained models as baselines; training from scratch with the target dataset; or fine-tuning from the pre-trained model. Although there are separate systems used for learning features from labelled and unlabelled data, there are few models combining all available information. Therefore, in this paper, we present a novel semi-supervised deep network training strategy that comprises a convolutional network and an autoencoder using a joint classification and reconstruction loss function. We show our network improves the learned feature embedding when including the unlabelled data in the training process. The results using the feature embedding obtained by our network achieve better classification accuracy when compared with competing methods, as well as offering good generalisation in the context of transfer learning. Furthermore, the proposed network ensemble and loss function is highly extensible and applicable in many recognition tasks.
Fernando Pereira dos Santos, Cemre Zor, Josef Kittler, Moacir Ponti
Neural Networks4
2019 Homogeneity Index as Stopping Criterion for Anisotropic Diffusion Filter
Fernando Pereira dos Santos, Moacir Ponti
CAIP (2)2
2019 Combining clustering and active learning for the detection and learning of new image classes
Luiz F. S. Coletta, Moacir Ponti, Eduardo R. Hruschka, Ayan Acharya, Joydeep Ghosh
Neurocomputing2
2019 Generalization of feature embeddings transferred from different video anomaly detection domains
abstract
Detecting anomalous activity in video surveillance often suffers from limited availability of training data. Transfer learning may close this gap, allowing to use existing annotated data from some source domain. However, analyzing the source feature space in terms of its potential for transfer of learning to another context is still to be investigated. This paper reports a study on video anomaly detection, focusing on the analysis of feature embeddings of pre-trained CNNs with the use of novel cross-domain generalization measures that allow to study how source features generalize for different target video domains. This generalization analysis represents not only a theoretical approach, can be useful in practice as a path to understand which datasets allow better transfer of knowledge. Our results confirm this, achieving better anomaly detectors for video frames and allowing analysis of transfer learning's positive and negative aspects.
Fernando Pereira dos Santos, Leo Sampaio Ferraz Ribeiro, Moacir Ponti
J. Vis. Commun. Image Represent.3
2018 Deep Manifold Alignment for Mid-Grain Sketch Based Image Retrieval
Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse
ACCV (3)3
2018 Sketching out the details: Sketch-based image retrieval using convolutional neural networks with multi-stage regression
abstract
We propose and evaluate several deep network architectures for measuring the similarity between sketches and photographs, within the context of the sketch based image retrieval (SBIR) task.We study the ability of our networks to generalize across diverse object categories from limited training data, and explore in detail strategies for weight sharing, pre-processing, data augmentation and dimensionality reduction.In addition to a detailed comparative study of network configurations, we contribute by describing a hybrid multi-stage training network that exploits both contrastive and triplet networks to exceed state of the art performance on several SBIR benchmarks by a significant margin.Datasets and models are available at www.cvssp.org.
Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse
Comput. Graph.3
2017 Deep Convolutional Neural Networks and Noisy Images
Tiago S. Nazaré, Gabriel B. Paranhos da Costa, Welinton A. Contato, Moacir Ponti
CIARP4
2017 Optical-flow features empirical mode decomposition for motion anomaly detection
abstract
In video data analysis of dynamic scenes, temporal characteristics of moving objects play an important role in decision-making. However, the temporal consistency of typical features used for video interpretation is low due to the overlap of the spectra of informative video signal component and the stochastic variations perturbing it. We propose a novel method for object motion anomaly detection in video designed to overcome this problem. It is based on empirical mode decomposition. We show in experiments on a benchmarking dataset that the deterministic component of an optical flow feature obtained using the proposed method is able to isolate the periodic behaviour of the motion from the stochastic values, facilitating much simpler analysis of the motion patterns and achieving impressive anomaly detection performance.
Moacir Ponti, Tiago S. Nazaré, Josef Kittler
ICASSP1
2017 Compact descriptors for sketch-based image retrieval using a triplet loss convolutional neural network
abstract
We present an efficient representation for sketch based image retrieval (SBIR) derived from a triplet loss convolutional neural network (CNN). We treat SBIR as a cross-domain modelling problem, in which a depiction invariant embedding of sketch and photo data is learned by regression over a siamese CNN architecture with half-shared weights and modified triplet loss function. Uniquely, we demonstrate the ability of our learned image descriptor to generalise beyond the categories of object present in our training data, forming a basis for general cross-category SBIR. We explore appropriate strategies for training, and for deriving a compact image descriptor from the learned representation suitable for indexing data on resource constrained e. g. mobile devices. We show the learned descriptors to outperform state of the art SBIR on the defacto standard Flickr15k dataset using a significantly more compact (56 bits per image, i. e. ≈ 105KB total) search index than previous methods. Datasets and models are available from the CVSSP datasets server at www.cvssp.org.
Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse
Comput. Vis. Image Underst.3
2017 An incremental linear-time learning algorithm for the Optimum-Path Forest classifier
abstract
We present a classification method with incremental capabilities based on the Optimum-Path Forest classifier (OPF). The OPF considers instances as nodes of a fully-connected training graph, arc weights represent distances between two feature vectors. Our algorithm includes new instances in an OPF in linear-time, while keeping similar accuracies when compared with the original quadratic-time model.
Moacir Ponti, Mateus Riva
Inf. Process. Lett.1
2017 A decision cognizant Kullback-Leibler divergence
abstract
In decision making systems involving multiple classifiers there is the need to assess classifier (in)congruence, that is to gauge the degree of agreement between their outputs. A commonly used measure for this purpose is the Kullback–Leibler (KL) divergence. We propose a variant of the KL divergence, named decision cognizant Kullback–Leibler divergence (DC-KL), to reduce the contribution of the minority classes, which obscure the true degree of classifier incongruence. We investigate the properties of the novel divergence measure analytically and by simulation studies. The proposed measure is demonstrated to be more robust to minority class clutter. Its sensitivity to estimation noise is also shown to be considerably lower than that of the classical KL divergence. These properties render the DC-KL divergence a much better statistic for discriminating between classifier congruence and incongruence in pattern recognition systems.
Moacir Ponti, Josef Kittler, Mateus Riva, Teófilo Emídio de Campos, Cemre Zor
Pattern Recognit.1
2016 One-Class to Multi-Class Model Update Using the Class-Incremental Optimum-Path Forest Classifier
abstract
Incremental learning capabilities of classifiers is a relevant topic, specially when dealing with scenarios such as data stream mining, concept drift and active learning. We investigate the capabilities of an incremental version of the Optimum-Path Forest classifier (OPF-CI) in the context of learning new classes and compare its behavior against Support Vector Machines and k Nearest Neighbours classifiers. The OPF-CI classifier is a parameter-free, graph-based approach to incremental training that runs in linear time with respect to the number of instances. Our results show OPF-CI keeps the running time low while producing an accuracy behavior similar to the other classifiers for increments of instances. Also, it is robust to variations on the order of the learned classes, demonstrating the applicability of the method.
Mateus Riva, Moacir Ponti, Teófilo Emídio de Campos
ECAI2
2016 Image quantization as a dimensionality reduction procedure in color and texture feature extraction
abstract
The image-based visual recognition pipeline includes a step that converts color images into images with a single channel, obtaining a color-quantized image that can be processed by feature extraction methods. In this paper we explore this step in order to produce compact features that can be used in retrieval and classification systems. We show that different quantization methods produce very different results in terms of accuracy. While compared with more complex methods, this procedure allows the feature extraction in order to achieve a significant dimensionality reduction, while preserving or improving system accuracy. The results indicate that quantization simplify images before feature extraction and dimensionality reduction, producing more compact vectors and reducing system complexity.
Moacir Ponti, Tiago S. Nazaré, Gabriela S. Thumé
Neurocomputing1
2015 Color description of low resolution images using fast bitwise quantization and border-interior classification
abstract
Image classification often require preprocessing and feature extraction steps that are directly related to the accuracy and speed of the whole task. In this paper we investigate color features extracted from low resolution images, assessing the influence of the resolution settings on the final classification accuracy. We propose a border-interior classification extractor with a logarithmic distance function in order to maintain the discrimination capability in different resolutions. Our study shows that the overall computational effort can be reduced in 98%. Besides, a fast bitwise quantization is performed for its efficiency on converting RGB images to one channel images. The contributions can benefit many applications, when dealing with a large number of images or in scenarios with limited network bandwidth and concerns with power consumption.
Moacir Ponti, Camila T. Picon
ICASSP1
2013 Green Coverage Detection on Sub-orbital Plantation Images Using Anomaly Detection
Gabriel B. Paranhos da Costa, Moacir Ponti
CIARP (2)2
2013 Hand-Raising Gesture Detection with Lienhart-Maydt Method in Videoconference and Distance Learning
Tiago S. Nazaré, Moacir Ponti
CIARP (2)2
2013 Segmentation of Low-Cost Remote Sensing Images Combining Vegetation Indices and Mean Shift
abstract
The development of low-cost remote sensing systems is important in small agriculture business, particularly in developing countries, to allow feasible use of images to gather information. However, images obtained through such systems with uncalibrated cameras have often illumination variations, shadows, and other elements that can hinder the analysis by image processing techniques. This letter investigates the combination of vegetation indices (color index of vegetation extraction, visual vegetation index, and excess green) and the mean-shift algorithm, based on the local density estimation in the color space on images acquired by a low-cost system. The objective is to detect green coverage, gaps, and degraded areas. The results showed that combining local density estimation and vegetation indices improves the segmentation accuracy when compared with the competing methods. It deals well with images in different conditions and with regions of imbalanced sizes, confirming the practical application of the low-cost system.
Moacir Ponti
IEEE Geosci. Remote. Sens. Lett.1
2012 Improving restoration of microscopy images using iterative prototypes and a sequence of support constraints
abstract
Images obtained by wide-field microscopy are specially corrupted by an out-of-focus blur on the axial direction. Constrained restoration algorithms are able to restore some of the frequencies lost on the acquisition process by using prior knowledge. We report an improved restoration algorithm using prototype images produced by the iterative Richardson-Lucy algorithm and a sequence of finite support constraint sets. The finite support shrinks as the nonlinear restoration is performed. Results showed an improved restoration on images with dense background, achieving a larger passband and restoration on axial direction. It encourages the development of new restoration algorithms based on the projection onto sequence of sets.
Moacir Ponti
ICIP1
2011 A Markov Random Field Model for Combining Optimum-Path Forest Classifiers Using Decision Graphs and Game Strategy Approach
Moacir Ponti, João Paulo Papa, Alexandre L. M. Levada
CIARP1
2011 Feature selection through gravitational search algorithm
abstract
In this paper we deal with the problem of feature selection by introducing a new approach based on Gravitational Search Algorithm (GSA). The proposed algorithm combines the optimization behavior of GSA together with the speed of Optimum-Path Forest (OPF) classifier in order to provide a fast and accurate framework for feature selection. Experiments on datasets obtained from a wide range of applications, such as vowel recognition, image classification and fraud detection in power distribution systems are conducted in order to asses the robustness of the proposed technique against Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA) and a Particle Swarm Optimization (PSO)-based algorithm for feature selection.
João Paulo Papa, Andre Pagnin, Silvana Artioli Schellini, André Augusto Spadotto, Rodrigo Capobianco Guido, Moacir Ponti, Giovani Chiachia, Alexandre X. Falcão
ICASSP6