David Menotti

dblp:10/2834 · also David Menoti, David Menotti Gomes · DBLP profile ↗
← Back
62ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0003-2430-2030ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 10 since 2021Security and privacy · 6 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 CEZSAR: A Contrastive Embedding Method for Zero-Shot Action Recognition
Valter Estevam, Rayson Laroca, Hélio Pedrini, David Menotti
ICPR (13)4
2026 ICPR 2026 Competition on Low-Resolution License Plate Recognition
Rayson Laroca, Valfride Nascimento, Donggun Kim 0004, Sanghyeok Chung, Subin Bae, Uihwan Seo, Seungsang Oh, Chi M. Phung, Minh G. Vo, Xingsong Ye, Yongkun Du, Zhineng Chen, Sunhee Heo, Hyangwoo Lee, Kihyun Na, Khanh V. Vu Nguyen, Sang T. Pham, Duc N. N. Phung, Trong P. Le, Vy N. Vo Tran, David Menotti
ICPR (16)22
2025 Beyond a Single Perspective: Neural Fusion of Lévy-Generated Super-Resolution Images for Robust Face Recognition
abstract
Despite significant advancements in facial recognition technology, these systems struggle in real-world surveillance scenarios, especially when dealing with low-resolution images. Super-resolution techniques have emerged as a natural solution to this issue, primarily using generative models such as generative adversarial networks or diffusion models. Despite their potential, these methods face critical limitations, including mode collapse and difficulty in preserving identity-specific features. To address these challenges, we propose Lévy Super-Resolution (LSR), a novel super-resolution approach leveraging Lévy processes as a noise source in a modified diffusion model. This stochastic process enhances diversity and mitigates mode collapse, enabling the generation of diverse super-resolved samples from a single low-resolution input. These samples are fused using an ensemble of neural networks trained with a triplet loss, producing robust face descriptors with reduced bias and improved identity preservation. To the best of our knowledge, LSR is the first method to utilize Lévy processes for super-resolution generation. We validate our approach on real and synthetic low-resolution samples from three standard face recognition datasets, and we demonstrate that our approach can surpass both general and face-specific state-of-the-art super-resolution (SR) methods. Our code is publicly available at https://github.com/marcelowds/lsr.
Marcelo dos Santos, João C. Neves 0001, David Menotti
ICMLA3
2025 Improving Small Drone Detection Through Multi-Scale Processing and Data Augmentation
abstract
Detecting small drones, often indistinguishable from birds, is crucial for modern surveillance. This work introduces a drone detection methodology built upon the medium-sized YOLOv11 object detection model. To enhance its performance on small targets, we implemented a multi-scale approach in which the input image is processed both as a whole and in segmented parts, with subsequent prediction aggregation. We also utilized a copy-paste data augmentation technique to enrich the training dataset with diverse drone and bird examples. Finally, we implemented a post-processing technique that leverages frame-to-frame consistency to mitigate missed detections. The proposed approach attained a top-3 ranking in the 8th WOSDETC Drone-vs-Bird Detection Grand Challenge, held at the 2025 International Joint Conference on Neural Networks (IJCNN), showcasing its capability to detect drones in complex environments effectively.
Rayson Laroca, Marcelo dos Santos, David Menotti
IJCNN3
2025 Dense video captioning using unsupervised semantic information
Valter Estevam, Rayson Laroca, Hélio Pedrini, David Menotti
J. Vis. Commun. Image Represent.4
2025 ORCNet: A Context-Based Network to Simultaneously Segment the Ocular Region Components
abstract
Accurate extraction of the Region of Interest is critical for successful ocular region-based biometrics. In this direction, we propose a new context-based segmentation approach, entitled Ocular Region Context Network (ORCNet), introducing a specific loss function, i.e., the Punish Context Loss (PC-Loss). The PC-Loss punishes the segmentation losses of a network by using a percentage difference value between the ground truth and the segmented masks. We obtain the percentage difference by taking into account Biederman’s semantic relationship concepts, in which we use three contexts (semantic, spatial, and scale) to evaluate the relationships of the objects in an image. Our proposal achieved promising results in the evaluated scenarios—iris, sclera, and ALL (iris + sclera) segmentations—, outperforming the literature baseline techniques. The ORCNet with ResNet-152 outperforms the best baseline (EncNet with ResNet-152) on average by 2. $$27\%$$ , 28. $$26\%$$ and 6. $$43\%$$ in terms of F-Score, Error Rate and Intersection Over Union, respectively. We also provide (for research purposes) 3191 manually labeled masks for the MICHE-I database, as another contribution of our work.
Diego Rafael Lucio, Luiz Antonio Zanlorensi, Yandre M. G. Costa, David Menotti
Neural Process. Lett.4
2024 SDFR: Synthetic Data for Face Recognition Competition
abstract
Large-scale face recognition datasets are collected by crawling the Internet and without individuals' consent, raising legal, ethical, and privacy concerns. With the recent advances in generative models, recently several works proposed generating synthetic face recognition datasets to mitigate concerns in web-crawled face recognition datasets. This paper presents the summary of the Synthetic Data for Face Recognition (SDFR) Competition held in conjunction with the 18th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2024) and established to investigate the use of synthetic data for training face recognition models. The SDFR competition was split into two tasks, allowing participants to train face recognition systems using new synthetic datasets and/or existing ones. In the first task, the face recognition backbone was fixed and the dataset size was limited, while the second task provided almost complete freedom on the model backbone, the dataset, and the training pipeline. The submitted models were trained on existing and also new synthetic datasets and used clever methods to improve training with synthetic data. The submissions were evaluated and ranked on a diverse set of seven benchmarking datasets. The paper gives an overview of the submitted face recognition models and reports achieved performance compared to baseline models trained on real and synthetic datasets. Furthermore, the evaluation of submissions is extended to bias assessment across different demography groups. Lastly, an outlook on the current state of the research in training face recognition models using synthetic data is presented, and existing problems as well as potential future directions are also discussed.
Hatef Otroshi-Shahreza, Christophe Ecabert, Anjith George, Alexander Unnervik, Sébastien Marcel, Nicolò Di Domenico, Guido Borghi, Davide Maltoni, Fadi Boutros, Julia Vogel, Naser Damer, Ángela Sánchez-Pérez, Enrique Mas-Candela, Jorge Calvo-Zaragoza, Bernardo Biesseck, Pedro Vidal 0001, Roger Granada, David Menotti, Ivan DeAndres-Tame, Simone Maurizio La Cava, Sara Concas, Pietro Melzi, Ruben Tolosana, Rubén Vera-Rodríguez, Gianpaolo Perelli, Giulia Orrù, Gian Luca Marcialis, Julian Fierrez
FG18
2024 Watchlist Challenge: 3rd Open-set Face Detection and Identification
abstract
In the current landscape of biometrics and surveillance, the ability to accurately recognize faces in uncontrolled settings is paramount. The Watchlist Challenge addresses this critical need by focusing on face detection and open-set identification in real-world surveillance scenarios. This paper presents a comprehensive evaluation of participating algorithms, using the enhanced UnConstrained College Students (UCCS) dataset with new evaluation protocols. In total, four participants submitted four face detection and nine open-set face recognition systems. The evaluation demonstrates that while detection capabilities are generally robust, closed-set identification performance varies significantly, with models pre-trained on large-scale datasets showing superior performance. However, open-set scenarios require further improvement, especially at higher true positive identification rates, i.e., lower thresholds.
Furkan Kasim, Terrance E. Boult, Rensso Mora Colque, Bernardo Biesseck, Rafael O. Ribeiro, Jan Schlüter, Tomás Repák, Rafael Henrique Vareto, David Menotti, William Robson Schwartz, Manuel Günther
IJCB9
2024 Tell me what you see: A zero-shot action recognition method based on natural language descriptions
Valter Estevam, Rayson Laroca, Hélio Pedrini, David Menotti
Multim. Tools Appl.4
2023 Leveraging Model Fusion for Improved License Plate Recognition
Rayson Laroca, Luiz Antonio Zanlorensi, Valter Estevam, Rodrigo Minetto, David Menotti
CIARP5
2023 Do We Train on Test Data? The Impact of Near-Duplicates on License Plate Recognition
abstract
This work draws attention to the large fraction of near-duplicates in the training and test sets of datasets widely adopted in License Plate Recognition (LPR) research. These duplicates refer to images that, although different, show the same license plate. Our experiments, conducted on the two most popular datasets in the field, show a substantial decrease in recognition rate when six well-known models are trained and tested under fair splits, that is, in the absence of duplicates in the training and test sets. Moreover, in one of the datasets, the ranking of models changed considerably when they were trained and tested under duplicate-free splits. These findings suggest that such duplicates have significantly biased the evaluation and development of deep learning-based models for LPR. The list of near-duplicates we have found and proposals for fair splits are publicly available for further research at https://raysonlaroca.github.io/supp/lpr-train-on-test/.
Rayson Laroca, Valter Estevam, Alceu S. Britto Jr., Rodrigo Minetto, David Menotti
IJCNN5
2023 Super-resolution of license plate images using attention modules and sub-pixel convolution layers
Valfride Nascimento, Rayson Laroca, Jorge de A. Lambert, William Robson Schwartz, David Menotti
Comput. Graph.5
2023 Conference on graphics, patterns and images
Hugo Proença 0001, David Menotti, Afonso Paiva 0001, Gladimir V. G. Baranoski
Pattern Recognit. Lett.2
2023 Exploring Bias in Sclera Segmentation Models: A Group Evaluation Approach
abstract
Bias and fairness of biometric algorithms have been key topics of research in recent years, mainly due to the societal, legal and ethical implications of potentially unfair decisions made by automated decision-making models. A considerable amount of work has been done on this topic across different biometric modalities, aiming at better understanding the main sources of algorithmic bias or devising mitigation measures. In this work, we contribute to these efforts and present the first study investigating bias and fairness of sclera segmentation models. Although sclera segmentation techniques represent a key component of sclera-based biometric systems with a considerable impact on the overall recognition performance, the presence of different types of biases in sclera segmentation methods is still underexplored. To address this limitation, we describe the results of a group evaluation effort (involving seven research groups), organized to explore the performance of recent sclera segmentation models within a common experimental framework and study performance differences (and bias), originating from various demographic as well as environmental factors. Using five diverse datasets, we analyze seven independently developed sclera segmentation models in different experimental configurations. The results of our experiments suggest that there are significant differences in the overall segmentation performance across the seven models and that among the considered factors, ethnicity appears to be the biggest cause of bias. Additionally, we observe that training with representative and balanced data does not necessarily lead to less biased results. Finally, we find that in general there appears to be a negative correlation between the amount of bias observed (due to eye color, ethnicity and acquisition device) and the overall segmentation performance, suggesting that advances in the field of semantic segmentation may also help with mitigating bias.
Matej Vitek, Abhijit Das 0001, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Jalil Nourmohammadi-Khiarak, Mohsen Akbari Shahpar, Meysam Asgari-Chenaghlu, Farhang Jaryani, Juan E. Tapia, Andres Valenzuela, Caiyong Wang, Yunlong Wang 0003, Zhaofeng He 0001, Zhenan Sun, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Kiran B. Raja, Gourav Gupta, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, S. V. Aruna Kumar, B. S. Harish, Umapada Pal 0001, Peter Peer, Vitomir Struc
IEEE Trans. Inf. Forensics Secur.5
2022 OCFR 2022: Competition on Occluded Face Recognition from Synthetically Generated Structure-Aware Occlusions
abstract
This work summarizes the IJCB Occluded Face Recognition Competition 2022 (IJCB-OCFR-2022) embraced by the 2022 International Joint Conference on Biometrics (IJCB 2022). OCFR-2022 attracted a total of 3 participating teams, from academia. Eventually, six valid submissions were submitted and then evaluated by the organizers. The competition was held to address the challenge of face recognition in the presence of severe face occlusions. The participants were free to use any training data and the testing data was built by the organisers by synthetically occluding parts of the face images using a well-known dataset. The submitted solutions presented innovations and performed very competitively with the considered baseline. A major output of this competition is a challenging, realistic, and diverse, and publicly available occluded face recognition benchmark with well defined evaluation protocols.
Pedro C. Neto, Fadi Boutros, João Ribeiro Pinto, Naser Damer, Ana Filipa Sequeira, Jaime S. Cardoso 0001, Messaoud Bengherabi, Abderaouf Bousnat, Sana Boucheta, Nesrine Hebbadj, Mustafa Ekrem Erakin, Ugur Demir, Hazim Kemal Ekenel, Pedro Vidal 0001, David Menotti
IJCB15
2022 A Decidability-Based Loss Function
abstract
Computer vision problems often use deep learning models to extract features (or embeddings) from images. Moreover, the loss function used during training strongly influences the quality of the generated embeddings. In this work, a loss function based on the decidability index is proposed to improve the quality of embeddings for the verification task. Our proposal, the D-loss, yields faster convergence for training and is parameter-free compared to Triplet-based loss approaches. Moreover, it avoids some disadvantages such as the use of hard samples and tricky parameter tuning. The proposed approach is compared against the Softmax (cross-entropy), Triplets Soft-Hard, and the Multi Similarity losses in four different benchmarks: MNIST, Fashion-MNIST, CIFAR10, and CASIA-IrisV4. The achieved results show the efficacy of the proposal when compared to other popular metrics in the literature. The D-loss computation, besides being simple, non-parametric and easy to implement, favors both the inter-class and intra-class scenarios.1
Pedro Silva 0004, Gladston J. P. Moreira, Vander L. S. Freitas, Rodrigo Silva 0001, David Menotti, Eduardo José da S. Luz
IJCNN5
2022 Global Semantic Descriptors for Zero-Shot Action Recognition
abstract
The success of Zero-Shot Action Recognition (ZSAR) methods is intrinsically related to the nature of semantic side information used to transfer knowledge, although this aspect has not been primarily investigated in the literature. This work introduces a new ZSAR method based on the relationships of actions-objects and actions-descriptive sentences. We demonstrate that representing all object classes using descriptive sentences generates an accurate object-action affinity estimation when a paraphrase estimation method is used as an embedder. We also show how to estimate probabilities over the set of action classes based only on a set of sentences without hard human labeling. In our method, the probabilities from these two global classifiers (i.e., which use features computed over the entire video) are combined, producing an efficient transfer knowledge model for action classification. Our results are state-of-the-art in the Kinetics-400 dataset and are competitive on UCF-101 under the ZSAR evaluation. Our code is available athttps://github.com/valterlej/objsentzsar
Valter Estevam, Rayson Laroca, Hélio Pedrini, David Menotti
IEEE Signal Process. Lett.4
2021 A multimodal LIBRAS-UFOP Brazilian sign language dataset of minimal pairs using a microsoft Kinect sensor
Lourdes Ramirez Cerna, Edwin Jonathan Escobedo Cardenas, Dayse Garcia Miranda, David Menotti, Guillermo Cámara Chávez
Expert Syst. Appl.4
2021 Zero-shot action recognition in videos: A survey
Valter Estevam, Hélio Pedrini, David Menotti
Neurocomputing3
2020 SSBC 2020: Sclera Segmentation Benchmarking Competition in the Mobile Environment
abstract
The paper presents a summary of the 2020 Sclera Segmentation Benchmarking Competition (SSBC), the 7th in the series of group benchmarking efforts centred around the problem of sclera segmentation. Different from previous editions, the goal of SSBC 2020 was to evaluate the performance of sclera-segmentation models on images captured with mobile devices. The competition was used as a platform to assess the sensitivity of existing models to i) differences in mobile devices used for image capture and ii) changes in the ambient acquisition conditions. 26 research groups registered for SSBC 2020, out of which 13 took part in the final round and submitted a total of 16 segmentation models for scoring. These included a wide variety of deep-learning solutions as well as one approach based on standard image processing techniques. Experiments were conducted with three recent datasets. Most of the segmentation models achieved relatively consistent performance across images captured with different mobile devices (with slight differences across devices), but struggled most with low-quality images captured in challenging ambient conditions, i.e., in an indoor environment and with poor lighting.
Matej Vitek, Abhijit Das 0001, Yann Pourcenoux, Alexandre Missler, C. Paumier, Sumanta Das, Ishita De Ghosh, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Junxing Hu, Yong He 0009, Caiyong Wang, Yunlong Wang 0003, Zhenan Sun, Dailé Osorio Roig, Christian Rathgeb, Christoph Busch 0001, Juan E. Tapia, Andres Valenzuela, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, Sabari Nathan, R. Suganya 0001, Vineet Mehta, Abhinav Dhall, Kiran B. Raja, Gourav Gupta, Jalil Nourmohammadi-Khiarak, Mohsen Akbari-Shahper, Farhang Jaryani, Meysam Asgari-Chenaghlu, Ritesh Vyas, Sristi Dakshit, Peter Peer, Umapada Pal 0001, Vitomir Struc
IJCB10
2020 Unconstrained Periocular Recognition: Using Generative Deep Learning Frameworks for Attribute Normalization
abstract
Ocular biometric systems working in unconstrained environments usually face the problem of small within-class compactness caused by the multiple factors that jointly degrade the quality of the obtained data. In this work, we propose an attribute normalization strategy based on deep learning generative frameworks, that reduces the variability of the samples used in pairwise comparisons, without reducing their discriminability. The proposed method can be seen as a preprocessing step that contributes for data regularization and improves the recognition accuracy, being fully agnostic to the recognition strategy used. As proof of concept, we consider the “eyeglasses” and “gaze” factors, comparing the levels of performance of five different recognition methods with/without using the proposed normalization strategy. Also, we introduce a new dataset for unconstrained periocular recognition, composed of images acquired by mobile devices, particularly suited to perceive the impact of “wearing eyeglasses” in recognition effectiveness. Our experiments were performed in two different datasets, and support the usefulness of our attribute normalization scheme to improve the recognition performance.
Luiz Antonio Zanlorensi, Hugo Proença 0001, David Menotti
ICIP3
2020 Agriculture Multispectral Uav Image Registration Using Salient Features and Mutual Information
abstract
Multimodal image registration has been studied for a long time as a necessary pre-processing step to extract relevant information from the studied images. In this direction, agriculture remote sensing has evolved to use multispectral sensors and faces challenges since the application of classic solutions is not suitable. This paper preliminarily explores the benefits of applying Mutual Information (MI) based on SIFT points for image registration to agriculture remote sensing multi-spectral evaluated on a self-developed public database of images through a fixed-wing Unmanned Aerial Vehicle (UAV) equipped with a multispectral sensor operating within parameters that would apply to crops inspection in real life. Our preliminary results have shown a marginal improvement of MI after registration highlighting that we may apply it to improve the registration of agriculture remotely sensed images. This small variation of MI shows that there is room for improvement.
Sergio Stempliuk, David Menotti
IGARSS2
2020 Deep Learning for Image-based Automatic Dial Meter Reading: Dataset and Baselines
abstract
Smart meters enable remote and automatic electricity, water and gas consumption reading and are being widely deployed in developed countries. Nonetheless, there is still a huge number of non-smart meters in operation. Image-based Automatic Meter Reading (AMR) focuses on dealing with this type of meter readings. We estimate that the Energy Company of Paraná (Copel), in Brazil, performs more than 850,000 readings of dial meters per month. Those meters are the focus of this work. Our main contributions are: (i) a public real-world dial meter dataset (shared upon request) called UFPR-ADMR; (ii) a deep learning-based recognition baseline on the proposed dataset; and (iii) a detailed error analysis of the main issues present in AMR for dial meters. To the best of our knowledge, this is the first work to introduce deep learning approaches to multidial meter reading, and perform experiments on unconstrained images. We achieved a 100.0% F1-score on the dial detection stage with both Faster R-CNN and YOLO, while the recognition rates reached 93.6% for dials and 75.25% for meters using Faster R-CNN (ResNeXt-101).
Gabriel Salomon 0002, Rayson Laroca, David Menotti
IJCNN3
2020 Video action recognition based on visual rhythm representation
Thierry Pinheiro Moreira, David Menotti, Hélio Pedrini
J. Vis. Commun. Image Represent.2
2019 Multi-task Learning for Low-Resolution License Plate Recognition
Gabriel Resende Gonçalves, Matheus Alves Diniz, Rayson Laroca, David Menotti, William Robson Schwartz
CIARP4
2018 Multimodal Feature Level Fusion based on Particle Swarm Optimization with Deep Transfer Learning
abstract
There are several biometric-based systems which rely on a single biometric modality, most of them focus on face, iris or fingerprint. Despite the good accuracies obtained with single modalities, these systems are more susceptible to attacks, i.e, spoofing attacks, and noises of all kinds, especially in non-cooperative (in-the-wild) environments. Since non-cooperative environments are becoming more and more common, new approaches involving multi-modal biometrics have received more attention. One challenge in multimodal biometric systems is how to integrate the data from different modalities. Initially, we propose a deep transfer learning optimized from a model trained for face recognition achieving outstanding representation for only iris modality. Our feature level fusion by means of features selection targets the use of the Particle Swarm Optimization (PSO) for such aims. In our pool, we have the proposed iris fine-tuned representation and a periocular one from previous work of us. We compare this approach for fusion in feature level against three basic function rules for matching at score level: sum, multi, and min. Results are reported for iris and periocular region (NICE.II competition database) and also in an open-world scenario. The experiments in the NICE.II competition databases showed that our transfer learning representation for iris modality achieved a new state-of-the-art, i.e., decidability of 2.22 and 14.56% of EER. We also yielded a new state-of-the-art result when the fusion at feature level by PSO is done on periocular and iris modalities, i.e., decidability of 3.45 and 5.55% of EER.
Pedro Silva 0004, Eduardo José da S. Luz, Luiz Antonio Zanlorensi, David Menotti, Gladston J. P. Moreira
CEC4
2018 QRS Detection in ECG Signal with Convolutional Network
Pedro Silva 0004, Eduardo José da S. Luz, Elizabeth Wanner, David Menotti, Gladston J. P. Moreira
CIARP4
2018 A Robust Real-Time Automatic License Plate Recognition Based on the YOLO Detector
abstract
Automatic License Plate Recognition (ALPR) has been a frequent topic of research due to many practical applications. However, many of the current solutions are still not robust in real-world situations, commonly depending on many constraints. This paper presents a robust and efficient ALPR system based on the state-of-the-art YOLO object detector. The Convolutional Neural Networks (CNNs) are trained and finetuned for each ALPR stage so that they are robust under different conditions (e.g., variations in camera, lighting, and background). Specially for character segmentation and recognition, we design a two-stage approach employing simple data augmentation tricks such as inverted License Plates (LPs) and flipped characters. The resulting ALPR approach achieved impressive results in two datasets. First, in the SSIG dataset, composed of 2,000 frames from 101 vehicle videos, our system achieved a recognition rate of 93.53% and 47 Frames Per Second (FPS), performing better than both Sighthound and OpenALPR commercial systems (89.80% and 93.03%, respectively) and considerably outperforming previous results (81.80%). Second, targeting a more realistic scenario, we introduce a larger public dataset1dataset, designed to ALPR. This dataset contains 150 videos and 4,500 frames captured when both camera and vehicles are moving and also contains different types of vehicles (cars, motorcycles, buses and trucks). In our proposed dataset, the trial versions of commercial systems achieved recognition rates below 70%. On the other hand, our system performed better, with recognition rate of 78.33% and 35 FPS.The UFPR-ALPR dataset is publicly available to the research community at https://web.inf.ufpr.br/vri/databases/ufpr-alpr/ subject to privacy restrictions.
Rayson Laroca, Evair Severo, Luiz Antonio Zanlorensi, Luiz Eduardo Soares de Oliveira, Gabriel Resende Gonçalves, William Robson Schwartz, David Menotti
IJCNN7
2018 A Benchmark for Iris Location and a Deep Learning Detector Evaluation
abstract
The iris is considered as the biometric trait with the highest unique probability. The iris location is an important task for biometrics systems, affecting directly the results obtained in specific applications such as iris recognition, spoofing and contact lenses detection, among others. This work defines the iris location problem as the delimitation of the smallest squared window that encompasses the iris region. In order to build a benchmark for iris location we annotate (iris squared bounding boxes) four databases from different biometric applications and make them publicly available to the community. Besides these 4 annotated databases, we include 2 others from the literature. We perform experiments on these six databases, five obtained with near infra-red sensors and one with visible light sensor. We compare the classical and outstanding Daugman iris location approach with two window based detectors: 1) a sliding window detector based on features from Histogram of Oriented Gradients (HOG) and a linear Support Vector Machines (SVM) classifier; 2) a deep learning based detector fine-tuned from YOLO object detector. Experimental results showed that the deep learning based detector outperforms the other ones in terms of accuracy and runtime (GPUs version) and should be chosen whenever possible.
Evair Severo, Rayson Laroca, Cides S. Bezerra, Luiz Antonio Zanlorensi, Daniel Weingaertner, Gladston J. P. Moreira, David Menotti
IJCNN7
2018 Robust automated cardiac arrhythmia detection in ECG beat signals
Victor Hugo C. de Albuquerque, Thiago M. Nunes, Danillo Roberto Pereira, Eduardo José da S. Luz, David Menotti, João Paulo Papa, João Manuel R. S. Tavares
Neural Comput. Appl.5
2018 Deep periocular representation aiming video surveillance
Eduardo José da S. Luz, Gladston J. P. Moreira, Luiz Antonio Zanlorensi, David Menotti
Pattern Recognit. Lett.4
2018 Learning Deep Off-the-Person Heart Biometrics Representations
abstract
Since the beginning of the new millennium, the electrocardiogram (ECG) has been studied as a biometric trait for security systems and other applications. Recently, with devices such as smartphones and tablets, the acquisition of ECG signal in the off-the-person category has made this biometric signal suitable for real scenarios. In this paper, we introduce the usage of deep learning techniques, specifically convolutional networks, for extracting useful representation for heart biometrics recognition. Particularly, we investigate the learning of feature representations for heart biometrics through two sources: on the raw heartbeat signal and on the heartbeat spectrogram. We also introduce heartbeat data augmentation techniques, which are very important to generalization in the context of deep learning approaches. Using the same experimental setup for six methods in the literature, we show that our proposal achieves state-of-the-art results in the two off-the-person publicly available databases.
Eduardo José da S. Luz, Gladston J. P. Moreira, Luiz Eduardo Soares de Oliveira, William Robson Schwartz, David Menotti
IEEE Trans. Inf. Forensics Secur.5
2017 Noisy Character Recognition Using Deep Convolutional Neural Networks
Sirlene Peixoto, Gabriel Resende Gonçalves, Andrea Bianchi, Alceu De S. Brito, William Robson Schwartz, David Menotti
CIARP6
2017 First-person action recognition through Visual Rhythm texture description
abstract
First-person action recognition is a recent problem in computer vision, where an observer wears body cameras to understand and recognize actions from the captured video sequences. Technological advances have made it possible to offer small wearable cameras that can be attached onto bike helmets, belts, animal halters, among other accessories. Examples of potential applications include sports, security, healthcare, visual lifelogging, among others. In this paper, we propose a novel approach to first-person action recognition that consists in encoding video appearance, shape and motion information as visual rhythms and describing them through texture analysis. Experiments are conducted on the DogCentric Activity and JPL First-Person Interaction datasets, showing accuracy improvement over the baselines.
Thierry Pinheiro Moreira, David Menotti, Hélio Pedrini
ICASSP2
2017 Combination techniques for hyperspectral image interpretation
abstract
In this work, we propose two main contributions to hyperspectral image interpretation. Firstly, while the traditional Weighted Linear Combination optimized by Genetic Algorithms (WLC-GA) [1] intends to give more discriminant power to those classification approaches contributing the most, we extend it to make a fine tuning over the class probabilities within the combination process. Then, we compare both methods (WLC-GA and its extension) with a more complex non-linear meta learning strategy called Stacked Generalization in which Support Vector Machines with Radial Basis Function kernel was used as combiner [2]. The experimental results, considering two widely used data sets, the Indian Pines and the Pavia University, are conducted in three different scenarios. Results show that both WLC-GA and its extended version achieve the best overall accuracy, and the proposed classification approach overcomes the accuracies of the other traditional ones used in this study.
Andrey Bicalho Santos, Arnaldo de Albuquerque Araújo, Jefersson A. dos Santos, William Robson Schwartz, David Menotti
IGARSS5
2017 Colorness index strategy for pixel fire segmentation
abstract
The use of computational resources of the area of image processing for early fire detection proves to be a good low-cost alternative to save human lives and ecological systems. In the literature, the general fire detection approaches use some color channel rules to determine if a pixel is a fire one. In this paper, we propose a new method of fire pixel segmentation using Colorness theory to red, yellow and brown indices combined with a reference probability table. The proposed method was evaluated by comparison with other 12 technical previews works. With a better efficiency than other methods in terms of true positive (82.05%) Accuracy (88.40%), F-Measure (84, 96%), the results show that our proposed method is more accurate to detect fire regions, indicating the effectiveness contribution of our fire probabilistic Colorness method.
Bruno Miguel Nogueira de Souza, Jacques Facon, David Menotti
IJCNN3
2017 Bias effect on predicting market trends with EMD
Dennis Carnelossi Furlaneto, Luiz Eduardo Soares de Oliveira, David Menotti, George D. C. Cavalcanti
Expert Syst. Appl.3
2016 Optimizing acceptance frontier using PSO and GA for multiple signature iris recognition
abstract
In the last three decades, the eye iris has been investigated as the most unique phenotype feature in the biometric literature. In the iris recognition literature, works achieving outstanding accuracies in very well behavior environments, i.e., when the subject's eye is in a well lightened environment and at a fixed distance for image acquisition, are known since a decade. Nonetheless, in noncooperative environments, where the image acquisition aiming iris location/segmentation is not straightforward, the iris recognition problem has many open issues and has been well researched in recent works. A promising strategy in this context employs a classification approach using multiple signature extraction. Representations are extracted from overlapping and different parts of the iris region. Aiming a robust and noisy invariant classification, an increasing set of acceptance thresholds (frontier) are required for dealing with multiple signatures for iris recognition in the iris matching process. Usually this frontier is estimated using brute force algorithms and a specific step resolution playing an important trade-off between runtime and accuracy of recognition in terms of false acceptance and false rejection rates, measured using half total error rate (HTER). In this sense, this work aims to use Genetic Algorithms (GA) and Particle Swarm Optimization (PSO) for finding such frontier. Moreover, we also employ a robust feature extraction technique proposed by us in a previous work. The experiments showed that the use of these evolutionary algorithms provides similar, if not better, effectiveness in very little runtime in a complex database well-known in the literature. Furthermore, it is shown that the frontier obtained by PSO is more stable than the one obtained by GA.
Gladston J. P. Moreira, Eduardo José da S. Luz, David Menotti
CEC3
2016 Improving automatic cardiac arrhythmia classification: Joining temporal-VCG, complex networks and SVM classifier
abstract
The classification of heartbeats using electrocardiogram (ECG) aiming arrhythmia detection is a well researched subject and still there are room for improvements concerning the recommended databases. In this sense, aiming to classify heartbeats for arrhythmia detection, we extend a previous ours proposal that uses vectorcardiogram, a bi-dimensional representation of two ECG leads, by incorporating the time component producing a three-dimensional representation, the temporal vectorcardiogram. Along with the new representation, also we apply complex networks to extract features from the temporal VCG. The new proposed features feed then a Support Vector Machines (SVM) classifier. The temporal VCG have increased in the global accuracy, and have better results classifying the N and S classes, when it is compared with the best result to our previous work with VCG. We conclude that new techniques to extract 3D features from the Temporal VCG could be an interesting research direction.
Gabriel Garcia, Gladston J. P. Moreira, Eduardo José da S. Luz, David Menotti
IJCNN4
2015 Denoising Autoencoder for Iris Recognition in Noncooperative Environments
Eduardo José da S. Luz, David Menotti
CIARP2
2015 Efficient Polynomial Implementation of Several Multithresholding Methods for Gray-Level Image Segmentation
David Menotti, Laurent Najman, Arnaldo de Albuquerque Araújo
CIARP1
2015 Fast and Accurate Gesture Recognition Based on Motion Shapes
Thierry Pinheiro Moreira, Marlon Fernandes de Alcântara, Hélio Pedrini, David Menotti
CIARP4
2015 A Multi-objective Approach for Building Hyperspectral Remote Sensed Image Classifier Combiners
Sandro Luiz Jailson Lopes Tinôco, David Menotti, Jefersson A. dos Santos, Gladston J. P. Moreira
EMO (2)2
2015 Hyperspectral image interpretation based on partial least squares
abstract
Remote sensed hyperspectral images have been used for many purposes and have become one of the most important tools in remote sensing. Due to the large amount of available bands, e.g., a few hundreds, the feature extraction step plays an important role for hyperspectral images interpretation. In this paper, we extend a well-know feature extraction method called Extended Morphological Profile (EMP) which encodes spatial and spectral information by using Partial Least Squares (PLS) to emphasize the importance of the more discriminative features. PLS is employed twice in our proposal, i.e., to the EMP features and to the raw spatial information, which are then concatenated to be further interpreted by the SVM classifier. Our experiments in two well-known data sets, the Indian Pines and Pavia University, have shown that our proposal outperforms the accuracy of classification methods employing EMP and other baseline feature extraction methods with different classifiers.
Andrey Bicalho Santos, Arnaldo de Albuquerque Araújo, William Robson Schwartz, David Menotti
ICIP4
2015 Deep Representations for Iris, Face, and Fingerprint Spoofing Detection
abstract
Biometrics systems have significantly improved person identification and authentication, playing an important role in personal, national, and global security. However, these systems might be deceived (or spoofed) and, despite the recent advances in spoofing detection, current solutions often rely on domain knowledge, specific biometric reading systems, and attack types. We assume a very limited knowledge about biometric spoofing at the sensor to derive outstanding spoofing detection systems for iris, face, and fingerprint modalities based on two deep learning approaches. The first approach consists of learning suitable convolutional network architectures for each domain, whereas the second approach focuses on learning the weights of the network via back propagation. We consider nine biometric spoofing benchmarks - each one containing real and fake samples of a given biometric modality and attack type - and learn deep representations for each benchmark by combining and contrasting the two learning approaches. This strategy not only provides better comprehension of how these approaches interplay, but also creates systems that exceed the best known results in eight out of the nine benchmarks. The results strongly indicate that spoofing detection systems based on convolutional networks can be robust to attacks already known and possibly adapted, with little effort, to image-based attacks that are yet to come.
David Menotti, Giovani Chiachia, Allan Pinto, William Robson Schwartz, Hélio Pedrini, Alexandre X. Falcão, Anderson Rocha 0001
IEEE Trans. Inf. Forensics Secur.1
2014 GPUs and Multicore CPUs Implementations of a Static Video Summarization
Suellen S. de Almeida, Edward Cayllahua, Arnaldo de Albuquerque Araújo, Guillermo Cámara Chávez, David Menotti
CIARP5
2014 An Adaptive Vehicle License Plate Detection at Higher Matching Degree
Raphael Felipe de Carvalho Prates, Guillermo Cámara Chávez, William Robson Schwartz, David Menotti
CIARP4
2014 A new approach for multiple instance learning based on a homogeneity bag operator
Alexandre W. C. Faria, David Menotti, André P. Lemos, Antônio de Pádua Braga
ESANN2
2014 An Optimized Sliding Window Approach to Pedestrian Detection
abstract
While a large number of surveillance cameras available nowadays provide a safe environment, the huge amount of data generated by them prevents a manual processing, requiring the application of automated methods to understand the scene. However, the majority of the currently available methods are still unable to process this amount of data in real time, mainly those focusing on pedestrian detection. To optimize pedestrian detection methods, this work proposes a novel approach that performs a random filtering supported by the Maximum Search Problem theorem to select a very small number from all possible detection windows. Although the random filtering is able to select regions that capture every person on an image, some windows can cover only parts of a person, diminishing the accuracy. To solve that, a regression is applied to adjust the windows to the person's location. The computational cost reduction comes from the fact that the proposed approach does not need to perform any processing while selecting windows, differently from cascades of rejection that must evaluate at least simple features for every window. The experiments performed using a pedestrian detection based on Partial Least Squares show that the approach is effective in both accuracy and computational cost reduction.
Victor C. de Melo, Samir Leao, David Menotti, William Robson Schwartz
ICPR3
2014 Evaluating the use of ECG signal in low frequencies as a biometry
Eduardo José da S. Luz, David Menotti, William Robson Schwartz
Expert Syst. Appl.2
2013 Fast pedestrian detection based on a partial least squares cascade
abstract
In applications such as surveillance, pedestrian detection can be seen as a filtering stage which will locate the objects of interest so that higher level tasks, such as recognition, re-identification, action and activity recognition, can be performed considering only those objects. Therefore, it is imperative that the pedestrian detection task presents low computational cost. Several methods have been proposed to detect pedestrians in images and videos. However, a remaining challenge is to detect pedestrians with high accuracy at a very low computational cost. Towards accomplishing the goal of reducing the costs for pedestrian detection, we propose a cascade of rejection based on Partial Least Squares (PLS) and the variable selection method Variable Importance in Projection (VIP) combined with the propagation of latent variables through the stages. Our results show that the method reduces the computational cost by increasing the number of rejected background samples in earlier stages of the cascade.
Victor C. de Melo, Samir Leao, Mario Fernando Montenegro Campos, David Menotti, William Robson Schwartz
ICIP4
2013 Ensemble of classifiers for remote sensed hyperspectral land cover analysis: An approach based on Linear Programming and Weighted Linear Combination
abstract
Hyperspectral images have been considered as one of the most important tool for remote sensed land cover analysis. Such images have information about materials on earth's surface expressed in many wavelengths that allow us to identify and classify those materials with more accuracy. In this work we used a combination of several classification methods in order to produce an accurate thematic map based on the remote sensed hyperspectral image classification. To perform the combination, three types of feature representation and two learning algorithms (Support Vector Machines (SVM) and Backpropagation Multilayer Perceptron Neural Network (MLP)) were used yielding six classification methods. Our approach proposal is based onWeighted Linear Combination (WLC), in which weights are found using Linear Programming (LP) - WLC-LP. Experiments are carried out using two well-known databases: Indian Pines, acquired by AVIRIS sensor; and Pavia University, acquired by ROSIS sensor. Results show the efficiency of our proposed approach which significantly reduces the time required to found optimal weights for the combiner compared to a previous approach based on Genetic Algorithm.
Sandro Luiz Jailson Lopes Tinôco, Haroldo G. Santos, David Menotti, Andrey Bicalho Santos, Jefersson A. dos Santos
IGARSS3
2013 ECG arrhythmia classification based on optimum-path forest
Eduardo José da S. Luz, Thiago M. Nunes, Victor Hugo C. de Albuquerque, João Paulo Papa, David Menotti
Expert Syst. Appl.5
2012 Combiner of classifiers using Genetic Algorithm for classification of remote sensed hyperspectral images
abstract
In the past few years, hyperspectral images have been considered as one of the most important tool in land cover classification due to its capability to obtain rich information of materials on earth surface. In this work we aim to produce an accurate thematic map for the remote sensed hyperspectral image classification problem, which is obtained using a combination of several classification methods. Three types of feature representation and two learning algorithms (Support Vector Machines (SVM) and Backpropagation Multilayer Perceptron Neural Network (MLP)) were used yielding six classification methods to perform the combination. Our combination proposal is based on Weighted Linear Combination (WLC), in which weights are found using a Genetic Algorithm (GA) - WLC-GA. Experiments were carried out with two well-known datasets: Indian Pines and Pavia University, and we observed that our proposed WLC-GA method achieves the highest accuracy among traditional Conscious Combiners, the widely used Majority Vote (MV) and Weighted Majority Vote (WMV), for both datasets.
Andrey Bicalho Santos, Arnaldo de Albuquerque Araújo, David Menotti
IGARSS3
2012 A methodology for photometric validation in vehicles visual interactive systems
Alexandre W. C. Faria, David Menotti, Gisele L. Pappa, Daniel S. D. Lara, Arnaldo de Albuquerque Araújo
Expert Syst. Appl.2
2011 Application of complex networks for automatic classification of damaging agents in soybean leaflets
abstract
Many of the difficulties in managing soybean tillage are related to the identification of insect/pests harmful to the plant, since tillage can be attacked by a wide range of such agents. By identifying the most common agents that cause damages to the leaflets, we can obtain more knowledge about appropriate strategies of control. The proposed work presents an automatic method for classification of the main agents that cause damages to soybean leaflets, i.e., beetles and caterpillars. Acquired images are preprocessed and the contours of the damages are taken. Each contour is modeled as a complex network. Features are extracted for each damage based on the connectivity and the joint degree of this network. These features are then used to train a SVM algorithm. In the experiments, we analyze thresholds which model the network and the proposed method reports accuracy greater than 90% for damaging agent classification.
Thiago L. G. Souza, Eduardo S. Mapa, Kayran dos Santos, David Menotti
ICIP4
2011 Towards an automatic vehicle access control system: License plate location
abstract
An automatic vehicle access control system (AVACS) can be divided into three steps: vehicle location, vehicle license plate (VLP) location, and VLP recognition. This paper presents a new method for VLP location based on the horizontal gradient, morphological operations, connected components analysis, and statistical measures. First, the horizontal gradient is acquired and a mean filter is applied on it. Morphological operations are then used to darken non-VLP high-valued regions (saliences). At this step, we work on an integer-valued image rather than on the binary one, improving salience detection. Finally, the image is binarized, then a connected component analysis is performed, and statistical measures are used to decide among the VLP candidates. Experiments show that, in a database of 722 images, our method correctly locates the VLP in 95% of the cases, outperforming previous approaches.
Pedro Ribeiro Mendes Júnior, José Maria Ribeiro Neves, Andréa Iabrudi Tavares, David Menotti
SMC4
2011 An overview of automatic event detection in soccer matches
abstract
Sports video analysis has received special attention from researchers due to its high popularity and general interest on semantic analysis. Hence, soccer videos represent an interesting field for research allowing many types of applications: indexing, summarization, players' behavior recognition and so forth. Many approaches have been applied for field extraction and recognition, arc and goalmouth detection, ball and players tracking, and high level techniques such as team tactics detection and soccer models definition. In this paper, we provide an hierarchy and we classify approaches into this hierarchy based on their analysis level, i.e., low, middle, and high levels. An overview of soccer event identification is presented and we discuss general issues related to it in order to provide relevant information about what has been done on soccer video processing.
Samuel de Sousa 0001, Arnaldo de Albuquerque Araújo, David Menotti
WACV3
2005 Lacunarity as a Texture Measure for Address Block Segmentation
Jacques Facon, David Menotti, Arnaldo de Albuquerque Araújo
CIARP2
2005 Statistical Hypothesis Testing and Wavelet Features for Region Segmentation
David Menotti, Díbio Leandro Borges, Arnaldo de Albuquerque Araújo
CIARP1
2004 Fractal-Based Approach for Segmentation of Address Block in Postal Envelopes
Luiz Felipe Eiterer, Jacques Facon, David Menotti
CIARP3
2003 Segmentation of Postal Envelopes for Address Block Location: an approach based on feature selection in wavelet space
abstract
This paper presents a segmentation algorithm based on feature selection in wavelet space. The aim is to automatically separate in postal envelopes the regions related to background, stamps, rubber stamps, and the address blocks. First, a typical image of a postal envelope is decomposed using Mallat algorithm and Haar basis. High frequency channel outputs are analyzed to locate salient points in order to separate the background. A statistical hypothesis test is taken to decide upon more consistent regions in order to clean out some noise left. The selected points are projected back to the original gray level image, where the evidence from the wavelet space is used to start a growing process to include the pixels more likely to belong to the regions of stamps, rubber stamps, and written area. Experiments are run using original postal envelopes from the Brazilian Post Office Agency, and here we report results on 440 images with many different layouts and backgrounds. 1.
David Menotti, Díbio Leandro Borges, Jacques Facon, Alceu S. Britto Jr.
ICDAR1