VLDB 2026 Research / reviewers in the wild / expert
Guillermo Cámara Chávez
dblp:147/6674 · also Guillermo Cámara
· DBLP profile ↗
24ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-authorSoftware engineering, systems software and programming languages · 5 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spatio-Temporal Sign Recognition with Multiscale Vision Transformers and Multimodal FusionabstractSign language is the primary means of communication for many people who are deaf or hard of hearing. The advancement of automatic sign language recognition systems is essential to enhance accessibility and reduce communication barriers. However, the lack of robust and accessible technological solutions still poses a significant obstacle to the full social inclusion of this population. In this context, video signal recognition systems have gained prominence, especially with the advancement of computer vision techniques based on deep learning. This study presents a multimodal recognition approach that combines RGB and depth video data through an attention-based fusion mechanism. The proposed architecture employs two pre-trained Multiscale Vision Transformers (MViT) to extract spatiotemporal features from each modality. These features are then integrated using an attention module that dynamically adjusts the contribution of each input stream. Experiments were conducted using the LIBRAS-UFOP dataset, which contains 56 signs performed by five different individuals, grouped into four linguistic categories. To evaluate the model’s performance, two evaluation protocols were employed: a random data split and a leave-one-signer-out strategy to assess the model’s ability to generalize to unseen users. The results show that while multimodal fusion provided modest improvements in the random split scenario, it achieved significantly higher accuracy when evaluated on unseen signers. These findings demonstrate the value of combining multiple modalities and leveraging attention mechanisms to address user variability, ultimately contributing to more robust and reliable sign language recognition in real-world applications. Graziela Silva Araújo, Luana Isabel Gonçalvez de Lima, Roberta B. Oliveira, Guillermo Cámara Chávez |
CLEI | 4 |
| 2021 | Convolutional Neural Networks Applied for Skin Lesion SegmentationabstractSkin cancer is one of the cancers that most aggravates the problem in public health. Among the types of cancer, melanoma is the most aggressive type. Its early diagnosis is essential to increase the possibility of adequate treatment, aiming to reduce the mortality rate. Dermatologists generally use manual methods to diagnose skin lesions. These methods, in addition to being time-consuming, as they are performed manually, can present different results for the same lesion when analyzed by different specialists. Therefore, an automated diagnosis may be necessary to deal with this issue as well as avoid invasive tests. For this, the task of segmenting the skin lesion in the dermoscopic image can be fundamental, as it is a basic task in the image analysis process. In the present work, a Convolutional Neural Network (CNN) model, based on the U-Net, is used to segment the lesion in dermoscopic images. This proposal achieved an accuracy of 0.949 and Jaccard of 0.833 for the 2017 ISIC base, and an accuracy of 0.954 and Jaccard of 0.850 for the 2018 ISIC base. The proposed model has a simpler architecture, in addition to requiring less computational resources. The experiments made it possible to observe that the proposed model results are promising compared with other CNN models presented in the literature. Graziela Silva Araújo, Guillermo Cámara Chávez, Roberta B. Oliveira |
CLEI | 2 |
| 2021 | A comparative study of WHO and WHEN prediction approaches for early identification of university students at dropout riskabstractReducing the students' dropout is one of the biggest challenges faced by educational institutions, especially in underdeveloped countries. Identification of the student with the highest risk of dropping out is generally used to apply corrective actions (WHO). Therefore, it is also important to determine WHEN a student will drop out, which is fundamental to planning preventive actions. In this work, we perform a study to quantitatively compare several approaches to address the early identification of dropout students in universities. We categorize our study into three main methods families, i.e., analytical methods, traditional classification methods, and probabilistic methods. The first is exploited at preprocessing step for selecting significant variables into the dropout identification task. The second uses machine learning models to classify students into dropout prone or non-dropout prone classes. The third family uses survival models to determine when the student would desert. To evaluate the predictive capacity of the classification models, the Kappa coefficient was incorporated into the usual machine learning metrics and shows that Kappa is handy for evaluating performance in unbalanced data. Similarly, in the survival models, the concordance index was applied to evaluate the predictive capacity. Our approach was applied over a real data set of Peruvian university graduate students to identify when and who will drop out. Daniel A. Gutierrez-Pachas, Germain García-Zanabria, Alex J. Cuadros-Vargas, Guillermo Cámara Chávez, Jorge Poco, Erick Gomez Nieto |
CLEI | 4 |
| 2021 | Ear Recognition In The Wild with Convolutional Neural NetworksabstractEar recognition has gained attention in recent years. The possibility of being captured from a distance, contactless, without the cooperation of the subject and not be affected by facial expressions makes ear recognition a captivating choice for surveillance and security applications, and even more in the current COVID-19 pandemic context where modalities like face recognition fail due to mouth and facial covering masks usage. Applying any deep learning (DL) algorithm usually demands a large amount of training data and appropriate network architectures, therefore we introduce a large-scale database and explore fine-tuning pre-trained convolutional neural networks (CNNs) looking for a robust representation of ear images taken under uncontrolled conditions. Taking advantage of the face recognition field, we built an ear dataset based on the VGGFace dataset and use the Mask-RCNN for ear detection. Besides, adapting the VGGFace model to the ear domain leads to a better performance than using a model trained for general image recognition. Experiments on the UERC dataset have shown that fine-tuning from a face recognition model and using a larger dataset leads to a significant improvement of around 9% compared to state-of-the-art methods on the ear recognition field. In addition, we have explored score-level fusion by combining matching scores of the fine-tuning models which leads to an improvement of around 4% more. Open-set and close-set experiments have been performed and evaluated using Rank-1 and Rank-5 recognition rate metrics. Solange Ramos-Cooper, Guillermo Cámara Chávez |
CLEI | 2 |
| 2021 | A multimodal LIBRAS-UFOP Brazilian sign language dataset of minimal pairs using a microsoft Kinect sensor
Lourdes Ramirez Cerna, Edwin Jonathan Escobedo Cardenas, Dayse Garcia Miranda, David Menotti, Guillermo Cámara Chávez |
Expert Syst. Appl. | 5 |
| 2020 | Real-Time Violence Detection in Videos Using Dynamic ImagesabstractThe problem of violence detection consists of identifying scenes that characterize violence in a video stream. The violent actions in question can be of the most diverse, from fights, pushes, and robberies to shots and explosions. Detecting the presence of violence is useful for classifying videos and films, blocking inappropriate content for specific audiences, and improving security personnel's performance responsible for areas under surveillance. This work proposes an approach based on the Dynamic Images method, using handcrafted and CNN features the Bag of Visual Words paradigm and a SVM classifier to detect violent actions that involve corporal struggle in the video streams of databases of literature. The proposed methods can achieve an average accuracy of 97.50% for the Hockey dataset, 99.80% for the Movies dataset, and 93.40% for the Crowd dataset. Besides, the identification of violence in each video was performed in of hundredths of a second. Also, the techniques proposed in this work have the advantage that they can be applied even in environments where computational resources are limited, and technologies such as GPU or parallel processing are not available. Ademir Rafael Marques Guedes, Guillermo Cámara Chávez |
CLEI | 2 |
| 2020 | Multimodal hand gesture recognition combining temporal and pose information based on CNN descriptors and histogram of cumulative magnitudes
Edwin Jonathan Escobedo Cardenas, Guillermo Cámara Chávez |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Hierarchical segmentation from a non-increasing edge observation attribute
Edward Cayllahua, Jean Cousty, Silvio Jamil Ferzoli Guimarães, Yukiko Kenmochi, Guillermo Cámara Chávez, Arnaldo de Albuquerque Araújo |
Pattern Recognit. Lett. | 5 |
| 2019 | Cross-Domain Interpolation for Unpaired Image-to-Image Translation
Jorge López, Antoni Mauricio, José Díaz 0002, Guillermo Cámara Chávez |
ICVS | 4 |
| 2019 | A Sequential Approach for Pain Recognition Based on Facial Representations
Antoni Mauricio, Fabio A. M. Cappabianco, Adriano Veloso, Guillermo Cámara Chávez |
ICVS | 4 |
| 2019 | Efficient Algorithms for Hierarchical Graph-Based Segmentation Relying on the Felzenszwalb-Huttenlocher DissimilarityabstractHierarchical image segmentation provides a region-oriented scale-space, i.e. a set of image segmentations at different detail levels in which the segmentations at finer levels are nested with respect to those at coarser levels. However, most image segmentation algorithms, among which a graph-based image segmentation method relying on a region merging criterion was proposed by Felzenszwalb–Huttenlocher in 2004, do not lead to a hierarchy. In order to cope with a demand for hierarchical segmentation, Guimarães et al. proposed in 2012 a method for hierarchizing the popular Felzenszwalb–Huttenlocher method, without providing an algorithm to compute the proposed hierarchy. This paper is devoted to providing a series of algorithms to compute the result of this hierarchical graph-based image segmentation method efficiently, based mainly on two ideas: optimal dissimilarity measuring and incremental update of the hierarchical structure. Experiments show that, for an image of size 321 × 481 pixels, the most efficient algorithm produces the result in half a second whereas the most naive one requires more than 4 h. Edward Cayllahua, Jean Cousty, Yukiko Kenmochi, Arnaldo de Albuquerque Araújo, Guillermo Cámara Chávez, Silvio Jamil Ferzoli Guimarães |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2018 | Hierarchy-Based Salient Regions: A Region Detector Based on Hierarchies of Partitions
Karla Otiniano-Rodríguez, Arnaldo de Albuquerque Araújo, Guillermo Cámara Chávez, Jean Cousty, Silvio Jamil Ferzoli Guimarães, Benjamin Perret |
CIARP | 3 |
| 2017 | Real-Time Brand Logo Recognition
Leonardo Bombonato, Guillermo Cámara Chávez, Pedro Silva 0004 |
CIARP | 2 |
| 2017 | Fusion of Deep Learning Descriptors for Gesture Recognition
Edwin Jonathan Escobedo Cardenas, Guillermo Cámara Chávez |
CIARP | 2 |
| 2017 | Abnormal Event Detection in Video Using Motion and Appearance Information
Neptalí Menejes Palomino, Guillermo Cámara Chávez |
CIARP | 2 |
| 2016 | Building semantic understanding beyond deep learning from sound and visionabstractDeep learning-based models have recently been widely successful at outperforming traditional approaches in several computer vision applications such as image classification, object recognition and action recognition. However, those models are not naturally designed to learn structural information that can be important to tasks such as human pose estimation and structured semantic interpretation of video events. In this paper, we demonstrate how to build structured semantic understanding of audio-video events by reasoning on multiple-label decisions of deep visual models and auditory models using Grenander's structures for imposing semantic consistency. The proposed structured model does not require joint training of the structural semantic dependencies and deep models. Instead they are independent components linked by Grenander's structures. Furthermore, we exploited Grenander's structures as a means to facilitate and enrich the model with fusion of multimodal sensory data; in particular, auditory features with visual features. Overall, we observed improvements in the quality of semantic interpretations using deep models and auditory features in combination with Grenander's structures, reflecting as numerical improvements of up to 11.5% and 12.3% in precision and recall, respectively. Fillipe D. M. de Souza, Sudeep Sarkar, Guillermo Cámara Chávez |
ICPR | 3 |
| 2015 | A robust gesture recognition using hand local data and skeleton trajectoryabstractIn this paper, we propose a new approach for dynamic hand gesture recognition using intensity, depth and skeleton joint data captured by KinectTMsensor. The proposed approach integrates global and local information of a dynamic gesture. First, we represent the skeleton 3D trajectory in spherical coordinates. Then, we extract the key frames corresponding to the points with more angular and distance difference. In each key frame, we calculate the spherical distance from the hands, wrists and elbows to the shoulder center, also we record the hands position changes to obtain the global information. Finally, we segment the hands and use SIFT descriptor on intensity and depth data. Then, Bag of Visual Words (BOW) approach is used to extract local information. The system was tested with the ChaLearn 2013 gesture dataset and our own Brazilian Sign Language dataset, achieving an accuracy of 88.39% and 98.28%, respectively. Edwin Jonathan Escobedo Cardenas, Guillermo Cámara Chávez |
ICIP | 2 |
| 2014 | GPUs and Multicore CPUs Implementations of a Static Video Summarization
Suellen S. de Almeida, Edward Cayllahua, Arnaldo de Albuquerque Araújo, Guillermo Cámara Chávez, David Menotti |
CIARP | 4 |
| 2014 | Detection of Groups of People in Surveillance Videos Based on Spatio-Temporal Clues
Rensso Mora Colque, Guillermo Cámara Chávez, William Robson Schwartz |
CIARP | 2 |
| 2014 | An Adaptive Vehicle License Plate Detection at Higher Matching Degree
Raphael Felipe de Carvalho Prates, Guillermo Cámara Chávez, William Robson Schwartz, David Menotti |
CIARP | 2 |
| 2014 | Spatial Pyramid Matching for Finger Spelling Recognition in Intensity Images
Samira Silva, William Robson Schwartz, Guillermo Cámara Chávez |
CIARP | 3 |
| 2011 | Color-Aware Local Spatiotemporal Features for Action Recognition
Fillipe Dias Moreira de Souza, Eduardo Valle, Guillermo Cámara Chávez, Arnaldo de Albuquerque Araújo |
CIARP | 3 |
| 2009 | MammoSVD: A content-based image retrieval system using a reference database of mammographiesabstractIn this paper, we present a content-based image retrieval (CBIR) system called MammoSVD. This CBIR system is developed based on breast density — fatty or dense, and the database used, from the IRMA project, provides images with the ground truth already set. Singular value decomposition (SVD) is proposed for the breast density characterization by the selection of the first singular values, in order to represent texture along with the dimensionality reduction. Support-vector machine (SVM) is used to perform the retrieval operation. Considering the first 10% of the retrieved images, the precision rate is 90%, indicating the potential of the implemented CBIR system. Júlia Epischina Engrácia de Oliveira, Ana Paula Brandão Lopes, Guillermo Cámara Chávez, Arnaldo de Albuquerque Araújo, Thomas M. Deserno |
CBMS | 3 |
| 2009 | Harris-SIFT Descriptor for Video Event Detection Based on a Machine Learning ApproachabstractVideo data is becoming increasingly important in many commercial and scientific areas with the advent of applications such as digital broadcasting, video-conferencing and multimedia processing tools, and with the development of the hardware and communications infrastructure necessary to support visual applications. The objective of this work is to propose a method for event detection in a video stream. We combine Harris-SIFT descriptor with motion information in order to detect human actions in video. We tested our method in KTH database and compared it to space-time interest points (STIP) descriptor. The results obtained achieved similar results to the STIP method. Guillermo Cámara Chávez, Arnaldo de Albuquerque Araújo |
ISM | 1 |