EDBT 2026 Demo / reviewers in the wild / expert
Jesus Olivares-Mercado
dblp:80/5031
· DBLP profile ↗
16ranked-venue papers
0as first author
8since 2021 · last 2024
0000-0002-0337-5364ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 6 since 2021Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PFMNet: Face Mask Recognition with Deformable Convolution Networks and Category AttentionabstractThe challenges posed by the COVID-19 pandemic underscored the critical importance of proper mask usage, highlighting the need for automated systems to monitor face mask-wearing conditions. In this paper, we introduce PFMNet, a novel architecture for recognizing the wearing status of face masks. PFMNet is inspired by the InternImage architecture and employs Deformable Convolution Networks (DCNs) to capture long-range dependencies crucial for accurate mask status determination. The significant challenge of class imbalance, particularly the scarcity of improperly worn mask samples, is addressed by integrating the Category Attention Block (CAB). CAB improves distinct regions, diversifies feature representations, and utilizes efficient global pooling to identify crucial areas, such as the human face, while reducing the computational cost. The performance of PFMNet was assessed using the publicly available PWMFD dataset, which had to be refined due to duplicate images and incorrect annotations. PFMNet was compared to three other state-of-the-art models: InternImage, ConvNext, and EfficientNet. It outperformed these models, achieving an accuracy of 99.39%. This places it ahead of the second-best model by a margin of 0.45%. The confusion matrices illustrate that PFMNet outperforms other models in all classes, particularly excelling in the “with mask” and “without mask” categories, resulting in the best overall performance. Ulises Arroyo-Rojas, Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Hiroki Takahashi |
SoMeT | 3 |
| 2024 | Transformation Approach for Safe Source Code Through the Application of a Large Language Model and Adaptation of a Generative Adversarial NetworkabstractIn the software development life cycle, the implementation of stringent security requirements is essential to promote the creation of robust and secure code, thereby avoiding the need for extensive post-implementation revisions. A wide variety of methodologies are commonly employed to examine source code authorship, ranging from adherence to strict standards and guidelines to the application of best practices. However, these reviews are often very laborious and demand a broad spectrum of specialized knowledge from various DevOps task groups to effectively address underlying vulnerabilities. To streamline and enhance the efficiency of the review process, advanced Machine Learning techniques are increasingly being adopted as a critical factor in improving the precision of transitions to secure code structures. This manuscript introduces an innovative transformation system that leverages the contextual adaptability provided by the renowned advanced language model, CodeBERT, integrated with a Generative Adversarial Network (GAN). This synergistic combination allows for the precise classification of insecure code segments in different programming languages and the subsequent generation of their secure counterparts. Empirical results confirm the system’s ability to detect up to 98.3% of insecure tokens and reconstruct secure versions with an accuracy of up to 95.67%. Aldo Hernandez-Suarez, Héctor M. Pérez Meana, Gabriel Sanchez-Perez, José Portillo-Portillo, Jesus Olivares-Mercado, Linda K. Toscano-Medina |
SoMeT | 5 |
| 2024 | Topic Modeling in the Darknet via Semi-Supervised Learning and Linguistic TransformersabstractIn recent years, the darknet, a hidden part of the deep web associated with illicit activities, has been the subject of study due to the myths and mysteries surrounding it. Contemporary research aims to uncover the true topics hidden within this network using thematic analysis techniques, which are essential for cybercrime prevention and legal action. However, the dynamic and anonymous nature of the darknet poses the challenge of effectively navigating the TOR protocol to obtain and analyze samples from hidden sites. This paper presents an innovative approach to studying the darknet. Assuming limited prior knowledge of the original topics, a contextual relation-comparison technique with TinyBERT, a large language model, is used to generate super topics from previously identified hidden sites. From these super topics, keywords with contextual scores and weights are extracted, serving as input for a sensor that navigates the TOR network and aggregates new hidden sites. These sites are processed through semi-supervised learning to form clusters of sub-topics. Labels for each sub-topic propagate based on their similarity to the main topics and are ultimately classified in a fine-tuning layer of TinyBERT. The results demonstrate the identification of twelve classes of sub-topics in the darknet, related to drugs, hacking, marketplaces, pornography, and other areas, with a classification accuracy of 95.45%. Aldo Hernandez-Suarez, Héctor M. Pérez Meana, Gabriel Sanchez-Perez, José Portillo-Portillo, Jesus Olivares-Mercado, Linda K. Toscano-Medina |
SoMeT | 5 |
| 2024 | Frame-Level Deepfake Detection on Explicit Content with ID-Unaware Binary ClassificationabstractThe rapid advancement in deepfake technology has enabled the creation of highly realistic fake images and videos, posing significant risks, especially in the context of explicit content. Such content, which often involves the alteration of an individual’s identity in sexually explicit material, can lead to defamation, harassment, and blackmail. This paper focuses on the detection of deepfakes in explicit content using a state-of-the-art ID-unaware Binary Classification method. We evaluate its effectiveness in real-world scenarios by analyzing three versions of the model with different backbones: ResNet34, EfficientNet-B3, and EfficientNet-B4. To facilitate this evaluation, we curated a dataset of 200 videos, consisting of 100 genuine videos and their corresponding deepfake counterparts, ensuring a direct comparison between genuine and altered content. Our analysis revealed a significant decrease in detection performance when applying the state-of-the-art method to explicit content. Specifically, the AUC score dropped from 93% on standard datasets such as FaceForensics++ to 62% on our explicit content dataset. Additionally, the accuracy for detecting deepfakes plummeted to around 25%, while the accuracy for genuine videos remained high at approximately 90%. We identified specific factors contributing to this decline, including unconventional makeup, lighting issues, and facial blurring due to camera distance. These findings underscore the challenges and the necessity for robust detection methods to address the unique problems posed by explicit content deepfakes, ultimately aiming to protect individuals from the potential harms associated with this technology. Miguel Jimenez-Martinez, Gibran Benitez-Garcia, Linda K. Toscano-Medina, Jesus Olivares-Mercado |
SoMeT | 4 |
| 2022 | TFM a Dataset for Detection and Recognition of Masked Faces in the WildabstractDroplet transmission is one of the leading causes of the spread of respiratory infections, such as coronavirus disease (COVID-19). The proper use of face masks is an effective way to prevent the transmission of such diseases. Nonetheless, different types of masks provide various degrees of protection. Hence, automatic recognition of face mask types may benefit the control access to facilities where a specific protection degree is required. In the last two years, several deep learning models have been proposed for face mask detection and properly wearing mask recognition. However, the current publicly available datasets do not consider the different mask types and occasionally lack real-world elements needed to train robust models. In this paper, we introduce a new dataset named TFM with sufficient size and variety to train and evaluate deep learning models for face mask detection and recognition. This dataset contains more than 135,000 annotated faces from about 100,000 photographs taken in the wild. We consider four mask types (cloth, respirators, surgical and valved) as well as unmasked faces, of which up to six can appear in a single image. The photographs were mined from Twitter within two years since the beginning of the COVID-19 pandemic. Thus, they include diverse scenes with real-world variations in background and illumination. With our dataset, the performance of four state-of-the-art object detection models is evaluated. The experimental results show that YOLOv5 can achieve about 90% of [email protected], demonstrating that the TFM dataset can be used to train robust models and may help the community step forward in detecting and recognizing masked faces in the wild. Our dataset and pre-trained models used in the evaluation will be available upon the publication of this paper. Gibran Benitez-Garcia, Hiroki Takahashi, Miguel Jimenez-Martinez, Jesus Olivares-Mercado |
MMAsia | 4 |
| 2022 | Twitter Face Image Mining for Recognition of Different Face Mask TypesabstractIn the current pandemic of coronavirus disease (COVID-19), an effective way to prevent the transmission and infection of the virus is the proper use of face masks. However, the different types of masks provide different degrees of protection. For instance, valved masks protect the user but do not help to stop the transmission. Hence, the automatic recognition of face mask types may benefit applications that control access to facilities where a certain facepiece is required. In this paper, we propose a Twitter mining framework to gather a large-scale dataset of masked faces suitable to train deep learning-based models for face mask recognition. We employ a keyword-based selection where non-face images are discarded by an efficient face detector (Retinaface). Finally, we train a state-of-the-art CNN architecture (ConvNeXt) for recognizing the wearing mask. We also present a brief analysis of more than two million image-based tweets acquired over two years since the beginning of the pandemic. The code of the proposed framework and a preliminary dataset of more than 10K faces (manually annotated into unmasked, surgical, cloth, respirators, and valved masks) are available on github.com/GibranBenitez/FaceMask Twitter. Ulises Arroyo-Rojas, Miguel Jimenez-Martinez, Gibran Benitez-Garcia, Jesus Olivares-Mercado, Hiroki Takahashi |
SoMeT | 4 |
| 2022 | FASSD-Net: Fast and Accurate Real-Time Semantic Segmentation for Embedded SystemsabstractRecent works of real-time semantic segmentation, remove or make use of light decoders from dense deep neural networks to achieve fast inference speed. This strategy helps to achieve real-time performance; however, the accuracy is significantly compromised in comparison to non-real-time methods. In this paper, we introduce two key modules aimed to design a high-performance decoder for real-time semantic segmentation, which also reduces the accuracy gap between real-time and non-real-time networks. The first module, Dilated Asymmetric Pyramidal Fusion (DAPF), is designed to increase the receptive field on the top of the last stage of the encoder, obtaining richer contextual features. The second module, Multi-resolution Dilated Asymmetric (MDA) module, fuses and refines detail and contextual information from multi-scale feature maps coming from early and deeper stages of the network. Both modules are designed to keep a low computational complexity by using asymmetric convolutions. With these modules, we propose a network entitled “FASSD-Net,” which is based on a light-weight CNN backbone. Running on a single Nvidia GTX 1080Ti, our model reaches 77.5% and 69.3% of mIoU, at 41 and 80 FPS on the Cityscapes and CamVid datasets, respectively. We present an extensive analysis of the accuracy-speed tradeoffs of three FASSD-Net variations on different embedded systems, demonstrating that a light version of our network can run on the low-power consumption Jetson Xavier NX, at 32 FPS reaching 74% of mIoU with full resolution ($1024\times 2048$). The source code and pre-trained models are available at github.com/GibranBenitez/FASSD-Net. Leonel Rosas-Arias, Gibran Benitez-Garcia, José Portillo-Portillo, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Keiji Yanai |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Fingerprint Recognition System Based on Bifurcation MinutiaesabstractNowadays, fingerprint is the biometric more implemented to authentication and recognition of people for governmental and private purposes. This paper aims present the implementation of a fingerprint recognition system based only on bifurcation minutiaes and singularities to create a template, the template obtained is stored and used on recognition and verification tasks. The evaluation of the proposed system shows that using the bifurcation minutiaes the system provides high results and a good performance, the results were obtained in recognition and verification ways and the processing time was measured via an user interface. Alberto Antonio Vargas Mata, Jesus Olivares-Mercado, Linda K. Toscano-Medina, Gabriel Sanchez-Perez, Héctor M. Pérez Meana |
SoMeT | 2 |
| 2020 | IPN Hand: A Video Dataset and Benchmark for Real-Time Continuous Hand Gesture RecognitionabstractContinuous hand gesture recognition (HGR) is an essential part of human-computer interaction with a wide range of applications in the automotive sector, consumer electronics, home automation, and others. In recent years, accurate and efficient deep learning models have been proposed for HGR. However, in the research community, the current publicly available datasets lack real-world elements needed to build responsive and efficient HGR systems. In this paper, we introduce a new benchmark dataset named IPN Hand with sufficient size, variety, and real-world elements able to train and evaluate deep neural networks. This dataset contains more than 4,000 gesture samples and 800,000 RGB frames from 50 distinct subjects. We design 13 different static and dynamic gestures focused on interaction with touchless screens. We especially consider the scenario when continuous gestures are performed without transition states, and when subjects perform natural movements with their hands as non-gesture actions. Gestures were collected from about 30 diverse scenes, with real-world variation in background and illumination. With our dataset, the performance of three 3D-CNN models is evaluated on the tasks of isolated and continuous realtime HGR. Furthermore, we analyze the possibility of increasing the recognition accuracy by adding multiple modalities derived from RGB frames, i.e., optical flow and semantic segmentation, while keeping the real-time performance of the 3D-CNN model. Our empirical study also provides a comparison with the publicly available nvGesture (NVIDIA) dataset. The experimental results show that the state-of-the-art ResNext-101 model decreases about 30% accuracy when using our real-world dataset, demonstrating that the IPN Hand dataset can be used as a benchmark, and may help the community to step forward in the continuous HGR. Our dataset and pre-trained models used in the evaluation are publicly available at github.com/GibranBenitez/IPN-hand. Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Keiji Yanai |
ICPR | 2 |
| 2020 | A Fast-RCNN Implementation for Human Silhouette Detection in Video SequencesabstractThe intention of this article is to implement a system of detection and segmentation of human silhouettes, the above mentioned tasks present a great challenge in security topics and innovation, in the last years and mainly on automated video surveillance systems, which require understanding the presence and human interaction in video sequences, e.g. Human Computer Interaction (HCI), Human Behaviour comprehension, Human fall detection, among others, but the most important is behavioural biometrics, this paper tackles the common step in these research areas: the Human silhouette extraction through the bounding box. To evaluate the proposed system, standardized databases where used and also proper videos are obtained trying to emulate real-world scenarios, where the quality and the distance are factors that have demonstrated challenges for the detection with computer vision and machine learning. Luis Brandon Garcia-Ortiz, Gabriel Sanchez-Perez, Aldo Hernandez-Suarez, Jesus Olivares-Mercado, Héctor M. Pérez Meana, José Portillo-Portillo |
SoMeT | 4 |
| 2020 | Comparison of Face Detection and Recognition Algorithms in Real-Time Video
Alejandra Sarahi Sanchez-Moreno, Héctor M. Pérez Meana, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Linda K. Toscano-Medina |
SoMeT | 3 |
| 2018 | Can Twitter API Be Bypassed? A New Methodology for Collecting Chronological Information Without RestrictionsabstractRetrieving information from social networks is a first and primordial step in many data analysis fields such as Natural Language Processing and Machine Learning. Important data science tasks rely on historical data gathering for further predictive results. Recent works use public platforms for collecting public streams of information like Twitter API, which allows querying chronological tweets from periods no longer than three weeks. In this paper, we present Twitter Scrapy, a new methodology for collecting historical tweets from time periods of arbitrary duration using web scraping techniques that bypass Twitter API restrictions. Aldo Hernandez-Suarez, Gabriel Sanchez-Perez, Linda K. Toscano-Medina, Rocio Toscano-Medina, Victor Martinez-Hernandez, Jesus Olivares-Mercado, Héctor M. Pérez Meana, Victor Sanchez |
SoMeT | 6 |
| 2018 | Change Detection for Video Sequences Based on Incremental Subspace LearningabstractThis paper proposes a novel methodology for change detection in video sequences, which consists in the use of projection of the first eigenvector over the current frame in the video sequence. These eigenvectors are computed using the Incremental Principal Component Analysis (IPCA), assuming that the incremental computation of the eigenvalues and eigenvectors is made using the incremental block approach considering only two frames i.e. the past and the current frames in each incremental block. The main contribution of this work, is the use of the idea that the first eigenvector projects the maximum variability in their data matrix and then by using the incremental block of two frames in the IPCA, the maximum variability in those images could be considered as the change between them; such that after the post-processing in the projected matrix, we are able to labeled the change between the past and the current frames. José Portillo-Portillo, Blas Hernandez-Sanabria, Héctor M. Pérez Meana, Gabriel Sanchez-Perez, Linda K. Toscano-Medina, Jesus Olivares-Mercado, Mariko Nakano-Miyatake, Luis Carlos Castro-Madrid, Victor Sanchez-Silva |
SoMeT | 6 |
| 2018 | A view-invariant gait recognition algorithm based on a joint-direct linear discriminant analysis
José Portillo-Portillo, Roberto Leyva, Victor Sanchez, Gabriel Sanchez-Perez, Héctor M. Pérez Meana, Jesus Olivares-Mercado, Linda K. Toscano-Medina, Mariko Nakano-Miyatake |
Appl. Intell. | 6 |
| 2016 | A cheating-prevention mechanism for hierarchical secret-image-sharing using robust watermarking
Angelina Espejel Trujillo, Mariko Nakano-Miyatake, Jesus Olivares-Mercado, Héctor M. Pérez Meana |
Multim. Tools Appl. | 3 |
| 2013 | A sub-block-based eigenphases algorithm with optimum sub-block size
Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Mariko Nakano-Miyatake, Héctor M. Pérez Meana |
Knowl. Based Syst. | 2 |