VLDB 2026 Research / reviewers in the wild / expert
Gibran Benitez-Garcia
dblp:123/4910
· DBLP profile ↗
19ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0003-4945-8314ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TKA-STAGCN: a skeleton-based graph convolutional network with temporal attention for baseball pitch type classification
Sergio Huesca-Flores, Gibran Benitez-Garcia, Oswaldo Juarez-Sandoval, Hiroki Takahashi, Mariko Nakano-Miyatake |
Multim. Tools Appl. | 2 |
| 2025 | Automatic and Interactive Annotation of Non-manual and Spatial Features in Pidgin Sign Japanese for SLR
Gibran Benitez-Garcia, Nobuko Kato, Yuhki Shiraishi, Hiroki Takahashi |
ACIVS | 1 |
| 2025 | FCR-PoseHRNet: Flexible Feature Realignment and Cross-Resolution Coordinate Refinement in PoseHRNet for 2D Human Pose Estimation
Zakir Ali, Gibran Benitez-Garcia, Hiroki Takahashi |
ACIVS | 2 |
| 2024 | Automated Annotation Assistance for Pidgin Sign Japanese in Sign Language RecognitionabstractJapanese Sign Language (JSL) and Manually Coded Japanese (MCJ) are the two main forms of sign language used in Japan. The former differs significantly from spoken Japanese, while the latter closely aligns with spoken syntax. Pidgin Sign Japanese (PSJ) is an intermediate form that heavily relies on nonmanual signals such as facial expressions and head movements for grammatical nuances. Current sign language recognition (SLR) systems predominantly focus on MCJ, neglecting the challenging properties of PSJ. This paper proposes an annotation assistance tool designed to automate the annotation of non-manual and spatial elements in PSJ. Our tool significantly reduces the manual effort required for annotation by using state-of-the-art methods for tracking human pose, hand, and face landmarks, along with recognizing facial action units (FAUs). Validation on a preliminary dataset of 30 videos containing over 90 instances of nonmanual elements demonstrated a $\mathbf{4 0 \%}$ reduction in annotation time, highlighting our proposal’s efficiency and effectiveness in handling the complexities of PSJ. Gibran Benitez-Garcia, Nobuko Kato, Yuhki Shiraishi, Hiroki Takahashi |
CW | 1 |
| 2024 | PFMNet: Face Mask Recognition with Deformable Convolution Networks and Category AttentionabstractThe challenges posed by the COVID-19 pandemic underscored the critical importance of proper mask usage, highlighting the need for automated systems to monitor face mask-wearing conditions. In this paper, we introduce PFMNet, a novel architecture for recognizing the wearing status of face masks. PFMNet is inspired by the InternImage architecture and employs Deformable Convolution Networks (DCNs) to capture long-range dependencies crucial for accurate mask status determination. The significant challenge of class imbalance, particularly the scarcity of improperly worn mask samples, is addressed by integrating the Category Attention Block (CAB). CAB improves distinct regions, diversifies feature representations, and utilizes efficient global pooling to identify crucial areas, such as the human face, while reducing the computational cost. The performance of PFMNet was assessed using the publicly available PWMFD dataset, which had to be refined due to duplicate images and incorrect annotations. PFMNet was compared to three other state-of-the-art models: InternImage, ConvNext, and EfficientNet. It outperformed these models, achieving an accuracy of 99.39%. This places it ahead of the second-best model by a margin of 0.45%. The confusion matrices illustrate that PFMNet outperforms other models in all classes, particularly excelling in the “with mask” and “without mask” categories, resulting in the best overall performance. Ulises Arroyo-Rojas, Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Hiroki Takahashi |
SoMeT | 2 |
| 2024 | Multimodal Hand Gesture Recognition Using Automatic Depth and Optical Flow Estimation from RGB VideosabstractTraditional Hand Gesture Recognition (HGR) approaches often rely on multiple sensors, such as RGB, depth, and infrared cameras, to capture comprehensive multimodal data. However, this increases hardware complexity and costs, limiting the widespread adoption of HGR systems. In this paper, we propose a Multi-modal HGR approach that leverages automatic depth estimation from RGB videos to enhance HGR performance while using only a single RGB camera. Our method integrates synthetic depth features, optical flow (OF), and RGB data through an early fusion strategy. We conduct extensive experiments using three ConvNet-based models for HGR: the 3D-CNN variants of ResNet and ResNeXt, as well as the efficient 2D-CNN-based Temporal Shift Module (TSM). Our findings indicate that the multimodal input combination of synthetic Depth, OF, and RGB modalities results in superior performance compared to models using solely RGB or RGB+OF inputs, with the ResNeXt-101 model exhibiting the highest accuracy. To validate our approach, we employ the IPN Hand dataset, which we have meticulously refined to correct temporal annotation inconsistencies and increased the number of gesture classes from 13 to 14 by separating the similar dynamics of four specific gestures. Furthermore, we compute high-quality OF and depth maps for all 800,000 frames in the dataset. The enhanced and multimodal data of the IPN Hand dataset will be soon available at github.com/GibranBenitez/IPN-hand. Gibran Benitez-Garcia, Hiroki Takahashi |
SoMeT | 1 |
| 2024 | Optimal Feature Extractor for Video Anomaly Detection in Public Transportation ApplicationsabstractVideo Anomaly Detection (VAD) is a well-established area of research with significant potential for enhancing video surveillance in urban public transportation. However, current VAD systems often propose powerful methodologies but overlook their use in extreme environments like public transportation, necessitating a balance between performance and computational efficiency. In this paper, we evaluate a key component in many VAD frameworks: feature extractors. We investigate five extractors: Inflated 3D ConvNets (I3D), 3D Convolutional Neural Networks (C3D), Unified Transformer (UniFormer) in Small (UniFormer-S) and Base (UniFormer-B) versions, and Temporal Shift Module (TSM). These are integrated into a VAD architecture employing Bidirectional Encoder Representations from Transformers (BERT) with Multiple Instance Learning (MIL), chosen for its modularity and clear separation between the feature extractor and anomaly detector module. UniFormer-S demonstrated a processing rate of 4.64 clips per second with a computational demand of 28.717 GFLOPs on edge devices like the Jetson Orin NX (8GB RAM, 20W power). On the UCF-Crime dataset, UniFormer-S with BERT + MIL achieves an AUC of 79.74%. These findings highlight the promise of UniFormer-S and the use of edge devices like the Jetson Orin NX in public transportation due to their balance of performance and efficiency. Jonathan Flores-Monroy, Gibran Benitez-Garcia, Mariko Nakano-Miyatake, Hiroki Takahashi |
SoMeT | 2 |
| 2024 | Frame-Level Deepfake Detection on Explicit Content with ID-Unaware Binary ClassificationabstractThe rapid advancement in deepfake technology has enabled the creation of highly realistic fake images and videos, posing significant risks, especially in the context of explicit content. Such content, which often involves the alteration of an individual’s identity in sexually explicit material, can lead to defamation, harassment, and blackmail. This paper focuses on the detection of deepfakes in explicit content using a state-of-the-art ID-unaware Binary Classification method. We evaluate its effectiveness in real-world scenarios by analyzing three versions of the model with different backbones: ResNet34, EfficientNet-B3, and EfficientNet-B4. To facilitate this evaluation, we curated a dataset of 200 videos, consisting of 100 genuine videos and their corresponding deepfake counterparts, ensuring a direct comparison between genuine and altered content. Our analysis revealed a significant decrease in detection performance when applying the state-of-the-art method to explicit content. Specifically, the AUC score dropped from 93% on standard datasets such as FaceForensics++ to 62% on our explicit content dataset. Additionally, the accuracy for detecting deepfakes plummeted to around 25%, while the accuracy for genuine videos remained high at approximately 90%. We identified specific factors contributing to this decline, including unconventional makeup, lighting issues, and facial blurring due to camera distance. These findings underscore the challenges and the necessity for robust detection methods to address the unique problems posed by explicit content deepfakes, ultimately aiming to protect individuals from the potential harms associated with this technology. Miguel Jimenez-Martinez, Gibran Benitez-Garcia, Linda K. Toscano-Medina, Jesus Olivares-Mercado |
SoMeT | 2 |
| 2024 | Attention-Based Multi-Scale, Context-Aware Feature Integration into PoseResNet for Coordinate Classification in 2D HPEabstractHuman Pose Estimation (HPE) is a critical Computer Vision task with applications ranging from video surveillance to medical rehabilitation. Despite recent advancements in Deep Learning, HPE still faces challenges such as occluded keypoints, variable lighting conditions, and high computational demands. To address these issues, we present the Attention-Based Multi-Scale, Context-Aware Feature Integration into PoseResNet for Coordinate Classification (AMSF-PRNetCC). Our framework enhances the traditional ResNet architecture by incorporating CoordConv2d layers, depthwise separable convolutions, and novel attention mechanisms including Spatial-Enhanced Channel Attention (SECA) and Squeeze-and-Excitation (SE). We introduce a Context-Aware Feature Pyramid Network (CAFPN) with Dual Mask Global Context Blocks (DMGCB) to efficiently handle multi-scale information. The model culminates in Multi-Layer Perceptron (MLP) stages for precise keypoint coordinate classification. Evaluations on the COCO dataset demonstrate that AMSF-PRNetCC significantly outperforms existing 2D HPE methods in both accuracy and computational efficiency. Our approach achieves state-of-the-art results while requiring fewer computational resources, marking a substantial advancement in the field of HPE. Zakir Ali, Sartaj Ahmed Salman, Gibran Benitez-Garcia, Hiroki Takahashi |
SoMeT | 3 |
| 2023 | Efficient 3Dconv Fusion of RGB and Optical Flow for Dynamic Hand Gesture Recognition and Localization
Gibran Benitez-Garcia, Hiroki Takahashi |
PSIVT | 1 |
| 2022 | TFM a Dataset for Detection and Recognition of Masked Faces in the WildabstractDroplet transmission is one of the leading causes of the spread of respiratory infections, such as coronavirus disease (COVID-19). The proper use of face masks is an effective way to prevent the transmission of such diseases. Nonetheless, different types of masks provide various degrees of protection. Hence, automatic recognition of face mask types may benefit the control access to facilities where a specific protection degree is required. In the last two years, several deep learning models have been proposed for face mask detection and properly wearing mask recognition. However, the current publicly available datasets do not consider the different mask types and occasionally lack real-world elements needed to train robust models. In this paper, we introduce a new dataset named TFM with sufficient size and variety to train and evaluate deep learning models for face mask detection and recognition. This dataset contains more than 135,000 annotated faces from about 100,000 photographs taken in the wild. We consider four mask types (cloth, respirators, surgical and valved) as well as unmasked faces, of which up to six can appear in a single image. The photographs were mined from Twitter within two years since the beginning of the COVID-19 pandemic. Thus, they include diverse scenes with real-world variations in background and illumination. With our dataset, the performance of four state-of-the-art object detection models is evaluated. The experimental results show that YOLOv5 can achieve about 90% of [email protected], demonstrating that the TFM dataset can be used to train robust models and may help the community step forward in detecting and recognizing masked faces in the wild. Our dataset and pre-trained models used in the evaluation will be available upon the publication of this paper. Gibran Benitez-Garcia, Hiroki Takahashi, Miguel Jimenez-Martinez, Jesus Olivares-Mercado |
MMAsia | 1 |
| 2022 | Twitter Face Image Mining for Recognition of Different Face Mask TypesabstractIn the current pandemic of coronavirus disease (COVID-19), an effective way to prevent the transmission and infection of the virus is the proper use of face masks. However, the different types of masks provide different degrees of protection. For instance, valved masks protect the user but do not help to stop the transmission. Hence, the automatic recognition of face mask types may benefit applications that control access to facilities where a certain facepiece is required. In this paper, we propose a Twitter mining framework to gather a large-scale dataset of masked faces suitable to train deep learning-based models for face mask recognition. We employ a keyword-based selection where non-face images are discarded by an efficient face detector (Retinaface). Finally, we train a state-of-the-art CNN architecture (ConvNeXt) for recognizing the wearing mask. We also present a brief analysis of more than two million image-based tweets acquired over two years since the beginning of the pandemic. The code of the proposed framework and a preliminary dataset of more than 10K faces (manually annotated into unmasked, surgical, cloth, respirators, and valved masks) are available on github.com/GibranBenitez/FaceMask Twitter. Ulises Arroyo-Rojas, Miguel Jimenez-Martinez, Gibran Benitez-Garcia, Jesus Olivares-Mercado, Hiroki Takahashi |
SoMeT | 3 |
| 2022 | Continuous Finger Gesture Spotting and Recognition Based on Similarities Between Start and End FramesabstractTouchless in-car devices controlled by single and continuous finger gestures can provide comfort and safety on driving while manipulating secondary devices. Recognition of finger gestures is a challenging task due to (i) similarities between gesture and non-gesture frames, and (ii) the difficulty in identifying the temporal boundaries of continuous gestures. In addition, (iii) the intraclass variability of gestures’ duration is a critical issue for recognizing finger gestures intended to control in-car devices. To address difficulties (i) and (ii), we propose a gesture spotting method where continuous gestures are segmented by detecting boundary frames and evaluating hand similarities between thestartandendboundaries of each gesture. Subsequently, we introduce a gesture recognition based on a temporal normalization of features extracted from the set of spotted frames, which overcomes difficulty (iii). This normalization enables the representation of any gesture with the same limited number of features. We ensure real-time performance by proposing an approach based on compact deep neural networks. Moreover, we demonstrate the effectiveness of our proposal with a second approach based on hand-crafted features performing in real-time, even without GPU requirements. Furthermore, we present a realistic driving setup to capture a dataset of continuous finger gestures, which includes more than 2,800 instances on untrimmed videos covering safety driving requirements. With this dataset, our both approaches can run at 53 fps and 28 fps on GPU and CPU, respectively, around 13 fps faster than previous works, while achieving better performance (at least 5% higher mean tIoU). Gibran Benitez-Garcia, Muhammad Haris 0002, Yoshiyuki Tsuda, Norimichi Ukita |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | FASSD-Net: Fast and Accurate Real-Time Semantic Segmentation for Embedded SystemsabstractRecent works of real-time semantic segmentation, remove or make use of light decoders from dense deep neural networks to achieve fast inference speed. This strategy helps to achieve real-time performance; however, the accuracy is significantly compromised in comparison to non-real-time methods. In this paper, we introduce two key modules aimed to design a high-performance decoder for real-time semantic segmentation, which also reduces the accuracy gap between real-time and non-real-time networks. The first module, Dilated Asymmetric Pyramidal Fusion (DAPF), is designed to increase the receptive field on the top of the last stage of the encoder, obtaining richer contextual features. The second module, Multi-resolution Dilated Asymmetric (MDA) module, fuses and refines detail and contextual information from multi-scale feature maps coming from early and deeper stages of the network. Both modules are designed to keep a low computational complexity by using asymmetric convolutions. With these modules, we propose a network entitled “FASSD-Net,” which is based on a light-weight CNN backbone. Running on a single Nvidia GTX 1080Ti, our model reaches 77.5% and 69.3% of mIoU, at 41 and 80 FPS on the Cityscapes and CamVid datasets, respectively. We present an extensive analysis of the accuracy-speed tradeoffs of three FASSD-Net variations on different embedded systems, demonstrating that a light version of our network can run on the low-power consumption Jetson Xavier NX, at 32 FPS reaching 74% of mIoU with full resolution ($1024\times 2048$). The source code and pre-trained models are available at github.com/GibranBenitez/FASSD-Net. Leonel Rosas-Arias, Gibran Benitez-Garcia, José Portillo-Portillo, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Keiji Yanai |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Masked Batch Normalization to Improve Tracking-Based Sign Language Recognition Using Graph Convolutional NetworksabstractSign language recognition is a fundamental technique to improve communication between native signers and speakers. Current state-of-the-art sign language recognition methods often apply deep neural network models to learn an optimized projection between sign language videos and sentences in an end-to-end manner. Generally, minibatch training using sequential data requires the addition of padding to equalize the varying lengths of sequences. However, this training strategy induces performance degradation if batch normalization is used in the models because batch normalization assumes the validity of all inputs. In this study, we propose masked batch normalization, which normalizes input features while masking dummy signals. We apply masked batch normalization to tracking-based sign language recognition models using graph convolutional networks. The performance of the proposed method is evaluated in isolated sign language word recognition and continuous sign language words recognition settings. To evaluate the proposed method, we use two types of sign language video datasets, WLASL including 2000 types of isolated words, and a JSL dataset including 275 types of videos of isolated words and 113 types of videos showing sentences. The evaluation results show that the proposed method improves the tracking-based sign language recognition models in both cases. Natsuki Takayama, Gibran Benitez-Garcia, Hiroki Takahashi |
FG | 2 |
| 2021 | Similarity Learning for CNN-Based ASL Alphabet RecognitionabstractSign language is an important communication way to convey information among the deaf community, and it is primarily used by people who have hearing or speech impairments. Besides, sign language represents a direct Human-Computer-Interaction (HCI) similar to voice commands. Therefore, the purpose of this study is to investigate and develop a system for American Sign Language (ASL) alphabet recognition using convolutional neural networks. Our proposal is based on semantic similarity learning using Siamese Convolutional Neural Network to reduce the intra-class variation and inter-class similarity among sign images in a Euclidean space. The results of the siamese architecture applied to the ASL alphabet dataset outperform previous works found in the literature. From these results, using t-SNE visualization, we demonstrate that our hypothesis is correct; the ASL recognition improves when increasing the similarity among encoding of the images belonging to the same class and reducing it otherwise. Atoany N. Fierro-Radilla, Karina Perez-Daniel, Gibran Benitez-Garcia, Pedro Najera Garcia, Ramona Fuentes Valdez |
SoMeT | 3 |
| 2020 | IPN Hand: A Video Dataset and Benchmark for Real-Time Continuous Hand Gesture RecognitionabstractContinuous hand gesture recognition (HGR) is an essential part of human-computer interaction with a wide range of applications in the automotive sector, consumer electronics, home automation, and others. In recent years, accurate and efficient deep learning models have been proposed for HGR. However, in the research community, the current publicly available datasets lack real-world elements needed to build responsive and efficient HGR systems. In this paper, we introduce a new benchmark dataset named IPN Hand with sufficient size, variety, and real-world elements able to train and evaluate deep neural networks. This dataset contains more than 4,000 gesture samples and 800,000 RGB frames from 50 distinct subjects. We design 13 different static and dynamic gestures focused on interaction with touchless screens. We especially consider the scenario when continuous gestures are performed without transition states, and when subjects perform natural movements with their hands as non-gesture actions. Gestures were collected from about 30 diverse scenes, with real-world variation in background and illumination. With our dataset, the performance of three 3D-CNN models is evaluated on the tasks of isolated and continuous realtime HGR. Furthermore, we analyze the possibility of increasing the recognition accuracy by adding multiple modalities derived from RGB frames, i.e., optical flow and semantic segmentation, while keeping the real-time performance of the 3D-CNN model. Our empirical study also provides a comparison with the publicly available nvGesture (NVIDIA) dataset. The experimental results show that the state-of-the-art ResNext-101 model decreases about 30% accuracy when using our real-world dataset, demonstrating that the IPN Hand dataset can be used as a benchmark, and may help the community to step forward in the continuous HGR. Our dataset and pre-trained models used in the evaluation are publicly available at github.com/GibranBenitez/IPN-hand. Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Keiji Yanai |
ICPR | 1 |
| 2020 | Fast and Accurate Real-Time Semantic Segmentation with Dilated Asymmetric ConvolutionsabstractRecent works have shown promising results applied to real-time semantic segmentation tasks. To maintain fast inference speed, most of the existing networks make use of light decoders, or they simply do not use them at all. This strategy helps to maintain a fast inference speed; however, their accuracy performance is significantly lower in comparison to non-real-time semantic segmentation networks. In this paper, we introduce two key modules aimed to design a high-performance decoder for real-time semantic segmentation for reducing the accuracy gap between real-time and non-real-time segmentation networks. Our first module, Dilated Asymmetric Pyramidal Fusion (DAPF), is designed to substantially increase the receptive field on the top of the last stage of the encoder, obtaining richer contextual features. Our second module, Multi-resolution Dilated Asymmetric (MDA) module, fuses and refines detail and contextual information from multi-scale feature maps coming from early and deeper stages of the network. Both modules exploit contextual information without excessively increasing the computational complexity by using asymmetric convolutions. Our proposed network entitled “FASSD-Net” reaches 78.8 % of mIoU accuracy on the Cityscapes validation dataset at 41.1 FPS on full resolution images (1024 x 2048). Besides, with a light version of our network, we reach 74.1 % of mIoU at 133.1 FPS (full resolution) on a single NVIDIA GTX 1080Ti card with no additional acceleration techniques. The source code and pre-trained models are available at github.com/GibranBenitez/FASSD- Net. Leonel Rosas-Arias, Gibran Benitez-Garcia, José Portillo-Portillo, Gabriel Sanchez-Perez, Keiji Yanai |
ICPR | 2 |
| 2013 | A sub-block-based eigenphases algorithm with optimum sub-block size
Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Mariko Nakano-Miyatake, Héctor M. Pérez Meana |
Knowl. Based Syst. | 1 |