VLDB 2026 Research / reviewers in the wild / expert
Jonathan Flores-Monroy
dblp:309/7273
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0002-2467-3600ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Optimal Feature Extractor for Video Anomaly Detection in Public Transportation ApplicationsabstractVideo Anomaly Detection (VAD) is a well-established area of research with significant potential for enhancing video surveillance in urban public transportation. However, current VAD systems often propose powerful methodologies but overlook their use in extreme environments like public transportation, necessitating a balance between performance and computational efficiency. In this paper, we evaluate a key component in many VAD frameworks: feature extractors. We investigate five extractors: Inflated 3D ConvNets (I3D), 3D Convolutional Neural Networks (C3D), Unified Transformer (UniFormer) in Small (UniFormer-S) and Base (UniFormer-B) versions, and Temporal Shift Module (TSM). These are integrated into a VAD architecture employing Bidirectional Encoder Representations from Transformers (BERT) with Multiple Instance Learning (MIL), chosen for its modularity and clear separation between the feature extractor and anomaly detector module. UniFormer-S demonstrated a processing rate of 4.64 clips per second with a computational demand of 28.717 GFLOPs on edge devices like the Jetson Orin NX (8GB RAM, 20W power). On the UCF-Crime dataset, UniFormer-S with BERT + MIL achieves an AUC of 79.74%. These findings highlight the promise of UniFormer-S and the use of edge devices like the Jetson Orin NX in public transportation due to their balance of performance and efficiency. Jonathan Flores-Monroy, Gibran Benitez-Garcia, Mariko Nakano-Miyatake, Hiroki Takahashi |
SoMeT | 1 |
| 2024 | Voice Gender Recognition Under Unconstrained Environments Using Fine-Tuned CNNsabstractAutomatic voice gender recognition (VGR) offers several real-world applications, including recommender system, human-robot interaction, and forensic application. VGR systems become challenging when these operate under unconstrained environments. In this study, we evaluate the performance of VGR systems using different fine-tuned pretrained Convolutional Neural Networks (CNNs), in which the speech signals under unconstrained environments are introduced as input data. First, preprocessing is applied to the original speech signal, which consists of noise attenuation based on low-pass filter and silence part removal based on sound amplitude. Then, the time-frequency features, such as Spectrogram, Mel-Spectrogram and Mel Frequency Cepstral Coefficients (MFCC) are extracted, which are converted into RGB images and processed by CNN models. Our research utilizes the VoxCeleb dataset, which is the largest video-audio dataset recorded under unconstrained environments. The results obtained by several fine-tuned CNN models provide higher accuracy compared with the state-of-the-art techniques on this topic. The best accuracy achieved is 98.58% using fine-tuned MobileNet, which is higher than the best accuracy provided by previous works. Jorge Jorrin-Coz, Mariko Nakano-Miyatake, Jonathan Flores-Monroy, Héctor M. Pérez Meana |
SoMeT | 3 |
| 2022 | Implementation of a CNN-Based Driver Drowsiness and Distraction Detector in Mobile DevicesabstractDrowsiness and driver distraction are considered the main causes of traffic accidents in the world. Considering this situation, this paper proposes two important modifications to our previously proposed driver drowsiness and distraction detector for real-time implementation on handheld mobile devices, such as smartphones. The first modification is due to a large variation in the capacity of mobile devices. To adapt the proposed system to a wide range of mobile devices, we present two automatic threshold calculations, which are used to differentiate driver drowsiness from normal blinking and dangerous driver distraction from normal short-term distraction. The second modification is related to the alarm during a continuous dangerous situation of the driver. We introduce a new algorithm to ensure the continuous activation of the alarm while the dangerous situation continues. These improvements perform as the general algorithm, since when it was implemented in mobile devices with low computational power, as well as in devices that do not have these limitations, the alarm activation times were not affected; On the other hand, it was possible to increase the accuracy originally given by the first system with respect to Ground Truth by almost 25% on average, resulting in alarm activations not being affected to a great extent by the natural errors that the convolutional neural networks (CNN) may cause, these improvements are shown and supported by the implementation in real time through video links provided in this work. Jonathan Flores-Monroy, Mariko Nakano-Miyatake, Héctor M. Pérez Meana, Enrique Escamilla Hernández, Gabriel Sanchez-Perez |
SoMeT | 1 |