VLDB 2026 Research / reviewers in the wild / expert
Rowel Atienza
dblp:11/388
· DBLP profile ↗
11ranked-venue papers
7as first author
7since 2021 · last 2023
0000-0002-8830-2534ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 27% 3D vision · 24% Deep learning architectures and training · 24% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › autoregressive model
autoregressive sequence modeling |
0.6 | 1 | 2022 | Scene Text Recognition with Permuted Autoregressive Sequence Models · ECCV (28) 2022 |
Computer vision › Image recognition and object detection
scene text recognition |
0.6 | 1 | 2022 | Scene Text Recognition with Permuted Autoregressive Sequence Models · ECCV (28) 2022 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.3 | 1 | 2018 | Fast Disparity Estimation Using Dense Networks · ICRA 2018 |
Machine learning › Deep learning architectures and training › convolutional neural network
dense network |
0.3 | 1 | 2018 | Fast Disparity Estimation Using Dense Networks · ICRA 2018 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.3 | 1 | 2018 | Fast Disparity Estimation Using Dense Networks · ICRA 2018 |
Computer vision › 3D vision
stereo vision |
0.3 | 1 | 2018 | Fast Disparity Estimation Using Dense Networks · ICRA 2018 |
Machine learning › Generative modeling
autoregressive model |
0.2 | 1 | 2022 | Scene Text Recognition with Permuted Autoregressive Sequence Models · ECCV (28) 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic reasoning |
0.1 | 1 | 2018 | Fast Disparity Estimation Using Dense Networks · ICRA 2018 |
Methods — techniques the papers use, named apart from their topics
permuted autoregressive sequence modeling · 0.6end-to-end training · 0.3dense convolutional network · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | EfficientSpeech: An On-Device Text to Speech ModelabstractState of the art (SOTA) neural text to speech (TTS) models can generate natural-sounding synthetic voices. These models are characterized by large memory footprints and substantial number of operations due to the long-standing focus on speech quality with cloud inference in mind. Neural TTS models are generally not designed to perform standalone speech syntheses on resource-constrained and no Internet access edge devices. In this work, an efficient neural TTS called EfficientSpeech that synthesizes speech on an ARM CPU in real-time is proposed. EfficientSpeech uses a shallow non-autoregressive pyramid-structure transformer forming a U-Network. EfficientSpeech has 266k parameters and consumes 90 MFLOPS only or about 1% of the size and amount of computation in modern compact models such as Mixer-TTS. EfficientSpeech achieves an average mel generation real-time factor of 104.3 on an RPi4. Human evaluation shows only a slight degradation in audio quality as compared to FastSpeech2. Rowel Atienza |
ICASSP | 1 |
| 2023 | Scene Text Recognition Models Explainability Using Local FeaturesabstractExplainable AI (XAI) is the study on how humans can be able to understand the cause of a model’s prediction. In this work, the problem of interest is Scene Text Recognition (STR) Explainability, using XAI to understand the cause of an STR model’s prediction. Recent XAI literatures on STR only provide a simple analysis and do not fully explore other XAI methods. In this study, we specifically work on data explainability frameworks, called attribution-based methods, that explains the important parts of an input data in deep learning models. However, integrating them into STR produces inconsistent and ineffective explanations, because they only explain the model in the global context. To solve this problem, we propose a new method, STRExp, to take into consideration the local explanations, i.e. the individual character prediction explanations. This is then benchmarked across different attribution-based methods on different STR datasets and evaluated across different STR models. Mark Ty, Rowel Atienza |
ICIP | 2 |
| 2023 | Fast Data Augmentation for Scene Text Recognition Using CUDAabstractScene Text Recognition (STR) is a task in computer vision that is used to read texts in natural scene images. STR currently suffers from data distribution shift due to the lack of large real datasets for training. Data augmentation is a method that has been used in multiple studies to address this issue. However, performing augmentation also introduces computational overhead during training. In this paper, we propose FastSTRAug, a CUDA-based library of 36 augmentation functions specifically designed for STR. When executed through varying image sizes, FastSTRAug is observed to be significantly faster over its serial counterpart in most functions, reaching up to 380x speedup on larger images. David Angelo Piscasio, Rowel Atienza |
TENCON | 2 |
| 2022 | Scene Text Recognition with Permuted Autoregressive Sequence Models
Darwin Bautista, Rowel Atienza |
ECCV (28) | 2 |
| 2022 | Depth Pruning with Auxiliary Networks for TinymlabstractPruning is a neural network optimization technique that sacrifices accuracy in exchange for lower computational requirements. Pruning has been useful when working with extremely constrained environments in tinyML. Unfortunately, special hardware requirements and limited study on its effectiveness on already compact models prevent its wider adoption. Depth pruning is a form of pruning that requires no specialized hardware but suffers from a large accuracy falloff. To improve this, we propose a modification that utilizes a highly efficient auxiliary network as an effective interpreter of inter-mediate feature maps. Our results show a parameter reduction of 93% on the MLPerfTiny Visual Wakewords (VWW) task and 28% on the Keyword Spotting (KWS) task with accuracy cost of 0.65% and 1.06% respectively. When evaluated on a Cortex-M0 microcontroller, our proposed method reduces the VWW model size by 4.7× and latency by 1.6× while counter intuitively gaining 1% accuracy. KWS model size on Cortex-M0 was also reduced by 1.2× and latency by 1.2× at the cost of 2.21% accuracy. Josen Daniel De Leon, Rowel Atienza |
ICASSP | 2 |
| 2022 | Improving Model Generalization by Agreement of Learned Representations from Data AugmentationabstractData augmentation reduces the generalization error by forcing a model to learn invariant representations given different transformations of the input image. In computer vision, on top of the standard image processing functions, data augmentation techniques based on regional dropout such as CutOut, MixUp, and CutMix and policy-based selection such as AutoAugment demonstrated state-of-the-art (SOTA) results. With an increasing number of data augmentation algorithms being proposed, the focus is always on optimizing the input-output mapping while not realizing that there might be an untapped value in the transformed images with the same label. We hypothesize that by forcing the representations of two transformations to agree, we can further reduce the model generalization error. We call our proposed method Agreement Maximization or simply AgMax. With this simple constraint applied during training, empirical results show that data augmentation algorithms can further improve the classification accuracy of ResNet50 on ImageNet by up to 1.5%, WideResNet40-2 on CIFAR10 by up to 0.7%, WideResNet40-2 on CIFAR100 by up to 1.6%, and LeNet5 on Speech Commands Dataset by up to 1.4%. Experimental results further show that unlike other regularization terms such as label smoothing, AgMax can take advantage of the data augmentation to consistently improve model generalization by a significant margin. On downstream tasks such as object detection and segmentation on PascalVOC and COCO, AgMax pre-trained models outperforms other data augmentation methods by as much as 1.0mAP (box) and 0.5mAP (mask). Code is available at https://github.com/roatienza/agmax. Rowel Atienza |
WACV | 1 |
| 2021 | Vision Transformer for Fast and Efficient Scene Text Recognition
Rowel Atienza |
ICDAR (1) | 1 |
| 2018 | Fast Disparity Estimation Using Dense NetworksabstractDisparity estimation is a difficult problem in stereo vision because the correspondence technique fails in images with textureless and repetitive regions. Recent body of work using deep convolutional neural networks (CNN) overcomes this problem with semantics. Most CNN implementations use an autoencoder method; stereo images are encoded, merged and finally decoded to predict the disparity map. In this paper, we present a CNN implementation inspired by dense networks to reduce the number of parameters. Furthermore, our approach takes into account semantic reasoning in disparity estimation. Our proposed network, called DenseMapNet, is compact, fast and can be trained end-to-end. DenseMapNet requires 290k parameters only and runs at 30Hz or faster on color stereo images in full resolution. Experimental results show that DenseMapNet accuracy is comparable with other significantly bigger CNN-based methods. Rowel Atienza |
ICRA | 1 |
| 2003 | Interactive skills using active gaze trackingabstractWe have incorporated interactive skills into an active gaze tracking system. Our active gaze tracking system can identify an object in a cluttered scene that a person is looking at. By following the user's 3-D gaze direction together with a zero-disparity filter, we can determine the object's position. Our active vision system also directs attention to a user by tracking anything with both motion and skin color. A Particle Filter fuses skin color and motion from optical flow techniques together to locate a hand or a face in an image. The active vision then uses stereo camera geometry, Kalman Filtering and position and velocity controllers to track the feature in real-time. These skills are integrated together such that they cooperate with each other in order to track the user's face and gaze at all times. Results and video demos provide interesting insights on how active gaze tracking can be utilized and improved to make human-friendly user interfaces. Rowel Atienza, Alexander Zelinsky |
ICMI | 1 |
| 2003 | Intuitive Human-Robot Interaction Through Active 3D Gaze Tracking
Rowel Atienza, Alexander Zelinsky |
ISRR | 1 |
| 2002 | Active Gaze Tracking for Human-Robot InteractionabstractIn our effort to make human-robot interfaces more user-friendly, we built an active gaze tracking system that can measure a person's gaze direction in real-time. Gaze normally tells which object in his/her surrounding a person is interested in. Therefore, it can be used as a medium for human-robot interaction like instructing a robot arm to pick a certain object a user is looking at. We discuss how we developed and put together algorithms for zoom camera calibration, low-level control of active head, face and gaze tracking to create an active gaze tracking system. Rowel Atienza, Alexander Zelinsky |
ICMI | 1 |