Edwin Arkel Rios

dblp:284/0597 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-3020-0719ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Switchable-Precision Universally Slimmable Networks
Chi-Jui Chen, Edwin Arkel Rios, Bo-Cheng Lai
IEEE Signal Process. Lett.2
2025 Cross-Layer Cache Aggregation for Token Reduction in Ultra-Fine-Grained Image Recognition
abstract
Ultra-fine-grained image recognition (UFGIR) is a challenging task that involves classifying images within a macro-category. While traditional FGIR deals with classifying different species, UFGIR goes beyond by classifying sub-categories within a species such as cultivars of a plant. In recent times the usage of Vision Transformer-based backbones has allowed methods to obtain outstanding recognition performances in this task but this comes at a significant cost in terms of computation specially since this task significantly benefits from incorporating higher resolution images. Therefore, techniques such as token reduction have emerged to reduce the computational cost. However, dropping tokens leads to loss of essential information for fine-grained categories, specially as the token keep rate is reduced. Therefore, to counteract the loss of information brought by the usage of token reduction we propose a novel Cross-Layer Aggregation Classification Head and a Cross-Layer Cache mechanism to recover and access information from previous layers in later locations. Extensive experiments covering more than 2000 runs across diverse settings including 5 datasets, 9 backbones, 7 token reduction methods, 5 keep rates, and 2 image sizes demonstrate the effectiveness of the proposed plug-and-play modules and allow us to push the boundaries of accuracy vs cost for UFGIR by reducing the kept tokens to extremely low ratios of up to 10% while maintaining a competitive accuracy to state-of-the-art models. Code is available at: https://github.com/arkel23/CLCA
Edwin Arkel Rios, Jansen Christopher Yuanda, Vincent Leon Ghanz, Cheng-Wei Yu, Bo-Cheng Lai, Min-Chun Hu 0001
ICASSP1
2025 Global-Local Similarity for Efficient Fine-Grained Image Recognition with Vision Transformers
abstract
Fine-Grained recognition involves the classification of images from subordinate macro-categories, and it is challenging due to small inter-class differences. To overcome this, most methods perform discriminative feature selection enabled by a feature extraction backbone followed by a high-level feature refinement step. Recently, many studies have shown the potential behind vision transformers as a backbone for fine-grained recognition, but their usage of its attention mechanism to select discriminative tokens can be computationally expensive. In this work, we propose a novel and computationally inexpensive metric to identify discriminative regions in an image. We compare the similarity between the global representation of an image given by the CLS token, a learnable token used by transformers for classification, and the local representation of individual patches. We select the regions with the highest similarity to obtain crops, which are forwarded through the same transformer encoder. Finally, high-level features of the original and cropped representations are further refined together in order to make more robust predictions. We demonstrate the effectiveness of our proposed method through comprehensive experiments, obtaining superior accuracy across multiple datasets. Code and checkpoints are available at: https://github.com/arkel23/GLSim.
Edwin Arkel Rios, Min-Chun Hu 0001, Bo-Cheng Lai
ISCAS1
2022 Anime Character Recognition using Intermediate Features Aggregation
abstract
In this work we study the problem of anime character recognition. Anime, refers to animation produced within Japan and work derived or inspired from it. We propose a novel Intermediate Features Aggregation classification head, which helps smooth the optimization landscape of Vision Transformers (ViTs) by adding skip connections between intermediate layers and the classification head, thereby improving relative classification accuracy by up to 28%. The proposed model, named as Animesion, is the first end-to-end framework for large-scale anime character recognition. We conduct extensive experiments using a variety of classification models, including CNNs and self-attention based ViTs. We also adapt its multimodal variation Vision-Language Transformer (ViLT), to incorporate external tag data for classification, without additional multimodal pre-training. Through our results we obtain new insights into the effects of how hyperparameters such as input sequence length, mini-batch size, and variations on the architecture, affect the transfer learning performance of Vi(L)Ts.
Edwin Arkel Rios, Min-Chun Hu 0001, Bo-Cheng Lai
ISCAS1
2022 DLPrPPG: Development and Design of Deep Learning Platform for Remote Photoplethysmography
abstract
This paper presents a comprehensive neural network-based development platform for remote photoplethysmography (rPPG). rPPG is a growing and popular research area, especially with the introduction of deep learning methods that can significantly improve its signal quality and heart rate prediction reliability. However, there are still many problems with the experimental methods in current studies, such as non-standardized and private data, different pre-processing methods, and incomplete or irreproducible experiment methodologies, among others. These problems prevent methods from being compared fairly and lead to lower reliability of the proposed experimental results, hindering progress in this area. For these reasons, we propose an open-source framework to facilitate the design and experimentation of deep learning-based rPPG development, and it’s made freely available on GitHub(DLPrPPG). Through our platform we provide ready-to-use implementations of CNN-AE, LSTM, GAN, and Transformer models, whose hyperparameters we can easily and quickly optimize, and efficiently compare in a fair fashion. From our experiments we show that if the parameters of different neural networks are optimized, the performance of older architectures can be on par or even outperform newer ones.
Bo-Rong Yan, Edwin Arkel Rios, Wen-Hsien Lee, Bo-Cheng Lai
ISCAS2
2021 Parametric Study of Performance of Remote Photopletysmography System
abstract
Remote photoplethysmography (RPPG) is a technique in which we measure sub-cutaneous variations in blood flow, usually through a camera, to obtain physiological signals. Studies involving RPPG have increased in the past few years due to its numerous applications including remote healthcare, anti-spoofing, among others. While there have been many studies on how to increase RPPG's accuracy of bio-markers predictions in a variety of settings, most of them are usually done using workstation computers, yet some of the most promising applications of RPPG probably would be on limited resources, low-power embedded systems. Therefore, we did an extensive study on the effects of one of the most important design parameters in RPPG systems, sliding window (SW) size, for a variety of algorithms, in order to quantify the trade-off between computational cost in time and accuracy in root-mean-squared-error (RMSE), using a standardized public database. We also studied how different face detection and region-of-interest selection affected these results. Finally, based on these, we came up with a new and simple metric that takes into account both computation and accuracy, as a means to design dynamic systems which make the best out of the available resources With correct tuning, we can use this metric to reduce computational costs by up to 47%.
Edwin Arkel Rios, Chih-Chieh Lai, Bo-Rong Yan, Bo-Cheng Lai
ISCAS1