EDBT 2026 Demo / reviewers in the wild / expert
Christian Tenllado
dblp:88/6041
· DBLP profile ↗
20ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0003-2348-4741ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 4Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 93% Parallel and multicore computing · 7% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU computing |
0.2 | 2 | 2013 | Polyhedral parallel code generation for CUDA · ACM Trans. Archit. Code Optim. 2013 Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus Lifting · IEEE Trans. Parallel Distributed Syst. 2008 |
Compilers and program optimization › code generation
GPU code generation |
0.2 | 1 | 2013 | Polyhedral parallel code generation for CUDA · ACM Trans. Archit. Code Optim. 2013 |
Compilers and program optimization
parallelizing compiler |
0.2 | 1 | 2013 | Polyhedral parallel code generation for CUDA · ACM Trans. Archit. Code Optim. 2013 |
Image and video processing
wavelet transform |
0.1 | 1 | 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus Lifting · IEEE Trans. Parallel Distributed Syst. 2008 |
GPUs and heterogeneous computing › GPU computing
discrete wavelet transform |
0.1 | 1 | 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus Lifting · IEEE Trans. Parallel Distributed Syst. 2008 |
Parallel and multicore computing › parallel computing
parallel implementation |
0.0 | 1 | 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus Lifting · IEEE Trans. Parallel Distributed Syst. 2008 |
Methods — techniques the papers use, named apart from their topics
polyhedral compilation · 0.3multilevel tiling · 0.3lifting scheme · 0.2filter bank scheme · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | COMPAD: A heterogeneous cache-scratchpad CPU architecture with data layout compaction for embedded loop-dominated applicationsabstractThe growing trend of pervasive computing has consolidated the everlasting need for power efficient devices. The conventional cache subsystem of general-purpose CPUs, while being able to adapt to many use cases, suffers from energy inefficiencies in some scenarios. It is well-known by now in the academic literature that the utilization of a scratchpad memory (SPM) can help reducing the overall energy consumption of embedded systems. This work proposes a hybrid cache-SPM architecture with support logic for semi-transparent data management and spatial locality improvement. Selected data are transferred and stored in the SPM in a compact form using dynamic layout transformation. As a second major contribution, we introduce a methodology to identify memory access sequences that make an inefficient use of the cache, marking them as candidates to be moved to an SPM of constrained space. The methodology does not require access to the source code of the target applications, relying on binary instrumentation and offline profiling. The resulting mapping policies have been tested on a simulated system, showing a mean memory dynamic energy reduction of 43% and a mean speed gain of 13% with a representative benchmark set. Tommaso Marinelli, José Ignacio Gómez, Christian Tenllado, Francky Catthoor |
J. Syst. Archit. | 3 |
| 2022 | Microarchitectural Exploration of STT-MRAM Last-level Cache Parameters for Energy-efficient DevicesabstractAs the technology scaling advances, limitations of traditional memories in terms of density and energy become more evident. Modern caches occupy a large part of a CPU physical size and high static leakage poses a limit to the overall efficiency of the systems, including IoT/edge devices. Several alternatives to CMOS SRAM memories have been studied during the past few decades, some of which already represent a viable replacement for different levels of the cache hierarchy. One of the most promising technologies is the spin-transfer torque magnetic RAM (STT-MRAM), due to its small basic cell design, almost absent static current and non-volatility as an added value. However, nothing comes for free, and designers will have to deal with other limitations, such as the higher latencies and dynamic energy consumption for write operations compared to reads. The goal of this work is to explore several microarchitectural parameters that may overcome some of those drawbacks when using STT-MRAM as last-level cache (LLC) in embedded devices. Such parameters include: number of cache banks, number of miss status handling registers (MSHRs) and write buffer entries, presence of hardware prefetchers. We show that an effective tuning of those parameters may virtually remove any performance loss while saving more than 60% of the LLC energy on average. The analysis is then extended comparing the energy results from calibrated technology models with data obtained with freely available tools, highlighting the importance of using accurate models for architectural exploration. Tommaso Marinelli, José Ignacio Gómez, Christian Tenllado, Manu Perumkunnil Komalan, Mohit Gupta 0004, Francky Catthoor |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | A Comparative Analysis on the Impact of Bank Contention in STT-MRAM and SRAM Based LLCsabstractSpin Transfer Torque Magnetic RAM (STT-MRAM) is being extensively considered as a promising replacement for Last Level Caches (LLC), due to its high density, low leakage and non-volatility. However, writes to STT-MRAM are energy intensive and have a high latency. While the high dynamic energy consumption during writes can be compensated by the low static energy consumption, the high latency results in performance degradation. This work shows that in contrast to SRAM-based LLCs, the performance degradation for STT-MRAM is primarily due to bank contention, when trying to satisfy a read request while the bank is being written. We holistically explore the effects of cache banking and cache contention on energy and performance in the LLC of mobile multicore systems, with in-order cores or with out-of-order cores. The detail of the analysis is enabled by highly accurate cache models, based on a 28nm SRAM industry compiler, and an in-house developed STT-MRAM compiler, which generates full STT-MRAM macro designs with silicon-validated MTJ stack and complete parasitic extraction at the 28nm node. Our results show that there is a clear difference in the energy-performance optimal banking configuration between STT-MRAM caches and SRAM caches. These low contention STT-MRAM cache designs with the optimal number of banks save at least 60% cache energy while losing at most single digit percentages in system performance compared to SRAM cache designs. This show an increased potential of using STT-MRAM as a replacement for SRAM in an LLC. Timon Evenblij, Christian Tenllado, Manu Perumkunnil Komalan, Francky Catthoor, Sushil Sakhare, Peter Debacker, Gouri Sankar Kar, Arnaud Furnémont, Nicolas Bueno, José Ignacio Gómez |
ICCD | 2 |
| 2018 | Main memory organization trade-offs with DRAM and STT-MRAM options based on gem5-NVMain simulation frameworksabstractCurrent main memory organizations in embedded and mobile application systems are DRAM dominated. The ever-increasing gap between today's processor and memory speeds makes the DRAM subsystem design a major aspect of computer system design. However, the limitations to DRAM scaling and other challenges like refresh provide undesired trade-offs between performance, energy and area to be made by architecture designers. Several emerging NVM options are being explored to at least partly remedy this but today it is very hard to assess the viability of these proposals because the simulations are not fully based on realistic assumptions on the NVM memory technologies and on the system architecture level. In this paper, we propose to use realistic, calibrated STT-MRAM models and a well calibrated cross-layer simulation and exploration framework, named SEAT, to better consider technologies aspects and architecture constraints. We will focus on general purpose/mobile SoC multi-core architectures. We will highlight results for a number of relevant benchmarks, representatives of numerous applications based on actual system architecture. The most energy efficient STT-MRAM based main memory proposal provides an average energy consumption reduction of 27% at the cost of 2x the area and the least energy efficient STT-MRAM based main memory proposal provides an average energy consumption reduction of 8% at the around the same area or lesser when compared to DRAM. Manu Perumkunnil Komalan, Hyungrock Oh, Matthias Hartmann, Sushil Sakhare, Christian Tenllado, José Ignacio Gómez, Gouri Sankar Kar, Arnaud Furnémont, Francky Catthoor, Sophiane Senni, David Novo, Abdoulaye Gamatié, Lionel Torres |
DATE | 5 |
| 2018 | A CPU-GPU Parallel Ant Colony Optimization Solver for the Vehicle Routing Problem
Antón Rey, Manuel Prieto 0001, José Ignacio Gómez, Christian Tenllado, J. Ignacio Hidalgo |
EvoApplications | 4 |
| 2017 | Cross-layer design and analysis of a low power, high density STT-MRAM for embedded systemsabstractSTT-MRAM (Spin Transfer Torque Magnetic Random Access Memory) has attracted considerable attention of late since it is the most promising logic compatible nonvolatile memory that is suitable for advanced logic nodes (N28 and beyond) in terms of endurance, speed and power. Embedded STT-MRAM has thus been proposed as a candidate for emerging low standby-power connectivity systems such IoT (Internet-of-Things) and wearables. We utilize the high performance CoFeB based perpendicular MTJ (pMTJ) device to realize a low power and highly dense STT-MRAM array for such systems. This study is carried out on the TSMC 28nm technology node and includes a complete cross-layer design and analysis framework ranging from device modeling to circuit design, layout and system implementation. The process variations and temperature (PT) impact on the MTJ for the STT-MRAM design (and correspondingly the total energy consumption and performance of the system) is also analyzed. We report a ∼85% reduction in the energy consumption compared to the baseline SRAM based system for near negligible performance penalty (<5%). Manu Perumkunnil Komalan, Sushil Sakhare, Trong Huynh Bao, Siddharth Rao, Christian Tenllado, José Ignacio Gómez, Gouri Sankar Kar, Arnaud Furnémont, Francky Catthoor |
ISCAS | 6 |
| 2015 | System level exploration of a STT-MRAM based level 1 data-cache
Manu Perumkunnil Komalan, Christian Tenllado, José Ignacio Gómez, Francisco Tirado, Francky Catthoor |
DATE | 2 |
| 2014 | Adaptive Mapping and Parameter Selection Scheme to Improve Automatic Code Generation for GPUs
Juan Carlos Juega, José Ignacio Gómez, Christian Tenllado, Francky Catthoor |
CGO | 3 |
| 2014 | Feasibility exploration of NVM based I-cache through MSHR enhancementsabstractSRAM based memory systems are plagued by a number of problems like sub-threshold leakage and susceptibility to read/write failure with dynamic voltage scaling schemes or low supply voltage. Non-Volatile Memory (NVM) technologies are being explored extensively nowadays to replace the conventional SRAM memories even for level 1 (L1) caches. These NVMs like Spin Torque Transfer RAM (STT-MRAM), Resistive-RAM (ReRAM) and Phase Change RAM (PRAM) are less hindered by leakage problems with technology scaling and consume lesser area. However, simple replacement of SRAM by NVMs is not a viable option due to their write related issues. The main focus of this paper is the exploration of write delay and write energy issues in a NVM based L1 Instruction cache (I-cache) for an ARM like single core system. We propose a NVM I-cache and extend its MSHR (Miss Status Handling Register) functionality to address the NVMs write related issues. According to our simulations, appropriate tuning of selective architecture parameters can reduce the performance penalty introduced by the NVM (∼45%) to extremely tolerable levels (∼1%) and show energy gains up to 35%. Furthermore, on configuring our modified NVM based system to occupy area comparable to the original SRAM-based configuration, it outperforms the SRAM baseline and leads to even more energy savings. Manu Perumkunnil Komalan, José Ignacio Gómez, Christian Tenllado, Praveen Raghavan, Matthias Hartmann, Francky Catthoor |
DATE | 3 |
| 2013 | Multi-level Clustering on Metric Spaces Using a Multi-GPU Platform
Ricardo J. Barrientos, José Ignacio Gómez, Christian Tenllado, Manuel Prieto 0001, Pavel Zezula |
Euro-Par | 3 |
| 2013 | Polyhedral parallel code generation for CUDAabstractThis article addresses the compilation of a sequential program for parallel execution on a modern GPU. To this end, we present a novel source-to-source compiler called PPCG. PPCG singles out for its ability to accelerate computations from any static control loop nest, generating multiple CUDA kernels when necessary. We introduce a multilevel tiling strategy and a code generation scheme for the parallelization and locality optimization of imperfectly nested loops, managing memory and exposing concurrency according to the constraints of modern GPUs. We evaluate our algorithms and tool on the entire PolyBench suite. Sven Verdoolaege, Juan Carlos Juega, Albert Cohen 0001, José Ignacio Gómez, Christian Tenllado, Francky Catthoor |
ACM Trans. Archit. Code Optim. | 5 |
| 2013 | System-level memory management based on statistical variability compensation for frame-based applicationsabstractProcess variability and dynamic domains increase the uncertainty of embedded systems and force designers to apply pessimistic designs, which become unnecessarily conservative and have a tremendous impact on both performance and energy consumption. In this context, developing uncertainty-aware design methodologies that take both variation at platform and at application level into account becomes a must. These methodologies should mitigate the effects derived from uncertainty, avoiding worst-case assumptions. In this article we propose a comprehensive methodology to tackle two forms of uncertainty: (1) process variation on the memory system, (2) application dynamism. A statistical model has been developed to deal with variability derived from fabrication process, whereas system scenarios are selected to cope with dynamic domains. Both sources of uncertainty are firstly tackled in combination at design time, to be refined later, at setup. As a result, at run time the platform can be successfully adapted to the current application behaviour as well as the current variations. Our simulations show that this methodology provides significant energy savings while still meeting strict timing constraints. Concepción Sanz, José Ignacio Gómez, Christian Tenllado, Manuel Prieto 0001, Francky Catthoor |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2012 | Range Query Processing in a Multi-GPU EnvironmentabstractSimilarity search has been widely studied in the last years, as it can be applied to several fields such as searching by content in multimedia objects, text retrieval or computational biology. These applications usually work on very large databases that are often indexed off-line to enable the acceleration of on-line searches. However, to maintain an acceptable throughput, it is essential to exploit the intrinsic parallelism of the algorithms used for the on-line query solving process, even with indexed databases. Therefore, many strategies have been proposed in the literature to parallelize these algorithms, both on shared and distributed memory multiprocessor systems. Lately, GPUs have also been used to implement brute-force approaches instead of using indexing structures, due to the difficulties introduced by the index in the efficient exploitation of the GPU resources. In this work we propose a Multi-GPU metric-space technique that efficiently exploits index data structures for similarity search in large databases, and show how it outperforms previous OpenMP and GPU brute-force strategies. Furthermore, our analysis covers the effects of the database size and its nature. Ricardo J. Barrientos, José Ignacio Gómez, Christian Tenllado, Manuel Prieto 0001, Mauricio Marín |
ISPA | 3 |
| 2012 | OpenIRS-UCM: an open-source multi-platform for interactive response systemsabstractInteractive Response Systems (IRS) have been gaining acceptance within the educational community in recent years and a clear proof is the growing number of commercial systems available today in the market. However, most solutions are based on systems which are closed, rigid and dependent on proprietary keypad or platform. We have developed OpenIRS-UCM, a free teaching tool for interactive polling that solves these drawbacks. It is an open source software so it allows the development of new functions by anybody. It has a friendly interface that anyone without high computer skills can use. It enables the coexistence of several commercial clickers simultaneously with smart-phones, tablets or other modern electronic devices. It is developed in Java, thus its use is not restricted to systems based on Microsoft Windows and it is independent of any proprietary software. Carlos García 0001, Fernando Castro, José Ignacio Gómez, Christian Tenllado, Daniel Chaver, José Antonio López Orozco |
ITiCSE | 4 |
| 2011 | kNN Query Processing in Metric Spaces Using GPUs
Ricardo J. Barrientos, José Ignacio Gómez, Christian Tenllado, Manuel Prieto 0001, Mauricio Marín |
Euro-Par (1) | 3 |
| 2010 | Improving face recognition by combination of natural and Gabor faces
Christian Tenllado, José Ignacio Gómez, Javier Setoain, Darío Mora, Manuel Prieto 0001 |
Pattern Recognit. Lett. | 1 |
| 2009 | Endmember Extraction from Hyperspectral Imagery using a Parallel Ensemble Approach with Consensus AnalysisabstractWe have explored in this paper a framework to test in a quantitative manner the stability of different endmember extraction and spectral unmixing algorithms based on the concept of Consensus Clustering. The idea is to investigate if the sensibility of those algorithms to the number of endmembers can be used to estimate this parameter itself. Preliminary results on synthetic data reveal that the proposed scheme, which can be implemented efficiently in parallel, can compete with state-of-the-art schemes. Fermin Ayuso, Javier Setoain, Manuel Prieto 0001, Christian Tenllado, Francisco Tirado, Javier Plaza, Antonio Plaza |
IGARSS (5) | 4 |
| 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus LiftingabstractThe widespread usage of the DiscreteWaveletTransform (DWT) has motivated the development of fastDWT algorithms and their tuning on all sorts of computersystems. Several studies have compared the performanceof the most popular schemes, known as Filter Bank(FBS) and Lifting (LS), and have always concluded thatLifting is the most efficient option. However, there isno such study on streaming processors such as modernGraphic Processing Units (GPUs). Current trends havetransformed these devices into powerful stream processorswith enough flexibility to perform intensive and complexfloating-point calculations. The opportunities opened upby these platforms, as well as the growing popularityof the DWT within the computer graphics field, make anew performance comparison of great practical interest.Our study indicates that FBS outperforms LS in currentgeneration GPUs. In our experiments, the actual FBS gainsrange between 10% and 140%, depending on the problemsize and the type and length of the wavelet filter. Moreover,design trends suggest higher gains in future generationGPUs. Christian Tenllado, Javier Setoain, Manuel Prieto 0001, Luis Piñuel, Francisco Tirado |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2007 | Parallel Morphological Endmember Extraction Using Commodity Graphics HardwareabstractSpatial/spectral algorithms have been shown in previous work to be a promising approach to the problem of extracting image end members from remotely sensed hyperspectral data. Such algorithms map nicely on high-performance systems such as massively parallel clusters and networks of computers. Unfortunately, these systems are generally expensive and difficult to adapt to onboard data processing scenarios, in which low-weight and low-power integrated components are highly desirable to reduce mission payload. An exciting new development in this context is the emergence of graphics processing units (GPUs), which can now satisfy extremely high computational requirements at low cost. In this letter, we propose a GPU-based implementation of the automated morphological end member extraction algorithm, which is used in this letter as a representative case study of joint spatial/spectral techniques for hyperspectral image processing. The proposed implementation is quantitatively assessed in terms of both end member extraction accuracy and parallel efficiency, using two generations of commercial GPUs from NVidia. Combined, these parts offer a thoughtful perspective on the potential and emerging challenges of implementing hyperspectral imaging algorithms on commodity graphics hardware. Javier Setoain, Manuel Prieto 0001, Christian Tenllado, Antonio Plaza, Francisco Tirado |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2002 | -D Wavelet Transform Enhancement on General-Purpose Microprocessors: Memory Hierarchy and SIMD Parallelism Exploitation
Daniel Chaver, Christian Tenllado, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
HiPC | 2 |