EDBT 2026 Demo / reviewers in the wild / expert
Francisco Tirado
dblp:00/3600 · also Francisco Tirado Fernández, José Francisco Tirado Fernández
· DBLP profile ↗
52ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0003-0974-2687ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 9Software engineering, systems software and programming languages · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Integrated circuit design · 18% GPUs and heterogeneous computing · 16% Processor architecture and microarchitecture · 15% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 19 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
out-of-order execution |
0.2 | 2 | 2009 | Replacing Associative Load Queues: A Timing-Centric Approach · IEEE Trans. Computers 2009 DMDC: Delayed Memory Dependence Checking through Age-Based Filtering · MICRO 2006 |
Energy-efficient computing › low-power design
low-power processor design |
0.1 | 2 | 2009 | Replacing Associative Load Queues: A Timing-Centric Approach · IEEE Trans. Computers 2009 DMDC: Delayed Memory Dependence Checking through Age-Based Filtering · MICRO 2006 |
Image and video processing
wavelet transform |
0.1 | 1 | 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus Lifting · IEEE Trans. Parallel Distributed Syst. 2008 |
GPUs and heterogeneous computing › GPU computing
discrete wavelet transform |
0.1 | 1 | 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus Lifting · IEEE Trans. Parallel Distributed Syst. 2008 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus Lifting · IEEE Trans. Parallel Distributed Syst. 2008 |
Integrated circuit design › digital circuit design
arithmetic circuit design |
0.1 | 1 | 2006 | Bit-Parallel Finite Field Multipliers for Irreducible Trinomials · IEEE Trans. Computers 2006 |
Integrated circuit design › finite field arithmetic
finite field multiplier |
0.1 | 1 | 2006 | Bit-Parallel Finite Field Multipliers for Irreducible Trinomials · IEEE Trans. Computers 2006 |
Integrated circuit design › digital circuit design › arithmetic circuit design
parallel multiplier |
0.1 | 1 | 2006 | Bit-Parallel Finite Field Multipliers for Irreducible Trinomials · IEEE Trans. Computers 2006 |
Parallel and multicore computing › locality optimization
data locality optimization |
0.0 | 1 | 2000 | Data Locality Exploitation in the Decomposition of Regular Domain Problems · IEEE Trans. Parallel Distributed Syst. 2000 |
High-performance computing
domain decomposition |
0.0 | 1 | 2000 | Data Locality Exploitation in the Decomposition of Regular Domain Problems · IEEE Trans. Parallel Distributed Syst. 2000 |
High-performance computing
performance optimization |
0.0 | 1 | 2000 | Data Locality Exploitation in the Decomposition of Regular Domain Problems · IEEE Trans. Parallel Distributed Syst. 2000 |
Parallel and multicore computing › parallel computing
parallel implementation |
0.0 | 1 | 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus Lifting · IEEE Trans. Parallel Distributed Syst. 2008 |
Cryptographic primitives and cryptanalysis › finite field arithmetic
binary field arithmetic |
0.0 | 1 | 2006 | Bit-Parallel Finite Field Multipliers for Irreducible Trinomials · IEEE Trans. Computers 2006 |
Energy-efficient computing
power management |
0.0 | 1 | 2006 | DMDC: Delayed Memory Dependence Checking through Age-Based Filtering · MICRO 2006 |
High-performance computing › numerical linear algebra › linear solver › iterative linear solvers
multigrid method |
0.0 | 1 | 1997 | Relationships Between Efficiency and Execution Time of Full Multigrid Methods on Parallel Computers · IEEE Trans. Parallel Distributed Syst. 1997 |
High-performance computing
parallel numerical algorithms |
0.0 | 1 | 1997 | Relationships Between Efficiency and Execution Time of Full Multigrid Methods on Parallel Computers · IEEE Trans. Parallel Distributed Syst. 1997 |
Electronic design automation › high-level synthesis
area estimation |
0.0 | 1 | 1996 | A method for area estimation of data-path in high level synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996 |
Electronic design automation
high-level synthesis |
0.0 | 1 | 1996 | A method for area estimation of data-path in high level synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996 |
Electronic design automation
design space exploration |
0.0 | 1 | 1996 | A method for area estimation of data-path in high level synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996 |
Methods — techniques the papers use, named apart from their topics
lifting scheme · 0.2filter bank scheme · 0.2complexity analysis · 0.1timing-based dependence checking · 0.1hash table · 0.1age-based filtering · 0.1performance analysis · 0.0decomposition comparison · 0.0experimental validation · 0.0analytical modeling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Portability and Performance Assessment of the Non-Negative Matrix Factorization Algorithm with OpenMP and SYCLabstractThe SYCL standard was released to improve code portability across heterogeneous environments. Intel released the oneAPI toolkit, which includes the Data-Parallel C++ (DPC++) compiler which is the Intel’s SYCL implementation. SYCL is designed to use a single source code to target multiple accelerators such as: multi-core CPUs, GPUs and even FPGAs. Additionally, the C/C++ compiler provided in the oneAPI toolkit supports OpenMP which also allows targeting codes on both CPU and GPU devices. In this paper, the performance of SYCL and OpenMP is evaluated using the well-known non-negative matrix factorization (NMF) algorithm. Three different NMF implementations are developed: baseline, SYCL and OpenMP versions to analyze the acceleration on CPU and GPU. Experimental results show that while the two programming models perform almost identically on CPU, on GPU, SYCL outperforms its OpenMP counterpart slightly. Youssef Faqir-Rhazoui, Carlos García 0001, Francisco Tirado |
CLEI | 3 |
| 2020 | Leveraging knowledge-as-a-service (KaaS) for QoS-aware resource management in multi-user video transcoding
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Francisco Tirado |
J. Supercomput. | 4 |
| 2015 | System level exploration of a STT-MRAM based level 1 data-cache
Manu Perumkunnil Komalan, Christian Tenllado, José Ignacio Gómez, Francisco Tirado, Francky Catthoor |
DATE | 4 |
| 2015 | NMF-mGPU: non-negative matrix factorization on multi-GPU systemsabstractBACKGROUND: In the last few years, the Non-negative Matrix Factorization ( NMF ) technique has gained a great interest among the Bioinformatics community, since it is able to extract interpretable parts from high-dimensional datasets. However, the computing time required to process large data matrices may become impractical, even for a parallel application running on a multiprocessors cluster. In this paper, we present NMF-mGPU, an efficient and easy-to-use implementation of the NMF algorithm that takes advantage of the high computing performance delivered by Graphics-Processing Units ( GPUs ). Driven by the ever-growing demands from the video-games industry, graphics cards usually provided in PCs and laptops have evolved from simple graphics-drawing platforms into high-performance programmable systems that can be used as coprocessors for linear-algebra operations. However, these devices may have a limited amount of on-board memory, which is not considered by other NMF implementations on GPU. RESULTS: NMF-mGPU is based on CUDA ( Compute Unified Device Architecture ), the NVIDIA's framework for GPU computing. On devices with low memory available, large input matrices are blockwise transferred from the system's main memory to the GPU's memory, and processed accordingly. In addition, NMF-mGPU has been explicitly optimized for the different CUDA architectures. Finally, platforms with multiple GPUs can be synchronized through MPI ( Message Passing Interface ). In a four-GPU system, this implementation is about 120 times faster than a single conventional processor, and more than four times faster than a single GPU device (i.e., a super-linear speedup). CONCLUSIONS: Applications of GPUs in Bioinformatics are getting more and more attention due to their outstanding performance when compared to traditional processors. In addition, their relatively low price represents a highly cost-effective alternative to conventional clusters. In life sciences, this results in an excellent opportunity to facilitate the daily work of bioinformaticians that are trying to extract biological meaning out of hundreds of gigabytes of experimental information. NMF-mGPU can be used "out of the box" by researchers with little or no expertise in GPU programming in a variety of platforms, such as PCs, laptops, or high-end GPU clusters. NMF-mGPU is freely available at https://github.com/bioinfo-cnb/bionmf-gpu . Edgardo Mejía-Roa, Daniel Tabas-Madrid, Javier Setoain, Carlos García 0001, Francisco Tirado, Alberto D. Pascual-Montano |
BMC Bioinform. | 5 |
| 2015 | Write-Aware Replacement Policies for PCM-Based SystemsabstractThe gap between processor and memory speeds is one of the greatest challenges that current designers face in order to develop more powerful computer systems. In addition, the scalability of the Dynamic Random Access Memory (DRAM) technology is very limited nowadays, leading one to consider new memory technologies as candidates for the replacement of conventional DRAM. Phase-Change Memory (PCM) is currently postulated as the prime contender due to its higher scalability and lower leakage. However, compared with DRAM, PCM also exhibits some drawbacks, like lower endurance or higher dynamic energy consumption and write latency, that need to be mitigated before it can be used as the main memory technology for the next generation of computers. This work addresses the PCM endurance constraint. For this purpose, we present an analysis of conventional cache replacement policies in terms of the amount of writebacks to main memory that they imply and we also propose some new replacement algorithms for the last-level cache (LLC) with the goal of cutting down the write traffic to memory and consequently, to increase PCM lifetime without degrading system performance. In this paper, we target general purpose processors provided with this kind of non-volatile main memory and we exhaustively evaluate our proposed policies in both single- and multi-core environments. Experimental results show that, on average, compared with a conventional Least Recently Used (LRU) algorithm, some of our proposals manage to reduce the amount of writes to main memory up to 20–30% depending on the scenario evaluated, which leads to memory endurance extensions of up to 20–45%, also reducing the energy consumption in the memory hierarchy by up to 9% and hardly degrading performance. Roberto Rodríguez-Rodríguez, Fernando Castro, Daniel Chaver, Rekai González-Alberquilla, Luis Piñuel, Francisco Tirado |
Comput. J. | 6 |
| 2013 | Reducing writes in phase-change memory environments by using efficient cache replacement policiesabstractPhase Change Memory (PCM) is currently postulated as the best alternative for replacing Dynamic Random Access Memory (DRAM) as the technology used for implementing main memories, thanks to its significant advantages such as good scalability and low leakage. However, PCM also presents some drawbacks compared to DRAM, like its lower endurance. This work presents a behavior analysis of conventional cache replacement policies in terms of the amount of writes to main memory. Besides, new last level cache (LLC) replacement algorithms are exposed, aimed at reducing the number of writes to PCM and hence increasing its lifetime, without significantly degrading system performance. Roberto Rodríguez-Rodríguez, Fernando Castro, Daniel Chaver, Luis Piñuel, Francisco Tirado |
DATE | 5 |
| 2013 | Non-negative matrix factorization on low-power architectures: a comparative studyabstractPower consumption is emerging as one of the main concerns in the High Performance Computing (HPC) field. Many bioinformatics applications require HPC techniques and parallel architectures to meet performance requirements, but at the same time they can be severely limited by energy consumption restrictions. In this paper, we perform an empirical study of an optimized implementation of the Nonnegative Matrix Factorization (NMF), that is widely used in many fields of bioinformatics. We target different types of architectures, including general-purpose, low-power embedded processors and specific-purpose architectures like graphics processors and digital signal processors. From our study, we gain insights in both performance and energy consumption for each one of them under given experimental conditions, and conclude that the most appropriate architecture is usually a trade-off between performance and power consumption for a given experiment and dataset. Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Francisco Tirado |
EuroMPI | 5 |
| 2013 | GPU-based acceleration of bio-inspired motion estimation modelabstractSUMMARY In this paper, we describe the specific and efficient implementation of a gradient‐based optical flow model. This scheme was particularized using a validated neuromorphic motion estimation system for the robust extraction of image velocity. This model contains many characteristics that enhanced the capability when compared with other optical flow gradient family algorithms. Our implementation was performed using specific graphic processing units designed in an ad hoc framework for this model, which could be reused in several low‐level machine‐vision approaches. Observed performance results indicate that these accelerators be highly recommended. Furthermore, the throughput obtained in comparison with a general CPU was analyzed for the accurateness of a system built with regard to other optical flow systems. Additionally, several visual examples, commonly used for testing motion estimation sequences, were shown to reveal implementation behavior features. Copyright © 2012 John Wiley & Sons, Ltd. Fermin Ayuso, Guillermo Botella Juan, Carlos García 0001, Manuel Prieto 0001, Francisco Tirado |
Concurr. Comput. Pract. Exp. | 5 |
| 2013 | Low complexity bit-parallel polynomial basis multipliers over binary fields for special irreducible pentanomials
José Luis Imaña, Román Hermida, Francisco Tirado |
Integr. | 3 |
| 2011 | Biclustering and classification analysis in gene expression using Nonnegative Matrix Factorization on multi-GPU systemsabstractA great interest has been given to the Nonnegative Matrix Factorization (NMF) technique due to its ability of extracting highly-interpretable parts from data sets. Gene expression analysis is one of the most popular applications of NMF in Bioinformatics. Nonetheless, its usage is hindered by the computational complexity when processing large data sets. In this paper, we present two parallel implementations of NMF. The first version uses CUDA on a Graphics Processing Unit (GPU). Large input matrices are iteratively blockwise transferred and processed. The second implementation distributes data among multiple GPUs synchronized through MPI (Message Passing Interface). When analyzing large data sets with two and four GPUs, it performs respectively, 2.3 and 4.13 times faster than the single-GPU version. This represents about 120 times faster than a conventional CPU. These super linear speedups are achieved when data portions assigned to each GPU are small enough to be transferred only once. Edgardo Mejía-Roa, Carlos García 0001, José Ignacio Gómez, Manuel Prieto 0001, Francisco Tirado, Rubén Nogales, Alberto D. Pascual-Montano |
ISDA | 5 |
| 2010 | Building efficient multi-threaded search nodesabstractSearch nodes are single-purpose components of large Web search engines and their efficient implementation is critical to sustain thousands of queries per second and guarantee individual query response times within a fraction of a second. Current technology trends indicate that search nodes ought to be implemented as multi-threaded multi-core systems. The straightforward solution that system designers can apply in this case is simply to follow standard practice by deploying one asynchronous thread per active query in the node and attaching each thread to a different core. Each concurrent thread is responsible for sequentially processing a single query at a time. The only potential source of read/write conflicts among threads are the accesses to the different application caches present in the search node. However, new Web applications pose much more demanding requirements in terms of read/write conflicts than recent past applications since now data updates must take place concurrently with query processing. Insisting on the same paradigm of concurrent threads now augmented with a transaction concurrency control protocol is a feasible solution. In this paper we propose a more efficient and much simpler solution which has the additional advantage of enabling a very efficient administration of application caches. We propose performing relaxed bulk-synchronous parallelism at multi-core level. Carolina Bonacic, Carlos García 0001, Mauricio Marín, Manuel Prieto 0001, Francisco Tirado |
CIKM | 5 |
| 2010 | Stack filter: Reducing L1 data cache power consumption
Rodrígo González-Alberquilla, Fernando Castro, Luis Piñuel, Francisco Tirado |
J. Syst. Archit. | 4 |
| 2009 | Endmember Extraction from Hyperspectral Imagery using a Parallel Ensemble Approach with Consensus AnalysisabstractWe have explored in this paper a framework to test in a quantitative manner the stability of different endmember extraction and spectral unmixing algorithms based on the concept of Consensus Clustering. The idea is to investigate if the sensibility of those algorithms to the number of endmembers can be used to estimate this parameter itself. Preliminary results on synthetic data reveal that the proposed scheme, which can be implemented efficiently in parallel, can compete with state-of-the-art schemes. Fermin Ayuso, Javier Setoain, Manuel Prieto 0001, Christian Tenllado, Francisco Tirado, Javier Plaza, Antonio Plaza |
IGARSS (5) | 5 |
| 2009 | Using age registers for a simple load-store queue filtering
Fernando Castro, Daniel Chaver, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
J. Syst. Archit. | 5 |
| 2009 | Replacing Associative Load Queues: A Timing-Centric ApproachabstractOne of the main challenges of modern processor design is the implementation of a scalable and efficient mechanism to detect memory access order violations as a result of out-of-order execution. Traditional age-ordered associative load queues are complex, inefficient, and power hungry. In this paper, we introduce two new dependence checking schemes with different design tradeoffs, but both explicitly rely on timing information as a primary instrument to rule out dependence violation. Our timing-centric designs operate at a fraction of the energy cost of an associative LQ and achieve the same functionality with an insignificant performance impact on average. Studies with parallel benchmarks also show that they are equally effective and efficient in a chip-multiprocessor environment. Fernando Castro, Regana Noor, Alok Garg, Daniel Chaver, Michael C. Huang 0001, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
IEEE Trans. Computers | 8 |
| 2008 | Exploiting Hybrid Parallelism in Web Search Engines
Carolina Bonacic, Carlos García 0001, Mauricio Marín, Manuel Prieto 0001, Francisco Tirado |
Euro-Par | 5 |
| 2008 | Applying speculation techniques to implement functional unitsabstractThis paper justifies the use of estimation and prediction of carries to increase the performance of functional units built with the replication of full adders while keeping a low area penalization. Adders and multipliers are the most representative modules in this group of functional units. The use of these design techniques allows the implementation of modules with performance improvements ranging from 20% to 50% with only an area overheads around 5%. These functional units are suitable for asynchronous circuits but they could also be introduced in synchronous circuits with speculative techniques. The basic idea consists in estimating the carry out from some parts of the functional units, allowing every part to operate independently and in parallel. These modules are connected to build bigger ones. Results from simulations show that for some applications it is possible to make predictions even more accurate that the bit-based estimation. Predictions have also the advantage they can be introduced in the multipliers design, whether estimators cannot. These predictions are similar to the ones used in the branch prediction in a processor. Alberto A. Del Barrio, María C. Molina, Jose Manuel Mendias, Esther Andres Perez, Román Hermida, Francisco Tirado |
ICCD | 6 |
| 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus LiftingabstractThe widespread usage of the DiscreteWaveletTransform (DWT) has motivated the development of fastDWT algorithms and their tuning on all sorts of computersystems. Several studies have compared the performanceof the most popular schemes, known as Filter Bank(FBS) and Lifting (LS), and have always concluded thatLifting is the most efficient option. However, there isno such study on streaming processors such as modernGraphic Processing Units (GPUs). Current trends havetransformed these devices into powerful stream processorswith enough flexibility to perform intensive and complexfloating-point calculations. The opportunities opened upby these platforms, as well as the growing popularityof the DWT within the computer graphics field, make anew performance comparison of great practical interest.Our study indicates that FBS outperforms LS in currentgeneration GPUs. In our experiments, the actual FBS gainsrange between 10% and 140%, depending on the problemsize and the type and length of the wavelet filter. Moreover,design trends suggest higher gains in future generationGPUs. Christian Tenllado, Javier Setoain, Manuel Prieto 0001, Luis Piñuel, Francisco Tirado |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2007 | Parallel Morphological Endmember Extraction Using Commodity Graphics HardwareabstractSpatial/spectral algorithms have been shown in previous work to be a promising approach to the problem of extracting image end members from remotely sensed hyperspectral data. Such algorithms map nicely on high-performance systems such as massively parallel clusters and networks of computers. Unfortunately, these systems are generally expensive and difficult to adapt to onboard data processing scenarios, in which low-weight and low-power integrated components are highly desirable to reduce mission payload. An exciting new development in this context is the emergence of graphics processing units (GPUs), which can now satisfy extremely high computational requirements at low cost. In this letter, we propose a GPU-based implementation of the automated morphological end member extraction algorithm, which is used in this letter as a representative case study of joint spatial/spectral techniques for hyperspectral image processing. The proposed implementation is quantitatively assessed in terms of both end member extraction accuracy and parallel efficiency, using two generations of commercial GPUs from NVidia. Combined, these parts offer a thoughtful perspective on the potential and emerging challenges of implementing hyperspectral imaging algorithms on commodity graphics hardware. Javier Setoain, Manuel Prieto 0001, Christian Tenllado, Antonio Plaza, Francisco Tirado |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2006 | DMDC: Delayed Memory Dependence Checking through Age-Based FilteringabstractOne of the main challenges of modern processor design is the implementation of a scalable and efficient mechanism to detect memory access order violations as a result of out-of-order execution of memory instructions. Traditional CAM-based associative queues can be very slow and energy hungry. In this paper we introduce two new management schemes. The first one is a filtering scheme based on simple age-tracking. This scheme can easily avoid 95-98% of associative load queue (LQ) searches using only a few registers. This translates into significant power savings. More importantly, however, this filtering makes our second scheme, delayed memory dependence checking (DMDC), practical. With a small hash table, DMDC completely avoids the need for an associative LQ and relies on indexing-based checking at the commit phase and hence cuts the energy spent on LQ by an average of 95%. At an average of about 0.3%, the performance impact is negligible. When the energy cost of the increased execution time is factored in, the processor still makes net energy savings of about 3-8%, depending on the configuration and the applications Fernando Castro, Luis Piñuel, Daniel Chaver, Manuel Prieto 0001, Michael C. Huang 0001, Francisco Tirado |
MICRO | 6 |
| 2006 | Biclustering of gene expression data by non-smooth non-negative matrix factorizationabstractBACKGROUND: The extended use of microarray technologies has enabled the generation and accumulation of gene expression datasets that contain expression levels of thousands of genes across tens or hundreds of different experimental conditions. One of the major challenges in the analysis of such datasets is to discover local structures composed by sets of genes that show coherent expression patterns across subsets of experimental conditions. These patterns may provide clues about the main biological processes associated to different physiological states. RESULTS: In this work we present a methodology able to cluster genes and conditions highly related in sub-portions of the data. Our approach is based on a new data mining technique, Non-smooth Non-Negative Matrix Factorization (nsNMF), able to identify localized patterns in large datasets. We assessed the potential of this methodology analyzing several synthetic datasets as well as two large and heterogeneous sets of gene expression profiles. In all cases the method was able to identify localized features related to sets of genes that show consistent expression patterns across subsets of experimental conditions. The uncovered structures showed a clear biological meaning in terms of relationships among functional annotations of genes and the phenotypes or physiological states of the associated conditions. CONCLUSION: The proposed approach can be a useful tool to analyze large and heterogeneous gene expression datasets. The method is able to identify complex relationships among genes and conditions that are difficult to identify by standard clustering algorithms. Pedro Carmona-Saez, Roberto D. Pascual-Marqui, Francisco Tirado, José María Carazo, Alberto D. Pascual-Montano |
BMC Bioinform. | 3 |
| 2006 | bioNMF: a versatile tool for non-negative matrix factorization in biologyabstractBACKGROUND: In the Bioinformatics field, a great deal of interest has been given to Non-negative matrix factorization technique (NMF), due to its capability of providing new insights and relevant information about the complex latent relationships in experimental data sets. This method, and some of its variants, has been successfully applied to gene expression, sequence analysis, functional characterization of genes and text mining. Even if the interest on this technique by the bioinformatics community has been increased during the last few years, there are not many available simple standalone tools to specifically perform these types of data analysis in an integrated environment. RESULTS: In this work we propose a versatile and user-friendly tool that implements the NMF methodology in different analysis contexts to support some of the most important reported applications of this new methodology. This includes clustering and biclustering gene expression data, protein sequence analysis, text mining of biomedical literature and sample classification using gene expression. The tool, which is named bioNMF, also contains a user-friendly graphical interface to explore results in an interactive manner and facilitate in this way the exploratory data analysis process. CONCLUSION: bioNMF is a standalone versatile application which does not require any special installation or libraries. It can be used for most of the multiple applications proposed in the bioinformatics field or to support new research using this method. This tool is publicly available at http://www.dacya.ucm.es/apascual/bioNMF. Alberto D. Pascual-Montano, Pedro Carmona-Saez, Monica Chagoyen, Francisco Tirado, José María Carazo, Roberto D. Pascual-Marqui |
BMC Bioinform. | 4 |
| 2006 | Bit-Parallel Finite Field Multipliers for Irreducible TrinomialsabstractA new formulation for the canonical basis multiplication in the finite fields GF(2/sup m/) based on the use of a triangular basis and on the decomposition of a product matrix is presented. From this algorithm, a new method for multiplication (named transpositional) applicable to general irreducible polynomials is deduced. The transpositional method is based on the computation of 1-cycles and 2-cycles given by a permutation defined by the coordinate of the product to be computed and by the cardinality of the field GF(2/sup m/). The obtained cycles define groups corresponding to subexpressions that can be shared among the different product coordinates. This new multiplication method is applied to five types of irreducible trinomials. These polynomials have been widely studied due to their low-complexity implementations. The theoretical complexity analysis of the corresponding bit-parallel multipliers shows that the space complexities of our multipliers match the best results known to date for similar canonical GF(2/sup m/) multipliers. The most important new result is the reduction, in two of the five studied trinomials, of the time complexity with respect to the best known results. José Luis Imaña, Juan Manuel Sánchez, Francisco Tirado |
IEEE Trans. Computers | 3 |
| 2006 | Low Complexity Bit-Parallel Multipliers Based on a Class of Irreducible PentanomialsabstractIn this paper, we consider the design of bit-parallel canonical basis multipliers over the finite field$GF(2^{m})$generated by a special type ofirreducible pentanomialthat is used as an irreducible polynomial in theAdvanced Encryption Standard(AES). Explicit formulas for the coordinates of the multiplier are given. The main advantage of our design is that some of the expressions obtained are common toanyirreducible polynomial, so our multiplier can be generalized to perform the multiplication overgeneral irreducible polynomials. Moreover, the obtained expressions can be easily converted to parameterizable code using hardware description languages. The theoretical complexity analysis also shows that our bit-parallel multipliers present a reduced number ofxorgates with respect to the best known results found in the literature. José Luis Imaña, Román Hermida, Francisco Tirado |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2005 | Load-Store Queue Management: an Energy-Efficient Design Based on a State-Filtering MechanismabstractModern microprocessors incorporate sophisticated techniques to allow early execution of loads without compromising program correctness. To do so, the structures that hold the memory instructions (load and store queues) implement several complex mechanisms to dynamically resolve the memory-based dependences. Our main objective in this paper is to design an efficient LQ-SQ structure, which saves energy without sacrificing much performance. We propose a new design that divides the load queue into two structures, a conventional associative queue and a simpler FIFO queue that does not allow associative searching. A dependence predictor predicts whether a load instruction has a memory dependence on any inflight store instruction. If so, the load is sent to the conventional associative queue. Otherwise, it is sent to the non-associative queue which can only detect dependence in an inexact and conservative way. In addition, the load will not check the store queue at execution time. These measures combined save energy consumption. We explore different predictor designs and runtime policies. Our experiments indicate that such a design can reduce the energy consumption in the load-store queue by 35-50% with an insignificant performance penalty of about 1%. When the energy cost of the increased execution time is factored in, the processor still makes net energy savings of about 3-4%. Fernando Castro, Daniel Chaver, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado, Michael C. Huang 0001 |
ICCD | 5 |
| 2005 | Energy-aware fetch mechanism: trace cache and BTB customizationabstractA highly-efficient fetch unit is essential not only to obtain good performance but also to achieve energy efficiency. However, existing designs are inflexible and depending on program behavior, can be either insufficient or an overkill. We introduce a phase-based adaptive fetch mechanism that can be dynamically adjusted based on feedback information of the program behavior. This design adds very little hardware complexity and relegates complex tasks to the software components. It is also very effective: saving 26.8% and 34.1% fetch energy on average compared with a conventional and a trace cache-based fetch unit, respectively. At the same time, performance is improved by 5.7% and 0.6%, respectively Daniel Chaver, Miguel A. Rojas, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado, Michael C. Huang 0001 |
ISLPED | 5 |
| 2004 | Adaptive Tuning of Reserved Space in an Appel Collector
José Manuel Velasco, Katzalin Olcoz, Francisco Tirado |
ECOOP | 3 |
| 2003 | Branch prediction on demand: an energy-efficient solutionabstractHigh-end processors typically incorporate complex branch predictors consisting of many large structures that together consume a notable fraction of total chip power (more than 10% in some cases). Depending on the applications, some of these resources may remain underused for long periods of time. We propose a methodology to reduce the energy consumption of the branch predictor by characterizing prediction demand using profiling and dynamically adjusting predictor resources accordingly. Specifically, we disable components of the hybrid direction predictor and resize the branch target buffer. Detailed simulations show that this approach reduces the energy consumption in the branch predictor by an average of 72% and up to 89% with virtually no impact on prediction accuracy and performance. Daniel Chaver, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado, Michael C. Huang 0001 |
ISLPED | 4 |
| 2003 | A parallel multigrid solver for viscous flows on anisotropic structured grids
Manuel Prieto 0001, Rubén S. Montero, Ignacio Martín Llorente, Francisco Tirado |
Parallel Comput. | 4 |
| 2002 | -D Wavelet Transform Enhancement on General-Purpose Microprocessors: Memory Hierarchy and SIMD Parallelism Exploitation
Daniel Chaver, Christian Tenllado, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
HiPC | 5 |
| 2001 | A Multigrid Solver for the Incompressible Navier-Stokes Equations on a Beowulf-Class SystemabstractThis paper presents an efficient parallel multigrid solver for speeding up the computation of a 3-D model that treats the flow of a viscous fluid over a solid obstacle. From a numerical point of view, the main interest of this simulation lies in exhibiting some basic difficulties that prevent optimal multigrid efficiencies from being achieved. As the computing platform we have used Coral, a Beowulf-class system based on a high performance GigaNet cLAN interconnect and Dual Pentium III nodes. The solver which has been devised taking both algorithmic and architectural issues into account, not only attains good convergence rates, but also achieves satisfactory efficiencies on the target cluster. Manuel Prieto 0001, Rubén S. Montero, Ignacio Martín Llorente, Francisco Tirado |
ICPP | 4 |
| 2001 | Parallel Multigrid for Anisotropic Elliptic Equations
Manuel Prieto 0001, R. Santiago, David Espadas, Ignacio Martín Llorente, Francisco Tirado |
J. Parallel Distributed Comput. | 5 |
| 2001 | Analysing value substitution and confidence estimation for value prediction
Luis Piñuel, Rafael A. Moreno, Francisco Tirado |
J. Syst. Archit. | 3 |
| 2000 | Impact of PE Mapping on Cray T3E Message-Passing Performance
Eduardo Huedo, Manuel Prieto 0001, Ignacio Martín Llorente, Francisco Tirado |
Euro-Par | 4 |
| 2000 | A Power Perspective of Value Speculation for Superscalar MicroprocessorsabstractPower consumption has become an important issue in the design of high-performance microprocessors. Pipeline activity in superscalar processors represents one of the most significant sources of power dissipation, due to the growing complexity of the out-of-order issue logic and the extra-work caused by speculative execution. Therefore, analyzing speculative techniques from a power-perspective seems to be essential for the design of power-efficient superscalar processors. Value prediction is one of the most recent techniques for increasing ILP through speculative execution. Although it has been proved to be an effective technique for improving performance, its hardware cost and power dissipation can represent two main limitations. The aim of this paper is to explore the main sources of power dissipation to be considered when value prediction is used, and to propose solutions to reduce this dissipation. Rafael A. Moreno, Luis Piñuel, Silvia Del Pino, Francisco Tirado |
ICCD | 4 |
| 2000 | Data Locality Exploitation in the Decomposition of Regular Domain ProblemsabstractThe aim of this paper is to study the effect of local memory hierarchy and communication network exploitation on message sending and the influence of this effect on the decomposition of regular applications. In particular, we have considered two different parallel computers, a Cray T3E-900 and an SGI Origin 2000. In both systems, the bandwidth reduction due to non-unit-stride memory access is quite significant and could be more important than the reduction due to contention in the network. These conclusions affect the choice of optimal decompositions for regular domains problems. Thus, although traditional 3D decompositions lead to lower inherent communication-to-computation ratios and could exploit more efficiently the interconnection network, lower dimensional decompositions are found to be more efficient due to the data decomposition effects on the spatial locality of the messages to be communicated. This increasing importance of local optimisations has also been shown using a well-known communication-computation overlapping technique which increases execution time, instead of reducing it as we could expect, due to poor cache memory exploitation. Manuel Prieto 0001, Ignacio Martín Llorente, Francisco Tirado |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1999 | Implementation of Hybrid Context Based Value Predictors Using Value Sequence Classification
Luis Piñuel, Rafael A. Moreno, Francisco Tirado |
Euro-Par | 3 |
| 1999 | Message Passing Evaluation and Analysis on Cray T3E and SGI Origin 2000 Systems
Manuel Prieto 0001, David Espadas, Ignacio Martín Llorente, Francisco Tirado |
Euro-Par | 4 |
| 1999 | Solution of Alternating-Line Processes on Modern Parallel ComputersabstractThe aim of this paper is the study of different methods for the solution of alternating-line problems, taking into account the evolution of architectural parameters on modern parallel computers, i.e. processors, memory hierarchy, and interconnection network performance. Three different kinds of solvers are studied: The Pipelined Gaussian Elimination scheme, the Matrix Transposition scheme, and the new one which is presented in this paper: the Mapping Transposition scheme, whose performance clearly betters, in many cases, that obtained by all the other methods, due to its better fitting to the characteristics of modern parallel computers. The experimental results have been obtained on a Cray T3E and on an SGI Origin 2000, up to 512 and 32 processors, respectively. David Espadas, Manuel Prieto 0001, Ignacio Martín Llorente, Francisco Tirado |
ICPP | 4 |
| 1999 | Unified data path allocation and BIST intrusion
Katzalin Olcoz, Francisco Tirado, Hortensia Mecha |
Integr. | 2 |
| 1999 | An environment to develop parallel code for solving partial differential equations based-problems
Manuel Prieto 0001, Ignacio Martín Llorente, Francisco Tirado |
J. Syst. Archit. | 3 |
| 1997 | Relationships Between Efficiency and Execution Time of Full Multigrid Methods on Parallel ComputersabstractThe large number of processing elements in current parallel systems necessitates the development of more comprehensive and realistic tools for the scalability analysis of algorithms on those architectures. This paper presents a simple analytical tool with which to study the scalability of parallel algorithm-architecture combinations. Our practical method studies separately execution time, efficiency, and memory usage in the accuracy-critical scaling model, where the problem size-input data set size-increases with the number of processors, which is the relevant one in many situations. The paper defines quantitative and qualitative measurements of the scalability and derives important relationships between execution time and efficiency. For example, results show that the best way to scale the system (to deteriorate as little as possible the properties of the system) is by maintaining constant execution time. These analytical results are verified with one candidate application for massive parallel computers: the full multigrid method. We study the scalability of a general d-dimensional full multigrid method on an r-dimensional mesh of processors. The analytical expressions are verified through experimental results obtained by implementing the full multigrid method on a Transputer-based machine and on the CRAY T3D. Ignacio Martín Llorente, Francisco Tirado |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1996 | Some Aspects About the Scalability of Scientific Applications on Parallel Architectures
Ignacio Martín Llorente, Francisco Tirado, Luis Vázquez Martínez |
Parallel Comput. | 2 |
| 1996 | A method for area estimation of data-path in high level synthesisabstractThis paper describes a new method to estimate the area of data paths generated during a High Level Synthesis (HLS) process, when the information concerning the circuit is not yet complete. Our method is more accurate and considers more factors than those used by other HLS systems of which we are aware. Our main concern is the interconnection area, often neglected by HLS systems, which has a strong influence on the final circuit area being optimized, as well as a high dependency on the technology used and on the circuit area itself. Predicting the area of a design layout with accuracy is important because it allows one to foresee whether the design will satisfy the area constraints, and will lend the allocator towards the best design among several possibilities with guarantees. Our estimations of the final standard-cell layout area are similar, or even more accurate, than those obtained following methods used by low-level design systems, which have much more information available. Due to the performance penalty their relatively high complexity will produce, these methods are unusable in an HLS system exploring a wide design space. Our estimation, on the contrary, has a low complexity and can be repeated time and again as the HLS design space is searched. Hortensia Mecha, Milagros Fernández, Francisco Tirado, Julio Septién, D. Motes, Katzalin Olcoz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1993 | Guidance for optimization-based synthesis tools
Daniel Mozos, Julio Septién, Francisco Tirado, Milagros Fernández |
Microprocess. Microprogramming | 3 |
| 1993 | Data path structures and heuristics for testable allocation in high level synthesis
Katzalin Olcoz, Francisco Tirado, Daniel Mozos, Julio Septién |
Microprocess. Microprogramming | 2 |
| 1993 | RT-level synthesis
Francisco Tirado |
Microprocess. Microprogramming | 1 |
| 1992 | Design control in a high level synthesis system
Daniel Mozos, Julio Septién, Francisco Tirado, Román Hermida |
Microprocess. Microprogramming | 3 |
| 1991 | A hardware allocator guided by cost functions
J. Septiéna, Daniel Mozos, Román Hermida, Francisco Tirado |
Microprocessing and Microprogramming | 4 |
| 1983 | RDBAS: A relational database system for non-experienced users
Román Hermida, José J. Ruz, Francisco Tirado, Alfredo Bautista |
Microprocessing and Microprogramming | 3 |
| 1982 | An architectural design for simultaneous microdiagnostic
Alfredo Bautista, Francisco Tirado, José J. Ruz, Román Hermida |
Microprocessing and Microprogramming | 2 |
| 1978 | A general purpose computer emulator
Emilio Luque, Lorenzo Moreno Ruiz, Francisco Tirado |
Euromicro Newsletter | 3 |