EDBT 2026 Demo / reviewers in the wild / expert
José María González-Linares
dblp:27/943
· DBLP profile ↗
23ranked-venue papers
2as first author
3since 2021 · last 2022
0000-0002-0545-5958ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 since 2021Artificial intelligence and machine learning · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | A Hybrid Piece-Wise Slowdown Model for Concurrent Kernel Execution on GPU
Bernabé López-Albelda, Francisco M. Castro, José María González-Linares, Nicolás Guil |
Euro-Par | 3 |
| 2022 | CAVLCU: an efficient GPU-based implementation of CAVLCabstractAbstract CAVLC (Context-Adaptive Variable Length Coding) is a high-performance entropy method for video and image compression. It is the most commonly used entropy method in the video standard H.264. In recent years, several hardware accelerators for CAVLC have been designed. In contrast, high-performance software implementations of CAVLC (e.g., GPU-based) are scarce. A high-performance GPU-based implementation of CAVLC is desirable in several scenarios. On the one hand, it can be exploited as the entropy component in GPU-based H.264 encoders, which are a very suitable solution when GPU built-in H.264 hardware encoders lack certain necessary functionality, such as data encryption and information hiding. On the other hand, a GPU-based implementation of CAVLC can be reused in a wide variety of GPU-based compression systems for encoding images and videos in formats other than H.264, such as medical images. This is not possible with hardware implementations of CAVLC, as they are non-separable components of hardware H.264 encoders. In this paper, we present CAVLCU, an efficient implementation of CAVLC on GPU, which is based on four key ideas. First, we use only one kernel to avoid the long latency global memory accesses required to transmit intermediate results among different kernels, and the costly launches and terminations of additional kernels. Second, we apply an efficient synchronization mechanism for thread-blocks (In this paper, to prevent confusion, a block of pixels of a frame will be referred to as simply block and a GPU thread block as thread-block.) that process adjacent frame regions (in horizontal and vertical dimensions) to share results in global memory space. Third, we exploit fully the available global memory bandwidth by using vectorized loads to move directly the quantized transform coefficients to registers. Fourth, we use register tiling to implement the zigzag sorting, thus obtaining high instruction-level parallelism. An exhaustive experimental evaluation showed that our approach is between 2.5 $$\times$$ × and 5.4 $$\times$$ × faster than the only state-of-the-art GPU-based implementation of CAVLC. Antonio Fuentes-Alventosa, Juan Gómez-Luna, José María González-Linares, Nicolás Guil, Rafael Medina Carnicer |
J. Supercomput. | 3 |
| 2022 | FlexSched: Efficient scheduling techniques for concurrent kernel execution on GPUs
Bernabé López-Albelda, Francisco M. Castro, José María González-Linares, Nicolás Guil |
J. Supercomput. | 3 |
| 2020 | Heuristics for concurrent task scheduling on GPUsabstractSummary Concurrent execution of tasks in GPUs can reduce the computation time of a workload by overlapping data transfer and execution commands. However, it is difficult to implement an efficient runtime scheduler that minimizes the workload makespan as many execution orderings should be evaluated. In this paper, we employ scheduling theory to build a model that takes into account the device capabilities, workload characteristics, constraints, and objective functions. In our model, GPU tasks scheduling is reformulated as a flow shop scheduling problem, which allow us to apply and compare well‐known heuristics already developed in the operations research field. In addition, we develop a new heuristic, specifically focused on executing GPU commands, that achieves better scheduling results than previous ones. It leverages on a precise GPU command execution model for both computation and data transfers to carry out more advantageous scheduling decisions. A comprehensive evaluation, showing the suitability and robustness of this new approach, is conducted in three different NVIDIA architectures (Kepler, Maxwell, and Pascal). Results confirm the proposed heuristic achieves the best results in more than 90% of the experiments. Furthermore, a comparison has been made with MPS (Multi‐Process Service), the NVIDIA API that deals with the execution of concurrent tasks, which shows that our solution obtains speed‐ups ranging from 1.15 to 1.20. Bernabé López-Albelda, A. J. Lázaro-Muñoz, José María González-Linares, Nicolás Guil |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | A tasks reordering model to reduce transfers overhead on GPUs
A. J. Lázaro-Muñoz, José María González-Linares, Juan Gómez-Luna, Nicolás Guil |
J. Parallel Distributed Comput. | 2 |
| 2016 | Configurable XOR Hash Functions for Banked Scratchpad Memories in GPUsabstractScratchpad memories in GPU architectures are employed as software-controlled caches to increase the effective GPU memory bandwidth. Through the use of well-known optimization techniques, such as privatization and tiling, they are properly exploited. Typically, they are banked memories which are addressed with a$\text{mod}(2^N)$bank indexing scheme. Although their bandwidth is fully exploited for linear memory accesses, their performance is burdened when non-unit strides appear in memory access patterns because they provoke bank conflicts. This paper explores the use of configurablebit-vectorandbitwiseXOR-based hash functions to evenly distribute memory addresses of the access patterns over the memory banks, reducing the number of bank conflicts. An exhaustive, but lightweight, search is used to configure bit-vector hash functions. Bitwise hash functions are configured with heuristics. Hardware and software implementations are carried out. For the hardware approach, the experimental results show 24 percent performance speed-up for 22 benchmarks on GPGPU-Sim, a Fermi-like simulator. Bank conflicts are reduced by 96 percent with bit-vector hash functions, and 97 percent with bitwise hash functions using our proposed Minimum Imbalance Heuristic. The software approach, using bit-vector hash functions, demonstrates 23 percent speed-up and 96 percent bank conflict reduction on a Fermi GPU, and 33 percent speed-up and 99 percent bank conflict reduction on a Kepler GPU. Gert-Jan van den Braak, Juan Gómez-Luna, José María González-Linares, Henk Corporaal, Nicolás Guil |
IEEE Trans. Computers | 3 |
| 2016 | In-Place Matrix Transposition on GPUsabstractMatrix transposition is an important algorithmic building block for many numeric algorithms such as FFT. With more and more algebra libraries offloading to GPUs, a high performance in-place transposition becomes necessary. Intuitively, in-place transposition should be a good fit for GPU architectures due to limited available on-board memory capacity and high throughput. However, direct application of CPU in-place transposition algorithms lacks the amount of parallelism and locality required by GPU to achieve good performance. In this paper we present our in-place matrix transposition approach for GPUs that is performed using elementary tile-wise transpositions. We propose low-level optimizations for the elementary transpositions, and find the best performing configurations for them. Then, we compare all sequences of transpositions that achieve full transposition, and detect which is the most favorable for each matrix. We present an heuristic to guide the selection of tile sizes, and compare them to brute-force search. We diagnose the drawback of our approach, and propose a solution using minimal padding. With fast padding and unpadding kernels, the overall throughput is significantly increased. Finally, we compare our method to another recent implementation. Juan Gómez-Luna, I-Jui Sung, Li-Wen Chang, José María González-Linares, Nicolás Guil, Wen-Mei W. Hwu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | Calculation of dense trajectory descriptors on a heterogeneous embedded architecture
Julián Ramos Cózar, Manuel J. Marín-Jiménez, José María González-Linares, Nicolás Guil, Juan Gómez-Luna |
J. Syst. Archit. | 3 |
| 2014 | In-place transposition of rectangular matrices on acceleratorsabstractMatrix transposition is an important algorithmic building block for many numeric algorithms such as FFT. It has also been used to convert the storage layout of arrays. With more and more algebra libraries offloaded to GPUs, a high performance in-place transposition becomes necessary. Intuitively, in-place transposition should be a good fit for GPU architectures due to limited available on-board memory capacity and high throughput. However, direct application of CPU in-place transposition algorithms lacks the amount of parallelism and locality required by GPUs to achieve good performance. In this paper we present the first known in-place matrix transposition approach for the GPUs. Our implementation is based on a novel 3-stage transposition algorithm where each stage is performed using an elementary tiled-wise transposition. Additionally, when transposition is done as part of the memory transfer between GPU and host, our staged approach allows hiding transposition overhead by overlap with PCIe transfer. We show that the 3-stage algorithm allows larger tiles and achieves 3X speedup over a traditional 4-stage algorithm, with both algorithms based on our high-performance elementary transpositions on the GPU. We also show our proposed low-level optimizations improve the sustained throughput to more than 20 GB/s. Finally, we propose an asynchronous execution scheme that allows CPU threads to delegate in-place matrix transposition to GPU, achieving a throughput of more than 3.4 GB/s (including data transfers costs), and improving current multithreaded implementations of in-place transposition on CPU. I-Jui Sung, Juan Gómez-Luna, José María González-Linares, Nicolás Guil, Wen-Mei W. Hwu |
PPoPP | 3 |
| 2013 | Simulation and architecture improvements of atomic operations on GPU scratchpad memoryabstractGPUs are increasingly used as compute accelerators. With a large number of cores executing an even larger number of threads, significant speed-ups can be attained for parallel workloads. Applications that rely on atomic operations, such as histogram and Hough transform, suffer from serialization of threads in case they update the same memory location. Previous work shows that reducing this serialization with software techniques can increase performance by an order of magnitude. We observe, however, that some serialization remains and still slows down these applications. Therefore, this paper proposes to use a hash function in both the addressing of the banks and the locks of the scratchpad memory. To measure the effects of these changes, we first implement a detailed model of atomic operations on scratchpad memory in GPGPU-Sim, and verify its correctness. Second, we test our proposed hardware changes. They result in a speed-up up to 4.9× and 1.8× on implementations utilizing the aforementioned software techniques for histogram and Hough transform applications respectively, with minimum hardware costs. Gert-Jan van den Braak, Juan Gómez-Luna, Henk Corporaal, José María González-Linares, Nicolás Guil |
ICCD | 4 |
| 2013 | An optimized approach to histogram computation on GPU
Juan Gómez-Luna, José María González-Linares, José Ignacio Benavides Benítez, Nicolás Guil |
Mach. Vis. Appl. | 2 |
| 2013 | Performance Modeling of Atomic Additions on GPU Scratchpad MemoryabstractGPU application implementations using scatter approaches will fall into write contention due to atomic updates of output elements, if these result from more than one input element. Colliding threads will be serialized, seriously harming performance. Dealing with these issues requires a proper understanding of the behavior of the scratchpad or shared memory under conflicting accesses caused by concurrent threads. Thus, this paper presents an exhaustive microbenchmark-based analysis of atomic additions in shared memory that quantifies the impact of access conflicts on latency and throughput. This analysis has led us to discover the lock mechanism that enables atomic updates to shared memory and to propose a performance model to estimate the latency penalties due to collisions by position or bank conflicts. Then, we have derived experiments from this model that show us the way to optimize applications using atomic operations. Position and bank conflicts can be diminished by replication and padding, respectively. The benefits of such techniques are illustrated with the optimization of two widely used voting processes: the centroid updating step in k-means clustering, and histogram calculation. Juan Gómez-Luna, José María González-Linares, José Ignacio Benavides Benítez, Nicolás Guil |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | Reducing Vocabulary Size in Human Action ClassificationabstractHuman action classification is an important task in computer vision. Bag-of-Words using spatio-temporal features and some classification algorithm is one of the most successful methods in this context. In this work we have studied the effect of reducing the vocabulary size using a video word ranking method. We have used the KTH dataset to obtain a vocabulary with more descriptive words and, at the same time, more compact and efficient. Results for different vocabulary sizes show an improvement of the recognition rate whilst reducing the number of words due to the fact that non-descriptive words are removed. Julián Ramos Cózar, Ruber Hernández, Yanio Heredia, José María González-Linares, Nicolás Guil |
KES | 4 |
| 2012 | Pixel-based background initialization using spatio-temporal restrictionsabstractA precise background detection is required for video surveillance applications in order to detect foreground objects. However, in the presence of cluttered scenes, standard techniques for background segmentation can fail. In this work we present a new technique for foreground detection that is able to detect the correct background in complex scenes. It works grouping neighbour pixels that fulfill some kind of spatial and temporal criteria. Spatial criterion is based on the appearance similarity of neighbour pixels while temporal criterion looks for the best temporal correlation in the whole video sequence. Initially, a set of seed points of the image are selected and both criteria are applied in an alternate way until all the pixels of the image have been visited and their background value has been calculated. Juan Villalba Espinosa, José María González-Linares, Julián Ramos Cózar, Nicolás Guil |
KES | 2 |
| 2012 | Performance models for asynchronous data transfers on consumer Graphics Processing Units
Juan Gómez-Luna, José María González-Linares, José Ignacio Benavides Benítez, Nicolás Guil |
J. Parallel Distributed Comput. | 2 |
| 2011 | Detection of logos in low quality videosabstractThis paper presents a novel framework for logo detection in low quality videos. Our method assumes the logo template is unknown in advance and exploits the property that logotype pixels appearance through several consecutive frames has a lower variance than the others. Segmentation is difficult to accomplish if logo continuity is broken. In this work we propose the use of both edge and appearance continuity to carry out the segmentation. By checking edge continuity, the video is split into sequences with stable content. Later, sequences with similar static content are merged in order to build a longer sequence. Next a Gaussian mixture is used to model the variance of the pixels values in the merged sequences. Finally, a threshold that allows identification of the logo pixels is calculated. The new method is compared with a state-of-the-art method, obtaining better results in both accuracy and false logo rejection. Julián Ramos Cózar, Pablo Nieto, José María González-Linares, Nicolás Guil, Yanio Hernández Heredia |
ISDA | 3 |
| 2009 | Parallelization of a Video Segmentation Algorithm on CUDA-Enabled Graphics Processing Units
Juan Gómez-Luna, José María González-Linares, José Ignacio Benavides Benítez, Nicolás Guil |
Euro-Par | 2 |
| 2008 | A TV-logo classification and learning systemabstractLogotypes superimposed to broadcasted videos supply important information for semantic video annotation, such as the content creator. In this work a novel logo classification and learning system for TV broadcast videos is presented. Logos are segmented from the video stream but scale change, position shift, clutter and noise makes difficult to classify and to recognize them. Several robust features that use edges and shape information have been selected, and a Bayesian network classifier is used to classify the logos. New logos are recognized as such for the first time they appear and passed to a semi-supervised learning system. The learning process clusters the set of new logos to group different instances of the same new logo. A logo model is obtained for each cluster that must be validated by a human to incorporate them into the classification system. Comprehensive tests with a set of 724 TV logos show the high performance of our classification and learning system. Pablo Nieto, Julián Ramos Cózar, José María González-Linares, Nicolás Guil |
ICIP | 3 |
| 2007 | Logotype detection to support semantic-based video annotation
Julián Ramos Cózar, Nicolás Guil, José María González-Linares, Emilio L. Zapata, Ebroul Izquierdo |
Signal Process. Image Commun. | 3 |
| 2006 | Video Cataloging Based on Robust Logotype DetectionabstractIn this paper a technique for video cataloging based on logo detection is shown. No a priori knowledge about shape or spatial-temporal location of logos is assumed. The method implements a new algorithm for online logo detection based on temporal and spatial segmentation of broadcasted videos. Temporal segmentation identifies constant luminance regions within video frames while spatial segmentation helps to refine previous segmented regions. In a final step, identified logos are searched in a database and classified into candidate or learnt logotypes. Learnt logos can be directly tracked through the video. Candidate logotypes are assigned to a cluster of similar logos. After a promotion process, all the candidate logos belonging to the same cluster are used to create a new learnt logotype. Julián Ramos Cózar, Nicolás Guil, José María González-Linares, Emilio L. Zapata |
ICIP | 3 |
| 2003 | An efficient 2D deformable objects detection and location algorithm
José María González-Linares, Nicolás Guil, Emilio L. Zapata |
Pattern Recognit. | 1 |
| 2000 | Deformable Shapes Detection by Stochastic OptimizationabstractA new approach to the detection of shapes under global deformations is presented. The algorithm is based in the combination of a generalized Hough transform (GHT) and an universal evolutionary global optimizer (UEGO). This method exploits the invariant characteristics to rotation, scale and displacement of the GHT to detect shapes deformed by a global deformation model, and without an initial positioning of the template. The GHT is used as an objective function for the UEGO, an optimizer that is able to find multiple optima with a low computational cost. José María González-Linares, Nicolás Guil, Emilio L. Zapata, Pilar Martínez Ortigosa, Inmaculada García |
ICIP | 1 |
| 1999 | Bidimensional shape detection using an invariant approach
Nicolás Guil, José María González-Linares, Emilio L. Zapata |
Pattern Recognit. | 2 |