Pablo Enfedaque

dblp:158/9704 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5Systems, architecture and hardware · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Image and video coding · 75% Image and video processing · 19% Audio and music processing · 6%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
GPUs and heterogeneous computing · 50% Processor architecture and microarchitecture · 25% Parallel and multicore computing · 25%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video coding
bit-plane coding
0.522017
GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017
Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016
GPUs and heterogeneous computing
GPU computing
0.522017
GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017
Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015
Parallel and multicore computing › parallel algorithms › parallel image processing
parallel image coding
0.212016
Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016
Processor architecture and microarchitecture
SIMD
0.212016
Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016
Image and video coding › transform coding
discrete wavelet transform
0.212015
Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015
Image and video processing
wavelet transform
0.212015
Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015
Image and video coding › image compression
wavelet-based image coding
0.112017
GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017
Audio and music processing
decorrelation
0.112015
Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015

Methods — techniques the papers use, named apart from their topics

thread-to-data mapping · 0.6memory management · 0.6context modeling · 0.5arithmetic coding · 0.5SIMD processing · 0.5register sharing · 0.4CUDA · 0.4
YearPublicationVenuePosition
2019 Iterative Joint Ptychography-Tomography with Total Variation Regularization
abstract
In order to determine the 3D structure of a thick sample, researchers have recently combined ptychography (for high resolution) and tomography (for 3D imaging) in a single experiment. 2-step methods are usually adopted for reconstruction, where the ptychography and tomography problems are often solved independently. In this paper, we provide a novel model and ADMM-based algorithm to jointly solve the ptychography-tomography problem iteratively, also employing total variation regularization. The proposed method permits large scan stepsizes for the ptychography experiment, requiring less measurements and being more robust to noise with respect to other strategies, while achieving higher reconstruction quality results.
Huibin Chang, Pablo Enfedaque, Stefano Marchesini
ICIP2
2019 Blind Ptychographic Phase Retrieval via Convergent Alternating Direction Method of Multipliers
abstract
Ptychography has risen as a reference X-ray imaging technique: it achieves resolutions of one billionth of a meter, macroscopic field of view, or the capability to retrieve chemical or magnetic contrast, among other features. A ptychographic reconstruction is normally formulated as a blind phase retrieval problem, where both the image (sample) and the probe (illumination) have to be recovered from phaseless measured data. In this article we address a nonlinear least squares model for the blind ptychography problem with constraints on the image and the probe by maximum likelihood estimation of the Poisson noise model. We formulate a variant model that incorporates the information of phaseless measurements of the probe to eliminate possible artifacts. Next, we propose a generalized alternating direction method of multipliers designed for the proposed nonconvex models with convergence guarantee under mild conditions, where their subproblems can be solved by fast elementwise operations. Numerically, the proposed algorithm outperforms state-of-the-art algorithms in both speed and image quality.
Huibin Chang, Pablo Enfedaque, Stefano Marchesini
SIAM J. Imaging Sci.2
2018 High Throughput Image Codec for High-Resolution Satellite Images
abstract
The growth in the use of satellite images has generated the need for their fast compression, processing, and distribution. JPEG2000 is a widespread standard for the compression and transmission of such images once they are in the ground. Despite its advanced features and excellent coding performance, JPEG2000 demands significant computational resources. This paper introduces a wavelet-based codec that uses the JPEG2000 framework, but replaces its most computationally demanding coding stage by a highly parallel engine. When executed in Graphics Processing Units to code high-resolution satellite images, the proposed codec achieves speed-ups of up to 8× when compared to the fastest implementation of JPEG2000 executed in a multi-core platform.
Carlos de Cea-Dominguez, Pablo Enfedaque, Juan C. Moure, Joan Bartrina-Rapesta, Francesc Aulí Llinàs
IGARSS2
2017 GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression
abstract
The fast compression of images is a requisite in many applications like TV production, teleconferencing, or digital cinema. Many of the algorithms employed in current image compression standards are inherently sequential. High performance implementations of such algorithms often require specialized hardware like field integrated gate arrays. Graphics Processing Units (GPUs) do not commonly achieve high performance on these algorithms because they do not exhibit fine-grain parallelism. Our previous work introduced a new core algorithm for wavelet-based image coding systems. It is tailored for massive parallel architectures. It is called bitplane coding with parallel coefficient processing (BPC-PaCo). This paper introduces the first high performance, GPUbased implementation of BPC-PaCo. A detailed analysis of the algorithm aids its implementation in the GPU. The main insights behind the proposed codec are an efficient thread-to-data mapping, a smart memory management, and the use of efficient cooperation mechanisms to enable inter-thread communication. Experimental results indicate that the proposed implementation matches the requirements for high resolution (4 K) digital cinema in real time, yielding speedups of 30× with respect to the fastest implementations of current compression standards. Also, a power consumption evaluation shows that our implementation consumes 40× less energy for equivalent performance than state-of-the-art methods.
Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure
IEEE Trans. Parallel Distributed Syst.1
2016 Bitplane Image Coding With Parallel Coefficient Processing
abstract
Image coding systems have been traditionally tailored for multiple instruction, multiple data (MIMD) computing. In general, they partition the (transformed) image in codeblocks that can be coded in the cores of MIMD-based processors. Each core executes a sequential flow of instructions to process the coefficients in the codeblock, independently and asynchronously from the others cores. Bitplane coding is a common strategy to code such data. Most of its mechanisms require sequential processing of the coefficients. The last years have seen the upraising of processing accelerators with enhanced computational performance and power efficiency whose architecture is mainly based on the single instruction, multiple data (SIMD) principle. SIMD computing refers to the execution of the same instruction to multiple data in a lockstep synchronous way. Unfortunately, current bitplane coding strategies cannot fully profit from such processors due to inherently sequential coding task. This paper presents bitplane image coding with parallel coefficient (BPC-PaCo) processing, a coding method that can process many coefficients within a codeblock in parallel and synchronously. To this end, the scanning order, the context formation, the probability model, and the arithmetic coder of the coding engine have been re-formulated. The experimental results suggest that the penalization in coding performance of BPC-PaCo with respect to the traditional strategies is almost negligible.
Francesc Aulí Llinàs, Pablo Enfedaque, Juan C. Moure, Victor Sanchez
IEEE Trans. Image Process.2
2015 Strategy of Microscopic Parallelism for Bitplane Image Coding
abstract
Recent years have seen the upraising of a new type of processors strongly relying on the Single Instruction, Multiple Data (SIMD) architectural principle. The main idea behind SIMD computing is to apply a flow of instructions to multiple pieces of data in parallel and synchronously. This permits the execution of thousands of operations in parallel, achieving higher computational performance than with traditional Multiple Instruction, Multiple Data (MIMD) architectures. The level of parallelism required in SIMD computing can only be achieved in image coding systems via microscopic parallel strategies that code multiple coefficients in parallel. Until now, the only way to achieve microscopic parallelism in bit plane coding engines was by executing multiple coding passes in parallel. Such a strategy does not suit well SIMD computing because each thread executes different instructions. This paper introduces the first bit plane coding engine devised for the fine grain of parallelism required in SIMD computing. Its main insight is to allow parallel coefficient processing in a coding pass. Experimental tests show coding performance results similar to those of JPEG2000.
Francesc Aulí Llinàs, Pablo Enfedaque, Juan C. Moure, Ian Blanes, Victor Sanchez
DCC2
2015 Strategies of SIMD Computing for Image Coding in GPU
abstract
The main difficulty to implement modern image coding systems in a GPU is that the algorithms employed in the core of the coding scheme are inherently sequential. We recently proposed bitplane image coding with parallel coefficient processing (BPC-PaCo), a coding scheme that, contrarily to most systems, permits the processing of multiple coefficients of the image in parallel. This enables the use of SIMD computing, ideal for its implementation in a GPU. This paper introduces and evaluates the GPU implementation of BPC-PaCo employing two different strategies that tradeoff computational throughput and compression efficiency. The proposed implementation is compared to the best CPU and GPU implementations of JPEG2000, the state-of-the-art image compression standard. Experimental results indicate that BPC-PaCo achieves a computational throughput that is an order of magnitude superior to that achieved with such implementations with a small reduction in coding efficiency.
Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure
HiPC1
2015 Implementation of the DWT in a GPU through a Register-based Strategy
abstract
The release of the CUDA Kepler architecture in March 2012 has provided Nvidia GPUs with a larger register memory space and instructions for the communication of registers among threads. This facilitates a new programming strategy that utilizes registers for data sharing and reusing in detriment of the shared memory. Such a programming strategy can significantly improve the performance of applications that reuse data heavily. This paper presents a register-based implementation of the Discrete Wavelet Transform (DWT), the prevailing data decorrelation technique in the field of image coding. Experimental results indicate that the proposed method is, at least, four times faster than the best GPU implementation of the DWT found in the literature. Furthermore, theoretical analysis coincide with experimental tests in proving that the execution times achieved by the proposed implementation are close to the GPU's performance limits.
Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure
IEEE Trans. Parallel Distributed Syst.1
2014 Evaluation of context models to code wavelet-transformed hyperspectral images
abstract
Context modeling is key in wavelet-based image coding schemes to achieve competitive coding performance. Commonly, context models are devised for a particular coding system and are employed for many different types of images. The aim of this work is to evaluate the suitability of three well-known context models for coding hyperspectral images, without focusing on a particular wavelet-based coding system. To do so, an entropy-based measure defined using the mechanisms utilized by modern image codecs is employed. The experimental results assess the appropriateness of the context models considering different coding rates and transform strategies. They reveal that some widely-used context models may not be as adequate as it is generally thought. The hints provided by this analysis may help to design simpler and more efficient wavelet-based codecs for hyperspectral images.
Francesc Aulí Llinàs, Pablo Enfedaque, Joan Serra-Sagristà, Victor Sanchez
ICIP2