VLDB 2026 Research / reviewers in the wild / expert
Joel Sole
dblp:174/2071
· DBLP profile ↗
36ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HIIF: Hierarchical Encoding based Implicit Image Function for Continuous Super-resolutionabstractRecent advances in implicit neural representations (INRs) have shown significant promise in modeling visual signals for various low-vision tasks including image super-resolution (ISR). INR-based ISR methods typically learn continuous representations, providing flexibility for generating high-resolution images at any desired scale from their low-resolution counterparts. However, existing INR-based ISR methods utilize multi-layer perceptrons for parameterization in the network; this does not take account of the hierarchical structure existing in local sampling points and hence constrains the representation capability. In this paper, we propose a new Hierarchical encoding based Implicit Image Function for continuous image super-resolution, HIIF, which leverages a novel hierarchical positional encoding that enhances the local implicit representation, enabling it to capture fine details at multiple scales. Our approach also embeds a multi-head linear attention mechanism within the implicit attention network by taking additional non-local information into account. Our experiments show that, when integrated with different backbone encoders, HIIF outperforms the state-of-the-art continuous image super-resolution methods by up to 0.17dB in PSNR. The source code of HIIF will be made publicly available at https://github.com/YuxuanJJ/HIIF. Yuxuan Jiang 0015, Ho Man Kwan, Tianhao Peng 0004, Ge Gao 0005, Fan Zhang 0017, Joel Sole, David Bull 0001 |
CVPR | 7 |
| 2025 | RTSR: A Real-Time Super-Resolution Model for AV1 Compressed ContentabstractSuper-resolution (SR) is a key technique for improving the visual quality of video content by increasing its spatial resolution while reconstructing fine details. SR has been employed in many applications including video streaming, where compressed low-resolution content is typically transmitted to end users and then reconstructed with a higher resolution and enhanced quality. To support real-time playback, it is important to implement fast SR models while preserving reconstruction quality; however, most existing solutions, in particular those based on complex deep neural networks, fail to do so. To address this issue, this paper proposes a low-complexity SR method, RTSR, designed to enhance the visual quality of compressed video content, focusing on resolution up-scaling from a) 360p to 1080p and from b) 540p to 4K. The proposed approach utilizes a Convolutional Neural Network (CNN)-based network architecture, which was optimized for AOMedia Video 1 (AV1SVT)-encoded content at various quantization levels based on a dual-teacher knowledge distillation method. This method was submitted to the AIM 2024 Video Super-Resolution Challenge, specifically targeting the Efficient/Mobile Real-Time Video SuperResolution competition. It achieved the best trade-off between complexity and coding performance (measured in PSNR, SSIM and VMAF) among all six submissions. The code will be available at https://github.com/YuxuanJJ/RTSR. Yuxuan Jiang 0015, Jakub Nawala, Chen Feng 0008, Fan Zhang 0017, Joel Sole, David Bull 0001 |
ISCAS | 6 |
| 2024 | BVI-AOM: A New Training Dataset for Deep Video Compression OptimizationabstractDeep learning is now playing an important role in enhancing the performance of conventional hybrid video codecs. These learning-based methods typically require diverse and representative training material for optimization in order to achieve model generalization and optimal coding performance. However, existing datasets either offer limited content variability or come with restricted licensing terms constraining their use to research purposes only. To address these issues, we propose a new training dataset, named BVI-AOM, which contains 956 uncompressed sequences at various resolutions from 270p to 2160p, covering a wide range of content and texture types. The dataset comes with more flexible licensing terms and offers competitive performance when used as a training set for optimizing deep video coding tools. The experimental results demonstrate that when used as a training set to optimize two popular network architectures for two different coding tools, the proposed dataset leads to additional bitrate savings of up to 0.29 and 2.98 percentage points in terms of PSNR-Y and VMAF, respectively, compared to an existing training dataset, BVI-DVC, which has been widely used for deep video coding. The BVI-AOM dataset is available at https://github.com/fan-aaron-zhang/bvi-aom. Jakub Nawala, Yuxuan Jiang 0015, Fan Zhang 0017, Joel Sole, David Bull 0001 |
VCIP | 5 |
| 2024 | Learned fractional downsampling network for adaptive video streaming
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Chao Chen 0006, Alan C. Bovik |
Signal Process. Image Commun. | 4 |
| 2023 | A debanding algorithm for AV2abstractBanding is a visually unpleasing artifact appearing in flat areas of encoded content that no video standard has fully addressed. We propose a normative debanding filter to tackle banding artifacts and have tested it as an inloop and post-loop filter in AVM. Debanding is achieved by introducing dithering on a frame level to the luma component. The proposed filter shows CAMBI gains for content with banding while not affecting other content. Although the added dithering has a minor negative impact on some objective metrics, subjective improvements in banding-prone content are (informally) observed. On the test set, encoding time increases on average by ~0.5%, while decoding time increases by around 0.5% for in-loop and 1.5% for post-loop. Joel Sole, Mariana Afonso |
DCC | 1 |
| 2022 | Banding vs. Quality: perceptual impact and objective assessmentabstractStaircase-like contours introduced to a video by quantization in flat areas, commonly known as banding, have been a longstanding problem in both video processing and quality assessment communities. The fact that even a relatively small change of the original pixel values can result in a strong impact on perceived quality makes banding especially difficult to be detected by objective quality metrics. In this paper, we study how banding annoyance compares to more commonly studied scaling and compression artifacts with respect to the overall perceptual quality. We further propose a simple combination of VMAF and the recently developed banding index, CAMBI, into a banding-aware video quality metric showing improved correlation with overall perceived quality. Lukas Krasula, Zhi Li 0001, Christos G. Bampis, Mariana Afonso, Nil Fons Miret, Joel Sole |
ICIP | 6 |
| 2021 | A Progressive Architecture for Learned Fractional DownsamplingabstractIn many image and video processing applications, the ability to resize by a fractional factor, such as from 1080p to 720p, is essential. However, conventional CNN layers can only be used to alter the resolution of their inputs with integer scale factors. In this paper, we propose a downsampling network architecture that progressively reconstructs residuals at different scales. In particular, the aforementioned problem is solved by combining an upsampling sub-network and a downsampling subnetwork, both with integer scale factor. As an application, we apply the proposed downsampling network to an adaptive bitrate video streaming scenario. We extensively evaluate with different video codecs and upsampling algorithms to show the generality of our model. Our experimental results show that improvements in coding efficiency over the conventional Lanczos downsampling and state-of-the-art methods are attained, measured in different perceptual video quality models on large-resolution test videos. Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Alan C. Bovik |
PCS | 4 |
| 2021 | VMAF-based Bitrate Ladder Estimation for Adaptive StreamingabstractIn HTTP Adaptive Streaming, video content is conventionally encoded by adapting its spatial resolution and quantization level to best match the prevailing network state and display characteristics. It is well known that the traditional solution, of using a fixed bitrate ladder, does not result in the highest quality of experience for the user. Hence, in this paper, we introduce a content-driven approach for estimating the bitrate ladder, based on spatio-temporal features extracted from the uncompressed content. The method implements a content-driven interpolation. It uses the extracted features to train a machine learning model to infer the curvature points of the Rate-VMAF curves in order to guide a set of initial encodings. We employ the VMAF quality metric as a means of perceptually conditioning the estimation. When compared to the generation of a reference ladder using exhaustive encoding, 76.63% the estimated ladder's Rate-VMAF points are identical to those of the reference ladder. The proposed method benefits from a significant (77.4%) reduction in the number of encodes required with only a small (1.04%) average Bj⊘ntegaard Delta Rate increase. Angeliki V. Katsenou, Fan Zhang 0017, Kyle Swanson, Mariana Afonso, Joel Sole, David Bull 0001 |
PCS | 5 |
| 2021 | CAMBI: Contrast-aware Multiscale Banding IndexabstractBanding artifacts are artificially-introduced contours arising from the quantization of a smooth region in a video. Despite the advent of recent higher quality video systems with more efficient codecs, these artifacts remain conspicuous, especially on larger displays. In this work, a comprehensive subjective study is performed to understand the dependence of the banding visibility on encoding parameters and dithering. We subsequently develop a simple and intuitive no-reference banding index called CAMBI (Contrast-aware Multiscale Banding Index) which uses insights from Contrast Sensitivity Function in the Human Visual System to predict banding visibility. CAMBI correlates well with subjective perception of banding while using only a few visually-motivated hyperparameters. Pulkit Tandon, Mariana Afonso, Joel Sole, Lukas Krasula |
PCS | 3 |
| 2021 | Perceptual Video Quality Prediction Emphasizing Chroma DistortionsabstractMeasuring the quality of digital videos viewed by human observers has become a common practice in numerous multimedia applications, such as adaptive video streaming, quality monitoring, and other digital TV applications. Here we explore a significant, yet relatively unexplored problem: measuring perceptual quality on videos arising from both luma and chroma distortions from compression. Toward investigating this problem, it is important to understand the kinds of chroma distortions that arise, how they relate to luma compression distortions, and how they can affect perceived quality. We designed and carried out a subjective experiment to measure subjective video quality on both luma and chroma distortions, introduced both in isolation as well as together. Specifically, the new subjective dataset comprises a total of 210 videos afflicted by distortions caused by varying levels of luma quantization commingled with different amounts of chroma quantization. The subjective scores were evaluated by 34 subjects in a controlled environmental setting. Using the newly collected subjective data, we were able to demonstrate important shortcomings of existing video quality models, especially in regards to chroma distortions. Further, we designed an objective video quality model which builds on existing video quality algorithms, by considering the fidelity of chroma channels in a principled way. We also found that this quality analysis implies that there is room for reducing bitrate consumption in modern video codecs by creatively increasing the compression factor on chroma channels. We believe that this work will both encourage further research in this direction, as well as advance progress on the ultimate goal of jointly optimizing luma and chroma compression in modern video encoders. Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2020 | Cross-Component Prediction in HEVCabstractVideo coding in the YCbCr color space has been widely used, since it is efficient for compression, but it can result in color distortion due to conversion error. Meanwhile, coding in the RGB color space maintains high color fidelity, having the drawback of a substantial bitrate increase with respect to YCbCr coding. Cross-component prediction (CCP) efficiently compresses video content by decorrelating color components while keeping high color fidelity. In this scheme, the chroma residual signal is predicted from the luma residual signal inside the coding loop. This paper gives a description of the CCP scheme from several points of view, from theoretical background to practical implementation. The proposed CCP scheme has been evaluated in standardization communities and adopted into H.265/High Efficiency Video Coding (HEVC) Range Extensions. The experimental results show significant coding performance improvements for both natural and screen content video, while the quality of all color components is maintained. The average coding gains for natural video are 17% and 5% bitrate reduction in the case of intra coding and 11% and 4% in the case of inter coding for RGB and YCbCr coding, respectively, while the average increment of encoding and decoding times in the HEVC reference software implementation are 10% and 4%, respectively. Woo-Shik Kim, Ali Khairat, Mischa Siekmann, Joel Sole, Jianle Chen, Marta Karczewicz, Tung Nguyen 0001, Detlev Marpe |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Content-gnostic Bitrate Ladder Prediction for Adaptive Video StreamingabstractA challenge that many video providers face is the heterogeneity of networks and display devices for streaming, as well as dealing with a wide variety of content with different encoding performance. In the past, a fixed bit rate ladder solution based on a "fitting all" approach has been employed. However, such a content-tailored solution is highly demanding; the computational and financial cost of constructing the convex hull per video by encoding at all resolutions and quantization levels is huge. In this paper, we propose a content-gnostic approach that exploits machine learning to predict the bit rate ranges for different resolutions. This has the advantage of significantly reducing the number of encodes required. The first results, based on over 100 HEVC-encoded sequences demonstrate the potential, showing an average Bjøntegaard Delta Rate (BDRate) loss of 0.51% and an average BDPSNR loss of 0.01 dB compared to the ground truth, while significantly reducing the number of pre-encodes required when compared to two other methods (by 81%-94%). Angeliki V. Katsenou, Joel Sole, David Bull 0001 |
PCS | 2 |
| 2016 | High Dynamic Range Video Coding with Backward CompatibilityabstractThis paper presents a method for efficient compression of high dynamic range (HDR) and wide color gamut (WCG) video data. The proposed solution consists of two major elements: a conventional video codec (e.g., HEVC) and pre-and post-processing steps applied prior to encoding and after decoding process, respectively. The proposed HDR/WCG video coding system can be configured to provide two configurations: (1) a non-backward compatible bitstream with improved HDR video quality and (2) a SDR backward compatible bitstream with balanced visual quality between the reconstructed signal by the SDR and the HDR receivers. The simulations conducted under the MPEG Common Test Conditions for HDR demonstrate that the compression efficiency of the proposed solution outperforms the anchor solution on objective metrics. Additionally, subjective evaluations conducted under MPEG revealed improved visual quality for the proposed method. Dmytro Rusanovskyy, Döne Bugdayci Sansli, Adarsh K. Ramasubramonian, Joel Sole, Marta Karczewicz |
DCC | 5 |
| 2016 | Overview of the Range Extensions for the HEVC Standard: Tools, Profiles, and PerformanceabstractThe Range Extensions (RExt) of the High Efficiency Video Coding (HEVC) standard have recently been approved by both ITU-T and ISO/IEC. This set of extensions targets video coding applications in areas including content acquisition, postproduction, contribution, distribution, archiving, medical imaging, still imaging, and screen content. In addition to the functionality of HEVC Version 1, RExt provide support for monochrome, 4:2:2, and 4:4:4 chroma sampling formats as well as increased sample bit depths beyond 10 bits per sample. This extended functionality includes new coding tools with a view to provide additional coding efficiency, greater flexibility, and throughput at high bit depths/rates. Improved lossless, near-lossless, and very high bit-rate coding is also a part of the RExt scope. This paper presents the technical aspects of HEVC RExt, including a discussion of RExt profiles, tools, applications, and provides experimental results for a performance comparison with previous relevant coding technology. When compared with the High 4:4:4 Predictive Profile of H.264/Advanced Video Coding (AVC), the corresponding HEVC 4:4:4 RExt profile provides up to ~25$ %, ~32%, and ~36% average bit-rate reduction at the same PSNR quality level for intra, random access, and low delay configurations, respectively. David Flynn, Detlev Marpe, Matteo Naccari, Tung Nguyen 0001, Chris Rosewarne, Karl Sharman, Joel Sole, Jizheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2015 | Intra Block Copy for HEVC Screen Content CodingabstractSummary form only given. Screen content videos increasingly gain the popularity due to the rapid advances in cloud and multimedia technologies, which in turn requires highly efficient screen content compression. A recent standard, namely SCC is under development in JCT-VC, Joint Collaborative Team on Video Coding between ISO/IEC and ITU-T. In SCC, the most efficient new coding tool is Intra block copy (Intra BC). In this paper, we describe the Intra BC that has been proposed by the authors and adopted in the SCC standard and reference software for coding of screen content. Different from the conventional Intra prediction method where the prediction signal is derived from the spatially neighboring samples, the Intra BC mode greatly improves the prediction efficiency by fully exploiting the redundancy of repetitive patterns which typically appear in screen content. Experimental results suggest that the Intra BC mode can improve the coding efficiency significantly for typical screen content video sequences with 43.2% bit rate reduction on average. Joel Sole, Ying Chen 0011, Vadim Seregin, Marta Karczewicz |
DCC | 2 |
| 2015 | Adaptive Color-Space Transform for HEVC Screen Content CodingabstractThis paper presents an in-loop adaptive color-space transform for the HEVC Screen Content Coding extension. In the proposed method, the prediction residual is adaptively converted into a different color space to reduce the cross-component redundancy. After the ACT, the signal is coded following the existing HEVC framework. To keep the complexity as low as possible, fixed color-space transforms that are easily implemented with shift and add operations are utilized. Significant coding gains are achieved by this method in the current HEVC Screen Content Coding reference software with no increase of decoding runtime. The proposed method has been adopted to the HEVC Screen Content Coding extension. Li Zhang 0006, Jianle Chen, Joel Sole, Marta Karczewicz, Xiaoyu Xiu, Ji-Zheng Xu |
DCC | 3 |
| 2014 | Color palette for screen content codingabstractWith the prevalence of high speed Internet access, emerging video applications such as remote desktop sharing, virtual desktop infrastructure, and wireless display require high compression efficiency of screen contents. However, traditional intra and inter video coding tools were designed primarily for natural contents. Screen contents have significantly different characteristics compared with nature contents, e.g. sharp edges, less or no noise, which makes those traditional coding tools less sufficient. In this research, a new color palette based video coding tool is presented. Different from traditionally intra and inter prediction that mainly removes redundancy between different coding units, palette coding targets at the redundancy of repetitive pixel values/patterns within the coding unit. In the palette coding mode, a lookup table named palette which maps pixel values into table indices (also called palette indices) is signaled first. Then the mapped indice for a coding unit (which we call index block) are coded with a novel three-mode run-length entropy coding. Some encoder-side optimization for palette coding is also presented in detail in this paper. Simulation has been performed using the common screen content coding test condition defined by JCT-VC and the results show that palette coding can effectively improve screen content coding efficiency for both lossless and lossy scenarios. Joel Sole, Marta Karczewicz, Rajan L. Joshi |
ICIP | 4 |
| 2014 | Cross component decorrelation for HEVC range extension standardabstractThis paper presents a new coding tool named cross component decorrelation in the emerging High Efficiency Video Coding Range Extension (HEVC RExt) standard. Color video is generally composed of three color components, e.g., RGB or YCbCr. It has been known for over a decade that the three color components have correlation among each other. Although global out-of-loop color space conversion, e.g., RGB-to-YCbCr, can reduce the cross component correlation, local correlation still exists in YCbCr signal. Many methods have been developed in the literature to exploit such redundancy to improve coding efficiency. However, existing methods introduce high computational or implementation complexity, which makes them never be included in mainstream video coding standard such as H.264/AVC and HEVC version 1. In this research, a new hardware friendly cross component decorrelation method is presented which reduces implementation cost while achieving significant BD-rate reduction. For example, based on JCT-VC common test condition for HEVC RExt standardization, the proposed method results in (17.3%, 18.1%, 16.6%) BD-rate reduction for the three color components of the RGB test sequences, in the case of All Intra configuration. Woo-Shik Kim, Jianle Chen, Joel Sole, Marta Karczewicz |
ICIP | 4 |
| 2013 | Scalable Video Coding Extension for HEVCabstractThis paper describes a scalable video codec that was submitted as a response to the joint call for proposals issued by ISO/IEC MPEG and ITU-T VCEG on HEVC scalable extension. The proposed codec uses a multi-loop decoding structure. Several inter-layer texture prediction methods are employed to remove the inter-layer redundancy. Inter-layer prediction is also used when coding enhancement layer syntax elements such as motion parameter and intra prediction mode, to further reduce bit overhead. Additionally, alternative transforms as well as adaptive coefficients scanning are used to code the prediction residues more efficiently. Experimental results are presented to demonstrate the effectiveness of the proposed scheme. When compared to HEVC single-layer coding, the additional rate overhead for the proposed scalable extension is 1.2% to 6.4% to achieve two layers of SNR and spatial scalability. Jianle Chen, Krishnakanth Rapaka, Xiang Li 0003, Vadim Seregin, Marta Karczewicz, Geert Van der Auwera, Joel Sole, Xianglin Wang, Chengjie Tu, Ying Chen 0011, Rajan L. Joshi |
DCC | 8 |
| 2012 | Transform coefficient coding in HEVCabstractITU-T VCEG and ISO/IEC MPEG have undertaken a joint standardization activity on video coding called High Efficiency Video Coding (HEVC). This paper describes the transform coefficient coding in the HEVC Test Model (WD5) and the motivations driving the design. The coefficient coding description encompasses the scan patterns and the coding methods of the last significant coefficient, significance map and coefficient level. Special focus is given to the method for coding the last significant coefficient in the block. Joel Sole, Rajan L. Joshi, Wei-Jung Chien, Marta Karczewicz |
PCS | 1 |
| 2012 | Transform Coefficient Coding in HEVCabstractThis paper describes transform coefficient coding in the draft international standard of High Efficiency Video Coding (HEVC) specification and the driving motivations behind its design. Transform coefficient coding in HEVC encompasses the scanning patterns and coding methods for the last significant coefficient, significance map, coefficient levels, and sign data. Special attention is paid to the new methods of last significant coefficient coding, multilevel significance maps, high-throughput binarization, and sign data hiding. Experimental results are provided to evaluate the performance of transform coefficient coding in HEVC. Joel Sole, Rajan L. Joshi, Tianying Ji, Marta Karczewicz, Gordon Clare, Félix Henry, Alberto Duenas |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Directional adaptive loop filter for video codingabstractIn this paper, a directional adaptive loop filter is proposed for video coding. It classifies the pixels of a reconstructed frame into multiple categories based on the local directional characteristics. A Wiener filter is estimated for each category by minimizing the average mean square error between the original and reconstructed pixels in that category. The filter supports are carefully designed by exploiting the directional information in order to reduce the amount of overhead. The experimental results show promising performance in various coding configurations for both subjective and objective quality. Peng Yin 0002, Qian Xu 0003, Joel Sole, Xiaoan Lu |
ICIP | 4 |
| 2011 | Classified quadtree-based adaptive loop filterabstractIn this paper, we propose a classified quadtree-based adaptive loop filter (CQALF) in video coding. Pixels in a picture are classified into two categories by considering the impact of the deblocking filter, the pixels that are modified and the pixels that are not modified by the deblocking filter. A wiener filter is carefully designed for each category and the filter coefficients are transmitted to decoder. For the pixels that are modified by the deblocking filter, the filter is estimated at encoder by minimizing the mean square error between the original input frame and a combined frame which is a weighted average of the reconstructed frames before and after the deblocking filter. For pixels that the deblocking filter does not modify, the filter is estimated by minimizing the mean square error between the original frame and the reconstructed frame. The proposed algorithm is implemented on top of KTA software and compatible with the quadtree-based adaptive loop filter. Compared with kta2.6rl anchor, the proposed CQALF achieves 10.05%, 7.55%, and 6.19% BD bitrate reduction in average for intra only, IPPP, and HB coding structures respectively. Qian Chen 0024, Peng Yin 0002, Xiaoan Lu, Joel Sole, Qian Xu 0003, Edouard François, Dapeng Oliver Wu |
ICME | 5 |
| 2010 | Compressive sensing with adaptive pixel domain reconstruction for block-based video codingabstractThis paper presents a new look at image/video compression from the compressive sensing's perspective. Quantization in video compression can be regarded as a subsampling process where the signal is mapped into predefined levels. We view the problem of signal reconstruction from its quantized signal vector as a compressive sensing recovery problem where the quantized coefficients are subsampled measurements. Based on this observation, we propose a novel method of image/video coding that employs an adaptive Total-Variation (TV) minimization in the pixel domain to recover the gradient-sparse image blocks from their quantized transform coefficients. We further increase the coding efficiency by encoding only a subset of the transform coefficients and discard the remaining ones. Experiment results show that the proposed framework is efficient with gradient-sparse video signals and outperform the video compression standard H.264/AVC by up to 7% of bitrate reduction. Thong T. Do, Xiaoan Lu, Joel Sole |
ICIP | 3 |
| 2010 | Simplified geometry-adaptive block partitioning for video codingabstractGeometry-adaptive block partitioning (GEO) can greatly enhance video coding efficiency but at the expense of significantly increased computational complexity. Instead of proposing fast searching algorithm for encoding only, this paper proposes to reduce the size of partitions for both the encoder and the decoder. The proposed scheme only searches partitions recognized as most valuable partitions, which is derived by analyzing different GEO partitions from the contribution to the coding efficiency. On the one hand, by only searching limited number of partitions, the encoder computation burden is much alleviated and the decoder needs to handle fewer cases. On the other hand, restricting the number of candidate partitions causes fewer overhead bits. Results obtained in the experiments show that compared to the original GEO, the proposed scheme can achieve similar coding efficiency at much lower complexity. Peng Yin 0002, Xiaoan Lu, Qian Xu 0003, Joel Sole |
ICIP | 6 |
| 2010 | Adaptive motion vector resolution with implicit signalingabstractIn all video coding standards, motion vectors are restricted to have the same resolution for the whole sequence. In this paper, we propose an adaptive motion vector resolution scheme to improve the coding efficiency. The scheme is based on an analytical model of motion compensated residue which studies the relationship between motion vector errors and the variance of the residue. With this model, the proposed scheme can adaptively select motion vector resolution with implicit signaling. Experimental results show that the proposed scheme can improve coding efficiency at slightly increased encoding complexity. Peng Yin 0002, Xiaoan Lu, Qian Xu 0003, Joel Sole |
ICIP | 6 |
| 2010 | Coupled pre-/post-processing filters for predictive video codingabstractPre-/post-processing filters obtained from lapped transforms achieve higher coding efficiency by exploiting transform block boundary correlation. Existing approaches however do not explicitly consider the impact of pre-filtering on prediction efficiency, thereby resulting in a lower overall rate-distortion performance. We present a framework where filters are designed with an objective function based on a prediction model. The post-filter is constructed so as to reverse the pre-filter processing at the decoder. We show how the pre-filters thus derived can be refined and local adaptation parameters obtained for better compression efficiency. To further enhance the performance, the mode decision is altered by separating the contribution of luma and chroma components in the rate-distortion model. The overall system performance is measured using a weighted Bjontegaard delta bitrate, showing an overall saving of 2.58% to 6.78% with respect to the “Key Technical Area” (KTA) codec. Additionally, the designed filters significantly reduce the blocking artifacts observed in KTA. Kiran Misra, Joel Sole, Xiaoan Lu, Peng Yin 0002, Qian Xu 0003 |
ICIP | 2 |
| 2010 | Video Coding Using a Simplified Block Structure and Advanced Coding TechniquesabstractThis paper describes a new video coding scheme based on a simplified block structure that significantly outperforms the coding efficiency of the ISO/IEC 14496-10 ITU-T H.264 advanced video coding (AVC) standard. Its conceptual design is similar to a typical block-based hybrid coder applying prediction and subsequent prediction error coding. The basic coding unit is an 8 × 8 block for inter, and an 8 × 8 or a 16 × 16 block for intra, instead of the usual 16 × 16 macroblock. No larger block sizes are considered for prediction and transform. Based on this simplified block structure, the coding scheme uses simple and fundamental coding tools with optimized encoding algorithms. In particular, the motion representation is based on a minimum partitioning with blocks sharing motion borders. In addition, compared to AVC, the new and improved coding techniques include: block-based intensity compensation, motion vector competition, adaptive motion vector resolution, adaptive interpolation filters, edge-based intra prediction and enhanced chrominance prediction, intra template matching, larger trans forms and adaptive switchable transforms selection for intra and inter blocks, and nonlinear and frame-adaptive de-noising loop filters. Finally, the entropy coder uses a generic flexible zero tree representation applied to both motion and texture data. Attention has also been given to algorithm designs that facilitate parallelization. Compared to AVC, the new coding scheme offers clear benefits in terms of subjective video quality at the same bit rate. Objective quality improvements are equally significant. At the same quality, an average bit-rate reduction of 31% compared to AVC is reported. Frank Bossen, Virginie Drugeon, Edouard François, Joël Jung, Sandeep Kanumuri, Matthias Narroschke, Hisao Sasai, Joel Sole, Yoshinori Suzuki, Thiow Keng Tan, Thomas Wedi, Steffen Wittmann, Peng Yin 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2010 | Selective Data Pruning-Based Compression Using High-Order Edge-Directed InterpolationabstractThis paper proposes a selective data pruning-based compression scheme to improve the rate-distortion relation of compressed images and video sequences. The original frames are pruned to a smaller size before compression. After decoding, they are interpolated back to their original size by an edge-directed interpolation method. The data pruning phase is optimized to obtain the minimal distortion in the interpolation phase. Furthermore, a novel high-order interpolation is proposed to adapt the interpolation to several edge directions in the current frame. This high-order filtering uses more surrounding pixels in the frame than the fourth-order edge-directed method and it is more robust. The algorithm is also considered for multiframe-based interpolation by using spatio-temporally surrounding pixels coming from the previous frame. Simulation results are shown for both image interpolation and coding applications to validate the effectiveness of the proposed methods. Dung Trung Vo, Joel Sole, Peng Yin 0002, Cristina Gomila, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2009 | Data pruning-based compression using high order edge-directed interpolationabstractThis paper proposes a data pruning-based compression scheme to improve the rate-distortion relation of compressed images and video sequences. The original frames are pruned to a smaller size before compression. After decoding, they are interpolated to their original size by an edge-directed interpolation. The data pruning is optimized to obtain the minimal distortion in the interpolation phase. Furthermore, a novel high order interpolation is proposed to adapt the interpolation to many edge directions. This high order filtering uses extra surrounding pixels and achieves more robust edge-directed image interpolation. Simulation results are shown for both image interpolation and coding applications. Dung Trung Vo, Joel Sole, Peng Yin 0002, Cristina Gomila, Truong Q. Nguyen |
ICASSP | 2 |
| 2009 | Joint sparsity-based optimization of a set of orthonormal 2-D separable block transformsabstractWe propose an iterative method for the optimization of a set of 2-D separable transforms for a given training data set. The method outputs orthornormal transforms, each one being optimal for a subset of the data with respect to a sparsity-based objective function. The vertical and horizontal directions of the transform may be different, thus allowing directional-adapted transforms (in contrast to the usual DCT). Additionally, we relate the reconstruction error and the sparsity cost terms through the quantization step. To prove the validity of our approach, experimental results concerning coding applications are provided. Joel Sole, Peng Yin 0002, Cristina Gomila |
ICIP | 1 |
| 2009 | Sparsity-based deartifacting filtering in video compressionabstractIn the last years, many sparsity based denoising approaches for image/video denoising have been proposed. Most of them exploit the image/video sparsity model under certain overcomplete basis. In this paper, we unify three sparsity-based denoising techniques and apply them to the problem of video compression artifacts removal. We compare and analyze the three techniques from the aspects of operation atom, transform dimensionality, and quantization impact. Based on the provided analysis, the paper may serve as a guideline to apply sparsity-based denoising techniques to related problems. Peng Yin 0002, Joel Sole, Cristina Gomila, Dapeng Oliver Wu |
ICIP | 4 |
| 2007 | Generalized Lifting Prediction Optimization Applied to Lossless Image CompressionabstractA useful tool to construct wavelet decompositions is the lifting scheme. The generalized lifting is an extension of the classical lifting scheme to introduce more flexibility and to permit the creation of new nonlinear and adaptive transforms. However, the design of generalized prediction and update steps is more involved. This letter proposes a generalized prediction design that minimizes the detail signal energy and entropy at the same time. Two algorithm variants are given. The fixed prediction uses the image class statistics to derive the optimal transform. If the statistics are unknown, the adaptive prediction extracts them from the image being coded. The resulting decompositions are applied to lossless image coding, reporting good results. The adaptive algorithm has no bookkeeping or side information requirements, yet its performance is close to the fixed prediction performance. Joel Sole, Philippe Salembier |
IEEE Signal Process. Lett. | 1 |
| 2006 | A Common Formulation for Interpolation, Prediction, and Update Lifting DesignabstractThe optimization of a quadratic objective function with linear constraints is useful for interpolation purposes. This formulation may be employed to derive an initial prediction in the lifting scheme domain in order to construct wavelet transforms. We modify the formulation to design final prediction and update lifting steps. The linear constraints relate wavelet bases and coefficients with the underlying signal. The objective function is the detail signal energy for the prediction lifting design and the gradient of the approximation signal for the update. To report concrete results and the power of the approach, we derive update steps using an auto-regressive image model that show better performance than the 5/3 wavelet for the compression of several image classes Joel Sole, Philippe Salembier |
ICASSP (2) | 1 |
| 2005 | Adaptive Generalized Prediction for Lifting SchemesabstractThe lifting scheme is a useful tool to create different types of wavelet decompositions, including adaptive and nonlinear. Generalized lifting is more flexible and can improve lifting results, but the design of generalized prediction and update steps remains difficult for a given application. A design strategy to optimize the prediction step according to the image statistics is established. The criterion aims to minimize the detail signal coefficients energy, The scheme is used for lossless compression of several classes of images without any book-keeping or side information requirements. Promising results are reported for certain classes of images. Joel Sole, Philippe Salembier |
ICASSP (2) | 1 |
| 2004 | Adaptive discrete generalized lifting for lossless compressionabstractA method for lossless coding using a multi-resolution scheme with an adaptive lifting step is introduced. To this goal, a generalized lifting scheme, where the sums included in classical lifting are generalized to include possibly non linear and/or adaptive operations, is proposed. The conditions for the reversibility of the scheme are analyzed. Although the framework is valid for any type of sampled signals, we focus in particular on the discrete case (signals represented with a finite number of bits). Finally, experiments with a generalized prediction step are reported They show the interest of the approach for lossless compression. Joel Sole, Philippe Salembier |
ICASSP (3) | 1 |