Peter Lambert

dblp:56/4700 · DBLP profile ↗
← Back
120ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-5313-4158ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 106 · 4 first-author · 12 since 2021Security and privacy · 4 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Artificial intelligence and machine learning · 3Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 3Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 TGIF2: extended text-guided inpainting forgery dataset and benchmark
abstract
Generative AI has made text-guided inpainting a powerful image editing tool, but at the same time a growing challenge for media forensics. Existing benchmarks, including our text-guided inpainting forgery (TGIF) dataset, show that image forgery localization (IFL) methods can localize manipulations in spliced images but struggle in fully regenerated (FR) images, while synthetic image detection (SID) methods can detect fully regenerated images but cannot perform localization. With new generative inpainting models emerging and the open problem of localization in FR images remaining, updated datasets and benchmarks are needed. We introduce TGIF2, an extended version of TGIF, that captures recent advances in text-guided inpainting and enables a deeper analysis of forensic robustness. TGIF2 augments the original dataset with edits generated by FLUX.1 models, as well as with random non-semantic masks. Using the TGIF2 dataset, we conduct a forensic evaluation spanning IFL and SID, including fine-tuning IFL methods on FR images and generative super-resolution attacks. Our experiments show that both IFL and SID methods degrade on FLUX.1 manipulations, highlighting limited generalization. Additionally, while fine-tuning improves localization on FR images, evaluation with random non-semantic masks reveals object bias. Furthermore, generative super-resolution significantly weakens forensic traces, demonstrating that common image enhancement operations can undermine current forensic pipelines. In summary, TGIF2 provides an updated dataset and benchmark, which enables new insights into the challenges posed by modern inpainting and AI-based image enhancements. TGIF2 is available at https://github.com/IDLabMedia/tgif-dataset .
Hannes Mareen, Dimitrios Karageorgiou, Paschalis Giakoumoglou, Peter Lambert, Symeon Papadopoulos, Glenn Van Wallendael
J. Inf. Secur.4
2026 POTR: Post-Training 3DGS Compression
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising contender to Neural Radiance Fields (NeRF) in 3D scene reconstruction and real-time novel view synthesis. 3DGS outperforms NeRF in training and inference speed but has substantially higher storage requirements. To remedy this downside, we propose POTR, a post-training 3DGS codec built on two novel techniques. First, POTR introduces a novel pruning approach that uses a modified 3DGS rasterizer to efficiently calculate every splat’s individual removal effect simultaneously. This technique results in 2-4× fewer splats than other post-training pruning techniques and as a result also significantly accelerates inference with experiments demonstrating 1.5-2× faster inference than other compressed models. Second, we propose a novel method to recompute lighting coefficients, significantly reducing their entropy without using any form of training. Our fast and highly parallel approach especially increases AC lighting coefficient sparsity, with experiments demonstrating increases from 70% to 97%, with minimal loss in quality. Finally, we extend POTR with a simple fine-tuning scheme to further enhance pruning, inference, and rate-distortion performance. Experiments demonstrate that POTR, even without fine-tuning, consistently outperforms all other post-training compression techniques in both rate-distortion performance and inference speed.
Bert Ramlot, Martijn Courteaux, Peter Lambert, Glenn Van Wallendael
IEEE Trans. Circuits Syst. Video Technol.3
2025 X265-PVMAF: A Real-Time Perceptual Video Quality Metric for HEVC Video Encoding
abstract
Real-time video encoding requires efficient and accurate quality metrics to optimize performance under strict computational and latency constraints. Traditional low-complexity metrics such as PSNR and SSIM often fall short in perceptual alignment, while accurate metrics such as VMAF are too computationally intensive for real-time deployment.We present x265-pVMAF, a low-complexity perceptual quality metric integrated into the x265 encoding loop. By leveraging machine learning and efficiently extracted encoder features, x265-pVMAF bridges the gap between computational efficiency and perceptual accuracy. It replicates VMAF predictions with a correlation of 0.99, while delivering a 37× speed-up. These results establish x265-pVMAF as a practical solution for real-time video quality assessment in next-generation encoding workflows.
Axel De Decker, Sangar Sivashanmugam, Jan De Cock, Hannes Mareen, Peter Lambert, Glenn Van Wallendael
ICIP5
2024 OpenDIBR: Open Real-Time Depth-Image-Based renderer of light field videos for VR
Julie Artois, Martijn Courteaux, Glenn Van Wallendael, Peter Lambert
Multim. Tools Appl.4
2024 A study on keyframe injection in three generations of video coding standards for fast channel switching and packet-loss repair
Hannes Mareen, Martijn Courteaux, Pieter-Jan Speelmans, Peter Lambert, Glenn Van Wallendael
Multim. Tools Appl.4
2023 Temporal Layer Injection for Fast Bitrate Ladder Creation in Video Live Streaming
abstract
Video streaming systems aim to provide high-quality video adapted to clients’ device and network conditions. For this purpose, adaptive streaming architectures encode video content at a variety of quality levels, organized in a bitrate ladder. However, compressing a video into multiple streams is resource-intensive, which may become especially problematic in live streaming applications with real-time demands. Therefore, this paper proposes a novel solution for fast bitrate ladder creation, and provides the requirements for implementation in the H.266/VVC standard. More specifically, the proposed method creates new intermediate Combined Streams by injecting the lowest temporal layers of a higher-quality Augmentation Stream in a lower-quality Base Stream. Since the lowest layers are used as reference by the remaining layers, this procedure indirectly increases the quality of the frames in those untouched remaining layers as well. We demonstrate that injecting more layers brings both the quality and bitrate closer to that of the Augmentation Stream. The disadvantage of the Combined Streams is that their quality fluctuates more than the quality of the source streams, and that they are compressed less efficiently, comparable to going from a slower to fast or faster preset in the VVenC encoder. Most importantly, their main advantage is that they were generated at no significant additional computational complexity. In this way, the proposed method is of great benefit when generating a bitrate ladder of video streams under constrained computational resources.
Hannes Mareen, Casper Haems, Tim Wauters, Filip De Turck, Peter Lambert, Glenn Van Wallendael
ISM5
2023 P-Frame Injection for Efficient Packet-Loss Repair in Ultra-Low-Latency Video Streaming
abstract
Applications providing ultra-low-latency video streaming to large audiences require fast and efficient packet-loss repair. Previous methods utilizing keyframe injection (such as the High Efficiency Streaming Protocol) have a low impact on the repaired stream quality, but at a cost of a significant bitrate spike during repair. In this paper, we propose injecting P-frames for packet-loss repair of ultra-low-latency streaming. We implemented (open-source) and evaluated our approach in both H.265/HEVC and H.266/VVC standards. Through extensive evaluations, we demonstrate that the proposed solution significantly reduces bitrate overhead while maintaining a similar or lower decrease in quality compared to existing packet-loss-repair techniques. Overall, the proposed approach offers a promising solution to ensure reliable packet-loss recovery, efficient resource utilization, and a high-quality streaming experience.
Hannes Mareen, Peter Lambert, Glenn Van Wallendael
VCIP2
2022 Fast and Blind Detection of Rate-Distortion-Preserving Video Watermarks
abstract
Forensic watermarking enables the tracing of digital pirates that leak copyright-protected multimedia. To prevent a negative impact on the video quality or bit rate, rate-distortion-preserving watermarking exists, which represents a watermark as compression artifacts. However, this method has two main disadvantages; the detection has a high complexity and it is non-blind. Although a method based on perceptual hashing exists that speeds up the detection of a fallback watermarking system, it decreases its robustness. Therefore, this paper proposes a novel fast detection method that has less impact on the robustness than related work. Our method optimized NS-DCT-DST hashes for rate-distortion-preserving watermarking, which are more robust to content-preserving attacks. Moreover, a blind version is proposed which does not require the original video for hash extraction. As such, the detection is experimentally measured to be up to 5700 times faster, at the cost of a modest decrease in robustness. In fact, the proposed method shows good robustness to content-preserving recompression attacks when using hashes that are as small as 432 bytes. This is much smaller than related work at comparable performance. In conclusion, this paper enables fast adversary tracing using watermarks that do not impact the video’s compression efficiency.
Hannes Mareen, Glenn Van Wallendael, Peter Lambert, Fouad Khelifi
ARES3
2022 Keyframe Insertion for Random Access and Packet-Loss Repair in H.264/AVC, H.265/HEVC, and H.266/VVC
abstract
Sending low-delay live video over error-prone channels comes with packet-loss-repair and random-access challenges. Existing solutions have a negative impact on end-users with reliable connections or users that do not switch channels. To minimize this impact, the keyframe-insertion technique extends a compression-efficient normal stream (NS) with a companion stream (CS) solely consisting of keyframes [1].
Hannes Mareen, Martijn Courteaux, Johan Vounckx, Peter Lambert, Glenn Van Wallendael
DCC4
2022 SILVR: a synthetic immersive large-volume plenoptic dataset
abstract
In six-degrees-of-freedom light-field (LF) experiences, the viewer's freedom is limited by the extent to which the plenoptic function was sampled. Existing LF datasets represent only small portions of the plenoptic function, such that they either cover a small volume, or they have limited field of view. Therefore, we propose a new LF image dataset "SILVR" that allows for six-degrees-of-freedom navigation in much larger volumes while maintaining full panoramic field of view. We rendered three different virtual scenes in various configurations, where the number of views ranges from 642 to 2226. One of these scenes (called Zen Garden) is a novel scene, and is made publicly available. We chose to position the virtual cameras closely together in large cuboid and spherical organisations (2.2m3 to 48m3), equipped with 180° fish-eye lenses. Every view is rendered to a color image and depth map of 2048px × 2048px. Additionally, we present the software used to automate the multiview rendering process, as well as a lens-reprojection tool that converts between images with panoramic or fish-eye projection to a standard rectilinear (i.e., perspective) projection. Finally, we demonstrate how the proposed dataset and software can be used to evaluate LF coding/rendering techniques (in this case for training NeRFs with instant-ngp). As such, we provide the first publicly-available LF dataset for large volumes of light with full panoramic field of view.
Martijn Courteaux, Julie Artois, Stijn De Pauw, Peter Lambert, Glenn Van Wallendael
MMSys4
2022 Mixed-Resolution HESP for More Efficient Fast Channel Switching and Packet-Loss Repair
abstract
Low-delay live streaming applications desire fast channel switching and packet-loss repair capabilities. However, existing methods that provide these capabilities have a negative impact on the stream of steady-state users. To minimize this impact, techniques such as the High Efficiency Streaming Protocol (HESP) utilize keyframe injection. Such techniques combine compression-efficient normal streams with corresponding companion streams that are used in case of random access or packet loss. Unfortunately, because a companion stream is needed for every normal stream, the distribution cost and encoding complexity are considerable costs. Additionally, injecting a companion keyframe into a normal stream causes a bitrate spike. Therefore, this paper evaluates the impact of utilizing mixed-resolution keyframe injection in the H.266/VVC standard. By providing a single companion stream for all normal streams of a bitrate ladder, the three mentioned downsides can be mitigated. We found that injecting a lower-resolution keyframe effectively reduces the bitrate spike, at the cost of only a modest quality loss. For dynamic video content, the quality impact reduces over time, and is less perceptible than traditional packet-loss repair using frame copy. In conclusion, mixed-resolution HESP can reduce the bitrate spike and computational/distribution overhead of HESP when enabling fast channel switching or packet-loss repair.
Glenn Van Wallendael, Peter Lambert, Pieter-Jan Speelmans, Hannes Mareen
PCS2
2022 Linking the cognitive load induced by route instruction types and building configuration during indoor route guidance, a usability study in VR
abstract
Every route instruction type (e.g. map, symbol, photo) induces a specific cognitive load. However, when these types are used at different decision points in a building, the building configuration of these points also influences the induced cognitive load. Therefore, the process of route guidance results in an interaction between the instruction type and the decision point, which determines the induced cognitive load. One way of reducing cognitive load during route guidance is by using adaptive systems that show specific route instruction types at specific decision points. Therefore, in this VR experiment, the usability of such an adaptive indoor route guidance system is tested by tracking the wayfinding and gaze behavior of the users. First, the difference in wayfinding and gaze behavior between all route instruction types is compared. Next, the building configuration at the decision points is quantified through the architectural theory of space syntax, and the correlation with the wayfinding and gaze behavior is determined. Our findings indicate that adapting the route instruction type does make a difference for the user.
Laure De Cock, Nico Van de Weghe, Kristien Ooms, Ignace P. Saenen, Niels Van Kets, Glenn Van Wallendael, Peter Lambert, Philippe De Maeyer
Int. J. Geogr. Inf. Sci.7
2022 Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919
abstract
Recently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community.
Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García
IEEE Trans. Multim.18
2021 An experimental study on the perceived quality of natively graded versus inverse tone mapped high dynamic range video content on television
Gonzalo Luzardo, Tine Vyvey, Jan Aelterman, Tom Paridaens, Glenn Van Wallendael, Peter Lambert, Sven Rousseaux, Hiêp Quang Luong, Wouter Durnez, Jan Van Looy, Wilfried Philips, Daniel Ochoa 0001
Multim. Tools Appl.6
2021 Camcording-Resistant Forensic Watermarking Fallback System Using Secondary Watermark Signal
abstract
Forensic watermarking is used to track down digital pirates after they illegally redistribute video content. Although existing algorithms often resist common signal processing attacks, they are not always robust against camcording attacks. As a solution in the state of the art, registration methods are used to align the attacked video to the original one. However, watermark detection still fails when the quality is sufficiently decreased or when exposed to targeted attacks. Therefore, this paper proposes a novel fallback system that aims to detect the watermark when traditional methods fail. More concretely, we demonstrate that a primary watermark embedded by a traditional scheme indirectly creates a secondary watermark signal during video encoding. This secondary watermark consists of compression artifacts and is detected by the fallback system. Additionally, the proposed system incorporates video registration to cope with camcording attacks. The experimental results indicate that the fallback system has a striking increase in robustness compared to the existing methods. For example, the observed false-negative rate for targeted attacks improves from 100% to 0%. Moreover, the fallback is camcording resistant even when the traditional method combined with registration is not. In conclusion, the proposed system can be used as a fallback when traditional detection fails.
Hannes Mareen, Martijn Courteaux, Johan De Praeter, Md. Asikuzzaman, Glenn Van Wallendael, Mark R. Pickering, Peter Lambert
IEEE Trans. Circuits Syst. Video Technol.7
2020 Random access prediction structures for light field video coding with MV-HEVC
Vasileios Avramelos, Johan De Praeter, Glenn Van Wallendael, Peter Lambert
Multim. Tools Appl.4
2020 Steered Mixture-of-Experts for Light Field Images and Video: Representation and Coding
abstract
Research in light field (LF) processing has heavily increased over the last decade. This is largely driven by the desire to achieve the same level of immersion and navigational freedom for camera-captured scenes as it is currently available for CGI content. Standardization organizations such as MPEG and JPEG continue to follow conventional coding paradigms in which viewpoints are discretely represented on 2-D regular grids. These grids are then further decorrelated through hybrid DPCM/transform techniques. However, these 2-D regular grids are less suited for high-dimensional data, such as LFs. We propose a novel coding framework for higher-dimensional image modalities, called Steered Mixture-of-Experts (SMoE). Coherent areas in the higher-dimensional space are represented by single higher-dimensional entities, called kernels. These kernels hold spatially localized information about light rays at any angle arriving at a certain region. The global model consists thus of a set of kernels which define a continuous approximation of the underlying plenoptic function. We introduce the theory of SMoE and illustrate its application for 2-D images, 4-D LF images, and 5-D LF video. We also propose an efficient coding strategy to convert the model parameters into a bitstream. Even without provisions for high-frequency information, the proposed method performs comparable to the state of the art for low-to-mid range bitrates with respect to subjective visual quality of 4-D LF images. In case of 5-D LF video, we observe superior decorrelation and coding performance with coding gains of a factor of 4x in bitrate for the same quality. At least equally important is the fact that our method inherently has desired functionality for LF rendering which is lacking in other state-of-the-art techniques: (1) full zero-delay random access, (2) light-weight pixel-parallel view reconstruction, and (3) intrinsic view interpolation and super-resolution.
Ruben Verhack, Thomas Sikora, Glenn Van Wallendael, Peter Lambert
IEEE Trans. Multim.4
2019 Improving relevant subjective testing for validation: Comparing machine learning algorithms for finding similarities in VQA datasets using objective measures
Ahmed Aldahdooh, Enrico Masala, Glenn Van Wallendael, Peter Lambert, Marcus Barkowsky
Signal Process. Image Commun.4
2019 Scalable Wavelet-Based Coding of Irregular Meshes With Interactive Region-of-Interest Support
abstract
This paper proposes a novel functionality in wavelet-based irregular mesh coding, which is interactive region-of-interest (ROI) support. The proposed approach enables the user to define the arbitrary ROIs at the decoder side and to prioritize and decode these regions at arbitrarily high-granularity levels. In this context, a novel adaptive wavelet transform for irregular meshes is proposed, which enables: 1) varying the resolution across the surface at arbitrarily fine-granularity levels and 2) dynamic tiling, which adapts the tile sizes to the local sampling densities at each resolution level. The proposed tiling approach enables a rate-distortion-optimal distribution of rate across spatial regions. When limiting the highest resolution ROI to the visible regions, the fine granularity of the proposed adaptive wavelet transform reduces the required amount of graphics memory by up to 50%. Furthermore, the required graphics memory for an arbitrary small ROI becomes negligible compared to rendering without ROI support, independent of any tiling decisions. Random access is provided by a novel dynamic tiling approach, which proves to be particularly beneficial for large models of over 106~ 107vertices. The experiments show that the dynamic tiling introduces a limited lossless rate penalty compared to an equivalent codec without ROI support. Additionally, rate savings up to 85% are observed while decoding ROIs of tens of thousands of vertices.
Jonas El Sayeh Khalil, Adrian Munteanu 0001, Peter Lambert
IEEE Trans. Circuits Syst. Video Technol.3
2019 A Scalable Architecture for Uncompressed-Domain Watermarked Videos
abstract
Video watermarking is a well-established technology to help identify digital pirates when they illegally re-distribute multimedia content. In order to provide every client with a unique, watermarked video, the traditional distribution architectures separately encode each watermarked video. However, since these encodings require a high amount of computational resources, such architectures do not scale well to a large number of users. Therefore, this paper proposes a novel architecture that uses fast encoders instead of traditional, full encoders. The fast encoders re-use the coding information from a single, previously-encoded, unwatermarked video in order to speed up the encodings of the watermarked videos. As a result, the complexity of a fast encoder is only a fraction of the complexity of a full encoder. Due to a high correlation of the re-used coding information with the optimal coding information, the compression efficiency and watermark robustness decrease only slightly. Most importantly, the proposed fast encoder speeds up the compression process with a factor of 115, resulting in a low complexity similar to that of a video decoder. Consequently, video distributors can use the proposed architecture to deliver high-quality watermarked videos on a large-scale without requiring an excessive amount of computational resources.
Hannes Mareen, Johan De Praeter, Glenn Van Wallendael, Peter Lambert
IEEE Trans. Inf. Forensics Secur.4
2018 Observer-Based Sliding Mode Control of a 6-DOF Quadrotor UAV
abstract
A sliding mode control (SMC) strategy is presented for a 6 degrees of freedom (DOF) quadrotor unmanned aerial vehicle (UAV), which achieves asymptotic position and attitude regulation. To overcome the practical limitations of velocity measurements, a sliding mode observer (SMO) is designed to estimate both the translational and rotational velocities. A rigorous Lyapunov-based analysis is provided to prove convergence to a desired set point in the presence of model uncertainties. Computer simulation results are presented, which demonstrate the effectiveness of the control law when applied to the complete nonlinear system dynamics.
Peter Lambert, Mahmut Reyhanoglu
IECON1
2018 Observer-Based Sliding Mode Control of a 2-DOF Helicopter System
abstract
A sliding mode control (SMC) strategy is presented for a 2 degrees of freedom (DOF) helicopter system, which achieves asymptotic attitude regulation to desired set points as well as trajectory tracking. A rigorous Lyapunov-based analysis is provided to prove convergence to the desired set points. To overcome the practical limitations of velocity measurements, a sliding mode observer is designed to estimate the angular velocities. The convergence of the estimated velocities to the actual velocities is proven via a rigorous Lyapunov-based analysis and illustrated through computer simulations. In addition, computer simulations are presented, to demonstrate the effectiveness of the control law when applied to the complete nonlinear system dynamics. Finally, the control strategy is experimentally tested on the Quanser 2-DOF AERO helicopter. The results of the observer-based SMC strategy are compared to a conventional PID-controller.
Peter Lambert, Mahmut Reyhanoglu
IECON1
2018 Traitor Tracing After Visible Watermark Removal
Hannes Mareen, Johan De Praeter, Glenn Van Wallendael, Peter Lambert
IWDW4
2018 Hard Real-Time, Pixel-Parallel Rendering of Light Field Videos Using Steered Mixture-of-Experts
abstract
Steered Mixture-of-Experts (SMoE) is a novel framework for the approximation, coding, and description of image modalities such as light field images and video. The future goal is to arrive at a representation for Six Degrees-of-Freedom (6DoF) image data. Previous research has shown the feasibility of real-time pixel-parallel rendering of static light field images. Each pixel is independently reconstructed by kernels that lay in its vicinity. The number of kernels involved forms the bottleneck on the achievable framerate. The goal of this paper is twofold. Firstly, we introduce pixel-level rendering of light field video, as previous work only rendered static content. Secondly, we investigate rendering using a predefined number of most significant kernels. As such, we can deliver hard real-time constraints by trading off the reconstruction quality.
Ignace P. Saenen, Ruben Verhack, Vasileios Avramelos, Glenn Van Wallendael, Peter Lambert
PCS5
2018 Progressive Modeling of Steered Mixture-of-Experts for Light Field Video Approximation
abstract
Steered Mixture-of-Experts (SMoE) is a novel framework for the approximation, coding, and description of image modalities. The future goal is to arrive at a representation for Six Degrees-of-Freedom (6DoF) image data. The goal of this paper is to introduce SMoE for 4D light field videos by including the temporal dimension. However, these videos contain vast amounts of samples due to the large number of views per frame. Previous work on static light field images mitigated the problem by hard subdividing the modeling problem. However, such a hard subdivision introduces visually disturbing block artifacts on moving objects in dynamic image data. We propose a novel modeling method that does not result in block artifacts while minimizing the computational complexity and which allows for a varying spread of kernels in the spatio-temporal domain. Experiments validate that we can progressively model light field videos with increasing objective quality up to 0.97 SSIM.
Ruben Verhack, Glenn Van Wallendael, Martijn Courteaux, Peter Lambert, Thomas Sikora
PCS4
2018 A Just Noticeable Difference Subjective Test for High Dynamic Range Images
abstract
High Dynamic Range (HDR) imaging captures a wide range of luminance existing in real-world scenes. Due to large luminance levels and higher brightness of HDR displays, artefacts can be more noticeable to the Human Visual System (HVS). In a first attempt to experimentally quantify those noticeable levels for HDR images, we pioneered in conducting an exhaustive and comprehensive Just Noticeable Difference (JND) subjective experiment of which the outcome is presented in this paper. Six distortions including JPEG, JPEG2000, noise, blur, contrast change, and quantization artefacts have been considered in the test. The distortions were applied to 10 HDR images in 100 distortion levels resulting a database of 6000 HDR test images. The subjects were asked to find the image JND location on each set of 100 images they had the freedom to explore. The effect of content features on the noticeable threshold selection is investigated per distortion type. Our results in some cases show a significant correlation between content features and JNDs. We are hoping that our results can contribute to further exploitation of a precise HVS model for HDR quality assessment and optimization of the coding and bit allocation in HDR compression.
Ayyoub Ahar, Saeed Mahmoudpour, Glenn Van Wallendael, Tom Paridaens, Peter Lambert, Peter Schelkens
QoMEX5
2018 AQUa: an adaptive framework for compression of sequencing quality scores with random access functionality
abstract
Motivation: The past decade has seen the introduction of new technologies that significantly lowered the cost of genome sequencing. As a result, the amount of genomic data that must be stored and transmitted is increasing exponentially. To mitigate storage and transmission issues, we introduce a framework for lossless compression of quality scores. Results: This article proposes AQUa, an adaptive framework for lossless compression of quality scores. To compress these quality scores, AQUa makes use of a configurable set of coding tools, extended with a Context-Adaptive Binary Arithmetic Coding scheme. When benchmarking AQUa against generic single-pass compressors, file sizes are reduced by up to 38.49% when comparing with GNU Gzip and by up to 6.48% when comparing with 7-Zip at the Ultra Setting, while still providing support for random access. When comparing AQUa with the purpose-built, single-pass, and state-of-the-art compressor SCALCE, which does not support random access, file sizes are reduced by up to 21.14%. When comparing AQUa with the purpose-built, dual-pass, and state-of-the-art compressor QVZ, which does not support random access, file sizes are larger by 6.42-33.47%. However, for one test file, the file size is 0.38% smaller, illustrating the strength of our single-pass compression framework. This work has been spurred by the current activity on genomic information representation (MPEG-G) within the ISO/IEC SC29/WG11 technical committee. Availability and implementation: The software is available on Github: https://github.com/tparidae/AQUa. Contact: [email protected].
Tom Paridaens, Glenn Van Wallendael, Wesley De Neve, Peter Lambert
Bioinform.4
2018 The crowd as a cameraman: on-stage display of crowdsourced mobile video at large-scale events
Steven Bohez, Glenn Daneels, Lander Van Herzeele, Niels Van Kets, Sam Decrock, Matthias De Geyter, Glenn Van Wallendael, Peter Lambert, Bart Dhoedt, Pieter Simoens, Steven Latré, Jeroen Famaey
Multim. Tools Appl.8
2017 Color prediction in image coding using Steered Mixture-of-Experts
abstract
We propose a novel approach for modeling and coding color in images and video. Luminance is linearly correlated with chrominance locally, as such we can predict color given the luma value. Using the Steered Mixture-of-Experts (SMoE) approach, the image is viewed as a stochastic process over 5 random variables including the 2-D pixel locations, 1 luminance and 2 chrominance values. We model this process as a continuous joint density function by fitting a K-modal 5-D Gaussian Mixture Model (GMM). As such, the chroma values are predicted as the expectation of the conditional density. To validate, the technique was integrated within JPEG showing PSNR gains in the lower bitrate regions. A deeper analysis of the tolerance of the activation function is given through recycling color models in video sequences, yielding a high quality reconstruction over a considerable range of frames.
Ruben Verhack, Simon Van De Keer, Glenn Van Wallendael, Thomas Sikora, Peter Lambert
ICASSP5
2017 3D Mesh coding with predefined region-of-interest
abstract
We introduce a novel functionality for wavelet-based irregular mesh codecs which allows for prioritizing at the encoding side a region-of-interest (ROI) over a background (BG), and for transmitting the encoded data such that the quality in these regions increases first. This is made possible by appropriately scaling wavelet coefficients. To improve the decoded geometry in the BG, we propose an ROI-aware inverse wavelet transform which only upscales the connectivity in the required regions. Results show clear bitrate and vertex savings. For a trivial front-back selection of the ROI and BG, rendering from the front saves up to 5 bits per vertex and up to 50% of the geometry, while appearing visually lossless.
Jonas El Sayeh Khalil, Adrian Munteanu 0001, Peter Lambert
ICIP3
2017 Steered mixture-of-experts for light field coding, depth estimation, and processing
abstract
The proposed framework, called Steered Mixture-of-Experts (SMoE), enables a multitude of processing tasks on light fields using a single unified Bayesian model. The underlying assumption is that light field rays are instantiations of a non-linear or non-stationary random process that can be modeled by piecewise stationary processes in the spatial domain. As such, it is modeled as a space-continuous Gaussian Mixture Model. Consequently, the model takes into account different regions of the scene, their edges, and their development along the spatial and disparity dimensions. Applications presented include light field coding, depth estimation, edge detection, segmentation, and view interpolation. The representation is compact, which allows for very efficient compression yielding state-of-the-art coding results for low bit-rates. Furthermore, due to the statistical representation, a vast amount of information can be queried from the model even without having to analyze the pixel values. This allows for “blind” light field processing and classification.
Ruben Verhack, Thomas Sikora, Lieven Lange, Rolf Jongebloed, Glenn Van Wallendael, Peter Lambert
ICME6
2017 AFRESh: an adaptive framework for compression of reads and assembled sequences with random access functionality
abstract
MOTIVATION: The past decade has seen the introduction of new technologies that lowered the cost of genomic sequencing increasingly. We can even observe that the cost of sequencing is dropping significantly faster than the cost of storage and transmission. The latter motivates a need for continuous improvements in the area of genomic data compression, not only at the level of effectiveness (compression rate), but also at the level of functionality (e.g. random access), configurability (effectiveness versus complexity, coding tool set …) and versatility (support for both sequenced reads and assembled sequences). In that regard, we can point out that current approaches mostly do not support random access, requiring full files to be transmitted, and that current approaches are restricted to either read or sequence compression. RESULTS: We propose AFRESh, an adaptive framework for no-reference compression of genomic data with random access functionality, targeting the effective representation of the raw genomic symbol streams of both reads and assembled sequences. AFRESh makes use of a configurable set of prediction and encoding tools, extended by a Context-Adaptive Binary Arithmetic Coding scheme (CABAC), to compress raw genetic codes. To the best of our knowledge, our paper is the first to describe an effective implementation CABAC outside of its' original application. By applying CABAC, the compression effectiveness improves by up to 19% for assembled sequences and up to 62% for reads. By applying AFRESh to the genomic symbols of the MPEG genomic compression test set for reads, a compression gain is achieved of up to 51% compared to SCALCE, 42% compared to LFQC and 44% compared to ORCOM. When comparing to generic compression approaches, a compression gain is achieved of up to 41% compared to GNU Gzip and 22% compared to 7-Zip at the Ultra setting. Additionaly, when compressing assembled sequences of the Human Genome, a compression gain is achieved up to 34% compared to GNU Gzip and 16% compared to 7-Zip at the Ultra setting. AVAILABILITY AND IMPLEMENTATION: A Windows executable version can be downloaded at https://github.com/tparidae/AFresh . CONTACT: [email protected].
Tom Paridaens, Glenn Van Wallendael, Wesley De Neve, Peter Lambert
Bioinform.4
2017 Scalable Feature-Preserving Irregular Mesh Coding
abstract
Abstract This paper presents a novel wavelet‐based transform and coding scheme for irregular meshes. The transform preserves geometric features at lower resolutions by adaptive vertex sampling and retriangulation, resulting in more accurate subsampling and better avoidance of smoothing and aliasing artefacts. By employing octree‐based coding techniques, the encoding of both connectivity and geometry information is decoupled from any mesh traversal order, and allows for exploiting the intra‐band statistical dependencies between wavelet coefficients. Improvements over the state of the art obtained by our approach are three‐fold: (1) improved rate–distortion performance over Wavemesh and IPR for both the Hausdorff and root mean square distances at low‐to‐mid‐range bitrates, most obvious when clear geometric features are present while remaining competitive for smooth, feature‐poor models; (2) improved rendering performance at any triangle budget, translating to a better quality for the same runtime memory footprint; (3) improved visual quality when applying similar limits to the bitrate or triangle budget, showing more pronounced improvements than rate–distortion curves.
Jonas El Sayeh Khalil, Adrian Munteanu 0001, Leon Denis, Peter Lambert, Rik Van de Walle
Comput. Graph. Forum4
2017 Video Encoder Architecture for Low-Delay Live-Streaming Events
abstract
Video-streaming events such as virtual classrooms and video conferences require a low delay between sender and receiver. In order to achieve this requirement, and to make full use of the bandwidth capacity of each receiver, each client can be provided with a personalized bitstream of which the bit rate is continuously adapted to his current network bandwidth capacity. However, such an approach requires an excessive amount of computationally complex video encoders. Therefore, this paper proposes an architecture based on coding information calculation (CIC) modules and residual encoder (RE) modules. The CIC modules calculate coding information for the video at certain bit rates whereas the RE modules use this information to skip all encoding steps of a traditional encoder, except for the encoding of the residual. By reducing the amount of bits used to encode the residual, the RE modules can then provide bitstreams with personalized bit rates for several users at the same time. Each CIC module has approximately the same computational complexity as a traditional encoder, whereas an RE module has the approximate complexity of a decoder. The proposed architecture was evaluated for the high efficiency video coding standard, showing that the system achieves its goal of drastically reducing the computational complexity of low-delay live-streaming with many participants and suggesting that using less than six CIC modules results in the best tradeoff between compression efficiency and computational complexity.
Johan De Praeter, Glenn Van Wallendael, Jürgen Slowack, Peter Lambert
IEEE Trans. Multim.4
2016 Leveraging CABAC for No-Reference Compression of Genomic Data with Random Access Support
abstract
In previous work, the authors developed a modular no-reference framework that compresses FASTA files by applying a predict-and-residue method, as used in video coding. We extended this framework with support for Context-Adaptive Binary Arithmetic Coding (CABAC), while at the same time preserving random access functionality and offering support for the full IUB/IUPAC nucleic acid codes alphabet.
Tom Paridaens, Jens Panneel, Wesley De Neve, Peter Lambert, Rik Van de Walle
DCC4
2016 Low Delay Complexity Constrained Encoding
abstract
Complex software appliances typically consist of multiple software processes running concurrently to exploit the available computational resources in the hardware. However, the computational complexity of these software processes is often variable and the processes can interfere with each other. This can be an issue for real-time applications with a fixed deadline like low delay video encoding. In the context of High Efficiency Video Coding (HEVC), a limited number of publications have focused on controlling the complexity of an HEVC video encoder. In this paper, we propose a technique to control complexity by deciding between 2Nx2N merge mode and full encoding, at different Coding Unit (CU) depths. Our results demonstrate fast convergence to a given complexity threshold after a maximum of 10 frames, and a limited loss in rate-distortion performance (on average 2.84% Bjontegaard delta rate for 60% complexity reduction).
Thijs Vermeir, Jürgen Slowack, Glenn Van Wallendael, Peter Lambert, Rik Van de Walle
DCC4
2016 A universal image coding approach using sparse steered Mixture-of-Experts regression
abstract
Our challenge is the design of a “universal” bit-efficient image compression approach. The prime goal is to allow reconstruction of images with high quality. In addition, we attempt to design the coder and decoder “universal”, such that MPEG-7-like low-and mid-level descriptors are an integral part of the coded representation. To this end, we introduce a sparse Mixture-of-Experts regression approach for coding images in the pixel domain. The underlying stochastic process of the pixel amplitudes are modelled as a 3-dimensional and multi-modal Mixture-of-Gaussians with K modes. This closed form continuous analytical model is estimated using the Expectation-Maximization algorithm and describes segments of pixels by local 3-D Gaussian steering kernels with global support. As such, each component in the mixture of experts steers along the direction of highest correlation. The conditional density then serves as the regression function. Experiments show that a considerable compression gain is achievable compared to JPEG for low bitrates for a large class of images, while forming attractive low-level descriptors for the image, such as the local segmentation boundaries, direction of intensity flow and the distribution of these parameters over the image.
Ruben Verhack, Thomas Sikora, Lieven Lange, Glenn Van Wallendael, Peter Lambert
ICIP5
2016 Real-time complexity constrained encoding
abstract
Complex software appliances can be deployed on hardware with limited available computational resources. This computational boundary puts an additional constraint on software applications. This can be an issue for real-time applications with a fixed time constraint such as low delay video encoding. In the context of High Efficiency Video Coding (HEVC), a limited number of publications have focused on controlling the complexity of an HEVC video encoder. In this paper, a technique is proposed to control complexity by deciding between 2N×2N merge mode and full encoding, at different Coding Unit (CU) depths. The technique is demonstrated in two encoders. The results demonstrate fast convergence to a given complexity threshold, and a limited loss in rate-distortion performance (on average 2.84% Bjontegaard delta rate for 40% complexity reduction).
Thijs Vermeir, Jürgen Slowack, Glenn Van Wallendael, Peter Lambert, Rik Van de Walle
ICIP4
2016 Multistream video encoder for generating multiple dynamic range bitstreams
abstract
High-dynamic-range (HDR) technology allows capturing of video content at a wider range of luminance than low-dynamic-range (LDR) video. The resulting video more closely resembles the scene as perceived by the human eye. However, displays currently support only a limited range of HDR. Therefore, both an HDR version and LDR version of a video should be encoded during content acquisition. This means that the cost of encoder hardware in cameras would double. As a solution, this paper proposes a multistream video encoder that allows generating an HDR and LDR version of the same HDR video footage at roughly the same computational complexity as a single encoder, effectively allowing encoding of two dynamic-range versions of the video with a negligible increase in cost. For the LDR version, this multistream encoder results in a bit rate overhead of only 11.6% for the same quality as a two-encoder solution.
Cedric Van Goethem, Johan De Praeter, Tom Paridaens, Glenn Van Wallendael, Peter Lambert
PCS5
2016 Perceptual quality of 4K-resolution video content compared to HD
abstract
With the introduction of 4K UHD video and display resolution, questions arise on the perceptual differences between 4K UHD and upsampled HD video content. In this paper, a striped pair comparison has been performed on a diverse set of 4K UHD video sources. The goal was to subjectively assess the perceived sharpness of 4K UHD and downscaled/upscaled HD video. A striped pair comparison has been applied in order to make the test as straightforward as possible for a non-expert participant population. Under these conditions and over this set of sequences, on average, on 54.8% of the sequences (17 out of 31), 4K UHD resolution content could be identified as being sharper compared to its HD down and upsampled alternative. The probabilities in which 4K UHD could be differentiated from downscaled/upscaled HD range from 83.3% for the easiest to assess sequence down to 39.7% for the most difficult sequence. Although significance tests demonstrate there is a positive sharpness difference from camera quality 4K UHD content compared to the HD downscaled/upscaled variations, it is very content dependent and all circumstances have been chosen in favor of the 4K UHD representation. The results of this test can contribute to the research process of developing metrics indicating visibility of high resolution features within specific content.
Glenn Van Wallendael, Paulien Coppens, Tom Paridaens, Niels Van Kets, Wendy Van den Broeck, Peter Lambert
QoMEX6
2016 Spatially misaligned HEVC transcoding with computational-complexity scalability
Johan De Praeter, Glenn Van Wallendael, Thijs Vermeir, Jürgen Slowack, Peter Lambert
J. Vis. Commun. Image Represent.5
2015 Fast simultaneous video encoder for adaptive streaming
abstract
Content providers create different versions of a video to accommodate different end-user devices and network conditions. However, each of these versions requires a resource intensive encoding process. To reduce the computational complexity of the encodings, this paper proposes a fast simultaneous encoder. This encoder takes a single video as input and creates a number of bit streams encoded with different parameters. Only one version of the video is created with a full encode, whereas encoding of the other versions is accelerated by exploiting the correlation with the fully encoded version using machine learning techniques. In a practical scenario, the fast simultaneous encoder achieves a complexity reduction of 67.3% with a bit rate increase of 5.2% compared to performing a full encode of each version.
Johan De Praeter, Antonio Jesús Díaz-Honrubia, Niels Van Kets, Glenn Van Wallendael, Jan De Cock, Peter Lambert, Rik Van de Walle
MMSP6
2015 Lossless image compression based on Kernel Least Mean Squares
abstract
This paper introduces a novel approach for coding luminance images using kernel-based adaptive filtering and context-adaptive arithmetic coding. This approach tackles the problem that is present in current image and video coders; these coders depend on assumptions of the image and are constrained by the linearity of their predictors. The efficacy of the predictors determines the compression gain. The goal is to create a generic image coder that learns and adapts to the characteristics of the signals and handles nonlinearity in the prediction. Results show that pixel luminance prediction using the Kernel Least Mean Squares (KLMS) yields a significant gain compared to the standard Least Mean Squares algorithm. By coding the residual using a Context-Adaptive Arithmetic Coder (CAAC), the codec is able to outperform the current industry standards of lossless image coding. An average bitrate reduction of more than 2.5% is found for the used test set.
Ruben Verhack, Lieven Lange, Peter Lambert, Rik Van de Walle, Thomas Sikora
PCS3
2014 Efficient transcoding for spatially misaligned compositions for HEVC
abstract
The visualization of (ultra) high-resolution compositions created from multiple input bitstreams requires several decoders at the receiving device. Therefore, not all devices can properly display such compositions. To address this problem, the input streams are decoded, merged into a single video, and re-encoded by a transcoder in the network. However, this approach requires a computationally complex re-encoding step. To reduce this complexity, information from the input bit-streams can be reused during transcoding. In HEVC, simply reusing the original encoding information is not compression efficient if the inserted content is not aligned with the grid of coded blocks. In this paper, we applied feature selection based on a decision tree, which was used in a fast HEVC transcoding model for misaligned content. The performance varies depending on the shift and average transform size in the original sequence, resulting in complexity reductions of up to 76%.
Johan De Praeter, Jan De Cock, Glenn Van Wallendael, Sebastiaan Van Leuven, Peter Lambert, Rik Van de Walle
ICIP5
2014 Lossy image coding in the pixel domain using a sparse steering kernel synthesis approach
abstract
Kernel regression has been proven successful for image de-noising, deblocking and reconstruction. These techniques lay the foundation for new image coding opportunities. In this paper, we introduce a novel compression scheme: Sparse Steering Kernel Synthesis Coding (SSKSC). This pre- and postprocessor for JPEG performs non-uniform sampling based on the smoothness of an image, and reconstructs the missing pixels using adaptive kernel regression. At the same time, the kernel regression reduces the blocking artifacts from the JPEG coding. Crucial to this technique is that non-uniform sampling is performed while maintaining only a small overhead for signalization. Compared to JPEG, SSKSC achieves a compression gain for low bits-per-pixel regions of 50% or more for PSNR and SSIM. A PSNR gain is typically in the 0.0-0.5 bpp range, and an SSIM gain can mostly be achieved in the 0.0-1.0 bpp range.
Ruben Verhack, Andreas Krutz, Peter Lambert, Rik Van de Walle, Thomas Sikora
ICIP3
2014 Fast channel switching for single-loop scalable HEVC
abstract
In an IPTV environment, different techniques exist to provide faster random access or equivalently, a faster channel switching experience, but none of them provide backward compatibility and limited overhead at the same time. In this paper, a technique is proposed to increase the random access frequency for the High Efficiency Video Coding (HEVC) standard by using a single-loop scalable version of it. This is achieved by encoding an enhancement layer with a higher frequency of random access points compared to the base layer. With this scalable configuration, a backward compatible base stream remains present. The proposed technique requires 12.4% less bitrate compared to sending the fast and slow switching streams in a non-scalable way. Compared to the almost standardized multi-loop scalable extension of HEVC, 5.5% less bitrate is needed on the core of the IPTV network and 12.5% less bitrate on the access network during steady state conditions. Moreover, the multi-loop scalable extension does not provide backward compatibility with legacy HEVC decoders.
Glenn Van Wallendael, Nicolas Staelens, Sebastiaan Van Leuven, Jan De Cock, Peter Lambert, Piet Demeester, Rik Van de Walle
ICIP5
2014 Multi-modal time-of-flight based fire detection
Steven Verstockt, Sofie Van Hoecke, Pieterjan De Potter, Peter Lambert, Charles-Frederik Hollemeersch, Bart Sette, Bart Merci, Rik Van de Walle
Multim. Tools Appl.4
2013 Fast transrating for high efficiency video coding based on machine learning
abstract
To incorporate the newly developed High Efficiency Video Coding (HEVC) standard in real-life network applications, efficient transrating algorithms are required. We propose a fast transrating scheme, based on the early prediction of the partition split-flags in P pictures. Using machine learning techniques, the correlation between co-located partitions at different quantizations is investigated. This results in a model which predicts the split-flag and gives the associated prediction accuracy so that the splitting process in the transcoder is optimized. At each partition depth, the model indicates whether the full rate-distortion cost evaluations should be performed at the current depth, or if the partition can be split immediately. Experimental results show that the proposed transcoder reduces the complexity of the transrating process by 76.04%, while maintaining the coding efficiency of a cascaded decoder-encoder.
Luong Pham Van, Jan De Cock, Glenn Van Wallendael, Sebastiaan Van Leuven, Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Peter Lambert, Rik Van de Walle
ICIP7
2013 Format-compliant encryption techniques for high efficiency video coding
abstract
When middlebox devices should be able to adapt an encrypted video stream in the network without having the decryption key, format-compliant partial encryption schemes should be applied. In this paper, we propose such encryption schemes for the recently standardized High Efficiency Video Coding (HEVC) standard. By encrypting specific syntax elements like the sign of the residual information, the sign of the motion vector (MV) difference, the MV prediction index, and the MV reference index, format compliance and the possibility for adaptation are offered. Scrambling performance gradually increases when shifting from encrypting the motion information to encrypting the residual sign and finally to the combination thereof. Applying all these techniques has a negligible impact on the compression efficiency.
Glenn Van Wallendael, Jan De Cock, Sebastiaan Van Leuven, Andras Boho, Peter Lambert, Bart Preneel, Rik Van de Walle
ICIP5
2013 Sampling in transform domain for improved QoE of 3D frame-compatible video coding
Jan De Cock, Peter Lambert, Rik Van de Walle
IM3
2013 Evaluation of full-reference objective video quality metrics on high efficiency video coding
Glenn Van Wallendael, Sebastiaan Van Leuven, Jan De Cock, Peter Lambert, Rik Van de Walle, Nicolas Staelens, Piet Demeester
IM4
2012 Parallel deblocking filtering in H.264/AVC using multiple CPUs and GPUs
abstract
Deblocking filtering in the H.264/AVC standard is a computationally complex process because of the filter's high content adaptivity. Furthermore, the deblocking filter introduces a significant number of data dependencies, making parallel processing not obvious. Our previous works analyzed the dependencies of the filter and proposed a massively-parallel implementation, specifically tailored for execution on a single GPU. In this paper, we extend this work by proposing a parallel processing scheme for accelerating deblocking filtering using multiple CPU cores or GPUs. This scheme allows for standard-compliant filtering, regardless of slice configuration. Results show that our multi-GPU implementation using our proposed scheme achieves faster-than real-time deblocking at over 3794 frames per second for 1080p video pictures by using three GPUs. A multi-core CPU implementation using 8 CPU cores allows 1080p deblocking filtering of up to 695 frames per second.
Bart Pieters, Charles-Frederik Hollemeersch, Jan De Cock, Wesley De Neve, Peter Lambert, Rik Van de Walle
ACM Multimedia5
2012 Real-Time Visualizations of Gigapixel Texture Data Sets Using HTML5
Charles-Frederik Hollemeersch, Bart Pieters, Aljosha Demeulemeester, Peter Lambert, Rik Van de Walle
MMM4
2012 Feedback-constrained Wyner-Ziv video coding
abstract
Distributed video coding (DVC) systems described in the literature often make use of a feedback channel to determine the rate. However, supporting such a feedback channel in practice may be difficult, particularly considering that current approaches are unable to incorporate constraints on feedback channel usage. Therefore, in this paper we propose a system in which the number of requests per Wyner-Ziv frame can be constrained to a fixed value. Our technique involves decoder-side rate estimation, modeling of its accuracy, and finally defining the number of Wyner-Ziv bits for each of the requests allowed. The experimental results indicate that the performance loss is limited when compared to a configuration operating without constraints on the feedback channel.
Jürgen Slowack, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
PCS3
2012 Efficient disparity vector prediction schemes with modified P frame for 2D camera arrays
Aykut Avci, Jan De Cock, Peter Lambert, Roel Beernaert, Jelle De Smet, Lawrence Bogaert, Youri Meuret, Hugo Thienpont, Herbert De Smet
J. Vis. Commun. Image Represent.3
2012 On the performance of scalable video coding for VBR TV channels transport in multiple resolutions and qualities
Zlatka Avramova, Danny De Vleeschauwer, Pedro Debevere, Sabine Wittevrongel, Peter Lambert, Rik Van de Walle, Herwig Bruneel
Multim. Tools Appl.5
2012 An enhanced fast mode decision model for spatial enhancement layers in scalable video coding
Sebastiaan Van Leuven, Glenn Van Wallendael, Koen De Wolf, Jan De Cock, Peter Lambert, Rik Van de Walle
Multim. Tools Appl.5
2012 Decoder-driven mode decision in a block-based distributed video codec
Stefaan Mys, Jürgen Slowack, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
Multim. Tools Appl.5
2012 Efficient adaptive-shape partitioning of video
Kenneth Vermeirsch, Jan De Cock, Stijn Notebaert, Peter Lambert, Joeri Barbarien, Adrian Munteanu 0001, Rik Van de Walle
Multim. Tools Appl.4
2012 Silhouette-based multi-sensor smoke detection - Coverage analysis of moving object silhouettes in thermal and visual registered images
Steven Verstockt, Chris Poppe, Sofie Van Hoecke, Charles-Frederik Hollemeersch, Bart Merci, Bart Sette, Peter Lambert, Rik Van de Walle
Mach. Vis. Appl.7
2012 Data-parallel intra decoding for block-based image and video coding on massively parallel architectures
Bart Pieters, Charles-Frederik Hollemeersch, Jan De Cock, Peter Lambert, Rik Van de Walle
Signal Process. Image Commun.4
2012 Efficient Low-Delay Distributed Video Coding
abstract
Distributed video coding (DVC) is a video coding paradigm that allows for a low-complexity encoding process by exploiting the temporal redundancies in a video sequence at the decoder side. State-of-the-art DVC systems exhibit a structural coding delay since exploiting the temporal redundancies through motion-compensated interpolation requires the frames to be decoded out of order. To alleviate this problem, we propose a system based on motion-compensated extrapolation that allows for efficient low-delay video coding with low complexity at the encoder. The proposed extrapolation technique first estimates the motion field between the two most recently decoded frames using the Lucas–Kanade algorithm. The obtained motion field is then extrapolated to the current frame using an extrapolation grid. The proposed techniques are implemented into a novel architecture featuring hybrid block-frequency Wyner–Ziv coding as well as mode decision. Results show that having references from both temporal directions in interpolation provides superior rate-distortion performance over a single temporal direction in extrapolation, as expected. However, the proposed extrapolation method is particularly suitable for low-delay coding as it performs better than H.264/AVC intra, and it is even able to outperform the interpolation-based DVC codec from DISCOVER for several sequences.
Jozef Skorupa, Jürgen Slowack, Stefaan Mys, Nikos Deligiannis, Jan De Cock, Peter Lambert, Christos Grecos, Adrian Munteanu 0001, Rik Van de Walle
IEEE Trans. Circuits Syst. Video Technol.6
2012 Distributed Video Coding With Feedback Channel Constraints
abstract
Many of the distributed video coding (DVC) systems described in the literature make use of a feedback channel from the decoder to the encoder to determine the rate. However, the number of requests through the feedback channel is often high, and as a result the overall delay of the system could be unacceptable in practical applications. As a solution, feedback-free DVC systems have been proposed, but the problem with these solutions is that they incorporate a difficult trade-off between encoder complexity and compression performance. Recognizing that a limited form of feedback may be supported in many video-streaming scenarios, in this paper we propose a method for constraining the number of feedback requests to a fixed maximum number of$N$requests for an entire Wyner-Ziv (WZ) frame. The proposed technique estimates the WZ rate at the decoder using information obtained from previously decoded WZ frames and defines the$N$requests by minimizing the expected rate overhead. Tests on eight sequences show that the rate penalty is less than 5% when only five requests are allowed per WZ frame (for a group of pictures of size four). Furthermore, due to improvements from previous work, the system is able to perform better than or similar to DISCOVER even when up to two requests per WZ frame are allowed. The practical usefulness of the proposed approach is studied by estimating end-to-end delay and encoder buffer requirements, indicating that DVC with constrained feedback can be an important solution in the context of video-streaming scenarios.
Jürgen Slowack, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
IEEE Trans. Circuits Syst. Video Technol.4
2012 A new approach to combine texture compression and filtering
Charles-Frederik Hollemeersch, Bart Pieters, Peter Lambert, Rik Van de Walle
Vis. Comput.3
2011 Region-adaptive probability model selection for the arithmetic coding of video texture
abstract
In video coding systems using adaptive arithmetic coding to compress texture information, the employed symbol probability models need to be retrained every time the coding process moves into an area with different texture. To avoid this inefficiency, we propose to replace the probability models used in the original coder with multiple switchable sets of probability models. We determine the model set to use in each spatial region in an optimal manner, taking into account the additional signaling overhead. Experimental results show that this approach, when applied to H.264/AVC's context-based adaptive binary arithmetic coder (CABAC), yields significant bit-rate savings, which are comparable to or higher than those obtained using alternative improvements to CABAC previously proposed in the literature.
Kenneth Vermeirsch, Joeri Barbarien, Peter Lambert, Rik Van de Walle
ICASSP3
2011 Ultra High Definition video decoding with Motion JPEG XR using the GPU
abstract
Many applications require real-time decoding of high-resolution video pictures, for example, quick editing of video sequences in video editing applications. To increase decoding speed, parallelism can be exploited, yet, block-based image and video coding standards are difficult to decode in parallel because of the high number of dependencies between blocks. This paper investigates the parallel decoding capabilities of the new JPEG XR image coding standard for use on the massively-parallel architecture of the GPU. The potential of parallelism of the hierarchical frequency coding scheme used in the standard is addressed and a parallel decoding scheme is described suitable for real-time decoding of Ultra High Definition (4320p) Motion JPEG XR video sequences. Our results show a decoding speed of up to 46 frames per second for Ultra High Definition (4320p) sequences with high-dynamic range (32-bit/4:2:0) luma and chroma components.
Bart Pieters, Jan De Cock, Charles-Frederik Hollemeersch, Jeroen Wielandt, Peter Lambert, Rik Van de Walle
ICIP5
2011 Intra-WZ quantization mismatch in distributed video coding
abstract
During the past decade, Distributed Video Coding (DVC) has emerged as a new video coding paradigm, shifting the complexity from the encoder - to the decoder-side. This paper addresses a problem of current DVC architectures that has not been studied in the literature so far, that is, the mismatch between the intra and Wyner-Ziv (WZ) quantization processes. Due to this mismatch, WZ rate is spent even for spatial regions that are accurately approximated by the side-information. As a solution, this paper proposes side-information generation using selective unidirectional motion compensation from temporally adjacent WZ frames. Experimental results show that the proposed approach yields promising WZ rate gains of up to 7% relative to the conventional method.
Jürgen Slowack, Jozef Skorupa, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001
ICIP3
2011 Multi-sensor fire detection using visual and time-of-flight imaging
abstract
This paper proposes a novel multi-sensor fire detection method based on ordinary video images and the amplitude images of a time-of-flight camera. Using this multi-modal in formation, flame regions can be detected very accurately. Regions with high accumulative amplitude differences and high values in all detail images of the amplitude image its discrete wavelet transform, are labeled as candidate flame regions. Simultaneously, moving objects in the visual images are also investigated. Objects which possess the experimentally found low-cost flame features are also labeled as candidate flame region. Finally, if one of the visual and amplitude candidate flame regions overlap, fire alarm is given. Experiments show that the proposed detector has an average flame detection rate of 92% with no false positive detections.
Steven Verstockt, Pieterjan De Potter, Sofie Van Hoecke, Peter Lambert, Rik Van de Walle
ICME4
2011 Improved intra mode signaling for HEVC
abstract
In the current development of HEVC, compression performance improved significantly compared to H.264/AVC for both inter pictures and intra pictures. With intra compression, the main reason for this improvement is the large in crease in intra prediction directions (up to 34). The downside of having a larger number of modes is that they increase the signaling overhead in the bitstream. In this paper, a low complexity intra mode prediction algorithm is proposed which improves the mode prediction accuracy. This is achieved by exploiting the correlation between the prediction directions of the neighboring prediction units and that of the encoded prediction unit. As a result, more efficient intra mode signaling can be achieved with minimal impact on encoder and decoder complexity. On average, 0.33% bitrate improvement is obtained by employing the proposed algorithm. For sequences that are encoded with a high number of directional intra modes, around 1% bitrate improvement is measured.
Glenn Van Wallendael, Sebastiaan Van Leuven, Jan De Cock, Peter Lambert, Rik Van de Walle, Joeri Barbarien, Adrian Munteanu 0001
ICME4
2011 Hybrid Path Planning for Massive Crowd Simulation on the GPU
Aljosha Demeulemeester, Charles-Frederik Hollemeersch, Pieter Mees, Bart Pieters, Peter Lambert, Rik Van de Walle
MIG5
2011 Compressed-Domain Shot Boundary Detection for H.264/AVC Using Intra Partitioning Maps
Sarah De Bruyne, Jan De Cock, Chris Poppe, Charles-Frederik Hollemeersch, Peter Lambert, Rik Van de Walle
MMM (1)5
2011 Immersive Video Conferencing Architecture Using Game Engine Technology
Chris Poppe, Charles-Frederik Hollemeersch, Sarah De Bruyne, Peter Lambert, Rik Van de Walle
MMM (2)4
2011 Motion-refined rewriting of H.264/AVC-coded video to SVC streams
Jan De Cock, Stijn Notebaert, Peter Lambert, Rik Van de Walle
J. Vis. Commun. Image Represent.3
2011 On the performance of scalable video coding for VBR TV channels transport in multiple resolutions and qualities
Zlatka Avramova, Danny De Vleeschauwer, Pedro Debevere, Sabine Wittevrongel, Peter Lambert, Rik Van de Walle, Herwig Bruneel
Multim. Tools Appl.5
2011 Quantizer offset selection for improved requantization transcoding
Stijn Notebaert, Jan De Cock, Kenneth Vermeirsch, Peter Lambert, Rik Van de Walle
Signal Process. Image Commun.4
2011 Parallel Deblocking Filtering in MPEG-4 AVC/H.264 on Massively Parallel Architectures
abstract
The deblocking filter in the MPEG-4 AVC/H.264 standard is computationally complex because of its high content adaptivity, resulting in a significant number of data dependencies. These data dependencies interfere with parallel filtering of multiple macroblocks (MBs) on massively parallel architectures. In this letter, we introduce a novel MB partitioning scheme for concurrent deblocking in the MPEG-4 AVC/H.264 standard, based on our idea of deblocking filter independency, a corrected version of the limited error propagation effect proposed in the letter. Our proposed scheme enables concurrent MB deblocking of luma samples with limited synchronization effort, independently of slice configuration, and is compliant with the MPEG-4 H.264/AVC standard. We implemented the method on the massively parallel architecture of the graphics processing unit (GPU). Experimental results show that our GPU implementation achieves faster-than real-time deblocking at 1309 frames per second for 1080p video pictures. Both software-based deblocking filters and state-of-the-art GPU-enabled algorithms are outperformed in terms of speed by factors up to 10.2 and 19.5, respectively, for 1080p video pictures.
Bart Pieters, Charles-Frederik Hollemeersch, Jan De Cock, Peter Lambert, Wesley De Neve, Rik Van de Walle
IEEE Trans. Circuits Syst. Video Technol.4
2010 Bitplane intra coding with decoder-side mode decision in distributed video coding
abstract
While distributed video coding (DVC) has emerged as a new video coding paradigm, the compression performance of current systems is still low compared to conventional solutions such as H.264/AVC. While the latter uses many coding modes and an efficient mode decision strategy for choosing the best mode, in DVC, only a limited number of modes has been developed so far. Since encoder-side mode decision in DVC increases encoder's complexity, in this paper, we introduce decoder-side mode decision choosing between bitplane WZ coding and bitplane intra coding. This strategy proves to be efficient, delivering rate gains up to 22% over DISCOVER, without increasing the complexity at the encoder.
Jürgen Slowack, Stefaan Mys, Jozef Skorupa, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001
ICIP4
2010 Compensating for Motion Estimation Inaccuracies in DVC
Jürgen Slowack, Jozef Skorupa, Stefaan Mys, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
ICISP5
2010 Multi-sensor Fire Detection by Fusing Visual and Non-visual Flame Features
Steven Verstockt, Alexander Vanoosthuyse, Sofie Van Hoecke, Peter Lambert, Rik Van de Walle
ICISP4
2010 Correlation modeling with decoder-side quantization distortion estimation for distributed video coding
abstract
Aiming for low-complexity encoding, distributed video coders still fail to achieve the performance of current industrial standards for video coding. One of most important problems in this area is the accurate modeling of the correlation between the predicted signal and the original video. In our previous work we showed that exploiting the quantization distortion can significantly improve the accuracy of a correlation estimator. In this paper we describe how the quantization distortion can be exploited purely at the decoder side without any performance penalty when compared to an encoder-aided system. As a result, the proposed correlation estimator delivers state-of-the-art modeling accuracy while neatly fitting the low-encoder-complexity characteristic of distributed video coding.
Jozef Skorupa, Jan De Cock, Jürgen Slowack, Stefaan Mys, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001
PCS5
2010 Infinitex: An interactive editing system for the production of large texture data sets
Charles-Frederik Hollemeersch, Bart Pieters, Aljosha Demeulemeester, Frederik Cornillie, Bert Van Semmertier, Erik Mannens, Peter Lambert, Piet Desmet, Rik Van de Walle
Comput. Graph.7
2010 Dyadic spatial resolution reduction transcoding for H.264/AVC
Jan De Cock, Stijn Notebaert, Kenneth Vermeirsch, Peter Lambert, Rik Van de Walle
Multim. Syst.4
2010 Requantization transcoding for H.264/AVC video coding
Jan De Cock, Stijn Notebaert, Peter Lambert, Rik Van de Walle
Signal Process. Image Commun.3
2010 Exploiting quantization and spatial correlation in virtual-noise modeling for distributed video coding
Jozef Skorupa, Jürgen Slowack, Stefaan Mys, Nikos Deligiannis, Jan De Cock, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
Signal Process. Image Commun.6
2010 Rate-distortion driven decoder-side bitplane mode decision for distributed video coding
Jürgen Slowack, Stefaan Mys, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
Signal Process. Image Commun.5
2010 Flexible distribution of complexity by hybrid predictive-distributed video coding
Jürgen Slowack, Jozef Skorupa, Stefaan Mys, Peter Lambert, Christos Grecos, Rik Van de Walle
Signal Process. Image Commun.4
2010 Noise- and compression-robust biological features for texture classification
Gaëtan Martens, Chris Poppe, Peter Lambert, Rik Van de Walle
Vis. Comput.3
2009 Multi-view Object Localization in H.264/AVC Compressed Domain
abstract
This paper presents a multi-view homography-based approach for object localization in H.264/AVC compressed video surveillance sequences. The proposed novel, low-complexity method is able to accurately localize moving objects on a ground plane using multiple camera data. Contrary to existing work that exploits motion vectors for object detection and tracking, our compressed domain multi-view object localization solely uses macroblock (MB) partition information. Foreground segmentation is performed on single view compressed video data using MB partition-based temporal differencing. Blob merging, convex hull fitting and noise removal are applied on the resulting foreground views to extract objects. Once relevant objects are found in single views, they are projected onto a ground plane by exploiting the homography constraint. Since projected foreground MB views of multiple cameras will only overlap on points where foreground intersects the ground plane, object locations can be extracted by detecting local maxima on the accumulated ground plane image.
Steven Verstockt, Sarah De Bruyne, Chris Poppe, Peter Lambert, Rik Van de Walle
AVSS4
2009 Transcoding of H.264/AVC to SVC with motion data refinement
abstract
In this paper, we present motion-refined transcoding of H.264/AVC streams to SVC in the transform domain. By accurately taking into account both rate and distortion in the different layers on the one hand, and the SVC inter-layer motion prediction mechanisms on the other hand, the proposed transcoding architecture is able to improve rate-distortion performance over existing approaches. We propose a multilayer control mechanism that trades off performance between the different layers, resulting in 0.5 dB gains in the output SVC base layer.
Jan De Cock, Stijn Notebaert, Kenneth Vermeirsch, Peter Lambert, Rik Van de Walle
ICIP4
2009 Estimating motion reliability to improve moving object detection in the H.264/AVC domain
abstract
This paper presents a new algorithm for moving object detection in the H.264/AVC compressed domain which relies on motion vector information. In contrast to other motion vector-based algorithms, special attention is paid to noisy motion vectors as they highly decrease the performance of these algorithms. We propose to estimate the reliability of motion vectors by comparing them with projected motion vectors from surrounding frames. As such, noisy motion vectors are localized. By combining this information with the magnitude of motion vectors, foreground objects are distinguished. Experimental results demonstrate that our algorithm achieves significantly better segmentation quality compared to other motion vector-based approaches.
Sarah De Bruyne, Chris Poppe, Steven Verstockt, Peter Lambert, Rik Van de Walle
ICME4
2009 Fast Channel Switching Based on SVC in an IPTV Environment
abstract
An IPTV network contains characteristics of both a unicast environment and a broadcast environment. For both environments, techniques to accelerate channel switching exist. In a unicast environment a bandwidth efficient implementation to accelerate channel switching can be obtained with non scalable video coding, while for a broadcast environment the same acceleration is more optimal using scalable video coding. For a combination of both environments, as is the case for an IPTV environment, none of the solutions is optimal. In this paper, we propose a compromise between scalable and non scalable video compression to adapt to the properties of an IPTV environment. One of the proposed configurations obtains bandwidth reductions on the access network of the IPTV network between 2.4% and 4.3% with a small increase between 0.1% and 1.6% on the core network of the IPTV network.
Glenn Van Wallendael, Peter Lambert, Rik Van de Walle
ISM2
2009 Leveraging the quantization offset for improved requantization transcoding of H.264/AVC video
abstract
Requantization transcoding is a method for reducing the bit rate of compressed video bitstreams. Most research on requantization is concerned with the architectural design, the selection of a suitable quantizer, or the reduction of requantization errors. In this paper, we propose to incorporate a new dimension in requantization transcoding: the quantization offset. We compare two requantization methods: increasing the quantization step size and decreasing the quantization offset. Furthermore, we propose a novel heuristic for requantization transcoding based on a theoretical rate-distortion analysis. The experimental results for H.264/AVC video show that requantization is improved in rate-distortion sense with gains up to 1 dB for open-loop requantization of B pictures compared to requantization with fixed quantization offset.
Stijn Notebaert, Jan De Cock, Kenneth Vermeirsch, Peter Lambert, Rik Van de Walle
PCS4
2009 Stopping criterions for turbo coding in a Wyner-Ziv video codec
abstract
Distributed video coding (DVC) targets video coding applications with low encoding complexity by generating a prediction of the video signal at the decoder. One of the most common architectures uses turbo codes to correct errors in this prediction. Unfortunately, a rigorous analysis of turbo coding in the context of DVC is missing. We have targeted one particular aspect of turbo coding: the stopping criterion. The stopping criterion indicates whether decoding was successful, i.e., whether the errors in the prediction signal have been corrected. In this paper we describe and compare several stopping criterions known from the field of channel coding and criterions currently used in DVC. As our results suggest the choice of the stopping criterion has a significant impact on the overall video-coding performance. Moreover, we have found that there are even better performing criterions than those currently used in DVC.
Jozef Skorupa, Jürgen Slowack, Stefaan Mys, Peter Lambert, Rik Van de Walle, Christos Grecos
PCS4
2009 Accounting for quantization noise in online correlation noise estimation for Distributed Video Coding
abstract
In Distributed Video Coding (DVC), compression is achieved by exploiting correlation between frames at the decoder, instead of at the encoder. More specifically, the decoder uses already decoded frames to generate side information Y for each Wyner-Ziv frame X, and corrects errors in Y using error correcting bits received from the encoder. For efficient use of these bits, the decoder needs information about the correlation between X available at the encoder and Y at the decoder. While several techniques for online estimation of correlation noise X - Y have been proposed, the quantization noise in Y has not been taken into account. As a solution, in this paper, we calculate the quantization noise of intra frames at the encoder and use this information at the decoder to improve the accuracy of the correlation noise estimation. Results indicate averageWyner-Ziv bit rate reductions up to 19.5% (Bjøntegaard delta) for coarse quantization.
Jürgen Slowack, Stefaan Mys, Jozef Skorupa, Peter Lambert, Rik Van de Walle, Christos Grecos
PCS4
2009 Evaluation of transform performance when using shape-adaptive partitioning in video coding
abstract
When combining non-rectangular (shape-adaptive) partitioning of inter pictures with a rectangular block transform, some of the transforms will be applied to a residual signal which originates from two different predictions. This is a condition that is never encountered in any of the existing video coding standards, and the impact of this on the performance of the transform coder has not been investigated. In this paper we investigate the effect of these mixed-signal blocks on the coding gain for various transforms, and we compare these against the optimal KLT gain. We find that, despite the transformed residual block being composed of two parts that were predicted from different areas of the reference picture, correlation within mixed blocks is very similar to that of normal blocks. The DCT is only marginally suboptimal w.r.t. KLT. KLT has practical issues that will reduce its coding gain or increase the signaling overhead: transform bases need to be quantized and transmitted, or a number of fixed bases needs to be chosen offline. Therefore we recommend DCT be used for all types of blocks.
Kenneth Vermeirsch, Jan De Cock, Stijn Notebaert, Peter Lambert, Rik Van de Walle
PCS4
2009 Moving object detection in the H.264/AVC compressed domain for video surveillance applications
Chris Poppe, Sarah De Bruyne, Tom Paridaens, Peter Lambert, Rik Van de Walle
J. Vis. Commun. Image Represent.4
2009 Introducing skip mode in distributed video coding
Stefaan Mys, Jürgen Slowack, Jozef Skorupa, Peter Lambert, Rik Van de Walle
Signal Process. Image Commun.4
2009 Architectures for Fast Transcoding of H.264/AVC to Quality-Scalable SVC Streams
abstract
The scalable extension of H.264/AVC (SVC) was recently standardized, and offers scalability at a minor penalty in rate-distortion efficiency when compared to single-layer H.264/AVC coding. In SVC, a scaled version of the original video sequence can easily be extracted by dropping layers from the stream. However, most of the video content nowadays is still produced in a single-layer format. While decoding and reencoding is a possible solution to introduce scalability in the existing bitstreams, this is an approach which requires a tremendous amount of time and effort. In this paper, we show that transcoding can be used to intelligently derive scalable bitstreams from existing single-layer streams. We focus on SNR scalability, and introduce techniques that are able to create multiple quality layers in the bitstreams. We also discuss bitstream rewriting from SVC to H.264/AVC, and examine how our newly proposed architectures can benefit from the changes that were introduced for bitstream rewriting. Architectures with different rate distribution flexibility and computational complexity are discussed. Rate-distortion performance of transcoding is shown to be comparable to that of reencoding at a fraction of the time needed for the latter.
Jan De Cock, Stijn Notebaert, Peter Lambert, Rik Van de Walle
IEEE Trans. Multim.3
2008 Advanced bitstream rewriting from H.264/AVC to SVC
abstract
In previous work, we introduced an H.264/AVC-to-SVC transcoder for creating SVC streams with multiple quality layers from a single-layer H.264/AVC stream. This architecture was able to restrain the error drift due to requantization of the residual coefficients. In this paper, we show that it is possible to further reduce the complexity, and completely eliminate the drift in the enhancement layer, by making use of the bitstream rewriting functionality in SVC. We propose different novel architectures, which are able to flexibly distribute the data among the different created layers.
Jan De Cock, Stijn Notebaert, Peter Lambert, Rik Van de Walle
ICIP3
2008 Efficient spatial resolution reduction transcoding for H.264/AVC
abstract
In this paper, we present a spatial resolution reduction transcoding architecture for H.264/AVC, which extends open-loop transcoding with a low-complexity compensation technique in the reduced-resolution domain. The proposed architecture removes visual artifacts from the transcoded sequence, while keeping complexity significantly lower than more traditional cascaded decoder-encoder architectures. The refinement step of the proposed architecture can be used to further improve rate-distortion performance, at the cost of additional complexity. In this way, a dynamic-complexity transcoder is rendered possible.
Jan De Cock, Stijn Notebaert, Kenneth Vermeirsch, Peter Lambert, Rik Van de Walle
ICIP4
2008 Improved dynamic rate shaping for H.264/AVC video streams
abstract
In this paper, dynamic rate shaping for H.264/AVC intra-coded video streams is investigated. An analysis of the distortion distinguishes different error components that lead to the degradation of the output video stream. Experimental results show that the propagation error due to intra prediction has a major impact on the visual quality. This results from the highly increased number of dependencies which are invoked by the H.264/AVC intra prediction. In order to eliminate the propagation error, we propose to use single-loop compensation techniques. Both objective and subjective results show that the compensation techniques highly improve the visual quality.
Stijn Notebaert, Jan De Cock, Kenneth Vermeirsch, Peter Lambert, Rik Van de Walle
ICIP4
2008 Extended Macroblock Bipartitioning Modes for H.264/AVC Inter Coding
abstract
We present a family of new macroblock partitions for H.264/AVC inter prediction. These modes allow a macroblock to be bipartitioned along a horizontal, vertical, or diagonal edge with an arbitrary location within the block. These new partitioning modes allow an encoder to adapt more closely to the local motion characteristics and improve its rate-distortion performance. In order to quantify the improvements we implemented the proposed partitioning modes in the context of the H.264/AVC video coding standard. Over half the coding efficiency of existing more general schemes is attained at a fraction of the number of added partitioning modes, and hence of encoder complexity, of those schemes. The results are consistent across different profiles, indicating that the proposed extension has good compatibility with the existing H.264/AVC tools for motion compensation and residual coding.
Kenneth Vermeirsch, Jan De Cock, Stijn Notebaert, Peter Lambert, Rik Van de Walle
ISM4
2008 Unsupervised texture segmentation and labeling using biologically inspired features
abstract
Due to the semantic gap, describing high-level semantic concepts with low-level visual features is a very challenging task. The classification of textures in scene images is intricate because of the high variation of the data. Therefore, the application of appropriate features is of utter importance. This paper presents biologically inspired features for texture segmentation and an unsupervised method to link those texture features with semantic concepts. The calculation of the features is inspired by the human visual system and corresponds to cell outputs in the first stage of the visual cortex. Analogously to the processing principles of the cortex, self-organizing maps are employed for classification. The performance of the texture segmentation and labeling is evaluated on textures from the Brodatz album and on a real-life scenery image dataset. For both methods, a high percentage of pixels is correctly classified.
Gaëtan Martens, Chris Poppe, Peter Lambert, Rik Van de Walle
MMSP3
2008 Increased flexibility in inter picture partitioning
abstract
To attain efficient coding of sequences with complex motion activity, modern video coding standards allow variable block sizes to be employed in temporal prediction. A block with complex motion can be partitioned into two equal-sized halves or into four quadrants. In this paper we study the impact of allowing blocks to be partitioned in two unequal subpartitions. Additionally, we allow block partitioning along diagonal edges. These provisions allows encoders to better adapt to the local characteristics of the motion activity in a video sequence. We verified experimentally that the presence of partitioning edges that do not coincide with transform boundaries does not negatively impact the decorrelating strength of the residual transform, so the proposed extended partitioning strategies can be applied regardless of the details of the residual coder. Implementing the proposed extended partitioning modes in an H.264/AVC coder at the macroblock and submacroblock level, we observe that coding efficiency gains are greatest in low-resolution sequences, where moving features in the video sequence tend to be more spatially localized. For CIF and QCIF sequences we achieve a bit rate reduction of about 3–6% over a wide fidelity range.
Kenneth Vermeirsch, Jan De Cock, Stijn Notebaert, Peter Lambert, Rik Van de Walle
MMSP4
2008 Optimizing user QoE through overlay routing, bandwidth management and dynamic transcoding
abstract
More and more, multimedia services are being accessed via fixed and mobile networks. These services are typically much more sensitive to packet loss, delay and/or congestion than traditional services. In particular, multimedia data is often time critical and, as a result, network issues are not well tolerated and significantly deteriorate the userpsilas quality of experience (QoE). We therefore propose a QoE optimization platform that is able to mitigate problems that might occur at any location in the delivery path from service provider to customer. More specifically, the distributed architecture supports overlay routing to circumvent erratic parts of the network core. In addition, it comprises proxy components that realize last mile optimization through automatic bandwidth management and the application of processing on multimedia flows. In this paper we introduce a transcoding service for this proxy component which enables the transformation of H.264/AVC video flows to an arbitrary bit rate. Through representative experimental results, we illustrate how this addition enhances the QoE optimization capabilities of the proposed platform by allowing the proxy component to compute more flexible and effective bandwidth distributions.
Maarten Wijnants 0001, Wim Lamotte, Bart De Vleeschauwer, Filip De Turck, Bart Dhoedt, Piet Demeester, Peter Lambert, Dieter Van de Walle, Jan De Cock, Stijn Notebaert, Rik Van de Walle
WOWMOM7
2008 A compressed-domain approach for shot boundary detection on H.264/AVC bit streams
Sarah De Bruyne, Davy Van Deursen, Jan De Cock, Wesley De Neve, Peter Lambert, Rik Van de Walle
Signal Process. Image Commun.5
2007 Improved Background Mixture Models for Video Surveillance Applications
Chris Poppe, Gaëtan Martens, Peter Lambert, Rik Van de Walle
ACCV (1)3
2007 Bridging the Gap: Transcoding from Single-Layer H.264/AVC to Scalable SVC Video Streams
Jan De Cock, Stijn Notebaert, Peter Lambert, Rik Van de Walle
ACIVS3
2007 Mixture Models Based Background Subtraction for Video Surveillance Applications
Chris Poppe, Gaëtan Martens, Peter Lambert, Rik Van de Walle
CAIP3
2007 Analysis of Prediction Mode Decision in Spatial Enhancement Layers in H.264/AVC SVC
Koen De Wolf, Davy De Schrijver, Wesley De Neve, Saar De Zutter, Peter Lambert, Rik Van de Walle
CAIP5
2007 XML-driven Exploitation of Combined Scalability in Scalable H.264/AVC Bitstreams
abstract
The heterogeneity in the contemporary multimedia environments requires a format-agnostic adaptation framework for the consumption of digital video content. Scalable bitstreams can be used in order to satisfy as many circumstances as possible. In this paper, the scalable extension on the H.264/AVC specification is used to obtain the parent bitstreams. The adaptation along the combined scalability axis of the bitstreams is done in a format-independent manner. Therefore, an abstraction layer of the bitstream is needed. In this paper, XML descriptions are used representing the high-level structure of the bitstreams by relying on the MPEG-21 bitstream syntax description language standard. The exploitation of the combined scalability is executed in the XML domain by implementing the adaptation process in a streaming transformation for XML (STX) stylesheet. The algorithm used in the transformation of the XML description is discussed in detail in this paper. From the performance measurements, one can conclude that the STX transformation in the XML domain and the generation of the corresponding adapted bitstream can be realized in real time.
Davy De Schrijver, Wesley De Neve, Koen De Wolf, Peter Lambert, Davy Van Deursen, Rik Van de Walle
ISCAS4
2007 Scalable, Wavelet-Based Video: From Server to Hardware-Accelerated Client
abstract
Video source, carrier and client diversification have led the video coding community to develop scalable video codecs supporting efficient decoding at varying resolution, frame rate and quality. Scalable video has several advantages over a nonscalable approach, but a large scale deployment is far from trivial and a lot of open questions remain. To resolve these, we developed a complete video delivery chain for scalable wavelet-based video. This includes a video server, a negotiation framework, a video scaling infrastructure and two scalable video clients, one pure software client and one real-time, hardware accelerated client. This paper describes the complete chain and identifies and quantifies the impact of using scalable video in every link of this chain.
Hendrik Eeckhaut, Harald Devos, Peter Lambert, Davy De Schrijver, Wim Van Lancker, Vincent Nollet, Prabhat Avasare, Tom Clerckx, Fabio Verdicchio, Mark Christiaens, Peter Schelkens, Rik Van de Walle, Dirk Stroobandt
IEEE Trans. Multim.3
2006 Requantization Transcoding in Pixel and Frequency Domain for Intra 16x16 in H.264/AVC
Jan De Cock, Stijn Notebaert, Peter Lambert, Davy De Schrijver, Rik Van de Walle
ACIVS3
2006 A Real-Time Content Adaptation Framework for Exploiting ROI Scalability in H.264/AVC
Peter Lambert, Davy De Schrijver, Davy Van Deursen, Wesley De Neve, Yves Dhondt, Rik Van de Walle
ACIVS1
2006 A Flexible Macroblock Scheme for Unequal Error Protection
abstract
This paper proposes an enhanced error protection scheme using flexible macroblock ordering in H.264/AVC. The algorithm uses a two-phase system. In the first phase, the importance of every macroblock is calculated based on its influence on the current frame and future frames. In the second phase, the macroblocks with the highest impact factor are grouped together in a separate slice group using the flexible macroblock ordering feature of H.264/AVC. By using an unequal error protection scheme, the slice group containing the most important macroblocks can be better protected than the other slice group. The proposed algorithm offers better concealment opportunities than the algorithms which are predefined for flexible macroblock ordering in H.264/AVC.
Yves Dhondt, Peter Lambert, Rik Van de Walle
ICIP2
2006 Flexible macroblock ordering in H.264/AVC
Peter Lambert, Wesley De Neve, Yves Dhondt, Rik Van de Walle
J. Vis. Commun. Image Represent.1
2006 Rate-distortion performance of H.264/AVC compared to state-of-the-art video codecs
abstract
In the domain of digital video coding, new technologies and solutions are emerging in a fast pace, targeting the needs of the evolving multimedia landscape. One of the questions that arises is how to assess these different video coding technologies in terms of compression efficiency. In this paper, several compression schemes are compared by means of peak signal-to-noise ratio (PSNR) and just noticeable difference (JND). The codecs examined are XviD 0.9.1 (conform to the MPEG-4 Visual Simple Profile), DivX 5.1 (implementing the MPEG-4 Visual Advanced Simple Profile), Windows Media Video 9, MC-EZBC and H.264/AVC AHM 2.0 (version JM 6.1 of the reference software, extended with rate control). The latter plays a key role in this comparison because the H.264/AVC standard can be considered as the de facto benchmark in the field of digital video coding. The obtained results show that H.264/AVC AHM 2.0 outperforms current proprietary and standards-based implementations in almost all cases. Another observation is that the choice of a particular quality metric can influence general statements about the relation between the different codecs.
Peter Lambert, Wesley De Neve, Peter De Neve, Ingrid Moerman, Piet Demeester, Rik Van de Walle
IEEE Trans. Circuits Syst. Video Technol.1
2004 Low-level behavioral analysis of the JVT/AVC decoder
abstract
H.264/AVC is a video codec developed by the Joint Video Team (JVT); a cooperation between the ITU-T VCEG (Video Coding Experts Group) and ISO/IEC MPEG (Moving Picture Experts Group). This new video coding standard has some new features that allow to get significant improvements in coding efficiency. This improved coding efficiency leads to an overall more complex algorithm which has high demands regarding memory usage and processing power. Complexity, however, is an abstract concept and cannot be measured in a simple manner. In this paper we present a method to obtain an accurate and more in-depth view on the internals of the JVT/AVC decoder. By decoding several bit streams having different encoding parameters, various program characteristics were measured. On these measurements, principal components analysis was performed to get a different view on these measurements. Our results show that the various encoding parameters have a clear impact on the low level behavior of the decoder. Moreover, our methodology allows us to give an explanation for the observed dissimilarities.
Peter Lambert, Lieven Eeckhout, Robbie De Sutter, Koen De Bosschere, Rik Van de Walle
VCIP1
2004 Fully scalable video coding in multicast applications
abstract
The increasing diversity of the characteristics of the terminals and networks that are used to access multimedia content through the internet introduces new challenges for the distribution of multimedia data. Scalable video coding will be one of the elementary solutions in this domain. This type of coding allows to adapt an encoded video sequence to the limitations of the network or the receiving device by means of very basic operations. Algorithms for creating fully scalable video streams, in which multiple types of scalability are offered at the same time, are becoming mature. On the other hand, research on applications that use such bitstreams is only recently emerging. In this paper, we introduce a mathematical model for describing such bitstreams. In addition, we show how we can model applications that use scalable bitstreams by means of definitions that are built on top of this model. In particular, we chose to describe a multicast protocol that is targeted at scalable bitstreams. This way, we will demonstrate that it is possible to define an abstract model for scalable bitstreams, that can be used as a tool for reasoning about such bitstreams and related applications.
Sam Lerouge, Robbie De Sutter, Peter Lambert, Rik Van de Walle
VCIP3
2004 Assessment of the compression efficiency of the MPEG-4 AVC specification
abstract
Video coding is used under the hood of a lot of multimedia applications, such as video conferencing, digital storage media, television broadcasting, and internet streaming. Recently, new standards-based and proprietary technologies have emerged. An interesting problem is how to evaluate these different video coding solutions in terms of delivered quality. In this paper, a PSNR-based approach is applied in order to compare the coding potential of H.264/AVC AHM 2.0 with the compression efficiency of XviD 0.9.1, DivX 5.05, Windows Media Video 9, and MC-EZBC. Our results show that MPEG-4-based tools, and in particular H.264/AVC, can keep step with proprietary solutions. The rate-distortion performance of MC-EZBC, a wavelet-based video codec, looks very promising too.
Wesley De Neve, Peter Lambert, Sam Lerouge, Rik Van de Walle
VCIP2