EDBT 2026 Demo / reviewers in the wild / expert
Adrian Munteanu 0001
dblp:87/4387
· DBLP profile ↗
143ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0001-7290-0428ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 123 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 3Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProtoSeg: A prototype-based point cloud instance segmentation method
Remco Royen, Leon Denis, Adrian Munteanu 0001 |
Signal Process. Image Commun. | 3 |
| 2026 | Efficient and Scalable Point Cloud Generation With Sparse Point-Voxel Diffusion ModelsabstractWe propose a novel point cloud U-Net diffusion architecture for 3-D generative modeling capable of generating high-quality and diverse 3-D shapes while maintaining fast generation times. Our network employs a dual-branch architecture, combining the high-resolution representations of points with the computational efficiency of sparse voxels. Our fastest variant outperforms all nondiffusion generative approaches on unconditional shape generation, the most popular benchmark for evaluating point cloud generative models, while our largest model achieves state-of-the-art results among diffusion methods, with a runtime approximately 70% of the previously state-of-the-art point-voxel diffusion (PVD), measured on the same hardware setting. Beyond unconditional generation, we perform extensive evaluations, including conditional generation on all categories of ShapeNet, demonstrating the scalability of our model to larger datasets, and implicit generation, which allows our network to produce high-quality point clouds on fewer timesteps, further decreasing the generation time. Finally, we evaluate the architecture's performance in point cloud completion and super-resolution. Our model excels in all tasks, establishing it as a state-of-the-art diffusion U-Net for point cloud generative modeling. The code is publicly available at https://github.com/JohnRomanelis/SPVD. Ioannis Romanelis, Vlassis Fotis, Athanasios P. Kalogeras, Christos Alexakos, Adrian Munteanu 0001, Konstantinos Moustakas |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | MeasureXpert: Automatic Anthropometric Measurement Extraction from Two Unregistered, Partial, Posed, and Dressed Body ScansabstractWhile automatic anthropometric measurement extraction has witnessed growth in recent years, effective, non-contact, and precise measurement methods for dressed humans in arbitrary poses are still lacking, limiting the widespread application of this technology. The occlusion caused by clothing and the adverse influence of posture on body shape significantly increase the complexity of this task. Additionally, current methods often assume the availability of a complete 3D body mesh in a canonical pose (e.g., ”A” or ”T” pose), which is not always the case in practice. To address these challenges, we propose MeasureXpert, a novel learningbased model that requires only two unregistered, partial, and dressed body scans as input, and accommodates entirely independent and arbitrary poses for each scan. MeasureXpert computes a comprehensive representation of the naked body shape by synergistically fusing features from the front- and back-view partial point clouds. The comprehensive representation obtained is mapped onto a 3D undressed body shape space, assuming a canonical posture and incorporating predefined measurement landmarks. A pointbased offset optimization is also developed to refine the reconstructed complete body shape, enabling accurate regression of measurement values. To train the proposed model, a new large-scale dataset, consisting of 300K samples, was synthesized. The proposed model was validated using two publicly available real-world datasets and was compared with different relevant methods. Extensive experimental results demonstrate that MeasureXpert achieves superior performance compared to the reference methods. The code and dataset are available at: MeasureXpertProject. Xinxin Dai, Pengpeng Hu, Vasile Palade, Adrian Munteanu 0001 |
ICCV | 5 |
| 2025 | Global-to-Local Color Correction with Full-Region Coverage for Multi-view Light Field ImagesabstractColor correction methods for multi-view images are typically divided into global-based and local-based approaches. Global methods perform global color mapping but fail to address local differences, leading to local color inconsistencies. Local methods focus on regional color mapping based on distributions or semantics but struggle with sparse semantic correspondences and are sensitive to lighting and noise. To address these, we propose a hybrid color correction method for light field images. First, a global color correction is applied to ensure overall correction. Moreover, an object matching algorithm is designed to match regions and calculate both global and local similarities between images to refine the corrections. Next, a local optimization module is introduced to optimize the adjustment of specific regions. Finally, gradient preservation is incorporated to maintain structural consistency. Experiments on our proposed dataset using a 3×3 light field camera array demonstrate that our method outperforms existing approaches. Yixu Huang, Rui Zhong 0005, Ségolène Rogge, Adrian Munteanu 0001 |
ICME | 4 |
| 2025 | G-SPVD: Image and Sketch Guided Point Cloud Generation with Sparse Point-Voxel Diffusion ModelsabstractWe propose a novel framework, Guided Sparse Point-Voxel Diffusion (G-SPVD), for Point Cloud generation guided from a single visual input - either an image or a rough hand-drawn sketch, both from an unknown viewing angle. G-SPVD combines a Vision Transformer with a Diffusion Model that iteratively forms a noisy set of points to match the requested input. Our quantitative evaluation demonstrates that our framework achieves state-of-the-art results compared to other methods in single-image reconstruction on the ShapeNet dataset. Moreover, despite the reduced information available in sketchbased inputs, our sketch-guided model still attains competitive reconstruction metrics. We present several qualitative results for both tasks to further illustrate the effectiveness of our method. Finally, we evaluate our method on unconditional generation, demonstrating that our model can generate shapes with quality and diversity on par with the current state-of-the-art. Our code will be released upon publication. Ioannis Romanelis, Vlassis Fotis, Adrian Munteanu 0001, Konstantinos Moustakas |
VCIP | 3 |
| 2025 | LF-GS: 3D Gaussian Splatting for View Synthesis of Multi-View Light Field Imagesabstract3D Gaussian Splatting (3D-GS) has emerged as a groundbreaking approach for view synthesis. However, when applied to light field image synthesis, the issue of a too narrow field of view (FOV) that leaves some areas uncovered, compounded by the problem of data sparsity, significantly compromises the quality of synthesized views using 3D-GS. To overcome these limitations, we present LF-GS, a specialized 3D-GS variant optimized for light field image synthesis. Our methodology incorporates two key innovations. First, by harnessing the unique advantage of light field sub-aperture images that provide dense geometric cues, our method enables the effective incorporation of enhanced depth and normal priors derived from light field images. This allows for more accurate depth than monocular depth estimation. Second, unlike other methods that struggle to control the generation of unreasonable Gaussians, we introduce adaptive regularization mechanisms. These mechanisms strategically regulate Gaussian opacity and spatial scale during optimization, thereby preventing model overfitting and preserving essential scene details. Comprehensive experiments on our newly constructed light field dataset demonstrate that LF-GS achieves significant quality improvements over 3D-GS. Yixu Huang, Rui Zhong 0005, Ségolène Rogge, Adrian Munteanu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Forecasting Traffic Progression in Terms of Semantically Interpretable States by Exploring Multiple Data RepresentationsabstractIn the rapidly evolving landscape of mobility modelling, the application of deep learning approaches introduces both opportunities and challenges. Such approaches, while powerful, yield opaque models that lack interpretability and adaptability to diverse traffic contexts. Addressing those challenges, a finite set of humanly-interpretable traffic states is exploited here for the purpose of facilitating the annotation of mobility data with meaningful labels such as congestion, free-flow, traffic build-up, etc. Such annotation unlocks a range of opportunities to integrate multiple complementary approaches for modelling state transition behaviour. Concretely, a novel hybrid modelling framework is introduced in this article leveraging multiple data representations (temporal, time-frequency and symbolic) with the aim to forecast traffic progression in terms of humanly-explicable state transitions. Three distinct modelling paradigms are subsequently explored: neural, neural-to-symbolic, and symbolic-to-neural, by demonstrating their potential to capture and forecast traffic dynamics on real-world mobility data. While the fully neural approach is undoubtedly the most accurate one, the two neuro-symbolic approaches offer a better trade-off between accuracy on one side and interpretability, probability calibration, and computational efficiency, on the other. This work illustrates the importance of tailored data representations in understanding and predicting complex mobility behaviour, highlighting the benefits of hybrid approaches in achieving interpretability and efficiency in traffic data analysis. Michiel Dhont, Adrian Munteanu 0001, Elena Tsiporkova |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | RT-GS2: Real-Time Generalizable Semantic Segmentation for 3D Gaussian Representations of Radiance Fields
Mihnea Bogdan Jurca, Remco Royen, Ion Giosan, Adrian Munteanu 0001 |
BMVC | 4 |
| 2024 | RESSCAL3D++: Joint Acquisition and Semantic Segmentation of 3D Point Cloudsabstract3D scene understanding is crucial for facilitating seamless interaction between digital devices and the physical world. Real-time capturing and processing of the 3D scene are essential for achieving this seamless integration. While existing approaches typically separate acquisition and processing for each frame, the advent of resolution-scalable 3D sensors offers an opportunity to overcome this paradigm and fully leverage the otherwise wasted acquisition time to initiate processing. In this study, we introduce VX-S3DIS, a novel point cloud dataset accurately simulating the behavior of a resolution-scalable 3D sensor. Additionally, we present RESSCAL3D++, an important improvement over our prior work, RESSCAL3D, by incorporating an update module and processing strategy. By applying our method to the new dataset, we practically demonstrate the potential of joint acquisition and semantic segmentation of 3D point clouds. Our resolution-scalable approach significantly reduces scalability costs from 2% to just 0.2% in mIoU while achieving impressive speed-ups of 15.6 to 63.9% compared to the non-scalable baseline. Furthermore, our scalable approach enables early predictions, with the first one occurring after only 7% of the total inference time of the baseline. The new VX-S3DIS dataset is available at https://github.com/remcoroyen/vx-s3dis. Remco Royen, Kostas Pataridis, Ward van der Tempel, Adrian Munteanu 0001 |
ICIP | 4 |
| 2024 | MonoKalman: Monocular Vehicle Pose Estimation with Kalman Filter-based temporal consistencyabstractIn the past few years, there has been a rising demand for monocular vehicle 6D pose estimation in traffic video, due to its applicability in emerging domains such as smart mobility and intelligent transportation systems. However, most of the existing approaches are image-based, which results in undesirable temporal pose inconsistencies of vehicles between consecutive frames when applied to video sequences. In this work, we present a Kalman filter-based post-processing method, named MonoKalman, that enhances the temporal consistencies of the 6D pose estimates of vehicles in traffic video. We compare our method with a state-of-the-art 6D pose estimation method on synthetic video data of traffic scenery. The experimental results indicate that our MonoKalman significantly outperforms the image-based baseline method and effectively reduces temporal pose artifacts, ensuring a more coherent and stable representation of 6D vehicle temporal poses in traffic video. To more effectively demonstrate MonoKalman’s enhancements over the baseline model, we design a graphical user interface. This interface offers users insights through detailed quantitative metrics and dynamic visualizations, allowing them to conduct and customize their experiments. Leandro Di Bella, Yangxintong Lyu, Bruno Cornelis, Adrian Munteanu 0001 |
MDM | 4 |
| 2024 | STRADA: Spatial-Temporal Dashboard for traffic forecastingabstractEfficiently visualizing Spatial-Temporal traffic data plays an important role nowadays in traffic monitoring. Interactive dashboards offering effective visualizations of spatial-temporal traffic data play a more prominent role in traffic monitoring. In this paper, we introduce a dashboard for visualizing traffic data. Specifically, our dashboard integrates spatial-temporal components for the time-series traffic data of Brussels, which is the first GNN-based traffic demonstration tool for Brussels. Furthermore, we provide an interface for displaying traffic prediction of deep-learning-based Spatial-Temporal Graph Neural Networks (STGNNs), which have demonstrated state of the art performance in Intelligent Transpiration Systems (ITS). In addition, we demonstrate two real-world use cases by using the proposed dashboard which provides the potential for a future tool to achieve intelligent transportation management. Seyed Mohamad Moghadas, Yangxintong Lyu, Bruno Cornelis, Adrian Munteanu 0001 |
MDM | 4 |
| 2024 | LAM3D: Leveraging Attention for Monocular 3D Object DetectionabstractSince the introduction of the self-attention mechanism and the adoption of the Transformer architecture for Computer Vision tasks, the Vision Transformer-based architectures gained a lot of popularity in the field, being used for tasks such as image classification, object detection and image segmentation. However, efficiently leveraging the attention mechanism in vision transformers for the Monocular 3D Object Detection task remains an open question. In this paper, we present LAM3D, a framework that Leverages self-Attention mechanism for Monocular 3D object Detection. To do so, the proposed method is built upon a Pyramid Vision Transformer v2 (PVTv2) as feature extraction backbone and 2D/3D detection machinery. We evaluate the proposed method on the KITTI 3D Object Detection Benchmark, proving the applicability of the proposed solution in the autonomous driving domain and outperforming reference methods. Moreover, due to the usage of self-attention, LAM3D is able to systematically outperform the equivalent architecture that does not employ self-attention. Diana-Alexandra Sas, Leandro Di Bella, Yangxintong Lyu, Florin Oniga, Adrian Munteanu 0001 |
MMSP | 5 |
| 2023 | W2H-Net: Fast Prediction of Waist-to-Hip Ratio from Single Partial Dressed Body Scans in Arbitrary Postures via Deep LearningabstractThe Waist-to-Hip Ratio (WHR) is an important indicator for health risk prediction, body fat distribution, body shape analysis, and physical fitness analysis. The conventional approach for obtaining the WHR entails manual measurement, which necessitates experienced anthropometrists to measure the waist and hip circumferences of a subject wearing tight clothing in a predetermined posture, and subsequently calculate the ratio based on the acquired measurements. WHR errors may be accumulated due to the anthropometrist’s subjectivity, as well as the person’s pose and attire during the measurement process. Non-contact anthropometric measurements using 3D scanning technology have shown promise in providing higher accuracy and faster measurement compared to traditional methods. However, they require complete undressed body scans as input, which is not always available. In this paper, we proposed, to the best of our knowledge, the first deep learning-based algorithm, dubbed W2H-Net, to predict the WHR directly from single partial dressed body scans in arbitrary postures. W2H-Net introduces a novel framework called Focus-Net to improve learning accuracy by selectively focusing on parts that require attention. W2H-Net provides a flexible, cost-effective, and privacy-preserving way to obtain accurate WHR measurements, which are crucial for predicting health risks associated with central obesity. Extensive experimental results can demonstrate the superiority of the proposed method. Xinxin Dai, Pengpeng Hu, Adrian Munteanu 0001 |
IJCB | 4 |
| 2023 | Measure4dhand: Dynamic Hand Measurement Extraction from 4D ScansabstractHand measurement is vital for hand-centric applications such as glove design, immobilization design, protective gear design, to name a few. Vision-based methods have been previously proposed but are limited in their ability to only extract hand dimensions in a static and standardized posture (open-palm hand). However, dynamic hand measurements should be considered when designing these wearable products since the interaction between hands and products cannot be ignored. Unfortunately, none of the existing methods are designed for measuring dynamic hands. To address this problem, we propose a user-friendly and fast method dubbed Measure4DHand, which automatically extracts dynamic hand measurements from a sequence of depth images captured by a single depth camera. Firstly, the ten dimensions of the hand are defined. Secondly, a deep neural network is developed to predict landmark sequences for the ten dimensions from partial point cloud sequences. Finally, a method is designed to calculate dimension values from landmark sequences. A novel synthetic dataset consisting of 234K hands in various shapes and poses, along with their corresponding ground truth landmarks, is proposed for training the proposed methods. The experiment based on real-world data captured by a Kinect illustrates the evolution of the ten dimensions during hand movement, while the mean ranges of variation are also reported, providing valuable information for the hand wearable product design. (The video abstract is available here.) Xinxin Dai, Pengpeng Hu, Vasile Palade, Adrian Munteanu 0001 |
ICIP | 5 |
| 2023 | RESSCAL3D: Resolution Scalable 3D Semantic Segmentation of Point CloudsabstractWhile deep learning-based methods have demonstrated outstanding results in numerous domains, some important functionalities are missing. Resolution scalability is one of them. In this work, we introduce a novel architecture, dubbed RESSCAL3D, providing resolution-scalable 3D semantic segmentation of point clouds. In contrast to existing works, the proposed method does not require the whole point cloud to be available to start inference. Once a low-resolution version of the input point cloud is available, first semantic predictions can be generated in an extremely fast manner. This enables early decision-making in subsequent processing steps. As additional points become available, these are processed in parallel. To improve performance, features from previously computed scales are employed as prior knowledge at the current scale. Our experiments show that RESSCAL3D is 31-62% faster than the non-scalable baseline while keeping a limited impact on performance. To the best of our knowledge, the proposed method is the first to propose a resolution-scalable approach for 3D semantic segmentation of point clouds based on deep learning. Remco Royen, Adrian Munteanu 0001 |
ICIP | 2 |
| 2023 | Occlusion-Aware 3D Priors for Deep Learning-Based ApplicationsabstractThis paper investigates the benefits of incorporating point visibility information of 3D point clouds within a deep learning framework, using occlusion-aware 3D priors. The presented methods for deriving the visibility of each point rely on ray-casting techniques, making the proposed solution generic and sensor independent. We demonstrate the benefits of integrating point visibility using two real-world applications. In a first application, a novel data augmentation technique is proposed leveraging occlusion-aware CAD 3D priors, resulting in state-of-the-art 3D vehicle detection. In a second application, we integrate visibility information into a vehicle pose estimation pipeline based on 3D priors. The presented techniques achieve state-of-the-art performance, significantly improving both translation and rotation accuracy on the Apollo3DCar dataset. Olivier Ducastel, Yangxintong Lyu, Leon Denis, Adrian Munteanu 0001 |
MMSP | 4 |
| 2023 | PoseNormNet: Identity-Preserved Posture Normalization of 3-D Body Scans in Arbitrary PosturesabstractThree-dimensional (3-D) human models accurately represent the shape of the subjects, which is key to many human-centric industrial applications, including fashion design, body biometrics extraction, and computer animation. These tasks usually require a high-fidelity human body mesh in a canonical posture (e.g., “A” pose or “T” pose). Although 3-D scanning technology is fast and popular for acquiring the subject's body shape, automatically normalizing the posture of scanned bodies is still under-researched. Existing methods highly rely on skeleton-driven animation technologies. However, these methods require carefully designed skeleton and skin weights, which is time-consuming and fails when the initial posture is complicated. In this article, a novel deep learning-based approach, dubbed PoseNormNet, is proposed to automatically normalize the postures of scanned bodies. The proposed algorithm provides strong operability since it does not require any rigging priors and works well for subjects in arbitrary postures. Extensive experimental results on both synthetic and real-world datasets demonstrate that the proposed method achieves state-of-the-art performance in both objective and subjective terms. Xinxin Dai, Pengpeng Hu, Adrian Munteanu 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Semantic Representation and Attention Alignment for Graph Information Bottleneck in Video SummarizationabstractEnd-to-end Long Short-Term Memory (LSTM) has been successfully applied to video summarization. However, the weakness of the LSTM model, poor generalization with inefficient representation learning for inputted nodes, limits its capability to efficiently carry out node classification within user-created videos. Given the power of Graph Neural Networks (GNNs) in representation learning, we adopted the Graph Information Bottle (GIB) to develop a Contextual Feature Transformation (CFT) mechanism that refines the temporal dual-feature, yielding a semantic representation with attention alignment. Furthermore, a novel Salient-Area-Size-based spatial attention model is presented to extract frame-wise visual features based on the observation that humans tend to focus on sizable and moving objects. Lastly, semantic representation is embedded within attention alignment under the end-to-end LSTM framework to differentiate indistinguishable images. Extensive experiments demonstrate that the proposed method outperforms State-Of-The-Art (SOTA) methods. Rui Zhong 0005, Rui Wang 0107, Wenjin Yao, Shi Dong 0004, Adrian Munteanu 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | Anet: A Deep Neural Network for Automatic 3D Anthropometric Measurement Extractionabstract3D Anthropometric measurement extraction is of paramount importance for several applications such as clothing design, online garment shopping, and medical diagnosis, to name a few. State-of-the-art 3D anthropometric measurement extraction methods estimate the measurements either through some landmarks found on the input scan or by fitting a template to the input scan using optimization-based techniques. Finding landmarks is very sensitive to noise and missing data. Template-based methods address this problem, but the employed optimization-based template fitting algorithms are computationally very complex and time-consuming. To address the limitations of existing methods, we propose a deep neural network architecture which fits a template to the input scan and outputs the reconstructed body as well as the corresponding measurements. Unlike existing template-based anthropocentric measurement extraction methods, the proposed approach does not need to transfer and refine the measurements from the template to the deformed template, thereby being faster and more accurate. A novel loss function, especially developed for 3D anthropometric measurement extraction is introduced. Additionally, two large datasets of complete and partial front-facing scans are proposed and used in training. This results in two models, dubbedAnet-completeandAnet-partial, which extract the body measurements from complete and partial front-facing scans, respectively. Experimental results on synthesized data as well as on real 3D scans captured by a photogrammetry-based scanner, an Azure Kinect sensor, and the very recent TrueDepth camera system demonstrate that the proposed approach systematically outperforms the state-of-the-art methods in terms of accuracy and robustness. Nastaran Nourbakhsh Kaashki, Pengpeng Hu, Adrian Munteanu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | MONO6D: Monocular Vehicle 6D Pose Estimation with 3D PriorsabstractPredicting the 6DoF pose of vehicles from a single view image without additional constraints remains an ill-posed problem. Current monocular approaches require expensive and time-consuming annotations of vehicle-specific feature points and/or the 2D-3D feature correspondences. In this paper, we propose a novel monocular approach for vehicle pose estimation in SE(3), dubbed Mono6D, that uses vehicle 3D priors provided by vehicle make-and-model recognition methods to estimate the 6D pose. The proposed method mainly consists of: 1) a two-separate-branch module to learn multi-modal representations; 2) a fusion schema to learn pose-specific representative embeddings. The experimental results show that the proposed method is superior to the state-of-the-art approaches in both objective and subjective terms. Yangxintong Lyu, Remco Royen, Adrian Munteanu 0001 |
ICIP | 3 |
| 2022 | Predicting high-fidelity human body models from impaired point clouds
Pengpeng Hu, Xinxin Dai, Adrian Munteanu 0001 |
Signal Process. | 4 |
| 2022 | 3DBodyNet: Fast Reconstruction of 3D Animatable Human Body Shape From a Single Commodity Depth CameraabstractKnowledge about individual body shape has numerous applications in various domains such as healthcare, fashion and personalized entertainment. Most of the depth based whole body scanners need multiple cameras surrounding the user and requiring the user to keep a canonical pose strictly during capturing depth images. These scanning devices are expensive and need professional knowledge for operation. In order to make 3D scanning as easy-to-use and fast as possible, there is a great demand to simplify the process and to reduce the hardware requirements. In this paper, we propose a deep learning algorithm, dubbed 3DBodyNet, to rapidly reconstruct the 3D shape of human bodies using a single commodity depth camera. As easy-to-use as taking a photo using a mobile phone, our algorithm only needs two depth images of the front-facing and back-facing bodies. The proposed algorithm has strong operability since it is insensitive to the pose and the pose variations between the two depth images. It can also reconstruct an accurate body shape for users under tight/loose clothing. Another advantage of our method is the ability to generate an animatable human body model. Extensive experimental results show that the proposed method enables robust and easy-to-use animatable human body reconstruction, and outperforms the state-of-the-art methods with respect to running time and accuracy. Pengpeng Hu, Edmond S. L. Ho, Adrian Munteanu 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | cREAtIve: reconfigurable embedded artificial intelligenceabstractcREAtIve targets the development of novel highly-adaptable embedded deep learning solutions for automotive and traffic monitoring applications, including position sensor processing, scene interpretation based on LiDAR, and object detection and classification in thermal images for traffic camera systems. These applications share the need for deep learning solutions tailored for deployment on embedded devices with limited resources and featuring high adaptability and robustness to changing environmental conditions. cREAtIve develops knowledge, tools and methods that enable hardware-efficient, adaptable, and robust deep learning. Poona Bahrebar, Leon Denis, Maxim Bonnaerens, Kristof Coddens, Joni Dambre, Wouter Favoreel, Illia Khvastunov, Adrian Munteanu 0001, Hung Nguyen-Duc, Stefan Schulte 0001, Dirk Stroobandt, Ramses Valvekens, Nick Van den Broeck, Geert Verbruggen |
CF | 8 |
| 2021 | Graph convolutional neural networks with node transition probability-based message passing and DropNode regularization
Tien Do Huu, Duc Minh Nguyen 0002, Giannis Bekoulis, Adrian Munteanu 0001, Nikos Deligiannis |
Expert Syst. Appl. | 4 |
| 2021 | MaskLayer: Enabling scalable deep learning solutions by training embedded feature sets
Remco Royen, Leon Denis, Quentin Bolsee, Pengpeng Hu, Adrian Munteanu 0001 |
Neural Networks | 5 |
| 2021 | Frame-Wise CNN-Based Filtering for Intra-Frame Quality Enhancement of HEVC VideosabstractThe paper proposes a novel frame-wise filtering method based on Convolutional Neural Networks (CNNs) for enhancing the quality of HEVC decoded videos. A novel deep neural network architecture is proposed for post-filtering the entire intra-coded videos. A novel scheme utilizing frame-size patches is employed for training the network. The proposed method filters the luminance channel separately from the pair of chrominance channels. A novel patch generation paradigm is proposed where, for each color channel, the corresponding mode map is generated based on the HEVC intra-prediction mode index and block segmentation. The proposed CNN-based filtering method is an alternative to the traditional HEVC built-in in-loop filtering module for intra-coded frames. Experimental results on standard test sequences show that the proposed method outperforms the HEVC standard with average BD-rate savings of 11.1% and an average BD-PSNR improvement of 0.602 dB. The average relative improvement in ΔPSNR is around 105% at QP = 42 and around 85% at QP = 32 compared with state-of-the-art machine-learning-based methods. Hongyue Huang, Ionut Schiopu, Adrian Munteanu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Learning to Estimate the Body Shape Under Clothing From a Single 3-D ScanabstractEstimating the 3-D human body shape and pose under clothing is important for many applications, including virtual try-on, noncontact body measurement, and avatar creation for virtual reality. Existing body shape estimation methods formulate this task as an optimization problem by fitting a parametric body model to a single dressed-human scan or a sequence of dressed-human meshes for a better accuracy. This is impractical for many applications that require fast acquisition, such as gaming and virtual try-on due to the expensive computation. In this article, we propose the first learning-based approach to estimate the human body shape under clothing from a single dressed-human scan, dubbed Body PointNet. The proposed Body PointNet operates directly on raw point clouds and predicts the undressed body in a coarse-to-fine manner. Due to the nature of the data—aligned paired dressed scans and undressed bodies; and genus-0 manifold meshes (i.e., single-layer surfaces)—we face a major challenge of lacking training data. To address this challenge, we propose a novel method to synthesize the dressed-human pseudoscans and corresponding ground truth bodies. A new large-scale dataset, dubbed body under virtual garments, is presented, employed for the learning task of body shape estimation from 3-D dressed-human scans. Comprehensive evaluations show that the proposed Body PointNet outperforms the state-of-the-art methods in terms of both accuracy and running time. Pengpeng Hu, Nastaran Nourbakhsh Kaashki, Vasile Teodor Dadarlat, Adrian Munteanu 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | Low-Rank Constrained Super-Resolution for Mixed-Resolution Multiview VideoabstractMultiview video allows for simultaneously presenting dynamic imaging from multiple viewpoints, enabling a broad range of immersive applications. This paper proposes a novel super-resolution (SR) approach to mixed-resolution (MR) multiview video, whereby the low-resolution (LR) videos produced by MR camera setups are up-sampled based on the neighboring HR videos. Our solution analyzes the statistical correlation of different resolutions between multiple views, and introduces a low-rank prior based SR optimization framework using local linear embedding and weighted nuclear norm minimization. The target HR patch is reconstructed by learning texture details from the neighboring HR camera views using local linear embedding. A low-rank constrained patch optimization solution is introduced to effectively restrain visual artifacts and the ADMM framework is used to solve the resulting optimization problem. Comprehensive experiments including objective and subjective test metrics demonstrate that the proposed method outperforms the state-of-the-art SR methods for MR multiview video. Shao-Ping Lu, Senmao Li, Gauthier Lafruit, Ming-Ming Cheng, Adrian Munteanu 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | A Study Of Prediction Methods Based On Machine Learning Techniques For Lossless Image CodingabstractIn recent years, a new research strategy for coding has emerged by exploring the advances brought by modern machine learning techniques. Novel hybrid coding solutions were proposed by replacing specific modules in conventional coding frameworks with more efficient modules based on innovative deep-learning (DL) methods. The paper studies first our recently proposed DL-based prediction methods for lossless image coding by analyzing their designs and employed ML concepts. A novel neural network architecture is proposed, based on a new structure of layers which proves to deliver an improved prediction compared to the reference designs. Context-tree based bit-plane coding is employed to encode the resulting prediction error. The experimental results reveal that the proposed codec reduces the lossless coding rate with 1.9% compared to state-of-the-art DL-based methods while having 4.95% less parameters. The performance gap of almost 50% compared to traditional codecs recommends the use of ML-based tools in the design of future standards for lossless image compression. Ionut Schiopu, Adrian Munteanu 0001 |
ICIP | 2 |
| 2020 | Low-Complexity Angular Intra-Prediction Convolutional Neural Network for Lossless HEVCabstractThe paper proposes a novel low-complexity Convolutional Neural Network (CNN) architecture for block-wise angular intra-prediction in lossless video coding. The proposed CNN architecture is designed based on an efficient patch processing layer structure. The proposed CNN-based prediction method is employed to process an input patch containing the causal neighborhood of the current block in order to directly generate the predicted block. The trained models are integrated in the HEVC video coding standard to perform CNN-based angular intra-prediction and to compete with the conventional HEVC prediction. The proposed CNN architecture contains a reduced number of parameters equivalent to only 37% of that of the state-of-the-art reference CNN architecture. Experimental results show that the inference runtime is also reduced by around 5.5% compared to that of the reference method. At the same time, the proposed coding systems yield 83% to 91% of the compression performance of the reference method. The results demonstrate the potential of structural and complexity optimizations in CNN-based intra-prediction for lossless HEVC. Hongyue Huang, Ionut Schiopu, Adrian Munteanu 0001 |
MMSP | 3 |
| 2020 | CNN-Based Intra-Prediction for Lossless HEVCabstractThe paper proposes a novel block-wise prediction paradigm based on Convolutional Neural Networks (CNNs) for lossless video coding. A deep neural network model which follows a multi-resolution design is employed for block-wise prediction. Several contributions are proposed to improve neural network training. A first contribution proposes a novel loss function formulation for an efficient network training based on a new approach for patch selection. Another contribution consists in replacing all HEVC-based angular intra-prediction modes with a CNN-based intra-prediction method, where each angular prediction mode is complemented by a CNN-based prediction mode using a specifically trained model. Another contribution consists in an efficient adaptation of the CNN-based intra-prediction residual for lossless video coding. Experimental results on standard test sequences show that the proposed coding system outperforms the HEVC standard with an average bitrate improvement of around 5%. To our knowledge, the paper is the first to replace all the traditional HEVC-based angular intra-prediction modes with an intra-prediction method based on modern Machine Learning techniques for lossless video coding applications. Ionut Schiopu, Hongyue Huang, Adrian Munteanu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Deep-Learning-Based Lossless Image CodingabstractThis paper proposes a novel approach for lossless image compression. The proposed coding approach employs a deep-learning-based method to compute the prediction for each pixel, and a context-tree-based bit-plane codec to encode the prediction errors. First, a novel deep learning-based predictor is proposed to estimate the residuals produced by traditional prediction methods. It is shown that the use of a deep-learning paradigm substantially boosts the prediction accuracy compared with the traditional prediction methods. Second, the prediction error is modeled by a context modeling method and encoded using a novel context-tree-based bit-plane codec. Codec profiles performing either one or two coding passes are proposed, trading off complexity for compression performance. The experimental evaluation is carried out on three different types of data: photographic images, lenslet images, and video sequences. The experimental results show that the proposed lossless coding approach systematically and substantially outperforms the state-of-the-art methods for each type of data. Ionut Schiopu, Adrian Munteanu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Deep Learning Based Angular Intra-Prediction for Lossless HEVC Video CodingabstractThis work proposes the fi rst block -wise prediction paradigm based on CNNs for lossless video coding. The propose prediction scheme improves the HEVC performance by replacing a set of 9 angular intra-prediction modes with an improved CNN -based prediction; these include m2, m6, m10, m14, m18, m22, m26, m30, m34. A causal neighborhood selecting a 16 x 16 block around the currently predicted block is used as input. The novel neural network model called Angular intra-Prediction Convolutional Neural Network (AP -CNN) is designed based on the U -net architecture and operates on three resolutions (16 x 16, 8 x 8,4 x 4). AP -CNN contains the following U-net structure: 10 convolutional layers (2 + 2, 2 + 2, 2) with (32, 64, 128) fi lters; 2 deconvolution layers with 32 and 64 fi lters; 2 filter concatenation layers. A final convolutional layer with one filter is used to compute a 16 x 16 block, which is further clipped out on the bottom-right corner to obtain the 4 x 4 output predicted block. AP-CNN uses a 3 x 3 window and ReLU and it employs the Adam optimizer with MSE loss function. The experimental assessment is carried out on the Y channel of two datasets: 15 HEVC Test Sequences on 8 -bit, and 7 TUT Sequences from Ultra Video Group (TUT-vSEQ). A model is trained for each of the 9 modes using a corresponding training set generated based on HEVC's optimal mode segmentation applied to 15 HD sequences from Xiph.org and a collection of RGB images from NYC Library. The size of each of the training sets varies between 6700 and 37300 batches, where one batch contains 500 samples (input blocks). Each AP -CNN model was trained during 20 epochs, and using a 90% -10% ratio for training -validation data splitting. Table 1 shows compression results for HEVC running under the lossless setting and with all intra-prediction, and for AP-CNN where Lossless HEVCIntra employs the CNN-based predictors. AP-CNN outperforms Lossless HEVC with an average bitrate improvement of around 0.85%. An increased performance is obtained on 1080p resolutions and above. The improved coding performance is due to AP-CNN's capability to combine linear and nonlinear CNN-based prediction models. Hongyue Huang, Ionut Schiopu, Adrian Munteanu 0001 |
DCC | 3 |
| 2019 | Plenoptic Camera Calibration Based on Sub-Aperture ImagesabstractAccurate and robust calibration methods for plenoptic cameras are necessary due to their increasing popularity as depth estimation devices. Existing calibration methods can be placed in two categories: microlens based and sub-aperture image based. The low pixel-count in the microlenses is a limitation for microlens-based methods. On the other hand, calibrating each sub-aperture image independently leads to misalignments, resulting in poor depth estimates due to the high sensitivity on the baseline and alignment. In this paper we propose a hybrid calibration approach which can calibrate sub-aperture images while taking both main lens and microlens geometric constraints into account. Our model is simple yet effective and offers a high degree of interpretability on the intrinsic/extrinsic parameters of the virtual sub-aperture cameras. The experimental evaluation against existing methods shows that our proposed method leads to more accurate measurements in 3D. Walid Darwish, Quentin Bolsee, Adrian Munteanu 0001 |
ICIP | 3 |
| 2019 | Depth Estimation with Occlusion Prediction in Light Field ImagesabstractThis paper addresses the problem of depth estimation in light field images by handling occlusions in a robust way. Previous methods determine the occlusion maps either by using the edge information in the center view or by employing various cues based on disparity cost. Here we propose to determine the occlusions based on both disparity cost and edge information. The proposed method first gets a collective response from all the depth cues of the different views. Then it determines the occluded pixels by relying on the edge information in this collective response. Based on this predicted occlusion map, the resulting depth map is regularized by standard graph-cut optimization and filtered using weighted median filtering. Experimental results on synthetic light field images demonstrate the superiority of the proposed method compared to the state-of-the-art both in quantitatively and qualitatively. Mrinmoy Ghorai, Adrian Munteanu 0001 |
ICIP | 2 |
| 2019 | Joint Registration of Multiple Point Sets with RefinementabstractThe recent advances in fast and affordable 3D scanners, such as the Microsoft Kinect, have triggered new developments in 3D reconstruction. However, the geometric alignment of multiple point clouds is still a very challenging task. This paper addresses the problem of registering multiple point sets by building upon the state-of-the-art Joint Registration of Multiple Point Clouds (JRMPC) algorithm. We advance over JRMPC by incorporating the surface normal orientation of each point in the Gaussian Mixture Models (GMM) employed by this method. We formally derive an expectation-conditional maximization algorithm that iteratively estimates the refined model and transformation parameters. Experiments performed on both synthetic, real indoor and outdoor data prove that incorporating the orientation information substantially improves the parameter estimation, allowing for a 50% average reduction of the registration error compared to JRMPC. Our method is compared to other pairwise approaches, demonstrating the best performance on three different realistic datasets. Léo Moulin, Ségolène Rogge, Adrian Munteanu 0001 |
ISM | 3 |
| 2019 | Consistent video projection on curved displays
Yangxintong Lyu, Shao-Ping Lu, Quentin Bolsee, Adrian Munteanu 0001 |
Signal Process. Image Commun. | 4 |
| 2019 | Scalable Wavelet-Based Coding of Irregular Meshes With Interactive Region-of-Interest SupportabstractThis paper proposes a novel functionality in wavelet-based irregular mesh coding, which is interactive region-of-interest (ROI) support. The proposed approach enables the user to define the arbitrary ROIs at the decoder side and to prioritize and decode these regions at arbitrarily high-granularity levels. In this context, a novel adaptive wavelet transform for irregular meshes is proposed, which enables: 1) varying the resolution across the surface at arbitrarily fine-granularity levels and 2) dynamic tiling, which adapts the tile sizes to the local sampling densities at each resolution level. The proposed tiling approach enables a rate-distortion-optimal distribution of rate across spatial regions. When limiting the highest resolution ROI to the visible regions, the fine granularity of the proposed adaptive wavelet transform reduces the required amount of graphics memory by up to 50%. Furthermore, the required graphics memory for an arbitrary small ROI becomes negligible compared to rendering without ROI support, independent of any tiling decisions. Random access is provided by a novel dynamic tiling approach, which proves to be particularly beneficial for large models of over 106~ 107vertices. The experiments show that the dynamic tiling introduces a limited lossless rate penalty compared to an equivalent codec without ROI support. Additionally, rate savings up to 85% are observed while decoding ROIs of tens of thousands of vertices. Jonas El Sayeh Khalil, Adrian Munteanu 0001, Peter Lambert |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Dictionary Learning-Based, Directional, and Optimized Prediction for Lenslet Image CodingabstractIn this paper, a novel approach to encode lenslet (LL) images is proposed. The method departs from traditional block-based coding structures and employs a hexagonal-shaped pixel cluster, called macro-pixel, as an elementary coding unit. A novel prediction mode based on dictionary learning is proposed, whereby macro-pixels are represented by a sparse linear combination of atoms from a generic dictionary. Additionally, an optimized linear prediction mode and a directional prediction mode specifically designed for macro-pixels are proposed. Rate-distortion optimization is utilized to select the best intra prediction mode for each macro-pixel. Experimental results on the light field image data set show that the proposed coding system outperforms HEVC and the state-of-the-art in LL image coding with an average peak signal to noise ratio gain of 3.33 and 1.41 dB, respectively, and with rate savings of 67.13% and 34.30%, respectively. Rui Zhong 0005, Ionut Schiopu, Bruno Cornelis, Shao-Ping Lu, Junsong Yuan 0001, Adrian Munteanu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | L-Infinite Predictive Coding of Depth
Wenqi Chang, Ionut Schiopu, Adrian Munteanu 0001 |
ACIVS | 3 |
| 2018 | Cnn-based Denoising of Time-Of-Flight Depth ImagesabstractThis paper is the first to propose a deep learning approach for denoising depth images produced by Time of Flight (ToF) cameras. The noise in ToF depth images is spatially nonstationary, and depends on the strength of the infrared signal coming back to the sensor. Existing ToF denoising methods do not capture the non-stationary nature of noise in such images. We propose a fairly simple, yet efficient Convolutional Neural Network (CNN) that makes use of both the depth and infrared information to denoise ToF depth images. To train the network, a novel methodology to generate a high amount of training samples is proposed. The experimental results demonstrate that the proposed CNN-based ToF denoising method substantially outperforms the state-of-the-art denoising methods, including wavelet shrinkage, least squares optimization, Bilateral Filtering and BM3D. Quentin Bolsee, Adrian Munteanu 0001 |
ICIP | 2 |
| 2018 | Synthesis of Shaking Video Using Motion Capture Data and Dynamic 3D Scene ModelingabstractImportant video processing methods such as video stabilization and deblurring often do not have ground-truth data available. This poses a great challenge in the development and parameter tunning of such methods. Synthetic shaken video is very useful to generate well-defined ground-truth datasets. Existing shaking video synthesis methods simulate shaky camera motion by performing 2D view warping using only a single 2D video, which does not always correspond to realistic 3D motions. In this paper, we introduce a novel shaking video synthesis approach. The proposed framework constructs the camera motion trajectory by making use of human motion information that is captured in the real-world. Moreover, we render the shaken video from man-made dynamic 3D scenes with detailed camera pose information. Our novel approach provides both accurate 2D visual content and camera motion trajectory in the 3D scene, which allows for evaluating the visual distortion as well as the offsets of the recovered camera trajectory. The proposed synthesis method of shaking video will benefit and ease future research on 3D-aware video stabilization. Shao-Ping Lu, Beerend Ceulemans, Miao Wang 0004, Adrian Munteanu 0001 |
ICIP | 5 |
| 2018 | Macro-Pixel Prediction Based on Convolutional Neural Networks for Lossless Compression of Light Field ImagesabstractThe paper introduces a novel macro-pixel prediction method based on Convolutional Neural Networks (CNN) for lossless compression of light field images. In the proposed method, each macro-pixel is predicted based on a volume of macro-pixels from its immediate causal neighborhood. The proposed deep neural network operates on these macro-pixel volumes and provides accurate macro-pixel prediction in light field images. The resulting macro-pixel residuals are encoded by a reference codec built based on the CALIC codec. A context modeling method for light field images is proposed. Experimental results on a large light field image dataset show that the proposed prediction method systematically and substantially outperforms state-of-the-art predictors. To our knowledge, the paper is the first to introduce deep-learning based prediction of macro-pixels, enabling efficient lossless compression of light field images. Ionut Schiopu, Adrian Munteanu 0001 |
ICIP | 2 |
| 2018 | CNN-based Prediction for Lossless Coding of Photographic ImagesabstractThe paper proposes a novel prediction paradigm in image coding based on Convolutional Neural Networks (CNN). A deep neural network is designed to provide accurate pixel-wise prediction based on a causal neighbourhood. The proposed CNN prediction method is trained on the high-activity areas in the image and it is incorporated in a lossless compression system for high-resolution photographic images. The system uses the proposed CNN-based prediction paradigm as well as LOCO-I, whereby the predictor selection is performed using a local entropy-based descriptor. The prediction errors are encoded using a CALIC-based reference codec. The experimental results show a good performance for the proposed prediction scheme compared to state-of-the-art predictors. To our knowledge, the paper is the first to introduce CNN-based prediction in image coding, and demonstrates the potential offered by machine learning methods in coding applications. Ionut Schiopu, Adrian Munteanu 0001 |
PCS | 3 |
| 2018 | Optimized wavelet-based texture representation and streaming for GPU texture mapping
Bob Andries, Jan Lemeire, Adrian Munteanu 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Robust Multiview Synthesis for Wide-Baseline Camera ArraysabstractIn many advanced multimedia systems, multiview content can offer more immersion compared to classical stereoscopy. The feeling of immersiveness is increased substantially by offering motion-parallax, as well as stereopsis. This drives both the so-called free-navigation and super-multiview technologies. However, it is currently still challenging to acquire, store, process, and transmit this type of content. This paper presents a novel multiview-interpolation framework for wide-baseline camera arrays. The proposed method comprises several novel components, including point cloud-based filtering, improved de-ghosting, multireference color blending, and depth-aware MRF-based disocclusion in painting. The method offers robustness against depth errors caused by quantization and smoothing across object boundaries. Furthermore, the available input color and depth are maximally exploited while preventing propagation of unreliable information to virtual viewpoints. The experimental results show that the proposed method outperforms the state-of-the-art View Synthesis Reference Software (VSRS 4.1) both in objective terms as well as subjectively, based on a visual assessment on a high-end light-field three-dimensional display. Beerend Ceulemans, Shao-Ping Lu, Gauthier Lafruit, Adrian Munteanu 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | 3D Mesh coding with predefined region-of-interestabstractWe introduce a novel functionality for wavelet-based irregular mesh codecs which allows for prioritizing at the encoding side a region-of-interest (ROI) over a background (BG), and for transmitting the encoded data such that the quality in these regions increases first. This is made possible by appropriately scaling wavelet coefficients. To improve the decoded geometry in the BG, we propose an ROI-aware inverse wavelet transform which only upscales the connectivity in the required regions. Results show clear bitrate and vertex savings. For a trivial front-back selection of the ROI and BG, rendering from the front saves up to 5 bits per vertex and up to 50% of the geometry, while appearing visually lossless. Jonas El Sayeh Khalil, Adrian Munteanu 0001, Peter Lambert |
ICIP | 2 |
| 2017 | Efficient directional and L1-optimized intra-prediction for light field image compressionabstractLight field images can be conveniently captured by consumer-level plenoptic cameras. However, as the resulting data rates are very high, providing efficient compression for this type of data is of critical importance. This remains an open problem which has recently attracted a lot of attention from the coding community. State-of-the-art compression systems prove to be inefficient when directly applied on this type of data due to the inherent spatial discontinuities in light field images. In this paper, a novel intra-prediction method for disk-shaped pixel clusters is proposed. An L1 minimization of the prediction residuals is performed followed by clustering of the predictors, leading to an optimized set of predictors for the macro-pixels. Furthermore, directional intra-prediction modes based on HEVC are devised for the macro-pixels. Experimental results obtained on the EPFL light field image dataset demonstrate that the proposed coding scheme yields an average of 3.22 dB and 1.45 dB gain in PSNR, and 59.6% and 30.88% average rate savings compared to HEVC and the state-of-the-art in light field image coding respectively. Rui Zhong 0005, Shizheng Wang, Bruno Cornelis, Yuanjin Zheng, Junsong Yuan 0001, Adrian Munteanu 0001 |
ICIP | 6 |
| 2017 | Scalable Feature-Preserving Irregular Mesh CodingabstractAbstract This paper presents a novel wavelet‐based transform and coding scheme for irregular meshes. The transform preserves geometric features at lower resolutions by adaptive vertex sampling and retriangulation, resulting in more accurate subsampling and better avoidance of smoothing and aliasing artefacts. By employing octree‐based coding techniques, the encoding of both connectivity and geometry information is decoupled from any mesh traversal order, and allows for exploiting the intra‐band statistical dependencies between wavelet coefficients. Improvements over the state of the art obtained by our approach are three‐fold: (1) improved rate–distortion performance over Wavemesh and IPR for both the Hausdorff and root mean square distances at low‐to‐mid‐range bitrates, most obvious when clear geometric features are present while remaining competitive for smooth, feature‐poor models; (2) improved rendering performance at any triangle budget, translating to a better quality for the same runtime memory footprint; (3) improved visual quality when applying similar limits to the bitrate or triangle budget, showing more pronounced improvements than rate–distortion curves. Jonas El Sayeh Khalil, Adrian Munteanu 0001, Leon Denis, Peter Lambert, Rik Van de Walle |
Comput. Graph. Forum | 2 |
| 2017 | Color correction for large-baseline multiview videoabstractColor misalignment correction is an important, yet unsolved problem , especially for multiview video captured by large disparity camera setups. In this paper, we introduce a robust large-baseline color correction method that preserves the original manifold structure of the input video. The manifold structure is extracted by locally linear embedding (LLE), aimed at linearly representing each pixel based on its neighbors, assuming that they are all clustered in a high-dimensional feature space. Besides the proposed manifold structure preservation constraint, the proposed method enforces spatio-temporal color consistencies and gradient preservation. The multiview color correction solution is obtained by solving a global optimization problem . Thorough objective and subjective experimental results demonstrate that our proposed approach significantly and systematically outperforms the state-of-the-art color correction methods on large-baseline multiview video data . Siqi Ye, Shao-Ping Lu, Adrian Munteanu 0001 |
Signal Process. Image Commun. | 3 |
| 2017 | Wavelet-Based L∞ Semi-regular Mesh CodingabstractPolygonal meshes are popular three-dimensional virtual representations employed in a wide range of applications. Users have very high expectations with respect to the accuracy of these virtual representations, fueling a steady increase in the processing power and performance of graphics processing hardware. This accuracy is closely related to how detailed the virtual representations are. The more detailed these representations become, the higher the amount of data that will need to be displayed, stored, or transmitted. Efficient compression techniques are of critical importance in this context. State-ofthe-art compression performance of semi-regular mesh coding systems has been achieved through the use of subdivision-based wavelet coding techniques. However, the vast majority of these codecs are optimized with respect to the L2distortion metric, i.e., the average error. This makes them unsuitable for applications where each input signal sample has a certain significance. To alleviate this problem, we propose to optimize the mesh codec with respect to the L∞metric, which allows for the control of the local reconstruction error. This paper proposes novel data-dependent formulations for the L∞distortion. The proposed L∞estimators are incorporated in a state-of-the-art wavelet-based semi-regular mesh codec. The resulting coding system offers scalability in L∞sense. The experiments demonstrate the advantages of L∞coding in providing a tight control on the local reconstruction error. Furthermore, the proposed data-dependent L∞approaches significantly improve estimation accuracy, reducing the classical low-rate gap between the estimated and actual L∞distortion observed for previous L∞estimators. Ruxandra-Marina Florea, Adrian Munteanu 0001, Shao-Ping Lu, Peter Schelkens |
IEEE Trans. Multim. | 2 |
| 2017 | Scalable texture compression using the wavelet transform
Bob Andries, Jan Lemeire, Adrian Munteanu 0001 |
Vis. Comput. | 3 |
| 2016 | The Rate Loss in Binary Source Coding with Decoder Side InformationabstractSummary form only given. Motivated by the correlation channel modeling problem in practical applications, such as distributed video coding, we study the binary source coding of a uniform source with side information, under asymmetric correlation channel assumptions. First, we consider the case where side information is available to both the encoder and decoder, and give an analytical formula for the rate-distortion bound. Then, we consider the side information to be available only to the decoder and present the derivation of the associated Wyner-Ziv rate-distortion bound. Most importantly, we characterize the evolution of the rate-loss suffered by Wyner-Ziv coding, for all possible binary asymmetric correlation channels. Andrei Sechelea, Adrian Munteanu 0001, Samuel Cheng 0001, Nikos Deligiannis |
DCC | 2 |
| 2016 | Efficient MRF-based disocclusion inpainting in multiview videoabstractView synthesis using depth image-based rendering generates virtual viewpoints of a 3D scene based on texture and depth information from a set of available cameras. One of the core components in view synthesis is image inpainting which performs the reconstruction of areas that were occluded in the available cameras but are visible from the virtual viewpoint. Inpainting methods based on Markov random fields (MRFs) have been shown to be very effective in inpainting large areas in images. In this paper, we propose a novel MRF-based in-painting method for multiview video. The proposed method steers the MRF optimization towards completion from background to foreground and exploits the available depth information in order to avoid bleeding artifacts. The proposed approach allows for efficiently filling-in large disocclusion areas and greatly accelerates execution compared to traditional MRF-based inpainting techniques. The experimental results show that view synthesis based on the proposed inpainting method systematically improves performance over the state-of-the-art in multiview view synthesis. Average PSNR gains up to 1.88 dB compared to the MPEG View Synthesis Reference software were observed. Beerend Ceulemans, Shao-Ping Lu, Gauthier Lafruit, Peter Schelkens, Adrian Munteanu 0001 |
ICME | 5 |
| 2016 | L1-optimized linear prediction for light field image compressionabstractThe advent of consumer-level plenoptic cameras has sparkled the interest towards the design of efficient compression techniques for light field images. State-of-the-art compression systems such as HEVC prove to be inefficient when directly applied on this type of data due to the inherent spatial discontinuities among neighboring microlens images. In this paper, a novel light field image compression system is proposed. The disk-shaped pixel clusters corresponding to each microlens in the light field image are efficiently predicted based on the neighboring disks. In this context, an optimized linear prediction design based on L1 minimization of the residuals is proposed. K-means clustering is employed on training data in order to determine the optimized set of predictors. The experimental results on an extensive set of light field images demonstrate that the proposed coding scheme yields an average of 2.93 dB and 3.22 dB gain in PSNR, and 52.67% and 57.27% average rate savings compared to HEVC and JPEG2000 respectively. Rui Zhong 0005, Shizheng Wang, Bruno Cornelis, Yuanjin Zheng, Junsong Yuan 0001, Adrian Munteanu 0001 |
PCS | 6 |
| 2016 | On the Rate-Distortion Function for Binary Source Coding With Side InformationabstractWe present an in-depth analysis of the problem of lossy compression of binary sources in the presence of correlated side information, where the correlation is given by a generic binary asymmetric channel and the Hamming distance is the distortion metric. Our analysis is motivated by systematic rate-distortion gains observed when applying asymmetric correlation models in Wyner-Ziv video coding. First, we derive for the first time the rate-distortion function for conventional predictive coding in the binary-asymmetric-correlation-channel scenario. Second, we propose a new bound for the case where the side information is only available at the decoder-Wyner-Ziv coding. We conjecture this bound to be tight. We show that the maximum rate needed to encode as well as the maximum rate-loss of Wyner-Ziv coding relative to predictive coding corresponds to uniform sources and symmetric correlations. Importantly, we show that the upper bound on the rate-loss established by Zamir is not tight and that the maximum value is actually significantly lower. Moreover, we prove that the only binary correlation channel that incurs no rate-loss for Wyner-Ziv coding compared with predictive coding is the Z-channel. Finally, we complement our analysis with new compression performance results obtained with our state-of-the-art Wyner-Ziv video coding system. Andrei Sechelea, Adrian Munteanu 0001, Samuel Cheng 0001, Nikos Deligiannis |
IEEE Trans. Commun. | 2 |
| 2015 | Efficient scalable compression of sparsely sampled imagesabstractAdvanced sparse sampling acquisition systems capture only scattered information from the continuous image domain. Unfortunately, conventional image encoders are not yet able to properly compress arbitrarily subsampled image data. This work introduces a system leveraging the JPEG 2000 image compression framework by enabling scalable compression of the selected image samples. Using a complete dictionary of CDF 9/7 wavelets, a minimum l1-norm compressed sensing solution is recovered which can be fed directly into the encoder, producing a bitstream that can be decoded with existing JPEG 2000-compliant implementations. Experiments on standard images with quasi-random subsampling demonstrate that the proposed system outperforms regular JPEG 2000 compression of stacked sample images and quad-tree based compression for point-clouds. We also demonstrate the robustness of the technique for images that infringe the sparsity prior of compressed sensing. Colas Schretter, David Blinder, Tim Bruylants, Peter Schelkens, Adrian Munteanu 0001 |
ICIP | 5 |
| 2015 | CDF 9/7 wavelets as sparsifying operator in compressive holographyabstractCompressive sensing is a mathematical framework, which seeks to capture the information of an object using as few measurements as possible. Recently, it has been applied to holography, where the most frequently used reconstruction method is l1-norm minimization with the Haar wavelet as the sparsifying operator. In this work, we promote the CDF 9/7 wavelet as the sparsifying operator. We demonstrate that the CDF 9/7 wavelet performs better than the Haar wavelet. David Blinder, Stijn Bettens, Heidi Ottevaere, Adrian Munteanu 0001, Peter Schelkens |
ICIP | 5 |
| 2015 | Color retargeting: Interactive time-varying color image composition from time-lapse sequencesabstractIn this paper, we present an interactive static image composition approach, namely color retargeting , to flexibly represent time-varying color editing effect based on time-lapse video sequences. Instead of performing precise image matting or blending techniques, our approach treats the color composition as a pixel-level resampling problem. In order to both satisfy the user’s editing requirements and avoid visual artifacts, we construct a globally optimized interpolation field. This field defines from which input video frames the output pixels should be resampled. Our proposed resampling solution ensures that (i) the global color transition in the output image is as smooth as possible, (ii) the desired colors/objects specified by the user from different video frames are well preserved, and (iii) additional local color transition directions in the image space assigned by the user are also satisfied. Various examples have been shown to demonstrate that our efficient solution enables the user to easily create time-varying color image composition results. Shao-Ping Lu, Guillaume Dauphin, Gauthier Lafruit, Adrian Munteanu 0001 |
Comput. Vis. Media | 4 |
| 2015 | Wavelet based volumetric medical image compressionabstractThe amount of image data generated each day in health care is ever increasing, especially in combination with the improved scanning resolutions and the importance of volumetric image data sets. Handling these images raises the requirement for efficient compression, archival and transmission techniques. Currently, JPEG 2000׳s core coding system, defined in Part 1, is the default choice for medical images as it is the DICOM-supported compression technique offering the best available performance for this type of data. Yet, JPEG 2000 provides many options that allow for further improving compression performance for which DICOM offers no guidelines. Moreover, over the last years, various studies seem to indicate that performance improvements in wavelet-based image coding are possible when employing directional transforms. In this paper, we thoroughly investigate techniques allowing for improving the performance of JPEG 2000 for volumetric medical image compression. For this purpose, we make use of a newly developed generic codec framework that supports JPEG 2000 with its volumetric extension (JP3D), various directional wavelet transforms as well as a generic intra-band prediction mode. A thorough objective investigation of the performance-complexity trade-offs offered by these techniques on medical data is carried out. Moreover, we provide a comparison of the presented techniques to H.265/MPEG-H HEVC, which is currently the most state-of-the-art video codec available. Additionally, we present results of a first time study on the subjective visual performance when using the aforementioned techniques. This enables us to provide a set of guidelines and settings on how to optimally compress medical volumetric images at an acceptable complexity level. Tim Bruylants, Adrian Munteanu 0001, Peter Schelkens |
Signal Process. Image Commun. | 2 |
| 2015 | Spatio-Temporally Consistent Color and Structure Optimization for Multiview Video Color CorrectionabstractWhen compared to conventional 2-D video, multiview video can significantly enhance the visual 3-D experience in 3-D applications by offering horizontal parallax. However, when processing images originating from different views, it is common that the colors between the different cameras are not well- calibrated . To solve this problem, a novel energy function -based color correction method for multiview camera setups is proposed to enforce that colors are as close as possible to those in the reference image but also that the overall structural information is well-preserved. The proposed system introduces a spatio-temporal correspondence matching method to ensure that each pixel in the input image gets bijectively mapped to a reference pixel. By combining this mapping with the original structural information, we construct a global optimization algorithm in a Laplacian matrix formulation and solve it using a sparse matrix solver. We further introduce a novel forward-reverse objective evaluation model to overcome the problem of lack of ground truth in this field. The visual comparisons are shown to outperform state-of-the-art multiview color correction methods, while the objective evaluation reports PSNR gains of up to 1.34 dB and SSIM gains of up to 3.2%, respectively. Shao-Ping Lu, Beerend Ceulemans, Adrian Munteanu 0001, Peter Schelkens |
IEEE Trans. Multim. | 3 |
| 2014 | Optimized quantization of wavelet subbands for high quality real-time texture compressionabstractThis paper proposes a new wavelet-based system for fixed-rate texture compression in 3D graphics applications. An analysis of the optimized quantization of the wavelet subbands is carried out, focusing on Lloyd-Max quantization of subband blocks as well as on quantization techniques employed in existing texture compression formats on GPUs. The results demonstrate that the proposed compression technique yields higher performance compared to existing GPU texture compression schemes and previously proposed transformed-based systems. Moreover, compared to conventional schemes, the wide range of quantization schemes results in a much wider range of available bitrates. Additionally, the proposed scheme offers real-time execution, being suitable for real time rendering applications. Bob Andries, Jan Lemeire, Adrian Munteanu 0001 |
ICIP | 3 |
| 2014 | Progressively refined wyner-ziv video coding for visual sensorsabstractWyner-Ziv video coding constitutes an alluring paradigm for visual sensor networks, offering efficient video compression with low complexity encoding characteristics. This work presents a novel hash-driven Wyner-Ziv video coding architecture for visual sensors, implementing the principles of successively refined Wyner-Ziv coding. To this end, so-called side-information refinement levels are constructed for a number of grouped frequency bands of the discrete cosine transform. The proposed codec creates side-information by means of an original overlapped block motion estimation and pixel-based multihypothesis prediction technique, specifically built around the pursued refinement strategy. The quality of the side-information generated at every refinement level is successively improved, leading to gradually enhanced Wyner-Ziv coding performance. Additionally, this work explores several temporal prediction structures, including a new hierarchical unidirectional prediction structure, providing both temporal scalability and low delay coding. Experimental results include a thorough evaluation of our novel Wyner-Ziv codec, assessing the impact of the proposed successive refinement scheme and the supported temporal prediction structures for a wide range of hash configurations and group of pictures sizes. The results report significant compression gains with respect to benchmark systems in Wyner-Ziv video coding (e.g., up to 42.03% over DISCOVER) as well as versus alternative state-of-the-art schemes refining the side-information. Nikos Deligiannis, Frederik Verbist, Jürgen Slowack, Rik Van de Walle, Peter Schelkens, Adrian Munteanu 0001 |
ACM Trans. Sens. Networks | 6 |
| 2013 | Reversible DCT-based lossy-to-lossless still image compressionabstractThis paper presents a novel still image compression scheme that extends the traditional JPEG standard with lossy-to-lossless compression support. The system follows a two-layer design approach which allows for backward compatibility with the conventional JPEG standard for the base layer and provides lossless compression when decoding the enhancement layer. The system employs several coding tools, including quadtree coding, spatial domain prediction, reversible discrete cosine transforms, and context-based arithmetic coding to efficiently encode losslessly the enhancement layer. Performance evaluations using a standard JPEG set of images show that the proposed system yields similar lossless compression performance to the state-of-the-art single-layer JPEG-LS standard while providing quality scalability and a JPEG-compatible base layer at the same time. Heng Chen 0003, Geert Braeckman, Adrian Munteanu 0001, Peter Schelkens |
ICIP | 3 |
| 2013 | Optimized segmentation of H.264/AVC video for HTTP adaptive streaming
Jan Lievens, Shahid M. Satti, Nikos Deligiannis, Peter Schelkens, Adrian Munteanu 0001 |
IM | 5 |
| 2013 | Real-time texture sampling and reconstruction with wavelet filtersabstractCurrently, the use of the 2D wavelet transform in texture compression for real-time texture mapping on the GPU is limited. The main cause of this is the lack of real-time texture filtering implementations which do not require specialized hardware. This work proposes a novel system to perform 2D wavelet reconstruction and bilinear texture filtering using a high performance GPU shader. The system is able to generate a performant GLSL shader for arbitrary wavelet filter configurations. This goes beyond earlier works in the literature proposing Haar wavelet and Discrete Cosine Transform (DCT) implementations on the GPU. We analyse the shader performance and run-time complexity for several wavelet filters. The experimental results show that filters longer than Haar are deployable on the GPU while maintaining accurate texture filtering and real-time performance. Bob Andries, Adrian Munteanu 0001, Jan Lemeire, Peter Schelkens |
MMSP | 2 |
| 2013 | Making Communication a First-Class Citizen in Multicore PartitioningabstractComputation-intensive image processing applications need to be implemented on multicore architectures. If they are to be executed efficiently on such platforms, the underlying data and/or functions should be partitioned and distributed among the processors. The optimal partitioning approach is the one which aims to minimize the inter-processor communication while maximizing the load balance. With the continuously increasing number of cores which exacerbates the demand for more complex memory hierarchies, non-uniform memory access, etc., on-chip communication has gained a significant role in taking advantage of the multicore chips. Therefore, making partitioning decisions just based on conventional performance results and without communication profiling is suboptimal. In this paper, we explore the behavior of a mesh decoder as a case study in terms of communication and computation, and propose models that allow early prediction of the application's behavior. Using these models, profiling the application for all of the input samples is not necessary anymore. As a result, communication- and computation-aware parallelization could be performed faster and easier. Poona Bahrebar, Ruxandra-Marina Florea, Wim Heirman, Leon Denis, Adrian Munteanu 0001, Dirk Stroobandt |
PDP | 5 |
| 2013 | Visually lossless screen content coding using HEVC base-layerabstractThis paper presents a novel two-layer coding framework targeting visually lossless compression of screen content video. The proposed framework employs the conventional HEVC standard for the base-layer. For the enhancement layer, a hybrid of spatial and temporal block-prediction mechanism is introduced to guarantee a small energy of the error-residual. Spatial prediction is generally chosen for dynamic areas, while temporal predictions yield better prediction for static areas in a video frame. The prediction residual is quantized based on whether a given block is static or dynamic. Run-length coding, Golomb based binarization and context-based arithmetic coding are employed to efficiently code the quantized residual and form the enhancement-layer. Performance evaluations using 4:4:4 screen content sequences show that, for visually lossless video quality, the proposed system significantly saves the bit-rate compared to the two-layer lossless HEVC framework. Geert Braeckman, Shahid M. Satti, Heng Chen 0003, Adrian Munteanu 0001, Peter Schelkens |
VCIP | 4 |
| 2013 | Probabilistic motion-compensated prediction in distributed video coding
Frederik Verbist, Nikos Deligiannis, Joeri Barbarien, Peter Schelkens, Adrian Munteanu 0001, Jan Cornelis 0001 |
Multim. Tools Appl. | 6 |
| 2013 | JPSearch: An answer to the lack of standardization in mobile image retrieval
Frederik Temmermans, Mario Döller, Iris Vanhamel, Bart Jansen 0001, Adrian Munteanu 0001, Peter Schelkens |
Signal Process. Image Commun. | 5 |
| 2012 | An optimization algorithm for scalable multiple description scalar quantizers
Shahid M. Satti, Nikos Deligiannis, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ISITA | 3 |
| 2012 | Feedback-constrained Wyner-Ziv video codingabstractDistributed video coding (DVC) systems described in the literature often make use of a feedback channel to determine the rate. However, supporting such a feedback channel in practice may be difficult, particularly considering that current approaches are unable to incorporate constraints on feedback channel usage. Therefore, in this paper we propose a system in which the number of requests per Wyner-Ziv frame can be constrained to a fixed value. Our technique involves decoder-side rate estimation, modeling of its accuracy, and finally defining the number of Wyner-Ziv bits for each of the requests allowed. The experimental results indicate that the performance loss is limited when compared to a configuration operating without constraints on the feedback channel. Jürgen Slowack, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle |
PCS | 4 |
| 2012 | Decoder-driven mode decision in a block-based distributed video codec
Stefaan Mys, Jürgen Slowack, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle |
Multim. Tools Appl. | 6 |
| 2012 | Efficient adaptive-shape partitioning of video
Kenneth Vermeirsch, Jan De Cock, Stijn Notebaert, Peter Lambert, Joeri Barbarien, Adrian Munteanu 0001, Rik Van de Walle |
Multim. Tools Appl. | 6 |
| 2012 | Efficient Low-Delay Distributed Video CodingabstractDistributed video coding (DVC) is a video coding paradigm that allows for a low-complexity encoding process by exploiting the temporal redundancies in a video sequence at the decoder side. State-of-the-art DVC systems exhibit a structural coding delay since exploiting the temporal redundancies through motion-compensated interpolation requires the frames to be decoded out of order. To alleviate this problem, we propose a system based on motion-compensated extrapolation that allows for efficient low-delay video coding with low complexity at the encoder. The proposed extrapolation technique first estimates the motion field between the two most recently decoded frames using the Lucas–Kanade algorithm. The obtained motion field is then extrapolated to the current frame using an extrapolation grid. The proposed techniques are implemented into a novel architecture featuring hybrid block-frequency Wyner–Ziv coding as well as mode decision. Results show that having references from both temporal directions in interpolation provides superior rate-distortion performance over a single temporal direction in extrapolation, as expected. However, the proposed extrapolation method is particularly suitable for low-delay coding as it performs better than H.264/AVC intra, and it is even able to outperform the interpolation-based DVC codec from DISCOVER for several sequences. Jozef Skorupa, Jürgen Slowack, Stefaan Mys, Nikos Deligiannis, Jan De Cock, Peter Lambert, Christos Grecos, Adrian Munteanu 0001, Rik Van de Walle |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2012 | Distributed Video Coding With Feedback Channel ConstraintsabstractMany of the distributed video coding (DVC) systems described in the literature make use of a feedback channel from the decoder to the encoder to determine the rate. However, the number of requests through the feedback channel is often high, and as a result the overall delay of the system could be unacceptable in practical applications. As a solution, feedback-free DVC systems have been proposed, but the problem with these solutions is that they incorporate a difficult trade-off between encoder complexity and compression performance. Recognizing that a limited form of feedback may be supported in many video-streaming scenarios, in this paper we propose a method for constraining the number of feedback requests to a fixed maximum number of$N$requests for an entire Wyner-Ziv (WZ) frame. The proposed technique estimates the WZ rate at the decoder using information obtained from previously decoded WZ frames and defines the$N$requests by minimizing the expected rate overhead. Tests on eight sequences show that the rate penalty is less than 5% when only five requests are allowed per WZ frame (for a group of pictures of size four). Furthermore, due to improvements from previous work, the system is able to perform better than or similar to DISCOVER even when up to two requests per WZ frame are allowed. The practical usefulness of the proposed approach is studied by estimating end-to-end delay and encoder buffer requirements, indicating that DVC with constrained feedback can be an important solution in the context of video-streaming scenarios. Jürgen Slowack, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2012 | Side-Information-Dependent Correlation Channel Estimation in Hash-Based Distributed Video CodingabstractIn the context of low-cost video encoding, distributed video coding (DVC) has recently emerged as a potential candidate for uplink-oriented applications. This paper builds on a concept of correlation channel (CC) modeling, which expresses the correlation noise as being statistically dependent on the side information (SI). Compared with classical side-information-independent (SII) noise modeling adopted in current DVC solutions, it is theoretically proven that side-information-dependent (SID) modeling improves the Wyner-Ziv coding performance. Anchored in this finding, this paper proposes a novel algorithm for online estimation of the SID CC parameters based on already decoded information. The proposed algorithm enables bit-plane-by-bit-plane successive refinement of the channel estimation leading to progressively improved accuracy. Additionally, the proposed algorithm is included in a novel DVC architecture that employs a competitive hash-based motion estimation technique to generate high-quality SI at the decoder. Experimental results corroborate our theoretical gains and validate the accuracy of the channel estimation algorithm. The performance assessment of the proposed architecture shows remarkable and consistent coding gains over a germane group of state-of-the-art distributed and standard video codecs, even under strenuous conditions, i.e., large groups of pictures and highly irregular motion content. Nikos Deligiannis, Joeri Barbarien, Adrian Munteanu 0001, Athanassios N. Skodras, Peter Schelkens |
IEEE Trans. Image Process. | 4 |
| 2011 | Distributed coding of endoscopic videoabstractTriggered by the challenging prerequisites of wireless capsule endoscopic video technology, this paper presents a novel distributed video coding (DVC) scheme, which employs an original hash-based side-information creation method at the decoder. In contrast to existing DVC schemes, the proposed codec generates high quality side-information at the decoder, even under the strenuous motion conditions encountered in endoscopic video. Performance evaluation using broad endoscopic video material shows that the proposed approach brings notable and consistent compression gains over various state-of-the-art video codecs at the additional benefit of vastly reduced encoding complexity. Nikos Deligiannis, Frederik Verbist, Joeri Barbarien, Jürgen Slowack, Rik Van de Walle, Peter Schelkens, Adrian Munteanu 0001 |
ICIP | 7 |
| 2011 | Intra-WZ quantization mismatch in distributed video codingabstractDuring the past decade, Distributed Video Coding (DVC) has emerged as a new video coding paradigm, shifting the complexity from the encoder - to the decoder-side. This paper addresses a problem of current DVC architectures that has not been studied in the literature so far, that is, the mismatch between the intra and Wyner-Ziv (WZ) quantization processes. Due to this mismatch, WZ rate is spent even for spatial regions that are accurately approximated by the side-information. As a solution, this paper proposes side-information generation using selective unidirectional motion compensation from temporally adjacent WZ frames. Experimental results show that the proposed approach yields promising WZ rate gains of up to 7% relative to the conventional method. Jürgen Slowack, Jozef Skorupa, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001 |
ICIP | 6 |
| 2011 | Improved intra mode signaling for HEVCabstractIn the current development of HEVC, compression performance improved significantly compared to H.264/AVC for both inter pictures and intra pictures. With intra compression, the main reason for this improvement is the large in crease in intra prediction directions (up to 34). The downside of having a larger number of modes is that they increase the signaling overhead in the bitstream. In this paper, a low complexity intra mode prediction algorithm is proposed which improves the mode prediction accuracy. This is achieved by exploiting the correlation between the prediction directions of the neighboring prediction units and that of the encoded prediction unit. As a result, more efficient intra mode signaling can be achieved with minimal impact on encoder and decoder complexity. On average, 0.33% bitrate improvement is obtained by employing the proposed algorithm. For sequences that are encoded with a high number of directional intra modes, around 1% bitrate improvement is measured. Glenn Van Wallendael, Sebastiaan Van Leuven, Jan De Cock, Peter Lambert, Rik Van de Walle, Joeri Barbarien, Adrian Munteanu 0001 |
ICME | 7 |
| 2010 | Modeling Wavelet Coefficients for Wavelet Subdivision Transforms of 3D Meshes
Shahid M. Satti, Leon Denis, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
ACIVS (1) | 3 |
| 2010 | Semi-regular remeshing with reduced remeshing errorabstractIn this paper we present a remeshing algorithm which drastically reduces the aliasing artifacts inherent in regularly-sampled remeshed objects. Starting from a semi-regular mesh, the proposed algorithm reduces the remeshing error and avoids aliasing by displacing vertices such that most samples of the original mesh are present in the remeshed model as well. Computational efficiency is provided by using a search-tree, which efficiently gathers vertices near a given point in a 3D space. Compared to the state-of-the-art semi-regular remesher, the proposed remesher drastically improves the visual quality of the high-frequency regions in remeshed objects. Additionally, the proposed algorithm yields a lower remeshing error, which is reflected by a significantly increased PSNR upper-bound in wavelet-based compression of such meshes. Leon Denis, Adrian Munteanu 0001, Peter Schelkens |
ICIP | 2 |
| 2010 | Bitplane intra coding with decoder-side mode decision in distributed video codingabstractWhile distributed video coding (DVC) has emerged as a new video coding paradigm, the compression performance of current systems is still low compared to conventional solutions such as H.264/AVC. While the latter uses many coding modes and an efficient mode decision strategy for choosing the best mode, in DVC, only a limited number of modes has been developed so far. Since encoder-side mode decision in DVC increases encoder's complexity, in this paper, we introduce decoder-side mode decision choosing between bitplane WZ coding and bitplane intra coding. This strategy proves to be efficient, delivering rate gains up to 22% over DISCOVER, without increasing the complexity at the encoder. Jürgen Slowack, Stefaan Mys, Jozef Skorupa, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001 |
ICIP | 7 |
| 2010 | Compensating for Motion Estimation Inaccuracies in DVC
Jürgen Slowack, Jozef Skorupa, Stefaan Mys, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle |
ICISP | 6 |
| 2010 | Efficient error control in 3D mesh codingabstractOur recently proposed wavelet-based L-infinite-constrained coding approach for meshes ensures that the maximum error between the vertex positions in the original and decoded meshes is guaranteed to be lower than a given upper bound. Instantiations of both L-2 and L-infinite coding approaches are demonstrated for MESHGRID, which is a scalable 3D object encoding system, part of MPEG-4 AFX. In this survey paper, we compare the novel L-infinite distortion estimator against the L-2 distortion estimator which is typically employed in 3D mesh coding systems. In addition, we show that, under certain conditions, the L-infinite estimator can be exploited to approximate the Hausdorff distance in real-time implementations. Dan C. Cernea, Adrian Munteanu 0001, Alin Alecu, Jan Cornelis 0001, Peter Schelkens, Francisco Morán |
MMSP | 2 |
| 2010 | Correlation modeling with decoder-side quantization distortion estimation for distributed video codingabstractAiming for low-complexity encoding, distributed video coders still fail to achieve the performance of current industrial standards for video coding. One of most important problems in this area is the accurate modeling of the correlation between the predicted signal and the original video. In our previous work we showed that exploiting the quantization distortion can significantly improve the accuracy of a correlation estimator. In this paper we describe how the quantization distortion can be exploited purely at the decoder side without any performance penalty when compared to an encoder-aided system. As a result, the proposed correlation estimator delivers state-of-the-art modeling accuracy while neatly fitting the low-encoder-complexity characteristic of distributed video coding. Jozef Skorupa, Jan De Cock, Jürgen Slowack, Stefaan Mys, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001 |
PCS | 8 |
| 2010 | Exploiting quantization and spatial correlation in virtual-noise modeling for distributed video coding
Jozef Skorupa, Jürgen Slowack, Stefaan Mys, Nikos Deligiannis, Jan De Cock, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle |
Signal Process. Image Commun. | 7 |
| 2010 | Rate-distortion driven decoder-side bitplane mode decision for distributed video coding
Jürgen Slowack, Stefaan Mys, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle |
Signal Process. Image Commun. | 6 |
| 2010 | Scalable Intraband and Composite Wavelet-Based Coding of Semiregular MeshesabstractThis paper proposes novel scalable mesh coding designs exploiting the intraband or composite statistical dependencies between the wavelet coefficients. A Laplacian mixture model is proposed to approximate the distribution of the wavelet coefficients. This model proves to be more accurate when compared to commonly employed single Laplacian or generalized Gaussian distribution models. Using the mixture model, we determine theoretically the optimal embedded quantizers to be used in scalable wavelet-based coding of semiregular meshes. In this sense, it is shown that the commonly employed successive approximation quantization is an acceptable, but in general, not an optimal solution. Novel scalable intraband and composite mesh coding systems are proposed, following an information-theoretic analysis of the statistical dependencies between the coefficients. The wavelet subbands are independently encoded using octree-based coding techniques. Furthermore, context-based entropy coding employing either intraband or composite models is applied. The proposed codecs provide both resolution and quality scalability. This lies in contrast to the state-of-the-art interband zerotree-based semiregular mesh coding technique, which supports only quality scalability. Additionally, the experimental results show that, on average, the proposed codecs outperform the interband state-of-the-art for both normal and nonnormal meshes. Finally, compared with a zerotree coding system, the proposed coding schemes are better suited for software/hardware parallelism, due to the independent processing of wavelet subbands. Leon Denis, Shahid M. Satti, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
IEEE Trans. Multim. | 3 |
| 2010 | Scalable L-Infinite Coding of MeshesabstractThe paper investigates the novel concept of local-error control in mesh geometry encoding. In contrast to traditional mesh-coding systems that use the mean-square error as target distortion metric, this paper proposes a new L-infinite mesh-coding approach, for which the target distortion metric is the L-infinite distortion. In this context, a novel wavelet-based L-infinite-constrained coding approach for meshes is proposed, which ensures that the maximum error between the vertex positions in the original and decoded meshes is lower than a given upper bound. Furthermore, the proposed system achieves scalability in L-infinite sense, that is, any decoding of the input stream will correspond to a perfectly predictable L-infinite distortion upper bound. An instantiation of the proposed L-infinite-coding approach is demonstrated for MESHGRID, which is a scalable 3D object encoding system, part of MPEG-4 AFX. In this context, the advantages of scalable L-infinite coding over L-2-oriented coding are experimentally demonstrated. One concludes that the proposed L-infinite mesh-coding approach guarantees an upper bound on the local error in the decoded mesh, it enables a fast real-time implementation of the rate allocation, and it preserves all the scalability features and animation capabilities of the employed scalable mesh codec. Adrian Munteanu 0001, Dan C. Cernea, Alin Alecu, Jan Cornelis 0001, Peter Schelkens |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2009 | Modeling the Correlation Noise in Spatial Domain Distributed Video CodingabstractThe paper thoroughly validates the proposed SID model on a broad dataset, showing significant accuracy improvements over SII models. Additionally, it is theoretically demonstrated that DVC systems that make SII assumptions suffer a performance penalty depending entirely on the correlation statistics of the video data. Nikos Deligiannis, Adrian Munteanu 0001, Tom Clerckx, Peter Schelkens, Jan Cornelis 0001 |
DCC | 2 |
| 2009 | On the side-information dependency of the temporal correlation in Wyner-Ziv video codingabstractCurrent models in Wyner-Ziv video coding consider the temporal correlation noise to be side-information independent (SII). This paper goes beyond this assumption and proposes a novel model, of which the parameters are side-information dependent (SID). The proposed model is experimentally validated showing remarkable accuracy improvement over the conventional SII model. Moreover, a novel SID technique for the accurate estimation of the correlation channel in video is introduced. The proposed technique enables the design of a novel pixel-domain Wyner-Ziv video coding system operating without a feedback channel. Preliminary experimental results show that the proposed codec achieves superior performance compared to the state-of-the-art in pixel-domain Wyner-Ziv coding. Nikos Deligiannis, Adrian Munteanu 0001, Tom Clerckx, Jan Cornelis 0001, Peter Schelkens |
ICASSP | 2 |
| 2009 | Context-conditioned composite coding of 3D meshes based on wavelets on surfacesabstractIn this paper, a novel wavelet-based composite mesh coding scheme is presented. In contrast to the state-of-the-art scalable interband mesh codec, the proposed codec relies on composite dependency models that capture both the interband and intraband statistical dependencies between wavelet coefficients. Additionally, each wavelet subband is processed independently, which allows for parallelized processing and for a progressive reconstruction of each mesh resolution level. Compared to the state-of-the-art, the proposed codec yields for almost all rate points superior compression performance in L2-sense. Furthermore, the generated bitstreams are near-optimal in rate-distortion sense, which eliminates the need of a post-compression rate-distortion optimization technique. Leon Denis, Shahid M. Satti, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
ICIP | 3 |
| 2009 | Overlapped Block Motion Estimation and Probabilistic Compensation with Application in Distributed Video CodingabstractIt has been recently demonstrated that, in distributed video coding (DVC) side-information dependent modeling of the correlation channel brings significant performance gains over side-information independent assumptions. In this letter, we present a novel technique enabling advanced side-information dependent estimation of the correlation channel at the decoder starting from a very coarse knowledge of it. The proposed technique triggers the design of a spatial-domain DVC codec which enables the suppression of the feedback channel. Experimental results show that the proposed codec achieves similar performance compared to the state-of-the-art in transform-domain Wyner-Ziv video coding while still operating at a fraction of its encoding complexity. Nikos Deligiannis, Adrian Munteanu 0001, Tom Clerckx, Jan Cornelis 0001, Peter Schelkens |
IEEE Signal Process. Lett. | 2 |
| 2009 | Combined Wavelet-Domain and Motion-Compensated Video Denoising Based on Video Codec Motion Estimation MethodsabstractIntegrating video coding and denoising is a novel processing paradigm, bringing mutual benefits to both video processing tools. In this paper, we propose a novel video denoising approach of which the main idea is reusing motion estimation resources from the video coding module for video denoising. In most cases, the motion fields produced by real-time video codecs cannot be directly employed in video denoising, since they, as opposed to noise filters, tolerate errors in the motion field. In order to solve this problem, we propose a novel motion-field filtering step that refines the accuracy of the motion estimates to a degree that is required for denoising. Additionally, a novel temporal filter is proposed that is robust against errors in the estimated motion field. Numerical results demonstrate that the proposed denoising scheme is of low-complexity and compares favorably to the state-of-the-art video denoising methods. Ljubomir Jovanov, Aleksandra Pizurica, Stefan Schulte 0001, Peter Schelkens, Adrian Munteanu 0001, Etienne E. Kerre, Wilfried Philips |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2008 | Applying Open-Loop Coding in Predictive Coding Systems
Adrian Munteanu 0001, Frederik Verbist, Jan Cornelis 0001, Peter Schelkens |
ACIVS | 1 |
| 2008 | Statistical L-infinite distortion estimation in scalable coding of meshesabstractThis paper investigates the novel concept of local error control in arbitrary mesh encoding, and proposes a new L-infinite mesh coding approach implementing this concept. In contrast to traditional mesh coding systems that use the mean-square error as distortion measure, the proposed approach employs the L-infinite distortion as target distortion metric. In this context, a novel wavelet-based L-infinite-constrained coding approach for meshes is proposed, which ensures that the maximum local error between the original and decoded meshes is lower than a given upper-bound. Additionally, the proposed system achieves scalability in L-infinite sense, that is, the L-infinite distortion upper-bound can be accurately estimated when decoding any layer from the input stream. Moreover, a distortion estimation approach is proposed, expressing the L-infinite distortion in the spatial domain as a statistical estimate of quantization errors produced in the wavelet domain. An instantiation of the proposed L-infinite coding approach is demonstrated for MESHGRID, which is a scalable 3D object coding system, part of MPEG-4 AFX. The proposed L-infinite coding approach guarantees that the maximum error is upper-bounded, it enables a fast real-time implementation of the rate-allocation, and it preserves all the scalability features and animation capabilities of the employed scalable mesh codec. Dan C. Cernea, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
MMSP | 2 |
| 2008 | Intra-frame video coding using an open-loop predictive coding approachabstractA novel coding approach, applying open-loop coding principles in predictive coding systems is proposed in this paper. The proposed approach is instantiated with an intra-frame video codec employing the transform and spatial prediction modes from H.264. Additionally, a novel rate-distortion model for open-loop predictive coding is proposed and experimentally validated. Optimally allocating rate based on the proposed model provides significant gains in comparison to a straightforward rate allocation not accounting for drift. Furthermore, the proposed open-loop predictive codec provides gains of up to 2.3 dB in comparison to an equivalent closed-loop intra-frame video codec employing the transform, prediction modes and rate-allocation from H.264. This indicates that, with appropriate drift compensation, open-loop predictive coding offers the possibility for further improving the compression performance in predictive coding systems. Frederik Verbist, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
MMSP | 2 |
| 2008 | Scalable Joint Source-Channel Coding for the Scalable Extension of H.264/AVCabstractThis paper proposes a novel joint source-channel coding (JSCC) methodology which minimizes the end-to-end distortion for the transmission over packet loss channels of scalable video encoded using SVC, the scalable extension of H.264/AVC. The proposed JSCC approach performs channel protection using low-density parity-check codes and relies on Lagrangian-based optimization techniques to derive the appropriate protection levels for each layer produced by the scalable source codec. Our JSCC approach for SVC can support spatial, temporal and quality scalability and can provide an optimized channel protection in any scalable setting. Experiments show that our JSCC methodology yields competitive results against state-of-the-art Lagrangian-based JSCC algorithms. Compared to the state-of-the-art, our approach significantly reduces the number of computations needed to derive the rate-distortion hulls. Moreover, the proposed approach constructs convex rate-distortion hulls for each frame, irrespective of the target rate. This allows the pre-computation of the convex rate-distortion hulls for typical packet loss channels, such that the extraction of a near-optimal JSCC allocation can be achieved on-the-fly for any target rate or packet-loss rate. We conclude that the proposed JSCC methodology provides optimized resilience against transmission errors in scalable video streaming over variable-bandwidth error-prone channels. Maryse R. Stoufs, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Scalable Joint Source and Channel Coding of MeshesabstractThis paper proposes a new approach for joint source and channel coding (JSCC) of meshes, simultaneously providing scalability and optimized resilience against transmission errors. An unequal error protection approach is followed, to cope with the different error-sensitivity levels characterizing the various resolution and quality layers produced by the input scalable source codec. The number of layers and the protection levels to be employed for each layer are determined by solving a joint source and channel coding problem. In this context, a novel fast algorithm for solving the optimization problem is conceived, enabling a real-time implementation of the JSCC rate-allocation. An instantiation of the proposed JSCC approach is demonstrated for MeshGrid, which is a scalable 3-D object representation method, part of MPEG-4 AFX. In this context, the L-inflnite distortion metric is employed, which is to our knowledge a unique feature in mesh coding. Numerical results show the superiority of the L-inflnite norm over the classical L-2 norm in a JSCC setting. One concludes that the proposed joint source and channel coding approach offers resilience against transmission errors, provides graceful degradation, enables a fast real-time implementation, and preserves all the scalability features and animation capabilities of the employed scalable mesh codec. Dan C. Cernea, Adrian Munteanu 0001, Alin Alecu, Jan Cornelis 0001, Peter Schelkens |
IEEE Trans. Multim. | 2 |
| 2007 | On Hybrid Directional Transform-Based Intra-band Image Coding
Alin Alecu, Adrian Munteanu 0001, Aleksandra Pizurica, Jan Cornelis 0001, Peter Schelkens |
ACIVS | 2 |
| 2007 | Analysis of the Statistical Dependencies in the Curvelet Domain and Applications in Image Compression
Alin Alecu, Adrian Munteanu 0001, Aleksandra Pizurica, Jan Cornelis 0001, Peter Schelkens |
ACIVS | 2 |
| 2007 | Distributed Video Coding with Shared Encoder/Decoder ComplexityabstractDistributed video coding is a coding paradigm that allows complexity to be shared between encoder and decoder. In this context, video coding systems have been developed with encoder complexities similar to H.263+ intra-coding, while obtaining compression performance comparable to H.263+ inter coding. The decoders in these systems typically employ motion-compensated frame interpolation or extrapolation to generate side-information. However, as motion complexity of the video sequence increases, such generators fail to provide reliable side information. This paper proposes a pixel-domain distributed video coding method, combining low-complexity encoder-side bitplane motion estimation with decoder-side motion-compensated frame interpolation. It is shown that such a system is more suitable for sequences with increased motion complexity, compared to codecs that employ motion estimation at the decoder only. Tom Clerckx, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
ICIP (6) | 2 |
| 2007 | Segmentation-Driven Direction-Adaptive Discrete Wavelet TransformabstractThis paper proposes a novel segmentation-driven direction-adaptive discrete wavelet transform (SD DADWT), wherein the adaptation of the directional wavelet bases is performed on the segments describing the natural geometry of the image. First, a multi-resolution segmentation of the image is performed, obtained through an Edgmentation procedure. The optimum lifting directions are then selected for each segment and at each resolution. The proposed SD DADWT retains the inherent advantages offered by a multiresolution representation of the geometric features in the image, and in the same time provides a sparse image representation via DADWT. Preliminary experimental results obtained in a coding application show that the visual quality of the reconstructed image can be further improved by applying a geometrically-oriented transform on segments that approximate the natural borders in the image. Adrian Munteanu 0001, Oana Maria Surdu, Jan Cornelis 0001, Peter Schelkens |
ICIP (1) | 1 |
| 2007 | Optimal Joint Source-Channel Coding using Unequal Error Protection for the Scalable Extension of H.264/MPEG-4 AVCabstractThis paper proposes an optimized joint source-channel coding methodology with unequal error protection for the transmission of video encoded with the recently developed scalable extension of H.264/MPEG-4 AVC. The proposed methodology uses a simplified Viterbi-based search method which significantly outperforms the classical exhaustive search method in terms of computational complexity, leading to a practically applicable solution at the expense of a minimal loss of optimality. Experimental results show the effectiveness of our protection methodology and illustrate its capability to provide graceful degradation in the presence of channel mismatches. Maryse R. Stoufs, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ICIP (6) | 2 |
| 2007 | Joint Source-Channel Coding for the Scalable Extension of H.264/MPEG-4 AVCabstractIn this paper, we propose a joint source-channel coding (JSCC) methodology which minimizes the end-to-end distortion for the transmission of H.264/MPEG-4 scalable video over packet loss channels. The proposed JSCC-approach employs low-density parity-check codes in order to provide channel protection and relies on Lagrangian-based optimization techniques to derive the appropriate protection levels for each layer produced by the scalable source codec. Experiments show that our JSCC methodology delivers competitive results to state-of-the-art Lagrangian-based algorithms. However, in contrast to the state-of-the-art, our approach significantly reduces the computational complexity. We conclude that the proposed JSCC methodology provides optimized resilience against transmission errors in scalable video streaming over variable-bandwidth error-prone channels. Maryse R. Stoufs, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
MMSP | 2 |
| 2006 | Complexity Scalability in Motion-Compensated Wavelet-Based Video Coding
Tom Clerckx, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
ACIVS | 2 |
| 2006 | Performing Deblocking in Video Coding Based on Spatial-Domain Motion-Compensated Temporal Filtering
Adrian Munteanu 0001, Joeri Barbarien, Jan Cornelis 0001, Peter Schelkens |
ACIVS | 1 |
| 2006 | Scalable and Channel-Adaptive Unequal Error Protection of Images with LDPC Codes
Adrian Munteanu 0001, Maryse R. Stoufs, Jan Cornelis 0001, Peter Schelkens |
ACIVS | 1 |
| 2006 | Information-Theoretic Analysis of Dependencies Between Curvelet CoefficientsabstractThis paper reports an information-theoretic analysis of the inter-scale, inter-orientation and inter-location dependencies that exist between curvelet coefficients. We show that the marginal statistics of these coefficients can be accurately modeled using generalized Gaussian density functions. Though generally decorrelated, we find that curvelets exhibit unusually high dependencies in intra-band local micro-neighborhoods, of a magnitude not found for instance in classical wavelets. Furthermore, dependencies are subject to and decrease with increasing orientation and location differences. Finally, we conclude that intra-band coefficient dependencies are stronger than either their inter-scale or inter-direction counterparts. Alin Alecu, Adrian Munteanu 0001, Aleksandra Pizurica, Wilfried Philips, Jan Cornelis 0001, Peter Schelkens |
ICIP | 2 |
| 2006 | JPEG2000. Part 10. Volumetric data encodingabstractThe joint photographic experts group (JPEG) committee (ISO/IEC JTC1/SC29/WG1) is currently pursuing the standardization of a three-dimensional extension of the JPEG-2000 standard (Parts 1 and 2) to support the encoding of volumetric data sets. This extension, Part 10 - extensions for three-dimensional data (JP3D), will support functionalities like resolution scalability, quality scalability and region-of-interest coding, while exploiting the entropy in the additional third dimension to improve the rate-distortion performance. In this paper, we give an overview of the markets and application areas targeted by JP3D, the imposed requirements and the algorithm under study Peter Schelkens, Adrian Munteanu 0001, Alexis Tzannes, Christopher M. Brislawn |
ISCAS | 2 |
| 2006 | Wavelet-based scalable L-infinity-oriented compressionabstractAmong the different classes of coding techniques proposed in literature, predictive schemes have proven their outstanding performance in near-lossless compression. However, these schemes are incapable of providing embedded L(infinity)-oriented compression, or, at most, provide a very limited number of potential L(infinity) bit-stream truncation points. We propose a new multidimensional wavelet-based L(infinity)-constrained scalable coding framework that generates a fully embedded L(infinity)-oriented bit stream and that retains the coding performance and all the scalability options of state-of-the-art L2-oriented wavelet codecs. Moreover, our codec instantiation of the proposed framework clearly outperforms JPEG2000 in L(infinity) coding sense. Alin Alecu, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
IEEE Trans. Image Process. | 2 |
| 2006 | Embedded Multiple Description Coding of VideoabstractReal-time delivery of video over best-effort error-prone packet networks requires scalable erasure-resilient compression systems in order to 1) meet the users' requirements in terms of quality, resolution, and frame-rate; 2) dynamically adapt the rate to the available channel capacity; and 3) provide robustness to data losses, as retransmission is often impractical. Furthermore, the employed erasure-resilience mechanisms should be scalable in order to adapt the degree of resiliency against transmission errors to the varying channel conditions. Driven by these constraints, we propose in this paper a novel design for scalable erasure-resilient video coding that couples the compression efficiency of the open-loop architecture with the robustness provided by multiple description coding. In our approach, scalability and packet-erasure resilience are jointly provided via embedded multiple description scalar quantization. Furthermore, a novel channel-aware rate-allocation technique is proposed that allows for shaping on-the-fly the output bit rate and the degree of resiliency without resorting to channel coding. As a result, robustness to data losses is traded for better visual quality when transmission occurs over reliable channels, while erasure resilience is introduced when noisy links are involved. Numerical results clearly demonstrate the advantages of the proposed approach over equivalent codec instantiations employing 1) no erasure-resilience mechanisms, 2) erasure-resilience with nonscalable redundancy, or 3) data-partitioning principles. Fabio Verdicchio, Adrian Munteanu 0001, Augustin Gavrilescu, Jan Cornelis 0001, Peter Schelkens |
IEEE Trans. Image Process. | 2 |
| 2005 | Robust Motion Vector Coding and Error Concealment in MCTF-Based Video CodingabstractError resilience is of paramount importance in video transmission over variable-bandwidth error-prone channels, such as wireless channels. In this paper, we investigate the influence of corrupted motion vectors in video coding based on motion compensated temporal filtering, and develop various error resilience and concealment mechanisms for this class of codecs. The experimental results show that our proposed motion vector coding technique significantly increases the robustness against transmission errors at the cost of less than 3% in terms of rate. It is also shown that our proposed spatial error-concealment mechanism leads to performance gains of up to 6 dB in comparison to a classical slicing-based approach employing no error concealment. Maryse R. Stoufs, Joeri Barbarien, Peter Schelkens, Jan Cornelis 0001, Adrian Munteanu 0001 |
ICASSP (2) | 5 |
| 2005 | Error-resilient video coding using motion compensated temporal filtering and embedded multiple description scalar quantizersabstractReal time delivery of video over best-effort networks requires compression systems that dynamically adapt the rate to the available channel capacity and exhibit robustness to data loss as retransmission is often impractical. Error-resilience, however, significantly lowers the coding performance when rigid design is performed based on a worst-case scenario. This paper presents an original scalable video coding scheme that couples the compression efficiency of the open-loop architecture with the robustness of multiple description source coding. The use of embedded multiple description quantization and a novel channel-aware rate-allocation allow for shaping on-the-fly the output bit-rate and the degree of resilience. As a result, robustness to data losses is traded for better visual quality when transmission occurs over reliable channels, while error-resilience is introduced when noisy links are involved. The advantage of our proposal is demonstrated in the context of packet-lossy networks. Fabio Verdicchio, Adrian Munteanu 0001, Augustin Gavrilescu, Jan Cornelis 0001, Peter Schelkens |
ICIP (3) | 2 |
| 2005 | Single-rate calculation of overcomplete discrete wavelet transforms for scalable coding applications
Yiannis Andreopoulos, Adrian Munteanu 0001, Geert Van der Auwera, Jan Cornelis 0001, Peter Schelkens |
Signal Process. | 2 |
| 2005 | Motion and texture rate-allocation for prediction-based scalable motion-vector coding
Joeri Barbarien, Adrian Munteanu 0001, Fabio Verdicchio, Yiannis Andreopoulos, Jan Cornelis 0001, Peter Schelkens |
Signal Process. Image Commun. | 2 |
| 2005 | Unconstrained motion compensated temporal filtering (UMCTF) for efficient and flexible interframe wavelet video coding
Deepak S. Turaga, Mihaela van der Schaar, Yiannis Andreopoulos, Adrian Munteanu 0001, Peter Schelkens |
Signal Process. Image Commun. | 4 |
| 2004 | Scalable motion vector codingabstractRecently proposed scalable wavelet-based video codecs using spatial-domain motion compensated temporal filtering (SDMCTF) offer competitive compression performance when compared to H.264 and generate embedded bit-streams supporting quality, resolution and temporal scalability. To be able to support a large range of bit-rates with optimal compression efficiency, these codecs require a quality-scalable motion vector coding technique. Such an algorithm based on the integer wavelet transform followed by embedded coding of the wavelet coefficients was proposed in the recent past. In this paper, we present a quality-scalable motion vector coding algorithm using median-based motion vector prediction. The compression performance of the proposed algorithm is compared to that of the wavelet-based technique and is found to be superior. Additionally, the proposed motion vector codec is incorporated into an SDMCTF-based video codec and the benefits of using quality-scalable motion vector representations are experimentally demonstrated. Joeri Barbarien, Adrian Munteanu 0001, Fabio Verdicchio, Yiannis Andreopoulos, Jan Cornelis 0001, Peter Schelkens |
ICIP | 2 |
| 2004 | A new family of embedded multiple description scalar quantizersabstractA new family of embedded multiple description scalar quantizers (EMDSQ) that support the progressive transmission of images over variable-bandwidth error-prone channels is proposed in this paper. A control mechanism that allows for tuning the redundancy between the two descriptions for each quantization level is also designed. The employed mechanism enables control of the tradeoff between the coding efficiency and error-resilience, and provides an increased robustness by improving the error resilience in the most important layers of the embedded bit-streams. Instantiations of the proposed family are incorporated in a wavelet-based embedded coding system, and the redundancy-control mechanism is practically demonstrated. Experimental results show that the proposed EMDSQ outperform the state-of-the-art multiple description uniform scalar quantizers (MDUSQ) previously proposed in the literature for error-resilient progressive image transmission. Augustin Gavrilescu, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
ICIP | 2 |
| 2004 | Scalable video coding based on motion-compensated temporal filtering: complexity and functionality analysisabstractVideo coding techniques yielding state-of-the-art compression performance require large amount of computational resources, hence practical implementations, which target a broad market, often tend to trade-off coding efficiency and flexibility for reduced complexity. Scalable video coding instead, not only provides seamless adaptation to bit-rate variation, but also allows the end user to trim down the resources he needs to perform real-time decoding by limiting the process to a subset of the original content. Hence, by choosing the quality, frame-rate and/or resolution of the reconstructed sequence, each decoder can meet its hardware limitations without affecting the encoding process of the media provider. This paper proposes a preliminary analysis of the memory-access behavior of a fully scalable video decoder and investigates the capability of selecting the operational settings in order to adapt to the available hardware resources on the target device. Fabio Verdicchio, Yiannis Andreopoulos, Tom Clerckx, Joeri Barbarien, Adrian Munteanu 0001, Jan Cornelis 0001, Peter Schelkens |
ICIP | 5 |
| 2004 | In-band motion compensated temporal filtering
Yiannis Andreopoulos, Adrian Munteanu 0001, Joeri Barbarien, Mihaela van der Schaar, Jan Cornelis 0001, Peter Schelkens |
Signal Process. Image Commun. | 2 |
| 2004 | On the optimality of embedded deadzone scalar-quantizers for wavelet-based L-infinite-constrained image codingabstractIn wavelet-based L/sub /spl infin//-constrained embedded coding, the bit-stream is truncated at the bit-rate that corresponds to a guaranteed, user-defined distortion bound. The letter analyzes the optimality of embedded deadzone scalar-quantizers for high-rate L/sub /spl infin//-constrained scalable wavelet-based image coding. A rate-distortion model applicable to the family of embedded deadzone scalar-quantizers is derived and experimentally validated. Conclusions are drawn regarding the optimal subband-quantizer instantiations. The optimal quantizers are employed in a coding algorithm that retains the coding performance and the flexibility options of wavelet-based codecs while allowing for a fully embedded L/sub /spl infin//-oriented bit-stream. Alin Alecu, Adrian Munteanu 0001, Jan Cornelis 0001, Steven Dewitte, Peter Schelkens |
IEEE Signal Process. Lett. | 2 |
| 2004 | MESHGRID-a compact, multiscalable and animation-friendly surface representationabstractMESHGRID is a novel, compact, multiscalable and animation-friendly surface representation method, which has been introduced in MPEG-4 . The MESHGRID representation attaches a description of the "global connectivity" between the vertices on the object's surface (i.e., the 3-D connectivity wireframe) to a regular 3-D grid of points (i.e., the reference grid). MESHGRID efficiently encodes the 3-D connectivity wireframe by using a new type of 3-D extension of Freeman chain-code. MESHGRID does not explicitly store the polygons of the surface, since the 3-D connectivity wireframe has particular connectivity properties allowing for the unambiguous derivation of the triangulation. The reference grid is a smooth vector field defined on a regular discrete 3-D space. This grid is efficiently compressed by using an embedded 3-D wavelet-based multiresolution intra-band coding algorithm. MESHGRID can be efficiently exploited for QoS since it allows for three types of scalability in both view-dependent and view-independent scenarios, including: 1) resolution scalability, i.e., the adaptation of the number of transmitted vertices; 2) shape precision, i.e., the adaptive reconstruction of the reference grid positions; and 3) vertex position scalability, i.e., the change of the precision of known vertex positions with respect to the reference grid. Furthermore, in addition to the classical vertex-based animation, MESHGRID also supports specific animation capabilities, such as: 1) rippling effects by changing the position of the vertices relative to corresponding reference grid points and 2) reshaping on a hierarchical basis of the regular reference grid and its attached vertices. Ioan Alexandru Salomie, Adrian Munteanu 0001, Augustin Gavrilescu, Gauthier Lafruit, Peter Schelkens, Rudi Deklerck, Jan Cornelis 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | On the Optimality of Embedded Deadzone Scalar-Quantizers for Wavelet-based L-infinite-constrained Image CodingabstractSummary form only given. Several methods for L/sub /spl infin//-distortion constrained compression have been proposed that target a set of fixed reconstruction-error bounds. Recently, a wavelet-based L/sub /spl infin//-constrained embedded image-coding technique was proposed that guarantees the required distortion bound while retaining the coding performance and scalability option state-of-the-art wavelet-based L/sub /spl infin//-oriented codecs. The optimality of embedded deadzone scalar-quantizers for high-rate L/sub /spl infin//-constrained scalable wavelet-based coding of images was analyzed. The optimal quantizers were employed in a coding algorithm that outperforms embedded L/sub /spl infin//-oriented wavelet-coders in terms of maximum absolute error. The algorithm was briefly described and coding results were provided. Alin Alecu, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001, Steven Dewitte |
DCC | 2 |
| 2003 | Fully-scalable wavelet video coding using in-band motion compensated temporal filteringabstractThis paper presents a novel fully-scalable wavelet video coding scheme that performs efficient open-loop motion compensated temporal filtering (MCTF) in the wavelet domain (in-band). Unlike the conventional spatial-domain MCTF (SDMCTF) schemes, which apply MCTF on the original image data and then encode the residual image using a critically-sampled wavelet transform, the framework presented here applies the in-band MCTF (IBMCTF) after the discrete wavelet transform (DWT) is performed in the spatial dimensions. To overcome the inefficiency of motion estimation (ME) in the wavelet domain, a complete-to-overcomplete DWT (CODWT) is performed. The proposed framework provides improved quality (SNR) and temporal scalability as compared with existing in-band closed-loop temporal prediction schemes with ODWT and improved spatial scalability as compared to SDMCTF. We present a thorough comparison between SDMCTF and the proposed IBMCTF in terms of coding efficiency and scalability. Furthermore, we describe several extensions that enable the filtering of the various bands to be performed independently, based on the resolution, sequence content, complexity requirements and desired scalability. Yiannis Andreopoulos, Mihaela van der Schaar, Adrian Munteanu 0001, Joeri Barbarien, Peter Schelkens, Jan Cornelis 0001 |
ICASSP (3) | 3 |
| 2003 | Embedded multiple description scalar quantizers for progressive image transmissionabstractRobust progressive image transmission over unreliable channels with variable bandwidth requires multiple description coding (MDC) systems that produce highly error-resilient embedded bitstreams. The proposed embedded multiple description scalar quantizers (EMDSQ) meet the desired features consisting of a high redundancy level, fine grain rate adaptation and progressive transmission of each description. Experimental results show that EMDSQ yield better rate-distortion performance in comparison to the multiple description uniform scalar quantizers; (MDUSQ) previously proposed in the literature. Moreover, the generalized form of EMDSQ targeting an arbitrary number of channels is proposed, which offers the possibility of designing realistic coders for practical multi-channel communication systems. Augustin Gavrilescu, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ICASSP (5) | 2 |
| 2003 | Spatio-temporal-SNR scalable wavelet coding with motion-compensated DCT base-layer architecturesabstractIt has been demonstrated recently that 3-D wavelet coding with motion-compensated temporal filtering (MCTF) provides a wide range of spatio-temporal-SNR scalability with state-of-the-art coding performance. However, the coding system is very different from the already standardized motion-compensated DCT (MC-DCT) video coders such as MPEG-2, MPEG-4, or H.26L. Nonetheless, the market acceptance of the new scalable technology will come much easier if backward compatibility to such previous standards would exist. In this paper, we present a new coder architecture where the base layer can be a standard MC-DCT coder, while the enhancement layer is an in-band MCTF codec operating in the overcomplete wavelet domain. First, we propose a simple extension of the scalable 3-D wavelet codec with a standard MC-DCT base layer. As a second step, to improve the performance of the proposed scalable coder over a wide range of bit-rates, we describe several small extensions to the standardized MC-DCT codec structure that can be applied for the base layer coding. Yiannis Andreopoulos, Mihaela van der Schaar, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ICIP (2) | 3 |
| 2003 | Motion vector coding for in-band motion compensated temporal filteringabstractRecently, a new wavelet-based video codec using in-band motion compensated temporal filtering (IBMCTF) was introduced This codec is fully scalable in resolution, quality aid frame-rate. In comparison to an equivalent video coding scheme based on spatial domain motion compensated temporal filtering (SDMCTF), its compression performance when decoding to lower resolutions is very promising. However, since the IBMCTF scheme is based on in-band motion estimation, considerably more motion vector data is generated than in the SDMCTF scheme. Efficient compression of these motion vectors is therefore of utmost importance. In this paper, several solutions for the compression of motion vectors generated by a video codec based on IBMCTF are presented and compared. Joeri Barbarien, Yiannis Andreopoulos, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ICIP (2) | 3 |
| 2003 | Control of the distortion variation in video coding systems based on motion compensated temporal filteringabstractThe paper proposes a new framework for the control of the distortion variation in video coding schemes based on motion-compensated temporal filtering (MCTF). The distortion in an arbitrary decoded frame at any temporal level in the MCTF pyramid is expressed as a function of the distortions in the reference frames at the same temporal level. The approach is formulated for the bi-directional unconstrained MCTF (UMCTF) scheme of Turaga et al., (2002), which does not include the update-lifting step. The proposed framework can be extended to the generalized form of MCTF by utilizing additional control parameters. Experimental results demonstrate the control of the distortion variation in video coding systems based on spatial-domain and wavelet-domain MCTF. One concludes that the proposed framework provides the means of controlling the tradeoff between the average distortion and the distortion variation in each group-of-pictures (GOPs) within the decoded sequence. Adrian Munteanu 0001, Yiannis Andreopoulos, Mihaela van der Schaar, Peter Schelkens, Jan Cornelis 0001 |
ICIP (2) | 1 |
| 2003 | Embedded multiple description scalar quantizers for progressive image transmissionabstractRobust progressive image transmission over unreliable channels with variable bandwidth requires multiple description coding (MDC) systems that produce highly error-resilient embedded bit-streams. The proposed embedded multiple description scalar quantizers (EMDSQ) meet the desired features consisting of a high redundancy level, fine grain rate adaptation and progressive transmission of each description. Experimental results show that EMDSQ yield better rate-distortion performance in comparison to the multiple description uniform scalar quantizers (MDUSQ) previously proposed in the literature. Moreover, the generalized form of EMDSQ targeting an arbitrary number of channels is proposed, which offers the possibility of designing realistic coders for practical multi-channel communication systems. Augustin Gavrilescu, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ICME | 2 |
| 2003 | Complete-to-overcomplete discrete wavelet transforms for scalable video coding with MCTF
Yiannis Andreopoulos, Mihaela van der Schaar, Adrian Munteanu 0001, Joeri Barbarien, Peter Schelkens, Jan Cornelis 0001 |
VCIP | 3 |
| 2003 | Wavelet Coding of Volumetric Medical DatasetsabstractSeveral techniques based on the three-dimensional (3-D) discrete cosine transform (DCT) have been proposed for volumetric data coding. These techniques fail to provide lossless coding coupled with quality and resolution scalability, which is a significant drawback for medical applications. This paper gives an overview of several state-of-the-art 3-D wavelet coders that do meet these requirements and proposes new compression methods exploiting the quadtree and block-based coding concepts, layered zero-coding principles, and context-based arithmetic coding. Additionally, a new 3-D DCT-based coding scheme is designed and used for benchmarking. The proposed wavelet-based coding algorithms produce embedded data streams that can be decoded up to the lossless level and support the desired set of functionality constraints. Moreover, objective and subjective quality evaluation on various medical volumetric datasets shows that the proposed algorithms provide competitive lossy and lossless compression results when compared with the state-of-the-art. Peter Schelkens, Adrian Munteanu 0001, Joeri Barbarien, Mihnea Galca, Xavier Giró-i-Nieto, Jan Cornelis 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2002 | Scalable wavelet video-coding with in-band prediction - implementation and experimental resultsabstractIn this paper we elaborate on a recently proposed approach for scalable video-coding based on in-band prediction in the overcomplete wavelet domain. It is shown that, through a new calculation scheme for the level-by-level complete to overcomplete discrete wavelet transform (DWT) that exploits certain symmetries, important reductions in the multiplication budget are obtained in comparison to the fastest-known algorithm of the literature. Based on the derived overcomplete transform-domain coefficients, a pixel-accurate motion estimation and compensation (ME/MC) algorithm is proposed, which provides a hybrid-coding framework that supports full scalability in resolution, quality and frame rate. To give an indication of the coding performance of such a system, some preliminary results are reported. Yiannis Andreopoulos, Geert Van der Auwera, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ICIP (3) | 3 |
| 2002 | Scalable wavelet video-coding with in-band prediction - the bottom-up overcomplete discrete wavelet transformabstractA new in-band motion compensation algorithm for wavelet-based video coding is proposed: the bottom-up prediction algorithm (BUP). The BUP algorithm overcomes the periodic shift-invariance of the discrete wavelet transform (DWT) and is formalized into new prediction rules using filtering operations. BUP is based on the relationships between the subbands of the shifted input signal and the subbands of the non-shifted reference signal, whereby the number of shifts is limited by the periodic shift-invariance of the DWT. We derive the algorithm for the 1-D DWT with three decomposition levels. The combination of all prediction rules of the BUP algorithm defines a new transform: the bottom-up overcomplete DWT or BUP ODWT, which is shift-invariant. The BUP ODWT calculates the overcomplete subbands by applying the prediction rules to the critically sampled subbands of a wavelet-transformed image. The envisaged application for the BUP algorithm is spatially scalable video coding. Geert Van der Auwera, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ICIP (3) | 2 |
| 2002 | Wavelet coding of volumetric medical datasetsabstractThis paper assesses the coding performance of several state-of-the-art 3D wavelet coders and compares them with newly proposed compression methods exploiting quadtree and block-based coding concepts, layered zero-coding principles and context-based arithmetic coding. The discussed wavelet-based coding algorithms produce losslessly compressed embedded data streams, supporting the desired progressive data transmission functionality. Moreover, objective and subjective quality evaluations on various medical volumetric datasets show that the newly proposed algorithms provide competitive lossy and lossless compression results compared to the state-of-the-art. Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001 |
ICIP (3) | 1 |
| 2002 | MESHGRID - a compact, multi-scalable and animation-friendly surface representationabstractMESHGRID is a novel, compact, multi-scalable and animation-friendly surface representation method, which has been introduced in MPEG-4. The MESHGRID representation attaches a description of the "global connectivity" between the vertices on the object's surface (i.e. the 3D connectivity wireframe) to a regular 3D grid of points (i.e. the reference-grid). The 3D connectivity wireframe is efficiently encoded by using a new type of 3D extension of Freeman chain-code. MESHGRID does not explicitly store the polygons of the surface, since the 3D connectivity wireframe has particular connectivity properties allowing for the unambiguous derivation of the triangulation. The reference-grid is a smooth vector field defined on a regular discrete 3D space. This grid is efficiently compressed by using an embedded 3D wavelet-based multi-resolution intra-band coding algorithm. MESHGRID allows for three types of scalability in both view-dependent and view-independent scenarios: resolution scalability, shape precision, and vertex position scalability. Furthermore, in addition to the classical vertex-based animation, MESHGRID supports specific animation capabilities, such as rippling effects and reshaping on a hierarchical basis of the regular reference-grid and its attached vertices. Ioan Alexandru Salomie, Adrian Munteanu 0001, Augustin Gavrilescu, Gauthier Lafruit, Peter Schelkens, Rudi Deklerck, Jan Cornelis 0001 |
ICIP (3) | 2 |
| 2000 | Evaluation of a Quincunx Wavelet Filter Design Approach for Quadtree-Based Embedded Image CodingabstractWe study the compression performance of the quincunx discrete wavelet transform (DWT) and we compare it with the dyadic DWT in terms of rate-distortion. The 2D non-separable quincunx wavelet filters are designed by making use of the transformations of variables technique of Tay and Kingsbury (1993) starting from the 1D biorthogonal (9,7)-taps filters. The applied transformation functions are respectively based on the Kaiser and Chebyshev windows, and on the Lagrange halfband filters. The novelties of this work are in the evaluation of these filters for image compression and the usage of an embedded coder based on the quadtree approach for coding the quincunx subbands. Geert Van der Auwera, Adrian Munteanu 0001, Jan Cornelis 0001 |
ICIP | 2 |
| 1999 | Wavelet-based compression of medical images: Protocols to improve resolution and quality scalability and region-of-interest coding
Peter Schelkens, Adrian Munteanu 0001, Jan Cornelis 0001 |
Future Gener. Comput. Syst. | 2 |
| 1999 | Wavelet image compression - the quadtree coding approachabstractPerfect reconstruction, quality scalability, and region-of-interest coding are basic features needed for the image compression schemes used in telemedicine applications. This paper proposes a new wavelet-based embedded compression technique that efficiently exploits the intraband dependencies and uses a quadtree-based approach to encode the significance maps. The algorithm produces a losslessly compressed embedded data stream, supports quality scalability, and permits region-of-interest coding. Moreover, experimental results obtained on various images show that the proposed algorithm provides competitive lossless/lossy compression results. The proposed technique is well suited for telemedicine applications that require fast interactive handling of large image sets, over networks with limited and/or variable bandwidth. Adrian Munteanu 0001, Jan Cornelis 0001, Geert Van der Auwera, Paul Dan Cristea |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 1999 | Wavelet Based Lossless Compression of Coronary Angiography ImagesabstractThe final diagnosis in coronary angiography has to be performed on a large set of original images. Therefore, lossless compression schemes play a key role in medical database management and telediagnosis applications. This paper proposes a wavelet-based compression scheme that is able to operate in the lossless mode. The quantization module implements a new way of coding of the wavelet coefficients that is more effective than the classical zerotree coding. The experimental results obtained on a set of 20 angiograms show that the algorithm outperforms the embedded zerotree coder, combined with the integer wavelet transform, by 0.38 bpp, the set partitioning coder by 0.21 bpp, and the lossless JPEG coder by 0.71 bpp. The scheme is a good candidate for radiological applications such as teleradiology and picture archiving and communications systems (PACS's). Adrian Munteanu 0001, Jan Cornelis 0001, Paul Dan Cristea |
IEEE Trans. Medical Imaging | 1 |
| 1998 | Video coding based on motion estimation in the wavelet detail imagesabstractThis work proposes a new block based motion estimation and compensation technique applied on the detail images of the wavelet pyramidal decomposition. The algorithm uses two matching criteria, namely the absolute difference and the absolute sum. For a wavelet decomposed one-dimensional step function, it is shown that for odd translations of the step, the absolute sum reaches a smaller minimum than the absolute difference. We also derive in this case a constraint on the highpass filter coefficients so that a zero prediction error can be reached by using the absolute sum. Although this cannot be easily generalized for an arbitrary signal profile, experimental results obtained with photorealistic image sequences indicate that the prediction error can be reduced with respect to techniques that only use the absolute difference as matching criterion. Geert Van der Auwera, Adrian Munteanu 0001, Gauthier Lafruit, Jan Cornelis 0001 |
ICASSP | 2 |
| 1998 | Progressive lossless coding of medical images
Paul Dan Cristea, Jan Cornelis 0001, Adrian Munteanu 0001 |
Future Gener. Comput. Syst. | 3 |