EDBT 2026 Demo / reviewers in the wild / expert
Joël Jung
dblp:84/579 · also Joel Jung
· DBLP profile ↗
40ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-3878-6454ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GeodesicPSIM: Predicting the Quality of Static Mesh With Texture Map via Geodesic Patch SimilarityabstractStatic meshes with texture maps have attracted considerable attention in both industrial manufacturing and academic research, leading to an urgent requirement for effective and robust objective quality evaluation. However, current model-based static mesh quality metrics (i.e., metrics that directly use the raw data of the static mesh to extract features and predict the quality) have obvious limitations: most of them only consider geometry information, while color information is ignored, and they have strict constraints for the meshes' geometrical topology. Other metrics, such as image-based and point-based metrics, are easily influenced by the prepossessing algorithms, e.g., projection and sampling, hampering their ability to perform at their best. In this paper, we propose Geodesic Patch Similarity (GeodesicPSIM), a novel model-based metric to accurately predict human perception quality for static meshes. After selecting a group keypoints, 1-hop geodesic patches are constructed based on both the reference and distorted meshes cleaned by an effective mesh cleaning algorithm. A two-step patch cropping algorithm and a patch texture mapping module refine the size of 1-hop geodesic patches and build the relationship between the mesh geometry and color information, resulting in the generation of 1-hop textured geodesic patches. Three types of features are extracted to quantify the distortion: patch color smoothness, patch discrete mean curvature, and patch pixel color average and variance. To the best of our knowledge, GeodesicPSIM is the first model-based metric especially designed for static meshes with texture maps. GeodesicPSIM provides state-of-the-art performance in comparison with image-based, point-based, and video-based metrics on a newly created and challenging database. We also prove the robustness of GeodesicPSIM by introducing different settings of hyperparameters. Ablation studies also exhibit the effectiveness of three proposed features and the patch cropping algorithm. The code is available at https://multimedia.tencent.com/resources/GeodesicPSIM. Qi Yang 0003, Joël Jung, Xiaozhong Xu, Shan Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | TDMD: A Database for Dynamic Color Mesh Quality Assessment StudyabstractDynamic colored meshes (DCM) are widely used in various applications. However, this kind of meshes may undergo different processes, such as compression or transmission, which can distort them and degrade their quality. To facilitate the development of objective metrics for DCMs and study the influence of typical distortions on their perception, we create the Tencent - Dynamic colored Mesh Database (TDMD) containing eight reference DCM objects with six typical distortions. Using processed video sequences (PVS) derived from the DCM, we conduct a large-scale subjective experiment that resulted in 303 distorted DCM samples with mean opinion scores, making the TDMD the largest available DCM database to our knowledge. This database enables us to study the impact of different types of distortion on human perception and offers recommendations for DCM compression and related tasks. Additionally, we have evaluated three types of state-of-the-art objective metrics on the TDMD, including image-based, point-based, and video-based metrics, on the TDMD. Our experimental results highlight the strengths and weaknesses of each metric, and we provide suggestions about the selection of metrics in practical DCM applications. Qi Yang 0003, Joël Jung, Timon Deschamps, Xiaozhong Xu, Shan Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Exploring the Influence of View and Camera Path Selection for Dynamic Mesh Quality AssessmentabstractWith the development of 3D mesh processing and applications, the quality assessment of dynamic mesh sequences attracts more and more attention. One prevalent strategy for performing the subjective experiment and designing objective quality metrics is to convert the 3D dynamic mesh into 2D images or videos via projection and to collect subjective scores or calculate objective indexes based on these images or videos. In this paper, we study the influence of the view, or camera path selection for the projection, for both subjective and objective dynamic mesh quality assessment, and compare the performance of image-based metrics and point-based metrics corresponding to the collected subjective scores. First, we use the dynamic mesh sequences proposed by MPEG as anchors and generate videos corresponding to different coding configurations and different camera paths. Then, we conduct subjective experiments to collect the ground truth of mean opinion scores. Besides, we calculate the state-of-the-art objective metric scores for each sequence. We analyze the differences between subjective scores with respect to different camera paths and the correlation between subjective scores and objective metrics. The results show that different camera paths tend to generate close subjective perceptions and that the selection of views can influence some objective metrics. Kaifa Yang, Qi Yang 0003, Joël Jung, Yiling Xu, Xiaozhong Xu, Shan Liu 0001 |
ICME | 3 |
| 2023 | TSMD: A Database for Static Color Mesh Quality Assessment StudyabstractStatic meshes with texture map are widely used in modern industrial and manufacturing sectors, attracting considerable attention in the mesh compression community due to its huge amount of data. To facilitate the study of static mesh compression algorithm and objective quality metric, we create the Tencent – Static Mesh Dataset (TSMD) containing 42 reference meshes with rich visual characteristics. 210 distorted samples are generated by the lossy compression scheme developed for the Call for Proposals on polygonal static mesh coding, released on June 23 by the Alliance for Open Media Volumetric Visual Media group. Using processed video sequences, a large-scale, crowdsourcing-based, subjective experiment was conducted to collect subjective scores from 74 viewers. The dataset undergoes analysis to validate its sample diversity and Mean Opinion Scores (MOS) accuracy, establishing its heterogeneous nature and reliability. State-of-the-art objective metrics are evaluated on the new dataset. Pearson and Spearman correlations around 0.75 are reported, deviating from results typically observed on less heterogeneous datasets, demonstrating the need for further development of more robust metrics. The TSMD, including meshes, PVSs, bitstreams, and MOS, is made publicly available at the following location: https://multimedia.tencent.com/resources/tsmd. Qi Yang 0003, Joël Jung, Haiqiang Wang, Xiaozhong Xu, Shan Liu 0001 |
VCIP | 2 |
| 2022 | Towards Joint Frame-Level and MOS Quality Predictions with Low-Complexity Objective ModelsabstractThe evaluation of the quality of gaming content, with low-complexity and low-delay approaches is a major challenge raised by the emerging gaming video streaming and cloud-gaming services. Considering two existing and a newly created gaming databases this paper confirms that some low-complexity metrics match well with subjective scores when considering usual correlation indicators. It is however argued such a result is insufficient: gaming content suffers from sudden large quality drops that these indicators do not capture. In addition to proposing three new low-complexity models based on various machine learning techniques, this paper introduces a new indicator to capture sudden quality variations and reports poor results for most of the models when applying this indicator. Consequently, an original way to train the models, using jointly the subjective scores and the frame level scores of a full-reference metric, is proposed. The high correlation through traditional indicators is preserved, while the efficiency on the new indicator is drastically improved. Joël Jung, Alexandre Giraud, Meijia Song, Songnan Li, Xiang Li 0003, Shan Liu 0001 |
ICASSP | 1 |
| 2022 | Immersive Video Coding: Should Geometry Information Be Transmitted as Depth Maps?abstractImmersive video often refers to multiple views with texture and scene geometry information, from which different viewports can be synthesized on the client side. To design efficient immersive video coding solutions, it is desirable to minimize bitrate, pixel rate and complexity. We investigate whether the classical approach of sending the geometry of a scene as depth maps is appropriate to serve this purpose. Previous work shows that bypassing depth transmission entirely and estimating depth at the client side improves the synthesis performance while saving bitrate and pixel rate. In order to understand if the encoder side depth maps contain information that is beneficial to be transmitted, we first explore a hybrid approach which enables partial depth map transmission using a block-based RD-based decision in the depth coding process. This approach reveals that partial depth map transmission may improve the rendering performance but does not present a good compromise in terms of compression efficiency. This led us to address the remaining drawbacks of decoder side depth estimation: complexity and depth map inaccuracy. We propose a novel system that takes advantage of high quality depth maps at the server side by encoding them into lightweight features that support the depth estimator at the client side. These features allow reducing the amount of data that has to be handled during decoder side depth estimation by 88%, which significantly speeds up the cost computation and the energy minimization of the depth estimator. Furthermore, −46.0% and −37.9% average synthesis BD-Rate gains are achieved compared to the classical approach with depth maps estimated at the encoder. Patrick Garus, Félix Henry, Joël Jung, Thomas Maugey, Christine Guillemot |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Overview and Efficiency of Decoder-Side Depth Estimation in MPEG Immersive VideoabstractThis paper presents the overview and rationale behind the Decoder-Side Depth Estimation (DSDE) mode of the MPEG Immersive Video (MIV) standard, using the Geometry Absent profile, for efficient compression of immersive multiview video. A MIV bitstream generated by an encoder operating in the DSDE mode does not include depth maps. It only contains the information required to reconstruct them in the client or in the cloud: decoded views and metadata. The paper explains the technical details and techniques supported by this novel MIV DSDE mode. The description additionally includes the specification on Geometry Assistance Supplemental Enhancement Information which helps to reduce the complexity of depth estimation, when performed in the cloud or at the decoder side. The depth estimation in MIV is a non-normative part of the decoding process, therefore, any method can be used to compute the depth maps. This paper lists a set of requirements for depth estimation, induced by the specific characteristics of the DSDE. The depth estimation reference software, continuously and collaboratively developed with MIV to meet these requirements, is presented in this paper. Several original experimental results are presented. The efficiency of the DSDE is compared to two MIV profiles. The combined non-transmission of depth maps and efficient coding of textures enabled by the DSDE leads to efficient compression and rendering quality improvement compared to the usual encoder-side depth estimation. Moreover, results of the first evaluation of state-of-the-art multiview depth estimators in the DSDE context, including machine learning techniques, are presented. Dawid Mieloch, Patrick Garus, Marta Milovanovic, Joël Jung, Jun Young Jeong, Smitha Lingadahalli Ravi, Basel Salahieh |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Patch Decoder-Side Depth Estimation In Mpeg Immersive VideoabstractThis paper presents a new approach for achieving bitrate and pixel rate reduction in the MPEG immersive video coding setting. We demonstrate that it is possible to avoid the transmission of some depth information in the Test Model for Immersive Video (TMIV) by estimating it at the receiver's side. Although the transmitted information in TMIV is considered as non-redundant, we show that it is possible to improve this algorithm. This method provides 3.4%, 9.0%, and 12.1% average BD-rate gain for natural content on high, medium, and low bitrate, respectively, with up to respectively 12.3%, 16.0%, and 18.4% peak reductions. Moreover, it preserves the perceptual quality as measured with MS-SSIM and VMAF metrics. Additionally, it decreases the pixel rate by 8.3% for each test sequence. Marta Milovanovic, Félix Henry, Marco Cagnazzo, Joël Jung |
ICASSP | 4 |
| 2021 | MPEG Immersive Video Coding StandardabstractThis article introduces the ISO/IEC MPEG Immersive Video (MIV) standard, MPEG-I Part 12, which is undergoing standardization. The draft MIV standard provides support for viewing immersive volumetric content captured by multiple cameras with six degrees of freedom (6DoF) within a viewing space that is determined by the camera arrangement in the capture rig. The bitstream format and decoding processes of the draft specification along with aspects of the Test Model for Immersive Video (TMIV) reference software encoder, decoder, and renderer are described. The use cases, test conditions, quality assessment methods, and experimental results are provided. In the TMIV, multiple texture and geometry views are coded as atlases of patches using a legacy 2-D video codec, while optimizing for bitrate, pixel rate, and quality. The design of the bitstream format and decoder is based on the visual volumetric video-based coding (V3C) and video-based point cloud compression (V-PCC) standard, MPEG-I Part 5. Jill M. Boyce, Renaud Doré, Adrian Dziembowski, Julien Fleureau, Joël Jung, Bart Kroon, Basel Salahieh, Vinod Kumar Malamal Vadakital, Lu Yu 0003 |
Proc. IEEE | 5 |
| 2019 | Compression Improvement via Reference Organization for 2D-multiview ContentabstractOne of the most challenging goals of future immersive services is to enable the observation of a scene from any viewpoint, thus making free-navigation possible under certain constraints. In order to provide such kind of services with smooth navigation, a huge amount of views should be available on the client's device. In particular, it is important for the case of 2D-multiview content, where cameras are positioned on a 2D grid in order to provide both horizontal and vertical parallax. This kind of content requires a large coding rate; therefore improving the compression performance of video encoders is especially relevant in this case. This paper studies how the encoder configuration affects the compression, by taking into account the spatial position of each camera. Four parameters are addressed in this work: coding order of the views, the number of reference lists, the number of reference pictures, and the ordering of pictures in the reference lists. An average of 12.0% bitrate saving is achieved for medium bitrate and 11.1% for low bitrate compared to the state of the art techniques. Pavel Nikitin, Marco Cagnazzo, Joël Jung |
ICASSP | 3 |
| 2019 | Bypassing Depth Maps Transmission For Immersive Video CodingabstractThis paper addresses several downsides of the system under development in MPEG-I for coding and transmission of immersive media. We present a solution, which enables Depth-Image-Based Rendering for immersive video applications, while lifting the requirement of transmitting depth information. Instead, we estimate the depth information on the client-side from the transmitted views. The approach leads to an impressive rate saving (37.3% in average). Preserving perceptual quality in terms of MS-SSIM of synthesized views, it yields to 24.6% rate reduction for the same quality of reconstructed views after residue transmission under the MPEG-I common test conditions. Simultaneously, the required pixel rate, i.e. the number of pixels processed per second by the decoder, is reduced by 50% for any test sequence. To the author's knowledge, this is the first time that such an approach is under consideration in the context of immersive video coding. Patrick Garus, Joël Jung, Thomas Maugey, Christine Guillemot |
PCS | 2 |
| 2019 | Flexible Motion Vector Resolution Prediction for Video CodingabstractThe latest video coding standard, High Efficiency Video Coding (HEVC), uses quarter-pixel motion vector (MV) resolution for motion compensation. The adaptation of MV resolution supported by progressive MV resolution (PMVR) brings further improvement to performance by progressively adjusting the resolution according to the distance between the MV and its predictor. However, progressive adjustment of resolution by PMVR does not consider the inherent characteristics of the coding block. In this paper, we propose several ways to improve PMVR. First, we show that the performance of PMVR is correlated with the spatiotemporal characteristics of the video sequence. Then, to cope with the limitations of PMVR, we propose a flexible framework for the adaptation of MV resolution using: 1) PU size and gradient; 2) PU size, gradient, and MV components; and 3) PU size and spatiotemporal characteristics of the frames. Finally, a smart motion estimation around multiple MV predictors is performed to take full advantage of the proposed scheme. The proposed tools are implemented on top of HM-16.6. Extensive experiments and comparison with HEVC show 1.3%, 2.7%, and 1.0% average BD-Rate savings for random access, low-delay P, and low-delay B configurations, respectively. Bappaditya Ray, Mohamed-Chaker Larabi, Joël Jung |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | A Low-Complexity Video Encoder for Equirectangular Projected 360 Video Contentabstract360- video is gaining a lot of interest because of the immersive feeling brought by such a technology. Several projection formats are used to represent this type of content. Equirectangular projection (ERP) is one of the most widely used projection scheme for 360 panoramic content. The main drawback of ERP is its latitude dependent sampling density unlike conventional 2D content. Consequently, conventional 2D codecs such as HEVC are not optimal for the coding of ERP projected 360 content. To cope with this dependency, this work proposes an adaptation of motion vector resolution and minimum width of the coding block depending on its latitude. Experimental results show up to 0.5% BD-rate savings for motion contained sequences with 15% encoding time reduction in random access configuration. Bappaditya Ray, Joël Jung, Mohamed-Chaker Larabi |
ICASSP | 2 |
| 2017 | A block level adaptive MV resolution for video codingabstractThe latest video coding standard, HEVC, uses quarter pixel motion vector (MV) resolution for motion compensation. The adaptation of MV resolution supported by PMVR (progressive MV resolution) brings further improvement of the performance, by progressively adjusting the resolution according to the distance between the MV and the MV predictor (MVP). In this work, we propose to improve PMVR by adapting MV resolution at the prediction unit (PU) level relying on its size and its average absolute gradient. We additionally perform a smarter motion estimation around multiple MV predictors to fully take advantage of the proposed scheme. Compared to HEVC reference software (HM-16.6), the proposed method provides 1.2%, 3.2% and 1.2% average BD rate savings respectively for random access (RA), low-delay P (LP) and low-delay (LD) configurations. Bappaditya Ray, Joël Jung, Mohamed-Chaker Larabi |
ICME | 2 |
| 2015 | Subjectie evaluation of Super Multi-View compressed contents on high-end light-field 3D displays
Antoine Dricot, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu, Frédéric Dufaux, Péter Tamás Kovács, Vamsi Kiran Adhikarla |
Signal Process. Image Commun. | 2 |
| 2014 | Smart decoder: A new paradigm for video codingabstractThe coding efficiency of the new video coding standard, High Efficiency Video Coding (HEVC), is strongly associated with better use of spatio-temporal redundancies thanks to an increased number of competing coding modes. However, this competition involves a massive increase in signaling bitrate which becomes a possible limit for the next generation of encoder. This paper proposes a new coding scheme that breaks with conventional approaches. It exploits a more complex decoder able to reproduce the choice of the encoder based on causal references, eliminating thus the need to signal coding modes and associated parameters. The general outline of this new codec and a proposed implementation are described in this paper. Experimental results under common test conditions report an average bitrate saving of 1.7% at the same quality compared to HEVC for a wide range of video sequences. D.-K. Vo-Nguyen, Joël Jung, Jean-Marc Thiesse, Marc Antonini |
ICASSP | 2 |
| 2014 | Full parallax super multi-view video codingabstractSuper Multi-View (SMV) video is a key enabler for future 3D video services that allows a glasses-free visualization and eliminates many causes of discomfort existing in current available 3D video technologies. SMV video content is composed of tens or hundreds of views, that can be aligned in horizontal only or both horizontal and vertical directions, providing respectively horizontal parallax or full parallax. This paper compares several coding schemes and coding orders, and proposes a coding structure that exploits inter-view correlations in the two directions, providing BD-rate gains up to 29.1% when compared to a basic anchor structure. Additionally, Neighboring Block Disparity Vector (NBDV) and Inter-View Motion Prediction (IVMP) coding tools are further improved to efficiently exploit coding structures in two dimensions, with BD-rate gains up to 4.2% reported over the reference 3D-HEVC encoder. Antoine Dricot, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu, Frédéric Dufaux |
ICIP | 2 |
| 2014 | Initialization, Limitation, and Predictive Coding of the Depth and Texture Quadtree in 3D-HEVCabstractThe 3D video extension of High Efficiency Video Coding (3D-HEVC) exploits texture-depth redundancies in 3D videos using intercomponent coding tools. It also inherits the same quadtree coding structure as HEVC for both components. The current software implementation of 3D-HEVC includes encoder shortcuts that speed up the quadtree construction process, but those are always accompanied by coding losses. Furthermore, since the texture and its associated depth represent the same scene, at the same time instant and view point, their quadtrees are closely linked. In this paper, an intercomponent tool is proposed in which this link is exploited to save both runtime and bits through a joint coding of the quadtrees. If depth is coded before the texture, the texture quadtree is initialized from the coded depth quadtree. Otherwise, the depth quadtree is limited to the coded texture quadtree. A 31% encoder runtime saving, a -0.3% gain for coded and synthesized views and a -1.8% gain for coded views are reported for the second method. Elie Gabriel Mora, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Modification of the merge candidate list for dependent views in 3D-HEVCabstractA test model for an HEVC-based 3D video coding standard (3D-HEVC) has recently been drafted. 3D-HEVC exploits inter-view redundancies by including disparity-compensated prediction (DCP) for efficient dependent view coding. It also uses the Merge coding mode to reduce the cost of motion / disparity parameters. However, the candidates in the Merge list are mostly temporal motion vectors. DCP does not often benefit from accurate predictors and is thus costly. Consequently, motion-compensated prediction (MCP) remains largely preferred. In this paper, we propose to reduce the cost of DCP by modifying the Merge candidate list to always include a disparity vector candidate. Two methods are proposed: the new candidate is either added in the secondary or in the primary list of candidates. The latter method, which achieves average bitrate reductions of 0.6% for dependent views, and 0.2% for coded and synthesized views, was adopted in both the 3D-HEVC working draft and software. Elie Gabriel Mora, Joël Jung, Marco Cagnazzo, Béatrice Pesquet-Popescu |
ICIP | 2 |
| 2013 | Modification of the disparity vector derivation process in 3D-HEVCabstractThe up-and-coming extension of HEVC for 3D video (3D-HEVC) includes various tools to exploit different redundancies in a 3D video signal. Inter-view redundancies are in particular exploited using Inter-View Motion Prediction (IVMP) and Inter-View Residual Prediction (IVRP). Both of these tools compensate disparity-wise the current prediction unit (PU) in order to find its corresponding PU in a base view, from which some prediction information for the current PU is retrieved. The disparity vector (DV) used for disparity compensation is currently derived using a neighboring search process (NBDV) for a DV across spatial and temporal neighbors. The first DV found is selected as the final DV used in IVMP and IVRP, with no guarantee of optimality. In this paper, the NBDV derivation process is changed: all found DVs from different neighbors are stored in a list. Redundant vectors in this list are removed, and a median computation on the remaining vectors is performed. The resulting DV is set as the DV used for IVMP. Average bitrate reductions of 0.6% and 0.8% for the two dependent views and 0.2% on synthesized views are reported with only a slight increase in encoder and decoder runtimes. Elie Gabriel Mora, Joël Jung, Béatrice Pesquet-Popescu, Marco Cagnazzo |
MMSP | 2 |
| 2012 | Block Merging for Quadtree-Based Partitioning in HEVCabstractThe joint development of the upcoming High Efficiency Video Coding (HEVC) standard by ITU-T Video Coding Experts Group and ISO/IEC Moving Picture Experts Group marks a new step in video compression capability. In technical terms, HEVC is a hybrid video-coding approach using quadtree-based block partitioning together with motion-compensated prediction. Even though a high degree of adaptability is achieved by quadtree-based block partitioning, this approach has certain intrinsic drawbacks, which may result in redundant sets of motion parameters being transmitted. Previous work has shown that those redundancies can effectively be removed by merging the leafs of a particular quadtree structure. Following this concept, a block merging algorithm for HEVC is now proposed. This algorithm generates a single motion parameter set for a whole region of contiguous motion-compensated blocks. In this paper, we describe the various components of the proposed block merging algorithm and, using experimental evidence, demonstrate their benefits in terms of coding efficiency. Philipp Helle, Simon Oudin, Benjamin Bross, Detlev Marpe, M. Oguz Bici, Kemal Ugur, Joël Jung, Gordon Clare, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2011 | A New Coding Mode for Hybrid Video Coders Based on Quantized Motion VectorsabstractThe rate allocation tradeoff between motion vectors and transform coefficients has a major importance when it comes to efficient video compression. This paper introduces a new coding mode for an H.264/AVC-like video coder, which improves the management of this resource allocation. The proposed technique can be used within any hybrid video encoder allowing a different coding mode for any macroblock. The key tool of the new mode is the lossy coding of motion vectors, obtained via quantization: while the transformed motion-compensated residual is computed with a high-precision motion vector, the motion vector itself is quantized before being sent to the decoder, in a rate/distortion optimized way. Several problems have to be faced with in order to get an efficient implementation of the coding mode, especially the coding and prediction of the quantized motion vectors, and the selection and encoding of the quantization steps. This new coding mode improves the performance of the hybrid video encoder over several sequences at different resolutions. Marie Andrée Agostini, Marco Cagnazzo, Marc Antonini, Guillaume Laroche, Joël Jung |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | Rate Distortion Data Hiding of Motion Vector Competition Information in Chroma and Luma Samples for Video CompressionabstractNew standardization activities have been recently launched by the JCT-VC experts group in order to challenge the current video compression standard H.264/AVC. Several improvements of this standard, previously integrated in the JM key technical area software, are already known and gathered in the high efficiency video coding test model. In particular, competition-based motion vector prediction has proved its efficiency. However, the targeted 50% bitrate saving for equivalent quality is not yet achieved. In this context, this paper proposes to reduce the signaling information resulting from this motion vector competition, by using data hiding techniques. As data hiding and video compression traditionally have contradictory goals, an advanced study of data hiding schemes is first performed. Then, an original way of using data hiding for video compression is proposed. The main idea of this paper is to hide the competition index into appropriately selected chroma and luma transform coefficients. To minimize the prediction errors, the transform coefficients modification is performed via a rate-distortion optimization. The proposed scheme is evaluated on several low and high resolution sequences. Objective improvements (up to 2.40% bitrate saving) and subjective assessment of the chroma loss are reported. Jean-Marc Thiesse, Joël Jung, Marc Antonini |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Data hiding of intra prediction information in chroma samples for video compressionabstractNew activities have been recently launched in order to challenge the H.264/AVC standard. Several improvements of this standard are already known, however the targeted 50% bitrate saving for equivalent quality is not yet achieved. In this context, a previous work proposes to use data hiding techniques to reduce the signaling information resulting from an improvement of Inter-coding. The main idea is to hide the indices into appropriately selected chroma and luma transform coefficients. To minimize the prediction errors, the modification is performed via a rate-distortion optimization. In this paper, the scheme is extended for Intra-coding and 4:4:4 sequences in order to explore the limits highlighted by the previous study and tackle some remaining issues. A different rate-distortion optimization built on the Pareto theory is also proposed. Resulting improvements (1.5% in average) for several sequences are reported and analyzed. Jean-Marc Thiesse, Joël Jung, Marc Antonini |
ICIP | 2 |
| 2010 | Motion vector forecast and mapping (MV-FMap) method for entropy coding based video codersabstractSince the finalization of the H.264/AVC standard and in order to meet the target set by both ITU-T and MPEG to define a new standard that reaches 50% bit rate reduction compared to H.264/AVC, many tools have efficiently improved the texture coding and the motion compensation accuracy. These improvements have resulted in increasing the proportion of bit rate allocated to motion information. Thus, the bit rate reduction of this information becomes a key subject of research. This paper proposes a method for motion vector coding based on an adaptive redistribution of motion vector residuals before entropy coding. Motion information is gathered to forecast a list of motion vector residuals which are redistributed to unexpected residuals of lower coding cost. Compared to H.264/AVC, this scheme provides systematic gain on tested sequences, and 2.3% in average, reaching up to 4.9% for a given sequence. Julien Le Tanou, Jean-Marc Thiesse, Joël Jung, Marc Antonini |
MMSP | 3 |
| 2010 | Data hiding of motion information in chroma and luma samples for video compressionabstract2010 appears to be the launching date for new compression activities intended to challenge the current video compression standard H.264/AVC. Several improvements of this standard are already known like competition-based motion vector prediction. However the targeted 50% bitrate saving for equivalent quality is not yet achieved. In this context, this paper proposes to reduce the signaling information resulting from this vector competition, by using data hiding techniques. As data hiding and video compression traditionally have contradictory goals, a study of data hiding is first performed. Then, an efficient way of using data hiding for video compression is proposed. The main idea is to hide the indices into appropriately selected chroma and luma transform coefficients. To minimize the prediction errors, the modification is performed via a rate-distortion optimization. Objective improvements (up to 2.3% bitrate saving) and subjective assess ment of chroma loss are reported and analyzed for several sequences. Jean-Marc Thiesse, Joël Jung, Marc Antonini |
MMSP | 2 |
| 2010 | Video Coding Using a Simplified Block Structure and Advanced Coding TechniquesabstractThis paper describes a new video coding scheme based on a simplified block structure that significantly outperforms the coding efficiency of the ISO/IEC 14496-10 ITU-T H.264 advanced video coding (AVC) standard. Its conceptual design is similar to a typical block-based hybrid coder applying prediction and subsequent prediction error coding. The basic coding unit is an 8 × 8 block for inter, and an 8 × 8 or a 16 × 16 block for intra, instead of the usual 16 × 16 macroblock. No larger block sizes are considered for prediction and transform. Based on this simplified block structure, the coding scheme uses simple and fundamental coding tools with optimized encoding algorithms. In particular, the motion representation is based on a minimum partitioning with blocks sharing motion borders. In addition, compared to AVC, the new and improved coding techniques include: block-based intensity compensation, motion vector competition, adaptive motion vector resolution, adaptive interpolation filters, edge-based intra prediction and enhanced chrominance prediction, intra template matching, larger trans forms and adaptive switchable transforms selection for intra and inter blocks, and nonlinear and frame-adaptive de-noising loop filters. Finally, the entropy coder uses a generic flexible zero tree representation applied to both motion and texture data. Attention has also been given to algorithm designs that facilitate parallelization. Compared to AVC, the new coding scheme offers clear benefits in terms of subjective video quality at the same bit rate. Objective quality improvements are equally significant. At the same quality, an average bit-rate reduction of 31% compared to AVC is reported. Frank Bossen, Virginie Drugeon, Edouard François, Joël Jung, Sandeep Kanumuri, Matthias Narroschke, Hisao Sasai, Joel Sole, Yoshinori Suzuki, Thiow Keng Tan, Thomas Wedi, Steffen Wittmann, Peng Yin 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Intra Coding With Prediction Mode Information InferenceabstractIn a typical competition-based coding, the pertinence of a prediction mode does not only depend on its own efficiency but also on the fact that it is complementary with the other modes. The method proposed in this paper to improve the intra coding of the H.264/AVC standard relies on this remark; it shows how the cost of signaling predictors that are quite similar can be avoided. Indeed, at low bitrates, the information related to the predictor signaling in intra coding reaches up to 25% of the total bitrate for the whole set of standard VCEG test sequences. In order to reduce this cost, a method reproducible at the decoder side is proposed to eliminate some predictors from the intra predictor set. The proposed method exploits the proximity of the predictors in the transform domain in order to obtain a representative and non-redundant set of predictors. Guillaume Laroche, Joël Jung, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Improving H.264 performances by quantization of motion vectorsabstractThe coding resources used for motion vectors (MVs) can attain quite high ratios even in the case of efficient video coders like H.264, and this can easily lead to suboptimal rate-distortion performance. In a previous paper, we proposed a new coding mode for H.264 based on the quantization of motion vectors (QMV). We only considered the case of 16 times 16 partitions for motion estimation and compensation. That method allowed us to obtain an improved trade-off in the resource allocation between vectors and coefficients, and to achieve better rate-distortion performances with respect to H.264. In this paper, we build on the proposed QMV coding mode, extending it to the case of macroblock partition into smaller blocks. This issue requires solving some problems mainly related to the motion vector coding. We show how this task can be performed efficiently in our framework, obtaining further improvements over the standard coding technique. Silvia Corrado, Marie Andrée Agostini, Marco Cagnazzo, Marc Antonini, Guillaume Laroche, Joël Jung |
PCS | 6 |
| 2008 | RD Optimized Coding for Motion Vector Predictor SelectionabstractThe H.264/MPEG4-AVC video coding standard has achieved a higher coding efficiency compared to its predecessors. The significant bitrate reduction is mainly obtained by efficient motion compensation tools, as variable block sizes, multiple reference frames, 1/4-pel motion accuracy and powerful prediction modes (e.g., SKIP and DIRECT). These tools have contributed to an increased proportion of the motion information in the total bit- stream. To achieve the performance required by the future ITU-T challenge, namely to provide a codec with 50% bitrate reduction compared to the current H.264, the reduction of this motion information cost is essential. This paper proposes a competing framework for better motion vector coding and SKIP mode. The predictors for the SKIP mode and the motion vector predictors are optimally selected by a rate-distortion criterion. These methods take advantage from the use of the spatial and the temporal redundancies in the motion vector fields, where the simple spatial median usually fails. An adaptation of the temporal predictors according to the temporal distances between motion vector fields is also described for multiple reference frames and B-slices options. These two combined schemes lead to a systematic bitrate saving on Baseline and High profile, compared to an H.264/MPEG4-AVC standard codec, which reaches up to 45%. Guillaume Laroche, Joël Jung, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | RD Optimized Coding for Motion Vector Predictor SelectionabstractThe H.264/MPEG4-AVC video coding standard has achieved a higher coding efficiency compared to its predecessors. The significant bit rate reduction is mainly obtained by efficient motion compensation tools, as variable block sizes, multiple reference frames, 1/4-pel motion accuracy and powerful prediction modes (e.g., SKIP and DIRECT). These tools have contributed to an increased proportion of the motion information in the total bit stream. To achieve the performance required by the future ITU-T challenge, namely to provide a codec with 50% bit rate reduction compared to the current H.264, the reduction of this motion information cost is essential. Guillaume Laroche, Joël Jung, Béatrice Pesquet-Popescu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Competition Based Prediction for Skip Mode Motion Vector Using Macroblock Classification for the H.264 JM KTA Software
Guillaume Laroche, Joël Jung, Béatrice Pesquet-Popescu |
ACIVS | 2 |
| 2007 | RD-optimized competition scheme for efficient motion predictionabstractH.264/MPEG4-AVC is the latest video codec provided by the Joint Video Team, gathering ITU-T and ISO/IEC experts. Technically there are no drastic changes compared to its predecessors H.263 and MPEG-4 part 2. It however significantly reduces the bitrate and seems to be progressively adopted by the market. The gain mainly results from the addition of efficient motion compensation tools, variable block sizes, multiple reference frames, 1/4-pel motion accuracy and powerful Skip and Direct modes. A close study of the bits repartition in the bitstream reveals that motion information can represent up to 40% of the total bitstream. As a consequence reduction of motion cost is a priority for future enhancements. This paper proposes a competition-based scheme for the prediction of the motion. It impacts the selection of the motion vectors, based on a modified rate-distortion criterion, for the Inter modes and for the Skip mode. Combined spatial and temporal predictors take benefit of temporal redundancies, where the spatial median usually fails. An average 7% bitrate saving compared to a standard H.264/MPEG4-AVC codec is reported. In addition, on the fly adaptation of the set of predictors is proposed and preliminary results are provided. Joël Jung, Guillaume Laroche, Béatrice Pesquet-Popescu |
VCIP | 1 |
| 2004 | Low-power H.264 video decoder with graceful degradationabstractIn this paper, we present a novel H.264 video decoder where memory transfers and energy consumption are significantly reduced. A power-friendly solution is indeed required in mobile applications, where autonomy is a key feature. Whereas low-power design is usually regarded as a pure architectural and implementation problem, our solution is based on an algorithmic approach. We first observe that in modern hardware architectures, computational complexity plays a second role in terms of energy dissipation, which is dominated by memory transfers. We thus identify the motion compensation module as the most consuming part of the standard H.264 decoder, due to its numerous memory accesses to the reference frame(s). Through an algorithmic modification that uses concurrent embedded compression techniques and relies on the data structure of the memory, we reduce memory transfers, and hence power dissipation, while maintaining the visual quality. Experimental results prove that the new method is close to the reference H.264 baseline decoder both in terms of objective and subjective measurements. In the meantime, memory transfers have been reduced by 55% in average, implying power savings which lengthen the battery life of the mobile device, increase the reliability of the chip and lower production costs. Arnaud Bourge, Joël Jung |
VCIP | 2 |
| 2004 | Power-scalable video encoder for mobile devices based on collocated motion estimationabstractIn this paper, a method for designing low-power video schemes is presented. Algorithms that imply a very low dissipation are required for new applications where the energy source is limited, e.g. mobile phones including a camera and video features. Whereas it can be observed that video standards are mainly designed around coding efficiency, we propose to take into account power consumption characteristics directly when designing the algorithm. More precisely, we give some guidelines for the design of low-power video codecs in the scope of modern hardware architectures and we introduce the notion of power scalability. We present an original encoder based on so-called 'Collocated Motion Estimation' designed using the proposed methodology. Experimental results show that we remain close to the coding efficiency of the reference H.264 baseline encoder while the power consumption is largely reduced in our solution. Moreoever this encoder is scalable in memory transfer and computational complexity. Joël Jung, Arnaud Bourge |
VCIP | 1 |
| 2003 | Optimal decoder for block-transform based video codersabstractIn this paper, we introduce a new decoding algorithm for DCT-based video encoders, such as Motion JPEG (M-JPEG), H26x, or MPEG. This algorithm considers not only the compression artifacts but also the ones due to transmission, acquisition or storage of the video. The novelty of our approach is to jointly tackle these two problems, using a variational approach. The resulting decoder is object-based, allowing independent and adaptive processing of objects and backgrounds, and considers available information provided by the bitstream, such as quantization steps, and motion vectors. Several experiments demonstrate the efficiency of the proposed method. Objective and subjective quality assessment methods are used to evaluate the improvement upon standard algorithms, such as the deblocking and deringing filters included in MPEG-4 postprocessing. Joël Jung, Marc Antonini, Michel Barlaud |
IEEE Trans. Multim. | 1 |
| 2002 | Novel approach for temporal filtering of MPEG distortionsabstractStandard DCT -based video coding techniques perform good results in terms of data compaction, making feasible the use of digital video in several application frameworks. The price to pay is the introduction of annoying visual distortions/artefacts in the reconstructed video. The lower the encoding bit-rate, the larger the number of artefacts. Post-processing is a practical solution that achieves a visual enhancement of the compressed images after decoding. Some of the artefacts, such as blocking (tiled-effect aspect) and ringing (ghost effect) have already been widely studied. Sandra Del Corso, Joël Jung |
ICASSP | 2 |
| 2002 | Robust wavelet-based arbitrary grid detection for MPEGabstractFrom the industrial point of view, image quality is a key-issue. Many post-processing algorithms have been proposed to improve visual quality after the MPEG decoder. Most of them need precise location of the 8/spl times/8 grid on which the blocking effect appears. However, in real-life applications, the blocking effect is rarely located on such a basic grid, due to the cascaded bit-rate or format transcoding, rescaling, etc., that occur during acquisition, compression, transmission and display of the video. Consequently, most of these methods see their efficiency largely reduced, or are simply useless. A grid detector is proposed, based on a fine modeling of blocking artifacts in the wavelet domain. It aims at providing essential information to any post-processing algorithm that requires the position of the grid. Several experiments and reliable subjective tests demonstrate the accuracy of the proposed grid detector, and highlight the added value it yields to a post-processing algorithm in terms of visual quality. Estelle Lesellier, Joël Jung |
ICIP (3) | 2 |
| 2001 | New object-based variational approach for MPEG-2 data recovery over lossy packet networks
Joël Jung, Marc Antonini, Michel Barlaud |
VCIP | 1 |
| 1998 | Optimal JPEG DecodingabstractThis paper introduces an optimal decoding scheme for the baseline Joint Photographers Expert Group (JPEG) standard. In particular, it deals with the minimization of a half-quadratic criterion which takes into account observed data, a priori knowledge of the solution, and precise spatial location of blocking artifacts. The method considers, at the same time, information in the original spatial domain, in the DCT domain, and in the spatio-frequency (DWT) domain. A model of the quantizer is also included. The method reduces blocking and ringing artifacts, resulting in improved peak signal-to-noise ratio performance as well as greater visual quality. Joël Jung, Marc Antonini, Michel Barlaud |
ICIP (1) | 1 |