EDBT 2026 Demo / reviewers in the wild / expert
Luce Morin
dblp:95/6318
· DBLP profile ↗
60ranked-venue papers
0as first author
20since 2021 · last 2026
0000-0001-8241-1425ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 18 since 2021Artificial intelligence and machine learning · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LISA: A New Subjective Test Protocol and Tool for Local Image Quality Assessment
Ewen Démézet, Meriem Outtas, Séverine Baudry, Luce Morin, Lu Zhang 0037 |
QoMEX | 4 |
| 2025 | A New Benchmark Database and Objective Metric for Light Field Image Quality EvaluationabstractLight Field Image (LFI) records both angular and spatial information and provides immersive experiences for observers by rendering a scene from multiple perspectives. To cope with the resolution limitations of capture hardware, LFI angular reconstruction and spatial super-resolution are two widely-used methods, but they can also induce some special types of distortions, especially when two methods are adopted in combination. To this end, new challenges have been brought in assessing the quality of these distorted LFIs. In this paper, firstly, we conduct subjective experiments to evaluate the distorted LFI quality and present a novel perceptual quality assessment database with the associated subjective quality scores. Specifically, the proposed database focuses on the distortions introduced by deep learning-based LFI angular reconstruction and spatial super-resolution methods, individually and multiplely. Besides, in the case of multiple distortions, the adoption order of two distortions is taken into consideration. Further, our database presents three types of LFIs that suffer from distortions: real-world, dense synthesis, and sparse synthesis. As a result, the quality of distorted LFIs was subjectively assessed by 32 valid observers using the Pairwise Comparison (PC) protocol. Secondly, we develop a novel objective No-Reference (NR) metric for LFI quality evaluation, based on the features extracted from spatial gradients, angular-spatial statistics, and binocular disparity. Finally, a benchmark of the proposed metric and numerous state-of-the-art quality assessment metrics on the proposed database is presented. Experimental results demonstrate the superiority of the proposed metric over most existing metrics in various aspects. The proposed database and metric will be publicly available athttps://github.com/ZhengyuZhang96/IETR-LFI. Shishun Tian, Jinjia Zhou, Luce Morin, Lu Zhang 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Convex Hull Prediction Methods for Bitrate Ladder Construction: Design, Evaluation, and ComparisonabstractHTTP adaptive streaming (HAS) has emerged as a prevalent approach for over-the-top (OTT) video streaming services due to its ability to deliver a seamless user experience. A fundamental component of HAS is the bitrate ladder, which comprises a set of encoding parameters (e.g., bitrate-resolution pairs) used to encode the source video into multiple representations. This adaptive bitrate ladder enables the client’s video player to dynamically adjust the quality of the video stream in real-time based on fluctuations in network conditions, ensuring uninterrupted playback by selecting the most suitable representation for the available bandwidth. The most straightforward approach involves using a fixed bitrate ladder for all videos, consisting of pre-determined bitrate-resolution pairs known as one-size-fits-all . Conversely, the most reliable technique relies on intensively encoding all resolutions over a wide range of bitrates to build the convex hull , thereby optimizing the bitrate ladder by selecting the representations from the convex hull for each specific video. Several techniques have been proposed to predict content-based ladders without performing a costly, exhaustive search encoding. This article provides a comprehensive review of various convex hull prediction methods, including both conventional and learning-based approaches. Furthermore, we conduct a benchmark study of several handcrafted- and deep learning (DL)-based approaches for predicting content-optimized convex hulls across multiple codec settings. The considered methods are evaluated on our proposed large-scale dataset, which includes 300 UHD video shots encoded with software and hardware encoders using three state-of-the-art video standards, including AVC/H.264, HEVC/H.265, and VVC/H.266, at various bitrate points. Our analysis provides valuable insights and establishes baseline performance for future research in this field ( Dataset URL : https://nasext-vaader.insa-rennes.fr/ietr-vaader/datasets/br_ladder ). Ahmed Telili, Wassim Hamidouche, Hadi Amirpour, Sid Ahmed Fezza, Christian Timmerer, Luce Morin |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | NeRV++: An Enhanced Implicit Neural Video RepresentationabstractNeural fields, also known as implicit neural representations (INRs), have shown a remarkable capability of representing, generating, and manipulating various data types, allowing for continuous data reconstruction at a low memory footprint. Though promising, INRs applied to video compression still need to improve their rate-distortion performance by a large margin, and require a huge number of parameters and long training iterations to capture high-frequency details, limiting their wider applicability. Resolving this problem remains a quite challenging task, which would make INRs more accessible in compression tasks. We take a step towards resolving these shortcomings by introducing neural representations for videos (NeRV)++, an enhanced implicit neural video representation (INVR), as more straightforward yet effective enhancement over the original NeRV decoder architecture, featuring separable conv2d residual blocks (SCRBs) that sandwiches the upsampling block (UB), and a bilinear interpolation skip layer for improved feature representation. NeRV++ allows videos to be directly represented as a function approximated by a neural network, and significantly enhance the representation capacity beyond current INR-based video codecs. We evaluate our method on UVG, MCL JVC, and Bunny datasets, achieving competitive results for video compression with INRs. This achievement narrows the gap to autoencoder-based video coding, marking a significant stride in INR-based video compression research. The code of NeRV++ is available on GitHub. Ahmed Ghorbel, Wassim Hamidouche, Luce Morin |
VCIP | 3 |
| 2023 | AICT: An Adaptive Image Compression TransformerabstractMotivated by the efficiency investigation of the Tranformer-based transform coding framework, namely SwinT-ChARM, we propose to enhance the latter, as first, with a more straightforward yet effective Tranformer-based channel-wise auto-regressive prior model, resulting in an absolute image compression transformer (ICT). Current methods that still rely on ConvNet-based entropy coding are limited in long-range modeling dependencies due to their local connectivity and an increasing number of architectural biases and priors. On the contrary, the proposed ICT can capture both global and local contexts from the latent representations and better parameterize the distribution of the quantized latents. Further, we leverage a learnable scaling module with a sandwich ConvNeXt-based pre/post-processor to accurately extract more compact latent representation while reconstructing higher-quality images. Extensive experimental results on benchmark datasets showed that the proposed adaptive image compression transformer (AICT) framework significantly improves the trade-off between coding efficiency and decoder complexity over the versatile video coding (VVC) reference encoder (VTM-18.0) and the neural codec SwinT-ChARM. Ahmed Ghorbel, Wassim Hamidouche, Luce Morin |
ICIP | 3 |
| 2023 | Self-Supervised Focus Measure Fusing for Depth Estimation from Computer-Generated HologramsabstractDepth from focus is a simple and effective methodology for retrieving the scene geometry from a hologram when used with the appropriate focus measure and patch size. However, fixing those parameters for every sample may not be the right choice, as different scenes can be composed with various types of textures. In this work, we propose a self-supervised learning methodology for fusing the depth maps produced using different focus measures with variable patch sizes applied to the holographic reconstruction volume. Experimental results show that fusing depth information produces more accurate and smoother depth maps, which can be directly used for alternative tasks such as motion estimation. Nabil Madali, Antonin Gilles, Patrick Gioia, Luce Morin |
ICIP | 4 |
| 2023 | Efficient Per-Shot Transformer-Based Bitrate Ladder Prediction for Adaptive Video StreamingabstractRecently, HTTP adaptive streaming (HAS) has become a standard approach for over-the-top (OTT)-based video streaming services due to its ability to provide smooth streaming. In HAS, stream representations are encoded to target a specific bitrate providing a wide range of operating bitrates known as the bitrate ladder. In the past, a fixed bitrate ladder approach for all videos has been widely used. However, such a method does not consider video content, which can vary considerably in motion, texture, and scene complexity. Moreover, building a per-title bitrate ladder based on an exhaustive encoding is quite expensive due to the large encoding parameter space. Thus, alternative solutions allowing accurate and efficient per-title bitrate ladder prediction are in great demand. On the other hand, self-attention-based architectures have achieved tremendous performance in large language models (LLMs) and particularly vision transformers (ViTs) in computer vision tasks. Therefore, this paper investigates ViT’s capabilities in building an efficient bitrate ladder without performing any encoding process. We provide the first in-depth analysis of the prediction accuracy and the complexity overhead induced by the ViTs model in predicting the bitrate ladder on a large and diverse video dataset. The source code of the proposed solution and the dataset will be made publicly available. Ahmed Telili, Wassim Hamidouche, Sid Ahmed Fezza, Luce Morin |
ICIP | 4 |
| 2023 | Blind Quality Assessment of Light Field Image Based on Spatio-Angular Textural VariationabstractLight Field Image Quality Assessment (LF-IQA) is vitally important to facilitate the development of immersive technologies. However, current state-of-the-art LF-IQA metrics still struggle to handle Light Field Image (LFI) with massive data in an efficient manner. To cope with this challenge, we propose a simple yet effective Blind LF-IQA metric based on Spatio-Angular Textural Variation, named SATV-BLiF. Given a distorted LFI, we first apply Local Binary Pattern (LBP) operator to measure the textural variation in the spatial and angular domains respectively. Then the generated spatial and angular textural matrices are merged and further transformed into statistical textural histogram features. Finally, Support Vector Regression (SVR) is employed to construct a nonlinear mapping function between the statistical textural histogram features and the perceptual quality score of the distorted LFI. Experimental results on three representative light field databases show that the proposed metric achieves state-of-the-art quality evaluation performance, while having much lower complexity than the existing No-Reference (NR) LF-IQA metrics. The code of the proposed SATV-BLiF metric is available at https://github.com/ZhengyuZhang96/SATV-BLiF. Shishun Tian, Wenbin Zou, Yuhang Zhang 0011, Luce Morin, Lu Zhang 0037 |
ICIP | 5 |
| 2023 | Towards Machine Perception Aware Image Quality AssessmentabstractOver the years, the objective of image and video compression has been to preserve perceived quality according to the Human Visual System (HVS) with minimal rate. Traditional encoders achieve this with the use of Rate-Distortion Optimization (RDO) techniques along with Image Quality Assessment (IQA) metrics that are correlated with human perception. Nowadays, a fast-growing number of applications fall within the realm of Video Coding for Machines (VCM), where the final recipient of compressed data is not a human but a machine performing a vision task. Recently, the lack of correlation between existing distortion measures and machine perception has been revealed, especially for RDO algorithms where distortion measures are computed on a local scale. In this paper, we propose a machine perception-aware metric designed to be incorporated into a standard-compliant Versatile Video Coding (VVC) encoder. Our proposed metric relies on a supervised training procedure as well as additional information available on the encoder side. In terms of correlation with machine perception, our metric significantly outperforms existing distortion measures in the literature. Alban Marie, Karol Desnos, Jinjia Zhou, Luce Morin, Lu Zhang 0037 |
MMSP | 5 |
| 2023 | Open access dataset of holographic videos for codec analysis and machine learning applicationsabstractDespite the growing interest for Holography, there is a lack of publicly available three-dimensional hologram sequences for the evaluation of video codecs with inter-frame compression mechanisms such as motion estimation and compensation. In this paper, we report the first large-scale dataset containing 18 holographic videos computed with three different resolutions and pixel pitches. By providing the color and depth images corresponding to each hologram frame, our dataset can be used in additional applications such as the validation of 3D scene geometry retrieval or deep learning-based hologram synthesis methods. Altogether, our dataset comprises 5400 pairs of RGB-D images and holograms, totaling more than 550 GB of data. Antonin Gilles, Patrick Gioia, Nabil Madali, Anas El Rhammad, Luce Morin |
QoMEX | 5 |
| 2023 | Evaluation of Image Quality Assessment Metrics for Semantic Segmentation in a Machine-to-Machine Communication ScenarioabstractImage and video compression aims at finding an optimal trade-off between rate and distortion. This is done through Rate-Distortion Optimization (RDO) in traditional en-coders with the use of Image Quality Assessment (IQA) metrics. While it is known that most IQA metrics are designed to be correlated with human perception, there is no evidence that this observation can be generalized in a Video Coding for Machines (VCM) context, where the receiver is not a human anymore but a machine. In this paper, we propose an evaluation protocol to measure the correlation level between conventional Full-Reference (FR) IQA metrics and machine perception through the semantic segmentation vision task. Experiments showed a relatively low correlation between them when measured on the block-level. This observation implies the need of RDO algorithms that are better suited for Machine-to-Machine (M2M) communications. In order to facilitate the emergence of IQA metrics that better reflect machine perception, the code and dataset used to perform this study is made freely available at https://github.com/albmarie/iqa_m2m_segmentation. Alban Marie, Karol Desnos, Luce Morin, Lu Zhang 0037 |
QoMEX | 3 |
| 2023 | EDDMF: An Efficient Deep Discrepancy Measuring Framework for Full-Reference Light Field Image Quality AssessmentabstractThe increasing demand for immersive experience has greatly promoted the quality assessment research of Light Field Image (LFI). In this paper, we propose an efficient deep discrepancy measuring framework for full-reference light field image quality assessment. The main idea of the proposed framework is to efficiently evaluate the quality degradation of distorted LFIs by measuring the discrepancy between reference and distorted LFI patches. Firstly, a patch generation module is proposed to extract spatio-angular patches and sub-aperture patches from LFIs, which greatly reduces the computational cost. Then, we design a hierarchical discrepancy network based on convolutional neural networks to extract the hierarchical discrepancy features between reference and distorted spatio-angular patches. Besides, the local discrepancy features between reference and distorted sub-aperture patches are extracted as complementary features. After that, the angular-dominant hierarchical discrepancy features and the spatial-dominant local discrepancy features are combined to evaluate the patch quality. Finally, the quality of all patches is pooled to obtain the overall quality of distorted LFIs. To the best of our knowledge, the proposed framework is the first patch-based full-reference light field image quality assessment metric based on deep-learning technology. Experimental results on four representative LFI datasets show that our proposed framework achieves superior performance as well as lower computational complexity compared to other state-of-the-art metrics. Shishun Tian, Wenbin Zou, Luce Morin, Lu Zhang 0037 |
IEEE Trans. Image Process. | 4 |
| 2022 | Deeblif: Deep Blind Light Field Image Quality Assessment by Extracting Angular and Spatial InformationabstractIn the era of immersive media, the high-dimensional Light Field Image (LFI) puts forward higher requirements for Light Field Image Quality Assessment (LF-IQA). However, currently most existing LF-IQA metrics still rely on sophisticated hand-crafted feature extraction, which fail to predict the quality of LFI accurately. In this paper, we propose a patch-based Deep Blind Light Field image quality assessment metric (abbreviated as DeeBLiF), by employing a two-stream Convolutional Neural Network (CNN) model specifically designed for extracting the angular and spatial information of LFI. Firstly, the spatio-angular patches are generated as input data, which effectively reflect the spatio-angular information of LFI. After that, a two-stream CNN model is exploited to extract the patch features and further predict the patch scores. Finally, all the patch scores are pooled into an overall quality score of LFI. Experimental results on the LFI dataset demonstrate that the proposed DeeBLiF outperforms the state-of-the-art LF-IQA metrics. The code will be publicly available at https://github.com/ZhengyuZhang96/DeeBLiF. Shishun Tian, Wenbin Zou, Luce Morin, Lu Zhang 0037 |
ICIP | 4 |
| 2022 | Video Coding for Machines: Large-Scale Evaluation of Deep Neural Networks Robustness to Compression Artifacts for Semantic SegmentationabstractIn the Video Coding for Machines (VCM) context where visual content is compressed before being transmitted to a vision task algorithm, appropriate trade-off between the compression level and the vision task performance must be chosen. In this paper, a Deep Neural Networks (DNN) based semantic segmentation algorithm robustness to compression artifacts is evaluated with a total of 1486 different coding configurations. Results indicate the importance of using an appropriate image resolution to overcome the block-partitioning limitations in existing compression algorithms, allowing 58.3%, 49.8%, 33.5% and 24.3% bitrate savings at equivalent prediction accuracy for JPEG, JM, x265 and VVenC, respectively. Surprisingly, JPEG can achieve 73.41% bitrate reduction with the inclusion of compressed images at training time over VVC Test Model (VTM) with a DNN trained on pristine data, which implies that DNN generalization ability must not be overlooked. Alban Marie, Karol Desnos, Luce Morin, Lu Zhang 0037 |
MMSP | 3 |
| 2022 | A Study of Conventional and Learning-Based Depth Estimators for Immersive Video TransmissionabstractObtaining an accurate depth map of a scene is very important for major applications like immersive video, robotics, autonomous driving, and many more. The different methods to estimate depths can be classified as conventional and learning-based methods. While these methods have been studied for their depth accuracy, less attention has been paid to studying their performance in the use case of depth image-based rendering (DIBR). Here we study and evaluate two conventional methods and five learning-based methods for a real-world use case of immersive video transmission in the context of MPEG-I. The user-requested views are synthesized using Test Model for Immersive Video (TMIV) from the depth maps obtained by all methods and original texture views. The synthesized images are compared with their original counterparts using various quality metrics. Smitha Lingadahalli Ravi, Marta Milovanovic, Luce Morin, Félix Henry |
MMSP | 3 |
| 2022 | Benchmarking Learning-based Bitrate Ladder Prediction Methods for Adaptive Video StreamingabstractHTTP adaptive streaming (HAS) is increasingly adopted by over-the-top (OTT)-based video streaming services, it allows clients to dynamically switch among various stream representations. Each of these representations is encoded to target a specific bitrate providing a wide range of operating bitrates known as the bitrate ladder. Several approaches with different levels of complexity are currently used to build such a bitrate ladder. The most straightforward method is to use a fixed bitrate ladder for all videos, which is a set of bitrate-resolution pairs, called “one-size-fits-all”, and the most complex is based on the intensive encoding of all resolutions over a wide bitrate range to construct the convex-hull. This latter is then used to obtain a per-title bitrate ladder. Recently, various methods relying on machine learning (ML) techniques have been proposed to predict content-based ladder without performing exhaustive search encoding. In this paper, we conduct a benchmark study of several handcrafted and deep learning (DL)-based approaches for predicting content-optimized bitrate ladder, which we believe provides baseline methods and will be useful for future research in this field. The obtained results, based on 200 video sequences compressed with the high-efficiency video coding (HEVC) encoder, reveal that the most efficient method predicts the bitrate ladder without performing any encoding process at the cost of a slight Bjøntegaard delta bitrate (BD-BR) loss of 1.43% compared to the exhaustive approach. The dataset and the source code of the considered methods are made publicly available at: https://github.com/atelili/Bitrate-Ladder-Benchmark. Ahmed Telili, Wassim Hamidouche, Sid Ahmed Fezza, Luce Morin |
PCS | 4 |
| 2021 | Model Selection CNN-based VVC Quality EnhancementabstractArtifact removal and filtering methods are inevitable parts of video coding. On one hand, new codecs and compression standards come with advanced in-loop filters and on the other hand, displays are equipped with high capacity processing units for post-treatment of decoded videos. This paper proposes a Convolutional Neural Network (CNN)-based post-processing algorithm for intra and inter frames of Versatile Video Coding (VVC) coded streams. Depending on the frame type, this method benefits from normative prediction signal by feeding it as an additional input along with reconstructed signal and a Quantization Parameter (QP)-map to the CNN. Moreover, an optional Model Selection (MS) strategy is adopted to pick the best trained model among available ones at the encoder side, and signal it to the decoder side. This MS strategy is applicable at both frame level and block level. The experiments under the Random Access (RA) configuration of the VVC Test Model (VTM-10.0) show that the proposed prediction-aware algorithm can bring an additional BD-BR gain of -1.3% compared to the method without the prediction information. Furthermore, the proposed MS scheme brings -0.5% more BD-BR gain on top of the prediction-aware method. Fatemeh Nasiri, Wassim Hamidouche, Luce Morin, Nicolas Dhollande, Gildas Cocherel |
PCS | 3 |
| 2021 | Plane-based Accurate Registration of Real-world Point CloudsabstractTraditional 3D point clouds registration algorithms, based on Iterative Closest Point (ICP), rely on point matching of large point clouds. In well-structured environments, such as buildings, planes can be segmented and used for registration, similarly to the classical point-based ICP approach. Using planes tremendously reduces the number of inputs.In this article, an efficient plane-based registration algorithm is presented. The optimal transformation is estimated through a two-step approach, successively performing robust plane-to-plane minimization and non-linear robust point-to-plane registration. Experiments on the Autonomous Systems Lab (ASL) benchmark dataset show that the proposed method enables to successfully register 100% of the scans from the three indoor sequences. Experiments also show that the proposed method is robust in large motion scenarios and more accurate than other state-of-the-art algorithms. Moreover, a new challenging dataset, LOOP’IN, is provided. It is composed of two loops in real-world indoor scenes, with a large number of scans captured with a 3D LiDAR. Tests led on this dataset show that the algorithm is able to register long sequences, to close loops and to build an incremental map of the explored environment. Ketty Favre, Muriel Pressigout, Éric Marchand, Luce Morin |
SMC | 4 |
| 2021 | Quality assessment of DIBR-synthesized views: An overview
Shishun Tian, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Ting Su 0004, Luce Morin, Olivier Déforges |
Neurocomputing | 6 |
| 2021 | Quality-Driven Variable Frame-Rate for Green Video Coding in Broadcast ApplicationsabstractThe Digital Video Broadcasting (DVB) has proposed to introduce the Ultra-High Definition services in three phases: UHD-1 phase 1, UHD-1 phase 2 and UHD-2. The UHD-1 phase 2 specification includes several new features such as High Dynamic Range (HDR) and High Frame-Rate (HFR). It has been shown in several studies that HFR (+100 fps) enhances the perceptual quality and that this quality enhancement is content-dependent. On the other hand, HFR brings several challenges to the transmission chain including codec complexity increase and bit-rate overhead, which may delay or even prevent its deployment in the broadcast echo-system. In this paper, we propose a Variable Frame Rate (VFR) solution to determine the minimum (critical) frame-rate that preserves the perceived video quality of HFR video. The frame-rate determination is modeled as a 3-class classification problem which consists in dynamically and locally selecting one frame-rate among three: 30, 60 and 120 frames per second. Two random forests classifiers are trained with a ground truth carefully built by experts for this purpose. The subjective results conducted on ten HFR video contents, not included in the training set, clearly show the efficiency of the proposed solution enabling to locally determine the lowest possible frame-rate while preserving the quality of the HFR content. Moreover, our VFR solution enables significant bit-rate savings and complexity reductions at both encoder and decoder sides. Glenn Herrou, Charles Bonnineau, Wassim Hamidouche, Patrick Dumenil, Jérôme Fournier, Luce Morin |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | A Plane-based Approach for Indoor Point Clouds RegistrationabstractIterative Closest Point (ICP) is one of the mostly used algorithms for 3D point clouds registration. This classical approach can be impacted by the large number of points contained in a point cloud. Planar structures, which are less numerous than points, can be used in well-structured man-made environment. In this paper we propose a registration method inspired by the ICP algorithm in a plane-based registration approach for indoor environments. This method is based solely on data acquired with a LiDAR sensor. A new metric based on plane characteristics is introduced to find the best plane correspondences. The optimal transformation is estimated through a two-step minimization approach, successively performing robust plane-to-plane minimization and non-linear robust point-to-plane registration. Experiments on the Autonomous Systems Lab (ASL) dataset show that the proposed method enables to successfully register 100 % of the scans from the three indoor sequences. Experiments also show that the proposed method is more robust in large motion scenarios than other state-of-the-art algorithms. Ketty Favre, Muriel Pressigout, Éric Marchand, Luce Morin |
ICPR | 4 |
| 2020 | Prediction-Aware Quality Enhancement of VVC Using CNNabstractThe upcoming video coding standard, Versatile Video Coding (VVC), has shown great improvement compared to its predecessor, High Efficiency Video Coding (HEVC), in terms of bitrate saving. Despite its substantial performance, compressed videos might still suffer from quality degradation at low bitrates due to coding artifacts such as blockiness, blurriness and ringing. In this work, we exploit Convolutional Neural Networks (CNN) to enhance quality of VVC coded frames after decoding in order to reduce low bitrate artifacts. The main contribution of this work is the use of coding information from the compressed bitstream. More precisely, the prediction information of intra frames is used for training the network in addition to the reconstruction information. The proposed method is applied on both luminance and chrominance components of intra coded frames of VVC. Experiments on VVC Test Model (VTM) show that, both in low and high bitrates, the use of coding information can improve the BD-rate performance by about 1% and 6% for luma and chroma components, respectively. Fatemeh Nasiri, Wassim Hamidouche, Luce Morin, Nicolas Dhollande, Gildas Cocherel |
VCIP | 3 |
| 2019 | Low-Complexity Scalable Encoder Based on Local Adaptation of the Spatial ResolutionabstractA two-layer low-complexity scalable encoding scheme based on local Adaptive Spatial Resolution (ASR) is proposed. This scheme relies on a block-level spatial resolution adaptation in the enhancement layer encoder. For each block, the optimal resolution is either obtained via a rate-distortion optimization followed by a decision refinement process or by a prediction via motion compensation exploiting the base layer motion vectors. The proposed architecture has been integrated over of the High Efficiency Video Coding (HEVC) reference software (HM16.12) which is used as a base layer encoder. Compared to SHVC, the scalable extension of HEVC, experimental results show bitrate savings of 0.76 % as well as encoding complexity reductions of 47 % for the whole scalable encoder and 96 % for the enhancement layer encoder. Glenn Herrou, Wassim Hamidouche, Luce Morin |
ICIP | 3 |
| 2019 | Optimal Adaptive Quantization Based on Temporal Distortion Propagation Model for HEVCabstractOptimal adaptive quantization is one of the key points to optimize the coding efficiency of video encoders. The latest block-based video compression standards, such as high-efficiency video coding (HEVC), extensively use predictive coding techniques that create dependencies between blocks and increase the complexity of optimal block quantizers search. Specifically, the motion compensation is responsible for a dependency network connecting all blocks of the same GOP together. In this paper, this dependency network is estimated by a temporal distortion propagation model and an accurate estimation of Inter and Skip modes probabilities. Optimal quantizers are then designed per block in order to achieve global optimization in terms of rate-distortion efficiency. By implementing the algorithm into the HEVC reference model (HM), we report -16.51% PSNR-based and -26.26% SSIM-based average bitrate savings compared to no adaptive quantization. The proposed algorithm outperforms several related methods from the state-of-the-art. Moreover, along with the demonstration of an optimal quantizer solution, we propose an in-depth analysis of the algorithm behavior. This analysis includes, among others, the relative distribution of rates between frames and the control of quantizers dynamic range. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin |
IEEE Trans. Image Process. | 5 |
| 2019 | A Benchmark of DIBR Synthesized View Quality Assessment Metrics on a New Database for Immersive Media ApplicationsabstractDepth-image-based rendering (DIBR) is a fundamental technology in several 3-D-related applications, such as free viewpoint video, virtual reality, and augmented reality. However, new challenges have also been brought in assessing the quality of DIBR-synthesized views since this process induces some new types of distortions, which are inherently different from the distortion caused by video coding. In this paper, we present a new DIBR-synthesized image database with the associated subjective scores. We also test the performances of the state-of-the-art objective quality metrics on this database. This paper focuses on the distortions only induced by different DIBR synthesis methods. Seven state-of-the-art DIBR algorithms, including inter-view synthesis and single-view-based synthesis methods, are considered in this database. The quality of synthesized views was assessed subjectively by 41 observers and objectively using 14 state-of-the-art objective metrics. Subjective test results show that the interview synthesis methods, having more input information, significantly outperform the single-view-based ones. Correlation results between the tested objective metrics and the subjective scores on this database reveal that further studies are still needed for a better objective quality metric dedicated to the DIBR-synthesized views. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
IEEE Trans. Multim. | 3 |
| 2018 | Low-Complexity Spatial Scalability Scheme Using HEVC for 4K and VR VideosabstractScalable video coding enables to compress the video at different formats within a single layered bitstream. SHVC, the scalable extension of the High Efficiency Video Coding (HEVC) standard, enables x2 spatial scalability, among other additional features. The closed-loop architecture of the SHVC codec is based on the use of multiple instances of the HEVC codec to encode the video layers, which considerably increases the encoding complexity. With the arrival of new immersive video formats, like 4K, 8K, High Frame Rate (HFR) and 360° videos, the quantity of data to compress is exploding, making the use of high-complexity coding algorithms unsuitable. In this paper, we propose a low-complexity scalable coding scheme based on the use of a single HEVC codec instance and a wavelet-based decomposition as preprocessing. The pre-encoding image decomposition relies on well-known simple Discrete Wavelet Transform (DWT) kernels, such as Haar or Le Gall 5/3. Compared to SHVC, the proposed architecture achieves a similar rate distortion performance with a coding complexity reduction of 50%. Glenn Herrou, Wassim Hamidouche, Luce Morin |
DCC | 3 |
| 2018 | Low Complexity Joint RDO of Prediction Units Couples for HEVC Intra CodingabstractHEVC is the latest block-based video compression standard, outperforming H.264/AVC by 50% bitrate savings for the same perceptual quality. An HEVC encoder provides Rate-Distortion optimization coding tools for block-wise compression. Because of complexity limitations, Rate-Distortion Optimization (RDO) is usually performed independently for each block, assuming coding efficiency losses to be negligible. In this paper, we propose an acceleration solution for the Intra coding scheme named Dual-JRDO, which takes advantage of Inter-Block dependencies related to both predictive coding and CABAC. The Dual-JRDO improves Intra coding efficiency at the expense of higher computational complexity. The acceleration of the Dual-JRDO scheme includes adaptive use of the Dual-JRDO model based on source analysis, short-listing and early decisions strategies. The proposed Fast Dual-JRDO reduces the original model complexity by 89.54%, while providing tractable computation for average R-D gains of -0.45% (up to -0.82%) in the HM16.12 reference software model. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
ICASSP | 5 |
| 2018 | Temporal Adaptive Quantization using Accurate Estimations of Inter and Skip ProbabilitiesabstractHybrid video coding systems use spatial and temporal predictions in order to remove redundancies within the video source signal. These predictions create coding-scheme-related dependencies, often neglected for sake of simplicity. The R-D Spatio-Temporal Adaptive Quantization (RDSTQ) solution uses such dependencies to achieve better coding efficiency. It models the temporal distortion propagation by estimating the probability of a Coding Unit (CU) to be Inter coded. Uased on this probability, each CU is given a weight depending on its relative importance compared to other CUs. However, the initial approach roughly estimates the Inter probability and does not take into account the Skip mode characteristics in the propagation. It induces important Target uitrate Deviation (TBD) compared to the reference target rate. This paper provides undeniable improvements of the original RDSTQ model in using a more accurate estimation of the Inter probability. Then a new analytical solution for local quantizers is obtained by introducing the Skip probability of a CU into the temporal distortion propagation model. The proposed solution brings -2.05% BD-BR gain in average over the RDSTQ at low rate, which corresponds to -13.54% BD-BR gain in average against no local quantization. Moreover, the TBD is reduced from 38% to 14%. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
PCS | 5 |
| 2018 | Wavelet Decomposition Pre-processing for Spatial Scalability Video Compression SchemeabstractScalable video coding enables to compress the video at different formats within a single layered bitstream. SHVC, the scalable extension of the High Efficiency Video Coding (HEVC) standard, enables x2 spatial scalability, among other additional features. The closed-loop architecture of the SHVC codec is based on the use of multiple instances of the HEVC codec to encode the video layers, which considerably increases the encoding complexity. With the arrival of new immersive video formats, like 4K, 8K, High Frame Rate (HFR) and 360° videos, the quantity of data to compress is exploding, making the use of high-complexity coding algorithms unsuitable. In this paper, we propose a lowcomplexity scalable coding scheme based on the use of a single HEVC codec instance and a wavelet-based decomposition as pre-processing. The pre-encoding image decomposition relies on well-known simple Discrete Wavelet Transform (DWT) kernels, such as Haar or Le Gall 5/3. Compared to SHVC, the proposed architecture achieves a similar rate distortion performance with a coding complexity reduction of 50%. Glenn Herrou, Wassim Hamidouche, Luce Morin |
PCS | 3 |
| 2018 | SC-IQA: Shift compensation based image quality assessment for DIBR-synthesized viewsabstractDepth-image-based-rendering (DIBR) has been used to generate the virtual views for Multi-view videos and Free-viewpoint videos. However, the quality assessment of DIBR-synthesized views is very challenging owing to the new types of distortions induced by inaccurate depth maps, dis-occlusions and image inpainting methods. There exist a large number of object shifts and geometric distortions in the synthesized view which the traditional 2D quality metrics may fail to assess. In this paper, we propose a shift compensation based image quality assessment metric (SC-IQA) for DIBR-synthesized views. Firstly, the global geometric shift is compensated roughly by an SURF + RANSAC homography approach. Then, a multi-resolution block matching method, which performs a more accurate matching, is used to precisely compensate the shift and penalize the local geometric distortion as well. In addition, a visual saliency map is also used as a weighting function. To calculate the final overall quality scores, only the worst blocks are utilized since the biggest distortions have the most effects on the overall perceptual quality. The results show that the proposed metric significantly outperforms the state-of-the-art synthesized view dedicated metrics and the conventional 2D IQA metrics. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
VCIP | 3 |
| 2018 | NIQSV+: A No-Reference Synthesized View Quality Assessment MetricabstractBenefiting from multi-view video plus depth and depth-image-based-rendering technologies, only limited views of a real 3-D scene need to be captured, compressed, and transmitted. However, the quality assessment of synthesized views is very challenging, since some new types of distortions, which are inherently different from the texture coding errors, are inevitably produced by view synthesis and depth map compression, and the corresponding original views (reference views) are usually not available. Thus the full-reference quality metrics cannot be used for synthesized views. In this paper, we propose a novel no-reference image quality assessment method for 3-D synthesized views (called NIQSV+). This blind metric can evaluate the quality of synthesized views by measuring the typical synthesis distortions: blurry regions, black holes, and stretching, with access to neither the reference image nor the depth map. To evaluate the performance of the proposed method, we compare it with four full-reference 3-D (synthesized view dedicated) metrics, five full-reference 2-D metrics, and three no-reference 2-D metrics. In terms of their correlations with subjective scores, our experimental results show that the proposed no-reference metric approaches the best of the state-of-the-art full reference and no-reference 3-D metrics; and outperforms the widely used no-reference and full-reference 2-D metrics significantly. In terms of its approximation of human ranking, the proposed metric achieves the best performance in the experimental test. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
IEEE Trans. Image Process. | 3 |
| 2017 | Inter-block dependencies consideration for intra coding in H.264/AVC and HEVC standardsabstractRecent MPEG video compression standards are still block-based: blocks of pixels are sequentially coded using spatial or temporal prediction schemes. For each block, a vector of coding parameters has to be selected. In order to limit the complexity of this decision, independence between blocks is assumed, and coding parameters are locally optimized to maximize the coding efficiency. Few studies have investigated the benefits of inter-block dependencies consideration using Joint Rate-Distortion Optimization (JRDO), especially in Intra coding. To the best of our knowledge, maximum achievable gains of such approaches have never been exhibited. In this paper, we propose two JRDO models performing joint optimization of multiple blocks applied to intra prediction mode decision. The proposed models have been evaluated in both H.264/AVC and HEVC standards. These two models enables a bitrate saving with respect to the classical RDO model up to -3.10% and -2.31% in H.264/AVC and HEVC, respectively. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
ICASSP | 5 |
| 2017 | NIQSV: A no reference image quality assessment metric for 3D synthesized viewsabstractThe popularity of 3D applications, such as Free View-point TV (FTV) and Multi-view Video plus Depth (MVD), induces a heavy requirement of synthesized views. However, the quality assessment of synthesized views is very challenging because the corresponding original views (reference views) are usually not available at both encoder and decoder sides. In this paper, we propose a new no-reference quality assessment model to evaluate the quality of 3D synthesized views, called NIQSV (No-reference Image Quality assessment of Synthesized Views). This metric is based on the hypothesis that a good quality image is composed of flat areas (objects) separated by sharp edges, and the quality estimation involves only a set of simple morphological operators. NIQSV integrates the distortions of all the components, and then uses an edge image to weight the final distortions since the distortions of synthesized views mainly happen around object edges. The experimental results show that the proposed metric outperforms traditional 2D metrics and ranks among the best of dedicated 3D synthesized and full reference metrics. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
ICASSP | 3 |
| 2017 | An automatized method to parameterize embedded stereo matching algorithms
Judicael Menant, Guillaume Gautier, Muriel Pressigout, Luce Morin, Jean-François Nezan |
J. Syst. Archit. | 4 |
| 2016 | Multi-reference combinatorial strategy towards longer long-term dense motion estimation
Pierre-Henri Conze, Philippe Robert, Tomás Crivelli, Luce Morin |
Comput. Vis. Image Underst. | 4 |
| 2015 | Complex modulation computer-generated hologram by a fast hybrid point-source/wave-field approachabstractWe propose a fast Computer-Generated Hologram (CGH) computation method based on a hybrid point-source/wave-field approach. Whereas previously proposed methods tried to reduce the computational complexity of the point-source or the wave-field approaches independently, our method uses the two approaches together and therefore takes advantages from both of them. The algorithm consists of three steps. First, the 3D scene is sliced into several depth layers parallel to the hologram plane. Then, for each layer, we compute the complex wave scattered by this layer either using a wave-field or a point-source approach according to a threshold criterion on the number of points within the layer. Finally, we sum up the complex waves scattered by all the depth layers in order to obtain the final CGH. Experimental results reveal that this combination of approaches does not produce any visible artifact and outperforms both the point-source and wave-field approaches. Antonin Gilles, Patrick Gioia, Rémi Cozot, Luce Morin |
ICIP | 4 |
| 2015 | A framework for view-dependent hologram representation and adaptive reconstructionabstractIn this paper, we present a framework for networked hologram adaptive transmission. Depending on the observer position, a subset of hologram data is transmitted and displayed. Wavelet decomposition and pruning of wavelet coefficients according to the user position allows local diffractive pattern extraction. The proposed framework has been validated on an experimental set-up involving a kinect sensor for viewer position estimation. Kartik Viswanathan, Patrick Gioia, Luce Morin |
ICIP | 3 |
| 2013 | Dense motion estimation between distant frames: Combinatorial multi-step integration and statistical selectionabstractAccurate estimation of dense point correspondences between two distant frames of a video sequence is a challenging task. To address this problem, we present a combinatorial multistep integration procedure which allows one to obtain a large set of candidate motion fields between the two distant frames by considering multiple motion paths across the video sequence. Given this large candidate set, we propose to perform the optimal motion vector selection by combining a global optimization stage with a new statistical processing. Instead of considering a selection only based on intrinsic motion field quality and spatial regularization, the statistical processing exploits the spatial distribution of candidates and introduces an intra-candidate quality based on forward-backward consistency. Experiments evaluate the effectiveness of our method for distant motion estimation in the context of video editing. Pierre-Henri Conze, Tomás Crivelli, Philippe Robert, Luce Morin |
ICIP | 4 |
| 2013 | Video/GIS registration system based on skyline matching methodabstractFor applications such as outdoor Augmented Reality (AR) or 3D City Model construction of updating, a real-time registration between a video sequence and a Geographic Information System (GIS) is required. In this work, we present a registration system using a GPS prior and a skyline matching method to estimate the camera pose from images. A skyline rectification method through vertical vanishing point detection is also developed to enable arbitrary camera orientation, with non null tilt. The proposed approach is robust, fully automatic and computationally inexpensive which makes it a possible solution for mobile device applications. The performance of our approach is demonstrated in several experimental evaluations. The possible refinement of camera location through a systematic method is also investigated. Luce Morin, Muriel Pressigout, Guillaume Moreau, Myriam Servières |
ICIP | 2 |
| 2012 | An edge-based structural distortion indicator for the quality assessment of 3D synthesized viewsabstract3D-TV applications require the generation of novel viewpoints through Depth-Image-Based-Rendering methods. These synthesized views need to be assessed by a reliable quality metric. Most of the proposed metrics are inspired from 2D commonly used quality metrics. Yet, the latter were originally designed to address 2D compression distortions which are different from the distortions related to DIBR processes. We propose an edge-based method that indicates the level of structural degradation in the synthesized image. The first results are encouraging since the correlation to subjective scores is higher than other tested metrics. Emilie Bosc, Patrick Le Callet, Luce Morin, Muriel Pressigout |
PCS | 3 |
| 2011 | Can 3D synthesized views be reliably assessed through usual subjective and objective evaluation protocols?abstractThis paper addresses the problem of evaluating virtual view synthesized images in the multi-view video context. As a matter of fact, view synthesis brings new types of distortion. The question refers to the ability of the traditional used objective metrics to assess synthesized views quality, considering the new types of artifacts. The experiments conducted to determine their reliability consist in assessing seven different view synthesis algorithms. Subjective and objective measurements have been performed. Results show that the most commonly used objective metrics can be far from human judgment depending on the artifact to deal with. Emilie Bosc, Martin Köppel, Romuald Pépion, Muriel Pressigout, Luce Morin, Patrick Ndjiki-Nya, Patrick Le Callet |
ICIP | 5 |
| 2011 | Object-based Layered Depth Images for improved virtual view synthesis in rate-constrained contextabstractLayered Depth Image (LDI) representations are attractive compact representations for multi-view videos. Any virtual viewpoint can be rendered from LDI by using view synthesis technique. However, rendering from classical LDI leads to annoying visual artifacts, such as cracks and disocclusions. Visual quality gets even worse after a DCT-based compression of the LDI, because of blurring effects on depth discontinuities. In this paper, we propose a novel object-based LDI representation, improving synthesized virtual views quality, in a rate-constrained context. Pixels from each LDI layer are reorganised to enhance depth continuity. Vincent Jantet, Christine Guillemot, Luce Morin |
ICIP | 3 |
| 2010 | Focus on visual rendering quality through content-based depth map codingabstractMulti-view video plus depth (MVD) data is a set of multiple sequences capturing the same scene at different viewpoints, with their associated per-pixel depth value. Overcoming this large amount of data requires an effective coding framework. Yet, a simple but essential question refers to the means assessing the proposed coding methods. While the challenge in compression is the optimization of the rate-distortion ratio, a widely used objective metric to evaluate the distortion is the Peak-Signal-to-Noise-Ratio (PSNR), because of its simplicity and mathematically easiness to deal with such purposes. This paper points out the problem of reliability, concerning this metric, when estimating 3D video codec performances. We investigated the visual performances of two methods, namely H.264/MVC and Locally Adaptive Resolution (LAR) method, by encoding depth maps and reconstructing existing views from those degraded depth images. The experiments revealed that lower coding efficiency, in terms of PSNR, does not imply a lower rendering visual quality and that LAR method preserves the depth map properties correctly. Emilie Bosc, Muriel Pressigout, Luce Morin |
PCS | 3 |
| 2010 | A polygon soup representation for multiview coding
Thomas Colleu, Stéphane Pateux, Luce Morin, Claude Labit |
J. Vis. Commun. Image Represent. | 3 |
| 2007 | GPS, GIS and Video Registration for Building Reconstructionabstract3D reconstruction of urban environments is a widely studied subject since several years, as it can lead to many useful applications: virtual navigation, augmented reality, architectural planification, etc. One of the most difficult problem nowadays in this context is the acquisition and treatment of very large scale data if precise reconstruction is aimed. In this paper we present a system for computing geo-referenced positions and orientations of images of buildings from non calibrated videos. Providing such information is a mandatory step to well conditioned large scale and precise 3D reconstruction of urban areas. Our method is based on the registration of multimodal datasets, namely GPS measures, video sequences and rough 3D models of buildings. Gaël Sourimant, Luce Morin, Kadi Bouatouch |
ICIP (6) | 2 |
| 2007 | 3-D Model-Based Frame Interpolation for Distributed Video Coding of Static ScenesabstractThis paper addresses the problem of side information extraction for distributed coding of videos captured by a camera moving in a 3-D static environment. Examples of targeted applications are augmented reality, remote-controlled robots operating in hazardous environments, or remote exploration by drones. It explores the benefits of the structure-from-motion paradigm for distributed coding of this type of video content. Two interpolation methods constrained by the scene geometry, based either on block matching along epipolar lines or on 3-D mesh fitting, are first developed. These techniques are based on a robust algorithm for sub-pel matching of feature points, which leads to semi-dense correspondences between key frames. However, their rate-distortion (RD) performances are limited by misalignments between the side information and the actual Wyner-Ziv (WZ) frames due to the assumption of linear motion between key frames. To cope with this problem, two feature point tracking techniques are introduced, which recover the camera parameters of the WZ frames. A first technique, in which the frames remain encoded separately, performs tracking at the decoder and leads to significant RD performance gains. A second technique further improves the RD performances by allowing a limited tracking at the encoder. As an additional benefit, statistics on tracks allow the encoder to adapt the key frame frequency to the video motion content. Matthieu Maitre, Christine Guillemot, Luce Morin |
IEEE Trans. Image Process. | 3 |
| 2006 | 3D Scene Modeling for Distributed Video CodingabstractThe compression efficiency of distributed video-coding (DVC) suffers from the necessity of transmitting a large number of key-frames which are intra-coded. This paper describes a new 3D model-based DVC approach which reduces the key- frame frequency. The decoder first recovers a 3D model from the key-frames. It then predicts the intermediate frames by projecting it onto 2D image planes and applying image-based rendering techniques. This paper also introduces a new quasi-DVC method relying on a limited point tracking at the encoder. It greatly improves the prediction PSNR, while only slightly increasing the encoder complexity. It also allows the encoder to adaptively select the key-frames based on the video motion-content. Matthieu Maitre, Christine Guillemot, Luce Morin |
ICIP | 3 |
| 2006 | Scalable and Efficient Video Coding Using 3-D ModelingabstractIn this paper, we present a three-dimensional (3D) model-based video coding scheme for streaming static scene video in a compact way but also enabling time and spatial scalability according to network or terminal capability and providing 3D functionalities. The proposed format is based on encoding the sequence of reconstructed models using second-generation wavelets, and efficiently multiplexing the resulting geometric, topological, texture, and camera motion binary representations. The wavelets decomposition can be adaptive in order to fit to images and scene contents. To ensure time scalability, this representation is based on a common connectivity for all 3D models, which also allows straightforward morphing between successive models ensuring visual continuity at no additional cost. The method proves to be better than previous methods for video encoding of static scenes, even better than state-of-the-art video coders such as H264 (also known as MPEG AVC). Another application of our approach are smoothing camera path for suppression of jitter from hand-held acquisition and the fast transmission and real-time visualization of virtual environments obtained by video capture, for virtual or augmented reality and interactive walk-through in photo-realistic 3D environments around the original camera path Raphaèle Balter, Patrick Gioia, Luce Morin |
IEEE Trans. Multim. | 3 |
| 2005 | Time-evolving 3D model representation for scalable video codingabstractThis paper presents an efficient and scalable coding scheme for transmitting a stream of 3D models extracted from a video of a static scene. As in classical model-based video coding, the geometry, connectivity, and texture of the 3D models have to be transmitted, as well as the camera position for each frame in the original video. The proposed method is based on exploiting the interrelations existing between each type of information, instead of coding them independently, allowing a better prediction between the media and of the next information in the stream. Scalability is achieved through the use of wavelet-based representations for both texture and geometry of the models. A consistent connectivity is built for all 3D models extracted from the video sequence in order to get a consistent representation of the sequence possibly evolving in time. This allows a more compact representation and straightforward geometric morphing between successive time representations of the model. Furthermore this leads to consistent wavelet decomposition for 3D models in the stream. Targeted applications include distant visualization of the original video at very low bitrate and interactive navigation in the extracted 3D scene on heterogeneous terminals. Raphaèle Balter, Patrick Gioia, Luce Morin |
ICIP (1) | 3 |
| 2004 | 3D Models Coding and Morphing for Efficient Video Compression
Franck Galpin, Raphaèle Balter, Luce Morin, Koichiro Deguchi |
CVPR (1) | 3 |
| 2003 | One-dimensional dense disparity estimation for three-dimensional reconstructionabstractWe present a method for fully automatic three-dimensional (3D) reconstruction from a pair of weakly calibrated images in order to deal with the modeling of complex rigid scenes. A two-dimensional (2D) triangular mesh model of the scene is calculated using a two-step algorithm mixing sparse matching and dense motion estimation approaches. The 2D mesh is iteratively refined to fit any arbitrary 3D surface. At convergence, each triangular patch corresponds to the projection of a 3D plane. The proposed algorithm relies first on a dense disparity field. The dense field estimation modelized within a robust framework is constrained by the epipolar geometry. The resulting field is then segmented according to homographic models using iterative Delaunay triangulation. In association with a weak calibration and camera motion estimation algorithm, this 2D planar model is used to obtain a VRML-compatible 3D model of the scene. Lionel Oisel, Étienne Mémin, Luce Morin, Franck Galpin |
IEEE Trans. Image Process. | 3 |
| 2001 | New disparity map estimation using higher order statisticsabstractThis paper presents a new algorithm of disparity map estimation. The originality of this method lies in the process of dense disparity map estimation using dynamic programming constrained by interest points and using the higher order statistics (HOS) criteria for matching noisy images. Experiments with noisy real images have validated our method and have clearly shown the improvement over the existing ones. The dense disparity map obtained is more reliable when compared to the similar second-order statistics (SOS)-based dynamic programming and HOS-based correlation methods. Mohammed Rziza, Driss Aboutajdine, Luce Morin, Ahmed Tamtaoui |
ICASSP | 3 |
| 2001 | Computed 3D models for very low bit-rate video coding
Franck Galpin, Luce Morin |
VCIP | 2 |
| 2000 | Geometric Driven Optical Flow Estimation and Segmentation for 3D Reconstruction
Lionel Oisel, Étienne Mémin, Luce Morin |
ECCV (2) | 3 |
| 2000 | Estimation and segmentation of a dense disparity map for 3D reconstructionabstractThis paper presents a new algorithm of disparity map segmentation in planar facets. The origins of this method lie in the process of dense disparity map estimation, using the dynamic programming subject to interest points previously extracted. The segmentation of this map uses the normal vector at each pixel surface. The matching of pixels between the two images by dynamic programming provides us with a scattered disparity map. So the densification of this map is achieved by matching contour points extracted between the two available images. Experiments with real images have validated our method and have clearly shown the improvement over the existing methods. The dense disparity map obtained is reliable when compared to classical methods. We also get a normal vector map segmented in contours and in homogeneous regions reflecting 3D planar facets. Mohammed Rziza, Ahmed Tamtaoui, Luce Morin, Driss Aboutajdine |
ICASSP | 3 |
| 2000 | Video Coding Using Streamed 3D RepresentationabstractWe present a global scheme for encoding/decoding natural video sequences with partial 3D models. This technique is based on a robust estimation of a constrained dense field and on the estimation of projection matrices for reference images. Then a sequential encoding of the sequence is performed in order to produce streamed 3D models. The paper presents the global scheme and focuses on the most original steps of the encoding method. Some results on real video sequences are given to validate this approach. Franck Galpin, Luce Morin |
ICIP | 2 |
| 1998 | Planar Facets Segmentation using a Multiresolution Dense Disparity Field EstimationabstractIn the present paper we propose a new algorithm for planar facets segmentation of sequences of uncalibrated images in order to recover 3D models of complex scenes. This is performed using a two-step algorithm. First a robust and regularized dense disparity field is computed under the epipolar geometry constraint. The resulting field is segmented according to homographic models using an iterative Delaunay triangulation. Lionel Oisel, Luce Morin, Étienne Mémin, Claude Labit |
ICIP (2) | 2 |
| 1998 | Using geometric properties for automatic object positioning
Boubakeur Boufama, Roger Mohr, Luce Morin |
Image Vis. Comput. | 3 |
| 1996 | Semi-local projective invariants for the recognition of smooth plane curves
Stefan Carlsson, Roger Mohr, Theo Moons, Luce Morin, Charlie Rothwell, Marc Van Diest, Luc Van Gool, Francoise Veillon, Andrew Zisserman |
Int. J. Comput. Vis. | 4 |
| 1991 | Relative positioning from geometric invariantsabstractThe author gives geometric constructive solutions for 3-D vision problems like positioning a point in space from two views. Using reference points in the scene, no calibration is needed. The method involves only simple geometric computation. From the experiments it is concluded that positioning 3-D points relatively to reference points, is easy and provides more reliable results than absolute positioning as is usually done.> Roger Mohr, Luce Morin |
CVPR | 2 |