VLDB 2026 Research / reviewers in the wild / expert
Touradj Ebrahimi
dblp:72/6965
· DBLP profile ↗
206ranked-venue papers
16as first author
27since 2021 · last 2025
0000-0002-9900-3687ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 179 · 13 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 22 · 5 since 2021Artificial intelligence and machine learning · 14 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 first-authorSecurity and privacy · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fine-Grained Subjective Visual Quality Assessment for High-Fidelity Compressed ImagesabstractAdvances in image compression, storage, and display technologies have made high-quality images and videos widely accessible. At this level of quality, distinguishing between compressed and original content becomes difficult, highlighting the need for assessment methodologies that are sensitive to even the smallest visual quality differences. Conventional subjective visual quality assessments often use absolute category rating scales, ranging from “excellent” to “bad”. While suitable for evaluating more pronounced distortions, these scales are inadequate for detecting subtle visual differences. The JPEG standardization project AIC is currently developing a subjective image quality assessment methodology for high-fidelity images. This paper presents the proposed assessment methods, a dataset of high-quality compressed images, and their corresponding crowdsourced visual quality ratings. It also outlines a data analysis approach that reconstructs quality scale values in just noticeable difference (JND) units. The assessment method uses boosting techniques on visual stimuli to help observers detect compression artifacts more clearly. This is followed by a rescaling process that adjusts the boosted quality values back to the original perceptual scale. This reconstruction yields a fine-grained, high-precision quality scale in JND units, providing more informative results for practical applications. The dataset and code to reproduce the results will be available at https://github.com/jpeg-aic/dataset-BTC-PTC-24. Michela Testolina, Mohsen Jenadeleh, Shima Mohammadi, Shaolin Su, João Ascenso, Touradj Ebrahimi, Jon Sneyers, Dietmar Saupe |
DCC | 6 |
| 2025 | Error Correction for DNA-Based Image StorageabstractDNA has been proposed as an alternative support for data storage due to its lower energy consumption and longer lifespan when compared to conventional techniques. DNA-based storage faces several challenges, particularly in managing errors that arise during synthesis and sequencing. The JPEG Committee has been working towards a DNA-based image coding standard and has proposed an initial prototype without including error-correction mechanisms. Inserted in the context of the JPEG DNA standardization, this paper presents two joint error-correcting pipelines for DNA-based image storage combining cyclic redundancy checks, Reed-Solomon codes, convolutional codes, and Raptor codes. The proposed pipelines are evaluated across various channel models and benchmarked against state-of-the-art, demonstrating enhanced coding performance, reduced complexity, and improved flexibility. The efficacy of the proposed systems has resulted in their inclusion in the JPEG DNA verification model. The results presented in this paper can serve as a reference to benchmark against future ideas on joint source-channel coding of images in DNA. Davi Lazzarotto, Michela Testolina, Touradj Ebrahimi |
ICIP | 3 |
| 2025 | Pose-Invariant Face Recognition via Feature-Space Pose FrontalizationabstractPose-invariant face recognition has become a challenging problem for modern AI-based face recognition systems. It aims at matching a profile face captured in the wild with a frontal face registered in a database. Existing methods perform face frontalization via either generative models or learning a pose robust feature representation. In this paper, a new method is presented to perform face frontalization and recognition within the feature space. First, a novel feature space pose frontalization module (FSPFM) is proposed to transform profile images with arbitrary angles into frontal counterparts. Second, a new training paradigm is proposed to maximize the potential of FSPFM and boost its performance. The latter consists of a pre-training and an attention-guided fine-tuning stage. Moreover, extensive experiments have been conducted on five popular face recognition benchmarks. Results show that not only our method outperforms the state-of-the-art in the pose-invariant face recognition task but also maintains superior performance in other standard scenarios. Nikolay Stanishev, Touradj Ebrahimi |
ICIP | 3 |
| 2025 | DNA-based image storage with substitution and indel error correction
Davi Lazzarotto, Frédéric Piguet, Touradj Ebrahimi |
PCS | 3 |
| 2025 | Fine-Grained HDR Image Quality Assessment From Noticeably Distorted to Very High FidelityabstractHigh dynamic range (HDR) and wide color gamut (WCG) technologies significantly improve color reproduction compared to standard dynamic range (SDR) and standard color gamuts, resulting in more accurate, richer, and more immersive images. However, HDR increases data demands, posing challenges for bandwidth efficiency and compression techniques. Advances in compression and display technologies require more precise image quality assessment, particularly in the high-fidelity range where perceptual differences are subtle. To address this gap, we introduce AIC-HDR2025, the first such HDR dataset, comprising 100 test images generated from five HDR sources, each compressed using four codecs at five compression levels. It covers the high-fidelity range, from visible distortions to compression levels below the visually lossless threshold. A subjective study was conducted using the JPEG AIC-3 test methodology, combining plain and boosted triplet comparisons. In total, 34,560 ratings were collected from 151 participants across four fully controlled labs. The results confirm that AIC-3 enables precise HDR quality estimation, with 95% confidence intervals averaging a width of 0.27 at 1 JND. In addition, several recently proposed objective metrics were evaluated based on their correlation with subjective ratings. The dataset is publicly available1. Mohsen Jenadeleh, Jon Sneyers, Davi Lazzarotto, Shima Mohammadi, Dominik Keller, Atanas Boev, Rakesh Rao Ramachandra Rao, António M. G. Pinheiro, Thomas Richter 0005, Alexander Raake, Touradj Ebrahimi, João Ascenso, Dietmar Saupe |
QoMEX | 11 |
| 2024 | XAIface: A Framework and Toolkit for Explainable Face RecognitionabstractArtificial intelligence-based face recognition solutions are becoming increasingly popular. Therefore, it is crucial to fully understand and explain how these technologies work in order to make them more effective and acceptable to society. This is the goal of the CHIST-ERA project XAIface, the final results of which are reported in this article: a framework and toolkit for improving AI decision explainability, in the context of automated face recognition, through several novel methods are presented. These methods are integrated into an end-to-end face recognition demonstrator system, which facilitates studying the impact of various influencing factors and system processes on recognition performance. By doing so, we can visually explain the decisions made by the face verification pipeline for specific instances in our test set using heatmaps and locally interpretable features. Furthermore, we offer a comprehensive explanation of the end-to-end model by examining the relationship between verification failures and misclassifications of soft biometric facial traits. Nélida Mirabet-Herranz, Naima Bousnina, Jonas Pfister, Chiara Galdi, Jean-Luc Dugelay, Werner Bailer, Touradj Ebrahimi, Paulo Lobato Correia, Fernando Pereira 0001, Felix Schmautzer, Erich Schweighofer |
CBMI | 9 |
| 2024 | Explainable Face Verification via Feature-Guided Gradient BackpropagationabstractRecent years have witnessed significant advancement in face recognition (FR) techniques, with their applications impacting people's lives including in security-sensitive areas. There is a growing need for reliable interpretation of decisions of such systems. Existing studies relying on various mechanisms have investigated the usage of saliency maps as an explanation approach, but suffer from different limitations. This paper first explores the spatial relationship between face image and its deep representation via gradient backpropagation. Then a new explanation approach called Feature-Guided Gradient Backpropagation (FGGB) has been conceived, which provides precise and insightful similarity and dissimilarity saliency maps to explain the “Accept” and “Reject” decision of an FR system. Extensive visual presentation and quantitative measurement have shown that FGGB achieves comparable results in similarity maps and superior performance in dissimilarity maps when compared to current state-of-the-art explainable face verification approaches. Zewei Xu, Touradj Ebrahimi |
FG | 3 |
| 2024 | Temporal Conditional Coding for Dynamic Point Cloud Geometry CompressionabstractPoint clouds allow for the representation of 3D multimedia content as a set of disconnected points in space. Their inherent irregular geometric nature poses a challenge to efficient compression, a critical operation for both storage and transmission. This paper proposes a VAE-inspired codec tailored for dynamic point cloud geometry compression, taking advantage of a temporal autoregressive hyperprior to enhance compression performance. Specifically, features derived from adjacent point cloud frames help build a hyperprior for conditional entropy coding. Sparse convolutions are leveraged to reach higher computational efficiency when compared to 3D dense convolutions. Remarkably, the proposed approach achieves an average 60.2% BD-rate gain against the contemporary V-PCC compression standard from MPEG. Davi Lazzarotto, Touradj Ebrahimi |
ICASSP | 3 |
| 2024 | An International Standard For Assessing Trustworthiness In MediaabstractThe proliferation of synthetic media generation technologies, such as generative AI, has led to a surge of media content generation and consumption. While this progress opens new opportunities, especially in creative industries, it also causes challenges, including piracy, fake media distribution, and concerns about trust and privacy. In the creative sector, media modifications are often part of the production pipelines and in many application domains, creators need or want to declare the type of modifications that were performed on the media asset. The cryptographically signed association of provenance information with the media asset itself provides a trust link between the owner or editor of a media asset and its consumers. The absence of such assertions may reveal the lack of trustworthiness in media assets or worse, the intention to hide the existence of manipulations. This paper describes the JPEG Trust framework (ISO/IEC 21617) that aims to establish trust in digital media creation, modification, annotation, distribution and consumption. The framework provides standardized protocols to extract indicators to assess trustworthiness, means to annotate media provenance, and securely link the assets and associated annotations together. Deepayan Bhowmik, Sabrina B. Caldwell, Jaime Delgado, Touradj Ebrahimi, Nikolaos Fotos, Xiaojun Gu, Ziyuan Hu, Xin Kang 0001, Fernando Pereira 0001, Leonard Rosenthol, Frederik Temmermans |
ICIP | 4 |
| 2024 | Towards the Detection of AI-Synthesized Human Face ImagesabstractOver the past years, image generation and manipulation have achieved remarkable progress due to the rapid development of generative AI based on deep learning. Recent studies have devoted significant efforts to address the problem of face image manipulation caused by deepfake techniques. However, the problem of detecting purely synthesized face images has been explored to a lesser extent. In particular, the recent popular Diffusion Models (DMs) have shown remarkable success in image synthesis. Existing detectors struggle to generalize between synthesized images created by different generative models. In this work, a comprehensive benchmark including human face images produced by Generative Adversarial Networks (GANs) and a variety of DMs has been established to evaluate both the generalization ability and robustness of state-of-the-art detectors. Then, the forgery traces introduced by different generative models have been analyzed in the frequency domain to draw various insights. The paper further demonstrates that a detector trained with frequency representation can generalize well to other unseen generative models. Touradj Ebrahimi |
ICIP | 2 |
| 2024 | Assessing objective quality metrics for JPEG and MPEG point cloud codingabstractAs applications using immersive media gained increased attention from both academia and industry, research in the field of point cloud compression has greatly intensified in recent years, leading to the development of the MPEG compression standards V-PCC and G-PCC, as well as the more recent JPEG Pleno learning-based point cloud coding. Each of the standards mentioned above is based on a different algorithm, introducing distinct types of degradation that may impair the quality of experience when lossy compression is applied. Although the impact on perceptual quality can be accurately evaluated during subjective quality assessment experiments, objective quality metrics also predict the visually perceived quality and provide similarity scores without human intervention. Nevertheless, their accuracy can be susceptible to the characteristics of the evaluated media as well as to the type and intensity of the added distortion. While the performance of multiple state-of-the-art objective quality metrics has already been evaluated through their correlation with subjective scores obtained in the presence of artifacts produced by the MPEG standards, no study has evaluated how metrics perform with the more recent JPEG Pleno point cloud coding. In this paper, a study is conducted to benchmark the performance of a large set of objective quality metrics in a subjective dataset including distortions produced by JPEG and MPEG codecs. The dataset also contains three different trade-offs between color and geometry compression for each codec, adding another dimension to the analysis. Performance indexes are computed over the entire dataset but also after splitting according to the codec and to the original model, resulting in detailed insights about the overall performance of each visual quality predictor as well as their cross-content and cross-codec generalization ability. Davi Lazzarotto, Michela Testolina, Touradj Ebrahimi |
QoMEX | 3 |
| 2024 | Towards Visual Saliency Explanations of Face VerificationabstractIn the past years, deep convolutional neural networks have been pushing the frontier of face recognition (FR) techniques in both verification and identification scenarios. Despite the high accuracy, they are often criticized for lacking explainability. There has been an increasing demand for understanding the decision-making process of deep face recognition systems. Recent studies have investigated the usage of visual saliency maps as an explanation, but they often lack a discussion and analysis in the context of face recognition. This paper concentrates on explainable face verification tasks and conceives a new explanation framework. Firstly, a definition of the saliency-based explanation method is provided, which focuses on the decisions made by the deep FR model. Secondly, a new model-agnostic explanation method named CorrRISE is proposed to produce saliency maps, which reveal both the similar and dissimilar regions of any given pair of face images. Then, an evaluation methodology is designed to measure the performance of general visual saliency explanation methods in face verification. Finally, substantial visual and quantitative results have shown that the proposed CorrRISE method demonstrates promising results in comparison with other state-of-the-art explainable face verification approaches. Zewei Xu, Touradj Ebrahimi |
WACV | 3 |
| 2024 | PRO-Face C: Privacy-Preserving Recognition of Obfuscated Face via Feature CompensationabstractThe advancement of face recognition technology has delivered substantial societal advantages. However, it has also raised global privacy concerns due to the ubiquitous collection and potential misuse of individuals’ facial data. This presents a notable paradox: while there is a societal demand for a robust face recognition ecosystem to ensure public security and convenience, an increasing number of individuals are hesitant to release their facial data. Numerous studies have endeavored to find such a utility-privacy trade-off, yet many struggle with the dilemma of prioritizing one at the expense of the other. In response to this challenge, this paper proposes PRO-Face C, a novel paradigm for privacy-preserving recognition of obfuscated faces via a dedicated feature compensation mechanism, aimed at optimizing the equilibrium between privacy preservation and utility maximization. The proposed approach is characterized by a specialized client-server architecture: the client transmits only obfuscated images to the server, which then performs identity recognition using a pre-trained model in conjunction with a suite of privacy-free complementary features. This framework facilitates accurate face identification while safeguarding the original facial appearance from explicit disclosure. Furthermore, the obfuscated image retains its visualization capability, crucial for image preview functionalities. To ensure the desired properties, we have developed an identity-guided feature compensation mechanism, complemented by several privacy-enhancing techniques. Extensive experiments conducted across multiple face datasets underscore the effectiveness of the proposed approach in diverse scenarios. Lin Yuan 0002, Xiao Pu 0002, Yan Zhang 0108, Yushu Zhang 0001, Xinbo Gao 0001, Touradj Ebrahimi |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2023 | Event Data Stream Compression Based on Point Cloud RepresentationabstractThe Dynamic Vision System (DVS) is a novel image acquisition system that works only when there is a brightness change in a pixel, resulting in a stream of events including timestamps, spatial coordinates and the sign of the brightness change (increase or decrease). Although DVS’s output data size is much smaller than conventional image systems, it still requires further compression, as the main applications of DVS are embedded systems with limited transmission and storage resources. In this paper, we propose a new method for lossless compression of event data streams based on point cloud representations. The event data stream is organized into a 3D point cloud to which a compression algorithm is applied. In addition, different generation strategies are devised in order to compare the compression performance of the proposed approach. Experimental results show an improved compression ratio of about 22% under lossless conditions. Touradj Ebrahimi |
ICIP | 2 |
| 2023 | On the Performance of Subjective Visual Quality Assessment Protocols for Nearly Visually Lossless Image CompressionabstractThe past decades have witnessed rapid growth in imaging as a major form of communication between individuals. Due to recent advances in capture, storage, delivery and display technologies, consumers demand improved perceptual quality while requiring reduced storage. In this context, research and innovation in lossy image compression have steered towards methods capable of achieving high compression ratios without compromising the perceived visual quality of images, and in some cases even enhancing the latter. Subjective visual quality assessment of images plays a fundamental role in defining quality as perceived by human observers. Although the field of image compression is constantly evolving towards efficient solutions for higher visual qualities, standardized subjective visual quality assessment protocols are still limited to those proposed in ITU-R Recommendation BT.500 and JPEG AIC standards. The number of comprehensive and in-depth studies where different protocols are compared is still insufficient. Moreover, previous works have not investigated the effectiveness of these methods on higher quality ranges, using recent image compression methods. In this paper, subjective visual scores collected from three subjective image quality assessment protocols, namely the Double Stimulus Continuous Quality Scale (DSCQS) and two test methods described in the JPEG AIC Part 2 standard, are compared between different laboratories under similar controlled conditions. The analysis of the experimental results has revealed that the DSCQS protocol is highly influenced by the quality of the reference images and experience of the subjects, while the JPEG AIC Part 2 specifications produce more stable results but are expensive and only suitable for a limited range of qualities. These emphasize the need for new robust subjective image quality assessment methodologies able to discriminate in the range of qualities generally demanded by consumers, i.e. from high to nearly visually lossless. Michela Testolina, Davi Lazzarotto, Rafael Rodrigues, Shima Mohammadi, João Ascenso, António M. G. Pinheiro, Touradj Ebrahimi |
ACM Multimedia | 7 |
| 2023 | Towards a Multiscale Point Cloud Structural Similarity MetricabstractPoint clouds are effective data structures for the representation of three-dimensional media and hence adopted in a wide range of practical applications. In many cases, the portrayed data is expected to be visualized by humans. After acquisition, point clouds may undergo different processing operations such as compression or denoising, potentially affecting their perceived quality. Although subjective experiments are still the most reliable form of assessing the intensity of degradation, they are expensive and time-consuming, pushing many systems to depend on objective metrics. Such algorithms are used to model the human visual system, and their performance is usually assessed through their correlation with subjective visual quality scores. In this paper, an objective quality metric capable of evaluating distortions between a reference and a distorted point cloud at multiple scales is presented. The proposed metric is based on the point cloud structural similarity metric (PointSSIM), which computes a score based on the difference between statistical estimators obtained on the distribution of the luminance attribute over local neighborhoods. A collection of PointSSIM scores is produced for multiple scales obtained through the voxelization of both models at different bit depth precisions. These scores are then pooled through a weighted sum, with the importance of each scale being defined through logistic fitting to subjective mean opinion scores, producing one MS-PointSSIM score. Three datasets were employed for fitting and performance assessment, demonstrating a clear advantage of the proposed metric when compared to the single-scale baseline. Moreover, the presented MS-PointSSIM is shown to be the best predictor according to the average Pearson correlation coefficient across the three datasets when compared to state-of-the-art metrics. Davi Lazzarotto, Touradj Ebrahimi |
MMSP | 2 |
| 2023 | Discriminative Deep Feature Visualization for Explainable Face RecognitionabstractDespite the huge success of deep convolutional neural networks in face recognition (FR) tasks, current methods lack explainability for their predictions because of their “black-box” nature. In recent years, studies have been carried out to give an interpretation of the decision of a deep FR system. However, the affinity between the input facial image and the extracted deep features has not been explored. This paper contributes to the problem of explainable face recognition by first conceiving a face reconstruction-based explanation module, which reveals the correspondence between the deep feature and the facial regions. To further interpret the decision of an FR model, a novel visual saliency explanation algorithm has been proposed. It provides insightful explanation by producing visual saliency maps that represent similar and dissimilar regions between input faces. A detailed analysis has been presented for the generated visual explanation to show the effectiveness of the proposed method. Zewei Xu, Touradj Ebrahimi |
MMSP | 3 |
| 2023 | JPEG AIC-3 Dataset: Towards Defining the High Quality to Nearly Visually Lossless Quality RangeabstractVisual data play a crucial role in modern society, and the rate at which images and videos are acquired, stored, and exchanged every day is rapidly increasing. Image compression is the key technology that enables storing and sharing of visual content in an efficient and cost-effective manner, by removing redundant and irrelevant information. On the other hand, image compression often introduces undesirable artifacts that reduce the perceived quality of the media. Subjective image quality assessment experiments allow for the collection of information on the visual quality of the media as perceived by human observers, and therefore quantifying the impact of such distortions. Nevertheless, the most commonly used subjective image quality assessment methodologies were designed to evaluate compressed images with visible distortions, and therefore are not accurate and reliable when evaluating images having higher visual qualities. In this paper, we present a dataset of compressed images with quality levels that range from high to nearly visually lossless, with associated quality scores in JND units. The images were subjectively evaluated by expert human observers, and the results were used to define the range from high to nearly visually lossless quality. The dataset is made publicly available to researchers, providing a valuable resource for the development of novel subjective quality assessment methodologies or compression methods that are more effective in this quality range. Michela Testolina, Vlad Hosu, Mohsen Jenadeleh, Davi Lazzarotto, Dietmar Saupe, Touradj Ebrahimi |
QoMEX | 6 |
| 2023 | Cross-resolution Face Recognition via Identity-Preserving Network and Knowledge DistillationabstractCross-resolution face recognition has become a challenging problem for modern deep face recognition systems. It aims at matching a low-resolution probe image with high-resolution gallery images registered in a database. Existing methods mainly leverage prior information from high-resolution images by either reconstructing facial details with super-resolution techniques or learning a unified feature space. To address this challenge, this paper proposes a new approach that enforces the network to focus on the discriminative information stored in the low-frequency components of a low-resolution image. A cross-resolution knowledge distillation paradigm is first employed as the learning framework. Then, an identity-preserving network, WaveResNet, and a wavelet similarity loss are designed to capture low-frequency details and boost performance. Finally, an image degradation model is conceived to simulate more realistic low-resolution training data. Consequently, extensive experimental results show that the proposed method consistently outperforms the baseline model and other state-of-the-art methods across a variety of image resolutions. Touradj Ebrahimi |
VCIP | 2 |
| 2022 | Latent Space Slicing for Enhanced Entropy Modeling In Learning-Based Point Cloud Geometry CompressionabstractThe growing adoption of point clouds as an imaging modality has stimulated the search for efficient solutions for compression. Learning-based algorithms have been reporting increasingly better performance and are drawing the attention from the research community and standardisation groups such as JPEG and MPEG. Learned autoencoder architectures based on 3D convolutional layers are popular solutions and have demonstrated higher performance when adopting latent space entropy modeling based on learned hyperpriors. We propose an enhanced entropy model that takes into account both the hyperprior and previously encoded latent features to estimate the mean and scale of compressed features. The obtained results show a large increase in performance, with a BD PSNR gain of 5.75dB when compared to the Octree coding module in G-PCC for the D2 PSNR metric. We also perform an ablation study to quantify the impact of network parameters in the performance of the model, drawing useful insights for future research. Nicolas Frank, Davi Lazzarotto, Touradj Ebrahimi |
ICASSP | 3 |
| 2022 | On the Impact of Spatial Rendering on Point Cloud Subjective Visual Quality AssessmentabstractImmersive imaging modalities have been receiving growing attention over the last years. In this context, point clouds demonstrated to be a competitive data representation format, mainly due to its compatibility with most acquisition devices such as LiDAR scans and depth cameras. On the other hand, the large associated data volume is a challenge to storage and transmission, resulting in the need for efficient point cloud compression methods. Multiple subjective studies have been conducted to assess the performance of such compression methods, mostly limiting the analysis to flat monitors or virtual and augmented reality headsets. In this paper, we investigate the impact of a novel eye-sensing light field display on several aspects of quality of experience, as well as on the subjective perception of compression artifacts. The two visualization devices have been observed to create distinct user experiences which lead to noticeable differences in subjective opinion. The advantages and disadvantages of each visualization strategy are underlined, based on rigorous statistical analysis. The subjectively annotated dataset is also released in order to foster future research. Davi Lazzarotto, Michela Testolina, Touradj Ebrahimi |
QoMEX | 3 |
| 2021 | On Block Prediction For Learning-Based Point Cloud CompressionabstractPoint clouds are among popular visual representations for immersive media. However, the vast amount of information generated during their acquisition requires effective compression for practical applications. Although relevant activities from standardization bodies have led to state-of-the-art compression using conventional methods, learning-based encoders have recently emerged as promising solutions with comparable performance while offering additional attractive features. Yet, there is still a large unexplored space for research that can lead to further advances. In this paper, we propose a block prediction module for bit-rate reduction of geometry-only point clouds. Our method exploits spatial redundancies at the decoding stage between block partitions in the point cloud, and predicts a query block using Generative Adversarial Networks. Results show performance improvements of the objective metrics at low bit-rates, after integration in a baseline auto-encoder architecture. Davi Lazzarotto, Evangelos Alexiou, Touradj Ebrahimi |
ICIP | 3 |
| 2021 | Open-Set Person Re-Identification Through Error Resilient Recurring Gallery BuildingabstractIn person re-identification, individuals must be correctly identified in images that come from different cameras or are captured at different points in time. In the open-set case, the above needs be achieved for people who have not been previously recognised. In this paper, we propose a universal method for building a multi-shot gallery of observed reference identities recurrently online. We perform L2-norm descriptor matching for gallery retrieval using descriptors produced by a generic closed-set re-identification system. Multi-shot gallery is continuously updated by replacing outliers with newly matched descriptors. Outliers are detected using the Isolation Forest algorithm, thus ensuring that the gallery is resilient to erroneous assignments, leading to improved re-identification results in the open-set case. Philine Witzig, Evgeniy Upenik, Touradj Ebrahimi |
ICIP | 3 |
| 2021 | Benchmarking of objective quality metrics for point cloud compressionabstractPoint cloud is a promising imaging modality for the representation of 3D media. The vast volume of data associated with it requires efficient compression solutions, with lossy algorithms leading to larger bit-rate savings at the expense of visual impairments. While conventional encoding approaches rely on efficient data structures, recent methods have incorporated deep learning for rate-distortion optimization, while inducing perceptual degradations of different natures. To measure the magnitude of such distortions, subjective or objective quality evaluation methodologies are employed. Lately, a remarkable amount of efforts has been devoted to the development of point cloud objective quality metrics, which have been reported to attain high prediction accuracy. However, their performance and generalization capabilities haven’t been evaluated yet in presence of artifacts from learning-based codecs. In this study, we tackle this matter by conducting the first crowdsourcing experiment for point cloud quality reported in the literature, in order to obtain subjective ratings for point cloud models whose topology and color attributes are encoded by both conventional and data-driven methods. Using the subjective scores as ground truth, the performance of a large pool of state-of-the-art quality metrics is rigorously benchmarked, drawing useful insights regarding their efficacy. Davi Lazzarotto, Evangelos Alexiou, Touradj Ebrahimi |
MMSP | 3 |
| 2021 | Performance Evaluation of Objective Image Quality Metrics on Conventional and Learning-Based Compression ArtifactsabstractLossy image compression is a popular, simple and effective solution to reduce the amount of data representing digital pictures. In most lossy compression methods, the reduced volume of data in bits is achieved at the expense of introducing visual artifacts in the picture. The perceptual quality impact of such artifacts can be assessed with expensive and time-consuming subjective image quality experiments or through objective image quality metrics. However, the faster and less resource demanding objective quality metrics are not always able to reliably predict the quality as perceived by human observers. In this paper, the performance of 14 objective image quality metrics is benchmarked against a dataset of compressed images labeled with their subjective quality scores. Moreover, the performance of the above objective quality metrics in predicting the subjective quality of images distorted by both conventional and learning-based lossy compression artifacts is assessed and conclusions are drawn. Michela Testolina, Evgeniy Upenik, João Ascenso, Fernando Pereira 0001, Touradj Ebrahimi |
QoMEX | 5 |
| 2021 | Large-Scale Crowdsourcing Subjective Quality Evaluation of Learning-Based Image CodingabstractLearning-based image codecs produce different compression artifacts, when compared to the blocking and blurring degradation introduced by conventional image codecs, such as JPEG, JPEG 2000 and HEIC. In this paper, a crowdsourcing based subjective quality evaluation procedure was used to benchmark a representative set of end-to-end deep learning-based image codecs submitted to the MMSP'2020 Grand Challenge on Learning-Based Image Coding and the JPEG AI Call for Evidence. For the first time, a double stimulus methodology with a continuous quality scale was applied to evaluate this type of image codecs. The subjective experiment is one of the largest ever reported including more than 240 pair-comparisons evaluated by 118 naïve subjects. The results of the benchmarking of learning-based image coding solutions against conventional codecs are organized in a dataset of differential mean opinion scores along with the stimuli and made publicly available. Evgeniy Upenik, Michela Testolina, João Ascenso, Fernando Pereira 0001, Touradj Ebrahimi |
VCIP | 5 |
| 2021 | JPEG XS - A New Standard for Visually Lossless Low-Latency Lightweight Image CodingabstractJoint Photographic Experts Group (JPEG) XS is a new International Standard from the JPEG Committee (formally known as ISO/International Electrotechnical Commission (IEC) JTC1/SC29/WG1). It defines an interoperable, visually lossless low-latency lightweight image coding that can be used for mezzanine compression within any AV market. Among the targeted use cases, one can cite video transport over professional video links (serial digital interface (SDI), internet protocol (IP), and Ethernet), real-time video storage, memory buffers, omnidirectional video capture and rendering, and sensor compression (for example, in cameras and the automotive industry). The core coding system is composed of an optional color transform, a wavelet transform, and a novel entropy encoder, processing groups of coefficients by coding their magnitude level and packing the magnitude refinement. Such a design allows for visually transparent quality at moderate compression ratios, scalable end-to-end latency that ranges from less than one line to a maximum of 32 lines of the image, and a low-complexity real-time implementation in application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), central processing unit (CPU), and graphics processing unit (GPU). This article details the key features of this new standard and the profiles and formats that have been defined so far for the various applications. It also gives a technical description of the core coding system. Finally, the latest performance evaluation results of recent implementations of the standard are presented, followed by the current status of the ongoing standardization process and future milestones. Antonin Descampe, Thomas Richter 0005, Touradj Ebrahimi, Siegfried Fößel, Joachim Keinert, Tim Bruylants, Pascal Pellegrin, Charles Buysschaert, Gaël Rouvroy |
Proc. IEEE | 3 |
| 2020 | Quality Evaluation Of Static Point Clouds Encoded Using MPEG CodecsabstractThis paper presents a quality evaluation study of point cloud codecs that have been recently standardised by the MPEG committee. In particular, a subjective experiment to assess their performance in terms of bitrate against visual quality is designed and realized in four independent laboratories. The experimental setup of each laboratory varies; yet, the obtained subjective scores exhibit high inter laboratory correlation, confirming that the adopted assessment protocol is robust to equipment selection and viewing conditions, ensuring reliability and facilitating repeatability. Our study confirms the superior compression performance of the MPEG V-PCC, when compared to MPEG G-PCC, in the case of static contents. Finally, results from a benchmark of the most popular objective quality metrics using the obtained subjective scores as ground truth, reveal that the point2plane with mean square error is the most accurate quality predictor, closely followed by the point2point also using mean square error as distance measure. Stuart W. Perry, Huy Phi Cong, Luís Alberto da Silva Cruz, João Prazeres, Manuela Pereira, António M. G. Pinheiro, Emil Dumic, Evangelos Alexiou, Touradj Ebrahimi |
ICIP | 9 |
| 2020 | PointXR: A Toolbox for Visualization and Subjective Evaluation of Point Clouds in Virtual RealityabstractIn this study, we explore the use of virtual reality to subjectively evaluate the visual quality of point cloud contents. To this aim, we develop the PointXR toolbox, a set of Unity applications that can host experiments under variants of interactive and passive evaluation protocols. An auxiliary tool to facilitate the configuration of the supported rendering schemes for point cloud visualization is provided as part of it. Our toolbox is employed to conduct two validating experiments in a virtual environment with 6 degrees of freedom. The purpose is to assess the performance of color encoders that are incorporated in the upcoming MPEG standard on point cloud compression. For this study, we convert a set of mesh models to point cloud contents, and form a high-quality cultural heritage repository, namely, PointXR dataset. A comparison between the adopted protocols and the codecs' performance is carried based on the ratings obtained from both experiments. Finally, interactivity patterns based on behavioral data that were recorded during the evaluations are extracted, and results are discussed. The PointXR toolbox, the PointXR dataset, and the experimental results are made publicly available. Evangelos Alexiou, Nanyang Yang, Touradj Ebrahimi |
QoMEX | 3 |
| 2019 | Inpainting in Omnidirectional Images for Privacy ProtectionabstractPrivacy protection is drawing more attention with the advances in image processing, visual and social media. Photo sharing is a popular activity, which also brings the concern of regulating permissions associated with shared content. This paper presents a method for protecting user privacy in omnidirectional media, by removing parts of the content selected by the user, in a reversible manner. Object removal is carried out using three different state-of-the-art inpainting methods, employed over the mask drawn in the viewport domain so that the geometric distortions are minimized. The perceived quality of the scene is assessed via subjective tests, comparing the proposed method against inpainting employed directly on the equirectangular image. Results on distinct contents indicate our object removal methodology on the viewport enhances perceived quality, thereby improves privacy protection as the user is able to hide objects with less distortion in the overall image. Evgeniy Upenik, Pinar Akyazi, Mehmet Tuzmen, Touradj Ebrahimi |
ICASSP | 4 |
| 2019 | Towards Modelling of Visual Saliency in Point Clouds for Immersive ApplicationsabstractModelling human visual attention is of great importance in the field of computer vision and has been widely explored for 3D imaging. Yet, in the absence of ground truth data, it is unclear whether such predictions are in alignment with the actual human viewing behavior in virtual reality environments. In this study, we work towards solving this problem by conducting an eye-tracking experiment in an immersive 3D scene that offers 6 degrees of freedom. A wide range of static point cloud models is inspected by human subjects, while their gaze is captured in real-time. The visual attention information is used to extract fixation density maps, that can be further exploited for saliency modelling. To obtain high quality fixation points, we devise a scheme that utilizes every recorded gaze measurement from the two eye-cameras of our set-up. The obtained fixation density maps together with the recorded gaze and head trajectories are made publicly available, to enrich visual saliency datasets for 3D models. Evangelos Alexiou, Peisen Xu, Touradj Ebrahimi |
ICIP | 3 |
| 2019 | Perceptual Quality Study on Deep Learning Based Image CompressionabstractRecently deep learning based image compression has made rapid advances with promising results based on objective quality metrics. However, a rigorous subjective quality evaluation on such compression schemes have rarely been reported. This paper aims at perceptual quality studies on learned compression. First, we build a general learned compression approach, and optimize the model. In total six compression algorithms are considered for this study. Then, we perform subjective quality tests in a controlled environment using high-resolution images. Results demonstrate learned compression optimized by MS-SSIM yields competitive results that approach the efficiency of state-of-the-art compression. The results obtained can provide a useful benchmark for future developments in learned image compression. Zhengxue Cheng, Pinar Akyazi, Heming Sun, Jiro Katto, Touradj Ebrahimi |
ICIP | 5 |
| 2019 | Saliency Driven Perceptual Quality Metric for Omnidirectional Visual ContentabstractThe problem of objectively measuring perceptual quality of omnidirectional visual content arises in many immersive imaging applications and particularly in compression. The interactive nature of this type of content limits the performance of earlier methods designed for static images or for video with a predefined dynamic. The non-deterministic impact must be addressed using statistical approach. One of the ways to describe, analyze and predict viewer interactions in omnidirectional imaging is through estimation of visual attention. We propose an objective metric to measure perceptual quality of omnidirectional visual content considering visual attention information. Evgeniy Upenik, Touradj Ebrahimi |
ICIP | 2 |
| 2019 | Light field compression using translation-assisted view estimationabstractLight field technology has recently been gaining traction in the research community. Several acquisition technologies have been demonstrated to properly capture light field information, and portable devices have been commercialized to the general public. However, new and efficient compression algorithms are needed to sensibly reduce the amount of data that needs to be stored and transmitted, while maintaining an adequate level of perceptual quality. In this paper, we propose a novel light field compression scheme that uses view estimation to recover the entire light field from a small subset of encoded views. Experimental results on a widely used light field dataset show that our method achieves good coding efficiency with average rate savings of 54.83% with respect to HEVC. Baptiste Hériard-Dubreuil, Irene Viola 0001, Touradj Ebrahimi |
PCS | 3 |
| 2019 | Exploiting user interactivity in quality assessment of point cloud imagingabstractPoint clouds are a new modality for representation of plenoptic content and a popular alternative to create immersive media. Despite recent progress in capture, display, storage, delivery and processing, the problem of a reliable approach to subjectively and objectively assess the quality of point clouds is still largely open. In this study, we extend the state of the art in projection-based objective quality assessment of point cloud imaging by investigating the impact of the number of viewpoints employed to assess the visual quality of a content, while discarding information that does not belong to the object under assessment, such as background color. Additionally, we propose assigning weights to the projected views based on interactivity information, obtained during subjective evaluation experiments. In the experiment that was conducted, human observers assessed a carefully selected collection of typical contents, subject to geometry and color degradations due to compression. The point cloud models were rendered using cubes as primitive elements with adaptive sizes based on local neighborhoods. Our results show that employing a larger number of projected views does not necessarily lead to better predictions of visual quality, while user interactivity information can improve the performance. Evangelos Alexiou, Touradj Ebrahimi |
QoMEX | 2 |
| 2019 | Point cloud quality evaluation: Towards a definition for test conditionsabstractRecently stakeholders in the area of multimedia representation and transmission have been looking at plenoptic technologies to improve immersive experience. Among these technologies, point clouds denote a volumetric information representation format with important applications in the entertainment, automotive and geographical mapping industries. There is some consensus that state-of-the-art solutions for efficient storage and communication of point clouds are far from satisfactory. This paper describes a study on point cloud quality evaluation, conducted in the context of JPEG Pleno to help define the test conditions of future compression proposals. A heterogeneous set of static point clouds in terms of number of points, geometric structure and represented scenarios were selected and compressed using octree-pruning and a projection-based method, with three different levels of degradation. The models were comprised of both geometrical and color information and were displayed using point sizes large enough to ensure observation of watertight surfaces. The stimuli under assessment were presented to the observers on 2D displays as animations, after defining suitable camera paths to enable visualization of the models in their entirety and realistic consumption. The experiments were carried out in three different laboratories and the subjective scores were used in a series of correlation studies to benchmark objective quality metrics and assess inter-laboratory consistency. Luís Alberto da Silva Cruz, Emil Dumic, Evangelos Alexiou, João Prazeres, Carlos Rafael Duarte, Manuela Pereira, António M. G. Pinheiro, Touradj Ebrahimi |
QoMEX | 8 |
| 2019 | An in-depth analysis of single-image subjective quality assessment of light field contentsabstractQuality assessment of light field images poses new questions and challenges, due to the enriched nature of the content and the possibilities it offers at the rendering step. Image-based rendering is conventionally used to showcase the increased capabilities of light field contents on traditional 2D screens. However, the range of possibilities for rendering parameters is virtually endless, which poses the problem of what rendered images should be used when performing visual quality assessments, as well as how to properly present them to subjects during quality evaluations. Single-image assessment has been used in the past to conduct subjective quality evaluations. Since this type of assessment generates a large number of stimuli to be evaluated, which increases the complexity, length, and cost of the test, it is fundamental to analyze whether the added strain on the evaluation procedure is compensated by statistically relevant results. In this paper, we analyze the results of a subjective evaluation campaign that used single-image assessment by means of statistical tools, to understand whether the advantages of evaluating light field contents through separately rendered images counterbalance the increase in complexity. In particular, we test whether different types of rendering lead to statistically different ratings, and if testing a variety of rendering parameters through single-image assessment is advisable. Results provide useful guidelines to designs more efficient subjective quality assessment for light field contents. Irene Viola 0001, Touradj Ebrahimi |
QoMEX | 2 |
| 2018 | Point Cloud Quality Assessment Metric Based on Angular SimilarityabstractThe rise of immersive technologies has been recently fuelled by emerging applications which employ advanced content representations. Among various alternatives, point clouds denote a promising solution which has recently drawn a significant amount of interest, as witnessed by the latest activities of standardization committees. However, subjective and objective quality assessments for this type of content still remain an open problem. In this paper, we introduce a simple yet efficient objective metric to capture perceptual degradations of a distorted point cloud. Correlation with subjective quality assessment scores carried out by human subjects shows the proposed metric to be superior to the state of the art in terms of predicting the visual quality of point clouds under realistic types of distortions, such as octree-based compression. Evangelos Alexiou, Touradj Ebrahimi |
ICME | 2 |
| 2018 | Benchmarking of Objective Quality Metrics for Colorless Point CloudsabstractRecent advances in depth sensing and display technologies, along with the significant growth of interest for augmented and virtual reality applications, lay the foundation for the rapid evolution of applications that provide immersive experiences. In such applications, advanced content representations are required in order to increase the engagement of the user with the displayed imageries. Point clouds have emerged as a promising solution to this aim, due to their efficiency in capturing, storing, delivering and rendering of 3D immersive contents. As in any type of imaging, the evaluation of point clouds in terms of visual quality is essential. In this paper, benchmarking results of the state-of-the-art objective metrics in geometry-only point clouds are reported and analyzed under two different types of geometry degradations, namely Gaussian noise and octree- based compression. Human ratings obtained from two subjective experiments are used as the ground truth. Our results show that most objective quality metrics perform well in the presence of noise, whereas one particular method has high predictive power and outperforms the others after octree-based encoding. Evangelos Alexiou, Touradj Ebrahimi |
PCS | 2 |
| 2018 | Comparison of Compression Efficiency between HEVC/H.265, VP9 and AV1 based on Subjective Quality AssessmentsabstractThe growing requirements for broadcasting and streaming of high quality video continue to trigger demands for codecs with higher compression efficiency. AV1 is the most recent open and royalty free video coding specification developed by Alliance for Open Media (AOMedia) with a declared ambition of becoming the most popular next generation video coding standard. Primary alternatives to AV1 are the VP9 and the HEVC/H.265 which are currently among the most popular and widespread video codecs used in applications. VP9 is also a royalty free and open specification similar to AV1, while HEVC/H.265 requires specific licensing terms for its use in commercial products and services. In this paper, we compare AV1 to VP9 and HEVC/H.265 from rate distortion point of view in a broadcasting use case scenario. Mutual comparison is performed by means of subjective evaluations carried out in a controlled environment using HD video content with typical bitrates ranging from low to high, corresponding to very low up to completely transparent quality. We then proceed with an in-depth analysis of advantages and drawbacks of each codec for specific types of content and compare the subjective comparisons and conclusions to those obtained by others in the state of the art as well to those measured by means of objective metrics such as PSNR. Pinar Akyazi, Touradj Ebrahimi |
QoMEX | 2 |
| 2018 | Point Cloud Subjective Evaluation Methodology based on 2D RenderingabstractPoint clouds are one of the most promising technologies for 3D content representation. In this paper, we describe a study on quality assessment of point clouds, degraded by octree-based compression on different levels. The test contents were displayed using Screened Poisson surface reconstruction, without including any textural information, and they were rated by subjects in a passive way, using a 2D image sequence. Subjective evaluations were performed in five independent laboratories in different countries, with the inter-laboratory correlation analysis showing no statistical differences, despite the different equipment employed. Benchmarking results reveal that the state-of-the-art point cloud objective metrics are not able to accurately predict the expected visual quality of such test contents. Moreover, the subjective scores collected from this experiment were found to be poorly correlated with subjective scores obtained from another test involving visualization of raw point clouds. These results suggest the need for further investigations on adequate point cloud representations and objective Quality assessment tools. Evangelos Alexiou, Touradj Ebrahimi, Marco V. Bernardo, Manuela Pereira, António M. G. Pinheiro, Luís Alberto da Silva Cruz, Carlos Duarte, Lovorka Gotal Dmitrovic, Emil Dumic, Dragan Matkovics, Athanassios N. Skodras |
QoMEX | 2 |
| 2018 | VALID: Visual quality Assessment for Light field Images DatasetabstractIn the last years, light field imaging has experienced a surge of popularity among the scientific community for its capability of rendering the 3D world in a more immersive way. In particular, several compression algorithms have been proposed to efficiently reduce the amount of data generated in the acquisition process, and different methodologies have been designed to reliably evaluate the visual quality of compressed contents. In this paper we propose a dataset for visual quality assessment of light field images (VALID). The dataset contains five contents compressed at various bitrates, using both off-the-shelf solutions and state-of-the-art algorithms. Results of objective quality evaluation using popular image metrics are included, as well as annotated subjective scores using three different methodologies and two types of visualization setups. The proposed dataset will help develop new objective metrics to predict visual quality, design new subjective assessment methodologies and compare them to existing ones, as well as produce novel analysis approaches to interpret the results. Irene Viola 0001, Touradj Ebrahimi |
QoMEX | 2 |
| 2018 | A Reliable and Reversible Image Privacy Protection Based on False ColorsabstractProtection of visual privacy has become an indispensable component of video surveillance systems due to pervasive use of video cameras for surveillance purposes. In this paper, we propose two fully reversible privacy protection schemes implemented within the JPEG architecture. In both schemes, privacy protection is accomplished by using false colors with the first scheme being adaptable to other privacy protection filters while the second is false color-specific. Both schemes support either a lossless mode in which the original unprotected content can be fully extracted or a lossy mode, which limits file size while still maintaining intelligibility. Our method is not region-of-interest (ROI)-based and can be applied on entire frames without compromising intelligibility. This frees the user from having to define ROIs and improves security as tracking ROIs under dynamic content may fail, exposing sensitive information. Our experimental results indicate the favorability of our method over other commonly used solutions to protect visual privacy. Serdar Çiftçi, Ahmet Oguz Akyüz, Touradj Ebrahimi |
IEEE Trans. Multim. | 3 |
| 2017 | MMHealth 2017: Workshop on Multimedia for Personal Health and Health CareabstractEver since the emergence of digitization, we've used the term multimedia to represent a combination of different kinds of media types, such as images, audio, and videos. As new sensing technologies emerge and are now becoming omnipresent in daily lives, the definition, role and significance of multimedia is changing. Multimedia now represents the means for communicating, cooperating, and also for monitoring numerous aspects of daily life, at various levels of granularity and application, ranging from personal to societal. With this shift, we have since moved from comprehending single media and its state toward comprehending media in terms of its use context. Susanne Boll, Touradj Ebrahimi, Cathal Gurrin, Laleh Jalali, Ramesh Jain 0001, Jochen Meyer 0001, Noel E. O'Connor |
ACM Multimedia | 2 |
| 2017 | Towards subjective quality assessment of point cloud imaging in augmented realityabstractRecently, there has been an increased interest in capture, processing and rendering of visual content in form of point clouds. Among other challenges, subjective and objective quality assessments of point clouds are still open problems. Most proposed subjective quality evaluation methodologies are variants or extensions of counter parts from conventional approaches such as those proposed in various ITU-R and ITU-T recommendations. A key issue with point cloud content is that of rendering and display devices which are thoroughly different from those in other modalities in addition to novel applications which depart from traditional display devices. In this paper, we propose a radically different approach to point cloud subjective quality assessment for point cloud by making use of augmented reality head mounted displays. Beside description of the approach, we show examples of implementation of the proposed methodology and draw conclusions regarding its advantages and drawbacks. Finally, the proposed approach is used in assessing the performance of widely used objective metrics to compute quality of point cloud contents when they undergo various types of distortions such as corruption by noise, simplification and compression. Evangelos Alexiou, Evgeniy Upenik, Touradj Ebrahimi |
MMSP | 3 |
| 2017 | On subjective and objective quality evaluation of point cloud geometryabstractPoint clouds have emerged as a promising solution for immersive representation of 3D contents. In this paper an interactive subjective quality assessment for point clouds is proposed. Quality assessment of geometry information in point clouds subject to realistic types of degradations is performed and correlation between state-of-the-art objective metrics and ground truth subjective scores is investigated. Preliminary results suggest that there is a need for more appropriate objective metrics, since the current solutions are not able to provide accurate predictions of quality for every type of degradations and contents. Evangelos Alexiou, Touradj Ebrahimi |
QoMEX | 2 |
| 2017 | Towards the need satisfaction in gaming: A comparison of different gaming platformsabstractRecent advances in Virtual Reality (VR) technologies have resulted in a wider availability of Head Mounted Displays (HMDs). However, it is still unclear if VR gaming offers a substantial added value to players. For this reason a comparison of gaming experiences on VR HMD to those on mobile and PC, two other popular gaming platforms, is performed by conducting a user study via two games available on all three platforms. We explore the QoE of gaming by investigating momentous dimensions using the Player Experience of Need Satisfaction (PENS) questionnaire. The results show higher Presence and Autonomy obtained by using HMD when compared to the two other platforms. However, these factors alone did not improve the Overall Quality. To take advantage of the new technology, satisfaction of all psychological needs, especially Competency, must be assured. Anne-Flore Perrin, Touradj Ebrahimi, Saman Zad Tootaghaj, Steven Schmidt 0001, Sebastian Möller 0001 |
QoMEX | 2 |
| 2017 | On the performance of objective metrics for omnidirectional visual contentabstractOmnidirectional image and video have gained popularity thanks to availability of capture and display devices for this type of content. Recent studies have assessed performance of objective metrics in predicting visual quality of omnidirectional content. These metrics, however, have not been rigorously validated by comparing their prediction results with ground-truth subjective scores. In this paper, we present a set of 360-degree images along with their subjective quality ratings. The set is composed of four contents represented in two geometric projections and compressed with three different codecs at four different bitrates. A range of objective quality metrics for each stimulus is then computed and compared to subjective scores. Statistical analysis is performed in order to assess performance of each objective quality metric in predicting subjective visual quality as perceived by human observers. Results show the estimated performance of the state-of-the-art objective metrics for omnidirectional visual content. Objective metrics specifically designed for 360-degree content do not outperform conventional methods designed for 2D images. Evgeniy Upenik, Martin Rerábek, Touradj Ebrahimi |
QoMEX | 3 |
| 2017 | Impact of interactivity on the assessment of quality of experience for light field contentabstractThe recent advances in light field imaging are changing the way in which visual content is captured, processed and consumed. Storage and delivery systems for light field images rely on efficient compression algorithms. Such algorithms must additionally take into account the feature-rich rendering for light field content. Therefore, a proper evaluation of visual quality is essential to design and improve coding solutions for light field content. Consequently, the design of subjective tests should also reflect the light field rendering process. This paper aims at presenting and comparing two methodologies to assess the quality of experience in light field imaging. The first methodology uses an interactive approach, allowing subjects to engage with the light field content when assessing it. The second, on the other hand, is completely passive to ensure all the subjects will have the same experience. Advantages and drawbacks of each approach are compared by relying on statistical analysis of results and conclusions are drawn. The obtained results provide useful insights for future design of evaluation techniques for light field content. Irene Viola 0001, Martin Rerábek, Touradj Ebrahimi |
QoMEX | 3 |
| 2017 | Context-Dependent Privacy-Aware Photo Sharing Based on Machine Learning
Lin Yuan 0002, Joël Theytaz, Touradj Ebrahimi |
SEC | 3 |
| 2017 | Privacy protection of tone-mapped HDR images using false coloursabstractHigh dynamic range (HDR) imaging has been developed for improved visual representation by capturing a wide range of luminance values. Owing to its properties, HDR content might lead to a larger privacy intrusion, requiring new methods for privacy protection. Previously, false colours were proved to be effective for assuring privacy protection for low dynamic range (LDR) images. In this work, the reliability of false colours when used for privacy protection of HDR images represented by tone‐mapping operators (TMOs) is studied. Two different TMO techniques are tested, a simple TMO based on the Gamma transform and a more complex local TMO. Moreover, two false colour palettes are also tested, and are applied to images that result from both TMOs and also to an LDR image that represents the centre exposure in the image sequence used to create the HDR image. The degree of privacy protection is analysed through both a subjective test using crowdsourcing and an objective test using face recognition algorithms. It is concluded that the application of the two studied false colour palettes reduces the recognition accuracy with respect to both tests. Serdar Çiftçi, Ahmet Oguz Akyüz, António M. G. Pinheiro, Touradj Ebrahimi |
IET Signal Process. | 4 |
| 2017 | Image privacy protection with secure JPEG transmorphingabstractThanks to advancements in smart mobile devices and social media platforms, sharing photos and experiences has significantly bridged the authors’ lives, allowing them to stay connected despite distance and other barriers. Most approaches to protect image visual privacy focus on encrypting or permuting image data, which generate unreadable image or highly distorted visual effect and therefore may not be in users best interest from both usage and perception perspectives. In this study, the authors propose secure JPEG transmorphing, a framework for protecting image visual privacy in a secure, reversible, and highly flexible and personalised manner. Secure JPEG transmorphing allows one to apply arbitrary regional visual manipulation on image regions of interests (ROIs), while secretly preserving the information about the original ROIs in application segments (APPn markers) of the visually obfuscated JPEG image. Objective and subjective experiments have been performed and results indicate that the proposed protection scheme provides near lossless image reconstruction, controllable level of file size expansion, good degree of privacy protection and especially better subjective pleasantness. Lin Yuan 0002, Touradj Ebrahimi |
IET Signal Process. | 2 |
| 2017 | Evaluation of JPEG XT for high dynamic range cameras
Philippe Hanhart, Touradj Ebrahimi |
Signal Process. Image Commun. | 2 |
| 2016 | Testbed for subjective evaluation of omnidirectional visual contentabstractOmni-directional visual content is a form of representing graphical and cinematic media content which provides subjects with the ability to freely change their direction of view. Along with virtual reality, omnidirectional imaging is becoming a very important type of the modern media content. This brings new challenges to the omnidirectional visual content processing, especially in the field of compression and quality evaluation. More specifically, the ability to assess quality of omnidirectional images in reliable manner is a crucial step to provide a rich quality of immersive experience. In this paper we introduce a testbed suitable for subjective evaluations of omnidirectional visual contents. We also show the results of a conducted pilot experiment to illustrate the applicability of the proposed testbed. Evgeniy Upenik, Martin Rerábek, Touradj Ebrahimi |
PCS | 3 |
| 2016 | Objective and subjective evaluation of light field image compression algorithmsabstractThis paper reports results of subjective and objective quality assessments of responses to a grand challenge on light field image compression. The goal of the challenge was to collect and evaluate new compression algorithms for light field images. In total seven proposals were received, out of which five were accepted for further evaluations. For objective evaluations, conventional metrics were used, whereas the double stimulus continuous quality scale method was selected to perform subjective assessments. Results show competitive performance among submitted proposals. However, in low bitrates, one proposal outperforms the others. Irene Viola 0001, Martin Rerábek, Tim Bruylants, Peter Schelkens, Fernando Pereira 0001, Touradj Ebrahimi |
PCS | 6 |
| 2016 | How to benchmark objective quality metrics from paired comparison data?abstractThe procedures commonly used to evaluate the performance of objective quality metrics rely on ground truth mean opinion scores and associated confidence intervals, which are usually obtained via direct scaling methods. However, indirect scaling methods, such as the paired comparison method, can also be used to collect ground truth preference scores. Indirect scaling methods have a higher discriminatory power and are gaining popularity, for example in crowdsourcing evaluations. In this paper, we present how the classification errors, an existing analysis tool, can also be used with subjective preference scores. Additionally, we propose a new analysis tool based on the receiver operating characteristic analysis. This tool can be used to further assess the performance of objective metrics based on ground truth preference scores. We provide a MATLAB script with an implementation of the proposed tools and we show one example of application of the proposed tools. Philippe Hanhart, Lukas Krasula, Patrick Le Callet, Touradj Ebrahimi |
QoMEX | 4 |
| 2016 | Subjective and objective evaluation of HDR video coding technologiesabstractThis paper reports the details and results of a subjective and objective quality evaluation assessing responses to an MPEG call for evidence (CfE) on high dynamic range (HDR) and wide color gamut video coding. Five HDR video contents, compressed at four bit rates by each proponent responding to the CfE, were used in the subjective assessments. To be able to evaluate the performance of objective quality metrics, the double stimulus impairment scale (DSIS) method was used for subjective assessments instead of previously published paired comparison to an anchor. Subjective results show evidence that coding efficiency can be improved in a statistically noticeable way over the HEVC anchor in terms of perceived quality. However, when compared to paired comparison, less statistically significant differences are observed because of the lower discrimination power of the DSIS method. The collected subjective scores were used as a ground truth to benchmark and analyze the performance of objective metrics. Results show that HDR-VDP-2 and PQ2VIFP have the highest correlation with subjective scores and outperform other investigated metrics. Philippe Hanhart, Martin Rerábek, Touradj Ebrahimi |
QoMEX | 3 |
| 2016 | Modeling immersive media experiences by sensing impact on subjects
Eleni Kroupi, Philippe Hanhart, Jong-Seok Lee, Martin Rerábek, Touradj Ebrahimi |
Multim. Tools Appl. | 5 |
| 2016 | Subject-Independent Odor Pleasantness Classification Using Brain and Peripheral SignalsabstractEnhanced sensation of reality from multimedia contents can be achieved by creating realistic multimedia environments, using visual, auditory, and olfactory information. Although the affective information from video and audio has been extensively studied, the olfactory sense has received less attention. A way to assess human experience from audio, video or odors, is by investigating physiological signals. In this study, 23 subjects experienced pleasant, unpleasant, and neutral odors while their electroencephalogram (EEG), and electrocardiogram (ECG) were recorded. Two independent three-class classifiers were trained and tested, using EEG or ECG features. The results reveal a significant increase in the classification performance when EEG features were used (Cohen's kappa k = 0.44 ± 0.14; p <; 0.001). The results also indicate that it is possible to automatically classify the perception of unpleasant odors using EEG signals, but the classification performance decreases significantly when classifying between pleasant and neutral odors. Among the EEG features, the Wasserstein distance metric estimated between trial and baseline power achieved the highest classification performance. Features from ECG signals did not result in a significantly non-random performance. Eleni Kroupi, Jean-Marc Vesin, Touradj Ebrahimi |
IEEE Trans. Affect. Comput. | 3 |
| 2015 | Functional connectivity from EEG signals during perceiving pleasant and unpleasant odorsabstractThe olfactory sense is strongly related to memory and emotional processes. Studies on the effects of odor perception from brain activity have been conducted by using different neuro-imaging techniques. In this paper, we analyse electroencephalography (EEG) of 23 subjects during perceiving pleasant and unpleasant odor stimuli. We describe the construction of brain functional connectivity networks measured by most commonly used models. We discuss the network-based features of functional connectivity, and design classifiers by applying different functional connectivity network features. Finally, we show that pleasant and unpleasant emotions from olfactory perceptions can be better classified if we see the brain as a nonlinear small-world network. By extracting appropriate features from functional connectivity networks, we manage to classify pleasant and unpleasant olfactory perceptions with an average Kappa value of 0.11 ± 0.17, which is significantly non-random. Eleni Kroupi, Touradj Ebrahimi |
ACII | 3 |
| 2015 | Privacy in mini-drone based video surveillanceabstractMini-drones are increasingly used in video surveillance. Their areal mobility and ability to carry video cameras provide new perspectives in visual surveillance which can impact privacy in ways that have not been considered in a typical surveillance scenario. To better understand and analyze them, we have created a publicly available video dataset of typical drone-based surveillance sequences in a car parking. Using the sequences from this dataset, we have assessed five privacy protection filters via a crowdsourcing evaluation. We asked crowdsourcing workers several privacy- and surveillance-related questions to determine the tradeoff between intelligibility of the scene and privacy, and we present conclusions of this evaluation in this paper. Margherita Bonetto, Pavel Korshunov, Giovanni Ramponi, Touradj Ebrahimi |
ICIP | 4 |
| 2015 | Rate-distortion evaluation for two-layer coding systemsabstractThe Bjøntegaard model is widely used to calculate the compression efficiency between different codecs. However, this model is not sufficient to investigate the impact on quality of the interaction of the base and enhancement layer bit rates when comparing two-layer coding systems. Therefore, in this paper, we propose an extension of the Bjøntegaard model from rate-distortion (R-D) curve fitting to rate-rate-distortion (R2-D) surface fitting. The proposed model uses a cubic surface as fitting function and a more complex characterization of the domain formed by the data points to compute a more realistic estimate of compression efficiency. The proposed model can be used to measure the compression efficiency of two-layer coding systems, as well as for other applications, e.g., to optimize the bit rate allocation between texture and depth in 3D video coding. Two examples of assessment of compression efficiency in JPEG XT are presented as illustrations. Philippe Hanhart, Touradj Ebrahimi |
ICIP | 2 |
| 2015 | Image transmorphing with JPEGabstractPicture-related applications are extremely popular because pictures present attractive and vivid information. Nowadays, people record everyday life, communicate with each other, and enjoy entertainment using various interesting imaging applications. In many cases, processed images need to be recovered to their original versions. However, most approaches require storage or transmission of both original and processed images separately, which result in increased bandwidth and storage resources to be used. In contrast, in this paper, we present a JPEG transmorphing algorithm, which converts an image to its processed version while preserving sufficient information about the original image in the processed image. It does this by inserting partial information about the original image in the application markers of the processed JPEG image file, so that the original image can be later recovered. Experiments are conducted and results show that the proposed method offers a number of attractive features and a good performance in many applications. Lin Yuan 0002, Touradj Ebrahimi |
ICIP | 2 |
| 2015 | Active crosstalk reduction system for multiview autostereoscopic displaysabstractMultiview autostereoscopic displays are considered as the future of 3DTV. However, these displays suffer from a high level of crosstalk, which negatively impacts quality of experience (QoE). In this paper, we propose a system to improve 3D QoE on multiview autostereoscopic displays. First, the display is characterized in terms of luminance distribution. Then, the luminance profiles are modeled using a limited set of parameters. A Kinect sensor is used to determine the viewer position in front of the display. Finally, the proposed system performs an intelligent on the fly allocation of the output views to minimize the perceived crosstalk. The user preference between 2D and 3D modes and the proposed system is evaluated. Results show that picture quality is significantly improved when compared to the standard 3D mode, for a similar depth perception and visual comfort. Philippe Hanhart, Carmelo di Nolfo, Touradj Ebrahimi |
ICME | 3 |
| 2015 | Opportunities and Challenges of Global Network CamerasabstractSince the introduction of consumer digital cameras, user-created multimedia content has become increasingly popular. Digital cameras, together with inexpensive editing tools, and free hosting sites have made multimedia an integral part of everyday life. Today, hundreds of hours video are uploaded to hosting sites every minute. Video-on-demand through wireless networks and smartphones have profoundly changed how people consume multimedia content. Meanwhile, the widely deployed network cameras can provide live views of many parts of the world. These cameras can provide rich sources creating multimedia content. This panel will explore the opportunities and discuss the challenges using global network cameras for creating multimedia contents and understanding the world. Every year, millions of network cameras are deployed. The data from some of these network cameras are publicly available, continuously streaming live views of national parks, city halls, streets, highways, and shopping malls. A person may see multiple tourist attractions through these cameras, without leaving home. Researchers may observe the weather in different cities. Using the data from the cameras, it is possible to observe natural disasters, such as volcano eruption or tsunami, at a safe distance. News reporters may obtain instant views of an unfolding riot without risking their lives. A spectator may watch a celebration parade from multiple locations using the street cameras. Despite the many promising applications, the opportunities of using global network cameras for creating multimedia content have not been fully exploited. Joanna Batstone, Touradj Ebrahimi, Tiejun Huang 0001, Yung-Hsiang Lu, Yonggang Wen 0001 |
ACM Multimedia | 2 |
| 2015 | Multimodal Dataset for Assessment of Quality of Experience in Immersive MultimediaabstractThis paper presents a novel multimodal dataset for the analysis of Quality of Experience(QoE) in emerging immersive multimedia technologies. In particular, the perceived Sense of Presence (SoP) induced by one-minute long video stimuli is explored with respect to content, quality, resolution, and sound reproduction and annotated with subjective scores. Furthermore, a complementary analysis of the acquired physiological signals, such as EEG, ECG, and respiration is carried out, aiming at an alternative evaluation of human experience while consuming immersive multimedia. Anne-Flore Perrin, Eleni Kroupi, Martin Rerábek, Touradj Ebrahimi |
ACM Multimedia | 5 |
| 2015 | Report on the evaluation of current and future image compression technologiesabstractThis document reports the conclusions, comments and recommendations resulting from the final panel discussion which happened at the Session on Evaluation of Current and Future Image Compression Technologies held at the Picture Coding Symposium, Cairns, Australia, in June 2015. Fernando Pereira 0001, Ralf Schaefer, Touradj Ebrahimi, Jörn Ostermann, Edward J. Delp |
PCS | 3 |
| 2015 | Prediction of asynchronous dimensional emotion ratings from audiovisual and physiological data
Fabien Ringeval, Florian Eyben, Eleni Kroupi, Anil Yüce, Jean-Philippe Thiran, Touradj Ebrahimi, Denis Lalanne, Björn W. Schuller |
Pattern Recognit. Lett. | 6 |
| 2015 | Introduction to the Special Section on Visual Computing in the Cloud: Fundamentals and ApplicationsabstractCloud computing involves a large number of terminals connected through a real-time high-speed network (such as the Internet). The adoption rates for private and hybrid cloud services increased to 40% in 2013, with computing shifting from on-premise infrastructure to the cloud. To keep pace with the ever-accelerating rate of innovation, companies are moving to the cloud. However, visual computing in the cloud brings great challenges, such as how to measure and then improve the quality of experience in cloud computing. This Special Section provides the image/video community a forum to present new academic research and industrial development in running visual computing services in the cloud. This Special Section aims to address fundamental and practical aspects of visual computing in the cloud, such as how to build cloud platforms that can cope with seemingly unlimited supply of content coming from traditional media sources as well as new media uploaded to the Internet (YouTube, Facebook, etc.); how to leverage cloud technology to build high-quality image/video browsing and delivery experiences for a global audience; how to ingest, encode, process, adapt, as well as protect contents and privacy of users; how to provide both on-demand and live-streaming capabilities; how to tag image/video and allow consumers to access the image/video contents with high availability; how to support image/video services in mobile devices; and how to perform real-time image/video analytics in the cloud, to mention a few among a diverse range of challenges. Jiangchuan Liu, Wenwu Zhu 0001, Touradj Ebrahimi, John G. Apostolopoulos, Xian-Sheng Hua 0001, Chuan Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Crowd-based quality assessment of multiview video plus depth codingabstractCrowdsourcing is becoming a popular cost effective alternative to lab-based evaluations for subjective quality assessment. However, crowd-based evaluations are constrained by the limited availability of display devices used by typical online workers, which makes the evaluation of 3D content a challenging task. In this paper, we investigate two possible approaches to crowd-based quality assessment of multiview video plus depth (MVD) content on 2D displays: by using a virtual view and by using a free-viewpoint video, which corresponds to a smooth camera motion during a time freeze. We conducted the crowdsourcing experiments using seven MVD sequences encoded at different bit rates with the upcoming 3D-AVC video coding standard. The results demonstrate high correlation with subjective evaluations performed using a stereoscopic monitor in a controlled laboratory environment. The analysis shows no statistically significant difference between the two approaches. Philippe Hanhart, Pavel Korshunov, Touradj Ebrahimi |
ICIP | 3 |
| 2014 | Scrambling-based tool for secure protection of JPEG imagesabstractJPEG scrambling tool is a flexible web-based tool with an intuitive and simple to use GUI and interface to secure visual information in a region of interest (ROI) of JPEG images. The tool demonstrates an efficient integration and use of security tools in JPEG image format by an example of scrambling privacy filter, which enables a variety of security services such as confidentiality, integrity verification, source authentication, and conditional access. Pavel Korshunov, Touradj Ebrahimi |
ICIP | 2 |
| 2014 | Towards optimal distortion-based visual privacy filtersabstractThe widespread usage of digital video surveillance systems has increased the concerns for privacy violation. Since video surveillance systems are invasive, it is a challenge to find an acceptable balance between privacy of the public under surveillance and security related features of the systems. Many privacy protection tools have been proposed for preserving privacy, ranging from such simple methods like blurring or pixelization to more advanced like scrambling and geometrical transform based filters. However, for a given filter implemented in a practical video surveillance system, it is necessary to know the strength with which the filter should be applied to protect privacy reliably. Assuming an automated surveillance system, this paper objectively investigates several privacy protection filters with varying strength degrees and determines their optimal strength values to achieve privacy protection. To this end, five privacy filters were applied to images from FERET dataset and the performance of three recognition algorithms was evaluated. The results show that different privacy protection filters influence the accuracy of different versions of face recognition differently and this influence depends both on the robustness of the recognition and the type of distortion filter. Pavel Korshunov, Touradj Ebrahimi |
ICIP | 2 |
| 2014 | Predicting subjective sensation of reality during multimedia consumption based on EEG and peripheral physiological signalsabstractSensation of reality refers to the ability of users to feel present in a multimedia experience. As 3D technologies target to provide more immersive and higher quality multimedia experiences, it is important to understand Quality of Experience (QoE) and sensation of reality. Recently, there have been efforts to measure brain activity in order to understand implicitly QoE for various multimedia contents. However, brain activity accounting for sensation of reality has not been adequately investigated. The goal of this paper is twofold. First, we investigate how various aspects, such as perceived quality, perceived depth, and content preference affect subjective sensation of reality through explicit subjective ratings. Second, we construct subjective classification systems to predict sensation of reality from multimedia experiences based on electroencephalography (EEG) and peripheral physiological signals such as heart rate and respiration. Eleni Kroupi, Philippe Hanhart, Jong-Seok Lee, Martin Rerábek, Touradj Ebrahimi |
ICME | 5 |
| 2014 | Impact of Ultra High Definition on Visual AttentionabstractUltra high definition (UHD) TV is rapidly replacing high definition (HD) TV but little is known of its effects on human visual attention. However, a clear understanding of this effect is important, since accurate models, evaluation methodologies, and metrics for visual attention are essential in many areas, including image and video compression, camera and displays manufacturing, artistic content creation, and advertisement. In this paper, we address this problem by creating a dataset of UHD resolution images with corresponding eye-tracking data, and we show that there is a statistically significant difference between viewing strategies when watching UHD and HD contents. Furthermore, by evaluating five representative computational models of visual saliency, we demonstrate the decrease in models' accuracies on UHD contents when compared to HD contents. Therefore, to improve the accuracy of computational models for higher resolutions, we propose a segmentation-based resolution-adaptive weighting scheme. Our approach demonstrates that taking into account information about resolution of the images improves the performance of computational models. Hiromi Nemoto, Philippe Hanhart, Pavel Korshunov, Touradj Ebrahimi |
ACM Multimedia | 4 |
| 2014 | Free-viewpoint video sequences: A new challenge for objective quality metricsabstractFree-viewpoint television is expected to create a more natural and interactive viewing experience by providing the ability to interactively change the viewpoint to enjoy a 3D scene. To render new virtual viewpoints, free-viewpoint systems rely on view synthesis. However, it is known that most objective metrics fail at predicting perceived quality of synthesized views. Therefore, it is legitimate to question the reliability of commonly used objective metrics to assess the quality of free-viewpoint video (FVV) sequences. In this paper, we analyze the performance of several commonly used objective quality metrics on FVV sequences, which were synthesized from decompressed depth data, using subjective scores as ground truth. Statistical analyses showed that commonly used metrics were not reliable predictors of perceived image quality when different contents and distortions were considered. However, the correlation improved when considering individual conditions, which indicates that the artifacts produced by some view synthesis algorithms might not be correctly handled by current metrics. Philippe Hanhart, Emilie Bosc, Patrick Le Callet, Touradj Ebrahimi |
MMSP | 4 |
| 2014 | Performance evaluation of the emerging JPEG XT image compression standardabstractThe upcoming JPEG XT is under development for High Dynamic Range (HDR) image compression. This standard encodes a Low Dynamic Range (LDR) version of the HDR image generated by a Tone-Mapping Operator (TMO) using the conventional JPEG coding as a base layer and encodes the extra HDR information in a residual layer. This paper studies the performance of the three profiles of JPEG XT (referred to as profiles A, B and C) using a test set of six HDR images. Four TMO techniques were used for the base layer image generation to assess the influence of the TMOs on the performance of JPEG XT profiles. Then, the HDR images were coded with different quality levels for the base layer and for the residual layer. The performance of each profile was evaluated using Signal to Noise Ratio (SNR), Feature SIMilarity Index (FSIM), Root Mean Square Error (RMSE), and CIEDE2000 color difference objective metrics. The evaluation results demonstrate that profiles A and B lead to similar saturation of quality at the higher bit rates, while profile C exhibits no saturation. Profiles B and C appear to be more dependent on TMOs used for the base layer compared to profile A. António M. G. Pinheiro, Karel Fliegel, Pavel Korshunov, Lukas Krasula, Marco V. Bernardo, Maria Pereira, Touradj Ebrahimi |
MMSP | 7 |
| 2014 | Exploring the Impact of Food Craving and Pleasure Technologies on Aesthetic Experience in Digital MediaabstractHumans are known to be good in creating pleasure technologies. In fact, some evolutionary psychologists believe that aesthetic experiences are biologically hardwired. This article reviews how the pleasure stimulus concept can be explored to enhance human experience in digital media. It is argued that the individual's consumption motivation is more than the by-products of the biological pleasure circuits. For instance, in daily life, one experiences various information processing, some of which does not emerge into one's explicit consciousness but relevantly contributes to one's experience. Although craving or desiring is mostly an explicit process, it can also manifest as an unconscious aspect of experience leading to an irrational or intrusive thoughts that can in turn alter or contribute to the aesthetic character of an experience. Using Quality of Experience assessment methodologies and extensive literature from a multidisciplinary standpoint, this article shows that intrusive mental concepts on food craving saliently affect the user's aesthetic experience and perception of quality in digital media. Wendy Ann Mansilla, Andrew Perkis, Touradj Ebrahimi |
Int. J. Hum. Comput. Interact. | 3 |
| 2014 | Calculation of average coding efficiency based on subjective quality scores
Philippe Hanhart, Touradj Ebrahimi |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Attention Driven Foveated Video Quality AssessmentabstractContrast sensitivity of the human visual system to visual stimuli can be significantly affected by several mechanisms, e.g., vision foveation and attention. Existing studies on foveation based video quality assessment only take into account static foveation mechanism. This paper first proposes an advanced foveal imaging model to generate the perceived representation of video by integrating visual attention into the foveation mechanism. For accurately simulating the dynamic foveation mechanism, a novel approach to predict video fixations is proposed by mimicking the essential functionality of eye movement. Consequently, an advanced contrast sensitivity function, derived from the attention driven foveation mechanism, is modeled and then integrated into a wavelet-based distortion visibility measure to build a full reference attention driven foveated video quality (AFViQ) metric. AFViQ exploits adequately perceptual visual mechanisms in video quality assessment. Extensive evaluation results with respect to several publicly available eye-tracking and video quality databases demonstrate promising performance of the proposed video attention model, fixation prediction approach, and quality metric. Junyong You, Touradj Ebrahimi, Andrew Perkis |
IEEE Trans. Image Process. | 2 |
| 2014 | EEG Correlates of Pleasant and Unpleasant Odor PerceptionabstractOlfaction-enhanced multimedia experience is becoming vital for strengthening the sensation of reality and the quality of user experience. One approach to investigate olfactory perception is to analyze the alterations in brain activity during stimulation with different odors. In this article, the changes in the electroencephalogram (EEG) when perceiving hedonically-different odors are studied. Results of within and across-subject analysis are presented. We show that EEG-based odor classification using brain activity is possible and can be used to automatically recognize odor pleasantness when a subject-specific classifier is trained. However, it is a challenging problem to design a generic classifier. Eleni Kroupi, Ashkan Yazdani, Jean-Marc Vesin, Touradj Ebrahimi |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2013 | Phase-Amplitude Coupling between EEG and EDA While Experiencing Multimedia ContentabstractEmotion is a dynamic process that affects social relationships and influences the mechanisms of rational thinking and decision making. The influence of emotion on reasoning has gained a lot of attention in computer science. One approach to incorporating emotions into computers is by extracting features from various physiological signals and inferring the corresponding emotions or emotional dimensions. However, physiological processes are dynamic rather than static, and physiological signals mainly interact with each other during emotional processes rather than act individually. This study suggests that dynamical interdependencies of various signals that emanate from either the peripheral nervous system or the central nervous system might be involved in the emotional experience. Specifically, EEG (electroencephalography) signals captured from the temporal lobe are coupled with a narrow-band skin conductance response (SCR) while subjects are watching music clips, indicating that signals from the auditory cortex interact with SCR during this experience. Eleni Kroupi, Jean-Marc Vesin, Touradj Ebrahimi |
ACII | 3 |
| 2013 | Using face morphing to protect privacyabstractThe widespread use of digital video surveillance systems has also increased the concerns for violation of privacy rights. Since video surveillance systems are invasive, it is a challenge to find an acceptable balance between privacy of the public under surveillance and the functionalities of the systems. Tools for protection of visual privacy available today lack either all or some of the important properties such as security of protected visual data, reversibility (ability to undo privacy protection), simplicity, and independence from the video encoding used. To overcome these shortcomings, in this paper, we propose a morphing-based privacy protection method and focus on its robustness, reversibility, and security properties. We morph faces from a standard FERET dataset and run face detection and recognition algorithms on the resulted images to demonstrate that morphed faces retain the likeness of a face, while making them unrecognizable, which ensures the protection of privacy. Our experiments also demonstrate the influence of morphing strength on robustness and security. We also show how to determine the right parameters of the method. Pavel Korshunov, Touradj Ebrahimi |
AVSS | 2 |
| 2013 | COST Actions and Digital Libraries: Between Sustaining Best Practices and Unleashing Further Potential
Matthew J. Driscoll, Ralph Stübner, Touradj Ebrahimi, Muriel Foulonneau, Andreas Nürnberger, Andrea Scharnhorst, Joie Springer |
TPDL | 3 |
| 2013 | Objective quality metrics for video scalabilityabstractScalable video coding is emerging as an efficient alternative to simulcast encoding to distribute the same video content simultaneously to many users having different terminals and network conditions. In order to select the best combination of video scalability options for a given network condition and content, the availability of objective metrics that can reliably predict the video quality of scalable video sequence is crucial. In this paper, we propose a performance evaluation study of a set of state of the art Full-Reference and No-Reference metrics, considering a public database of test video sequences and related subjective quality annotations. The results demonstrate that most of the considered metrics show lack of robustness as predictors of subjective quality when both spatial and temporal quality distortions occur. Adrien Besson, Francesca De Simone, Touradj Ebrahimi |
ICIP | 3 |
| 2013 | Paired comparison-based subjective quality assessment of stereoscopic images
Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi |
Multim. Tools Appl. | 3 |
| 2013 | Comparative Study of Trust Modeling for Automatic Landmark TaggingabstractMany images uploaded to social networks are related to travel, since people consider traveling to be an important event in their life. However, a significant amount of travel images on the Internet lack proper geographical annotations or tags. In many cases, the images are tagged manually. One way to make this time-consuming manual tagging process more efficient is to propagate tags from a small set of tagged images to the larger set of untagged images automatically. In this paper, we present a system for automatic geotag propagation in images based on the similarity between image content (famous landmarks) and its context (associated geotags). In such a scenario, however, an incorrect or a spam tag can damage the integrity and reliability of the automated propagation system. Therefore, for reliable geotags propagation, we suggest adopting a user trust model based on social feedback from the users of the photo-sharing system. We compare this socially-driven approach with other user trust models via experiments and subjective testing on an image database of various famous landmarks. Results demonstrate that relying on user feedback is more efficient, since the number of propagated tags more than doubles without loss of accuracy compared to using other models or propagating without trust modeling. Peter Vajda, Pavel Korshunov, Touradj Ebrahimi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2012 | Automatic Webcam-Based Human Heart Rate Measurements Using Laplacian Eigenmap
Yonghong Tian 0001, Yaowei Wang 0001, Touradj Ebrahimi, Tiejun Huang 0001 |
ACCV (2) | 4 |
| 2012 | Video quality metric based on fixation prediction and foveal imagingabstractThis paper proposes a full-reference video quality metric based on foveated vision mechanism. Due to a non-uniform distribution of photo-receptors on the retina, the human visual system (HVS) has the highest resolution around the fixation point of eyes and dramatically decreases away from this point. Two key factors in the foveated vision are fixation point and retinal eccentricities of different visual objects. Based on an advanced video attention model in quality assessment scenarios, eye fixations are predicted from the attention map using a winner-takes-all (WTA) neural network. Four quality features describing distortions on luminance, spatial and temporal activities, as well as chrominance are derived between foveated representations of reference and distorted video frames. These quality features are then combined together by an appropriate spatiotemporal pooling scheme to build a video quality metric. Experimental results with respect to publicly available video quality databases demonstrate that the proposed quality model outperforms a previously proposed foveated video quality metric as well as state-of-the-art video quality models. Junyong You, Touradj Ebrahimi, Andrew Perkis |
ICIP | 2 |
| 2012 | Subjective Crosstalk Assessment Methodology for Auto-stereoscopic DisplaysabstractCross talk is one of the most annoying distortions in the visualization stage of stereoscopic systems. Specifically, both pattern and amount of cross talk in multi-view auto-stereoscopic displays are more complex because of viewing angle dependability, when compared to cross talk in 2-view stereoscopic displays. Regarding system cross talk there are objective measures to assess it in auto-stereoscopic displays. However, in addition to system cross talk, cross talk perceived by users is also impacted by scene content. Moreover, some cross talk is arguably beneficial in auto-stereoscopic displays. Therefore, in this paper, we further assess how cross talk is perceived by users with various scene contents and different viewing positions using auto-stereoscopic displays. In particular, the proposed subjective cross talk assessment methodology is realistic without restriction of the users viewing behavior and is not limited to the specific technique used in auto-stereoscopic displays. The test was performed on a slanted parallax barrier based auto-stereoscopic display. The subjective cross talk assessment results show their consistence to the system cross talk meanwhile more scene content and viewing position related cross talk perception information is provided. This knowledge can be used to design new cross talk perception metrics. Liyuan Xing, Jie Xu 0036, Kim Skildheim, Andrew Perkis, Touradj Ebrahimi |
ICME | 5 |
| 2012 | Visual Contrast Sensitivity Guided Video Quality AssessmentabstractContrast sensitivity is an important characteristic of the human visual system (HVS), which is widely used in image and video signal processing. For visual quality assessment, static spatio-temporal frequency based contrast sensitivity function (CSF) is often used, while the contrast sensitivity can also be affected by smooth pursuit of eyes when tracking attentive regions in the field of view. This paper proposes to tune CSF based on an attention map derived from a visual attention model. The tuned CSF formulated by spatial frequency, temporal velocity, and visual attention map is used to filter video signals in order to construct a quality metric according to the difference of the filtered signals between a reference video and its distorted version. Experimental results demonstrate that the proposed attention tuned spatio-velocity CSF outperforms the traditional spatio-temporal CSF in evaluating the perceived video quality. Junyong You, Liyuan Xing, Andrew Perkis, Touradj Ebrahimi |
ICME | 4 |
| 2012 | Subjective study of privacy filters in video surveillanceabstractExtensive adoption of video surveillance, affecting many aspects of the daily life, alarms the concerned public about the increasing invasion into personal privacy. Therefore, to address privacy issues, many tools have been proposed for protection of personal privacy in image and video. However, little is understood regarding the effectiveness of such tools and especially their impact on the underlying surveillance tasks. In this paper, we propose a subjective evaluation methodology to analyze the tradeoff between the preservation of privacy offered by these tools and the intelligibility of activities under video surveillance. As an example, the proposed method is used to compare several commonly employed privacy protection techniques, such as blurring, pixelization, and masking applied to indoor surveillance video. The results show that, for the test material under analysis, the pixelization filter provides the best performance in terms of balance between privacy protection and intelligibility. Pavel Korshunov, Claudia Araimo, Francesca De Simone, Carmelo Velardo, Jean-Luc Dugelay, Touradj Ebrahimi |
MMSP | 6 |
| 2012 | Geotag propagation in social networks based on user trust model
Peter Vajda, Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi |
Multim. Tools Appl. | 5 |
| 2012 | DEAP: A Database for Emotion Analysis ;Using Physiological SignalsabstractWe present a multimodal data set for the analysis of human affective states. The electroencephalogram (EEG) and peripheral physiological signals of 32 participants were recorded as each watched 40 one-minute long excerpts of music videos. Participants rated each video in terms of the levels of arousal, valence, like/dislike, dominance, and familiarity. For 22 of the 32 participants, frontal face video was also recorded. A novel method for stimuli selection is proposed using retrieval by affective tags from the last.fm website, video highlight detection, and an online assessment tool. An extensive analysis of the participants' ratings during the experiment is presented. Correlates between the EEG signal frequencies and the participants' ratings are investigated. Methods and results are presented for single-trial classification of arousal, valence, and like/dislike ratings using the modalities of EEG, peripheral physiological signals, and multimedia content analysis. Finally, decision fusion of the classification results from different modalities is performed. The data set is made publicly available and we encourage other researchers to use it for testing their own affective state estimation methods. Sander Koelstra, Christian Mühl, Mohammad Soleymani 0001, Jong-Seok Lee, Ashkan Yazdani, Touradj Ebrahimi, Thierry Pun, Anton Nijholt, Ioannis Patras |
IEEE Trans. Affect. Comput. | 6 |
| 2012 | Affect recognition based on physiological changes during the watching of music videosabstractAssessing emotional states of users evoked during their multimedia consumption has received a great deal of attention with recent advances in multimedia content distribution technologies and increasing interest in personalized content delivery. Physiological signals such as the electroencephalogram (EEG) and peripheral physiological signals have been less considered for emotion recognition in comparison to other modalities such as facial expression and speech, although they have a potential interest as alternative or supplementary channels. This article presents our work on: (1) constructing a dataset containing EEG and peripheral physiological signals acquired during presentation of music video clips, which is made publicly available, and (2) conducting binary classification of induced positive/negative valence, high/low arousal, and like/dislike by using the aforementioned signals. The procedure for the dataset acquisition, including stimuli selection, signal acquisition, self-assessment, and signal processing is described in detail. Especially, we propose a novel asymmetry index based on relative wavelet entropy for measuring the asymmetry in the energy distribution of EEG signals, which is used for EEG feature extraction. Then, the classification systems based on EEG and peripheral physiological signals are presented. Single-trial and single-run classification results indicate that, on average, the performance of the EEG-based classification outperforms that of the peripheral physiological signals. However, the peripheral physiological signals can be considered as a good alternative to EEG signals in the case of assessing a user's preference for a given music video clip (like/dislike) since they have a comparable performance to EEG signals while being more easily measured. Ashkan Yazdani, Jong-Seok Lee, Jean-Marc Vesin, Touradj Ebrahimi |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2012 | Assessment of Stereoscopic Crosstalk PerceptionabstractStereoscopic three-dimensional (3-D) services do not always prevail when compared with their two-dimensional (2-D) counterparts, though the former can provide more immersive experience with the help of binocular depth. Various specific 3-D artefacts might cause discomfort and severely degrade the Quality of Experience (QoE). In this paper, we analyze one of the most annoying artefacts in the visualization stage of stereoscopic imaging, namely, crosstalk, by conducting extensive subjective quality tests. A statistical analysis of the subjective scores reveals that both scene content and camera baseline have significant impacts on crosstalk perception, in addition to the crosstalk level itself. Based on the observed visual variations during changes in significant factors, three perceptual attributes of crosstalk are summarized as the sensorial results of the human visual system (HVS). These are shadow degree, separation distance, and spatial position of crosstalk. They are classified into two categories: 2-D and 3-D perceptual attributes, which can be described by a Structural SIMilarity (SSIM) map and a filtered depth map, respectively. An objective quality metric for predicting crosstalk perception is then proposed by combining the two maps. The experimental results demonstrate that the proposed metric has a high correlation (over 88%) when compared with subjective quality scores in a wide variety of situations. Liyuan Xing, Junyong You, Touradj Ebrahimi, Andrew Perkis |
IEEE Trans. Multim. | 3 |
| 2011 | EEG Correlates of Different Emotional States Elicited during Watching Music Videos
Eleni Kroupi, Ashkan Yazdani, Touradj Ebrahimi |
ACII (2) | 3 |
| 2011 | Audio-visual synchronization recovery in multimedia contentabstractThis paper proposes a method recovering audio-visual synchronization of multimedia content. It exploits the correlation between the acoustic and the visual signals in order to estimate the audio-visual drift existing in the content. By shifting the audio signal relative to the visual signal, the estimation of the drift is obtained by searching for the shift producing the maximal audio-visual correlation. We consider two correlation measures, namely, mutual information and canonical correlation, and compare their performance. Experimental results demonstrate that the method using the canonical correlation is effective in recovering the audio-visual synchronization for both speech and non-speech sequences. Jong-Seok Lee, Touradj Ebrahimi |
ICASSP | 2 |
| 2011 | Objective metrics for quality of experience in stereoscopic imagesabstractStereoscopic Quality of Experience (QoE) is the result of a complex combination of different influencing factors. Previously we had investigated the effect of factors such as scene content, camera baseline, screen size and viewing position on stereoscopic QoE using subjective tests. In this paper, we propose two objective metrics for predicting stereoscopic QoE using bottom-up and top-down approaches, respectively. Specifically, the bottom-up metric is based on characterizing the significant factors of QoE directly, which are scene content, camera baseline, screen size and crosstalk level. While the top-down metric interprets QoE from its perceptual attributes, including crosstalk perception and perceived depth. These perceptual attributes are modeled by their individual relationship with the significant factors and then combined linearly to build the top-down metric. Both proposed metrics have been validated against our own database and a publicly available database, showing a high correlation (over 86%) with the subjective scores of stereoscopic QoE. Liyuan Xing, Junyong You, Touradj Ebrahimi, Andrew Perkis |
ICIP | 3 |
| 2011 | Social game epitome versus automatic visual analysisabstractWith the rapid growth of digital photography, sharing of photos with friends and family has become very popular. When people share their photos, they usually organize them in albums according to events or places. To tell the story of some important events in one's life, it is desirable to have an efficient summarization tool which can help people to get a quick overview of an album containing huge number of photos. In this paper, we analyze an approach for photo album summarization through a novel social game “Epitome” as a Facebook application. This social game can collect research data and, at the same time, it provides a collage or a cover photo of the user's photo album, while, at the same time, the user enjoys playing the game. As a benchmark comparison to this game, we performed automatic visual analysis considering several state-of-the-art features. Peter Vajda, Lutz Goldmann, Touradj Ebrahimi |
ICME | 4 |
| 2011 | Visual attention tuned spatio-velocity contrast sensitivity for video quality assessmentabstractContrast sensitivity is an important characteristic of the human visual system (HVS), which is widely used in image and video signal processing. For visual quality assessment, static spatio-temporal frequency based contrast sensitivity function (CSF) is often used, while the contrast sensitivity can also be affected by smooth pursuit of eyes when tracking attentive regions in the field of view. This paper proposes to tune CSF based on an attention map derived from a visual attention model. The tuned CSF formulated by spatial frequency, temporal velocity, and visual attention map is used to filter video signals in order to construct a quality metric according to the difference of the filtered signals between a reference video and its distorted version. Experimental results demonstrate that the proposed attention tuned spatio-velocity CSF outperforms the traditional spatio-temporal CSF in evaluating the perceived video quality. Junyong You, Touradj Ebrahimi, Andrew Perkis |
ICME | 2 |
| 2011 | A new analysis method for paired comparison and its application to 3D quality assessmentabstractAmong various subjective quality evaluation methodologies, paired comparison has the advantage of improved simplicity of the subjects' evaluation task due to simplified rating scales and direct comparison of two stimuli. Thus, it may lead to more reliable results when individual quality levels are difficult to define, quality differences between stimuli are small or multiple quality factors are involved. This paper proposes a new method to analyze results of paired comparison-based subjective tests. By assuming that ties convey information about significant differences between two stimuli being compared, the confidence intervals for the quality scores are estimated using a maximum likelihood criterion, which enables us to intuitively examine the significance of quality score differences. We describe the complete test methodology including the test procedure, outlier detection and score analysis applied to quality assessment of 3D images acquired using varying camera distances. Experimental results demonstrate the usefulness of the proposed analysis method, as well as the enhanced quality discriminability of the paired comparison methodology in comparison to the conventional single stimulus methodology. Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi |
ACM Multimedia | 3 |
| 2011 | Implicit experiences as a determinant of perceptual quality and aesthetic appreciationabstractSince the Dadaist refusal of the conventional standards in art, followed by Fluxus` rejection of art as a commodity, and recently, the popularity of Internet and technology in art, artworks have become difficult to recognize as artworks in themselves. Modern works of art are no longer readily only seen today, more often fully experienced. The processing of an aesthetic experience needs a new understanding in terms of the changing context of art and the experiential perspective of art recipients. In the multimedia arena, the valid assumption is that, evaluations of aesthetic experiences are mostly based on the accessible information on the surface of the medium. Several research groups in psychology question the singularity of exterior-level assumptions demonstrating that there are modulating factors that affect aesthetic experiences and one of these is implicit experience. In this paper, we review the significance of empirical aesthetics in psychology and from the artistic point of view combined with a technical and experiential perspective. We also discuss our approach of considering the implications of implicit experiences to the modelling of Quality of Experience (QoE) where this is used as a measurement of perception and aesthetic judgment of contemporary and modern works of art. Wendy Ann Mansilla, Andrew Perkis, Touradj Ebrahimi |
ACM Multimedia | 3 |
| 2011 | Modeling motion visual perception for video quality assessmentabstractContrast sensitivity of Human Visual System (HVS) plays an important role in perceiving visual stimuli, and consequently, it has a significant impact on the perceived video quality. This paper proposes a visual perception model based on foveated vision and motion perception. The reference and the distorted video sequences are processed by the visual perception model to generate the perceived stimuli in HVS. The perceived difference of the processed sequences is measured in spatial and temporal domains considering the visual sensitivity. An advanced pooling scheme is proposed based on the visual attention mechanism, eye movement type, and influence of temporal quality variation, in order to estimate the perceived video quality. Experimental results demonstrate that the proposed metric significantly outperforms state-of-the-art quality models with respect to a combined eye-tracking and subjective video quality assessment data set. Junyong You, Touradj Ebrahimi, Andrew Perkis |
ACM Multimedia | 2 |
| 2011 | Motion parallax based restitution of 3D images on legacy consumer mobile devicesabstractWhile 3D display technologies are already widely available for cinema and home or corporate use, only a few portable devices currently feature 3D display capabilities. Moreover, the large majority of 3D display solutions rely on binocular perception. In this paper, we study the alternative methods for restitution of 3D images on conventional 2D displays and analyze their respective performance. This particularly includes the extension of wiggle stereoscopy for portable devices which relies on motion parallax as an additional depth cue. The goal of this paper is to compare two different 3D display techniques, the anaglyph method which provides binocular depth cues and a method based on motion parallax, and to show that the motion parallax based approach to present 3D images on consumer 2D portable screen is an equivalent way in comparison to the above mentioned and well-known anaglyph method. The subsequently conducted subjective quality tests show that viewers even prefer wiggle over anaglyph stereoscopy mainly due to a better color reproduction and a comparable depth perception. Martin Rerábek, Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi |
MMSP | 4 |
| 2011 | Efficient video coding based on audio-visual focus of attention
Jong-Seok Lee, Francesca De Simone, Touradj Ebrahimi |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | Towards high efficiency video coding: Subjective evaluation of potential coding technologies
Francesca De Simone, Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi |
J. Vis. Commun. Image Represent. | 4 |
| 2011 | Subjective Quality Evaluation via Paired Comparison: Application to Scalable Video CodingabstractScalable video coding is a powerful solution for content delivery in many interactive multimedia services due to its adaptability to varying terminal and network constraints. In order to successfully exploit such adaptability, it is necessary to understand users' preference among various scalability options and consequently develop an optimal bit rate adaptation strategy. In this paper, we present a study of subjective quality assessment of scalable video coding, which investigates the influence of the combination of scalability options on perceived quality with the goal of providing guidelines for an adaptive strategy that selects the optimal combination for a given bandwidth constraint. In particular, the study is based on paired comparison of stimuli that is suitable for our goal due to its simplicity and easiness. We propose a new method, called Paired Evaluation via Analysis of Reliability (PEAR), which analyzes paired comparison results and produces not only quality scores but also intuitive measures of confidence of the scores for significance analysis. Results and analysis of extensive subjective tests for two different scalable video codecs and high definition contents are described, from which general consistent conclusions are drawn. The video and subjective data used in the paper are publicly available to the research community. Jong-Seok Lee, Francesca De Simone, Touradj Ebrahimi |
IEEE Trans. Multim. | 3 |
| 2011 | Balancing Attended and Global Stimuli in Perceived Video Quality AssessmentabstractThe visual attention mechanism plays a key role in the human perception system and it has a significant impact on our assessment of perceived video quality. In spite of receiving less attention from the viewers, unattended stimuli can still contribute to the understanding of the visual content. This paper proposes a quality model based on the late attention selection theory, assuming that the video quality is perceived via two mechanisms: global and local quality assessment. First we model several visual features influencing the visual attention in quality assessment scenarios to derive an attention map using appropriate fusion techniques. The global quality assessment as based on the assumption that viewers allocate their attention equally to the entire visual scene, is modeled by four carefully designed quality features. By employing these same quality features, the local quality model tuned by the attention map considers the degradations on the significantly attended stimuli. To generate the overall video quality score, global and local quality features are combined by a content adaptive linear fusion method and pooled over time, taking the temporal quality variation into consideration. The experimental results have been compared to results from appropriate eye tracking and video quality assessment experiments, demonstrating promising performance. Junyong You, Jari Korhonen, Andrew Perkis, Touradj Ebrahimi |
IEEE Trans. Multim. | 4 |
| 2010 | Temporal synchronization in stereoscopic video: Influence on quality of experience and automatic asynchrony detectionabstractIn this paper, we analyze the influence of temporal asynchrony on the subjective quality of stereoscopic video. Based on our recently created 3D video database, different levels of asynchrony were simulated and a comprehensive subjective test was conducted to determine the associated degradations in quality of experience. Furthermore, we developed a method to detect asynchrony between left and right video streams based on canonical correlation analysis. Experiments demonstrate the robustness of this method with respect to different amounts of asynchrony and scene depth, which makes it suitable to predict quality of experience or automatic resynchronization. Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi |
ICIP | 3 |
| 2010 | Spatial noise shaping using convex optimization for perceptual image codingabstractIn this paper we propose a new convex optimization framework for precise spatial noise shaping. The effectiveness of this new technique is demonstrated in the application of perceptual coding of images. A modified JPEG 2000 codec is implemented using the proposed new framework and compared with existing perceptual coding algorithms. Results of subjective tests show that the new framework can provide a significant improvement in bitrate savings compared to the best performing wavelet domain technique. The algorithm allows much more precise control of distortion than existing spatial domain techniques and is fully compliant with part 1 of the JPEG 2000 standard. Mark R. Pickering, Junyong You, Touradj Ebrahimi, Andrew Perkis |
ICIP | 3 |
| 2010 | A perceptual quality metric for stereoscopic crosstalk perceptionabstractCompared to metrics proposed to assess the quality of two-dimensional (2D) images, there are very few metrics devoted to quality assessment of stereoscopic presentations. Crosstalk is one of the most annoying distortions in the visualization stage of stereoscopic imaging technology. This paper proposes a perceptual quality metric which takes characteristics of stereoscopic images into account for predicting quality levels of crosstalk perception in stereoscopic images, based on an understanding of three main factors, crosstalk level, camera baseline and scene content. The experimental results demonstrate that the proposed metric has Pearson correlation of 87.7% when compared to the ground truth results from the subjective experiments on the crosstalk perception, which is much better than the traditional 2D metrics without integrating 3D depth information. Liyuan Xing, Junyong You, Touradj Ebrahimi, Andrew Perkis |
ICIP | 3 |
| 2010 | Implicit retrieval of salient images using Brain Computer InterfaceabstractSpace missions are often equipped with several high definition sensors that can autonomously collect a potentially enormous amount of data. The bottleneck in retrieving these often precious datasets is the onboard data storing capability and the communication bandwidth, which limit the amount of data that can be sent back to Earth. In this paper, we propose a method based on the analysis of brain electrical activity to identify the scientific interest of experts towards a given image in a large set of images. Such a method can be used to efficiently create an abundant training set (images and whether they are scientifically interesting) with a considerably faster image presentation rate that can go beyond expert consciousness, with less interrogation time for experts and relatively high performance. Ashkan Yazdani, Jean-Marc Vesin, Dario Izzo, Christos Ampatzis, Touradj Ebrahimi |
ICIP | 5 |
| 2010 | A framework for the validation of privacy protection solutions in video surveillanceabstractThe issue of privacy protection in video surveillance has drawn a lot of interest lately. However, thorough performance analysis and validation is still lacking, especially regarding the fulfillment of privacy-related requirements. In this paper, we put forward a framework to assess the capacity of privacy protection solutions to hide distinguishing facial information and to conceal identity. We then conduct rigorous experiments to evaluate the performance of face recognition algorithms applied to images altered by privacy protection techniques. Results show the ineffectiveness of naïve privacy protection techniques such as pixelization and blur. Conversely, they demonstrate the effectiveness of more sophisticated scrambling techniques to foil face recognition. Frédéric Dufaux, Touradj Ebrahimi |
ICME | 2 |
| 2010 | Gesture and touch controlled video player interface for mobile devicesabstractToday, mobile communication devices allow users to access a wide variety of multimedia contents and services. In order to improve user experience and device usability, the design of interfaces and interaction techniques for mobile devices have focused on new modalities, other than those used for desktop computers. In this paper, we describe a novel gesture controlled video player interface for mobile devices. The results of a usability study confirm that users would definitely like to adopt the major part of the proposed features. Furthermore, the responsiveness and reliability of the interface has been studied. Measured response times have been found to be within acceptable boundaries and the number of unrecognised haptic controls is limited. Shelley Buchinger, Ewald Hotop, Helmut Hlavacs, Francesca De Simone, Touradj Ebrahimi |
ACM Multimedia | 5 |
| 2010 | Subjective evaluation of scalable video coding for content distributionabstractThis paper investigates the influence of the combination of the scalability parameters in scalable video coding (SVC) schemes on the subjective visual quality. We aim at providing guidelines for an adaptation strategy of SVC that can select the optimal scalability options for resource-constrained networks. Extensive subjective tests are conducted by using two different scalable video codecs and high definition contents. The results are analyzed with respect to five dimensions, namely, codec, content, spatial resolution, temporal resolution, and frame quality. Jong-Seok Lee, Francesca De Simone, Naeem Ramzan, Zhijie Zhao, Engin Kurutepe, Thomas Sikora, Jörn Ostermann, Ebroul Izquierdo, Touradj Ebrahimi |
ACM Multimedia | 9 |
| 2010 | Chroma space: affective colors in interactive 3d worldabstractWe have developed an installation called Chroma Space to serve as a platform for experimenting the novel usage of affective colors in an interactive synthetic scenario. Chroma Space demonstrates the effective impacts of using a stylistic approach to address emotional sensations, by making colors move in space. We conducted a study to assess the impact of exposure to achromatic colors to different interactive 3D scenarios. The results suggest that presentation of color has an emotional impact to viewers. We argue that designers of synthetic 3D environments should consider the application of stylistic or thematic use of color on screen to increase the emotional attractiveness of an application. Wendy Ann Mansilla, Jordi Puig, Andrew Perkis, Touradj Ebrahimi |
ACM Multimedia | 4 |
| 2010 | Encoder and decoder side global and local motion estimation for Distributed Video CodingabstractIn this paper, we propose a new Distributed Video Coding (DVC) architecture where motion estimation is performed both at the encoder and decoder, effectively combining global and local motion models. We show that the proposed approach improves significantly the quality of Side Information (SI), especially for sequences with complex motion patterns. In turn, it leads to rate-distortion gains of up to 1 dB when compared to the state-of-the-art DISCOVER DVC codec. Frédéric Dufaux, Touradj Ebrahimi |
MMSP | 2 |
| 2010 | An objective metric for assessing quality of experience on stereoscopic imagesabstractMost quality models for stereoscopic presentations are dedicated to measuring quality degradation caused by compression artefacts. However, non-compression distortions induced during acquisition and presentation usually have significant influence on 3D viewing experience. In this paper, we propose an objective metric for viewing experience assessment by taking camera baseline and binocular distortion crosstalk into consideration. In particular, the proposed metric is based on our previous work on both subjective evaluation and objective assessment of crosstalk perception. Results on a publicly available stereoscopic quality database demonstrate that the proposed metric can achieve more than 87% correlation with subjective assessment of viewing experience. Liyuan Xing, Junyong You, Touradj Ebrahimi, Andrew Perkis |
MMSP | 3 |
| 2010 | Subjective evaluation of stereoscopic crosstalk perceptionabstractWhile the causes and nature of crosstalk, as well as crosstalk reduction techniques have been extensively studied, it is still difficult to eliminate. Perceptually, crosstalk is one of the most annoying distortions in the visualization stage of stereoscopic imaging. Therefore, to understand how users perceive crosstalk is of fundamental importance to improve the quality of 3D presentations. In this paper, we aim at analyzing the impact of crosstalk level, camera baseline and scene content on users' perception of crosstalk. Extensive subjective tests are conducted and the opinion scores are statistically analyzed and discussed. The results indicate that crosstalk level, camera baseline, as well as scene content all have major impacts on the perception of crosstalk. We also show that these three factors correlate with each other in terms of impact on the crosstalk perception. Furthermore, we propose a content descriptor for crosstalk perception (CDCP) and show its effectiveness. Liyuan Xing, Touradj Ebrahimi, Andrew Perkis |
VCIP | 2 |
| 2009 | Two-Level Bimodal Association for Audio-Visual Speech Recognition
Jong-Seok Lee, Touradj Ebrahimi |
ACIVS | 2 |
| 2009 | Towards Generic Detection of Unusual Events in Video SurveillanceabstractIn this paper, we consider the challenging problem of unusual event detection in video surveillance systems. The proposed approach makes a step toward generic and automatic detection of unusual events in terms of velocity and acceleration. At first, the moving objects in the scene are detected and tracked. A better representation of moving objects trajectories is then achieved by means of appropriate pre-processing techniques. A supervised support vector machine method is then used to train the system with one or more typical sequences, and the resulting model is then used for testing the proposed method with other typical sequences (different scenes and scenarios). Experimental results are shown to be promising. The presented approach is capable of determining similar unusual events as in the training sequences. Frédéric Dufaux, Thien M. Ha, Touradj Ebrahimi |
AVSS | 4 |
| 2009 | Video coding based on audio-visual attentionabstractThis paper proposes an efficient video coding method based on audio-visual attention, which is motivated by the fact that cross-modal interaction significantly affects humans' perception of multimedia content. First, we propose an audio-visual source localization method to locate the sound source in a video sequence. Then, its result is used for applying spatial blurring to video frames in order to reduce redundant high-frequency information and achieve coding efficiency. We demonstrate the effectiveness of the proposed method for H.264/AVC coding along with the results of a subjective evaluation. Jong-Seok Lee, Francesca De Simone, Touradj Ebrahimi |
ICME | 3 |
| 2009 | Analysis of the Limits of Graph-Based Object Duplicate DetectionabstractSeveral applications require accurate and efficient object duplicate detection methods, such as automatic video and image tag propagation, video surveillance, and high level image or video search. In this paper, we explore the limits of our recently proposed graph-based object duplicate detection method. The dependency of the performance with respect to the number of training images is assessed and the optimal detection parameters are determined. Furthermore, the differences among various object classes are analyzed. In this way, this paper provides an in-depth analysis of the graph based object duplicate detection method. Peter Vajda, Lutz Goldmann, Touradj Ebrahimi |
ISM | 3 |
| 2009 | Quality of multimedia experience: past, present and futureabstractThis talk starts by defining what is Quality of Experience. It then provides an overview of state of the art in Quality of Experience in multimedia systems. It will finally conclude by presenting challenges and trends that need to be further addressed. Touradj Ebrahimi |
ACM Multimedia | 1 |
| 2009 | Multimodal person search combining information fusion and relevance feedbackabstractWith the increasing amount of multimedia data, efficient tools for search and retrieval are needed. Since people are naturally one of the most interesting objects within these documents, a system for multimodal person search and retrieval has been developed. It combines the audiovisual analysis of persons with the query by example paradigm and relevance feedback to provide an efficient tool for searching multimedia data. For the relevance feedback, one and two class approaches are considered and compared to each other. Multimodal fusion techniques are used to exploit the complementary character of the audio and video information. The experimental results prove that multimodal person search and retrieval is feasible and more efficient than manual exploration. Lutz Goldmann, Amjad Samour, Touradj Ebrahimi, Thomas Sikora |
MMSP | 3 |
| 2009 | Efficient video coding in H.264/AVC by using audio-visual informationabstractThis paper proposes an efficient video coding method which utilizes audio-visual information, based on the observation that sound-emitting regions in a video sequence attract observer's attention. The regions responsible for the sound are identified by an audio-visual source localization algorithm. Then, the result is used for encoding different regions in the scene with different quality in such a way that a region far from the sound source is coded with a lesser quality than the sound-emitting regions. This is implemented by assigning different quantization parameter values for different regions in H.264/AVC. Experimental results demonstrate the effectiveness of the proposed approach. Jong-Seok Lee, Touradj Ebrahimi |
MMSP | 2 |
| 2009 | Error-resilient scalable compression based on distributed video coding
Mourad Ouaret, Frédéric Dufaux, Touradj Ebrahimi |
Signal Process. Image Commun. | 3 |
| 2008 | Towards Fully Automatic Image Segmentation Evaluation
Lutz Goldmann, Tomasz Adamek, Peter Vajda, Mustafa Karaman, Roland Mörzinger, Eric Galmar, Thomas Sikora, Noel E. O'Connor, Thien Ha-Minh, Touradj Ebrahimi, Peter Schallauer, Benoit Huet |
ACIVS | 10 |
| 2008 | H.264/AVC video scrambling for privacy protectionabstractIn this paper, we address the problem of privacy in video surveillance systems. More specifically, we consider the case of H.264/AVC which is the state-of-the-art in video coding. We assume that regions of interest (ROI), containing privacy-sensitive information, have been identified. The content of these regions are then concealed using scrambling. More specifically, we introduce two region-based scrambling techniques. The first one pseudo-randomly flips the sign of transform coefficients during encoding. The second one is performing a pseudo-random permutation of transform coefficients in a block. The flexible macroblock ordering (FMO) mechanism of H.264/AVC is exploited to discriminate between the ROI which are scrambled and the background which remains clear. Experimental results show that both techniques are able to effectively hide private information in ROI, while the scene remains comprehensible. Furthermore, the loss in coding efficiency stays small, whereas the required additional computational complexity is negligible. Frédéric Dufaux, Touradj Ebrahimi |
ICIP | 2 |
| 2008 | Improved side information generation with iterative decoding and frame interpolation for Distributed Video CodingabstractDistributed Video Coding (DVC) is a new paradigm in video coding, which is receiving a lot of interests nowadays. Side Information (SI) generation is a key function in the DVC decoder, and plays a key-role in determining the performance of the codec. This paper proposes an improved side information generation scheme, which exploits both spatial and temporal correlations in the sequences. Partially decoded Wyner-Ziv (WZ) frames, based on initial SI by Motion Compensation Temporal Interpolation (MCTI), are exploited to improve the performance of the whole SI generation. In addition, an enhanced temporal frame interpolation is proposed, including motion vector refinement and smoothing, optimal compensation mode selection, and a new matching criterion for motion estimation. Simulation results show that the proposed scheme can achieve up to 2.3 dB improvement in Rate Distortion (RD) performance for video with high motion, when compared to state-of-the-art DVC. Shuiming Ye, Mourad Ouaret, Frédéric Dufaux, Touradj Ebrahimi |
ICIP | 4 |
| 2008 | Hybrid spatial and temporal error concealment for distributed video codingabstractDistributed video coding (DVC) is based on a new paradigm in coding, which has received many interests recently. This paper proposes a hybrid spatial and temporal error concealment scheme to conceal errors in Wyner-Ziv (WZ) frames. We first use a spatial concealment based on edge directed filter. This step is exploited to improve the performance of subsequent temporal concealment. An enhanced temporal concealment based on motion compensated temporal interpolation is proposed, including motion vector refinement and smoothing, optimal compensation mode selection, and a new matching criterion for motion estimation. Simulation results show that the objective qualities as well as the perceptual qualities of the corrupted sequences are significantly improved by the hybrid error concealment, outperforming both spatial and temporal concealments alone. Shuiming Ye, Mourad Ouaret, Frédéric Dufaux, Touradj Ebrahimi |
ICME | 4 |
| 2008 | Distributed Video Coding: Selecting the most promising application scenarios
Fernando Pereira 0001, Christine Guillemot, Touradj Ebrahimi, Riccardo Leonardi, Sven Klomp |
Signal Process. Image Commun. | 4 |
| 2008 | Scrambling for Privacy Protection in Video Surveillance SystemsabstractIn this paper, we address the problem of privacy protection in video surveillance. We introduce two efficient approaches to conceal regions of interest (ROIs) based on transform-domain or codestream-domain scrambling. In the first technique, the sign of selected transform coefficients is pseudorandomly flipped during encoding. In the second method, some bits of the codestream are pseudorandomly inverted. We address more specifically the cases of MPEG-4 as it is today the prevailing standard in video surveillance equipment. Simulations show that both techniques successfully hide private data in ROIs while the scene remains comprehensible. Additionally, the amount of noise introduced by the scrambling process can be adjusted. Finally, the impact on coding efficiency performance is small, and the required computational complexity is negligible. Frédéric Dufaux, Touradj Ebrahimi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Efficient Rotation-Discriminative Template Matching
David Marimon, Touradj Ebrahimi |
CIARP | 2 |
| 2007 | Codec-Independent Scalable Distributed Video CodingabstractIn this paper, we introduce novel schemes for scalable distributed video coding (DVC), dealing with temporal, spatial and quality scalabilities. More specifically, conventional coding is used to obtain a base layer. DVC is then applied to generate enhancement layers. The side information is generated either temporally by motion compensated interpolation, or spatially by a spatial bi-cubic interpolation. Note that this scalable DVC approach is independent from the codec used to encode or decode the base layer. Simulation results show that most of the proposed schemes outperform non-scalable DVC, in addition to enabling the scalability features. Mourad Ouaret, Frédéric Dufaux, Touradj Ebrahimi |
ICIP (3) | 3 |
| 2007 | Recent advances in brain-computer interfacesabstractA brain-computer interface (BCI) is a communication system that translates brain activity into commands for a computer or other devices. In other words, a BCI allows users to act on their environment by using only brain activity, without using peripheral nerves and muscles. The major goal of BCI research is to develop systems that allow disabled users to communicate with other persons, to control artificial limbs, or to control their environment. To achieve this goal, many aspects of BCI systems are currently being investigated. Research areas include evaluation of invasive and noninvasive technologies to measure brain activity, evaluation of control signals (i.e. patterns of brain activity that can be used for communication), development of algorithms for translation of brain signals into computer commands, and the development of new BCI applications. In this paper we give an overview of the aspects of BCI research mentioned above and highlight recent developments and open problems. Touradj Ebrahimi |
MMSP | 1 |
| 2007 | Feature point tracking combining the Interacting Multiple Model filter and an efficient assignment algorithmabstractAn algorithm for feature point tracking is proposed. The Interacting Multiple Model (IMM) filter is used to estimate the state of a feature point. The problem of data association, i.e. establishing which feature point to use in the state estimator, is solved by an assignment algorithm. A track management method is also developed. In particular a track continuation method and a track quality indicator are presented. The evaluation of the tracking system on real sequences shows that the IMM filter combined with the assignment algorithm outperforms the Kalman filter, used with the Nearest Neighbour (NN) filter, in terms of data association performance and robustness to sudden feature point manoeuvre. David Marimon, Yousri Abdeljaoued, Bruno Palacios, Touradj Ebrahimi |
VCIP | 4 |
| 2007 | Particle filter-based camera tracker fusing marker and feature point cuesabstractThis paper presents a video-based camera tracker that combines marker-based and feature point-based cues within a particle filter framework. The framework relies on their complementary performances. On the one hand, marker-based trackers can robustly recover camera position and orientation when a reference (marker) is available but fail once the reference becomes unavailable. On the other hand, filter-based camera trackers using feature point cues can still provide predicted estimates given the previous state. However, the trackers tend to drift and usually fail to recover when the reference reappears. Therefore, we propose a fusion where the estimate of the filter is updated from the individual measurements of each cue. The particularity of the fusion filter is to manipulate different sorts of cues in a single framework. The framework keeps a single motion model and its prediction is corrected by one cue at a time. More precisely, the marker-based cue is selected when the reference is available whereas the feature point-based cue is selected otherwise. The filter's state is updated by switching between two different likelihood distributions. Each likelihood distribution is adapted to the type of measurement (cue). Evaluations on real cases show that the fusion of these two approaches outperforms the individual tracking results. David Marimon, Yannick Maret, Yousri Abdeljaoued, Touradj Ebrahimi |
VCIP | 4 |
| 2007 | Watermarked 3-D Mesh Quality AssessmentabstractThis paper addresses the problem of assessing distortions produced by watermarking 3D meshes. In particular, a new methodology for subjective evaluation of the quality of 3D objects is proposed and implemented. Two objective metrics derived from measures of surface roughness are then proposed and their efficiency to predict the perceptual impact of 3D watermarking is assessed and compared with the state of the art. Results obtained show good correlations between the proposed objective metrics and subjective assessments by human observers Massimiliano Corsini, Elisa Drelie Gelasca, Touradj Ebrahimi, Mauro Barni |
IEEE Trans. Multim. | 3 |
| 2006 | Spatial filters for the classification of event-related potentials
Ulrich Hoffmann 0001, Jean-Marc Vesin, Touradj Ebrahimi |
ESANN | 3 |
| 2006 | A Novel Replica Detection System using Binary Classifiers, R-Trees, and PCAabstractReplica detection is a prerequisite for the discovery of copyright infringement and detection of illicit content. For this purpose, content-based systems can be an efficient alternative to watermarking. Rather than imperceptibly embedding a signal, content-based systems rely-on image similarity. Certain content-based systems use adaptive classifiers to detect replicas. In such systems, a suspect image is tested against every original, which can become computationally prohibitive as the number of original images grows. In this paper, we propose using R-tree indexing to decrease the necessary number of comparisons and rapidly select the most likely originals. Experimental results show that the proposed system performs very satisfactorily and that up to 99.3% of the originals can be discarded before applying the binary classifiers. Yannick Maret, Spiros Nikolopoulos, Frédéric Dufaux, Touradj Ebrahimi, Nikos Nikolaidis 0001 |
ICIP | 4 |
| 2006 | The emerging JPEG-2000 security (JPSEC) standardabstractThe emergence of digital imaging applications is accelerating the need for security of digital imagery. The emerging international standard ISO/IEC JPEG-2000 security (JPSEC) is designed to provide security for digital imagery, and in particular digital imagery coded with the JPEG-2000 image coding standard. This paper provides an overview of the JPSEC standard, including a description of its basic architecture and examples of its use. John G. Apostolopoulos, Susie J. Wee, Frédéric Dufaux, Touradj Ebrahimi, Qibin Sun, Zhishou Zhang |
ISCAS | 4 |
| 2006 | JPEG2000 image coding system theory and applicationsabstractJPEG2000, the new standard for still image coding, provides a new framework and an integrated toolbox to better address increasing needs for compression. It offers a wide range of functionalities such as lossless and lossy coding, embedded lossy to lossless coding, progression by resolution and quality, high compression efficiency, error resilience and region-of-interest (ROI) coding. Comparative results have shown that JPEG2000 is indeed superior to established image compression standards. Overall, the JPEG2000 standard offers the richest set of features in a very efficient way and within a unified algorithm. The price of this is its additional complexity, but this should not be perceived as a disadvantage, as the technology evolves rapidly Athanassios N. Skodras, Touradj Ebrahimi |
ISCAS | 2 |
| 2006 | Adaptive image replica detection based on support vector classifiers
Yannick Maret, Frédéric Dufaux, Touradj Ebrahimi |
Signal Process. Image Commun. | 3 |
| 2005 | Surveillance video for mobile devicesabstractIn this paper, we present a video encoding scheme that uses object-based adaptation to deliver surveillance video to mobile devices. The method relies on a set of complementary video adaptation strategies and generates content that matches various appliance and network resources. Prior to encoding, some of the adaptation strategies exploit video object segmentation and selective filtering in order to improve the perceived quality. Moreover, object segmentation enables the generation of automatic summaries and of simplified versions of the monitored scene. The performance of individual adaptation strategies is assessed using an objective video quality metric, which is also used to select the strategy that provides maximum value for the user under a given set of constraints. We demonstrate the effectiveness of the scheme on standard surveillance test sequences and realistic mobile client resource profiles. Olivier Steiger, Touradj Ebrahimi, Andrea Cavallaro |
AVSS | 2 |
| 2005 | Objective evaluation of the perceptual quality of 3D watermarkingabstractIn this paper an objective metric to measure the perceptual quality of watermarked 3D meshes is presented. The metric, which is based on a black-box approach, relies on the measurement of the roughness of 3D meshes before and after the insertion of the watermark. To calibrate the metric and to validate it, a set of psychovisual experiments has been carried out. Due to the lack of prior work in this field, a new methodology for the subjective evaluation of the quality of watermarked 3D objects is introduced. The validity of the proposed metric has been tested against a number of different 3D watermarking algorithms, showing an excellent match with the subjective evaluation of the quality stemming from the psychovisual experiments. Elisa Drelie Gelasca, Touradj Ebrahimi, Massimiliano Corsini, Mauro Barni |
ICIP (1) | 2 |
| 2005 | Evaluating Perceptually Prefiltered VideoabstractPerceptual prefiltering is the process of enhancing relevant portions of an image or of a video, and of simplifying contextual information in order to improve the perceived quality or the compression ratio. In this paper, we discuss the results of subjective quality evaluation experiments performed to assess the impact of perceptual prefiltering on video coding and we propose an objective quality metric that mimics the behavior of human observers. The predicted performance of the proposed metric is consistent with the subjective evaluation scores. Experimental results demonstrate that perceptual prefiltering leads to quality improvements by up to 10% at low bit rates Olivier Steiger, Touradj Ebrahimi, Andrea Cavallaro |
ICME | 2 |
| 2005 | Tracking Video Objects in Cluttered BackgroundabstractWe present an algorithm for tracking video objects which is based on a hybrid strategy. This strategy uses both object and region information to solve the correspondence problem. Low-level descriptors are exploited to track object's regions and to cope with track management issues. Appearance and disappearance of objects, splitting and partial occlusions are resolved through interactions between regions and objects. Experimental results demonstrate that this approach has the ability to deal with multiple deformable objects, whose shape varies over time. Furthermore, it is very simple, because the tracking is based on the descriptors, which represent a very compact piece of information about regions, and they are easy to define and track automatically. Finally, this procedure implicitly provides one with a description of the objects and their track, thus enabling indexing and manipulation of the video content. Andrea Cavallaro, Olivier Steiger, Touradj Ebrahimi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | Semantic video analysis for adaptive content delivery and automatic descriptionabstractWe present an encoding framework which exploits semantics for video content delivery. The video content is organized based on the idea of main content message. In the work reported in this paper, the main content message is extracted from the video data through semantic video analysis, an application-dependent process that separates relevant information from non relevant information. We use here semantic analysis and the corresponding content annotation under a new perspective: the results of the analysis are exploited for object-based encoders, such as MPEG-4, as well as for frame-based encoders, such as MPEG-1. Moreover, the use of MPEG-7 content descriptors in conjunction with the video is used for improving content visualization for narrow channels and devices with limited capabilities. Finally, we analyze and evaluate the impact of semantic video analysis in video encoding and show that the use of semantic video analysis prior to encoding sensibly reduces the bandwidth requirements compared to traditional encoders not only for an object-based encoder but also for a frame-based encoder. Andrea Cavallaro, Olivier Steiger, Touradj Ebrahimi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | Feature point extraction using scale-space representationabstractAn algorithm for feature point extraction is presented. It is based on a scale-space representation of the image as well as a system for tracking across scales. Using synthetic and real images, it is shown that the proposed algorithm produces stable and well-localized feature points estimates two essential properties for video applications. Yousri Abdeljaoued, Touradj Ebrahimi |
ICIP | 2 |
| 2004 | Annoyance of spatio-temporal artifacts in segmentation quality assessmentabstractThis paper describes the results of a series of subjective experiments that investigated the annoyance caused by the most common artifacts present in segmented video sequences. Various types of artifacts were inserted into a reference segmented video, considered as ideal, and shown to our test subjects. The artifacts varied in their location, size, appearance and duration. Annoyance of segmentation artifacts are found to be tied up with their intrinsic characteristics (e.g., size, position) but only weakly related to the video content. The results identify the characteristics that should be taken into account in the design of a perceptually driven objective metric. Elisa Drelie Gelasca, Touradj Ebrahimi, Mylène C. Q. Farias, Marco Carli, Sanjit K. Mitra |
ICIP | 2 |
| 2004 | Error-resilient video coding performance analysis of motion JPEG2000 and MPEG-4abstractThe new Motion JPEG 2000 standard is providing with some compelling features. It is based on an intra-frame wavelet coding, which makes it very well suited for wireless applications. Indeed, the state-of-the-art wavelet coding scheme achieves very high coding efficiency. In addition, Motion JPEG 2000 is very resilient to transmission errors as frames are coded independently (intra coding). Furthermore, it requires low complexity and introduces minimal coding delay. Finally, it supports very efficient scalability. In this paper, we analyze the performance of Motion JPEG 2000 in error-prone transmission. We compare it to the well-known MPEG-4 video coding scheme, in terms of coding efficiency, error resilience and complexity. We present experimental results which show that Motion JPEG 2000 outperforms MPEG-4 in the presence of transmission errors. Frédéric Dufaux, Touradj Ebrahimi |
VCIP | 2 |
| 2004 | Cast shadow segmentation using invariant color features
Elena Salvador, Andrea Cavallaro, Touradj Ebrahimi |
Comput. Vis. Image Underst. | 3 |
| 2004 | Perceptual blur and ringing metrics: application to JPEG2000
Pina Marziliano, Frédéric Dufaux, Stefan Winkler 0001, Touradj Ebrahimi |
Signal Process. Image Commun. | 4 |
| 2003 | Moving Object Detection Between Multiple and Color ImagesabstractThere are several publications dedicated to the description and analysis of change detection between two gray-value images. This paper introduces new methods to detect moving objects between multiple images and to detect changes between color images or any type of multispectral images. We are not aware of methods giving the possibility of detecting color changes and changes between multiple frames. All the proposed change detectors are based on the Gramian determinant, which provides low computational cost and is easy to implement. These features are very important due to the additional complexity of change detection between multiple as well as color images. Emrullah Durucan, Touradj Ebrahimi |
AVSS | 2 |
| 2003 | Secure JPEG 2000-JPSECabstractThe Joint Photographic Experts Group (JPEG) has recently rolled out a new still image coding standard called JPEG 2000. This standard integrates an efficient image compression scheme with functionalities required by multimedia applications, such as progressiveness up to lossless coding, region of interest coding, and error resiliency. Security is a concern in many applications, and therefore also a desired functionality. This paper provides readers with insights and examples of how to combine security solutions with JPEG 2000 compression. Tools for JPEG 2000 compressed image integrity, access control, and copyright protection are presented. They can be either applied to a JPEG 2000 codestream or directly integrated into the coding/decoding operations, resulting in a fully compliant JPEG 2000 image. Touradj Ebrahimi, Raphaël Grosbois |
ICASSP (4) | 1 |
| 2003 | Correlative exploration of EEG signals for direct brain-computer communicationabstractIn this study we present a method for classifying EEG signals based on the information content of their correlative time-frequency-space representation (CTFSR). A support vector machine (SVM) kernel is proposed that can be calculated in the time domain while it computes a similarity measure in the CTFSR space. This classification method is used in a brain-computer interface (BCI) application. The use of the SVM approach allows us to propose a simple strategy for adapting the BCI to possible long term variations in the brain activity. Gary Garcia Molina, Touradj Ebrahimi, Jean-Marc Vesin |
ICASSP (5) | 2 |
| 2003 | MPEG-based personalized content deliveryabstractIn this paper, we present a personalized multimedia content delivery system dealing with both user preferences and terminal/network capabilities. In order to ease interoperability with third-party applications, content annotation, user preferences handling and terminal/network capabilities description are managed by MPEG-7 and MPEG-21 standards. Offline content adaptation and annotation tools are proposed for content preparation. Content delivery is handled by client-server architecture. This architecture is built around custom personalization and adaptation tools, and a commercial content streamer. The proposed system has been used to provide universal multimedia access within the EC-funded R&D project PERSEO. Olivier Steiger, David Marimon, Touradj Ebrahimi |
ICIP (3) | 3 |
| 2003 | Semantic segmentation and description for video transcodingabstractWe present an automatic content-based video transcoding algorithm, which is based on how humans perceive visual information. The transcoder support multiple video objects and their description. First the video is decomposed into meaningful objects through semantic segmentation. Then the transcoder adapts its behavior to code relevant (foreground) and non-relevant objects differently. Both objects-based and frame-based encoders are combined with semantic segmentation. Experimental results show that the use of semantics and description prior to transcoding reduces the bandwidth requirements and makes it possible to adapt the video representation to limited network and terminal device capabilities still retaining the essential information. Andrea Cavallaro, Olivier Steiger, Touradj Ebrahimi |
ICME | 3 |
| 2003 | Object-based video: extraction tools, evaluation metrics, and applications
Andrea Cavallaro, Touradj Ebrahimi |
VCIP | 2 |
| 2003 | Intuitive strategy for parameter setting in video segmentation
Elisa Drelie Gelasca, Elena Salvador, Touradj Ebrahimi |
VCIP | 3 |
| 2003 | Non-linear subdivision using local spherical coordinates
Nicolas Aspert, Touradj Ebrahimi, Pierre Vandergheynst |
Comput. Aided Geom. Des. | 2 |
| 2003 | Special Issue on Image security: secure imaging - is it necessary?
Touradj Ebrahimi, Benoît Macq |
Signal Process. Image Commun. | 1 |
| 2002 | Objective evaluation of segmentation quality using spatio-temporal contextabstractIn this paper, we propose an automatic method for the objective evaluation of segmentation results. The method is based on computing the deviation of the segmentation results from a reference segmentation. The discrepancy between two results is weighted based on spatial and temporal contextual information, by taking into account the way humans perceive visual information. The metric is useful for applications where the final judge of the quality is a human observer or the results of segmentation are otherwise processed in a human-like fashion. The proposed evaluation has been applied both to automatically provide a ranking among different segmentation algorithms and to optimally set the parameters of a given algorithm. Andrea Cavallaro, Elisa Drelie Gelasca, Touradj Ebrahimi |
ICIP (3) | 3 |
| 2002 | A no-reference perceptual blur metricabstractWe present a no-reference blur metric for images and video. The blur metric is based on the analysis of the spread of the edges in an image. Its perceptual significance is validated through subjective experiments. The novel metric is near real-time, has low computational complexity and is shown to perform well over a range of image content. Potential applications include optimization of source coding, network resource management and autofocus of an image capturing device. Pina Marziliano, Frédéric Dufaux, Stefan Winkler 0001, Touradj Ebrahimi |
ICIP (3) | 4 |
| 2002 | MESH: measuring errors between surfaces using the Hausdorff distanceabstractThis paper proposes an efficient method to estimate the distance between discrete 3D surfaces represented by triangular 3D meshes. The metric used is based on an approximation of the Hausdorff distance, which has been appropriately implemented in order to reduce unnecessary computation and memory usage. Results show that when compared to similar tools, a significant gain in both memory and speed can be achieved. Nicolas Aspert, Diego Santa Cruz, Touradj Ebrahimi |
ICME (1) | 3 |
| 2002 | Accurate video object segmentation through change detectionabstractWe propose an algorithm for the accurate extraction of video objects from color sequences. The semantics defining the video objects is motion, and the extraction algorithm is based on change detection. The color difference between frames is modeled so as to separate the contributions caused by sensor noise and illumination variations from those caused by meaningful objects. Sensor noise is eliminated by using a probability-based classification, and local illumination variations are removed using a knowledge-based approach that is formulated as a hypothesize-and-test scheme. Experimental results show that the proposed method provides accurate contours of multiple deformable objects, thus providing a reliable input to object-based applications such as those supported by the MPEG-4 and MPEG-7 standards. Andrea Cavallaro, Touradj Ebrahimi |
ICME (1) | 2 |
| 2002 | Multiple video object tracking in complex scenesabstractWe present an automatic video object tracking algorithm capable of dealing with multiple simultaneous objects. The tracking is based on interactions between high-level and low-level image analysis results. The high-level result is a partition defining video objects, and the low-level result is a partition formed by homogeneous regions. For each region, a set of characteristic descriptors is produced. These region descriptors, and not regions themselves, are used to track the regions (and thus the objects) along time. Track management issues such as appearance and disappearance of objects, splitting and partial occlusions are resolved through interactions between regions and objects. Defining the tracking based on the parts of objects, identified by region segmentation, has led to a flexible technique that exploits the nature of the video object tracking problem. Experimental results show that the proposed method is able to track multiple rigid and deformable objects in indoor and outdoor scenes. Andrea Cavallaro, Olivier Steiger, Touradj Ebrahimi |
ACM Multimedia | 3 |
| 2002 | Compression of parametric surfaces for efficient 3D model coding
Diego Santa Cruz, Touradj Ebrahimi |
VCIP | 2 |
| 2002 | MPEG-7 description of generic video objects for scene reconstruction
Olivier Steiger, Andrea Cavallaro, Touradj Ebrahimi |
VCIP | 3 |
| 2002 | Coding of 3D virtual objects with NURBS
Diego Santa Cruz, Touradj Ebrahimi |
Signal Process. | 2 |
| 2002 | JPEG 2000 performance evaluation and assessment
Diego Santa Cruz, Raphaël Grosbois, Touradj Ebrahimi |
Signal Process. Image Commun. | 3 |
| 2002 | JPEG 2000
Touradj Ebrahimi, Charilaos A. Christopoulos, Daniel T. Lee |
Signal Process. Image Commun. | 1 |
| 2001 | Shadow identification and classification using invariant color modelsabstractA novel approach to shadow detection is presented. The method is based on the use of invariant color models to identify and to classify shadows in digital images. The procedure is divided into two levels: first, shadow candidate regions are extracted; then, by using the invariant color features, shadow candidate pixels are classified as self shadow points or as cast shadow points. The use of invariant color features allows a low complexity of the classification stage. Experimental results show that the method succeeds in detecting and classifying shadows within the environmental constrains assumed as hypotheses, which are less restrictive than state-of-the-art methods with respect to illumination conditions and the scene's layout. Elena Salvador, Andrea Cavallaro, Touradj Ebrahimi |
ICASSP | 3 |
| 2001 | MPEG-7 cameraabstractAn MPEG-7 camera extends the capabilities of conventional cameras by analyzing its scene in order to generate a content-based description according to the recently approved MPEG-7 standard. This gives to the camera a large variety of current and potential applications, such as surveillance, augmented reality, and virtual display. This paper provides an overview of what is meant by an MPEG-7 camera, discusses the above mentioned applications, and provides an implementation example of such a camera using existing hardware products. Touradj Ebrahimi, Yousri Abdeljaoued, Rosa M. Figueras i Ventura, Òscar Divorra Escoda |
ICIP (3) | 1 |
| 2001 | Watermarking in the JPEG 2000 domainabstractThere are several advantages to combine, at one end, the image coding and watermark insertion operations and, at the other end, the image decoding and watermark extraction. We describe a spread-spectrum-based watermarking technique in the framework of the JPEG 2000 still image compression. We also show that, by re-using the JPEG 2000 wavelet domain for watermark embedding, the proposed watermarking scheme exhibits a high robustness with respect to attacks which may occur in many applications. Raphaël Grosbois, Touradj Ebrahimi |
MMSP | 2 |
| 2001 | Video object extraction based on adaptive background and statistical change detection
Andrea Cavallaro, Touradj Ebrahimi |
VCIP | 2 |
| 2001 | Change detection and background extraction by linear algebraabstractChange detection plays a very important role in real-time image analysis, e.g., detection of intruders. One key issue is robustness to varying illumination conditions. We propose two techniques for change detection that have been developed to deal with variations in illumination and background, with real-time capabilities. The foundations of these techniques are based on a vector model of images and on the exploitation of the concepts of linear dependence and linear independence. Furthermore, the techniques are compatible with physical photometry. A detailed description of the proposed detector and three state-of-the art change detectors is also provided. For the purposes of comparison, an evaluation procedure is presented consisting of both objective and subjective parts. This evaluation procedure results in a final performance value for each detector analyzed. Emrullah Durucan, Touradj Ebrahimi |
Proc. IEEE | 2 |
| 2001 | JPEG2000: The upcoming still image compression standard
Athanassios N. Skodras, Charilaos A. Christopoulos, Touradj Ebrahimi |
Pattern Recognit. Lett. | 3 |
| 2000 | An Analytical Study of JPEG 2000 FunctionalitiesabstractJPEG 2000, the new ISO/ITU-T standard for still image coding, is about to be finished. Other new standards have been recently introduced, namely JPEG-LS and MPEG-4 VTC. This paper compares the set of features offered by JPEG 2000, and how well they are fulfilled, versus JPEG-LS and MPEG-4 VTC, as well as the older but widely used JPEG and the PNG. The study concentrates on the set of supported features, although lossless and lossy progressive compression efficiency results are also reported. Each standard, and the principles of the algorithms behind them, are also described. As the results show, JPEG 2000 supports the widest set of features among the evaluated standards, while providing superior rate-distortion performance. Diego Santa Cruz, Touradj Ebrahimi |
ICIP | 2 |
| 2000 | MPEG-4 natural video coding - An overview
Touradj Ebrahimi, Caspar Horne |
Signal Process. Image Commun. | 1 |
| 1999 | Towards Second Generation Watermarking SchemesabstractThe digital watermarking schemes of today use pixels (samples in the case of audio), frequency or other transform coefficients to embed the information. The drawback of such schemes is that the watermark is not embedded in the perceptually significant portions of the data. We refer to such techniques as first generation watermarking schemes. In this paper we introduce the concept of second generation watermarking schemes which, unlike first generation watermarking schemes, employ the notion of data features. We propose a scheme based on point features in images using a scale interaction technique based on 2D continuous wavelets. The features are used to compute a Voronoi partition of the image. The watermark is embedded in each segment using spread spectrum watermarking. In the recovery process the same features are detected, and again used to partition the image. Then the watermark is extracted from each segment separately. Martin Kutter, Sushil K. Bhattacharjee, Touradj Ebrahimi |
ICIP (1) | 3 |
| 1999 | Region of interest coding in JPEG2000 for interactive client/server applicationsabstractThis paper presents an approach which allows an efficient on-the-fly decoding of regions of interest of an already encoded image without need for a complete decoding/encoding process. The basic method is compatible with the wavelet based compression scheme adopted for the JPEG2000 standard and has been already adopted in a previous verification model. This technique is specially of value in interactive client/server applications linked through narrowband networks, in which the client can request the server to tailor more efficiently the transmission of the desired information at little processing cost. Simulation results show significant improvement in reduction of transmission time and enhanced flexibility at the expense of a very small complexity and bit rate overhead. Diego Santa Cruz, Touradj Ebrahimi, Mathias Larsson Carlander, Joel Askelöf, Charilaos A. Christopoulos |
MMSP | 2 |
| 1999 | Prolog to - High-performance compression of visual information-a tutorial review- part I: still pictures
Olivier Egger, Pascal Fleury, Touradj Ebrahimi, Murat Kunt |
Proc. IEEE | 3 |
| 1999 | High-performance compression of visual information-a tutorial review. I. Still picturesabstractDigital images have become an important source of information in the modern world of communication systems. In their raw form, digital images require a tremendous amount of memory. Many research efforts have been devoted to the problem of image compression in the last two decades. Two different compression categories must be distinguished: lossless and lossy. Lossless compression is achieved if no distortion is introduced in the coded image. Applications requiring this type of compression include medical imaging and satellite photography. For applications such as video telephony or multimedia applications, some loss of information is usually tolerated in exchange for a high compression ratio. In this two-part paper, the major building blocks of image coding schemes are overviewed. Part I covers still image coding, and Part II covers motion picture sequences. In this first part, still image coding schemes have been classified into predictive, block transform, and multiresolution approaches. Predictive methods are suited to lossless and low-compression applications. Transform-based coding schemes achieve higher compression ratios for lossy compression but suffer from blocking artifacts at high-compression ratios. Multiresolution approaches are suited for lossy as well for lossless compression. At lossy high-compression ratios, the typical artifact visible in the reconstructed images is the ringing effect. New applications in a multimedia environment drove the need for new functionalities of the image coding schemes. For that purpose, second-generation coding techniques segment the image into semantically meaningful pairs. Therefore, parts of these methods have been adapted to work for arbitrarily shaped regions. In order to add another functionality, such as progressive transmission of the information, specific quantization algorithms must he defined. A final step in the compression scheme is achieved by the codeword assignment. Finally, coding results are presented which compare state-of-the-art techniques for lossy and lossless compression. The different artifacts of each technique are highlighted and discussed. Also, the possibility of progressive transmission is illustrated. Olivier Egger, Pascal Fleury, Touradj Ebrahimi, Murat Kunt |
Proc. IEEE | 3 |
| 1998 | Progressive Content-Based Shape Compression for Retrieval of Binary Images
Corinne Le Buhan Jordan, Touradj Ebrahimi, Murat Kunt |
Comput. Vis. Image Underst. | 2 |
| 1998 | Visual data compression for multimedia applicationsabstractThe compression of visual information in the framework of multimedia applications is discussed. To this end, major approaches to compress still as well as moving pictures are reviewed. The most important objective in any compression algorithm is that of compression efficiency. High-compression coding of still pictures can be split into three categories: waveform, second-generation, and fractal coding techniques. Each coding approach introduces a different artifact at the target bit rates. The primary objective of most ongoing research in this field is to mask these artifacts as much as possible to the human visual system. Video-compression techniques have to deal with data enriched by one more component, namely, the temporal coordinate. Either compression techniques developed for still images can be generalized for three-dimensional signals (space and time) or a hybrid approach can be defined based on motion compensation. The video compression techniques can then be classified into the following four classes: waveform, object-based, model-based, and fractal coding techniques. This paper provides the reader with a tutorial on major visual data-compression techniques and a list of references for further information as the details of each method. Touradj Ebrahimi, Murat Kunt |
Proc. IEEE | 1 |
| 1998 | Video segmentation based on multiple features for interactive multimedia applicationsabstractWe present a scheme for interactive video segmentation. A key feature of the system is the distinction between two levels of segmentation, namely, regions and object segmentation. Regions are homogeneous areas of the images, which are extracted automatically by the computer. Semantically meaningful objects are obtained through user interaction by grouping of regions according to the specific application. This splitting relieves the computer of ill-posed semantic problems, and allows a higher level of flexibility of the method. The extraction of regions is based on the multidimensional analysis of several image features by a spatially constrained fuzzy C-means algorithm. The local level of reliability of the different features is taken into account in order to adaptively weight the contribution of each feature to the segmentation process. Results on the extraction of regions as well as on the tracking of spatiotemporal objects are presented. Roberto Castagno, Touradj Ebrahimi, Murat Kunt |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | A simple and efficient binary shape coding technique based on bitmap representationabstractThis article presents a technique based on the JBIG algorithm for binary shape coding in both lossless and lossy modes. Because it is applied directly to the bitmap representing the shape information, it bypasses the overhead in computation of an intermediate contour representation and its associated conversions. This leads to a simpler algorithm which is more suitable for a larger class of shape data. In addition a mechanism is proposed which allows a rate control for lossy coding mode. Frank Bossen, Touradj Ebrahimi |
ICASSP | 2 |
| 1997 | Streaming of Photo-Realistic Texture Mapped on 3D SurfaceabstractWe present a novel technique for efficient coding of texture to be mapped on 3D landscapes. The technique enables to stream the data across the network using a back-channel. The use of wavelet and discrete cosine transforms is investigated and compared. This technology has been proposed has a tool for the emerging MPEG-4 standard. Stefan Horbelt, Frederic D. Jordan, Touradj Ebrahimi |
ICIP (2) | 3 |
| 1997 | Scalable Shape Representation for Content-Based Visual Data CompressionabstractTwo major classes of shape coding methods are reviewed, namely bitmap coding and contour coding, and in particular their scalable extensions. In addition to their absolute compression efficiency, we analyze their performance in the framework of a complete image/video coding scheme, and show that they bring complementary functionalities depending on the targeted application. Corinne Le Buhan Jordan, F. Bossan, Touradj Ebrahimi |
ICIP (1) | 3 |
| 1997 | MPEG-4 video verification model: A video encoding/decoding algorithm based on content representation
Touradj Ebrahimi |
Signal Process. Image Commun. | 1 |
| 1997 | Dynamic approach to visual data compressionabstractThis paper presents the Swiss Federal Institute of Technology (EPFL) proposal to MPEG-4 video coding standardization activity. The proposed technique is based on a novel approach to audio-visual data compression entitled dynamic coding. The newly born multimedia environment supports a plethora of applications which cannot be covered adequately by a single compression technique. Dynamic coding offers the opportunity to combine several compression techniques and segmentation strategies. Given a particular application, these two degrees of freedom can be constrained and assembled in order to produce a particular profile which meets the set of specifications dictated by the application. The basic principles of this approach are presented together with the data representation system. The major characteristics of dynamic coding are reviewed, along with simulation results showing the performance of such an approach in a very low bit-rate video coding environment. Emmanuel Reusens, Touradj Ebrahimi, Corinne Le Buhan Jordan, Roberto Castagno, Vincent Vaerman, Laurent Piron, Carmen de Sola Fabregas, Sushil K. Bhattacharjee, Frank Bossen, Murat Kunt |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | Dynamic coding of visual informationabstractThis paper introduces a novel approach to visual data compression. The approach, named dynamic coding, consists of an effective competition between several representation models used for describing data portions. The image data is represented as the union of several regions each approximated by a representation model locally appropriate. The dynamic coding concept leads to attractive features such as genericness, flexibility, and openness and is therefore particularly suited to a multimedia environment in which many types of applications are involved. Dynamic coding is a general proposal to visual data compression and many variations on the same theme may be designed. They differ by the particular procedure by which the data is segmented into objects and the local representation model selected. As an illustrative example, a video compression scheme based on the principles of dynamic coding is presented. This compression algorithm performs a joint optimization of the segmentation (restricted to a so-called generalized quadtree partition) together with the representation models associated with each data segment. Four representation models are competing namely, fractal, motion compensation, text and graphics, and background modes. Optimality is defined with respect to a rate-distortion tradeoff and the optimization procedure leads to a multicriterion segmentation. Emmanuel Reusens, Touradj Ebrahimi, Murat Kunt |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1996 | Arbitrarily-shaped wavelet packets for zerotree codingabstractIn order to satisfy the needs of new multimedia applications, the problem of content-based video coding has to be addressed. A new approach of object interior coding is proposed. It is based on an arbitrarily-shaped subband transform followed by a generalized embedded zerotree wavelet algorithm. It is shown that the proposed technique achieves good compression results and has additional properties such as being computationally efficient, keeping the same dimensionality in the transformed domain, being perfect reconstruction, and allowing a perfect rate control. In addition a lossless mode can be defined by using an appropriate filter bank. Olivier Egger, Touradj Ebrahimi, Murat Kunt |
ICASSP | 2 |
| 1996 | Advanced imaging systems curricula at EPFLabstractThis paper describes the current status of the image system engineering curriculum at the Swiss Federal Institute of Technology at Lausanne (Ecole Polytechnique Federale de Lausanne-EPFL). The responsibility of this curriculum is with the Signal Processing Laboratory of EPFL. Touradj Ebrahimi, Murat Kunt |
ICIP (1) | 1 |
| 1996 | Image quality prediction for bitrate allocationabstractIn image coding, the choice of a good image coding algorithm is very dependent on the image content. Based on this fact, dynamic coding algorithms have been designed. They try to find an optimal coding scheme for each image segment. They rely on an exhaustive search of the best coding algorithm. Evaluation of all algorithms is computationally very intensive and strongly limits the number of considered algorithms for a given application. Therefore, current standards rely on a single coding algorithm. This paper investigates a way to predict the coding quality from the image content. This prediction is based on a neural network. The coding quality is computed from image region features. Those features are easy and fast to compute, and are common to the whole set of considered coding algorithms. Therefore, the choice of the best algorithm can be based on those predicted coding qualities, and does not require the computation of all coding algorithms. The system is also fast enough to be used for dynamic bitrate allocation, and a simple algorithm to do this is proposed. Pascal Fleury, Julien Reichel, Touradj Ebrahimi |
ICIP (3) | 3 |
| 1996 | Dynamic video coding-an overviewabstractIn this paper, we present an overview of the dynamic coding approach, together with recent developments carried out in this framework. Dynamic coding is a general approach to the problem of visual data representation in the context of multimedia. This approach consists of a dynamic combination of multiple representation models and segmentation strategies. Given an application, these two degrees of freedom are assembled so as to yield a specific profile which meets the specifications dictated by the application. The data is represented as the union of data segments, each described within a locally appropriate representation model. In order to illustrate this approach, a video compression system, based on the principles of dynamic coding, is proposed in the context of video-telephone/conference applications. This algorithm has been submitted to the MPEG-4 committee as a proposal for the first round of tests in November 1995. Recent developments have been added: in particular, a procedure enabling the generation of an object-oriented scalable bitstream is presented here. In order to reduce the blocking artifacts which are noticeable at high compression ratios, a post-processing technique is proposed. Emmanuel Reusens, Roberto Castagno, Corinne Le Buhan Jordan, Laurent Piron, Touradj Ebrahimi, Murat Kunt |
ICIP (2) | 5 |
| 1996 | Matching error based criterion of region merging for joint motion estimation and segmentation techniquesabstractThis paper describes a region merging method for joint motion estimation and segmentation of digital video sequences. The region merging criterion is based on the measure of the matching error for a region when applying a previously estimated motion to it. A region adjacency graph is used for data representation, which allows a scan independent processing and gives a high-level view. The method is simple-shape object-oriented and starts from a block-based segmentation. The aim of the proposed technique is to define simple shaped objects in a scene using motion information and a simple test. Markus Schütz, Touradj Ebrahimi |
ICIP (2) | 2 |
| 1995 | Progressive video coding for storage applications
Olivier Egger, Touradj Ebrahimi |
ICIP (3) | 2 |
| 1995 | A Region Based Motion Compensated Video Codec for Very Low Bitrate ApplicationsabstractA motion compensated video coding technique is discussed where square macro-blocks in the coding process are replaced by arbitrary shaped regions. Different components of this technique are presented in detail, such as: segmentation, region shape coding, motion estimation, and mode decisions. Simulation results show good quality images with sharper details achieved at bitrates as low as 9.6 Kb/s. Touradj Ebrahimi, Homer H. Chen, Barry G. Haskell |
ISCAS | 1 |
| 1995 | New trends in very low bitrate video codingabstractThe interest in very low birate video coding has increased considerably with the creation of new services and applications requiring this class of bitrates in their function. This paper aims in presenting the current problems and requirements, as well as the efforts in very low bitrate video coding. It also provides a number of solutions in order to solve these problems and to fulfill desired requirements in very, low bitrate coding. The main emphasis was given to an approach called "second generation" which has already shown its potential to efficiently code still images at higher compression factors. In particular, three techniques will be discussed in details, because of their ability to successfully achieve very low bitrate video coding with a reasonably good quality. The simulation results obtained by these techniques under the exact same conditions are presented and the merits and drawbacks of each one are pointed out throughout the paper.> Touradj Ebrahimi, Emmanuel Reusens |
Proc. IEEE | 1 |
| 1994 | A New Technique for Motion Field Segmentation and Coding for Very Low Birate Video Coding ApplicationsabstractMotion estimation is a key issue in video coding. In very low bitrate applications, the amount of the side information for the motion field represents an important portion of the total bitrate. This paper presents a joint motion estimation, segmentation and coding technique, which tries to reduce the motion side-information while providing a similar or smaller prediction error when compared to more classical motion estimation techniques.> Touradj Ebrahimi |
ICIP (2) | 1 |
| 1993 | Perceptually derived localized linear operators: Application to image sequence compression
Touradj Ebrahimi |
Signal Process. | 1 |
| 1992 | Application of an optimally localized and fast wavelet transform in image compressionabstractThe general subband decomposition problem is discussed. A fast and optimally localized biorthogonal wavelet transformation, suitable for image compression applications, is proposed. The performance of this transformation is compared to that of the discrete cosine transform in the context of still image compression.> Touradj Ebrahimi, Murat Kunt |
ICASSP | 1 |
| 1990 | Video coding using a pyramidal Gabor expansionabstractA compression technique based on an expansion is presented. The elementary functions of the expansion form a class of pyramidal Gabor functions covering the frequency domain in octave bands. The image sequence is coded by differentially coding selected coefficients of this expansion. Simulation results show sequences reconstructed with good quality for a bit rate less than 64 Kbit/s. Touradj Ebrahimi, Todd R. Reed, Murat Kunt |
VCIP | 1 |