VLDB 2026 Research / reviewers in the wild / expert
Mylène C. Q. Farias
dblp:96/2885 · also Mylène Christine Queiroz de Farias
· DBLP profile ↗
70ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0002-1957-9943ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 5 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 12 · 8 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RegionGate-CT: A Region-Gated Framework for Anatomy-Grounded CT Report Generation
Md Mustafizur Rahman, Sana Alamgeer, Mylène C. Q. Farias |
AIME (1) | 3 |
| 2026 | Convolutions Need Registers Too: HVS-Inspired Dynamic Attention for Video Quality Assessment
Mayesha Maliha Rahman Mithila, Mylène C. Q. Farias |
MMSys | 2 |
| 2026 | Class-Conditioned Gaussian Mixture Modeling for Imbalanced Time Series Quantification
Md. Shahriar Kabir, Mayesha Maliha R. Mithila, Anne H. H. Ngu, Mylène C. Q. Farias, Byron Gao |
PAKDD (3) | 4 |
| 2026 | DAL-PCQA: Enabling Distortion-Level and Language-Driven Reasoning for Point Cloud Quality Assessment
Swarna Chakraborty, Gabriel De Castro Araújo, Syeda Tasmi Faria, Marcelo M. Carvalho, Mylène C. Q. Farias |
QoMEX | 5 |
| 2026 | A Hybrid Subjective Quality Assessment Framework for Light Field Coding
Saeed Mahmoudpour, Mylène C. Q. Farias, Shengyang Zhao |
QoMEX | 2 |
| 2025 | MS-SCANet: A Multiscale Transformer-Based Architecture with Dual Attention for No-Reference Image Quality AssessmentabstractWe present the Multi-Scale Spatial Channel Attention Network (MS-SCANet), a transformer-based architecture designed for no-reference image quality assessment (IQA). MS-SCANet features a dual-branch structure that processes images at multiple scales, effectively capturing both fine and coarse details, an improvement over traditional single-scale methods. By integrating tailored spatial and channel attention mechanisms, our model emphasizes essential features while minimizing computational complexity. A key component of MS-SCANet is its cross-branch attention mechanism, which enhances the integration of features across different scales, addressing limitations in previous approaches. We also introduce two new consistency loss functions, Cross-Branch Consistency Loss and Adaptive Pooling Consistency Loss, which maintain spatial integrity during feature scaling, outforming conventional linear and bilinear techniques. Extensive evaluations on datasets like KonIQ-10k, LIVE, LIVE Challenge, and CSIQ show that MS-SCANet consistently surpasses state-of-the-art methods, offering a robust framework with stronger correlations with subjective human scores. Mayesha Maliha R. Mithila, Mylène C. Q. Farias |
ICASSP | 2 |
| 2025 | NASSBLiF: No-Reference Light Field Image Quality Assessment Via Neighborhood Attention and Scale SwinabstractImmersive media, such as Light Field images, enhance user experience by offering a high-fidelity spatial-angular representation of virtual environments. Challenges in transmitting and compressing content can cause noticeable artifacts, reducing quality. This highlights the need for efficient image quality assessment methods tailored to LF, given its large data requirements. Therefore, we introduce an efficient novel no-reference LF image quality assessment model utilizing neighborhood attention to extract spatial features and a Scale Swin Transformer to extract angular features. Due to their capability for extracting long-range dependencies and their linear computational cost relative to image size, both attention mechanisms prove highly effective in feature extraction for LF images. The proposed metric was evaluated on three datasets and outperformed state-of-the-art metrics. Myllena Prado, Mylène C. Q. Farias |
ICIP | 2 |
| 2025 | MT-DPCQA: A Multimodal Time-aware Learning Approach for No-Reference Dynamic Point Cloud Quality AssessmentabstractAs the use of dynamic point clouds (DPCs) expands in immersive media settings including augmented and virtual reality, it has become more important than ever to have precise and scalable methods for quality evaluation. However, most existing objective Point Cloud Quality Assessment (PCQA) methods focus on static content and fail to capture the temporal dynamics and multimodal perceptual cues inherent in dynamic scenarios. In this work, we propose a no-reference dynamic PCQA framework that integrates both geometric and visual modalities with global temporal modeling for perceptually aligned quality prediction. For the 3D modality, we extract localized spatio-temporal features using a time-aware point cloud encoder that incorporates the normalized frame index as an additional input channel. In parallel, we generate two complementary projections per frame and extract visual features using a pre-trained convolutional network. A dynamic gating network adaptively weights the contributions of the two modalities at each time step. These weighted features are fused and passed to a temporal transformer, which captures long-range temporal dependencies to regress the final quality score. Comprehensive tests on benchmark datasets reveal that our approach surpasses existing full-reference and no-reference PCQA techniques, demonstrating its efficacy in assessing the quality of dynamic point clouds. Swarna Chakraborty, Mylène C. Q. Farias |
ACM Multimedia | 2 |
| 2025 | ViGBLiF: A Graph-Based Approach to No-Reference Light Field Image Quality AssessmentabstractImmersive media aims to provide a compelling sense of being physically present in a virtual environment. Light Field images are a unique form of immersive media that enhance user interaction, offering an advanced spatial-angular depiction of the surroundings. However, the transmission and compression of the Light Field content can lead to noticeable artifacts that diminish the quality. Therefore, it is essential to have quality assessment methods tailored specifically for the Light Field content. We introduce ViGBLiF, an innovative model for assessing the quality of light-field images without reference, using the Epipolar image light-field format. ViGBLiF employs Vision Transformers to obtain pixel-level spatial-angular attributes from Epipolar image patches, using Graph Convolutional Layers to enable the depiction of structural dependencies. This graph-based approach improves the prediction of image quality. Evaluations conducted on three benchmark datasets reveal that the suggested ViGBLiF performs on par with, or even surpasses, leading metrics. Myllena Prado, Mylène C. Q. Farias |
QoMEX | 2 |
| 2025 | SSRT: Intra- and cross-view attention for stereo image super-resolution
Qixue Yang, Yi Zhang 0033, Damon M. Chandler, Mylène C. Q. Farias |
Multim. Tools Appl. | 4 |
| 2025 | SynFlowMap: A synchronized optical flow remapping for video motion magnification
Jonathan A. Lima, Cristiano Jacques Miosso, Mylène C. Q. Farias |
Signal Process. Image Commun. | 3 |
| 2025 | A 360-degree Video Player for Dynamic Video Editing Applicationsabstract360-degree videos introduce unique challenges to storytelling, especially if filmmakers want to draw the user’s attention to specific regions of the surrounding sphere to ensure the proper flow of narrative at any given time. For instance, at scene cuts, a user might lose track of a story if he/she looks toward a direction that prevents an intended Region of Interest (RoI) to appear in the viewport of the user’s Head-mounted Display (HMD). Because of that, new editing techniques for 360-degree videos are being investigated, and the use of dynamic (i.e., real time) alignment of intended RoI across scene cuts is receiving growing attention. To support research on this topic, we introduce 360EAVP, an open source web-based application for streaming and visualization of 360-degree edited videos on HMDs. The 360EAVP is built on top of the VR DASH Tile Player, which is based on cube mapping projection, and adds new capabilities to handle dynamic video editing. Specifically, 360EAVP introduces (1) real-time track of a user’s viewport direction on the HMD; (2) support for dynamic editing via “snap-change” or “fade-rotation” effects; (3) visibility mapping of the user’s Field of View (FoV) with respect to the player’s cube mapping projection (for purposes of video tile requests during streaming); (4) incorporation of editing timing information into the operation of an ABR algorithm; (5) viewport prediction module based on linear regression or ridge regression models; and (6) video playback data collection and log module. To showcase the use of 360EAVP, we present the results of a small-scale subjective experiment that aimed to evaluate the impact of the “snap-change” and “fade-rotation” editing techniques on user’s Quality of Experience (QoE), comfort, and head movement. Our findings indicate that these editing techniques do not compromise the overall QoE; in fact, in certain scenarios, an improvement is observed. In addition, some subjects significantly reduced their head movements with the editing techniques implemented. Gabriel De Castro Araújo, Henrique Domingues Garcia, Mylène C. Q. Farias, Ravi Prakash 0001, Marcelo M. Carvalho |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Perception-Driven Point Cloud Quality Assessment Through Projections and Deep Structure SimilarityabstractPoint Clouds (PCs) have gained considerable interest as a potential format for representing tridimensional (3D) data in various applications, including augmented reality and autonomous vehicles. In the realm of multimedia, where applications and technologies such as 3DTV revolve around human interaction, it is crucial to assess how humans perceive the visual quality of PCs. To address this need, research on Point Cloud Quality Assessment (PCQA) metrics, incorporating aspects of the human visual perception, has gained prominence. This research aims to facilitate PC quality enhancement, optimize PC registration pipelines, and improve the performance of PC codecs. However, assessing the quality of PCs poses inherent challenges due to its irregularity and sparsity, necessitating the identification of mathematical relationships within PCs for accurate quality predictions. In this work, we adopt a projection-based approach, representing a PC as a structured set of bidimensional (2D) projections-regular images derived from the 3D structure of the PC. We leverage the power of Neural Networks (NNs) and Machine Learning (ML) to create a perceptual-driven, full-reference metric for PCQA. We assess these images using established Image Quality Assessment (IQA) methods, widely acknowledged in the state-of-the-art (SOTA) for traditional 2D (non-immersive) visual content. The generated scores are then deposited into a vector, serving as input for a regressor model to predict the final quality score. Our results demonstrate the competitiveness of our model compared to SOTA metrics. Arthur H. S. Carvalho, Pedro Garcia Freitas, Mateus Gonçalves, Johann Homonnai, Mylène C. Q. Farias |
MMSP | 5 |
| 2024 | A Subjective Test Framework for JPEG Pleno Quality AssessmentabstractThe Joint Photographic Experts Group (JPEG) is currently addressing the challenges in assessing plenoptic image quality by developing new standards for both subjective and objective assessment. This process entails revisiting existing recommendations and establishing a visual quality assessment (QA) framework that considers the unique aspects of plenoptic data. The focus is currently on the light field modality, with efforts to gather expert contributions and develop tools that cater both subjective and objective QA, ensuring that the evolving requirements of plenoptic imaging QA are met. This paper presents the JPEG Pleno subjective test tool, a pivotal element in JPEG’s standardization endeavor, designed to facilitate a range of subjective QA experiments that steer the decisions during a standardization process. Shengyang Zhao, Saeed Mahmoudpour, Mylène C. Q. Farias, Carla L. Pagliari, Peter Schelkens |
QoMEX | 3 |
| 2024 | 360Align: An Open Dataset and Software for Investigating QoE and Head Motion in 360° Videos with Alignment EditsabstractThis paper presents the resources utilized for and gathered from a subjective QoE experiment conducted with 45 participants who watched 360° videos processed with offline alignment edits. The aim of these edits is to redirect the assumed field of view to a specified region of interest by rotating the 360° video around the horizon. The user experiment involved alignment edits employing both gradual and instant rotation. We employed the Double Stimulus method, whereby participants evaluated each original-processed video pair, resulting in a dataset with 5400 comfort and sense of presence ratings. During video consumption, head motion was recorded from the Meta Quest 2 device. The resulting dataset, containing the original and processed videos, is made publicly accessible. The accompanying web application, developed for the execution of the experiment, is released alongside scripts for the evaluation of head rotation data in a public repository. A cross-analysis of QoE and HM behavior provides insight into the efficacy of alignment edits for attention alignment within the same scene. This comprehensive set of experimental resources establishes a foundation for further research of 360° videos processed with alignment edits. Lucas S. Althoff, Alessandro Rodrigues e Silva, Marcelo M. Carvalho, Mylène C. Q. Farias |
IMX | 4 |
| 2023 | Using a Diverse Neural Network to Predict the Quality of Light Field ImagesabstractLight Fields (LF) capture both angular and spatial information of light rays traveling through free space, resulting in a richer representation of a scene from multiple viewpoints. However, the high dimensionality of LF data poses a challenge for compression and transmission algorithms, which can degrade visual quality. To address this issue, we propose a novel no-reference LF Image Quality Assessment (LF-IQA) method that accurately predicts the quality of complex LF content. Our method is based on a diverse neural network architecture that includes Convolutional Neural Network (CNN) blocks, Atrous Convolutional Layers (ACLs), and Long Short-Term Memory (LSTM) layers. The model architecture consists of two streams, each containing CNN, ACL, and LSTM layers, which take the horizontal and vertical epipolar plane LF images as input. CNN blocks extract basic input features, while ACL blocks extract high-level features. The LSTM layers are used to capture long-term dependencies and relationships among distortion-related features. Finally, the outputs of both streams are concatenated and fed into the regression block for quality prediction. Our results demonstrate that the proposed LF-IQA method is robust and outperforms current state-of-the-art methods, even for complex LF content. Sana Alamgeer, André H. M. Costa, Mylène C. Q. Farias |
MMSP | 3 |
| 2023 | A two-stream cnn based visual quality assessment method for light field images
Sana Alamgeer, Mylène C. Q. Farias |
Multim. Tools Appl. | 2 |
| 2023 | A survey on visual quality assessment methods for light fields
Sana Alamgeer, Mylène C. Q. Farias |
Signal Process. Image Commun. | 2 |
| 2023 | Point cloud quality assessment: unifying projection, geometry, and texture similarity
Pedro Garcia Freitas, Rafael Diniz, Mylène C. Q. Farias |
Vis. Comput. | 3 |
| 2022 | Light Field Image Quality Assessment with Dense Atrous ConvolutionsabstractUnlike regular images that represent only light intensities, Light Field (LF) contents carry information about the intensity of light in a scene, including the direction in which light rays are traveling in space. This allows for a richer representation of our world, but requires large amounts of data that need to be processed and compressed before being transmitted to the viewer. Since these techniques may introduce distortions, the design of Light Field Image Quality Assessment (LF-IQA) methods is important. Currently, most LF-IQA methods use traditional 2D image quality assessment techniques or rely on low-level spatial features. In this paper, we propose a novel no-reference LF-IQA method that takes into account both LF angular and spatial information. The proposed method is made up of two processing streams with identical blocks of Convolutional Neural Network (CNN), Atrous Convolution layers (ACL), and a regression block for quality prediction. The results show that the method is robust and outperforms current state-of-the-art methods. Sana Alamgeer, Mylène C. Q. Farias |
ICIP | 2 |
| 2022 | PIES-ME '22: 1st Workshop on Photorealistic Image and Environment Synthesis for Multimedia ExperimentsabstractPhotorealistic media aim to faithfully represent the world, creating an experience that is perceptually indistinguishable from a real world experience. In the past few years, this area has grown significantly, with new multimedia areas emerging, such as light fields, point clouds, ultra-high definition, high frame rate, high dynamic range imaging, and novel 3D audio and sound field technologies. In spite of all the advances done so far, there are several technological challenges to overcome. In particular, research in this field typically requires the use of big datasets, software tools, and powerful infrastructures. Among these, the availability of meaningful datasets, with a diverse and high-quality content, is of significant importance and, to date, most available datasets are limited and do not provide researchers adequate tools to advance the area of photorealistic applications. To help advancing research efforts in this area, this workshop aims to engage experts and researchers on the synthesis of photorealistic images and/or virtual environments, particularly in the form of public datasets, software tools, or infrastructures, for not only multimedia systems research, but also for other fields, such as machine learning, robotics, computer vision, mixed reality, and virtual reality. Ravi Prakash 0001, Mylène C. Q. Farias, Marcelo M. Carvalho, Ryan P. McMahan |
ACM Multimedia | 2 |
| 2022 | Light field image quality assessment method based on deep graph convolutional neural network: research proposalabstractThis paper contains the research proposal of Sana Alamgeer that was presented at the MMSys 2022 doctoral symposium. Unlike regular images that represent only light intensities, Light Field (LF) contents carry information about the intensity of light in a scene, including the direction light rays are traveling in space. This allows for a richer representation of our world, but requires large amounts of data that need to be processed and compressed before being transmitted to the viewer. Since these techniques may introduce distortions, the design of Light Field Image Quality Assessment (LF-IQA) methods is important. The majority of LF-IQA methods based on traditional Convolutional Neural Network (CNN) have limitations, i.e. they are unable to increase the receptive field of a neuron-pixel to model non-local image features. In this work, we propose a novel no-reference LF-IQA method that is based on Deep Graph Convolutional Neural Network (GCNN). Our method not only takes into account both LF angular and spatial information, but also learns the order of pixel information. Specifically, the method is composed of one input layer that takes a pair of graphs and their corresponding subjective quality scores as labels, 4 GCNN layers, fully connected layers, and a regression block for quality prediction. Our aim is to develop the quality prediction method with maximum accuracy for distorted LF content. Sana Alamgeer, Muhammad Irshad, Mylène C. Q. Farias |
MMSys | 3 |
| 2022 | Deep Learning-Based Light Field Image Quality Assessment Using Frequency Domain InputsabstractLight Field (LF) cameras capture angular and spa-tial information and, consequently, require a large amount of resources in memory and bandwidth. To reduce these requirements, LF contents generally need to undergo compression and transmission protocols. Since these techniques may introduce distortions, the design of Light-Field Image Quality Assessment (LFI - IQA) methods are important to monitor the quality of the LF image (LFI) content at the user side. The majority of the existing LFI-IQA methods work in the spatial domain, where it is more difficult to analyze changes in the spatial and angular domains. In this work, we present a novel NR LFI-IQA, which is based on a Deep Neural Network that uses Frequency domain inputs (DNNF-LFIQA). The proposed method predicts the quality of an LF image by taking as input the Fourier magnitude spectrum of LF contents, represented as horizontal and ver-tical Epipolar Plane Images (EPI)s. Specifically, DNNF-LFIQA is composed of two processing streams (streaml and stream2) that take as inputs the horizontal and the vertical epipolar plane images in the frequency domain. Both streams are composed of identical blocks of convolutional neural networks (CNNs), with their outputs being combined using two fusion blocks. Finally, the fused feature vector is fed to a regression block to generate the quality prediction. Results show that the proposed method is fast, robust, and accurate. Sana Alamgeer, Mylène C. Q. Farias |
QoMEX | 2 |
| 2022 | On the Performance of Temporal Pooling Methods for Quality Assessment of Dynamic Point CloudsabstractPoint Clouds (PCs) are collections of points distributed in the 3D space, containing attributes such as color, normals, transparency, and specularity. Dynamic Point Clouds (DPCs) correspond to sequences of points in the 3D space that vary over time like pixels vary over time in a conventional video. Dynamic PCs are a suitable way to represent volumetric videos that can be used in augmented or virtual reality applications. This representation, however, requires a large number of points to achieve a high quality of experience and needs to be compressed before storage and transmission. Therefore, reliable quality metrics are needed in order to automatically estimate the perceptual quality of dynamic PC contents. Since currently there are several quality assessment metrics for static PC, a possible approach solution consists of using temporal pooling functions to combine the quality scores predicted for each of the frames. In this paper, we study the effects of different temporal pooling strategies on the performance of dynamic PC quality assessment methods. Our experimental tests were performed using a recent publicly-available database, demonstrating the efficiency of the evaluated temporal pooling models. More specifically, the work provides a recipe on how to apply a temporal pooling function to combine frame-based quality predictions generated with texture-based static PC quality assessment methods to estimate the quality of dynamic PCs. Pedro Garcia Freitas, Mateus Gonçalves, Johann Homonnai, Rafael Diniz, Mylène C. Q. Farias |
QoMEX | 5 |
| 2022 | See hear now: is audio-visual QoE now just a fusion of audio and video metrics?abstractSingle-modal audio/speech and video quality models have reached high levels of performance. Although traditional algorithms are still preferred for many practical applications, advances in machine learning (ML) and deep learning techniques have exceeded their performance in several scientific comparisons. However, audio-visual (AV) models have received signifi-cantly less attention and development. Despite the acknowledged challenge that multimodal interaction poses to the AV problem, traditional AV models generally rely on simple fusion techniques of individual audio and video predictions. Consequently, the impact of recent advances in single-modal quality assessment models on SOTA (state-of-the-art) AV quality models merits attention. This paper presents a revised and updated benchmark for AV quality assessment with particular focus on new speech quality metrics. Three AV datasets were used to test audio, video, and AV quality metrics. For audio and video, the best performing metrics were selected to build simple late-fusion models using their raw predictions. The fused models were then compared to the SOTA AV models. Results show that a simple fusion strategy produces accurate AV quality predictions (LCC and SCC greater than 0.90) with low error rates (RMSE lower than 0.33). These results highlight the influence of advances in speech quality for AV quality assessment. Helard Becerra Martinez, Andrew Hines, Mylène C. Q. Farias |
QoMEX | 3 |
| 2022 | A Spendable Cold Wallet from QR VideoabstractHot/cold wallet refers to a widely used paradigm to enhance the security level of cryptocurrency applications that was proposed on Bitcoin Improvement Proposal 32. In a nutshell, after performing an initial setup in which the hot wallet receives partial information of the cold wallet in order to hierarchically generate (transaction receiving) addresses, the cold wallet stays offline, whereas the hot wallet is kept online. The initial transferred information enables the hot wallet to generate receiving addresses for both wallets, but it can only spend its own funds, i.e., it cannot spend the funds in the cold wallet. This design conveniently mimics money storage in daily life: pocket money is kept in a less safe location, e.g., a regular wallet, while life savings are kept in a more safe environment, e.g., banking account. Note that the funds that land in offline addresses cannot be spent if the cold wallet is kept permanently offline. We propose a protocol and a technical solution to spend funds from a cold wallet without physically connecting it to any network. We designed and implemented a prototype for a system based on Optical Camera Communication (OCC) in a screen to camera setting, which can receive messages from a computer screen at the rate of over 150kB per second. Our system consists of a sequence of QR codes – a QR video. Our solution minimizes the possible attack vectors, including malware, by relying on optical communication yet providing a larger bandwidth than regular QR code based solutions. Rafael Dowsley, Mylène C. Q. Farias, Mario Larangeira, Anderson Nascimento, Jot Virdee |
SECRYPT | 2 |
| 2022 | Point cloud quality assessment based on geometry-aware texture descriptors
Rafael Diniz, Pedro Garcia Freitas, Mylène C. Q. Farias |
Comput. Graph. | 3 |
| 2021 | Color and Geometry Texture Descriptors for Point-Cloud Quality AssessmentabstractPoint Clouds (PCs) have recently been adopted as the preferred data structure for representing 3D visual contents. Examples of Point Cloud (PC) applications range from 3D representations of small objects up to large scenes, both still or dynamic in time. PC adoption triggered the development of new coding, transmission, and display methodologies that culminated in new international standards for PC compression. Along with these, in the last couple of years, novel methods have been developed for evaluating the visual quality of PC contents. This paper presents a new objective full-reference visual quality assessment metric for static PC contents, named BitDance, which uses color and geometry texture descriptors. The proposed method first extracts the statistics of color and geometry information of the reference and test PCs. Then, it compares the color and geometry statistics and combines them to estimate the perceived quality of the test PC. Using publicly available PC quality assessment datasets, we show that the proposed PC quality assessment metric performs very well when compared to state-of-the-art quality metrics. In particular, the method performs well for different types of PC datasets, including the ones where both geometry and color are not degraded with similar intensities. BitDance is a low complexity algorithm, with an optimized C++ source code that is available for download at github.com/rafael2k/bitdance-pc_metric. Rafael Diniz, Pedro Garcia Freitas, Mylène C. Q. Farias |
IEEE Signal Process. Lett. | 3 |
| 2020 | Multi-Distance Point Cloud Quality AssessmentabstractThe popularity of smartphones, virtual reality headsets, and head-mounted devices is fomenting immersive applications that employ realistic representations of the real world. Among these, Point Cloud (PC) contents have recently gained prominence in academia and industry. However, there is some consensus that PC objective quality assessment methods are still an open problem. In this paper, we introduce a PC quality metric based on multiple distances between reference and test PCs. These distances are computed in both still points and texture spaces. Distances computed in still points space consider the direct differences between reference and test points. Distances in the texture space are measured after computing the Local Binary Pattern (LBP) descriptor of PCs. Since PCs are not equally geometrically distributed, we adapted the LBP descriptor to make the nearest points as the neighborhood pixels of the descriptor. The difference between the LBP statistics of reference and test PCs is used to assess the quality of the test PC. Experimental results show the proposed method performs well when compared with the state-of-the-art Point Cloud Quality Assessment (PCQA) methods. Rafael Diniz, Pedro Garcia Freitas, Mylène C. Q. Farias |
ICIP | 3 |
| 2020 | Local Luminance Patterns for Point Cloud Quality AssessmentabstractIn recent years, there has been an increase in the popularity of Point Clouds (PC) as the preferred data structure for representing 3D visual contents. Examples of PC applications range from 3D representations of small objects up to large maps. The advent of PC adoption triggered the development of new coding, transmission, and presentation methodologies. And, along with these, novel methods for evaluating the visual quality of PC contents. This paper presents a new objective full-reference visual quality metric for PC contents, which uses a proposed descriptor entitled Local Luminance Patterns (LLP). It extracts the statistics of the luminance information of reference and test PCs and compares their statistics to assess the perceived quality of the test PC. The proposed PC quality assessment method can be applied to both large and small scale PCs. Using publicly available PC quality datasets, we compared the proposed method with current state-of-the-art PC quality metrics, obtaining competing results. Rafael Diniz, Pedro Garcia Freitas, Mylène C. Q. Farias |
MMSP | 3 |
| 2020 | Hybrid Motion Magnification based on Same-Frame Optical Flow ComputationsabstractMotion magnification refers to the ability of amplifying small movements in a video in order to reveal important information about the observed scene. In the past, several motion magnification methods have been proposed, but most of them have the disadvantage of introducing annoying visual artifacts in the video. In this paper, we propose a method that analyses the optical flow between each original frame and the corresponding motion-magnified frame and, then, synthesizes a new motion-magnified video by remapping the original video using the generated optical flow map. The proposed approach is able to eliminate the artifacts that appear in Eulerian methods. Also, it is able to amplify the motion by an extra factor of 2 and to invert the motion direction. Jonathan A. Lima, Cristiano Jacques Miosso, Mylène C. Q. Farias |
MMSP | 3 |
| 2020 | Towards a Point Cloud Quality Assessment Model using Local Binary PatternsabstractThe proliferation of devices such as mobile phones, virtual reality headsets, and head-mounted displays has increased the popularity of immersive applications that deliver realistic representations of the real world. Among the technologies that enable such applications, the point cloud (PC) technology seems to be one of the most mature alternatives, gaining prominence in academia, industry, and standardization committees. Although PC technologies have been used in entertainment, automotive, and geographical location industries, the design of objective quality assessment methods for PC contents is still an open problem. In this paper, we introduce a texture-based objective quality assessment method for PC contents. The method analyzes the texture of the PC content using the Local Binary Pattern (LBP) descriptor. Unlike points in still (2D) images, the points in a PC are not equally distributed in space. Therefore, we adapted the LBP descriptor to allow processing a PC point and its neighboring points. The statistics of the LBP outputs, for both reference and test PCs, are computed and compared to obtain a quality estimate for the test (impaired) PC content. Experimental results show that the proposed PC quality metric has a good correlation with subjective quality scores, outperforming state-of-the-art PC quality metrics. Rafael Diniz, Pedro Garcia Freitas, Mylène C. Q. Farias |
QoMEX | 3 |
| 2020 | How Deep is Your Encoder: An Analysis of Features Descriptors for an Autoencoder-Based Audio-Visual Quality MetricabstractThe development of audio-visual quality assessment models poses a number of challenges in order to obtain accurate predictions. One of these challenges is the modelling of the complex interaction that audio and visual stimuli have and how this interaction is interpreted by human users. The No-Reference Audio-Visual Quality Metric Based on a Deep Autoencoder (NAViDAd) deals with this problem from a machine learning perspective. The metric receives two sets of audio and video features descriptors and produces a low-dimensional set of features used to predict the audio-visual quality. A basic implementation of NAViDAd was able to produce accurate predictions tested with a range of different audio-visual databases. The current work performs an ablation study on the base architecture of the metric. Several modules are removed or re-trained using different configurations to have a better understanding of the metric functionality. The results presented in this study provided important feedback that allows us to understand the real capacity of the metric's architecture and eventually develop a much better audio-visual quality metric. Helard Becerra Martinez, Andrew Hines, Mylène C. Q. Farias |
QoMEX | 3 |
| 2020 | Perceptual quality assessment of 3D videos with stereoscopic degradations
Alessandro Rodrigues e Silva, Mylène C. Q. Farias |
Multim. Tools Appl. | 2 |
| 2020 | Isotropic and anisotropic filtering norm-minimization: A generalization of the TV and TGV minimizations using NESTA
Jonathan A. Lima, Felipe Batista da Silva, Ricardo von Borries, Cristiano Jacques Miosso, Mylène C. Q. Farias |
Signal Process. Image Commun. | 5 |
| 2020 | Comparison between Digital Tone-Mapping Operators and a Focal-Plane Pixel-Parallel Circuit
Gustavo M. S. Nunes, Fernanda D. V. R. Oliveira, Mylène C. Q. Farias, José Gabriel R. C. Gomes, Antonio Petraglia, Jorge Fernández-Berni, Ricardo Carmona-Galán, Ángel Rodríguez-Vázquez |
Signal Process. Image Commun. | 3 |
| 2020 | Image quality assessment using BSIF, CLBP, LCP, and LPQ operators
Pedro Garcia Freitas, Luísa Peixoto da Eira, Samuel Soares Santos, Mylène C. Q. Farias |
Theor. Comput. Sci. | 4 |
| 2020 | Study of Subjective and Objective Quality Assessment of Audio-Visual SignalsabstractThe topics of visual and audio quality assessment (QA) have been widely researched for decades, yet nearly all of this prior work has focused only on single-mode visual or audio signals. However, visual signals rarely are presented without accompanying audio, including heavy-bandwidth video streaming applications. Moreover, the distortions that may separately (or conjointly) afflict the visual and audio signals collectively shape user-perceived quality of experience (QoE). This motivated us to conduct a subjective study of audio and video (A/V) quality, which we then used to compare and develop A/V quality measurement models and algorithms. The new LIVE-SJTU Audio and Video Quality Assessment (A/V-QA) Database includes 336 A/V sequences that were generated from 14 original source contents by applying 24 different A/V distortion combinations on them. We then conducted a subjective A/V quality perception study on the database towards attaining a better understanding of how humans perceive the overall combined quality of A/V signals. We also designed four different families of objective A/V quality prediction models, using a multimodal fusion strategy. The different types of A/V quality models differ in both the unimodal audio and video quality prediction models comprising the direct signal measurements and in the way that the two perceptual signal modes are combined. The objective models are built using both existing state-of-the-art audio and video quality prediction models and some new prediction models, as well as quality-predictive features delivered by a deep neural network. The methods of fusing audio and video quality predictions that are considered include simple product combinations as well as learned mappings. Using the new subjective A/V database as a tool, we validated and tested all of the objective A/V quality prediction models. We will make the database publicly available to facilitate further research. Xiongkuo Min, Guangtao Zhai, Jiantao Zhou 0001, Mylène C. Q. Farias, Alan C. Bovik |
IEEE Trans. Image Process. | 4 |
| 2019 | A No-Reference Autoencoder Video Quality MetricabstractIn this work, we introduce the No-reference Autoencoder VidEo (NAVE) quality metric, which is based on a deep au-toencoder machine learning technique. The metric uses a set of spatial and temporal features to estimate the overall visual quality, taking advantage of the autoencoder ability to produce a better and more compact set of features. NAVE was tested on two databases: the UnB-AVQ database and the LiveNetflix-II database. Results show that the method is able to estimate the perceived video quality with a good correlation performance and a small error, when compared to currently available no-reference and full-reference video quality objective metrics. Helard Becerra Martinez, Mylène C. Q. Farias, Andrew Hines |
ICIP | 2 |
| 2019 | A taxonomy and dataset for 360° videosabstractIn this paper, we propose a taxonomy for 360° videos that categorizes videos based on moving objects and camera motion. We gathered and produced 28 videos based on the taxonomy, and recorded viewport traces from 60 participants watching the videos. In addition to the viewport traces, we provide the viewers' feedback on their experience watching the videos, and we also analyze viewport patterns on each category. Afshin TaghaviNasrabadi, Aliehsan Samiei, Anahita Mahzari, Ryan P. McMahan, Ravi Prakash 0001, Mylène C. Q. Farias, Marcelo M. Carvalho |
MMSys | 6 |
| 2019 | A (2, 2) XOR-based visual cryptography scheme without pixel expansion
Max E. Vizcarra Melgar, Mylène C. Q. Farias |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | High density two-dimensional color code
Max E. Vizcarra Melgar, Mylène C. Q. Farias |
Multim. Tools Appl. | 2 |
| 2019 | A framework for computationally efficient video quality assessment
Welington Y. L. Akamine, Pedro Garcia Freitas, Mylène C. Q. Farias |
Signal Process. Image Commun. | 3 |
| 2019 | Bio-inspired optimization algorithms for real underwater image restoration
Camilo Sánchez-Ferreira, Leandro dos Santos Coelho, Helon V. H. Ayala, Mylène C. Q. Farias, Carlos H. Llanos |
Signal Process. Image Commun. | 4 |
| 2018 | Blind image quality assessment based on multiscale salient local binary patternsabstractDue to the rapid development of multimedia technologies, over the last decades image quality assessment (IQA) has become an important topic. As a consequence, a great research effort has been made to develop computational models that estimate image quality. Among the possible IQA approaches, blind IQA (BIQA) is of fundamental interest as it can be used in most multimedia applications. BIQA techniques measure the perceptual quality of an image without using the reference (or pristine) image. This paper proposes a new BIQA method that uses a combination of texture features and saliency maps of an image. Texture features are extracted from the images using the local binary pattern (LBP) operator at multiple scales. To extract the salient of an image, i.e. the areas of the image that are the main attractors of the viewers' attention, we use computational visual attention models that output saliency maps. These saliency maps can be used as weighting functions for the LBP maps at multiple scales. We propose an operator that produces a combination of multiscale LBP maps and saliency maps, which is called the multiscale salient local binary pattern (MSLBP) operator. To define which is the best model to be used in the proposed operator, we investigate the performance of several saliency models. Experimental results demonstrate that the proposed method is able to estimate the quality of impaired images with a wide variety of distortions. The proposed metric has a better prediction accuracy than state-of-the-art IQA methods. Pedro Garcia Freitas, Sana Alamgeer, Welington Y. L. Akamine, Mylène C. Q. Farias |
MMSys | 4 |
| 2018 | Combining audio and video metrics to assess audio-visual quality
Helard Becerra Martinez, Mylène C. Q. Farias |
Multim. Tools Appl. | 2 |
| 2018 | Using multiple spatio-temporal features to estimate video quality
Pedro Garcia Freitas, Welington Y. L. Akamine, Mylène C. Q. Farias |
Signal Process. Image Commun. | 3 |
| 2018 | No-Reference Image Quality Assessment Using Orthogonal Color Planes PatternsabstractThis paper proposes a new general-purpose no-reference image quality assessment (NR-IQA) method based on color texture analysis. Specifically, the proposed method uses the statistics of the orthogonal color planes pattern (OCPP) descriptor to characterize image quality. The OCPP descriptor, proposed in this paper, is an extension of the local binary pattern operator that incorporates color information. To make NR-IQA methods more generic, that is, more sensitivity to different types of degradation (e.g., color and contrast degradation), it is important to take into consideration the color information. In the proposed NR-IQA method, we use the statistics of the OCPP descriptor as an input vector to a regression algorithm, which models the nonlinear relationship between the OCPP statistics and the subjective opinion scores. Experimental results show that proposed IQA method is quite efficient when compared to popular state-of-the-art NR-IQA methods. Pedro Garcia Freitas, Welington Y. L. Akamine, Mylène C. Q. Farias |
IEEE Trans. Multim. | 3 |
| 2017 | Per-pixel mirror-based method for high-speed video acquisition
Jonathan A. Lima, Cristiano Jacques Miosso, Mylène C. Q. Farias |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | No-reference image quality assessment based on statistics of Local Ternary PatternabstractIn this paper, we propose a new no-reference image quality assessment (NR-IQA) method that uses a machine learning technique based on Local Ternary Pattern (LTP) descriptors. LTP descriptors are a generalization of Local Binary Pattern (LBP) texture descriptors that provide a significant performance improvement when compared to LBP. More specifically, LTP is less susceptible to noise in uniform regions, but no longer rigidly invariant to gray-level transformation. Due to its insensitivity to noise, LTP descriptors are not able to detect milder image degradation. To tackle this issue, we propose a strategy that uses multiple LTP channels to extract texture information. The prediction algorithm uses the histograms of these LTP channels as features for the training procedure. The proposed method is able to blindly predict image quality, i.e., the method is no-reference (NR). Results show that the proposed method is considerably faster than other state-of-the-art no-reference methods, while maintaining a competitive image quality prediction accuracy. Pedro Garcia Freitas, Welington Y. L. Akamine, Mylène C. Q. Farias |
QoMEX | 3 |
| 2016 | Annoyance models for videos with spatio-temporal artifactsabstractAlthough compression and transmission artifacts are likely to appear simultaneously in digital videos, their annoyance has been traditionally studied and modeled in isolation. So, while blockiness, blurriness, and packet-loss metrics exist, hardly any attempt has been made at modeling their joint impact on visual perception. In this paper, we evaluate the perceptual impact of those three artifacts on video quality. Based on data from three different experiments, in which a pool of participants evaluated videos impaired with packet-loss, blurriness and blockiness (in isolation and in combination), we analyze how the different artifacts combine to produce annoyance and propose several models for predicting the annoyance of videos impaired with combinations of packet-loss, blurriness and blockiness. Alexandre F. Silva, Mylène C. Q. Farias, Judith Redi |
QoMEX | 2 |
| 2016 | Image restoration for through-the-earth communicationsabstractThis article describes image transmission in Through-The-Earth Communications (TTE), which are characterized by non-Gaussian impulsive atmospheric noise. We analyze the transmission of compressed and uncompressed image files. We verify the possibility of transmitting uncompressed analog images, which can be restored even with signal noise ratio (SNR) equal to −10 dB, and, based on the Structural Similarity (SSIM) metric, we verify the existence of an SNR threshold, at which we should switch between compressed and uncompressed images. Savio Oliveira de Almeida Neves, Lucas Sousa e Silva, Mylène C. Q. Farias, André Noll Barreto |
WCNC | 3 |
| 2016 | Detecting tampering in audio-visual content using QIM watermarking
Ronaldo Rigoni, Pedro Garcia Freitas, Mylène C. Q. Farias |
Inf. Sci. | 3 |
| 2016 | Recent developments in visual quality monitoring by key performance indicatorsabstractIn addition to traditional Quality of Service (QoS), Quality of Experience (QoE) poses a real challenge for Internet service providers, audio-visual services, broadcasters and new Over-The-Top (OTT) services. Therefore, objective audio-visual metrics are frequently being dedicated in order to monitor, troubleshoot, investigate and set benchmarks of content applications working in real-time or off-line. The concept proposed here, Monitoring of Audio Visual Quality by Key Performance Indicators (MOAVI), is able to isolate and focus investigation, set-up algorithms, increase the monitoring period and guarantee better prediction of perceptual quality. MOAVI artefacts Key Performance Indicators (KPI) are classified into four categories, based on their origin: capturing, processing, transmission, and display. In the paper, we present experiments carried out over several steps with four experimental set-ups for concept verification. The methodology takes into the account annoyance visibility threshold. The experimental methodology is adapted from International Telecommunication Union – Telecommunication Standardization Sector (ITU-T) Recommendations: P.800, P.910 and P.930. We also present the results of KPI verification tests. Finally, we also describe the first implementation of MOAVI KPI in a commercial product: the NET-MOZAIC probe. Net Research, LLC, currently offers the probe as a part of NET-xTVMS Internet Protocol Television (IPTV) and Cable Television (CATV) monitoring system. Mikolaj Leszczuk, Mateusz Hanusiak, Mylène C. Q. Farias, Emmanuel Wyckens, George Heston |
Multim. Tools Appl. | 3 |
| 2016 | Hiding color watermarks in halftone images using maximum-similarity binary patterns
Pedro Garcia Freitas, Mylène C. Q. Farias, Aletéia P. F. Araújo |
Signal Process. Image Commun. | 2 |
| 2016 | Enhancing inverse halftoning via coupled dictionary training
Pedro Garcia Freitas, Mylène C. Q. Farias, Aletéia P. F. Araújo |
Signal Process. Image Commun. | 2 |
| 2016 | Perceptual Annoyance Models for Videos With Combinations of Spatial and Temporal ArtifactsabstractUnderstanding the perceptual impact of compression artifacts in video is one of the keys for designing better coding schemes and appropriate visual quality control chains. Although compression and transmission artifacts, such as blockiness, blurriness, and packet-loss, appear simultaneously in digital videos, traditionally they have been studied in isolation. In this paper, we report the results of three subjective quality assessment experiments aimed at studying perceptual characteristics of a set of artifacts common in digital videos. With this goal, first, we study the annoyance of each of three artifacts (blockiness, blurriness, and packet-loss) in isolation and then in combination. Based on the subjective evaluations, we design several models of the annoyance caused by the joint presence of these three artifacts on digital video. Alexandre F. Silva, Mylène C. Q. Farias, Judith Redi |
IEEE Trans. Multim. | 2 |
| 2015 | Multi-objective differential evolution algorithm for underwater image restorationabstractUnderwater image processing area has been considered an important topic within the last decades with important achievements. This kind of images are essentially characterized by their poor visibility because light is exponentially attenuated as it travels in the water and the scenes result poorly contrasted and hazy. On the other hand, image restoration takes into account the influence of the environment on the image in order to achieve an image with an improved quality. This technique consist of inverting the physical model of image formation. That model contains parameters which represent variables such as coefficients of absorption, scattering, among others. In this case, the quality of the restored image depends on the correct estimation of these parameters. In this work, an approach based on evolutionary optimization algorithms is proposed, for restoring underwater images by estimating the model parameters, and using two metrics for quality assessment. The degradation in the images has been simulated by using an image formation model. Results show that image restoration based on a Multi-Objective Differential Evolution (MODE) algorithm achieves images with good contrast and sharpness, being even better than the original image. Camilo Sánchez-Ferreira, Helon V. H. Ayala, Leandro dos Santos Coelho, Daniel M. Muñoz Arboleda, Mylène C. Q. Farias, Carlos H. Llanos |
CEC | 5 |
| 2015 | Improved performance of inverse halftoning algorithms via coupled dictionariesabstractInverse halftoning techniques are known to introduce visible distortions (typically, blurring or noise) into the reconstructed image. To reduce the severity of these distortions, we propose a novel training approach for inverse halftoning algorithms. The proposed technique uses a coupled dictionary (CD) to match distorted and original images via a sparse representation. This technique enforces similarities of sparse representations between distorted and non-distorted images. Results show that the proposed technique can improve the performance of different inverse halftone approaches. Images reconstructed with the proposed approach have a higher quality, showing less blur, noise, and chromatic aberrations. Pedro Garcia Freitas, Mylène C. Q. Farias, Aletéia P. F. Araújo |
ICME | 2 |
| 2013 | On the impact of packet-loss impairments on visual attention mechanismsabstractThis paper reports the results of a psychometric experiment aimed at investigating the visibility and annoyance of packet-loss artifacts, with a focus on understanding to what extent their presence influences viewing behavior. The rationale behind this study is that packet loss artifacts might “distract” visual attention, creating visual saliency themselves, thereby becoming more visible. In turn, higher artifact visibility might impact their annoyance. Our study involved seven videos, compressed at a very high bitrate (thus free of spatial artifacts), impaired by discarding packets at different packet loss ratios. We tracked the observers' eye-movements while (1) they assessed the annoyance of impaired videos and (2) looked freely at pristine videos. Our results show that the viewing behavior significantly changes from pristine to impaired videos. This change is related to both properties of the video and to the annoyance of the artifacts presented. Judith Redi, Ingrid Heynderickx, Bruno Macchiavello, Mylène C. Q. Farias |
ISCAS | 4 |
| 2005 | Detectability and Annoyance of Synthetic Blockiness, Blurriness, Noisiness, and Ringing in Video SequencesabstractSynthetic artifacts offer many advantages for experimental research on video quality because of the degree of control that the researchers have with respect to the amplitude, distribution, and mixture of artifacts. We have developed algorithms for synthetically generating four types of artifacts commonly found in digital videos: blockiness, blurriness, noisiness, and ringing. In this paper, we inserted these artifacts in short video sequences and performed a psychophysical experiment where we measured the probability of detection and the mean annoyance values of these artifacts as a function of their total squared error (TSE). The results show that, although the different artifacts looked different and affected the videos differently, there is no consistent difference between either their visibility thresholds or mid-annoyance TSE. Mid-annoyance values are positively correlated with the visibility threshold, and the relation can be described by a linear function. Mylène C. Q. Farias, John M. Foley, Sanjit K. Mitra |
ICASSP (2) | 1 |
| 2005 | Quality assessment using data hiding on perceptually important areasabstractIn this paper, we present a no-reference video quality metric that blindly estimates the quality of a video. The proposed approach makes use of a data hiding technique to embed a fragile mark into perceptually important areas of the video frame. To estimate the importance of an area, we take into account three perceptual features that are known to attract visual attention: motion, contrast, and color. At the receiver, the mark is extracted from the perceptually important areas of the decoded video. Then, a quality measure of the video is obtained by computing the degradation of the extracted mark. Simulation results indicate that the proposed video quality metric outperforms standard peak signal to noise ratio (PSNR) in estimating the perceived quality of a video. Additionally, results from a subjective experiment show that the metric output values increase monotonically with the mean annoyance scores gathered from the human observers. Marco Carli, Mylène C. Q. Farias, Elisa Drelie Gelasca, Roberto Tedesco, Alessandro Neri 0001 |
ICIP (3) | 2 |
| 2005 | No-reference video quality metric based on artifact measurementsabstractIn this paper we present a no-reference video quality metric based on individual measurements of three artifacts: blockiness, blurriness, and noisiness. The set of artifact metrics (physical strength measurements) was designed to be simple enough to be used in real-time applications. The metrics are tested using a proposed procedure that uses synthetic artifacts and subjective data obtained from previous experiments. The technique has the advantage of allowing us to test each metric on videos which contain only the desired artifact signal or a combination of artifact signals. Models for the overall annoyance based on a combination of the artifact metrics using both a Minkowski metric and a linear model are developed. Both models present a very good correlation with the data and show no statistical difference in their performances. Mylène C. Q. Farias, Sanjit K. Mitra |
ICIP (3) | 1 |
| 2005 | A robust error concealment technique using data hiding for image and video transmission over lossy channelsabstractA robust error concealment scheme using data hiding which aims at achieving high perceptual quality of images and video at the end-user despite channel losses is proposed. The scheme involves embedding a low-resolution version of each image or video frame into itself using spread-spectrum watermarking, extracting the embedded watermark from the received video frame, and using it as a reference for reconstruction of the parent image or frame, thus detecting and concealing the transmission errors. Dithering techniques have been used to obtain a binary watermark from the low-resolution version of the image/video frame. Multiple copies of the dithered watermark are embedded in frequencies in a specific range to make it more robust to channel errors. It is shown experimentally that, based on the frequency selection and scaling factor variation, a high-quality watermark can be extracted from a low-quality lossy received image/video frame. Furthermore, the proposed technique is compared to its two-part variant where the low-resolution version is encoded and transmitted as side information instead of embedding it. Simulation results show that the proposed concealment technique using data hiding outperforms existing approaches in improving the perceptual quality, especially in the case of higher loss probabilities. Chowdary Adsumilli, Mylène C. Q. Farias, Sanjit K. Mitra, Marco Carli |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Detectability and annoyance of synthetic blurring and ringing in video sequencesabstractOne approach for studying digital video impairments is to work with synthetic artifacts which look like real impairments, yet are simpler, purer and easier to describe. We have created synthetic ringing and blurring and inserted them in short video sequences. In a psychophysical experiment, we measured the probability of detection and the annoyance value of these artifacts as a function of their total squared error. Although ringing occurs only near edges and blurring can occur over wide areas of the images, there is no consistent difference between either the thresholds or mid-annoyance strengths. There are, on the other hand, interactions between the specific video and artifact type in the determination of these values. Mid-annoyance strength was found to be highly correlated with threshold. Also, we combined ringing and blurring to produce mixed artifacts. Their thresholds and mid-annoyance strengths tend to be intermediate between those of the individual artifacts. Their annoyance value is well predicted by a weighted sum of the annoyance values for blurring and ringing with weights of approximately 0.6 and 0.4, respectively. Mylène C. Q. Farias, Sanjit K. Mitra, John M. Foley |
ICASSP (3) | 1 |
| 2004 | Annoyance of spatio-temporal artifacts in segmentation quality assessmentabstractThis paper describes the results of a series of subjective experiments that investigated the annoyance caused by the most common artifacts present in segmented video sequences. Various types of artifacts were inserted into a reference segmented video, considered as ideal, and shown to our test subjects. The artifacts varied in their location, size, appearance and duration. Annoyance of segmentation artifacts are found to be tied up with their intrinsic characteristics (e.g., size, position) but only weakly related to the video content. The results identify the characteristics that should be taken into account in the design of a perceptually driven objective metric. Elisa Drelie Gelasca, Touradj Ebrahimi, Mylène C. Q. Farias, Marco Carli, Sanjit K. Mitra |
ICIP | 3 |
| 2003 | A hybrid constrained unequal error protection and data hiding scheme for packet video transmissionabstractA novel hybrid scheme with constrained unequal error protection (UEP) and data hiding is proposed; it maximizes the perceptual quality of the video at the end user to compensate for the effects of channel losses. The technique involves: (1) implementing a forcing function which weighs the objective perceptual quality of the video frame based on a hidden mark signal to give optimum protection level in the packet; (2) utilizing a data hiding mechanism to embed second level wavelet approximation coefficients of the frame in itself. An optimum UEP in the transmitted packets is sought using a constrained optimization approach. Simulation results show that the proposed technique outperforms existing non-adaptive error concealment approaches in improving the perceptual quality, especially for higher loss probabilities. Chowdary Adsumilli, Mylène C. Q. Farias, Marco Carli, Sanjit K. Mitra |
ICASSP (5) | 2 |
| 2003 | Perceptual contributions of blocky, blurry and noisy artifacts to overall annoyanceabstractIn this paper, we create synthetic artifacts that are perceived to be predominantly blocky, blurry or noisy. We present them alone or in various combinations and have subjects rate the perceived strength of each artifact and the overall annoyance of the combined artifacts. We found that a simple linear model with no interactions predicted how the perceived artifacts combine to determine overall annoyance. The estimated coefficients indicate that, when the artifacts are equated in perceived strength, blurriness contributes the most to annoyance followed by noise and then blockiness, although the relative weights of blockiness and noise vary with the video content. Mylène C. Q. Farias, Sanjit K. Mitra, John M. Foley |
ICME | 1 |
| 2003 | An accurate billing mechanism for multimedia communicationsabstractA novel billing mechanism is presented in this paper to provide telecommunication service users and advertisers with an effective billing mechanism based on the actual amount of advertisement data being transmitted/displayed. Experimental results show the effectiveness of the proposed system in terms of computational cost and performance. José Gabriel R. C. Gomes, Mylène C. Q. Farias, Sanjit K. Mitra, Marco Carli |
ICME | 2 |
| 2002 | A comparison between an objective quality measure and the mean annoyance values of watermarked videosabstractA comparison between an objective quality measure and the perceived mean annoyance values of watermarked videos is presented. A psychophysical experiment has been performed to measure the detection threshold and mean annoyance values of several watermarked videos, using two different marks. The results of this experiment were then compared with an objective quality measure, obtained through a tracing watermarking system. An estimation of the detection threshold of the watermarked videos was found. Mylène C. Q. Farias, Marco Carli, Sanjit K. Mitra, Alessandro Neri 0001 |
ICIP (3) | 1 |