Junyong You

dblp:65/3964 · DBLP profile ↗
← Back
43ranked-venue papers
31as first author
13since 2021 · last 2026
0000-0002-4288-5244ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 28 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing aesthetic image generation with reinforcement learning guided prompt optimization in stable diffusion
abstract
• An attention-driven image aesthetic assessment model has been proposed in a deep network consisting of three modules: spatial, channel and color. • A reinforcement learning framework for prompt improvement for increased aesthetic scores of images generated by Stable Diffusion. • Large language models for prompt editing to improve the aesthetics of generated images, guiding the actions in a reinforcement learning process. Generative models, e.g., stable diffusion, excel at producing compelling images but remain highly dependent on crafted prompts. Refining prompts for specific objectives, especially aesthetic quality, is time-consuming and inconsistent. We propose a novel approach that leverages LLMs to enhance prompt refinement process for stable diffusion. First, we propose a model to predict aesthetic image quality, examining various aesthetic elements in spatial, channel, and color domains. Reinforcement learning is employed to refine the prompt, starting from a rudimentary version and iteratively improving them with LLM’s assistance. This iterative process is guided by a policy network updating prompts based on interactions with the generated images, with a reward function measuring aesthetic improvement and adherence to the prompt. Our experimental results demonstrate that this method significantly boosts the visual quality of generated images when using these refined prompts. Beyond image synthesis, this approach provides a broader framework for improving prompts across diverse applications with the support of LLMs.
Junyong You, Bin Hu 0016
J. Vis. Commun. Image Represent.1
2025 Enhancing Visual Aesthetics in Stable Diffusion: A Reinforcement Learning Approach
abstract
Generative models such as stable diffusion have recently achieved significant success in producing high-quality images conditioned on textual inputs. It is difficult to control the quality of generated images, such as aesthetic quality, after an image generative model has been trained. In this paper, we pro-pose a novel method that leverages reinforcement learning (RL) to enhance the aesthetic appeal of images produced by stable diffusion models. An aesthetic assessment model has been first developed by modelling influence factors of image aesthetics in a deep network and then trained on publicly available datasets. The assessment model produces an aesthetic score of an image serving as the reward function in the RL framework. By reframing the denoising process of the stable diffusion model as a sequential decision-making problem, the intermediate denoising steps can be formulated as actions in a Markov decision process (MDP). The proximal policy optimization (PPO) algorithm can be applied in the RL framework to optimize the network parameters (i.e., U-Net) in the stable diffusion model, aiming to maximize the expected aesthetic reward. Through extensive experiments, we demonstrate that the proposed method can significantly improve the aesthetic quality of the generated images while maintaining their diversity and adherence to input prompts.
Junyong You, Bin Hu 0016
SMC1
2024 Gated Transformer Representing Region Importance for Image Quality Assessment
abstract
Deep neural networks, particularly convolutional neural networks (CNNs), have shown significant promise in image quality assessment (IQA), yet the underlying workings of these models in IQA remain partially unexplored. This study unveils a novel positionally masked transformer, shedding light on how various regions of an image influence its overall quality. Surprisingly, the findings reveal that half of an image may exert only a marginal influence on image quality, while the remaining half proves vital. This observation has been extended to other CNN-based IQA models, unearthing a consistent pattern where specific image regions significantly shape overall quality. In a stride to understand these phenomena, three semantic measures: saliency, frequency, and objectness, have been identified, exhibiting a strong correlation with the importance of image regions in IQA. Building upon these insights, a new gated operation has been proposed, representing the fluctuating significance of regions in image quality. A gate, integrable into a transformer encoder for IQA, serves to pinpoint the crucial spatial regions, enhancing their impact by amplifying attention weights. The resulting gated transformer has been rigorously tested on publicly available IQA datasets, demonstrating exceptional performance and reinforcing the innovative nature of this approach. The success of this study paves the way for more intricate and insightful analyses of IQA.
Junyong You, Jari Korhonen
IJCNN1
2023 On the Explainable Detection of Stress Levels Using Heart Rate Variability Based Deep Neural Networks
abstract
This paper presents one of the first explorations of transparency and explainability of Heart Rate Variability (HRV) based deep learning models designed for stress detection. We employed Shapley additive explanations (SHAP) as an explainable AI (XAI) method, and cross-validated the results with saliency maps, which provides valuable insights into the main contributing factors for decision-making process of these deep models.
Debasish Ghose, Jari Korhonen, Junyong You, Soumya P. Dash
HealthCom4
2023 Half of an Image is Enough for Quality Assessment
abstract
Deep networks show promising performance in image quality assessment (IQA), whereas few studies have investigated how a deep model works. In this work, a positional masked transformer for IQA is first developed, based on which we observe that half of an image might contribute trivially to image quality, whereas the other half is crucial. Such observation is generalized to that half of the image regions can dominate image quality in several CNN-based IQA models. Motivated by this observation, three semantic measures (saliency, frequency, objectness) are then derived, showing high accordance with importance degree of image regions in IQA.
Junyong You, Jari Korhonen
ICIP1
2023 Apples and Oranges? Assessing Image Quality over Content Recognition
abstract
Image recognition and quality assessment are two important viewing tasks, while potentially following different vis-ual mechanisms. This paper investigates if the two tasks can be performed in a multitask learning manner. A sequential spatial-channel attention module is proposed to simulate the visual attention and contrast sensitivity mechanisms that are crucial for con-tent recognition and quality assessment. Spatial attention is shared between content recognition and quality assessment, while channel attention is solely for quality assessment. Such attention module is integrated into Transformer to build a uniform model for the two viewing tasks. The experimental results have demonstrated that the proposed uniform model can achieve promising performance for both quality assessment and content recognition tasks.
Junyong You, Zheng Zhang 0063
ISCAS1
2023 Fast Accurate Fish Recognition with Deep Learning Based on a Domain-Specific Large-Scale Fish Dataset
Zhaoqi Chu, Jari Korhonen, Xiangrong Liu, Juan Liu 0003, Lvping Fang, Weidi Yang, Debasish Ghose, Junyong You
MMM (1)11
2022 Efficient Transformer with Locally Shared Attention for Video Quality Assessment
abstract
Transformer has shown outstanding performance in time-series data processing, which can definitely facilitate quality assessment of video sequences. However, the quadratic time and memory complexities of Transformer potentially impede its application to long video sequences. In this work, we study a mechanism of sharing attention across video clips in video quality assessment (VQA) scenario. Consequently, an efficient architecture based on integrating shared multi-head attention (MHA) into Transformer is proposed for VQA, which greatly ease the time and memory complexities. A long video sequence is first divided into individual clips. The quality features derived by an image quality model on each frame in a clip are aggregated by a shared MHA layer. The aggregated features across all clips are then fed into a global Transformer encoder for quality prediction at sequence level. The proposed model with a lightweight architecture demonstrates promising performance in no-reference VQA (NR-VQA) modelling on publicly available databases. The source code can be found at https://github.com/junyongyou/lagt_vqa.
Junyong You
ICIP1
2022 Explore Spatial and Channel Attention in Image Quality Assessment
abstract
As a subjective concept, perceived image quality is heavily affected by visual mechanisms, e.g., selective attention and contrast sensitivity. This work proposes a lightweight attention module in image quality assessment (IQA) to simulate spatial attention and contrast sensitivity mechanisms. The attention module can extract essential information from a CNN backbone for image quality perception using two sequential attention blocks: a spatial block for mimicking selective attention in spatial domain and a channel block for contrast sensitivity. Experimental results on two large-scale publicly available IQA datasets have demonstrated promising performance of the proposed approach. The source code can be found at https://github.com/junyongyou/sca_iqa.
Junyong You
ICIP1
2022 Attention integrated hierarchical networks for no-reference image quality assessment
abstract
Quality assessment of natural images is influenced by perceptual mechanisms, e.g., attention and contrast sensitivity, and quality perception can be generated in a hierarchical process. This paper proposes an architecture of Attention Integrated Hierarchical Image Quality networks (AIHIQnet) for no-reference quality assessment. AIHIQnet consists of three components: general backbone network, perceptually guided neck network, and head network. Multi-scale features extracted from the backbone network are fused to simulate image quality perception in a hierarchical manner. The attention and contrast sensitivity mechanisms modelled by an attention module capture essential information for quality perception. Considering that image rescaling potentially affects perceived quality, appropriate pooling methods in the non-convolution layers in AIHIQnet are employed to accept images with arbitrary resolutions. Comprehensive experiments on publicly available databases demonstrate outstanding performance of AIHIQnet compared to state-of-the-art models. Ablation experiments were performed to investigate the variants of the proposed architecture and reveal importance of individual components.
Junyong You, Jari Korhonen
J. Vis. Commun. Image Represent.1
2021 Transformer For Image Quality Assessment
abstract
Transformer has become the new standard method in natural language processing (NLP), and it also attracts research interests in computer vision area. In this paper we investigate the application of Transformer in Image Quality (TRIQ) assessment. Following the original Transformer encoder employed in Vision Transformer (ViT), we propose an architecture of using a shallow Transformer encoder on the top of a feature map extracted by convolution neural networks (CNN). Adaptive positional embedding is employed in the Transformer encoder to handle images with arbitrary resolutions. Different settings of Transformer architectures have been investigated on publicly available image quality databases. We have found that the proposed TRIQ architecture achieves outstanding performance. The implementation of TRIQ is published on Github (https://github.com/junyongyou/triq).
Junyong You, Jari Korhonen
ICIP1
2021 Reproducibility Companion Paper: Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial Features
abstract
Blind natural video quality assessment (BVQA), also known as no-reference video quality assessment, is a highly active research topic. In our recent contribution titled "Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial Features" published in ACM Multimedia 2020, we proposed a two-level video quality model employing statistical temporal features and spatial features extracted by a deep convolutional neural network (CNN) for this purpose. At the time of publishing, the proposed model (CNN-TLVQM) achieved state-of-the-art results in BVQA. In this paper, we describe the process of reproducing the published results by using CNN-TLVQM on two publicly available natural video quality datasets.
Jari Korhonen, Yicheng Su, Junyong You, Steven Alexander Hicks, Cise Midoglu
ACM Multimedia3
2021 Long Short-term Convolutional Transformer for No-Reference Video Quality Assessment
abstract
No-reference video quality assessment has not been widely benefited from deep learning, mainly due to the complexity, diversity and particularity of modelling spatial and temporal characteristics in quality assessment scenario. Image quality assessment (IQA) performed on video frames plays a key role in NR-VQA. A perceptual hierarchical network (PHIQNet) with an integrated attention module is first proposed that can appropriately simulate the visual mechanisms of contrast sensitivity and selective attention in IQA. Subsequently, perceptual quality features of video frames derived from PHIQNet are fed into a long short-term convolutional Transformer (LSCT) architecture to predict the perceived video quality. LSCT consists of CNN formulating quality features in video frames within short-term units that are then fed into Transformer to capture the long-range dependence and attention allocation over temporal units. Such architecture is in line with the intrinsic properties of VQA. Experimental results on publicly available video quality databases have demonstrated that the LSCT architecture based on PHIQNet significantly outperforms state-of-the-art video quality models.
Junyong You
ACM Multimedia1
2020 Attention Boosted Deep Networks For Video Classification
abstract
Video classification can be performed by summarizing image contents of individual frames into one class by deep neural networks, e.g., CNN and LSTM. Human interpretation of video content is influenced by the attention mechanism. In other words, video class can be more attentively decided by certain information than others. In this paper, we propose to integrate the attention mechanism into deep networks for video classification. The proposed framework employs 2D CNN networks with ImageNet pretrained weights to extract features of video frames that are then fed to a bidirectional LSTM network for video classification. An attention block has been developed that can be added after the LSTM network in the proposed framework. Several different 2D CNN architectures have been tested in the experiments. The results with respect to two publicly available datasets have demonstrated that integrating attention can boost the performance of deep networks in video classification compared to not applying the attention block. We also found out that applying attention to the LSTM outputs on the VGG19 architecture provides the highest classification accuracy in the proposed framework.
Junyong You, Jari Korhonen
ICIP1
2020 Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial Features
abstract
Due to the wide range of different natural temporal and spatial distortions appearing in user generated video content, blind assessment of natural video quality is a challenging research problem. In this study, we combine the hand-crafted statistical temporal features used in a state-of-the-art video quality model and spatial features obtained from convolutional neural network trained for image quality assessment via transfer learning. Experimental results on two recently published natural video quality databases show that the proposed model can predict subjective video quality more accurately than the publicly available video quality models representing the state-of-the-art. The proposed model is also competitive in terms of computational complexity.
Jari Korhonen, Yicheng Su, Junyong You
ACM Multimedia3
2019 Deep Neural Networks for No-Reference Video Quality Assessment
abstract
Video quality assessment (VQA) is a challenging task due to the complexity of modeling perceived quality characteristics in both spatial and temporal domains. A novel no-reference (NR) video quality metric (VQM) is proposed in this paper based on two deep neural networks (NN), namely 3D convolution network (3D-CNN) and a recurrent NN composed of long short-term memory (LSTM) units. 3D-CNNs are utilized to extract local spatiotemporal features from small cubic clips in video, and the features are then fed into the LSTM networks to predict the perceived video quality. Such design can elaborately tackle the issue of insufficient training data whilst also efficiently capture perceptive quality features in both spatial and temporal domains. Experimental results with respect to two publicly available video quality datasets have demonstrate that the proposed quality metric outperforms the other compared NR quality metrics.
Junyong You, Jari Korhonen
ICIP1
2019 Weather Data Integrated Mask R-CNN for Automatic Road Surface Condition Monitoring
abstract
Monitoring road surface conditions plays a crucial role in driving safety and road maintenance, especially in winter seasons. Traditional methodologies often employ manual inspection and expensive instruments, e.g., NIR cameras. However, image analysis based on normal cameras can provide an economical and efficient solution for road surface monitoring. This paper presents an automatic classification model of road surface conditions using a deep learning approach based on road images and weather measurement. A modified mask R-CNN model has been developed by integrating weather data based on transfer learning. Experimental results with respect to manual judgment of road surface conditions have demonstrated very high accuracy of the developed model.
Junyong You
VCIP1
2016 Perceptual contrast sensitivity based video quality assessment in DCT domain
abstract
Video quality assessment can be performed by comparing distorted video with the undistorted version by taking human vision system (HVS) into account. A perceptual vision sensitivity model is developed in this paper by systematically integrating visual attention and foveation mechanism into contrast sensitivity function (CSF). The model can accurately estimate a critical frequency beyond which the HVS cannot perceive contrast changes. Subsequently, a video quality metric is proposed by applying the new sensitivity model to a similarity measure between distorted video and reference one in the DCT domain. Experimental results with respect to publicly available video quality databases have demonstrated promising performance of the proposed quality metric.
Junyong You
ICIP1
2014 Quality of dirty: A decision making assessment methodology for automatic license plate recognition under dirtied conditions
abstract
Automatic license plate recognition (ALPR) from vehicle images plays an important role in intelligent traffic management. A critical problem is that the cameras can be dirtied such that the captured images are unrecognizable. This paper defines a new concept of Quality of Dirty (QoD) of captured images indicating when the dirtied cameras need to be cleaned. An accurate image QoD metric is proposed based on relevant image features and support vector regression (SVR), and it can handle a tricky issue of differentiating captured imaged containing dirtied vehicles from images captured by a dirtied camera. A subjective assessment to collect ground-truth QoD scores of real captured images has also been conducted. Experiments have demonstrated that the proposed image QoD metric achieves high accuracy when predicting the dirtied degree of cameras in the tunnel environment.
Junyong You, Hans Christian Bolstad, Andrew Perkis
ICIP1
2014 Attention Driven Foveated Video Quality Assessment
abstract
Contrast sensitivity of the human visual system to visual stimuli can be significantly affected by several mechanisms, e.g., vision foveation and attention. Existing studies on foveation based video quality assessment only take into account static foveation mechanism. This paper first proposes an advanced foveal imaging model to generate the perceived representation of video by integrating visual attention into the foveation mechanism. For accurately simulating the dynamic foveation mechanism, a novel approach to predict video fixations is proposed by mimicking the essential functionality of eye movement. Consequently, an advanced contrast sensitivity function, derived from the attention driven foveation mechanism, is modeled and then integrated into a wavelet-based distortion visibility measure to build a full reference attention driven foveated video quality (AFViQ) metric. AFViQ exploits adequately perceptual visual mechanisms in video quality assessment. Extensive evaluation results with respect to several publicly available eye-tracking and video quality databases demonstrate promising performance of the proposed video attention model, fixation prediction approach, and quality metric.
Junyong You, Touradj Ebrahimi, Andrew Perkis
IEEE Trans. Image Process.1
2013 Enhancing coded video quality with perceptual foveation driven bit allocation strategy
abstract
Contrast sensitivity plays an important role in visual perception when viewing external stimuli, e.g., video, and it has been taken into account in development of advanced video coding algorithms. This paper proposes a perceptual foveation model based on accurate prediction of video fixations and modeling of contrast sensitivity function (CSF). Consequently, an adaptive bit allocation strategy in H.264/AVC video compression is proposed by considering visible frequency threshold of the human visual system (HVS). A subjective video quality assessment together with objective quality metrics have been performed and demonstrated that the proposed perceptual foveation driven bit allocation strategy can significantly improve the perceived quality of coded video compared with standard coding scheme and another visual attention guided coding approach.
Junyong You, Xue-Cheng Tai
VCIP1
2012 Video quality metric based on fixation prediction and foveal imaging
abstract
This paper proposes a full-reference video quality metric based on foveated vision mechanism. Due to a non-uniform distribution of photo-receptors on the retina, the human visual system (HVS) has the highest resolution around the fixation point of eyes and dramatically decreases away from this point. Two key factors in the foveated vision are fixation point and retinal eccentricities of different visual objects. Based on an advanced video attention model in quality assessment scenarios, eye fixations are predicted from the attention map using a winner-takes-all (WTA) neural network. Four quality features describing distortions on luminance, spatial and temporal activities, as well as chrominance are derived between foveated representations of reference and distorted video frames. These quality features are then combined together by an appropriate spatiotemporal pooling scheme to build a video quality metric. Experimental results with respect to publicly available video quality databases demonstrate that the proposed quality model outperforms a previously proposed foveated video quality metric as well as state-of-the-art video quality models.
Junyong You, Touradj Ebrahimi, Andrew Perkis
ICIP1
2012 Video Gaze Prediction: Minimizing Perceptual Information Loss
abstract
Automatic detection of visually interesting regions and gaze points plays an important role in many video applications. Due to limited ability of the human visual system (HVS) when processing visual stimuli at any instant, a natural function of gaze changes is to collect as much information as possible to form an accurate understanding of the visual scene. This paper proposes an automatic gaze prediction algorithm by modeling such function. An improved foveal imaging model is developed by taking visual attention and temporal visual characteristics into account. Gaze changes are predicted based on minimizing perceptual information loss due to the foveated vision mechanism. Experimental results against a video eye-tracking database demonstrate a promising performance of the proposed gaze prediction algorithm.
Junyong You
ICME1
2012 Visual Contrast Sensitivity Guided Video Quality Assessment
abstract
Contrast sensitivity is an important characteristic of the human visual system (HVS), which is widely used in image and video signal processing. For visual quality assessment, static spatio-temporal frequency based contrast sensitivity function (CSF) is often used, while the contrast sensitivity can also be affected by smooth pursuit of eyes when tracking attentive regions in the field of view. This paper proposes to tune CSF based on an attention map derived from a visual attention model. The tuned CSF formulated by spatial frequency, temporal velocity, and visual attention map is used to filter video signals in order to construct a quality metric according to the difference of the filtered signals between a reference video and its distorted version. Experimental results demonstrate that the proposed attention tuned spatio-velocity CSF outperforms the traditional spatio-temporal CSF in evaluating the perceived video quality.
Junyong You, Liyuan Xing, Andrew Perkis, Touradj Ebrahimi
ICME1
2012 Scale- and Rotation-Invariant Local Binary Pattern Using Scale-Adaptive Texton and Subuniform-Based Circular Shift
abstract
This paper proposes an effective scale- and rotation-invariant local binary pattern (LBP) feature for texture classification. A circular neighboring set of an image pixel is defined as a scale-adaptive texton by taking into account the fundamental local structure property of the pixel. The scale space of a texture image is derived by the Laplacian of the Gaussian and then employed to determine the optimal scale of each pixel reflecting the characteristic length of the corresponding structure and determining the radius of the scale-adaptive texton. Different pixels have different optimal scales, resulting in the scale invariance. Contrary to the traditional LBP features that usually ignore global spatial information, the proposed method also defines subuniform patterns of each uniform pattern to improve the discrimination. For each uniform pattern, the subuniform pattern with the maximum statistical value is defined as the dominant orientation subuniform pattern. It is moved to the first column, and the others are circularly shifted. Experimental results demonstrate a good discrimination capability of the proposed scale- and rotation-invariant LBP in texture classification. Particularly, the LBP based on the scale-adaptive texton is promising to be powerful for texture description and scale-invariant texture classification, and the circular shift subuniform LBP can further improve the performance in the rotation-invariant texture classification.
Zhi Li 0003, Guizhong Liu, Yang Yang 0042, Junyong You
IEEE Trans. Image Process.4
2012 Assessment of Stereoscopic Crosstalk Perception
abstract
Stereoscopic three-dimensional (3-D) services do not always prevail when compared with their two-dimensional (2-D) counterparts, though the former can provide more immersive experience with the help of binocular depth. Various specific 3-D artefacts might cause discomfort and severely degrade the Quality of Experience (QoE). In this paper, we analyze one of the most annoying artefacts in the visualization stage of stereoscopic imaging, namely, crosstalk, by conducting extensive subjective quality tests. A statistical analysis of the subjective scores reveals that both scene content and camera baseline have significant impacts on crosstalk perception, in addition to the crosstalk level itself. Based on the observed visual variations during changes in significant factors, three perceptual attributes of crosstalk are summarized as the sensorial results of the human visual system (HVS). These are shadow degree, separation distance, and spatial position of crosstalk. They are classified into two categories: 2-D and 3-D perceptual attributes, which can be described by a Structural SIMilarity (SSIM) map and a filtered depth map, respectively. An objective quality metric for predicting crosstalk perception is then proposed by combining the two maps. The experimental results demonstrate that the proposed metric has a high correlation (over 88%) when compared with subjective quality scores in a wide variety of situations.
Liyuan Xing, Junyong You, Touradj Ebrahimi, Andrew Perkis
IEEE Trans. Multim.2
2011 Objective metrics for quality of experience in stereoscopic images
abstract
Stereoscopic Quality of Experience (QoE) is the result of a complex combination of different influencing factors. Previously we had investigated the effect of factors such as scene content, camera baseline, screen size and viewing position on stereoscopic QoE using subjective tests. In this paper, we propose two objective metrics for predicting stereoscopic QoE using bottom-up and top-down approaches, respectively. Specifically, the bottom-up metric is based on characterizing the significant factors of QoE directly, which are scene content, camera baseline, screen size and crosstalk level. While the top-down metric interprets QoE from its perceptual attributes, including crosstalk perception and perceived depth. These perceptual attributes are modeled by their individual relationship with the significant factors and then combined linearly to build the top-down metric. Both proposed metrics have been validated against our own database and a publicly available database, showing a high correlation (over 86%) with the subjective scores of stereoscopic QoE.
Liyuan Xing, Junyong You, Touradj Ebrahimi, Andrew Perkis
ICIP2
2011 Audiovisual quality fusion based on relative multimodal complexity
abstract
In multimodal presentations the perceived audiovisual quality assessment is significantly influenced by the content of both the audio and visual tracks. Based on our earlier subjective quality test for finding the optimal trade-off between audio and video quality, this paper proposes a novel method for relative multimodal complexity analysis to derive the fusion parameter in objective audiovisual quality metrics. Audio and video qualities are first estimated separately using advanced quality models, and then they are combined into the overall audiovisual quality using a linear fusion. Based on carefully designed auditory and visual features, the relative complexity analysis model across sensory modalities is proposed for deriving the fusion parameter. Experimental results have demonstrated that the content adaptive fusion parameter can improve the prediction accuracy of objective audiovisual quality metrics, compared to the fusion parameters obtained from the subjective quality tests using other known optimization methods.
Junyong You, Jari Korhonen, Ulrich Reiter
ICIP1
2011 Visual attention tuned spatio-velocity contrast sensitivity for video quality assessment
abstract
Contrast sensitivity is an important characteristic of the human visual system (HVS), which is widely used in image and video signal processing. For visual quality assessment, static spatio-temporal frequency based contrast sensitivity function (CSF) is often used, while the contrast sensitivity can also be affected by smooth pursuit of eyes when tracking attentive regions in the field of view. This paper proposes to tune CSF based on an attention map derived from a visual attention model. The tuned CSF formulated by spatial frequency, temporal velocity, and visual attention map is used to filter video signals in order to construct a quality metric according to the difference of the filtered signals between a reference video and its distorted version. Experimental results demonstrate that the proposed attention tuned spatio-velocity CSF outperforms the traditional spatio-temporal CSF in evaluating the perceived video quality.
Junyong You, Touradj Ebrahimi, Andrew Perkis
ICME1
2011 Modeling motion visual perception for video quality assessment
abstract
Contrast sensitivity of Human Visual System (HVS) plays an important role in perceiving visual stimuli, and consequently, it has a significant impact on the perceived video quality. This paper proposes a visual perception model based on foveated vision and motion perception. The reference and the distorted video sequences are processed by the visual perception model to generate the perceived stimuli in HVS. The perceived difference of the processed sequences is measured in spatial and temporal domains considering the visual sensitivity. An advanced pooling scheme is proposed based on the visual attention mechanism, eye movement type, and influence of temporal quality variation, in order to estimate the perceived video quality. Experimental results demonstrate that the proposed metric significantly outperforms state-of-the-art quality models with respect to a combined eye-tracking and subjective video quality assessment data set.
Junyong You, Touradj Ebrahimi, Andrew Perkis
ACM Multimedia1
2011 Balancing Attended and Global Stimuli in Perceived Video Quality Assessment
abstract
The visual attention mechanism plays a key role in the human perception system and it has a significant impact on our assessment of perceived video quality. In spite of receiving less attention from the viewers, unattended stimuli can still contribute to the understanding of the visual content. This paper proposes a quality model based on the late attention selection theory, assuming that the video quality is perceived via two mechanisms: global and local quality assessment. First we model several visual features influencing the visual attention in quality assessment scenarios to derive an attention map using appropriate fusion techniques. The global quality assessment as based on the assumption that viewers allocate their attention equally to the entire visual scene, is modeled by four carefully designed quality features. By employing these same quality features, the local quality model tuned by the attention map considers the degradations on the significantly attended stimuli. To generate the overall video quality score, global and local quality features are combined by a content adaptive linear fusion method and pooled over time, taking the temporal quality variation into consideration. The experimental results have been compared to results from appropriate eye tracking and video quality assessment experiments, demonstrating promising performance.
Junyong You, Jari Korhonen, Andrew Perkis, Touradj Ebrahimi
IEEE Trans. Multim.1
2010 Spatial and temporal pooling of image quality metrics for perceptual video quality assessment on packet loss streams
abstract
Video streaming through bandwidth-limited channels often suffer from packet losses. Therefore, perceptual quality assessment on video sequences with packet losses is a critical issue in digital video communications. This paper analyzes several image quality metrics and evaluates their applications using spatial and temporal pooling schemes in perceptual video quality assessment for video streams with packet losses. Several approaches using Minkowski summation and averages over different distorted spatial regions and temporal frames to pool the spatial and temporal qualities are evaluated. The experimental results with respect to the subjective video quality measurements demonstrate that the subjects are more sensitive to the most annoying spatial regions and temporal segments when assessing the video quality of the lossy streams.
Junyong You, Jari Korhonen, Andrew Perkis
ICASSP1
2010 On the relationship between perceptual impact of source and channel distortions in video sequences
abstract
It is known that peak signal-to-noise ratio (PSNR) can be used for assessing the relative qualities of distorted video sequences meaningfully only if the compared sequences contain similar types of distortions. In this paper, we propose a model for rough assessment of the bias in PSNR results, when video sequences with both channel and source distortion are compared against video sequences with source distortion only. The proposed method can be used to compare the relative perceptual quality levels of video sequences with different distortion types more reliably than using plain PSNR.
Jari Korhonen, Ulrich Reiter, Junyong You
ICIP3
2010 Spatial noise shaping using convex optimization for perceptual image coding
abstract
In this paper we propose a new convex optimization framework for precise spatial noise shaping. The effectiveness of this new technique is demonstrated in the application of perceptual coding of images. A modified JPEG 2000 codec is implemented using the proposed new framework and compared with existing perceptual coding algorithms. Results of subjective tests show that the new framework can provide a significant improvement in bitrate savings compared to the best performing wavelet domain technique. The algorithm allows much more precise control of distortion than existing spatial domain techniques and is fully compliant with part 1 of the JPEG 2000 standard.
Mark R. Pickering, Junyong You, Touradj Ebrahimi, Andrew Perkis
ICIP2
2010 A perceptual quality metric for stereoscopic crosstalk perception
abstract
Compared to metrics proposed to assess the quality of two-dimensional (2D) images, there are very few metrics devoted to quality assessment of stereoscopic presentations. Crosstalk is one of the most annoying distortions in the visualization stage of stereoscopic imaging technology. This paper proposes a perceptual quality metric which takes characteristics of stereoscopic images into account for predicting quality levels of crosstalk perception in stereoscopic images, based on an understanding of three main factors, crosstalk level, camera baseline and scene content. The experimental results demonstrate that the proposed metric has Pearson correlation of 87.7% when compared to the ground truth results from the subjective experiments on the crosstalk perception, which is much better than the traditional 2D metrics without integrating 3D depth information.
Liyuan Xing, Junyong You, Touradj Ebrahimi, Andrew Perkis
ICIP2
2010 Asymmetric multi-view video coding based on chrominance reconstruction
abstract
Three-dimensional video (3DV) technology is becoming increasingly popular, as it can provide high quality and immersive experience to end users. Huge amount of data for storage and transmission is an important problem to be solved. In this paper, an asymmetric MVC method is proposed. Color correction is first performed as a preprocessing step to provide consistent color information among views. Then, all color corrected views are classified into color views and non-color views. The chrominance information in non-color views is all discarded and only preserved in color view in MVC codec. Thus, a large amount of coding bitrate can be saved. At the decoder, a chrominance reconstruction algorithm is presented to achieve accurate color reconstruction for those non-color views. Experimental results show that the proposed method can achieve large bitrate saving against the results compressed with the original JMVM codec. Moreover, the proposed method can obtain better reconstruction quality without noticeable quality degradation.
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Junyong You
ICME4
2010 Attention modeling for video quality assessment: Balancing global quality and local quality
abstract
This paper proposes to evaluate video quality by balancing two quality components: global quality and local quality. The global quality is a result from subjects allocating their attention equally to all regions in a frame and all frames in a video. It is evaluated by image quality metrics (IQM) with averaged spatiotemporal pooling. The local quality is derived from visual attention modeling and quality variations over frames. Saliency, motion, and contrast information are taken into account in modeling visual attention, which is then integrated into IQMs to calculate the local quality of a video frame. The local quality of a video sequence is calculated by pooling local quality values over all frames with a temporal pooling scheme derived from the known relationship between perceived video quality and the frequency of temporal quality variations. The overall quality of a distorted video is a weighted average between the global quality and the local quality. Experimental results demonstrate that the combination of the global quality and local quality outperforms both sole global quality and local quality, as well as other quality models, in video quality assessment. In addition, the proposed video quality modeling algorithm can improve the performance of image quality metrics on video quality assessment compared to the normal averaged spatiotemporal pooling scheme.
Junyong You, Jari Korhonen, Andrew Perkis
ICME1
2010 An objective metric for assessing quality of experience on stereoscopic images
abstract
Most quality models for stereoscopic presentations are dedicated to measuring quality degradation caused by compression artefacts. However, non-compression distortions induced during acquisition and presentation usually have significant influence on 3D viewing experience. In this paper, we propose an objective metric for viewing experience assessment by taking camera baseline and binocular distortion crosstalk into consideration. In particular, the proposed metric is based on our previous work on both subjective evaluation and objective assessment of crosstalk perception. Results on a publicly available stereoscopic quality database demonstrate that the proposed metric can achieve more than 87% correlation with subjective assessment of viewing experience.
Liyuan Xing, Junyong You, Touradj Ebrahimi, Andrew Perkis
MMSP2
2010 A semantic framework for video genre classification and event analysis
Junyong You, Guizhong Liu, Andrew Perkis
Signal Process. Image Commun.1
2010 Perceptual-based quality assessment for audio-visual services: A survey
Junyong You, Ulrich Reiter, Miska M. Hannuksela, Moncef Gabbouj, Andrew Perkis
Signal Process. Image Commun.1
2009 An objective video quality metric based on spatiotemporal distortion
abstract
This paper proposes an objective video quality metric based on an analysis of spatial and temporal distortions. Spatial quality features extracted from the spatiotemporal region of reference and distorted videos are used to express the spatial distortion. Temporal distortion, caused by frame freezing resulting from a packet loss, is derived from the spatial distortion before and after the frozen frames. The overall quality is predicted according to the weighted combination of qualities over all the temporal regions. The experimental results with respect to the subjective measurements demonstrate the fast computation and promising performance of the proposed model compared with existing methods.
Junyong You, Miska M. Hannuksela, Moncef Gabbouj
ICIP1
2009 Perceptual quality assessment based on visual attention analysis
abstract
Most existing quality metrics do not take the human attention analysis into account. Attention to particular objects or regions is an important attribute of human vision and perception system in measuring perceived image and video qualities. This paper presents an approach for extracting visual attention regions based on a combination of a bottom-up saliency model and semantic image analysis. The use of PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural SIMilarity) in extracted attention regions is analyzed for image/video quality assessment, and a novel quality metric is proposed which can exploit the attributes of visual attention information adequately. The experimental results with respect to the subjective measurement demonstrate that the proposed metric outperforms the current methods.
Junyong You, Andrew Perkis, Miska M. Hannuksela, Moncef Gabbouj
ACM Multimedia1
2007 A Multiple Visual Models Based Perceptive Analysis Framework for Multilevel Video Summarization
abstract
In this paper, we propose a generic framework to human perception analysis in video understanding based on multiple visual cues. Video features that prominently influence human perception, such as motion, contrast, special scenes, and statistical rhythm, are first extracted and modeled. A perception curve that corresponds to human perception change is then constructed from these individual models using linear or priority based fusion approach. As an important application of the perceptive analysis framework, a feasible scheme for video summarization is implemented in order to demonstrate the validity, robustness, and generality of the proposed framework. The frames that correspond to the peak points in these individual models and the fusion curve are extracted as multilevel summarizations that include video keywords, keyframes, and dynamic segments. The subjective evaluations from a supplementary volunteer study on video summarizations indicate that the analysis framework is effective and offer a promising approach to semantic video management, access, and understanding
Junyong You, Guizhong Liu, Hongliang Li 0001
IEEE Trans. Circuits Syst. Video Technol.1