EDBT 2026 Demo / reviewers in the wild / expert
Jari Korhonen
dblp:15/6505
· DBLP profile ↗
54ranked-venue papers
22as first author
12since 2021 · last 2025
0000-0003-4354-5310ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 20 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorComputer networks · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NERO: Explainable Out-of-Distribution Detection with Neuron-Level Relevance in Gastrointestinal Imaging
Anju Chhetri, Jari Korhonen, Prashnna Gyawali, Binod Bhattarai |
MICCAI (10) | 2 |
| 2024 | High Resolution Image Quality DatabaseabstractWith technology for digital photography and high resolution displays rapidly evolving and gaining popularity, there is a growing demand for blind image quality assessment (BIQA) models for high resolution images. Unfortunately, the publicly available large scale image quality databases used for training BIQA models contain mostly low or general resolution images. Since image resizing affects image quality, we assume that the accuracy of BIQA models trained on low resolution images would not be optimal for high resolution images. Therefore, we created a new high resolution image quality database (HRIQ), consisting of 1120 images with resolution of 2880 × 2160 pixels. We conducted a subjective study to collect the subjective quality ratings for HRIQ in a controlled laboratory setting, resulting in accurate MOS at high resolution. To demonstrate the importance of a high resolution image quality database for training BIQA models to predict mean opinion scores (MOS) of high resolution images accurately, we trained and tested several traditional and deep learning based BIQA methods on different resolution versions of our database. The database is publicly available in https://github.com/jarikorhonen/hriq. Jari Korhonen |
ICASSP | 3 |
| 2024 | Gated Transformer Representing Region Importance for Image Quality AssessmentabstractDeep neural networks, particularly convolutional neural networks (CNNs), have shown significant promise in image quality assessment (IQA), yet the underlying workings of these models in IQA remain partially unexplored. This study unveils a novel positionally masked transformer, shedding light on how various regions of an image influence its overall quality. Surprisingly, the findings reveal that half of an image may exert only a marginal influence on image quality, while the remaining half proves vital. This observation has been extended to other CNN-based IQA models, unearthing a consistent pattern where specific image regions significantly shape overall quality. In a stride to understand these phenomena, three semantic measures: saliency, frequency, and objectness, have been identified, exhibiting a strong correlation with the importance of image regions in IQA. Building upon these insights, a new gated operation has been proposed, representing the fluctuating significance of regions in image quality. A gate, integrable into a transformer encoder for IQA, serves to pinpoint the crucial spatial regions, enhancing their impact by amplifying attention weights. The resulting gated transformer has been rigorously tested on publicly available IQA datasets, demonstrating exceptional performance and reinforcing the innovative nature of this approach. The success of this study paves the way for more intricate and insightful analyses of IQA. Junyong You, Jari Korhonen |
IJCNN | 3 |
| 2024 | 3DTA: No-Reference 3D Point Cloud Quality Assessment With Twin AttentionabstractPoint clouds are rapidly gaining popularity in many practical applications, and point cloud quality assessment (PCQA) is an important research topic that helps us measure and improve the visual experience in applications using point clouds. Research on full-reference (FR) PCQAs has recently made impressive progress, and research on no-reference (NR) PCQAs has also gradually increased. However, the performance of the prior NR PCQA methods still suffers from weak generalization ability and lower accuracy than the FR metrics in general. In this work, we propose a two-stage sampling method that can reasonably represent a whole point cloud, making it possible to efficiently calculate the point cloud quality. For quality prediction, we designed a twin-attention-based transformer PCQA model (3DTA), which uses the data of the two-stage sampling method as input and directly outputs the predicted quality score. Our model is accurate and widely applicable, and it has a simple and flexible structure. Experimental results show that in most cases, the proposed 3DTA model substantially outperforms the benchmark NR methods. The accuracy of the proposed method is competitive even against that of the FR method, which makes 3DTA a strong candidate for the PCQA task, regardless of the reference availability. The code of the proposed model is publicly available athttps://github.com/philox12358/3DTA-PCQA. Linxia Zhu, Xu Wang 0006, Honglei Su, Huan Yang 0001, Hui Yuan 0001, Jari Korhonen |
IEEE Trans. Multim. | 7 |
| 2023 | On the Explainable Detection of Stress Levels Using Heart Rate Variability Based Deep Neural NetworksabstractThis paper presents one of the first explorations of transparency and explainability of Heart Rate Variability (HRV) based deep learning models designed for stress detection. We employed Shapley additive explanations (SHAP) as an explainable AI (XAI) method, and cross-validated the results with saliency maps, which provides valuable insights into the main contributing factors for decision-making process of these deep models. Debasish Ghose, Jari Korhonen, Junyong You, Soumya P. Dash |
HealthCom | 3 |
| 2023 | Half of an Image is Enough for Quality AssessmentabstractDeep networks show promising performance in image quality assessment (IQA), whereas few studies have investigated how a deep model works. In this work, a positional masked transformer for IQA is first developed, based on which we observe that half of an image might contribute trivially to image quality, whereas the other half is crucial. Such observation is generalized to that half of the image regions can dominate image quality in several CNN-based IQA models. Motivated by this observation, three semantic measures (saliency, frequency, objectness) are then derived, showing high accordance with importance degree of image regions in IQA. Junyong You, Jari Korhonen |
ICIP | 3 |
| 2023 | Fast Accurate Fish Recognition with Deep Learning Based on a Domain-Specific Large-Scale Fish Dataset
Zhaoqi Chu, Jari Korhonen, Xiangrong Liu, Juan Liu 0003, Lvping Fang, Weidi Yang, Debasish Ghose, Junyong You |
MMM (1) | 3 |
| 2023 | No-Reference Point Cloud Quality Assessment via Weighted Patch Quality PredictionabstractWith the rapid development of 3D vision applications based on point clouds, point cloud quality assessment (PCQA) is becoming an important research topic.However, the prior PCQA methods ignore the effect of local quality variance across different areas of the point cloud.To take an advantage of the quality distribution imbalance, we propose a no-reference point cloud quality assessment (NR-PCQA) method with local area correlation analysis capability, denoted as COPP-Net.More specifically, we split a point cloud into patches, generate texture and structure features for each patch, and fuse them into patch features to predict patch quality.Then, we gather the features of all the patches of a point cloud for correlation analysis, to obtain the correlation weights.Finally, the predicted qualities and correlation weights for all the patches are used to derive the final quality score.Experimental results show that our method outperforms the state-of-the-art benchmark NR-PCQA methods.The source code for the proposed COPP-Net can be found at https://github.com/philox12358/COPP-Net. Honglei Su, Jari Korhonen |
SEKE | 3 |
| 2022 | Attention integrated hierarchical networks for no-reference image quality assessmentabstractQuality assessment of natural images is influenced by perceptual mechanisms, e.g., attention and contrast sensitivity, and quality perception can be generated in a hierarchical process. This paper proposes an architecture of Attention Integrated Hierarchical Image Quality networks (AIHIQnet) for no-reference quality assessment. AIHIQnet consists of three components: general backbone network, perceptually guided neck network, and head network. Multi-scale features extracted from the backbone network are fused to simulate image quality perception in a hierarchical manner. The attention and contrast sensitivity mechanisms modelled by an attention module capture essential information for quality perception. Considering that image rescaling potentially affects perceived quality, appropriate pooling methods in the non-convolution layers in AIHIQnet are employed to accept images with arbitrary resolutions. Comprehensive experiments on publicly available databases demonstrate outstanding performance of AIHIQnet compared to state-of-the-art models. Ablation experiments were performed to investigate the variants of the proposed architecture and reveal importance of individual components. Junyong You, Jari Korhonen |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Spatio-Temporal Difference Descriptor for Skeleton-Based Action RecognitionabstractIn skeletal representation, intra-frame differences between body joints, as well as inter-frame dynamics between body skeletons contain discriminative information for action recognition. Conventional methods for modeling human skeleton sequences generally depend on motion trajectory and body joint dependency information, thus lacking the ability to identify the inherent differences of human skeletons. In this paper, we propose a spatio-temporal difference descriptor based on a directional convolution architecture that enables us to learn the spatio-temporal differences and contextual dependencies between different body joints simultaneously. The overall model is built on a deep symmetric positive definite (SPD) metric learning architecture designed to learn discriminative manifold features with the well-designed non-linear mapping operation. Experiments on several action datasets show that our proposed method achieves up to 3% accuracy improvement over state-of-the-art methods. Chongyang Ding, Kai Liu 0021, Jari Korhonen, Eugeniy Belyaev |
AAAI | 3 |
| 2021 | Transformer For Image Quality AssessmentabstractTransformer has become the new standard method in natural language processing (NLP), and it also attracts research interests in computer vision area. In this paper we investigate the application of Transformer in Image Quality (TRIQ) assessment. Following the original Transformer encoder employed in Vision Transformer (ViT), we propose an architecture of using a shallow Transformer encoder on the top of a feature map extracted by convolution neural networks (CNN). Adaptive positional embedding is employed in the Transformer encoder to handle images with arbitrary resolutions. Different settings of Transformer architectures have been investigated on publicly available image quality databases. We have found that the proposed TRIQ architecture achieves outstanding performance. The implementation of TRIQ is published on Github (https://github.com/junyongyou/triq). Junyong You, Jari Korhonen |
ICIP | 2 |
| 2021 | Reproducibility Companion Paper: Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial FeaturesabstractBlind natural video quality assessment (BVQA), also known as no-reference video quality assessment, is a highly active research topic. In our recent contribution titled "Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial Features" published in ACM Multimedia 2020, we proposed a two-level video quality model employing statistical temporal features and spatial features extracted by a deep convolutional neural network (CNN) for this purpose. At the time of publishing, the proposed model (CNN-TLVQM) achieved state-of-the-art results in BVQA. In this paper, we describe the process of reproducing the published results by using CNN-TLVQM on two publicly available natural video quality datasets. Jari Korhonen, Yicheng Su, Junyong You, Steven Alexander Hicks, Cise Midoglu |
ACM Multimedia | 1 |
| 2020 | Blind Natural Image Quality Prediction Using Convolutional Neural Networks And Weighted Spatial PoolingabstractTypically, some regions of an image are more relevant for its perceived quality than the others. On the other hand, subjective image quality is also affected by low level characteristics, such as sensor noise and sharpness. This is why image rescaling, as often used in object recognition, is not a feasible approach for producing input images for convolutional neural networks (CNN) used for blind image quality prediction. Generally, convolution layer can accept images of arbitrary resolution as input, whereas fully connected (FC) layer only can accept a fixed length feature vector. To solve this problem, we propose weighted spatial pooling (WSP) to aggregate spatial information of any size of weight map, which can be used to replace global average pooling (GAP). In this paper, we present a blind image quality assessment (BIQA) method based on CNN and WSP. Our experimental results show that the prediction accuracy of the proposed method is competitive against the state-of-the-art image quality assessment methods. Yicheng Su, Jari Korhonen |
ICIP | 2 |
| 2020 | Attention Boosted Deep Networks For Video ClassificationabstractVideo classification can be performed by summarizing image contents of individual frames into one class by deep neural networks, e.g., CNN and LSTM. Human interpretation of video content is influenced by the attention mechanism. In other words, video class can be more attentively decided by certain information than others. In this paper, we propose to integrate the attention mechanism into deep networks for video classification. The proposed framework employs 2D CNN networks with ImageNet pretrained weights to extract features of video frames that are then fed to a bidirectional LSTM network for video classification. An attention block has been developed that can be added after the LSTM network in the proposed framework. Several different 2D CNN architectures have been tested in the experiments. The results with respect to two publicly available datasets have demonstrated that integrating attention can boost the performance of deep networks in video classification compared to not applying the attention block. We also found out that applying attention to the LSTM outputs on the VGG19 architecture provides the highest classification accuracy in the proposed framework. Junyong You, Jari Korhonen |
ICIP | 2 |
| 2020 | Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial FeaturesabstractDue to the wide range of different natural temporal and spatial distortions appearing in user generated video content, blind assessment of natural video quality is a challenging research problem. In this study, we combine the hand-crafted statistical temporal features used in a state-of-the-art video quality model and spatial features obtained from convolutional neural network trained for image quality assessment via transfer learning. Experimental results on two recently published natural video quality databases show that the proposed model can predict subjective video quality more accurately than the publicly available video quality models representing the state-of-the-art. The proposed model is also competitive in terms of computational complexity. Jari Korhonen, Yicheng Su, Junyong You |
ACM Multimedia | 1 |
| 2020 | Motion JPEG Decoding via Iterative Thresholding and Motion-Compensated DeflickeringabstractThis paper studies the problem of decoding video sequences compressed by Motion JPEG (M-JPEG) at the best possible perceived video quality. We consider decoding of M-JPEG video as signal recovery from incomplete measurements known in compressive sensing. We take all quantized nonzero Discrete Cosine Transform (DCT) coefficients as measurements and the remaining zero coefficients as data that should be recovered. The output video is reconstructed via iterative thresholding algorithm, where Video Block Matching and 4-D filtering (VBM4D) is used as thresholding operator. To reduce non-linearities in the measurements caused by the quantization in JPEG, we propose to apply spatio-temporal pre-filtering before measurements calculation and recovery. Since temporal inconsistencies of the residual coding artifacts lead to strong flickering in recovered video, we also propose to apply motion-compensated deflickering filter as a post-filter. Experimental results show that the proposed approach provides 0.44-0.51 dB average improvement in Peak Signal to Noise Ratio (PSNR), as well as lower flickering level compared to the state-of-the-art method based on Coefficient Graph Laplacians (COGL). We have also conducted a subjective comparison study, indicating that the proposed approach outperforms state-of-the-art methods in terms of subjective video quality. Eugeniy Belyaev, Linlin Bie, Jari Korhonen |
MMSP | 3 |
| 2019 | Assessing Personally Perceived Image Quality via Image Features and Collaborative FilteringabstractDuring the past few years, different methods for optimizing the camera settings and post-processing techniques to improve the subjective quality of consumer photos have been studied extensively. However, most of the research in the prior art has focused on finding the optimal method for an average user. Since there is large deviation in personal opinions and aesthetic standards, the next challenge is to find the settings and post-processing techniques that fit to the individual users' personal taste. In this study, we aim to predict the personally perceived image quality by combining classical image feature analysis and collaboration filtering approach known from the recommendation systems. The experimental results for the proposed method show promising results. As a practical application, our work can be used for personalizing the camera settings or post-processing parameters for different users and images. Jari Korhonen |
CVPR | 1 |
| 2019 | Deep Neural Networks for No-Reference Video Quality AssessmentabstractVideo quality assessment (VQA) is a challenging task due to the complexity of modeling perceived quality characteristics in both spatial and temporal domains. A novel no-reference (NR) video quality metric (VQM) is proposed in this paper based on two deep neural networks (NN), namely 3D convolution network (3D-CNN) and a recurrent NN composed of long short-term memory (LSTM) units. 3D-CNNs are utilized to extract local spatiotemporal features from small cubic clips in video, and the features are then fed into the LSTM networks to predict the perceived video quality. Such design can elaborately tackle the issue of insufficient training data whilst also efficiently capture perceptive quality features in both spatial and temporal domains. Experimental results with respect to two publicly available video quality datasets have demonstrate that the proposed quality metric outperforms the other compared NR quality metrics. Junyong You, Jari Korhonen |
ICIP | 2 |
| 2019 | Optimizing the Parameters for Post-Processing Consumer Photos via Machine LearningabstractPhoto sharing in social media is a part of everyday life for many, as inexpensive cameras integrated in smartphones are widely available. Unfortunately, low cost consumer devices are often prone to capture artifacts, and this is why there is a growing demand for automatic post-processing to enhance the image quality. Due to the wide range of distortions in non-professional photography, automatic selection of the post-processing methods and parameters is a challenging problem. In this paper, we present a subjective study based on rank-ordering method, comparing the subjective preferences between photos processed with different parameters for image sharpening and denoising. The subjective results are used as a basis to derive the ground truth values for the post-processing parameters for different photos. Then, we apply a pre-trained convolutional neural network (CNN) to extract a set of features from photos, used as input to a regression model to predict the optimal post-processing parameters. Test results show that the learning-based approach can predict post-processing parameters with a satisfactory accuracy. Linlin Bie, Xu Wang 0006, Jari Korhonen |
ICTAI | 3 |
| 2019 | Two-Level Approach for No-Reference Consumer Video Quality AssessmentabstractSmartphones and other consumer devices capable of capturing video content and sharing it on social media in nearly real time are widely available at a reasonable cost. Thus, there is a growing need for no-reference video quality assessment (NR-VQA) of consumer produced video content, typically characterized by capture impairments that are qualitatively different from those observed in professionally produced video content. To date, most of the NR-VQA models in prior art have been developed for assessing coding and transmission distortions, rather than capture impairments. In addition, the most accurate NR-VQA methods known in prior art are often computationally complex, and therefore impractical for many real life applications. In this paper, we propose a new approach for learning-based video quality assessment, based on the idea of computing features in two levels so that low complexity features are computed for the full sequence first, and then high complexity features are extracted from a subset of representative video frames, selected by using the low complexity features. We have compared the proposed method against several relevant benchmark methods using three recently published annotated public video quality databases, and our results show that the proposed method can predict subjective video quality more accurately than the benchmark methods. The best performing prior method achieves nearly similar accuracy, but at substantially higher computational cost. Jari Korhonen |
IEEE Trans. Image Process. | 1 |
| 2018 | Subjective Assessment of Post-Processing Methods for Low Light Consumer PhotosabstractConsumer photos taken in low light conditions often suffer from substantial undesired capture artifacts, such as shakiness and sensor noise. In this paper, we use rank ordering method to assess the subjective preferences among different postprocessing methods used to alleviate capture artifacts. The results show that most users prefer sharpened photos, even in the presence of substantial sensor noise. However, there are also systematic differences in individual preferences between users. Therefore, user preferences need to be considered in addition to the image characteristics, when selecting the post-processing algorithms and parameters for photo quality enhancement. Linlin Bie, Xu Wang 0006, Jari Korhonen |
QoMEX | 3 |
| 2018 | Learning-based Prediction of Packet Loss Artifact Visibility in Networked VideoabstractThis In this paper, we study the problem of detecting packet loss distortion and estimating the perceived visibility of such distortion in decoded video. Our analysis is based on the features of the decoded video signal, and we assume that no information about actual packet losses is available from the underlying network or video decoder. First, we present a full-reference method for assessing packet loss visibility at the macroblock, frame and sequence levels. Second, we propose a no-reference method for detecting defected frames, based on spatiotemporal features and machine learning. Experimental results show that the proposed no-reference method achieves a high correlation with the full-reference method at both sequence and frame level. At sequence level, the no-reference method can also predict the subjective quality ratings at high accuracy. Jari Korhonen |
QoMEX | 1 |
| 2017 | Predicting personal preferences in subjective video quality assessmentabstractIn this paper, we study the problem of predicting the visual quality of a specific test sample (e.g. a video clip) experienced by a specific user, based on the ratings by other users for the same sample and the same user for other samples. A simple linear model and algorithm is presented, where the characteristics of each test sample are represented by a set of parameters, and the individual preferences are represented by weights for the parameters. According to the validation experiment performed on public visual quality databases annotated with raw individual scores, the proposed model can predict the scores by individuals more accurately than the average score for the respective sample computed from the scores given by other individuals. In many cases, the proposed algorithm also outperforms more generic Parametric Probabilistic Matrix Factorization (PPMF) technique developed for collaborative filtering in recommendation systems. Jari Korhonen |
QoMEX | 1 |
| 2016 | Modeling the Quality of Videos Displayed With Local Dimming Backlight at Different Peak White and Ambient Light LevelsabstractThis paper investigates the impact of ambient light and peak white (maximum brightness of a display) on the perceived quality of videos displayed using local backlight dimming. Two subjective tests providing quality evaluations are presented and analyzed. The analyses of variance show significant interactions of the factors peak white and ambient light with the perceived quality. Therefore, we proceed to predict the subjective quality grades with objective measures. The rendering of the frames on liquid crystal displays with light emitting diodes backlight at various ambient light and peak white levels is computed using a model of the display. Widely used objective quality metrics are applied based on the rendering models of the videos to predict the subjective evaluations. As these predictions are not satisfying, three machine learning methods are applied: partial least square regression, elastic net, and support vector regression. The elastic net method obtains the best prediction accuracy with a spearman rank order correlation coefficient of 0.71, and two features are identified as having a major influence on the visual quality. Claire Mantel, Jacob Søgaard, Soren Bech, Jari Korhonen, Jesper Melgaard Pedersen, Søren Forchhammer |
IEEE Trans. Image Process. | 4 |
| 2015 | Improving image fidelity by luma-assisted chroma subsamplingabstractChroma subsampling is commonly used for digital representations of images and video sequences. The basic rationale behind chroma subsampling is that the human visual system is less sensitive to color variations than luma variations. Therefore, chroma data can be coded in lower resolution than luma data, without noticeable loss in perceived image quality. In this paper, we compare different upsampling methods for chroma data and show that by using advanced upsampling schemes the fidelity of the reconstructed image can be significantly improved. We also present an adaptive upsampling method that uses full resolution luma information to assist chroma upsampling. Experimental results show that in the presence of compression noise, the proposed technique steadily outperforms advanced non-assisted upsampling. Jari Korhonen |
ICME | 1 |
| 2015 | No-Reference Video Quality Assessment Using Codec AnalysisabstractA no-reference (NR) video quality assessment (VQA) method is presented for videos distorted by H.264/Advanced Video Coding (AVC) and MPEG-2. The assessment is performed without access to the bitstream. Instead, we analyze and estimate coefficients based on decoded pixels. The approach involves distinguishing between the two types of videos, estimating the level of quantization used in the I-frames, and exploiting this information to assess the video quality. To do this for H.264/AVC, the distribution of the discrete cosine transform-coefficients after intra-prediction and deblocking are modeled. To obtain VQA features for H.264/AVC, we propose a novel estimation method of the quantization in H.264/AVC videos without bitstream access, which can also be used for peak signal-to-noise ratio estimation. The results from the MPEG-2 and H.264/AVC analysis are mapped to a perceptual measure of video quality by support vector regression. For validation purposes, the proposed method was tested on two databases. In both cases, a good performance compared with state of the art full, reduced, and NR VQA algorithms was achieved. Jacob Søgaard, Søren Forchhammer, Jari Korhonen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Modeling the Subjective Quality of Highly Contrasted Videos Displayed on LCD With Local Backlight DimmingabstractLocal backlight dimming is a technology aiming at both saving energy and improving visual quality on television sets. As the rendition of the image is specified locally, the numerical signal corresponding to the displayed image needs to be computed through a model of the display. This simulated signal can then be used as input to objective quality metrics. The focus of this paper is on determining which characteristics of locally backlit displays influence quality assessment. A subjective experiment assessing the quality of highly contrasted videos displayed with various local backlight-dimming algorithms is set up. Subjective results are then compared with both objective measures and objective quality metrics using different display models. The first analysis indicates that the most significant objective features are temporal variations, power consumption (probably representing leakage), and a contrast measure. The second analysis shows that modeling of leakage is necessary for objective quality assessment of sequences displayed with local backlight dimming. Claire Mantel, Soren Bech, Jari Korhonen, Søren Forchhammer, Jesper Melgaard Pedersen |
IEEE Trans. Image Process. | 3 |
| 2013 | No-Reference Video Quality Assessment using MPEG analysisabstractWe present a method for No-Reference (NR) Video Quality Assessment (VQA) for decoded video without access to the bitstream. This is achieved by extracting and pooling features from a NR image quality assessment method used frame by frame. We also present methods to identify the video coding and estimate the video coding parameters for MPEG-2 and H.264/AVC which can be used to improve the VQA. The analysis differs from most other video coding analysis methods since it is without access to the bitstream. The results show that our proposed method is competitive with other recent NR VQA methods for MPEG-2 and H.264/AVC. Jacob Søgaard, Søren Forchhammer, Jari Korhonen |
PCS | 3 |
| 2013 | Modeling the color image and video quality on liquid crystal displays with backlight dimmingabstractObjective image and video quality metrics focus mostly on the digital representation of the signal. However, the display characteristics are also essential for the overall Quality of Experience (QoE). In this paper, we use a model of a backlight dimming system for Liquid Crystal Display (LCD) and show how the modeled image can be used as an input to quality assessment algorithms. For quality assessment, we propose an image quality metric, based on Peak Signal-to-Noise Ratio (PSNR) computation in the CIE L*a*b* color space. The metric takes luminance reduction, color distortion and loss of uniformity in the resulting image in consideration. Subjective evaluations of images generated using different backlight dimming algorithms and clipping strategies show that the proposed metric estimates the perceived image quality more accurately than conventional PSNR. Jari Korhonen, Claire Mantel, Nino Burini, Søren Forchhammer |
VCIP | 1 |
| 2013 | Frame rate versus spatial quality: Which video characteristics do matter?abstractSeveral studies have shown that the relationship between perceived video quality and frame rate is dependent on the video content. In this paper, we have analyzed the content characteristics and compared them against the subjective results derived from preference decisions between different spatial and temporal quality levels. We also propose simple yet powerful metrics for characterizing spatial and temporal properties of a video sequence, and demonstrate how these metrics can be applied for evaluating the relative impact of spatial and temporal quality on the perceived overall quality. Jari Korhonen, Ulrich Reiter, Anna Ukhanova |
VCIP | 1 |
| 2013 | Enhancing perceived quality of compressed images and video with anisotropic diffusion and fuzzy filtering
Ehsan Nadernejad, Jari Korhonen, Søren Forchhammer, Nino Burini |
Signal Process. Image Commun. | 2 |
| 2012 | Image dependent energy-constrained local backlight dimmingabstractIn this work, we consider and propose two extensions to an optimization-based image dependent backlight dimming algorithm. The first extension introduces error weighting based on human perception of luminance, aiming to improve the perceived image quality; the second extension adds an adjustable term for power consumption to the cost function, allowing flexible power management. Experimental results show that the proposed solution can achieve better results than other algorithms at several power consumption levels. Nino Burini, Ehsan Nadernejad, Jari Korhonen, Søren Forchhammer, Xiaolin Wu 0001 |
ICIP | 3 |
| 2012 | Objective assessment of the impact of frame rate on video qualityabstractIn this paper, we present a novel objective quality metric that takes the impact of frame rate into account. The proposed metric uses PSNR, frame rate and a content dependent parameter that can easily be obtained from spatial and temporal activity indices. The results have been validated on data from a subjective quality study, where the test subjects have been choosing the preferred path from the lowest quality to the best quality, at each step making a choice in favor of higher frame rate or lower distortion. A comparison with other relevant objective metrics shows that the proposed metric on average provides a more precise correlation with the subjective results. Anna Ukhanova, Jari Korhonen, Søren Forchhammer |
ICIP | 2 |
| 2011 | Audiovisual quality fusion based on relative multimodal complexityabstractIn multimodal presentations the perceived audiovisual quality assessment is significantly influenced by the content of both the audio and visual tracks. Based on our earlier subjective quality test for finding the optimal trade-off between audio and video quality, this paper proposes a novel method for relative multimodal complexity analysis to derive the fusion parameter in objective audiovisual quality metrics. Audio and video qualities are first estimated separately using advanced quality models, and then they are combined into the overall audiovisual quality using a linear fusion. Based on carefully designed auditory and visual features, the relative complexity analysis model across sensory modalities is proposed for deriving the fusion parameter. Experimental results have demonstrated that the content adaptive fusion parameter can improve the prediction accuracy of objective audiovisual quality metrics, compared to the fusion parameters obtained from the subjective quality tests using other known optimization methods. Junyong You, Jari Korhonen, Ulrich Reiter |
ICIP | 2 |
| 2011 | Congestion control in wireless links based on selective delivery of erroneous packets
Jari Korhonen, Andrew Perkis, Ulrich Reiter |
Signal Process. Image Commun. | 1 |
| 2011 | Balancing Attended and Global Stimuli in Perceived Video Quality AssessmentabstractThe visual attention mechanism plays a key role in the human perception system and it has a significant impact on our assessment of perceived video quality. In spite of receiving less attention from the viewers, unattended stimuli can still contribute to the understanding of the visual content. This paper proposes a quality model based on the late attention selection theory, assuming that the video quality is perceived via two mechanisms: global and local quality assessment. First we model several visual features influencing the visual attention in quality assessment scenarios to derive an attention map using appropriate fusion techniques. The global quality assessment as based on the assumption that viewers allocate their attention equally to the entire visual scene, is modeled by four carefully designed quality features. By employing these same quality features, the local quality model tuned by the attention map considers the degradations on the significantly attended stimuli. To generate the overall video quality score, global and local quality features are combined by a content adaptive linear fusion method and pooled over time, taking the temporal quality variation into consideration. The experimental results have been compared to results from appropriate eye tracking and video quality assessment experiments, demonstrating promising performance. Junyong You, Jari Korhonen, Andrew Perkis, Touradj Ebrahimi |
IEEE Trans. Multim. | 2 |
| 2010 | Spatial and temporal pooling of image quality metrics for perceptual video quality assessment on packet loss streamsabstractVideo streaming through bandwidth-limited channels often suffer from packet losses. Therefore, perceptual quality assessment on video sequences with packet losses is a critical issue in digital video communications. This paper analyzes several image quality metrics and evaluates their applications using spatial and temporal pooling schemes in perceptual video quality assessment for video streams with packet losses. Several approaches using Minkowski summation and averages over different distorted spatial regions and temporal frames to pool the spatial and temporal qualities are evaluated. The experimental results with respect to the subjective video quality measurements demonstrate that the subjects are more sensitive to the most annoying spatial regions and temporal segments when assessing the video quality of the lossy streams. Junyong You, Jari Korhonen, Andrew Perkis |
ICASSP | 2 |
| 2010 | On the relationship between perceptual impact of source and channel distortions in video sequencesabstractIt is known that peak signal-to-noise ratio (PSNR) can be used for assessing the relative qualities of distorted video sequences meaningfully only if the compared sequences contain similar types of distortions. In this paper, we propose a model for rough assessment of the bias in PSNR results, when video sequences with both channel and source distortion are compared against video sequences with source distortion only. The proposed method can be used to compare the relative perceptual quality levels of video sequences with different distortion types more reliably than using plain PSNR. Jari Korhonen, Ulrich Reiter, Junyong You |
ICIP | 1 |
| 2010 | Attention modeling for video quality assessment: Balancing global quality and local qualityabstractThis paper proposes to evaluate video quality by balancing two quality components: global quality and local quality. The global quality is a result from subjects allocating their attention equally to all regions in a frame and all frames in a video. It is evaluated by image quality metrics (IQM) with averaged spatiotemporal pooling. The local quality is derived from visual attention modeling and quality variations over frames. Saliency, motion, and contrast information are taken into account in modeling visual attention, which is then integrated into IQMs to calculate the local quality of a video frame. The local quality of a video sequence is calculated by pooling local quality values over all frames with a temporal pooling scheme derived from the known relationship between perceived video quality and the frequency of temporal quality variations. The overall quality of a distorted video is a weighted average between the global quality and the local quality. Experimental results demonstrate that the combination of the global quality and local quality outperforms both sole global quality and local quality, as well as other quality models, in video quality assessment. In addition, the proposed video quality modeling algorithm can improve the performance of image quality metrics on video quality assessment compared to the normal averaged spatiotemporal pooling scheme. Junyong You, Jari Korhonen, Andrew Perkis |
ICME | 2 |
| 2009 | Loss-distortion estimation for practical H.264/AVC streamsabstractEven though several unequal erasure protection (UEP) schemes have been proposed for video streaming, it is still a challenge to define the optimal relative protection level for different data units. In this paper, we study loss-distortion modeling in realistic video streaming scenarios. The main observation is that when video streams with complex hierarchical structures and error resilience features are considered, the relative significance of different units cannot be reliably estimated from mutual dependencies between units as suggested in related research. Therefore, data units can be classified reliably into different perceptual priority classes via analysis by synthesis only. For estimating the co-impact of multiple losses, we propose a simple loss-distortion model. The model can be used for optimizing UEP in practical streaming systems. Jari Korhonen, Andrew Perkis |
ICME | 1 |
| 2009 | Comparison of unequal erasure protection schemes for video streamingabstractSeveral different schemes for unequal error/erasure protection (UEP) in video streaming have been proposed during the past years. However, it is not a trivial task to define the optimal relative protection levels for different media units (ie. the basic units of decoding handled individually by media decoder) in practical streaming applications. In this paper, we use a theoretical loss-distortion model to evaluate the performance of different UEP schemes in realistic video streaming scenarios with fixed redundancy overhead budget. Our results confirm the observation that equal error/erasure protection (EEP) is desired when the packet loss rate is low compared to the redundancy overhead. The higher the packet loss rate, the more segregating UEP should be used to achieve optimal performance. The results also suggest that comparable performances can be obtained by using two protection levels only (protected part and unprotected part) instead of more complex UEP schemes. Jari Korhonen, Andrew Perkis |
MMSP | 1 |
| 2009 | Battery life of mobile peers with UMTS and WLAN in a Kademlia-based P2P overlayabstractWe evaluate the battery life of mobile devices that act as full-fledged peer nodes in a Kademlia DHT based P2P overlay network. The motivation is to find out how long a mobile peer is able to function in a UMTS or WLAN access network, and how the different parameter settings affect this battery life; this is interesting as mobile access to P2P networks is expected to become common in the near future. The majority of the peers in an overlay are simulated on a server array, while the power measurements are conducted on actual mobile devices. The variable overlay parameters are the number of peers, resource lookup activity, and the level of churn. The chosen values of parameters represent a relatively high amount of activity. In UMTS the measured battery life is approximately 3 hours and in WLAN it is 5 to 10 (most often around 8) hours. We also provide power measurements on sending and receiving UDP packets in UMTS and WLAN, for approximating the power consumption of network protocols without protocol-specific measurements. Otso Kassinen, Erkki Harjula, Jari Korhonen, Mika Ylianttila |
PIMRC | 3 |
| 2009 | Flexible forward error correction codes with application to partial media data recovery
Jari Korhonen, Pascal Frossard |
Signal Process. Image Commun. | 1 |
| 2008 | Sparse FEC codes for flexible media protectionabstractIn this paper, we study block codes that are optimized to recover some lost source data even in case when full recovery is not possible. Conventionally, block codes designed for packet erasure networks are aimed to recover all the lost source packets, assuming that the amount of lost data does not exceed the redundancy overhead. Unfortunately, this approach leads to poor performance if the fraction of lost data even occasionally exceeds the limit for full recovery capability. Recovery of part of the data may prove to be beneficial, especially when media data packets are unequal in importance. We present a short linear block code design that improves the performance of traditional minimum distance separable (MDS) codes by reducing the fluctuation of the residual packet loss rate. These new codes also lead to a flexible design for unequal error protection of the media packets. Jari Korhonen, Pascal Frossard |
ICME | 1 |
| 2006 | Unacceptability of instantaneous errors in mobile television: from annoying audio to videoabstractAs in many digital telecommunications systems, the received data streams over Digital Video Broadcasting for Handhelds (DVB-H) may contain bursty transmission errors. The bursty error characteristics affect the end users' perceived audiovisual quality. This study examined the perceived unacceptability of instantaneous but noticeable audio, visual and audiovisual errors. The erroneous streams were generated from four popular television contents by applying three simulated error patterns with different error rates (1.7%, 6.9%, 13.8%) and error burst durations. Instantaneous unacceptability of errors was evaluated by 30 participants with simplified continuous assessment while watching the program content. The results show that with the two lowest error rates the audio errors were more unacceptable than video errors and with the highest error rate the visual and audiovisual errors become the most unacceptable. Satu Jumisko-Pyykkö, Vinod Kumar Malamal Vadakital, Jari Korhonen |
Mobile HCI | 3 |
| 2006 | Generic forward error correction of short frames for IP streaming applications
Jari Korhonen, Ye Wang 0007 |
Multim. Tools Appl. | 1 |
| 2005 | Optimization of source and channel coding for voice over IPabstractVoice over Internet protocol (VoIP) applications must typically choose a tradeoff between the bits allocated for forward error correcting (FEC) and that for the source coding to achieve the best speech quality at a given packet loss rate. In this paper, we present a new scheme to optimize the speech quality subject to the bandwidth constraints and the packet loss rate. The scheme adopts adaptive multi-rate (AMR) speech codec along with a FEC scheme based on exclusive OR (XOR) operations. Retransmission is also taken into account if the round trip time (RTT) is within a certain limit. We use a simplified E-model as objective metric. Subjective listening tests show that our scheme improves the perceptual speech quality significantly compared to the non-adaptive baseline speech transmission system. Jari Korhonen, Ye Wang 0007 |
ICME | 2 |
| 2005 | Power-efficient streaming for mobile terminalsabstractWireless Network Interface (WNI) is one of the most critical components for power efficiency in multimedia streaming to mobile devices. A common strategy to save power is to switch WNI to active mode only when network activity is expected. In streaming systems, this approach is problematic because data are typically received continuously. One solution is to transmit data packets as bursts, which leaves WNI more time between bursts in standby mode. However, that subjects bursty transmission in high peak rates, which leaves it prone to congestion. In this paper, we study theoretically and empirically the impact of burst length and peak transmission rate for observed packet loss and delay characteristics as well as potential energy savings in a Wireless Local Area Network (WLAN) environment. We outline and implement a test system with adaptive burst length to achieve improved trade-off between power efficiency and congestion tolerance. Jari Korhonen, Ye Wang 0007 |
NOSSDAV | 1 |
| 2005 | Effect of packet size on loss rate and delay in wireless linksabstractTransmitting large packets over wireless networks helps to reduce header overhead, but may have an adverse effect on loss rate due to corruptions in a radio link. Packet loss in lower layers, however, is typically hidden from the upper protocol layers by link or MAC layer protocols. For this reason, errors in the physical layer are observed by the application as higher variance in end-to-end delay rather than increased packet loss rate. We study the effect of packet size on loss rate and delay characteristics in a wireless real-time application. We derive an analytical model for the dependency between packet length and delay characteristics. We validate our theoretical analysis through experiments in an ad hoc network using WLAN technologies. We show that careful design of packetization schemes in the application layer may significantly improve radio link resource utilization in delay sensitive media streaming under difficult wireless network conditions. Jari Korhonen, Ye Wang 0007 |
WCNC | 1 |
| 2005 | Toward bandwidth-efficient and error-robust audio streaming over lossy packet networks
Jari Korhonen, Ye Wang 0007, David Isherwood |
Multim. Syst. | 1 |
| 2004 | A framework for robust and scalable audio streamingabstractWe propose a framework to achieve bandwidth efficient, error robust and bitrate scalable audio streaming. Our approach is compatible with most audio compression format. The main contributions of this paper include: 1) the proposal of a Multi-Stage Interleaving (MSI) strategy which translates packet loss into loss of separate frequency components that are less perceptually significant; and 2) the design of a Layered Unequal-Sized Packetization (LUSP) scheme which enables bitrate scalability and prioritized packet transmission. The combination of the proposed MSI and LUSP allows the use of a set of simple yet effective methods of error concealment in the compressed domain. Our approach offers significant advantages over existing methods in terms of memory consumption (a savings of over 40 times in the sample MP3 implementation), and computational complexity, which are critical issues for battery-powered small devices. Ye Wang 0007, Wendong Huang, Jari Korhonen |
ACM Multimedia | 3 |
| 2003 | Schemes for error resilient streaming of perceptually coded audioabstractThis paper presents novel extensions to our earlier system for streaming perceptually coded audio over error prone channels such as Mobile IP. To improve error robustness while maintaining bandwidth efficiency, the new extensions combine the strength of an error resilient coding scheme in the sender, a prioritized packet transport scheme in the network and a compressed domain error concealment strategy in the terminal. Different concealment methods are used for each part of the coded audio data according to their perceptual importance and statistical characteristics. In our current implementation, we employed MPEG-2 Advanced Audio Coding (AAC) encoded bitstreams and an RTP/UDP-based test system for performance evaluation. Simulation results have shown that our improved streaming system is more robust against packet losses in comparison with conventional methods. Jari Korhonen, Ye Wang 0007 |
ICASSP (5) | 1 |
| 2003 | Schemes for error resilient streaming of perceptually coded audioabstractThis paper presents novel extensions to our earlier system for streaming perceptually coded audio over error prone channels such as mobile IP. To improve error robustness while maintaining bandwidth efficiency, the new extensions combine the strength of an error resilient coding scheme in the sender, prioritized packet transport scheme in the network and a compressed domain error concealment strategy in the terminal. Different concealment methods are used for each part of the coded audio data according to their perceptual importance and statistical characteristics. In our current implementation, we employed MPEG-2 advanced audio coding (AAC) encoded bitstreams and an RTP/UDP-based test system for performance evaluation. Simulation results have shown that our improved streaming system is more robust against packet losses in comparison with conventional methods. Jari Korhonen, Ye Wang 0007 |
ICME | 1 |
| 2002 | Error robustness scheme for perceptually coded audio based on interframe shuffling of samplesabstractHigh-quality audio streaming over IP networks is a significant part of multimedia traffic in telecommunications networks in the future. Although there are a number of error concealment and correction methods developed for real-time multimedia streaming over unreliable packet-switched networks, many effective codec dependent error robustness schemes have not been utilized for the state-of-art audio compression standards developed mainly to compress audio for storage media. This paper describes how the idea of redistributing adjacent audio samples to different packets can be applied to the perceptual audio codecs, such as MP3 and AAC. The experiences with testing the concept using AAC are explained as well. The proposed approach is especially applicable with semi-reliable transport protocols and future networks providing flexible support for prioritized data traffic. Jari Korhonen |
ICASSP | 1 |