Marius Pedersen

dblp:18/7169 · DBLP profile ↗
← Back
38ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0001-9797-5821ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Few-Shot Supervised Contrastive Learning for Image/Video Distortion Classification
Riestiya Zain Fadillah, Seyed Ali Amirshahi, Marius Pedersen, Azeddine Beghdadi
ICPR (7)3
2026 Loss function assessment: Towards learning hard-to-learn classes
abstract
Deep learning models often struggle with hard-to-learn classes, which arise not only from class imbalance or long-tailed distributions but also from intrinsic complexity, particularly in medical applications. Existing loss functions typically target long-tailed and imbalanced datasets, yet they often fail to address the challenges posed by classes that remain difficult to learn beyond frequency effects. Moreover, their expected impact is difficult to assess prior to costly model training. We propose a Loss Function Assessment (LFA) framework that evaluates loss functions, quantifying how their gradients contribute to learning between positive and negative labels when one is predicted more poorly, providing a diagnostic of learning dynamics. As a use case, we derive the Focal balanced Exponential Cross Entropy (F-ECE) loss from LFA analysis, combining exponential weighting with focal balancing to illustrate how the framework can inform new loss designs. We validate LFA on widely used losses, including Binary Cross Entropy, Focal Loss, Asymmetric Loss, and F-ECE, across CIFAR10-LT and ImageNet-LT (long-tailed distributions), CelebA (attribute-specific hard classes) and CAD-CAP (medical dataset). F-ECE achieves up to 1.95% recall improvement on CelebA, with consistent gains in F1-score and balanced accuracy across all datasets. Code and protocols are available at: https://github.com/Bozhao-Liu/LFA-demo .
Bozhao Liu, Marius Pedersen, Kiran B. Raja
Neurocomputing2
2026 Predicting image quality score distribution using deep neural networks
abstract
Blind objective image quality assessment methods typically predict a single Mean Opinion Score (MOS) to represent perceived image quality. However, MOS discards valuable information about observer variability, as different score distributions can produce the same mean value. Predicting the full quality score distribution provides a richer and more realistic representation of perceptual quality. In this work, we propose an Image Quality Distribution Network (IQDN) in five configurations: a custom convolutional neural network trained from scratch and four transfer learning variants based on ResNet50, VGG16, Xception, and DenseNet121 backbones. The models were trained to predict five-bin normalized quality score histograms on three datasets: KonIQ-10k, CID2013, and a newly introduced NAP540 dataset. Performance was evaluated using multiple point-wise error and distribution similarity metrics. Results show that the Image Quality Distribution Network with Xception backbone consistently performs better than existing state-of-the-art approaches in terms of Earth Mover’s Distance, capturing score distributions more effectively. Furthermore, a hybrid loss function combining Kullback–Leibler (KL) divergence and Huber loss improved performance across datasets. The predicted distributions also reconstruct MOS that closely correlate with ground-truth MOS. Finally, incorporating a quality-aware backbone demonstrates strong potential, particularly in cross-dataset testing scenarios. • Introduce the Image Quality Distribution Network (IQDN), a blind image quality assessment framework designed to predict full quality score distributions rather than single Mean Opinion Score (MOS) values. • Present a systematic evaluation of IQDN architectures on multiple image quality datasets, using a diverse set of distributional and point-wise error metrics. • Analyze the impact of different architectural configurations on the prediction performance, as well as cross-dataset generalization. • The proposed networks showed very good performance, both in predicting the distribution shape and in the Mean Opinion Score (MOS).
Nikola Plavac, Seyed Ali Amirshahi, Marius Pedersen, Sophie Triantaphillidou
J. Vis. Commun. Image Represent.3
2025 Chasing Shadows: Solving Deepfake Detection Benchmarks Using Irrelevant Features Only
abstract
The emergence of Deepfake technology poses significant threats, particularly regarding misinformation and privacy. To mitigate these threats, Deepfake benchmarks play an important role in developing and testing reliable Deepfake detection algorithms. Consequently, it is crucial that these benchmarks do not possess serious biases that hinder the robustness and generalizability of Deepfake detectors during training and subsequently distort their true reliability during operation. This work investigates inherent biases in various Deepfake detection benchmark datasets by training simple classification models based on soft-biometric facial properties that do not contain Deepfake-related clues, i.e., decoy features. These mirage models reach up to 87.42% (balanced) accuracy on benchmark datasets using irrelevant decoy features alone for this task. As large parts of the performance of state-of-the-art models could also be achieved through exploiting benchmark biases, this raises the question of the unbiased performance of Deepfake detectors and their general reliability. Our analysis includes various Deepfake detection benchmarks and analyzes soft-biometric properties in determining their contribution to “solving” these benchmarks. Our findings underscore the need for more unbiased benchmarks beyond simply balancing demographic groups to enable future work on developing reliable solutions.
Philipp Terhörst, Marius Pedersen, Kiran B. Raja
FG3
2025 Uncertainty Quantification in Video Distortion Classification Under Dataset Shift
Riestiya Zain Fadillah, Seyed Ali Amirshahi, Marius Pedersen, Azeddine Beghdadi
ICANN (2)3
2025 Improving Pseudo-Labels Selection Using Domain Priors for Semi-Supervised Detection in Capsule Endoscopy
abstract
The unavailability of expert annotations for learning deep models for pathology detection in Wireless Capsule Endoscopy (WCE) has led to an increase in explorations of semi-supervised learning. Semi-supervised models reduce the dependency on large-scale annotations. However, the resulting quality of representations relies on the quality of pseudo-labels generated from unlabeled data. In this work, we develop domain-specific augmentations in conjunction with weighted box fusion and active sampling to better select the unlabeled samples and reduce annotation cost in WCE. We demonstrate the improvement in performance using the proposed idea on two different datasets including SEE-AI and Kvasir-Capsule, with multiple standard metrics like average precision and its variants. The experimental results reveal that domain-specific augmentations with active sampling can help in selecting the most informative unlabeled samples, making it possible to improve semi-supervised models. With an annotation cost of only 35% of the actively selected unlabeled data, our method performs better than the state-of-the-art models on the same datasets for both fully supervised Faster-RCNN and Semi-supervised Unbiased Teacher and Active Teacher. We achieved a gain of +3.09% in AP50 on SEE-AI and +3.37% in AP75 on Kvasir-Capsule over the Active Teacher model. Our code is available at https://github.com/agossouema2011/SSOD_With_DTA_WCE.
Bidossessi Emmanuel Agossou, Marius Pedersen, Anuja Vats, Kiran B. Raja
ICIP2
2025 Uncertainty-Aware Regularization for Image-to-Image Translation
abstract
The importance of quantifying uncertainty in deep networks has become paramount for reliable real-world applications. In this paper, we propose a method to improve uncertainty estimation in medical Image-to-Image (I2I) translation. Our model integrates aleatoric uncertainty and employs Uncertainty-Aware Regularization (UAR) inspired by simple priors to refine uncertainty estimates and enhance reconstruction quality. We show that by leveraging simple priors on parameters, our approach captures more robust uncertainty maps, effectively refining them to indicate precisely where the network encounters difficulties, while being less affected by noise. Our experiments demonstrate that UAR not only improves translation performance, but also provides better uncertainty estimations, particularly in the presence of noise and artifacts. We validate our approach using two medical imaging datasets, showcasing its effectiveness in maintaining high confidence in familiar regions while accurately identifying areas of uncertainty in novel/ambiguous scenarios.
Anuja Vats, Ivar Farup, Marius Pedersen, Kiran B. Raja
WACV3
2024 CAPTIV8: A Comprehensive Large Scale Capsule Endoscopy Dataset For Integrated Diagnosis
abstract
Limited access to high-quality medical data poses a significant obstacle to automated diagnoses in medical modalities like Wireless Capsule Endoscopy (WCE), hindering potential advancements in automated medical diagnoses. This study presents a meticulously curated WCE dataset CAPTIV8, focused on the large colon and its pathologies, including Ulcerative Colitis (UC). Comprising a total of 1352 short video segments, totaling more than 200,000 frames with high mucosal visibility, the dataset features eight distinct types of pathology, along with signs of UC, accompanied by clinician-assigned text descriptions. To enhance its medical utility, the dataset integrates overlapping diagnoses from three diagnostic modalities: traditional and capsule endoscopy, and histology. Key attributes such as cleansing scores, text reports, capsule camera calibration and localization data have been incorporated to broaden its applicability in medical and artificial intelligence research. Designed for a wide spectrum of research challenges, from basic classification tasks to 3D reconstruction, CAPTIV8 aims to advance the incorporation of automated solutions in WCE diagnosis. The dataset can be accessed here:https:// dataverse.no/dataset.xhtml?persistentId=doi:10.18710/BSXNA1.
Anuja Vats, Pål Anders Floor, Ahmed Kedir Mohammed, Marius Pedersen, Oistein Hovde
ICIP5
2024 SSP-Net: A Siamese-Based Structure-Preserving Generative Adversarial Network for Unpaired Medical Image Enhancement
abstract
Recently, unpaired medical image enhancement is one of the important topics in medical research. Although deep learning-based methods have achieved remarkable success in medical image enhancement, such methods face the challenge of low-quality training sets and the lack of a large amount of data for paired training data. In this article, a dual input mechanism image enhancement method based on Siamese structure (SSP-Net) is proposed, which takes into account the structure of target highlight (texture enhancement) and background balance (consistent background contrast) from unpaired low-quality and high-quality medical images. Furthermore, the proposed method introduces the mechanism of the generative adversarial network to achieve structure-preserving enhancement by jointly iterating adversarial learning. Experiments comprehensively illustrate the performance in unpaired image enhancement of the proposed SSP-Net compared with other state-of-the-art techniques.
Guoxia Xu, Hao Wang 0003, Marius Pedersen, Meng Zhao 0001, Hu Zhu
IEEE Trans. Comput. Biol. Bioinform.3
2024 Terrain-Informed Self-Supervised Learning: Enhancing Building Footprint Extraction From LiDAR Data With Limited Annotations
abstract
Estimating building footprint maps from geospatial data is vital in urban planning, development, disaster management, and various other applications. Deep learning methodologies have gained prominence in building segmentation maps, offering the promise of precise footprint extraction without extensive post-processing. However, these methods face challenges in generalization and label efficiency, particularly in remote sensing, where obtaining accurate labels can be both expensive and time-consuming. To address these challenges, we propose terrain-aware self-supervised learning, tailored to remote sensing, using digital elevation models from LIght Detection and Ranging (LiDAR) data. We propose to learn a model to differentiate between bare Earth and superimposed structures enabling the network to implicitly learn domain-relevant features without the need for extensive pixel-level annotations. We test the effectiveness of our approach by evaluating building segmentation performance on test datasets with varying label fractions. Remarkably, with only 1% of the labels (equivalent to 25 labeled examples), our method improves over ImageNet pretraining, showing the advantage of leveraging unlabeled data for feature extraction in the domain of remote sensing. The performance improvement is more pronounced in few-shot scenarios and gradually closes the gap with ImageNet pretraining as the label fraction increases. We test on a dataset characterized by substantial distribution shifts (including resolution variation and labeling errors) to demonstrate the generalizability of our approach. When compared to other baselines, including ImageNet pretraining and more complex architectures, our approach consistently performs better, demonstrating the efficiency and effectiveness of self-supervised terrain-aware feature learning.
Anuja Vats, David Völgyes, Martijn Vermeer, Marius Pedersen, Kiran B. Raja, Daniele Fantin, Jacob Alexander Hay
IEEE Trans. Geosci. Remote. Sens.4
2023 This Changes to That : Combining Causal and Non-Causal Explanations to Generate Disease Progression in Capsule Endoscopy
abstract
The need to understand the decision-making mechanisms of deep learning networks has led to a growing effort in exploring both modal-dependent and model-agnostic research methods. Although both of these ideas provide transparency for automated decision making, most methodologies focus on either using the modal-gradients (model- dependent) or ignoring the model internal states and reasoning with a model's behavior/outcome (model-agnostic) to instances. In this work, we propose a unified explanation approach that given an instance combines both model-dependent and agnostic explanations to produce an explanation set. The generated explanations are not only consistent in the neighborhood of a sample but can highlight causal relationships between image content and the outcome. We use the Wireless Capsule Endoscopy (WCE) domain to illustrate the effectiveness of our explanations. The saliency maps generated by our approach are competitive on the softmax information score.
Anuja Vats, Ahmed Kedir Mohammed, Marius Pedersen, Nirmalie Wiratunga
ICASSP3
2023 Enhancement of clustering techniques by coupling clustering tree and neural network: Application to brain tumour segmentation
abstract
Abstract Currently, no classical clustering algorithm is efficient on its own. The predefined number of clusters required for their operation does not consistently produce satisfactory segmentation results. They exhibit cluster instability, are vulnerable to the local optimum trap, and are sensitive to noise and imaging artefacts. Most contributions designed to overcome these drawbacks incorporate prior knowledge such as cluster label information and statistic measures that demand minimal labelled training data. Although these approaches improve the segmentation accuracy, they tend to diminish the advantages of clustering algorithms over the supervised learning methods. This study proposes a shift from the use of a predefined number of clusters to a clustering tree‐based method for performance enhancement of classical clustering algorithms. The proposed method is a three‐stage algorithm. It begins with the extraction of low‐level features from a clustering tree. Clustering trees are sets of labelled clusters of an image at multiple clustering resolutions. The second stage extracts high‐level features by coupling the clustering tree to a single‐layer feedforward neural network. The third stage is the classification stage, where the basic model of a neural network extracts the tumour from a high‐level feature map. Because neither of the neural networks requires training, the proposed method is both fully unsupervised and fully automated and retains all its advantages over supervised methods. A performance evaluation using FLAIR MRI images of brain tumour patients from the BRATS2015 and BRATS2020 databases demonstrates significant performance enhancement over four classical clustering algorithms and two of the four proposed techniques were comparable to deep learning methods.
Michael Osadebey, Marius Pedersen, Meeta Kalra, Dag Waaler, Nizar Bouguila
Expert Syst. J. Knowl. Eng.2
2023 Proposal of a new fidelity measure between computed image quality and observers quality scores accounting for scores variability
Pedro Latorre-Carmona, Rafael Huertas, Marius Pedersen, Samuel Morillas
J. Vis. Commun. Image Represent.3
2023 Learning the Distribution-Based Temporal Knowledge With Low Rank Response Reasoning for UAV Visual Tracking
abstract
In recent years, the constraint based correlation filter has shown good performance in unmanned aerial vehicle (UAV) tracking, which gains a lot popularity in many intelligence transportation applications. In this work, a distribution-based temporal knowledge driven method is proposed to leverage the temporal translation property in UAV tracking. Instead of focusing on the traditional issues in the correlation filter, we provide a new method of learning parametric distribution on temporal knowledge by Wasserstein distance which is successfully embedded to solve the problem of temporal degeneration in learning process of tracking. Furthermore, we approximate optimal response reasoning with low-rank constraint over response consistency. Furthermore, the proposed method is solved by a simple iterative scheme with alternating direction multiplication ADMM algorithm. We demonstrate the superior tracking performance in several public standard UAV tracking benchmarks compared with state-of-the-art algorithms.
Guoxia Xu, Hao Wang 0003, Meng Zhao 0001, Marius Pedersen, Hu Zhu
IEEE Trans. Intell. Transp. Syst.4
2022 FCFusion: Fractal Componentwise Modeling With Group Sparsity for Medical Image Fusion
abstract
Multimodal image fusion is the process of combing relevant biological information that can be used for automated industrial application. In this article, we present a novel framework combining fractal constraint with group sparsity to achieve the optimal fusion quality. First, we adopt the idea of patch division and componentwise separation to perceive the fractal characteristics across multimodality sources. Then, to preserve the spatial information against the redundancy of component-entanglement, the group sparsity is proposed. A dual variable weighting rule is inherently embedded to mitigate the overfitting across the component penalty. Furthermore, the alternating direction method of multipliers is conducted to the proposed model optimization. The experiments show that our model has a better performance in quantitative visual quality and qualitative evaluation analysis. Finally, a real segmentation application of positron emission tomography/computed tomography image fusion proves the effectiveness of our algorithm.
Guoxia Xu, Xiaoxue Deng, Xiaokang Zhou, Marius Pedersen, Lucia Cimmino, Hao Wang 0003
IEEE Trans. Ind. Informatics4
2021 Learning More for Free - A Multi Task Learning Approach for Improved Pathology Classification in Capsule Endoscopy
Anuja Vats, Marius Pedersen, Ahmed Kedir Mohammed, Oistein Hovde
MICCAI (7)2
2021 SEENS: Nuclei segmentation in Pap smear images with selective edge enhancement
Meng Zhao 0001, Hao Wang 0003, Xiaokang Wang 0001, Hongning Dai, Xuguo Sun, Marius Pedersen
Future Gener. Comput. Syst.8
2021 The Role of Subsurface Scattering in Glossiness Perception
abstract
This study investigates the potential impact of subsurface light transport on gloss perception for the purposes of broadening our understanding of visual appearance in computer graphics applications. Gloss is an important attribute for characterizing material appearance. We hypothesize that subsurface scattering of light impacts the glossiness perception. However, gloss has been traditionally studied as a surface-related quality and the findings in the state-of-the-art are usually based on fully opaque materials, although the visual cues of glossiness can be impacted by light transmission as well. To address this gap and to test our hypothesis, we conducted psychophysical experiments and found that subjects are able to tell the difference in terms of gloss between stimuli that differ in subsurface light transport but have identical surface qualities and object shape. This gives us a clear indication that subsurface light transport contributes to a glossy appearance. Furthermore, we conducted additional experiments and found that the contribution of subsurface scattering to gloss varies across different shapes and levels of surface roughness. We argue that future research on gloss should include transparent and translucent media and to extend the perceptual models currently limited to surface scattering to more general ones inclusive of subsurface light transport.
Davit Gigilashvili, Zeyu Wang 0003, Marius Pedersen, Jon Yngve Hardeberg, Holly E. Rushmeier
ACM Trans. Appl. Percept.4
2020 Prediction Of Chromaticvisualmaskingwithdeeplearning
abstract
Visual masking is a well-studied phenomenon that has been exploited for signal compression, computer graphics and data hiding. Among the different types of visual masking, chromatic masking has received very little attention despite its importance and proven potential for the aforementioned applications. In this paper, we ask whether a deep neural network can learn to predict the detection thresholds in a chromatic masking paradigm. For that, a CNN model was trained and evaluated using a dataset made of 480 image patches for which chromatic thresholds were registered in terms of log-Gabor targets, as well as Root Mean Square (RMS) error. Experimental results show the superiority of the proposed approach.
Aladine Chetouani, Marius Pedersen, Steven Le Moan
ICIP2
2020 High-Level Visual Masking of Image Compression Artefacts
abstract
We present the results of a subjective experiment where we measured detection thresholds for 2°wide noise targets placed in 23 different natural scenes. Unlike previous studies on visual masking, we focus particularly on dissociating cases of low-level and high-level masking. That is, cases where the target is not perceived predominantly due to limits of either early or late vision. To that end, we exploit the change blindness paradigm and analyse detection rates, times and primed subjective ratings of target visibility. Our results are of significance for developing advanced models of human vision for signal quality/fidelity assessment, particularly in the context of compression.
Steven Le Moan, Marius Pedersen, Aladine Chetouani
ICIP2
2020 PS-DeVCEM: Pathology-sensitive deep learning model for video capsule endoscopy based on weakly labeled data
abstract
We propose a novel pathology-sensitive deep learning model (PS-DeVCEM) for frame-level anomaly detection and multi-label classification of different colon diseases in video capsule endoscopy (VCE) data. Our proposed model is capable of coping with the key challenge of colon apparent heterogeneity caused by several types of diseases. Our model is driven by attention-based deep multiple instance learning and is trained end-to-end on weakly labeled data using video labels instead of detailed frame-by-frame annotation. This makes it a cost-effective approach for the analysis of large capsule video endoscopy repositories. Other advantages of our proposed model include its capability to localize gastrointestinal anomalies in the temporal domain within the video frames, and its generality, in the sense that abnormal frame detection is based on automatically derived image features. The spatial and temporal features are obtained through ResNet50 and residual Long short-term memory (residual LSTM) blocks, respectively. Additionally, the learned temporal attention module provides the importance of each frame to the final label prediction. Moreover, we developed a self-supervision method to maximize the distance between classes of pathologies. We demonstrate through qualitative and quantitative experiments that our proposed weakly supervised learning model gives a superior precision and F1-score reaching, 61.6% and 55.1%, as compared to three state-of-the-art video analysis methods respectively. We also show our model’s ability to temporally localize frames with pathologies, without frame annotation information during training. Furthermore, we collected and annotated the first and largest VCE dataset with only video labels. The dataset contains 455 short video segments with 28,304 frames and 14 classes of colorectal diseases and artifacts. Dataset and code supporting this publication will be made available on our home page.
Ahmed Kedir Mohammed, Ivar Farup, Marius Pedersen, Sule Yildirim Yayilgan, Oistein Hovde
Comput. Vis. Image Underst.3
2019 On the Use of a Convolutional Neural Network to Predict Perceptual Quality of Images without Reference for Different Viewing Distances
abstract
A plethora of image quality metrics have been proposed in the literature. These metrics aims to estimate the perceptual image quality automatically. One important key aspect that the perceived quality is dependent on is the viewing distance from the observer to the image. In this study, we propose to consider this information by estimating the quality of a given image without a reference image for different viewing distances. For that, a Convolutional Neural Network (CNN) model was used in this study. Relevant patches are first selected from the image and they are then used as inputs to the CNN. The selection is here based on saliency information. The used CNN is composed of two outputs that correspond to the predicted subjective scores for two viewing distances (50 cm and 100 cm). Our method was evaluated using the Colourlab Image Database: Image Quality (CID:IQ) that provides subjective scores at two different viewing distances. The obtained results show the efficiency of our method.
Aladine Chetouani, Marius Pedersen
ICIP2
2019 Subjective Image Fidelity Assessment: Effect of the Spatial Distance Between Stimuli
abstract
Understanding how we perceive visual quality is important in a range of applications such as streaming or cross-media reproduction. Despite current perception models showing high correlations with recorded mean opinion scores, the factors influencing visual quality are still not fully understood, particularly when it comes to memory. We designed and carried out a study to compare quality assessment for two different levels of reliance on visual short-term memory. We found that assessments based mostly on memory tend to be more positive for compression, blur or gamut mapping distortions. Our results further highlight the role of memory in subjective quality assessment and visual masking.
Steven Le Moan, Marius Pedersen
ICIP2
2019 Plant Leaves Region Segmentation in Cluttered and Occluded Images Using Perceptual Color Space and K-means-Derived Threshold with Set Theory
abstract
Presence of clutters and occluding objects within agricultural farm environments challenges accurate segmentation of plant leaves, a prerequisite for an effective machine-vision-based automation of agricultural tasks. In this paper, we propose a plant leaves segmentation method that can be integrated into vision-based robotic harvester and quality inspection systems. The proposed method combines the discriminatory power of color-based technique with the simplicity and computational efficiency of threshold-based technique. Clutters and occluding objects are eliminated by infinitesimal angular displacement of the threshold image, followed by the application of set theory. Performance evaluation shows that the proposed method demonstrate strong robust features and computational efficiency.
Michael Osadebey, Marius Pedersen, Dag Waaler
INDIN2
2018 Y-Net: A deep Convolutional Neural Network for Polyp Detection
Ahmed Kedir Mohammed, Sule Yildirim Yayilgan, Marius Pedersen, Ivar Farup, Oistein Hovde
BMVC3
2018 Measuring the Effect of High-Level Visual Masking in Subjective Image Quality Assessment with Priming
abstract
Despite recent advances in subjective image quality research, many fundamental questions are still unanswered. Although we understand early vision fairly well, little is known about late vision and how the two interact with each other. In this paper, we look at one particular limit in that interaction: high-level visual masking, which is best illustrated by the change blindness paradigm. We carried out a user study designed specifically to measure the influence of high-level masking by means of priming. Results suggest a significant influence of high-level masking in image fidelity assessment at the 95% confidence level for half of the participants, with an average magnitude over three times that of intra-observer variability.
Steven Le Moan, Marius Pedersen
ICIP2
2017 Sparse Coded Handcrafted and Deep Features for Colon Capsule Video Summarization
abstract
Capsule endoscopy, which uses a wireless camera to take images of the digestive track, is emerging as an alternative to traditional wired colonoscopy. A single examination produces a sequence of approximately 50,000 frames. These sequences are manually reviewed, which is time consuming and typically takes about 45-90 minutes and requires the undivided concentration of the reviewer. In this paper, we propose a novel capsule video summarization framework using sparse coding and dictionary learning in feature space. Video frames are clustered into superframes based on power spectral density, and cluster representative frames are used for video summarization. Handcrafted and deep features that are extracted for representative frames are sparse coded using a learned dictionary. Sparse coded features are later used for training SVM classifier. The proposed method was compared with state-of-the-art methods based on sensitivity and specificity. The achieved results show that our proposed framework provides robust capsule video summarization without losing informative segments.
Ahmed Kedir Mohammed, Sule Yildirim Yayilgan, Marius Pedersen, Oistein Hovde, Faouzi Alaya Cheikh
CBMS3
2017 Can no-reference image quality metrics assess visible wavelength iris sample quality?
abstract
The overall performance of iris recognition systems is affected by the quality of acquired iris sample images. Due to the development of imaging technologies, visible wavelength iris recognition gained a lot of attention in the past few years. However, iris sample quality of unconstrained imaging conditions is a more challenging issue compared to the traditional near infrared iris biometrics. Therefore, measuring the quality of such iris images is essential in order to have good quality samples for iris recognition. In this paper, we investigate whether general purpose no-reference image quality metrics can assess visible wavelength iris sample quality.
Xinwei Liu 0001, Marius Pedersen, Christophe Charrier, Patrick Bours
ICIP2
2017 Evidence of change blindness in subjective image fidelity assessment
abstract
Change blindness is a striking phenomenon which basically means that we can look without seeing. It originates from a faulty communication between early vision (the eye) and visual working memory (the brain). In this paper, we present evidence that this faulty communication needs to be accounted for in image fidelity assessment (also known as full-reference image quality assessment). We designed a user study to analyse participants opinions based on how much they have to rely on their visual working memory in order to give fidelity score. Results demonstrate that significantly more severe judgments were made when reliance on visual short-term memory was minimal, suggesting limitations in the observers' ability to notice image differences in the typical pairwise comparison setup. Furthermore, a comparison of the efficiency of six state-of-the-art image fidelity assessment models (so-called metrics) reveals that five of them perform significantly better at predicting results obtained when reliance on memory is minimal.
Steven Le Moan, Marius Pedersen
ICIP2
2017 No-reference quality measure in brain MRI images using binary operations, texture and set analysis
abstract
The authors propose a new application‐specific, post‐acquisition quality evaluation method for brain magnetic resonance imaging (MRI) images. The domain of a MRI slice is regarded as universal set. Four feature images; greyscale, local entropy, local contrast and local standard deviation are extracted from the slice and transformed into the binary domain. Each feature image is regarded as a set enclosed by the universal set. Four qualities attribute; lightness, contrast, sharpness and texture details are described by four different combinations of feature sets. In an ideal MRI slice, the four feature sets are identically equal. Degree of distortion in real MRI slice is quantified by fidelity between the sets that describe a quality attribute. Noise is the fifth quality attribute and is described by the slice Euler number region property. Total quality score is the weighted sum of the five quality scores. The authors' proposed method addresses current challenges in image quality evaluation. It is simple, easy‐to‐use and easy‐to‐understand. Incorporation of binary transformation in the proposed method reduces computational and operational complexity of the algorithm. They provide experimental results that demonstrate efficacy of their proposed method on good quality images and on common distortions in MRI images of the brain.
Michael Osadebey, Marius Pedersen, Douglas L. Arnold, Katrina Wendel-Mitoraj
IET Image Process.2
2016 The influence of short-term memory in subjective image quality assessment
abstract
Aiming at understanding the role of short-term memory in subjective image quality assessment, we report and compare results from two pair-comparison methods: stimuli shown side-by-side versus stimuli shown one after the other. Our results suggest that there is a significant chance that an observer will make different quality assessments in the two setups.
Steven Le Moan, Marius Pedersen, Ivar Farup, Jana Blahová
ICIP2
2015 Evaluation of 60 full-reference image quality metrics on the CID: IQ
abstract
Image quality metrics have become very popular and new metrics are proposed continuously. They have usually been developed with the goal of correlating with subjective image quality assessment. We perform an extensive evaluation of 60 state-of-the-art image quality metrics, including well-known metrics, such as SSIM, multiscale SSIM, VIF, MSE, S-DEE, CID, MAD, S-CIELAB, SHAME, VSNR, and PSNR. Evaluation is performed on the on the Colourlab Image Database: Image Quality (CID:IQ), a database consisting of 690 images where the subjective data has been collected at two different viewing distances. The performance of the image quality metrics is assessed in terms of correlation to subjective data.
Marius Pedersen
ICIP1
2014 Image registration for quality assessment of projection displays
abstract
In the full reference metric based image quality assessment of projection displays, it is critical to achieve accurate and fully automatic image registration between the captured projection and its reference image in order to establish a subpixel level mapping. The preservation of geometrical order as well as the intensity and chromaticity relationships between two consecutive pixels must be maximized. The existing camera based image registration methods do not meet this requirement well. In this paper, we propose a markerless and view independent method to use an un-calibrated camera to perform the task. The proposed method including three main components: feature extraction, feature expansion and geometric correction, and it can be implemented easily in a fully automatic fashion. The experimental results of both simulation and the one conducted in the field demonstrate that the proposed method is able to achieve image registration accuracy higher than 91% in a dark projection room and above 85% with ambient light lower than 30 Lux.
Marius Pedersen, Jon Yngve Hardeberg, Jean-Baptiste Thomas
ICIP2
2014 A Variational Approach for Denoising Hyperspectral Images Corrupted by Poisson Distributed Noise
abstract
Poisson distributed noise, such as photon noise is an important noise source in multi- and hyperspectral images. We propose a variational based denoising approach, that accounts the vectorial structure of a spectral image cube, as well as the poisson distributed noise. For this aim, we extend an approach for monochromatic images, by a regularisation term, that is spectrally and spatially adaptive and preserves edges. In order to take the high computational complexity into account, we derive a Split Bregman optimisation for the proposed model. The results show the advantages of the proposed approach compared to a marginal approach on synthetic and real data. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Ferdinand Deger, Alamin Mansouri, Marius Pedersen, Jon Yngve Hardeberg, Yvon Voisin
ICISP3
2014 CID: IQ - A New Image Quality Database
Xinwei Liu 0001, Marius Pedersen, Jon Yngve Hardeberg
ICISP2
2014 Exposure Fusion Algorithm Based on Perceptual Contrast and Dynamic Adjustment of Well-Exposedness
Pablo Martínez-Cañada, Marius Pedersen
ICISP2
2012 Spatial pooling for measuring color printing quality attributes
Mingming Gong, Marius Pedersen
J. Vis. Commun. Image Represent.2
2012 Measuring perceptual contrast in digital images
Gabriele Simone, Marius Pedersen, Jon Yngve Hardeberg
J. Vis. Commun. Image Represent.2