VLDB 2026 Research / reviewers in the wild / expert
Stefan Winkler 0001
dblp:57/4176
· DBLP profile ↗
72ranked-venue papers
9as first author
13since 2021 · last 2025
0000-0003-4305-8408ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 9 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 since 2021Human-computer interaction and ubiquitous computing · 6Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic DialoguesabstractAutomatic Speech Recognition (ASR) systems are pivotal in transcribing speech into text, yet the errors they introduce can significantly degrade the performance of downstream tasks like summarization. This issue is particularly pronounced in clinical dialogue summarization, a low-resource domain where supervised data for fine-tuning is scarce, necessitating the use of ASR models as black-box solutions. Employing conventional data augmentation for enhancing the noise robustness of summarization models is not feasible either due to the unavailability of sufficient medical dialogue audio recordings and corresponding ASR transcripts. To address this challenge, we propose MEDSAGE, an approach for generating synthetic samples for data augmentation using Large Language Models (LLMs). Specifically, we leverage the in-context learning capabilities of LLMs and instruct them to generate ASR-like errors based on a few available medical dialogue examples with audio recordings. Experimental results show that LLMs can effectively model ASR noise, and incorporating this noisy data into the training process significantly improves the robustness and accuracy of medical dialogue summarization systems. This approach addresses the challenges of noisy ASR outputs in critical applications, offering a robust solution to enhance the reliability of clinical dialogue summarization. Kuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu, Vijay Prakash Dwivedi, Thanh-Tung Nguyen, Xiaoxue Gao, Nancy F. Chen, Stefan Winkler 0001 |
AAAI | 9 |
| 2024 | MaxEnt Loss: Constrained Maximum Entropy for Calibration under Out-of-Distribution ShiftabstractWe present a new loss function that addresses the out-of-distribution (OOD) network calibration problem. While many objective functions have been proposed to effectively calibrate models in-distribution, our findings show that they do not always fare well OOD. Based on the Principle of Maximum Entropy, we incorporate helpful statistical constraints observed during training, delivering better model calibration without sacrificing accuracy. We provide theoretical analysis and show empirically that our method works well in practice, achieving state-of-the-art calibration on both synthetic and real-world benchmarks. Our code is available at https://github.com/dexterdley/MaxEnt-Loss. Yuan Rong Dexter Neo, Stefan Winkler 0001, Tsuhan Chen |
AAAI | 2 |
| 2024 | A Comprehensive Survey of Sentence Representations: From the BERT Epoch to the CHATGPT Era and BeyondabstractAbhinav Ramesh Kashyap, Thanh-Tung Nguyen, Viktor Schlegel, Stefan Winkler, See-Kiong Ng, Soujanya Poria. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Abhinav Ramesh Kashyap, Thanh-Tung Nguyen, Viktor Schlegel, Stefan Winkler 0001, See-Kiong Ng, Soujanya Poria |
EACL (1) | 4 |
| 2024 | Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?abstractState-of-the-art Large Language Models (LLMs) are accredited with an increasing number of different capabilities, ranging from reading comprehension over advanced mathematical and reasoning skills to possessing scientific knowledge.In this paper we focus on multi-hop reasoning-the ability to identify and integrate information from multiple textual sources.Given the concerns with the presence of simplifying cues in existing multi-hop reasoning benchmarks, which allow models to circumvent the reasoning requirement, we set out to investigate whether LLMs are prone to exploiting such simplifying cues.We find evidence that they indeed circumvent the requirement to perform multi-hop reasoning, but they do so in more subtle ways than what was reported about their fine-tuned pre-trained language model (PLM) predecessors.We propose a challenging multi-hop reasoning benchmark by generating seemingly plausible multi-hop reasoning chains that ultimately lead to incorrect answers.We evaluate multiple open and proprietary state-of-the-art LLMs and show that their multi-hop reasoning performance is affected, as indicated by up to 45% relative decrease in F1 score when presented with such seemingly plausible alternatives.We also find that-while LLMs tend to ignore misleading lexical cues-misleading reasoning paths indeed present a significant challenge.The code and data are made available at https: //github.com/zawedcvg/Are-Large- Language-Models-Attentive-Readers. Neeladri Bhuiya, Viktor Schlegel, Stefan Winkler 0001 |
EMNLP | 3 |
| 2024 | Efficient Training of Self-Supervised Speech Foundation Models on a Compute BudgetabstractDespite their impressive success, training foundation models remains computationally costly. This paper investigates how to efficiently train speech foundation models with self-supervised learning (SSL) under a limited compute budget. We examine critical factors in SSL that impact the budget, including model architecture, model size, and data size. Our goal is to make analytical steps toward understanding the training dynamics of speech foundation models. We benchmark SSL objectives in an entirely comparable setting and find that other factors contribute more significantly to the success of SSL. Our results show that slimmer model architectures outperform common small architectures under the same compute and parameter budget. We demonstrate that the size of the pre-training data remains crucial, even with data augmentation during SSL training, as performance suffers when iterating over limited data. Finally, we identify a trade-off between model size and data size, highlighting an optimal model size for a given compute budget. Andy T. Liu, Yi-Cheng Lin, Stefan Winkler 0001, Hung-yi Lee |
SLT | 4 |
| 2023 | Transfer Learning for Cloud Image ClassificationabstractCloud image classification has been extensively studied in the literature, as it has several radio-meteorological and remote sensing applications. Recently, images from ground-based sky imagers (GSIs) are being widely used because of their high temporal and spatial resolution and low infrastructure cost as compared to satellites. To classify sky/cloud images obtained from such GSIs, this paper1examines the application of transfer learning using the standard VGG-16 architecture. The paper further analyzes the importance of adjusting the number of neurons in the top dense layers to improve the performance of the model. The reasons for the same are traced by conducting extensive experiments on multiple datasets exhibiting varied properties. Navya Jain, Yee Hui Lee, Stefan Winkler 0001, Soumyabrata Dev |
IGARSS | 4 |
| 2022 | Trusted Media Challenge Dataset and User StudyabstractThe emergence of fake media that can be easily created by technology has the potential to generate potent misinformation causing harm to both society and individuals. To tackle the issue, we have organized the Trusted Media Challenge (TMC) to explore how Artificial Intelligence (AI) technologies could be leveraged to combat fake media. To enable further research, we are releasing the dataset from the TMC, consists of 4,380 fake and 2,563 real videos, with various video and audio manipulation methods employed to produce different types of fake media. We have also carried out a user study to demonstrate the quality of the TMC dataset and to compare the performance of humans and AI models. The results show that the TMC dataset can fool human participants in many cases. The TMC dataset is available for research purposes upon request via [email protected] Sheng Lun Benjamin Chua, Stefan Winkler 0001, See-Kiong Ng |
CIKM | 3 |
| 2022 | Image Data Augmentation with Unpaired Image-to-Image Camera Model TranslationabstractMany image datasets are built from web searches, with images taken by various cameras. The variance of camera sources can lead to different camera signals and colors within images of the same class, which may impede neural networks from fitting the data. To generalize neural networks to different camera sources, we propose an augmentation method using unpaired image-to-image translation to transfer training images into another camera model domain. Our approach utilizes CycleGAN to create a translation mapping between two different camera models. We show that such a mapping can be applied to any image as a form of data augmentation and is able to outperform traditional color-based transformations. Additionally, this approach can be further enhanced with geometric transformations. Chi Fa Foo, Stefan Winkler 0001 |
ICIP | 2 |
| 2022 | Mixed Membership Generative Adversarial NetworksabstractGANs are designed to learn a single distribution, though multiple distributions can be modeled by treating them separately. However, this naive implementation does not consider overlapping distributions. We propose Mixed Membership Generative Adversarial Networks (MMGAN) analogous to mixed-membership models that model multiple distributions and discover their commonalities and particularities. Each data distribution is modeled as a mixture over a common set of generator distributions, and mixture weights are automatically learned from the data. Mixture weights can give insight into common and unique features of each data distribution. We evaluate our proposed MMGAN and show its effectiveness on MNIST and Fashion-MNIST with various settings. Yasin Yazici, Bruno Lecouat, Kim-Hui Yap, Stefan Winkler 0001, Georgios Piliouras, Vijay Chandrasekhar 0001, Chuan-Sheng Foo |
ICIP | 4 |
| 2022 | Recognition of Advertisement Emotions With Application to Computational AdvertisingabstractAdvertisements (ads) often contain strong emotions to capture audience attention and convey an effective message. Still, little work has focused on affect recognition (AR) from ads employing audiovisual or user cues. This work (1) compiles an affective video ad dataset which evokes coherent emotions across users; (2) explores the efficacy of content-centric convolutional neural network (CNN) features for ad AR vis-ã-vis handcrafted audio-visual descriptors; (3) examines user-centric ad AR from Electroencephalogram (EEG) signals, and (4) demonstrates how better affect predictions facilitate effective computational advertising via a study involving 18 users. Experiments reveal that (a) CNN features outperform handcrafted audiovisual descriptors for content-centric AR; (b) EEG features encode ad-induced emotions better than content-based features; (c) Multi-task learning achieves optimal ad AR among a slew of classifiers and (d) Pursuant to (b), EEG features enable optimized ad insertion onto streamed video compared to content-based or manual insertion, maximizing ad recall and viewing experience. Abhinav Shukla, Shruti Shriya Gullapuram, Harish Katti, Mohan Kankanhalli, Stefan Winkler 0001, Subramanian Ramanathan |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | PersEmoN: A Deep Network for Joint Analysis of Apparent Personality, Emotion and Their RelationshipabstractApparent personality and emotion analysis are both central to affective computing. Existing works solve them individually. In this paper we investigate if such high-level affect traits and their relationship can be jointly learned from face images in the wild. To this end, we introducePersEmoN, an end-to-end trainable and deep Siamese-like network. It consists of two convolutional network branches, one for emotion and the other for apparent personality. Both networks share their bottom feature extraction module and are optimized within a multi-task learning framework. Emotion and personality networks are dedicated to their own annotated dataset. Furthermore, an adversarial-like loss function is employed to promote representation coherence among heterogeneous dataset sources. Based on this, we also explore the emotion-to-apparent-personality relationship. Extensive experiments demonstrate the effectiveness ofPersEmoN. Le Zhang 0001, Songyou Peng, Stefan Winkler 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Morphset: Augmenting Categorical Emotion Datasets With Dimensional Affect Labels Using Face MorphingabstractEmotion recognition and understanding is a vital component in human-machine interaction. Dimensional models of affect such as those using valence and arousal have advantages over traditional categorical ones due to the complexity of emotional states in humans. However, dimensional emotion annotations are difficult and expensive to collect, therefore they are not as prevalent in the affective computing community. To address these issues, we propose a method to generate synthetic images from existing categorical emotion datasets using face morphing as well as dimensional labels in the circumplex space with full control over the resulting sample distribution, while achieving augmentation factors of at least 20x or more. Vassilios Vonikakis, Yuan Rong Dexter Neo, Stefan Winkler 0001 |
ICIP | 3 |
| 2021 | ADGD'21: 1st Workshop on Synthetic Multimedia - Audiovisual Deepfake Generation and DetectionabstractDeepfakes, i.e.synthetic or "fake" media content generated using deep learning, are a double-edged sword. On one hand, they pose new threats and risks in the form of scams, fraud, disinformation, social manipulation, or celebrity porn. On the other hand, deepfakes have just as many meaningful and beneficial applications - they allow us to create and experience things that no longer exist, or that have never existed, enabling numerous exciting applications in entertainment, education, and even privacy. Stefan Winkler 0001, Abhinav Dhall, Pavel Korshunov |
ACM Multimedia | 1 |
| 2020 | Identity-Invariant Facial Landmark Frontalization For Facial Expression AnalysisabstractWe propose a frontalization technique for 2D facial landmarks, designed to aid in the analysis of facial expressions. It employs a new normalization strategy aiming to minimize identity variations, by displacing groups of facial landmarks to standardized locations. The technique operates directly on 2D landmark coordinates, does not require additional feature extraction and as such is computationally light. It achieves considerable improvement over a reference approach, justifying its use as an efficient preprocessing step for facial expression analysis based on geometric features. Vassilios Vonikakis, Stefan Winkler 0001 |
ICIP | 2 |
| 2020 | Empirical Analysis Of Overfitting And Mode Drop In Gan TrainingabstractWe examine two key questions in GAN training, namely overfitting and mode drop, from an empirical perspective. We show that when stochasticity is removed from the training procedure, GANs can overfit and exhibit almost no mode drop. Our results shed light on important characteristics of the GAN training procedure. They also provide evidence against prevailing intuitions that GANs do not memorize the training set, and that mode dropping is mainly due to properties of the GAN objective rather than how it is optimized during training. Yasin Yazici, Chuan-Sheng Foo, Stefan Winkler 0001, Kim-Hui Yap, Vijay Chandrasekhar 0001 |
ICIP | 3 |
| 2020 | HelipadCat: Categorised Helipad Image Dataset and Detection MethodabstractWe present HelipadCat, a dataset of aerial images of helipads, together with a method to identify and locate such helipads from the air. Based on the FAA's database of US airports, we create the first dataset of helipads, including a classification by visual helipad shape and features, which we make available to the research community. The dataset includes nearly 6,000 images with 12 different categories.We then train several Mask-RCNN models based on ResNet101 using our dataset. Image augmentation is applied according to learned augmentation policies. We characterize the performance of the models on HelipadCat and pick the best-performing configuration. We further evaluate that model on the metropolitan area of Manila and show that it is able to detect helipads successfully, with their exact geographical coordinates, in another country. To reduce false positives, the bounding boxes are filtered by confidence score, size, and the presence of shadows. Dataset and code are available for download. Jonas Bitoun, Stefan Winkler 0001 |
TENCON | 2 |
| 2019 | The Unusual Effectiveness of Averaging in GAN Training
Yasin Yazici, Chuan-Sheng Foo, Stefan Winkler 0001, Kim-Hui Yap, Georgios Piliouras, Vijay Chandrasekhar 0001 |
ICLR (Poster) | 3 |
| 2019 | CloudSegNet: A Deep Network for Nychthemeron Cloud Image SegmentationabstractWe analyze clouds in the earth's atmosphere using ground-based sky cameras. An accurate segmentation of clouds in the captured sky/cloud image is difficult, owing to the fuzzy boundaries of clouds. Several techniques have been proposed, which use color as the discriminatory feature for cloud detection. In the existing literature, however, analysis of daytime and nighttime images is considered separately, mainly because of differences in image characteristics and applications. In this letter, we propose a lightweight deep-learning architecture called CloudSegNet. It is the first that integrates daytime and nighttime (also known as nychthemeron) image segmentation in a single framework and achieves state-of-the-art results on public databases. Soumyabrata Dev, Atul Nautiyal, Yee Hui Lee, Stefan Winkler 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | A Data-Driven Approach for Accurate Rainfall PredictionabstractIn recent years, there has been growing interest in using precipitable water vapor (PWV) derived from global positioning system (GPS) signal delays to predict rainfall. However, the occurrence of rainfall is dependent on a myriad of atmospheric parameters. This paper proposes a systematic approach to analyze various parameters that affect precipitation in the atmosphere. Different ground-based weather features such as Temperature, Relative Humidity, Dew Point, Solar Radiation, PWV along with Seasonal and Diurnal variables are identified, and a detailed feature correlation study is presented. While all features play a significant role in rainfall classification, only a few of them, such as PWV, Solar Radiation, Seasonal, and Diurnal features, stand out for rainfall prediction. Based on these findings, an optimum set of features are used in a data-driven machine learning algorithm for rainfall prediction. The experimental evaluation using a 4-year (2012-2015) database shows a true detection rate of 80.4%, a false alarm rate of 20.3%, and an overall accuracy of 79.6%. Compared to the existing literature, our method significantly reduces the false alarm rates. Shilpa Manandhar, Soumyabrata Dev, Yee Hui Lee, Yu Song Meng, Stefan Winkler 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | EEG-based Evaluation of Cognitive Workload Induced by Acoustic Parameters for Data SonificationabstractData Visualization has been receiving growing attention recently, with ubiquitous smart devices designed to render information in a variety of ways. However, while evaluations of visual tools for their interpretability and intuitiveness have been commonplace, not much research has been devoted to other forms of data rendering, \eg, sonification. This work is the first to automatically estimate the cognitive load induced by different acoustic parameters considered for sonification in prior studies~\citeferguson2017evaluation,ferguson2018investigating. We examine cognitive load via (a) perceptual data-sound mapping accuracies of users for the different acoustic parameters, (b) cognitive workload impressions explicitly reported by users, and (c) their implicit EEG responses compiled during the mapping task. Our main findings are that (i) low cognitive load-inducing (ıe, more intuitive) acoustic parameters correspond to higher mapping accuracies, (ii) EEG spectral power analysis reveals higher α band power for low cognitive load parameters, implying a congruent relationship between explicit and implicit user responses, and (iii) Cognitive load classification with EEG features achieves a peak F1-score of 0.64, confirming that reliable workload estimation is achievable with user EEG data compiled using wearable sensors. Maneesh Bilalpur, Mohan Kankanhalli, Stefan Winkler 0001, Subramanian Ramanathan |
ICMI | 3 |
| 2018 | Systematic Study of Weather Variables for Rainfall DetectionabstractNumerous weather parameters affect the occurrence and amount of rainfall. Therefore, it is important to study these parameters and their interdependency. In this paper, different weather and time-related variables - relative humidity, solar radiation, temperature, dew point, day-of-year, and time-of-day are analyzed systematically using Principal Component Analysis (PCA). We found that four principal components explain a cumulative variance of 85%. The first two principal components are applied to distinguish rain and no-rain scenarios as well. We conclude that all 7 variables have similar contribution towards rainfall detection. Shilpa Manandhar, Soumyabrata Dev, Yee Hui Lee, Stefan Winkler 0001, Yu Song Meng |
IGARSS | 4 |
| 2018 | A Data-Driven Approach to Detect Precipitation from Meteorological Sensor DataabstractPrecipitation is dependent on a myriad of atmospheric conditions. In this paper, we study how certain atmospheric parameters impact the occurrence of rainfall. We propose a data-driven, machine-learning based methodology to detect precipitation using various meteorological sensor data. Our approach achieves a true detection rate of 87.4% and a moderately low false alarm rate of 32.2%. Shilpa Manandhar, Soumyabrata Dev, Yee Hui Lee, Yu Song Meng, Stefan Winkler 0001 |
IGARSS | 5 |
| 2018 | Give Me One Portrait Image, I Will Tell You Your Emotion and PersonalityabstractPersonality and emotion are both central to affective computing. Existing works address them individually. In this demo we investigate if such high-level affect traits and their relationship can be jointly learned from face images in the wild. To this end, we introduce an end-to-end trainable and deep Siamese-like network. At inference time, our system can take one portrait photo as input and predict one's Big-Five apparent personality as well as emotion attributes. With such a system, we also demonstrate the feasibility of inferring the apparent personality directly fro emotion. Songyou Peng, Le Zhang 0001, Stefan Winkler 0001, Marianne Winslett |
ACM Multimedia | 3 |
| 2018 | ASCERTAIN: Emotion and Personality Recognition Using Commercial SensorsabstractWe present ASCERTAIN-a multimodal databaASe for impliCit pERsonaliTy and Affect recognitIoN using commercial physiological sensors. To our knowledge, ASCERTAIN is the first database to connect personality traits and emotional states via physiological responses. ASCERTAIN contains big-five personality scales and emotional self-ratings of 58 users along with their Electroencephalogram (EEG), Electrocardiogram (ECG), Galvanic Skin Response (GSR) and facial activity data, recorded using off-the-shelf sensors while viewing affective movie clips. We first examine relationships between users' affective ratings and personality scales in the context of prior observations, and then study linear and non-linear physiological correlates of emotion and personality. Our analysis suggests that the emotion-personality relationship is better captured by non-linear rather than linear statistics. We finally attempt binary emotion and personality trait recognition using physiological features. Experimental results cumulatively confirm that personality differences are better revealed while comparing user responses to emotionally homogeneous videos, and above-chance recognition is achieved for both affective and personality dimensions. Subramanian Ramanathan, Julia Wache, Mojtaba Khomami Abadi, Radu L. Vieriu, Stefan Winkler 0001, Nicu Sebe |
IEEE Trans. Affect. Comput. | 5 |
| 2017 | BAFT: Binary affine feature transformabstractWe introduce BAFT, a fast binary and quasi affine invariant local image feature. It combines the affine invariance of Harris Affine feature descriptors with the speed of binary descriptors such as BRISK and ORB. BAFT derives its speed and precision from sampling local image patches in a pattern that depends on the second moment matrix of the same image patch. This approach results in a fast but discriminative descriptor, especially for image pairs with large perspective changes. Our evaluation on 40 different image pairs shows that BAFT increases the area under the precision/recall curve (AUC) compared to traditional descriptors for the majority of image pairs. In addition we show that this improvement comes with a very low performance penalty compared to the similar ORB descriptor. The BAFT source code is available for download. Jonas Toft Arnfred, Viet Dung Nguyen, Stefan Winkler 0001 |
ICIP | 3 |
| 2017 | Nighttime sky/cloud image segmentationabstractImaging the atmosphere using ground-based sky cameras is a popular approach to study various atmospheric phenomena. However, it usually focuses on the daytime. Nighttime sky/cloud images are darker and noisier, and thus harder to analyze. An accurate segmentation of sky/cloud images is already challenging because of the clouds' non-rigid structure and size, and the lower and less stable illumination of the night sky increases the difficulty. Nonetheless, nighttime cloud imaging is essential in certain applications, such as continuous weather analysis and satellite communication. In this paper, we propose a superpixel-based method to segment nighttime sky/cloud images. We also release the first nighttime sky/cloud image segmentation database to the research community. The experimental results show the efficacy of our proposed algorithm for nighttime images. Soumyabrata Dev, Florian M. Savoy, Yee Hui Lee, Stefan Winkler 0001 |
ICIP | 4 |
| 2017 | Stereoscopic cloud base reconstruction using high-resolution whole sky imagersabstractCloud base height and volume estimation is needed in meteorology and other applications. We have deployed a pair of custom-designed Whole Sky Imagers, which capture stereo pictures of the sky at regular intervals. Using these images, we propose a method to create rough 3D models of the base of clouds, using feature point detection, matching, and triangulation. The novelty of our method lies in the fact that it locates the cloud base in all three dimensions, instead of only estimating cloud base height. For validation, we compare the results with measurements from weather radar. Florian M. Savoy, Soumyabrata Dev, Yee Hui Lee, Stefan Winkler 0001 |
ICIP | 4 |
| 2017 | Rough-Set-Based Color Channel SelectionabstractColor channel selection is essential for accurate segmentation of sky and clouds in images obtained from ground-based sky cameras. Most prior works in cloud segmentation use threshold-based methods on color channels selected in an ad hoc manner. In this letter, we propose the use of rough sets for color channel selection in visible-light images. Our proposed approach assesses color channels with respect to their contribution for segmentation and identifies the most effective ones. Soumyabrata Dev, Florian M. Savoy, Yee Hui Lee, Stefan Winkler 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | A Probabilistic Approach to People-Centric Photo Selection and SequencingabstractWe present a crowdsourcing (CS) study to examine how specific attributes probabilistically affect the selection and sequencing of images from personal photo collections. Thirteen image attributes are explored, including seven people-centric properties. We first propose a novel dataset shaping technique based on mixed integer linear programming (MILP) to identify a subset of photos in which the attributes of interest are uniformly distributed and minimally correlated. Shaping enables the synthesis of compact, balanced, and representative datasets for CS, and facilitates effective learning of the selection likelihood of an image as well as its relative position in a sequence, given its attributes. We further present an ILP-based slideshow creation framework to select and arrange (a subset of) appealing images from a personal photo library. Quantitative and qualitative evaluations confirm that our method outperforms regression-based and greedy approaches for photo selection and sequencing, generating slideshows similar in quality to those created by humans. Vassilios Vonikakis, Subramanian Ramanathan, Jonas Toft Arnfred, Stefan Winkler 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Sparse Code Filtering for Action Pattern Mining
Wei Wang 0108, Yan Yan 0002, Liqiang Nie, Stefan Winkler 0001, Nicu Sebe |
ACCV (2) | 5 |
| 2016 | Shaping datasets: Optimal data selection for specific target distributions across dimensionsabstractThis paper presents a method for dataset manipulation based on Mixed Integer Linear Programming (MILP). The proposed optimization can narrow down a dataset to a particular size, while enforcing specific distributions across different dimensions. It essentially leverages the redundancies of an initial dataset in order to generate more compact versions of it, with a specific target distribution across each dimension. If the desired target distribution is uniform, then the effect is balancing: all values across all different dimensions are equally represented. Other types of target distributions can also be specified, depending on the nature of the problem. The proposed approach may be used in machine learning, for shaping training and testing datasets, or in crowdsourcing, for preparing datasets of a manageable size. Vassilios Vonikakis, Subramanian Ramanathan, Stefan Winkler 0001 |
ICIP | 3 |
| 2016 | COVERAGE - A novel database for copy-move forgery detectionabstractWe present COVERAGE - a novel database containing copy-move forged images and their originals with similar but genuine objects. COVERAGE is designed to highlight and address tamper detection ambiguity of popular methods, caused by self-similarity within natural images. In COVERAGE, forged-original pairs are annotated with (i) the duplicated and forged region masks, and (ii) the tampering factor/similarity metric. For benchmarking, forgery quality is evaluated using (i) computer vision-based methods, and (ii) human detection performance. We also propose a novel sparsity-based metric for efficiently estimating forgery quality. Experimental results show that (a) popular forgery detection methods perform poorly over COVERAGE, and (b) the proposed sparsity based metric best correlates with human detection performance. We release the COVERAGE database to the research community. Bihan Wen, Subramanian Ramanathan, Tian-Tsong Ng, Xuanjing Shen, Stefan Winkler 0001 |
ICIP | 6 |
| 2016 | Group happiness assessment using geometric features and dataset balancingabstractThis paper presents the techniques employed in our team's submissions to the 2016 Emotion Recognition in the Wild contest, for the sub-challenge of group-level emotion recognition. The objective of this sub-challenge is to estimate the happiness intensity of groups of people in consumer photos. We follow a predominately bottom-up approach, in which the individual happiness level of each face is estimated separately. The proposed technique is based on geometric features derived from 49 facial points. These features are used to train a model on a subset of the HAPPEI dataset, balanced across expression and headpose, using Partial Least Squares regression. The trained model exhibits competitive performance for a range of non-frontal poses, while at the same time offering a semantic interpretation of the facial distances that may contribute positively or negatively to group-level happiness. Various techniques are explored in combining these estimations in order to perform group-level prediction, including the distribution of expressions, significance of a face relative to the whole group, and mean estimation. Our best submission achieves an RMSE of 0.8316 on the competition test set, which compares favorably to the RMSE of 1.30 of the baseline. Vassilios Vonikakis, Yasin Yazici, Viet Dung Nguyen, Stefan Winkler 0001 |
ICMI | 4 |
| 2016 | Estimation of solar irradiance using ground-based whole sky imagersabstractGround-based whole sky imagers (WSIs) can provide localized images of the sky of high temporal and spatial resolution, which permits fine-grained cloud observation. In this paper, we show how images taken by WSIs can be used to estimate solar radiation. Sky cameras are useful here because they provide additional information about cloud movement and coverage, which are otherwise not available from weather station data. Our setup includes ground-based weather stations at the same location as the imagers. We use their measurements to validate our methods. Soumyabrata Dev, Florian M. Savoy, Yee Hui Lee, Stefan Winkler 0001 |
IGARSS | 4 |
| 2016 | Geo-referencing and stereo calibration of ground-based Whole Sky Imagers using the sun trajectoryabstractGround-based Whole Sky Imagers (WSIs) are now commonly used for cloud observations. Upon deployment, they may not be exactly level or precisely face north. This significantly affects subsequent processing of the images, especially for applications where two or more imagers are required, e.g. 3D volumetric cloud reconstruction. We present a method to remove this mis-alignment using the sun position in images captured over a whole day. Coupled with precise coordinates of the device locations, this method also improves the geo-referencing accuracy of the captured images. We detect the sun in the images and compute the corresponding 3D vectors using the lens calibration function. These vectors are compared to the actual sun direction. The mismatch between the two sets of vectors is then corrected using a 3D rotation matrix. The method can also be applied to other celestial bodies, such as stars or the moon. Florian M. Savoy, Soumyabrata Dev, Yee Hui Lee, Stefan Winkler 0001 |
IGARSS | 4 |
| 2016 | Mean opinion score (MOS) revisited: methods and applications, limitations and alternatives
Robert C. Streijl, Stefan Winkler 0001, David S. Hands |
Multim. Syst. | 2 |
| 2016 | A general framework for image feature matching without geometric constraints
Jonas Toft Arnfred, Stefan Winkler 0001 |
Pattern Recognit. Lett. | 2 |
| 2016 | Category Specific Dictionary Learning for Attribute Specific Feature SelectionabstractAttributes, as mid-level features, have demonstrated great potential in visual recognition tasks due to their excellent propagation capability through different categories. However, existing attribute learning methods are prone to learning the correlated attributes. To discover the genuine attribute specific features, many feature selection methods have been proposed. However, these feature selection methods are implemented at the level of raw features that might be very noisy, and these methods usually fail to consider the structural information in the feature space. To address this issue, in this paper, we propose a label constrained dictionary learning approach combined with a multilayer filter. The feature selection is implemented at dictionary level, which can better preserve the structural information. The label constrained dictionary learning suppresses the intra-class noise by encouraging the sparse representations of intra-class samples to lie close to their center. A multilayer filter is developed to discover the representative and robust attribute specific bases. The attribute specific bases are only shared among the positive samples or the negative samples. The experiments on the challenging Animals with Attributes data set and the SUN attribute data set demonstrate the effectiveness of our proposed method. Wei Wang 0108, Yan Yan 0002, Stefan Winkler 0001, Nicu Sebe |
IEEE Trans. Image Process. | 3 |
| 2015 | Fast-match: Fast and robust feature matching on large imagesabstractToday's cameras produce images that often exceed 10 megapixels. Yet computing and matching local features for images of this size can easily take 20 seconds or more using optimized matching algorithms. This is much too slow for interactive applications and much too expensive for large scale image operations. We introduce Fast-Match, an algorithm designed to match large images efficiently without compromising matching accuracy. It derives its speed from only computing features in those parts of the image that can be confidently matched. Fast-Match is an order of magnitude faster than the popular Ratio-Match, yet often doubles matching precision for difficult image pairs. Jonas Toft Arnfred, Stefan Winkler 0001 |
ICIP | 2 |
| 2015 | Categorization of cloud image patches using an improved texton-based approachabstractWe propose a modified texton-based classification approach that integrates both color and texture information for improved classification results. We test our proposed method for the task of cloud classification on SWIMCAT, a large new database of cloud images taken with a ground-based sky imager, with very good results. We perform an extensive evaluation, comparing different color components, filter banks, and other parameters to understand their effect on classification accuracy. Finally, we release the SWIMCAT dataset that was created for the task of cloud categorization. Soumyabrata Dev, Yee Hui Lee, Stefan Winkler 0001 |
ICIP | 3 |
| 2015 | Multi-level semantic labeling of Sky/cloud imagesabstractSky/cloud images captured by ground-based Whole Sky Imagers (WSIs) are extensively used now-a-days for various applications. In this paper, we learn the semantics of sky/cloud images, which allows an automatic annotation of pixels with different class labels. We model the various labels/classes with a continuous-valued multi-variate distribution. Using a set of training images, the distributions for different labels are learnt, and subsequently used for labeling test images. We also present a method to determine the number of clusters. Our proposed approach is the first for multi-class sky-cloud image annotation and achieves very good results. Soumyabrata Dev, Yee Hui Lee, Stefan Winkler 0001 |
ICIP | 3 |
| 2015 | On the utility of canonical correlation analysis for domain adaptation in multi-view headpose estimationabstractThe utility of canonical correlation analysis (CCA) for domain adaptation (DA) in the context of multi-view head pose estimation is examined in this work. We consider the three problems studied in [1], where different DA approaches are explored to transfer head pose-related knowledge from an extensively labeled source dataset to a sparsely labeled target set, whose attributes are vastly different from the source. CCA is found to benefit DA for all the three problems, and the use of a covariance profile-based diagonality score (DS) also improves classification performance with respect to a nearest neighbor (NN) classifier. Anoop Kolar Rajagopal, Subramanian Ramanathan, Vassilios Vonikakis, K. R. Ramakrishnan, Stefan Winkler 0001 |
ICIP | 5 |
| 2015 | PET: An eye-tracking dataset for animal-centric Pascal object classesabstractWe present PET- the Pascal animal classes Eye Tracking database. Our database comprises eye movement recordings compiled from forty users for the bird, cat, cow, dog, horse and sheep trainval sets from the VOC 2012 image set. Different from recent eye-tracking databases such as [1, 2], a salient aspect of PET is that it contains eye movements recorded for both the free-viewing and visual search task conditions. While some differences in terms of overall gaze behavior and scanning patterns are observed between the two conditions, a very similar number of fixations are observed on target objects for both conditions. As a utility application, we show how feature pooling around fixated locations enables enhanced (animal) object classification accuracy. Syed Omer Gilani, Subramanian Ramanathan, Yan Yan 0002, David Melcher, Nicu Sebe, Stefan Winkler 0001 |
ICME | 6 |
| 2015 | Deep Learning for Emotion Recognition on Small Datasets using Transfer Learning
Hongwei Ng, Viet Dung Nguyen, Vassilios Vonikakis, Stefan Winkler 0001 |
ICMI | 4 |
| 2015 | Implicit User-centric Personality Recognition Based on Physiological Responses to Emotional VideosabstractWe present a novel framework for recognizing personality traits based on users' physiological responses to affective movie clips. Extending studies that have correlated explicit/implicit affective user responses with Extraversion and Neuroticism traits, we perform single-trial recognition of the big-five traits from Electrocardiogram (ECG), Galvanic Skin Response (GSR), Electroencephalogram (EEG) and facial emotional responses compiled from 36 users using off-the-shelf sensors. Firstly, we examine relationships among personality scales and (explicit) affective user ratings acquired in the context of prior observations. Secondly, we isolate physiological correlates of personality traits. Finally, unimodal and multimodal personality recognition results are presented. Personality differences are better revealed while analyzing responses to emotionally homogeneous (e.g., high valence, high arousal) clips, and significantly above-chance recognition is achieved for all five traits. Julia Wache, Subramanian Ramanathan, Mojtaba Khomami Abadi, Radu L. Vieriu, Nicu Sebe, Stefan Winkler 0001 |
ICMI | 6 |
| 2015 | Design of low-cost, compact and weather-proof whole sky imagers for high-dynamic-range capturesabstractGround-based whole sky imagers are popular for monitoring cloud formations, which is necessary for various applications. We present two new Wide Angle High-Resolution Sky Imaging System (WAHRSIS) models, which were designed especially to withstand the hot and humid climate of Singapore. The first uses a fully sealed casing, whose interior temperature is regulated using a Peltier cooler. The second features a double roof design with ventilation grids on the sides, allowing the outside air to flow through the device. Measurements of temperature inside these two devices show their ability to operate in Singapore weather conditions. Unlike our original WAHRSIS model, neither uses a mechanical sun blocker to prevent the direct sunlight from reaching the camera; instead they rely on high-dynamic-range imaging (HDRI) techniques to reduce the glare from the sun. Soumyabrata Dev, Florian M. Savoy, Yee Hui Lee, Stefan Winkler 0001 |
IGARSS | 4 |
| 2015 | Cloud base height estimation using high-resolution whole sky imagersabstractFine scale cloud monitoring using ground-based imagers is becoming popular for a variety of applications and domains. We present a framework for cloud base height estimation using two such imagers; our method is based on stereoscopic scene flow. We demonstrate the feasibility of our approach and use computer-generated images with controlled cloud height to validate the accuracy of our method. Florian M. Savoy, Joseph Chadi Lemaitre, Soumyabrata Dev, Yee Hui Lee, Stefan Winkler 0001 |
IGARSS | 5 |
| 2015 | Inferring Painting Style with Multi-Task Dictionary Learning
Gaowen Liu, Yan Yan 0002, Elisa Ricci 0001, Yi Yang 0001, Yahong Han, Stefan Winkler 0001, Nicu Sebe |
IJCAI | 6 |
| 2015 | Jointly Estimating Interactions and Head, Body Pose of Interactors from Distant Social ScenesabstractWe present joint estimation of F-formations and head, body pose of interactors in a social scene captured by surveillance cameras. Unlike prior works that have focused on (a) discovering F-formations based on head pose and position cues, or (b) jointly learned head and body pose of individuals based on anatomic constraints, we exploit positional and pose cues characterizing interactors and interactions to jointly infer both (a) and (b). We show how the joint inference framework benefits both F-formation and head, body pose estimation accuracy via experiments on two social datasets. Subramanian Ramanathan, Jagannadan Varadarajan, Elisa Ricci 0001, Oswald Lanz, Stefan Winkler 0001 |
ACM Multimedia | 5 |
| 2014 | Systematic study of color spaces and components for the segmentation of sky/cloud imagesabstractSky/cloud imaging using ground-based Whole Sky Imagers (WSI) is a cost-effective means to understanding cloud cover and weather patterns. The accurate segmentation of clouds in these images is a challenging task, as clouds do not possess any clear structure. Several algorithms using different color models have been proposed in the literature. This paper presents a systematic approach for the selection of color spaces and components for optimal segmentation of sky/cloud images. Using mainly principal component analysis (PCA) and fuzzy clustering for evaluation, we identify the most suitable color components for this task. Soumyabrata Dev, Yee Hui Lee, Stefan Winkler 0001 |
ICIP | 3 |
| 2014 | A data-driven approach to cleaning large face datasetsabstractLarge face datasets are important for advancing face recognition research, but they are tedious to build, because a lot of work has to go into cleaning the huge amount of raw data. To facilitate this task, we describe an approach to building face datasets that starts with detecting faces in images returned from searches for public figures on the Internet, followed by discarding those not belonging to each queried person. We formulate the problem of identifying the faces to be removed as a quadratic programming problem, which exploits the observations that faces of the same person should look similar, have the same gender, and normally appear at most once per image. Our results show that this method can reliably clean a large dataset, leading to a considerable reduction in the work needed to build it. Finally, we are releasing the FaceScrub dataset that was created using this approach. It consists of 141,130 faces of 695 public figures and can be obtained from http://vintage.winklerbros.net/facescrub.html. Hongwei Ng, Stefan Winkler 0001 |
ICIP | 2 |
| 2013 | Impact of image appeal on visual attention during photo triagingabstractImage appeal is determined by factors such as exposure, white balance, motion blur, scene perspective, and semantics. All these factors influence the selection of the best image(s) in a typical photo triaging task. This paper presents the results of an exploratory study on how image appeal affected selection behavior and visual attention patterns of 11 users, who were assigned the task of selecting the best photo from each of 40 groups. Images with low appeal were rejected, while highly appealing images were selected by a majority. Images with higher appeal attracted more visual attention, and users spent more time exploring them. A comparison of user eye fixation maps with three state-of-the-art saliency models revealed that these differences are not captured by the models. Syed Omer Gilani, Subramanian Ramanathan, Huang Hua, Stefan Winkler 0001, Shih-Cheng Yen |
ICIP | 4 |
| 2013 | Stereo/multiview picture quality: Overview and recent advances
Stefan Winkler 0001, Dongbo Min |
Signal Process. Image Commun. | 1 |
| 2012 | Brush-and-drag: a multi-touch interface for photo triagingabstractDue to the convenience of taking pictures with various digital cameras and mobile devices, people often end up with multiple shots of the same scene with only slight variations. To enhance photo triaging, which is a very common photowork activity, we propose an effective and easy-to-use brush-and-drag interface that allows the user to interactively explore and compare photos within a broader scene context. First, we brush to mark an area of interest on a photo with our finger(s); our tailored segmentation engine automatically determines corresponding image elements among the photos. Then, we can drag the segmented elements from different photos across the screen to explore them simultaneously, and further perform simple finger gestures to interactively rank photos, select favorites for sharing, or to remove unwanted ones. This novel interaction method was implemented on a consumer-level tablet computer and demonstrated to offer effective interactions in a user study. Seon Joo Kim, Hongwei Ng, Stefan Winkler 0001, Peng Song 0001, Chi-Wing Fu |
Mobile HCI | 3 |
| 2012 | System for creating slideshows based on people and their emotionsabstractWe demonstrate a system for the automatic creation of slideshows from photo collections, based on a user-specified group of people, their emotions and other image similarity criteria such as color, timeline and scene characteristics. The main objective of the system is to form meaningful image sequences for slideshows, with minimal user input. The intuitive interface allows the user to provide a personalized definition of image similarity, resulting in different image sequences, and making the system highly user-adaptable and flexible. Vassilios Vonikakis, Stefan Winkler 0001 |
ACM Multimedia | 2 |
| 2012 | Emotion-based sequence of family photosabstractThis paper presents a method for the automatic creation of slideshows from family photo collections based on the emotions of a given group of people. The user specifies the desired person(s) to be included in the slideshow. A natural image sequence is formed based on people's emotions and several other, user-defined image similarity attributes, in order to form meaningful slideshow transitions. This process makes use of a new image dissimilarity function, which can integrate various attribute combinations and preferences, making the system highly user-adaptable and flexible. Vassilios Vonikakis, Stefan Winkler 0001 |
ACM Multimedia | 2 |
| 2011 | Effects of Rain Attenuation on Satellite Video TransmissionabstractHeavy convective rain events are often experienced in tropical countries such as Singapore. The operation of high-speed satellite transmission in the Ka-band is therefore susceptible to attenuation. We present the setup of a high-speed link via the WINDS satellite using an ultra small aperture terminals (USAT). Video streaming is performed over the high-speed link so as to investigate the effects of rainfall on the signal strength of the link and the video quality. The video streaming is performed at two locations 40km apart in order to examine site diversity as a mitigation technique. Yee Hui Lee, Stefan Winkler 0001 |
VTC Spring | 2 |
| 2010 | Special issue on Image and Video Quality Assessment
Stefan Winkler 0001, Frédéric Dufaux, Dominique Barba, Vittorio Baroncini |
Signal Process. Image Commun. | 1 |
| 2010 | Immersive Multiplayer Games With Tangible and Physical InteractionabstractIn this paper, we present a new immersive multiplayer game system developed for two different environments, namely, virtual reality (VR) and augmented reality (AR). To evaluate our system, we developed three game applications-a first-person-shooter game (for VR and AR environments, respectively) and a sword game (for the AR environment). Our immersive system provides an intuitive way for users to interact with the VR or AR world by physically moving around the real world and aiming freely with tangible objects. This encourages physical interaction between players as they compete or collaborate with other players. Evaluation of our system consists of users' subjective opinions and their objective performances. Our design principles and evaluation results can be applied to similar immersive game applications based on AR/VR. Jefry Tedjokusumo, Steven Zhiying Zhou, Stefan Winkler 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2008 | A Hybrid Framework for 3-D Human Motion TrackingabstractIn this paper, we present a hybrid framework for articulated 3-D human motion tracking from multiple synchronized cameras with potential uses in surveillance systems. Although the recovery of 3-D motion provides richer information for event understanding, existing methods based on either deterministic search or stochastic sampling lack robustness or efficiency. We therefore propose a hybrid sample-and-refine framework that combines both stochastic sampling and deterministic optimization to achieve a good compromise between efficiency and robustness. Similar motion patterns are used to learn a compact low-dimensional representation of the motion statistics. Sampling in a low-dimensional space is implemented during tracking, which reduces the number of particles drastically. We also incorporate a local optimization method based on simulated physical force/moment into our framework, which further improves the optimality of the tracking. Experimental results on several real human motion sequences show the accuracy and robustness of our method, which also has a higher sampling efficiency than most particle filtering-based methods. Bingbing Ni, Ashraf A. Kassim, Stefan Winkler 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Perceived Audiovisual Quality of Low-Bitrate Multimedia ContentabstractThis paper studies the quality of multimedia content at very low bitrates. We carried out subjective experiments for assessing audiovisual, audio-only, and video-only quality. We selected content and encoding parameters that are typical of mobile applications. Our focus were the MPEG-4 AVC (a.k.a. H.264) and AAC coding standards. Based on these data, we first analyze the influence of video and audio coding parameters on quality. We investigate the optimal trade-off between bits allocated to audio and to video under global bitrate constraints. Finally, we explore models for the interactions between audio and video in terms of perceived audiovisual quality Stefan Winkler 0001, Christof Faller |
IEEE Trans. Multim. | 1 |
| 2004 | Segmentation-driven perceptual quality metricsabstractWe present a full-reference and a no-reference perceptual video quality metric that incorporate both low-level and high-level aspects of vision. Low-level aspects include color perception, contrast sensitivity, masking as well as artifact analysis. High-level aspects take into account the cognitive behavior of an observer when watching a video by means of semantic segmentation. Using the special case of semantic face segmentation, we evaluate the proposed segmentation-driven perceptual quality metrics using a range of test sequences and demonstrate an improvement of their prediction performance. Andrea Cavallaro, Stefan Winkler 0001 |
ICIP | 2 |
| 2004 | Perceptual blur and ringing metrics: application to JPEG2000
Pina Marziliano, Frédéric Dufaux, Stefan Winkler 0001, Touradj Ebrahimi |
Signal Process. Image Commun. | 3 |
| 2003 | Video quality evaluation for mobile streaming applications
Stefan Winkler 0001, Frédéric Dufaux |
VCIP | 1 |
| 2002 | A no-reference perceptual blur metricabstractWe present a no-reference blur metric for images and video. The blur metric is based on the analysis of the spread of the edges in an image. Its perceptual significance is validated through subjective experiments. The novel metric is near real-time, has low computational complexity and is shown to perform well over a range of image content. Potential applications include optimization of source coding, network resource management and autofocus of an image capturing device. Pina Marziliano, Frédéric Dufaux, Stefan Winkler 0001, Touradj Ebrahimi |
ICIP (3) | 3 |
| 2002 | Vision-model-based impairment metric to evaluate blocking artifacts in digital videoabstractIn this paper investigations are conducted to simplify and refine a vision-model-based video quality metric without compromising its prediction accuracy. Unlike other vision-model-based quality metrics, the proposed metric is parameterized using subjective quality assessment data recently provided by the Video Quality Experts Group. The quality metric is able to generate a perceptual distortion map for each and every video frame. A perceptual blocking distortion metric (PBDM) is introduced which utilizes this simplified quality metric. The PBDM is formulated based on the observation that blocking artifacts are noticeable only in certain regions of a picture. A method to segment blocking dominant regions is devised, and perceptual distortions in these regions are summed up to form an objective measure of blocking artifacts. Subjective and objective tests are conducted and the performance of the PBDM is assessed by a number of measures such as the Spearman rank-order correlation, the Pearson correlation, and the average absolute error The results show a strong correlation between the objective blocking ratings and the mean opinion scores on blocking artifacts. Zhenghua Yu, Hong Ren Wu, Stefan Winkler 0001, Tao Chen 0044 |
Proc. IEEE | 3 |
| 2002 | A vision-based masking model for spread-spectrum image watermarkingabstractWe present a perceptual model for hiding a spread-spectrum watermark of variable amplitude and density in an image. The model takes into account the sensitivity and masking behavior of the human visual system by means of a local isotropic contrast measure and a masking model. We compare the insertion of this watermark in luminance images and in the blue channel of color images. We also evaluate the robustness of such a watermark with respect to its embedding density. Our results show that this approach facilitates the insertion of a more robust watermark while preserving the visual quality of the original. Furthermore, we demonstrate that the maximum watermark density generally does not provide the best detection performance. Martin Kutter, Stefan Winkler 0001 |
IEEE Trans. Image Process. | 2 |
| 2000 | Video Quality Experts Group: current results and future directions
Ann M. Rohaly, Philip J. Corriveau, John M. Libert, Arthur A. Webster, Vittorio Baroncini, John G. Beerends, Jean-Louis Blin, Laura Contin, Takahiro Hamada, David Harrison, Andries P. Hekstra, Jeffrey Lubin, Yukihiro Nishida, Ricardo Nishihara, John C. Pearson, Antonio F. Pessoa, Neil Pickford, Alexander Schertz, Massimo Visca, Andrew B. Watson, Stefan Winkler 0001 |
VCIP | 21 |
| 1999 | Computing Isotropic Local Contrast from Oriented Pyramid DecompositionsabstractWorking with contrast instead of luminance can facilitate numerous image processing and analysis tasks. Unfortunately, a common definition of contrast suitable for all situations does not exist. We review existing contrast definitions for natural images and propose a new isotropic contrast measure, which is computed from oriented filters. We investigate some of its properties and apply it to natural images. Stefan Winkler 0001, Pierre Vandergheynst |
ICIP (4) | 1 |
| 1999 | Visual quality assessment using a contrast gain control modelabstractMuch of the work on visual quality assessment has been devoted to gray-level images; metrics taking into account color information and the temporal component are still relatively rare. This paper presents a quality metric for color video which is based on a recent contrast gain control model of the human visual system. It is used to assess the quality of MPEG-coded sequences and exhibits a behavior that is consistent with subjective ratings. Stefan Winkler 0001 |
MMSP | 1 |
| 1999 | Issues in vision modeling for perceptual video quality assessment
Stefan Winkler 0001 |
Signal Process. | 1 |
| 1998 | A Perceptual Distortion Metric for Digital Color Images
Stefan Winkler 0001 |
ICIP (3) | 1 |