EDBT 2026 Demo / reviewers in the wild / expert
Pavel Korshunov
dblp:84/53
· DBLP profile ↗
29ranked-venue papers
16as first author
12since 2021 · last 2025
0000-0002-2244-1830ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 16 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Security and privacy · 5 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic ClaimsabstractMisinformation can be countered with factchecking, but the process is costly and slow.Identifying checkworthy claims is the first step, where automation can help scale fact-checkers' efforts.However, detection methods struggle with content that is (1) multimodal, (2) from diverse domains, and (3) synthetic.We introduce HINTSOFTRUTH, a public dataset for multimodal checkworthiness detection with 27K real-world and synthetic image/claim pairs.The mix of real and synthetic data makes this dataset unique and ideal for benchmarking detection methods.We compare fine-tuned and prompted Large Language Models (LLMs).We find that well-configured lightweight textbased encoders perform comparably to multimodal models but the former only focus on identifying non-claim-like content.Multimodal LLMs can be more accurate but come at a significant computational cost, making them impractical for large-scale applications.When faced with synthetic data, multimodal models perform more robustly. Michiel van der Meer, Pavel Korshunov, Sébastien Marcel, Lonneke van der Plas |
ACL (1) | 2 |
| 2025 | Investigation of accuracy and bias in face recognition trained with synthetic dataabstractSynthetic data has emerged as a promising alternative for training face recognition (FR) models, offering advantages in scalability, privacy compliance, and potential for bias mitigation. However, critical questions remain on whether both high accuracy and fairness can be achieved with synthetic data. In this work, we evaluate the impact of synthetic data on bias and performance of FR systems. We generate balanced face dataset, FairFaceGen, using two state of the art text-to-image generators, Flux.1-dev and Stable Diffusion v3.5 (SD35), and combine them with several identity augmentation methods, including Arc2Face and four IP-Adapters. By maintaining equal identity count across synthetic and real datasets, we ensure fair comparisons when evaluating FR performance on standard (LFW, AgeDB-30, etc.) and challenging IJB-B/C benchmarks and FR bias on Racial Faces in-the-Wild (RFW) dataset. Our results demonstrate that although synthetic data still lags behind the real datasets in the generalization on IJB-B/C, demographically balanced synthetic datasets, especially those generated with SD35, show potential for bias mitigation. We also observe that the number and quality of intra-class augmentations significantly affect FR accuracy and fairness. These findings provide practical guidelines for constructing fairer FR systems using synthetic data. Pavel Korshunov, Ketan Kotwal, Christophe Ecabert, Vidit Amir Mohammadi, Sébastien Marcel |
IJCB | 1 |
| 2025 | FantasyID: A Dataset for Detecting Digital Manipulations in ID-DocumentsabstractAdvancements in image generation led to the availability of easy-to-use tools for malicious actors to create forged images. These tools pose a serious threat to the widespread Know Your Customer (KYC) applications, requiring robust systems for detection of the forged Identity Documents (IDs). To facilitate the development of the detection algorithms, in this paper, we propose a novel publicly available (including commercial use) dataset, FantasyID, which mimics real-world IDs but without tampering with legal documents and, compared to previous public datasets, it does not contain generated faces or specimen watermarks. FantasyID contains ID cards with diverse design styles, languages, and faces of real people. To simulate a realistic KYC scenario, the cards from FantasyID were printed and captured with three different devices, constituting the bonafide class. We have emulated digital forgery/injection attacks that could be performed by a malicious actor to tamper the IDs using the existing generative tools. The current state-of-the-art forgery detection algorithms, such as TruFor, MMFusion, UniFD, and FatFormer, are challenged by FantasyID dataset. It especially evident, in the evaluation conditions close to practical, with the operational threshold set on validation set so that false positive rate is at 10%, leading to false negative rates close to 50% across the board on the test set. The evaluation experiments demonstrate that FantasyID dataset is complex enough to be used as an evaluation benchmark for detection algorithms. Pavel Korshunov, Amir Mohammadi, Vidit Vidit, Christophe Ecabert, Sébastien Marcel |
IJCB | 1 |
| 2025 | Identity-Preserving Aging and De-Aging of Faces in the StyleGAN Latent SpaceabstractFace aging or de-aging with generative AI has gained significant attention for its applications in such fields like forensics, security, and media. However, most state of the art methods rely on conditional Generative Adversarial Networks (GANs), Diffusion-based models, or Visual Language Models (VLMs) to age or de-age faces based on predefined age categories and conditioning via loss functions, fine-tuning, or text prompts. The reliance on such conditioning leads to complex training requirements, increased data needs, and challenges in generating consistent results. Additionally, identity preservation is rarely taken into account or evaluated on a single face recognition system without any control or guarantees on whether identity would be preserved in a generated aged/de-aged face. In this paper, we propose to synthesize aged and de-aged faces via editing latent space of StyleGAN2 using a simple support vector modeling of aging/de-aging direction and several feature selection approaches. By using two state-of-the-art face recognition systems, we empirically find the identity preserving subspace within the StyleGAN2 latent space, so that an apparent age of a given face can changed while preserving the identity. We then propose a simple yet practical formula for estimating the limits on aging/de-aging parameters that ensures identity preservation for a given input face. Using our method and estimated parameters we have generated a public dataset of synthetic faces at different ages that can be used for benchmarking cross-age face recognition, age assurance systems, or systems for detection of synthetic images. Our code and dataset are available at the project page https://www.idiap.ch/paper/agesynth/ Luis S. Luevano, Pavel Korshunov, Sébastien Marcel |
IJCB | 2 |
| 2024 | Vulnerability of Face age Verification to Replay AttacksabstractPresentation attacks on biometric systems have long created significant security risks. The increase in the adoption of age verification systems, which ensure that only age-appropriate content is consumed online, raises the question of vulnerability of such systems to replay presentation attacks. In this paper, we analyze the vulnerability of face age verification to simple replay attacks and assess whether presentation attack detection (PAD) systems created for biometrics can be effective at detecting similar attacks on age verification. We used three types of attacks captured with iPhone 12, Galaxy S9, and Huawei Mate 30 phones from iPad Pro, which replayed the images from a commonly used UTKFace dataset of faces with true age labels. We evaluated four state of the art face age verification algorithms, including simple classification, distribution-based, regression via classification, and adaptive distribution approaches. We show that these algorithms are vulnerable to the attacks, since the accuracy of age verification on replayed images is only a couple of percentage points different compared to when the original images are used, which means an age verification system cannot distinguish attacks from bona fide images. Using two state of the art presentation attack detection systems, DeepPixBiS and CDCN, trained to detect similar attacks on biometrics, we demonstrate that they struggle to detect both: the types of attacks that are possible in age verification scenario and the type of bona fide images that are commonly used. These results highlight the need for the development of age verification specific attack detection systems for age verification to become practical. Pavel Korshunov, Anjith George, Gökhan Özbulak, Sébastien Marcel |
ICASSP | 1 |
| 2023 | Vulnerability of Automatic Identity Recognition to Audio-Visual DeepfakesabstractThe task of deepfakes detection is far from being solved by speech or vision researchers. Several publicly available databases of fake synthetic video and speech were built to aid the development of detection methods. However, existing databases typically focus on visual or voice modalities and provide no proof that their deepfakes can in fact impersonate any real person. In this paper, we present the first realistic audio-visual database of deepfakes SWAN-DF, where lips and speech are well synchronized and video have high visual and audio qualities. We took the publicly available SWAN dataset of real videos with different identities to create audio-visual deepfakes using several models from DeepFaceLab and blending techniques for face swapping and HiFiVC, DiffVC, YourTTS, and FreeVC models for voice conversion. From the publicly available speech dataset LibriTTS, we also created a separate database of only audio deepfakes LibriTTS-DF using several latest text to speech methods: YourTTS, Adaspeech, and TorToiSe. We demonstrate the vulnerability of a state of the art speaker recognition system, such as ECAPA-TDNN-based model from SpeechBrain, to the synthetic voices. Similarly, we tested face recognition system based on the MobileFaceNet architecture to several variants of our visual deepfakes. The vulnerability assessment show that by tuning the existing pretrained deepfake models to specific identities, one can successfully spoof the face and speaker recognition systems in more than 90% of the time and achieve a very realistic looking and sounding fake video of a given person. Pavel Korshunov, Philip N. Garner, Sébastien Marcel |
IJCB | 1 |
| 2022 | Custom Attribution Loss for Improving Generalization and Interpretability of Deepfake DetectionabstractThe simplicity and accessibility of tools for generating deepfakes pose a significant technical challenge for their detection and filtering. Many of the recently proposed methods for deeptake detection focus on a ‘blackbox’ approach and therefore suffer from the lack of any additional information about the nature of fake videos beyond the fake or not fake labels. In this paper, we approach deepfake detection by solving the related problem of attribution, where the goal is to distinguish each separate type of a deepfake attack. We design a training approach with customized Triplet and ArcFace losses that allow to improve the accuracy of deepfake detection on several publicly available datasets, including Google and Jigsaw, FaceForensics++, HifiFace, DeeperForensics, Celeb-DF, DeepfakeTIMIT, and DF-Mobio. Using an example of Xception net as an underlying architecture, we also demonstrate that when trained for attribution, the model can be used as a tool to analyze the deepfake space and to compare it with the space of original videos. Pavel Korshunov, Anubhav Jain 0002, Sébastien Marcel |
ICASSP | 1 |
| 2022 | Are GAN-based morphs threatening face recognition?abstractMorphing attacks are a threat to biometric systems where the biometric reference in an identity document can be altered. This form of attack presents an important issue in applications relying on identity documents such as border security or access control. Research in generation of face morphs and their detection is developing rapidly, however very few datasets with morphing attacks and open-source detection toolkits are publicly available. This paper bridges this gap by providing two datasets and the corresponding code for four types of morphing attacks: two that rely on facial landmarks based on OpenCV and FaceMorpher, and two that use StyleGAN 2 to generate synthetic morphs. We also conduct extensive experiments to assess the vulnerability of four state-of-the-art face recognition systems, including FaceNet, VGG-Face, ArcFace, and ISV. Surprisingly, the experiments demonstrate that, although visually more appealing, morphs based on StyleGAN 2 do not pose a significant threat to the state to face recognition systems, as these morphs were outmatched by the simple morphs that are based facial landmarks. Eklavya Sarkar, Pavel Korshunov, Laurent Colbois, Sébastien Marcel |
ICASSP | 2 |
| 2022 | Face Anthropometry Aware Audio-visual Age VerificationabstractProtection of minors against destructive content or illegal advertising is an important problem, which is now under increasing societal and legislative pressure. The latest advancements in an automated age verification is a possible solution to this problem. There are however limitations of the current state of the art age verification methods, specifically, the lack of approaches focusing on video-based or even solely audio-based approaches, since the image domain is the one with the majority of publicly available datasets. In this paper, we consider the problem of age verification as a multimodal problem by proposing and evaluating several audio- and image-based models and their combinations. To that end, we annotated a set of publicly available videos with age labels, with a special focus on the children age labels. We also propose a new training strategy based on the adaptive label distribution learning (ALDL), which is driven by facial anthropometry and age-based skin degradation. This adaptive approach demonstrates the best accuracy when evaluated across several test databases. Pavel Korshunov, Sébastien Marcel |
ACM Multimedia | 1 |
| 2021 | Subjective and Objective Evaluation of Deepfake VideosabstractPractically anyone can now generate a realistic looking deepfake video. It is clear that the online prevalence of such fake videos will erode the societal trust in video evidence even further. To counter the looming threat, many methods to detect deepfakes were recently proposed by the research community. However, it is still unclear how realistic deep-fake videos are for an average person and whether the algorithms are significantly better than humans at detecting them. Therefore, this paper, presents a subjective study, which, using 60 naïve subjects, evaluates how hard it is for humans to see if a video is a deepfake or not. For the study, 120 videos (60 deepfakes and 60 originals) were manually selected from the Facebook database used in Kaggle’s Deepfake Detection Challenge 2020. The results of the subjective evaluation were compared with two state of the art deepfake detection methods, based on Xception and EfficientNet (B4 variant) neural network models pre-trained on two other public databases: Google and Jiqsaw subset from FaceForensics++ and Celeb-DF v2 dataset. The experiments demonstrate that while the human perception is very different from the perception of a machine, both successfully but in different ways are fooled by deepfakes. Specifically, algorithms struggle to detect the deepfake videos that humans find to be very easy to spot. Pavel Korshunov, Sébastien Marcel |
ICASSP | 1 |
| 2021 | ADGD'21: 1st Workshop on Synthetic Multimedia - Audiovisual Deepfake Generation and DetectionabstractDeepfakes, i.e.synthetic or "fake" media content generated using deep learning, are a double-edged sword. On one hand, they pose new threats and risks in the form of scams, fraud, disinformation, social manipulation, or celebrity porn. On the other hand, deepfakes have just as many meaningful and beneficial applications - they allow us to create and experience things that no longer exist, or that have never existed, enabling numerous exciting applications in entertainment, education, and even privacy. Stefan Winkler 0001, Abhinav Dhall, Pavel Korshunov |
ACM Multimedia | 4 |
| 2021 | Improving Generalization of Deepfake Detection by Training for AttributionabstractRecent advances in automated video and audio editing tools, generative adversarial networks (GANs), and social media allow the creation and fast dissemination of high-quality tampered videos, which are commonly called deepfakes. Typically, in these videos, a face is automatically swapped with the face of another person. The simplicity and accessibility of tools for generating deepfakes pose a significant technical challenge for their detection and filtering. In response to the threat, several large datasets of deepfake videos and various methods to detect them were proposed recently. However, the proposed methods suffer from the problem of over-fitting on the training data and the lack of generalization across different databases and generative approaches. In this paper, we approach deepfake detection by solving the related problem of attribution, where the goal is to distinguish each separate type of a deepfake attack. Using publicly available datasets from Google and Jigsaw, FaceForensics++, Celeb-DF, DeepfakeTIMIT, and our own large database DF-Mobio, we demonstrate that an XceptionNet and EfficientNet based models trained for an attribution task generalize better to unseen deepfakes and different datasets, compared to the same models trained for a typical binary classification task. We also demonstrate that by training for attribution with a triplet-loss, the generalization in cross-database scenario improves even more, compared to the binary system, while the performance on the same database degrades only marginally. Anubhav Jain 0002, Pavel Korshunov, Sébastien Marcel |
MMSP | 2 |
| 2020 | Pyannote.Audio: Neural Building Blocks for Speaker DiarizationabstractWe introduce pyannote.audio, an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be combined and jointly optimized to build speaker diarization pipelines. pyannote.audio also comes with pre-trained models covering a wide range of domains for voice activity detection, speaker change detection, overlapped speech detection, and speaker embedding - reaching state-of-the-art performance for most of them. Hervé Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, Marie-Philippe Gill |
ICASSP | 5 |
| 2017 | Long-Term Spectral Statistics for Voice Presentation Attack DetectionabstractAutomatic speaker verification systems can be spoofed through recorded, synthetic, or voice converted speech of target speakers. To make these systems practically viable, the detection of such attacks, referred to as presentation attacks, is of paramount interest. In that direction, this paper investigates two aspects: 1) a novel approach to detect presentation attacks where, unlike conventional approaches, no speech signal modeling related assumptions are made, rather the attacks are detected by computing first-order and second-order spectral statistics and feeding them to a classifier, and 2) generalization of the presentation attack detection systems across databases. Our investigations on ASVspoof 2015 challenge database and AVspoof database show that, when compared to the approaches based on conventional short-term spectral features, the proposed approach with a linear discriminative classifier yields a better system, irrespective of whether the spoofed signal is replayed to the microphone or is directly injected into the system software process. Cross-database investigations show that neither the short-term spectral processing-based approaches nor the proposed approach yield systems which are able to generalize across databases or methods of attack. Thus, revealing the difficulty of the problem and the need for further resources and research. Hannah Muckenhirn, Pavel Korshunov, Mathew Magimai-Doss, Sébastien Marcel |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Cross-Database Evaluation of Audio-Based Spoofing Detection SystemsabstractSince automatic speaker verification (ASV) systems are highly vulnerable to spoofing attacks, it is important to develop mechanisms that can detect such attacks. To be practical, however, a spoofing attack detection approach should have (i) high accuracy, (ii) be well-generalized for practical attacks, and (iii) be simple and efficient. Several audio-based spoofing detection methods have been proposed recently but their evaluation is limited to less realistic databases containing homogeneous data. In this paper, we consider eight existing presentation attack detection (PAD) methods and evaluate their performance using two major publicly available speaker databases with spoofing attacks: AVspoof and ASVspoof. We first show that realistic presentation attacks (speech is replayed to PAD system) are significantly more challenging for the considered PAD methods compared to the so called `logical access' attacks (speech is presented to PAD system directly). Then, via a cross-database evaluation, we demonstrate that the existing methods generalize poorly when different databases or different types of attacks are used for training and testing. The results question the efficiency and practicality of the existing PAD systems, as well as, call for creation of databases with larger variety of realistic speech presentation attacks. Pavel Korshunov, Sébastien Marcel |
INTERSPEECH | 1 |
| 2015 | Privacy in mini-drone based video surveillanceabstractMini-drones are increasingly used in video surveillance. Their areal mobility and ability to carry video cameras provide new perspectives in visual surveillance which can impact privacy in ways that have not been considered in a typical surveillance scenario. To better understand and analyze them, we have created a publicly available video dataset of typical drone-based surveillance sequences in a car parking. Using the sequences from this dataset, we have assessed five privacy protection filters via a crowdsourcing evaluation. We asked crowdsourcing workers several privacy- and surveillance-related questions to determine the tradeoff between intelligibility of the scene and privacy, and we present conclusions of this evaluation in this paper. Margherita Bonetto, Pavel Korshunov, Giovanni Ramponi, Touradj Ebrahimi |
ICIP | 2 |
| 2014 | Crowd-based quality assessment of multiview video plus depth codingabstractCrowdsourcing is becoming a popular cost effective alternative to lab-based evaluations for subjective quality assessment. However, crowd-based evaluations are constrained by the limited availability of display devices used by typical online workers, which makes the evaluation of 3D content a challenging task. In this paper, we investigate two possible approaches to crowd-based quality assessment of multiview video plus depth (MVD) content on 2D displays: by using a virtual view and by using a free-viewpoint video, which corresponds to a smooth camera motion during a time freeze. We conducted the crowdsourcing experiments using seven MVD sequences encoded at different bit rates with the upcoming 3D-AVC video coding standard. The results demonstrate high correlation with subjective evaluations performed using a stereoscopic monitor in a controlled laboratory environment. The analysis shows no statistically significant difference between the two approaches. Philippe Hanhart, Pavel Korshunov, Touradj Ebrahimi |
ICIP | 2 |
| 2014 | Scrambling-based tool for secure protection of JPEG imagesabstractJPEG scrambling tool is a flexible web-based tool with an intuitive and simple to use GUI and interface to secure visual information in a region of interest (ROI) of JPEG images. The tool demonstrates an efficient integration and use of security tools in JPEG image format by an example of scrambling privacy filter, which enables a variety of security services such as confidentiality, integrity verification, source authentication, and conditional access. Pavel Korshunov, Touradj Ebrahimi |
ICIP | 1 |
| 2014 | Towards optimal distortion-based visual privacy filtersabstractThe widespread usage of digital video surveillance systems has increased the concerns for privacy violation. Since video surveillance systems are invasive, it is a challenge to find an acceptable balance between privacy of the public under surveillance and security related features of the systems. Many privacy protection tools have been proposed for preserving privacy, ranging from such simple methods like blurring or pixelization to more advanced like scrambling and geometrical transform based filters. However, for a given filter implemented in a practical video surveillance system, it is necessary to know the strength with which the filter should be applied to protect privacy reliably. Assuming an automated surveillance system, this paper objectively investigates several privacy protection filters with varying strength degrees and determines their optimal strength values to achieve privacy protection. To this end, five privacy filters were applied to images from FERET dataset and the performance of three recognition algorithms was evaluated. The results show that different privacy protection filters influence the accuracy of different versions of face recognition differently and this influence depends both on the robustness of the recognition and the type of distortion filter. Pavel Korshunov, Touradj Ebrahimi |
ICIP | 1 |
| 2014 | Impact of Ultra High Definition on Visual AttentionabstractUltra high definition (UHD) TV is rapidly replacing high definition (HD) TV but little is known of its effects on human visual attention. However, a clear understanding of this effect is important, since accurate models, evaluation methodologies, and metrics for visual attention are essential in many areas, including image and video compression, camera and displays manufacturing, artistic content creation, and advertisement. In this paper, we address this problem by creating a dataset of UHD resolution images with corresponding eye-tracking data, and we show that there is a statistically significant difference between viewing strategies when watching UHD and HD contents. Furthermore, by evaluating five representative computational models of visual saliency, we demonstrate the decrease in models' accuracies on UHD contents when compared to HD contents. Therefore, to improve the accuracy of computational models for higher resolutions, we propose a segmentation-based resolution-adaptive weighting scheme. Our approach demonstrates that taking into account information about resolution of the images improves the performance of computational models. Hiromi Nemoto, Philippe Hanhart, Pavel Korshunov, Touradj Ebrahimi |
ACM Multimedia | 3 |
| 2014 | Survey of web-based crowdsourcing frameworks for subjective quality assessmentabstractThe popularity of the crowdsourcing for performing various tasks online increased significantly in the past few years. The low cost and flexibility of crowdsourcing, in particular, attracted researchers in the field of subjective multimedia evaluations and Quality of Experience (QoE). Since online assessment of multimedia content is challenging, several dedicated frameworks were created to aid in the designing of the tests, including the support of the testing methodologies like ACR, DCR, and PC, setting up the tasks, training sessions, screening of the subjects, and storage of the resulted data. In this paper, we focus on the web-based frameworks for multimedia quality assessments that support commonly used crowdsourcing platforms such as Amazon Mechanical Turk and Microworkers. We provide a detailed overview of the crowdsourcing frameworks and evaluate them to aid researchers in the field of QoE assessment in the selection of frameworks and crowdsourcing platforms that are adequate for their experiments. Tobias Hoßfeld, Matthias Hirth, Pavel Korshunov, Philippe Hanhart, Bruno Gardlo, Christian Keimel, Christian Timmerer |
MMSP | 3 |
| 2014 | Performance evaluation of the emerging JPEG XT image compression standardabstractThe upcoming JPEG XT is under development for High Dynamic Range (HDR) image compression. This standard encodes a Low Dynamic Range (LDR) version of the HDR image generated by a Tone-Mapping Operator (TMO) using the conventional JPEG coding as a base layer and encodes the extra HDR information in a residual layer. This paper studies the performance of the three profiles of JPEG XT (referred to as profiles A, B and C) using a test set of six HDR images. Four TMO techniques were used for the base layer image generation to assess the influence of the TMOs on the performance of JPEG XT profiles. Then, the HDR images were coded with different quality levels for the base layer and for the residual layer. The performance of each profile was evaluated using Signal to Noise Ratio (SNR), Feature SIMilarity Index (FSIM), Root Mean Square Error (RMSE), and CIEDE2000 color difference objective metrics. The evaluation results demonstrate that profiles A and B lead to similar saturation of quality at the higher bit rates, while profile C exhibits no saturation. Profiles B and C appear to be more dependent on TMOs used for the base layer compared to profile A. António M. G. Pinheiro, Karel Fliegel, Pavel Korshunov, Lukas Krasula, Marco V. Bernardo, Maria Pereira, Touradj Ebrahimi |
MMSP | 3 |
| 2013 | Using face morphing to protect privacyabstractThe widespread use of digital video surveillance systems has also increased the concerns for violation of privacy rights. Since video surveillance systems are invasive, it is a challenge to find an acceptable balance between privacy of the public under surveillance and the functionalities of the systems. Tools for protection of visual privacy available today lack either all or some of the important properties such as security of protected visual data, reversibility (ability to undo privacy protection), simplicity, and independence from the video encoding used. To overcome these shortcomings, in this paper, we propose a morphing-based privacy protection method and focus on its robustness, reversibility, and security properties. We morph faces from a standard FERET dataset and run face detection and recognition algorithms on the resulted images to demonstrate that morphed faces retain the likeness of a face, while making them unrecognizable, which ensures the protection of privacy. Our experiments also demonstrate the influence of morphing strength on robustness and security. We also show how to determine the right parameters of the method. Pavel Korshunov, Touradj Ebrahimi |
AVSS | 1 |
| 2013 | Comparative Study of Trust Modeling for Automatic Landmark TaggingabstractMany images uploaded to social networks are related to travel, since people consider traveling to be an important event in their life. However, a significant amount of travel images on the Internet lack proper geographical annotations or tags. In many cases, the images are tagged manually. One way to make this time-consuming manual tagging process more efficient is to propagate tags from a small set of tagged images to the larger set of untagged images automatically. In this paper, we present a system for automatic geotag propagation in images based on the similarity between image content (famous landmarks) and its context (associated geotags). In such a scenario, however, an incorrect or a spam tag can damage the integrity and reliability of the automated propagation system. Therefore, for reliable geotags propagation, we suggest adopting a user trust model based on social feedback from the users of the photo-sharing system. We compare this socially-driven approach with other user trust models via experiments and subjective testing on an image database of various famous landmarks. Results demonstrate that relying on user feedback is more efficient, since the number of propagated tags more than doubles without loss of accuracy compared to using other models or propagating without trust modeling. Peter Vajda, Pavel Korshunov, Touradj Ebrahimi |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Subjective study of privacy filters in video surveillanceabstractExtensive adoption of video surveillance, affecting many aspects of the daily life, alarms the concerned public about the increasing invasion into personal privacy. Therefore, to address privacy issues, many tools have been proposed for protection of personal privacy in image and video. However, little is understood regarding the effectiveness of such tools and especially their impact on the underlying surveillance tasks. In this paper, we propose a subjective evaluation methodology to analyze the tradeoff between the preservation of privacy offered by these tools and the intelligibility of activities under video surveillance. As an example, the proposed method is used to compare several commonly employed privacy protection techniques, such as blurring, pixelization, and masking applied to indoor surveillance video. The results show that, for the test material under analysis, the pixelization filter provides the best performance in terms of balance between privacy protection and intelligibility. Pavel Korshunov, Claudia Araimo, Francesca De Simone, Carmelo Velardo, Jean-Luc Dugelay, Touradj Ebrahimi |
MMSP | 1 |
| 2011 | Video quality for face detection, recognition, and trackingabstractMany distributed multimedia applications rely on video analysis algorithms for automated video and image processing. Little is known, however, about the minimum video quality required to ensure an accurate performance of these algorithms. In an attempt to understand these requirements, we focus on a set of commonly used face analysis algorithms. Using standard datasets and live videos, we conducted experiments demonstrating that the algorithms show almost no decrease in accuracy until the input video is reduced to a certain critical quality, which amounts to significantly lower bitrate compared to the quality commonly acceptable for human vision. Since computer vision percepts video differently than human vision, existing video quality metrics, designed for human perception, cannot be used to reason about the effects of video quality reduction on accuracy of video analysis algorithms. We therefore investigate two alternate video quality metrics, blockiness and mutual information, and show how they can be used to estimate the critical video qualities for face analysis algorithms. Pavel Korshunov, Wei Tsang Ooi |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2010 | Reducing Frame Rate for Object Tracking
Pavel Korshunov, Wei Tsang Ooi |
MMM | 1 |
| 2006 | Rate-accuracy tradeoff in automated, distributed video surveillance systemsabstractKeywords: Video Surveillance ; Video Analysis ; Video Features ; Rate-Accuracy Tradeoff Reference EPFL-CONF-176952 Record created on 2012-05-09, modified on 2017-05-10 Pavel Korshunov |
ACM Multimedia | 1 |
| 2005 | Critical video quality for distributed automated video surveillanceabstractLarge-scale distributed video surveillance systems pose new scalability challenges. Due to the large number of video sources in such systems, the amount of bandwidth required to transmit video streams for monitoring often strains the capability of the network. On the other hand, large-scale surveillance systems often rely on computer vision algorithms to automate surveillance tasks. We observe that these surveillance tasks present an opportunity for trade-off between the accuracy of the tasks and the bit rate of the video being sent. This paper shows that there exists a sweet spot, which we term critical video quality that can be used to reduce video bit rate without significantly affecting the accuracy of the surveillance tasks. We demonstrate this point by running extensive experiments on standard face detection and face tracking algorithms. Our experiments show that face detection works equally well even if the quality of compression is significantly reduced, and face tracking still works even if the frame rate is reduced to 6 frames per second. We further develop a prototype video surveillance system to demonstrate this idea. Our evaluation shows that we can achieve up to 29 times reduction in video bit rate when detecting faces and 16 times reduction when tracking faces. This paper also proposes a formal rate-accuracy optimization framework which can be used to determine appropriate encoding parameters in distributed video surveillance systems that are subjected to either bandwidth constraints or accuracy constraints. Pavel Korshunov, Wei Tsang Ooi |
ACM Multimedia | 1 |