EDBT 2026 Demo / reviewers in the wild / expert
Hervé Bredin
dblp:59/4116
· DBLP profile ↗
50ranked-venue papers
13as first author
16since 2021 · last 2026
0000-0002-3739-925XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 13 first-author · 14 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design choices for PixIT-based speaker-attributed ASR: Team ToTaTo at the NOTSOFAR-1 challenge
Joonas Kalda, Séverin Baroudi, Martin Lebourdais, Clément Pagés, Ricard Marxer, Tanel Alumäe, Hervé Bredin |
Comput. Speech Lang. | 7 |
| 2025 | On the Use of Self-Supervised Representation Learning for Speaker Diarization and SeparationabstractSelf-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in lowresource settings, over the past few years. Despite this, evaluations on tasks such as Speaker Diarization and Speech Separation remain limited. This paper investigates the quality of recent selfsupervised speech representations on these two speaker identityrelated tasks, highlighting gaps in the current literature that stem from limitations in the existing benchmarks-particularly the lack of diversity in evaluation datasets and variety in downstream systems associated to both diarization and separation. Séverin Baroudi, Hervé Bredin, Joseph Razik, Ricard Marxer |
ASRU | 2 |
| 2025 | Diarization-Guided Multi-Speaker EmbeddingsabstractInternational audience Joonas Kalda, Clément Pagés, Tanel Alumäe, Hervé Bredin |
INTERSPEECH | 4 |
| 2024 | Specializing Self-Supervised Speech Representations for Speaker SegmentationabstractInternational audience Séverin Baroudi, Thomas Pellegrini, Hervé Bredin |
INTERSPEECH | 3 |
| 2024 | TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024abstractInternational audience Joonas Kalda, Tanel Alumäe, Martin Lebourdais, Hervé Bredin, Séverin Baroudi, Ricard Marxer |
INTERSPEECH | 4 |
| 2024 | Gryannote open-source speaker diarization labeling tool
Clément Pages, Hervé Bredin |
INTERSPEECH | 2 |
| 2024 | On the calibration of powerset speaker diarization modelsabstractEnd-to-end neural diarization models have usually relied on a multilabel-classification formulation of the speaker diarization problem. Recently, we proposed a powerset multiclass formulation that has beaten the state-of-the-art on multiple datasets. In this paper, we propose to study the calibration of a powerset speaker diarization model, and explore some of its uses. We study the calibration in-domain, as well as out-of-domain, and explore the data in low-confidence regions. The reliability of model confidence is then tested in practice: we use the confidence of the pretrained model to selectively create training and validation subsets out of unannotated data, and compare this to random selection. We find that top-label confidence can be used to reliably predict high-error regions. Moreover, training on low-confidence regions provides a better calibrated model, and validating on low-confidence regions can be more annotation-efficient than random regions. Alexis Plaquet, Hervé Bredin |
INTERSPEECH | 2 |
| 2024 | Multi-latency look-ahead for streaming speaker segmentationabstractInternational audience Bilal Rahou, Hervé Bredin |
INTERSPEECH | 2 |
| 2023 | Brouhaha: Multi-Task Training for Voice Activity Detection, Speech-to-Noise Ratio, and C50 Room Acoustics EstimationabstractMost automatic speech processing systems register degraded performance when applied to noisy or reverberant speech. But how can one tell whether speech is noisy or reverberant? We propose Brouhaha, a neural network jointly trained to extract speech/non-speech segments, speech-to-noise ratios, and C50 room acoustics from single-channel recordings. Brouhaha is trained using a data-driven approach in which noisy and reverberant audio segments are synthesized. We first evaluate its performance and demonstrate that the proposed multi-task regime is beneficial. We then present two scenarios illustrating how Brouhaha can be used on naturally noisy and reverberant data: 1) to investigate the errors made by a speaker diarization model (pyannote.audio); and 2) to assess the reliability of an automatic speech recognition model (Whisper from OpenAI). Both our pipeline and a pretrained model are open source and shared with the speech community. Marvin Lavechin, Marianne Métais, Hadrien Titeux, Alodie Boissonnet, Jade Copet, Morgane Rivière, Elika Bergelson, Alejandrina Cristià, Emmanuel Dupoux, Hervé Bredin |
ASRU | 10 |
| 2023 | pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipeabstractInternational audience Hervé Bredin |
INTERSPEECH | 1 |
| 2023 | BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language modelsabstractInternational audience Marvin Lavechin, Yaya Sy, Hadrien Titeux, María Andrea Cruz Blandón, Okko Johannes Räsänen, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 6 |
| 2023 | Powerset multi-class cross entropy loss for neural speaker diarizationabstractInternational audience Alexis Plaquet, Hervé Bredin |
INTERSPEECH | 2 |
| 2022 | Bazinga! A Dataset for Multi-Party Dialogues StructuringabstractWe introduce a dataset built around a large collection of TV (and movie) series. Those are filled with challenging multi-party dialogues. Moreover, TV series come with a very active fan base that allows the collection of metadata and accelerates annotation. With 16 TV and movie series, Bazinga! amounts to 400+ hours of speech and 8M+ tokens, including 500K+ tokens annotated with the speaker, addressee, and entity linking information. Along with the dataset, we also provide a baseline for speaker diarization, punctuation restoration, and person entity recognition. The results demonstrate the difficulty of the tasks and of transfer learning from models trained on mono-speaker audio or written text, which is more widely available. This work is a step towards better multi-party dialogue structuring and understanding. Bazinga! is available at hf.co/bazinga. Because (a large) part of Bazinga! is only partially annotated, we also expect this dataset to foster research towards self- or weakly-supervised learning methods. Paul Lerner, Juliette Bergoënd, Camille Guinaudeau, Hervé Bredin, Benjamin Maurice, Sharleyne Lefevre, Martin Bouteiller, Aman Berhe, Léo Galmant, Ruiqing Yin, Claude Barras |
LREC | 4 |
| 2022 | Continual Self-Supervised Domain Adaptation for End-to-End Speaker DiarizationabstractIn conventional domain adaptation for speaker diarization, a large collection of annotated conversations from the target domain is required. In this work, we propose a novel continual training scheme for domain adaptation of an end-to-end speaker diarization system, which processes one conversation at a time and benefits from full self-supervision thanks to pseudo-labels. The qualities of our method allow for autonomous adaptation (e.g. of a voice assistant to a new house-hold), while also avoiding permanent storage of possibly sensitive user conversations. We experiment extensively on the 11 domains of the DIHARD III corpus and show the effectiveness of our approach with respect to a pre-trained base-line, achieving a relative 17% performance improvement. We also find that data augmentation and a well-defined target domain are key factors to avoid divergence and to benefit from transfer. Juan Manuel Coria, Hervé Bredin, Sahar Ghannay, Sophie Rosset |
SLT | 2 |
| 2021 | Overlap-Aware Low-Latency Online Speaker Diarization Based on End-to-End Local SegmentationabstractWe propose to address online speaker diarization as a combination of incremental clustering and local diarization applied to a rolling buffer updated every 500ms. Every single step of the proposed pipeline is designed to take full advantage of the strong ability of a recently proposed end-to-end overlap-aware segmentation to detect and separate overlapping speakers. In particular, we propose a modified version of the statistics pooling layer (initially introduced in the x-vector architecture) to give less weight to frames where the segmentation model predicts simultaneous speakers. Furthermore, we derive cannot-link constraints from the initial segmentation step to prevent two local speakers from being wrongfully merged during the incremental clustering step. Finally, we show how the latency of the proposed approach can be adjusted between 500ms and 5s to match the requirements of a particular use case, and we provide a systematic analysis of the influence of latency on the overall performance (on AMI, DIHARD and VoxConverse). Juan Manuel Coria, Hervé Bredin, Sahar Ghannay, Sophie Rosset |
ASRU | 2 |
| 2021 | End-To-End Speaker Segmentation for Overlap-Aware ResegmentationabstractSpeaker segmentation consists in partitioning a conversation between one or more speakers into speaker turns. Usually addressed as the late combination of three sub-tasks (voice activity detection, speaker change detection, and overlapped speech detection), we propose to train an end-to-end segmentation model that does it directly. Inspired by the original end-to-end neural speaker diarization approach (EEND), the task is modeled as a multi-label classification problem using permutation-invariant training. The main difference is that our model operates on short audio chunks (5 seconds) but at a much higher temporal resolution (every 16ms). Experiments on multiple speaker diarization datasets conclude that our model can be used with great success on both voice activity detection and overlapped speech detection. Our proposed model can also be used as a post-processing step, to detect and correctly assign overlapped speech regions. Relative diarization error rate improvement over the best considered baseline (VBx) reaches 17% on AMI, 13% on DIHARD 3, and 13% on VoxConverse. Hervé Bredin, Antoine Laurent |
Interspeech | 1 |
| 2020 | Pyannote.Audio: Neural Building Blocks for Speaker DiarizationabstractWe introduce pyannote.audio, an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be combined and jointly optimized to build speaker diarization pipelines. pyannote.audio also comes with pre-trained models covering a wide range of domains for voice activity detection, speaker change detection, overlapped speech detection, and speaker embedding - reaching state-of-the-art performance for most of them. Hervé Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, Marie-Philippe Gill |
ICASSP | 1 |
| 2020 | Overlap-Aware Diarization: Resegmentation Using Neural End-to-End Overlapped Speech DetectionabstractWe address the problem of effectively handling overlapping speech in a diarization system. First, we detail a neural Long Short-Term Memory- based architecture for overlap detection. Secondly, detected overlap regions are exploited in conjunction with a frame-level speaker posterior matrix to make two-speaker assignments for overlapped frames in the resegmentation step. The overlap detection module achieves state-of-the-art performance on the AMI, DIHARD, and ETAPE corpora. We apply overlap-aware resegmentation on AMI, resulting in a 20% relative DER reduction over the baseline system. While this approach is by no means an end-all solution to overlap-aware diarization, it reveals promising directions for handling overlap. Latané Bullock, Hervé Bredin, L. Paola García-Perera |
ICASSP | 2 |
| 2020 | An Open-Source Voice Type Classifier for Child-Centered Daylong RecordingsabstractInternational audience Marvin Lavechin, Ruben Bousbib, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià |
INTERSPEECH | 3 |
| 2020 | End-to-End Domain-Adversarial Voice Activity DetectionabstractInternational audience Marvin Lavechin, Marie-Philippe Gill, Ruben Bousbib, Hervé Bredin, L. Paola García-Perera |
INTERSPEECH | 4 |
| 2019 | LSTM Based Similarity Measurement with Spectral Clustering for Speaker DiarizationabstractInternational audience Qingjian Lin, Ruiqing Yin, Ming Li 0026, Hervé Bredin, Claude Barras |
INTERSPEECH | 4 |
| 2018 | Neural Speech Turn Segmentation and Affinity Propagation for Speaker DiarizationabstractInternational audience Ruiqing Yin, Hervé Bredin, Claude Barras |
INTERSPEECH | 2 |
| 2017 | TristouNet: Triplet loss for speaker turn embeddingabstractTristouNet is a neural network architecture based on Long Short-Term Memory recurrent networks, meant to project speech sequences into a fixed-dimensional euclidean space. Thanks to the triplet loss paradigm used for training, the resulting sequence embeddings can be compared directly with the euclidean distance, for speaker comparison purposes. Experiments on short (between 500ms and 5s) speech turn comparison and speaker change detection show that TristouNet brings significant improvements over the current state-of-the-art techniques for both tasks. Hervé Bredin |
ICASSP | 1 |
| 2017 | pyannote.metrics: A Toolkit for Reproducible Evaluation, Diagnostic, and Error Analysis of Speaker Diarization SystemsabstractInternational audience Hervé Bredin |
INTERSPEECH | 1 |
| 2017 | Combining Speaker Turn Embedding and Incremental Structure Prediction for Low-Latency Speaker DiarizationabstractInternational audience Guillaume Wisniewski, Hervé Bredin, Gregory Gelly, Claude Barras |
INTERSPEECH | 2 |
| 2017 | Speaker Change Detection in Broadcast TV Using Bidirectional Long Short-Term Memory NetworksabstractInternational audience Ruiqing Yin, Hervé Bredin, Claude Barras |
INTERSPEECH | 2 |
| 2017 | Multimodal person discovery in broadcast TV: lessons learned from MediaEval 2015
Johann Poignant, Hervé Bredin, Claude Barras |
Multim. Tools Appl. | 2 |
| 2016 | Post-Hoc Interactive Analytics of Errors in the Context of a Person Discovery TaskabstractPart of the research effort in automatic person discovery in multimedia content consists in analyzing the errors made by algorithms. However exploring the space of models relating algorithmic errors in person discovery to intrinsic properties of associated shots (e.g. person facing the camera) - coined as post-hoc analysis in this paper - requires data curation and statistical model tuning, which can be cumbersome. In this paper we present a visual and interactive tool that facilitates this exploration. A case study is conducted with multimedia researchers to validate the tool. Real data obtained from the MediaEval person discovery task was used for this experiment. Our approach yielded novel insight that was completely unsuspected previously. Pierrick Bruneau, Mickaël Stefas, Johann Poignant, Hervé Bredin, Claude Barras |
ISM | 4 |
| 2016 | The CAMOMILE Collaborative Annotation Platform for Multi-modal, Multi-lingual and Multi-media Documents
Johann Poignant, Mateusz Budnik, Hervé Bredin, Claude Barras, Mickaël Stefas, Pierrick Bruneau, Gilles Adda, Laurent Besacier, Hazim Kemal Ekenel, Gil Francopoulo, Javier Hernando, Joseph Mariani, Ramon Morros, Georges Quénot, Sophie Rosset, Thomas Tamisier |
LREC | 3 |
| 2016 | Benchmarking multimedia technologies with the CAMOMILE platform: the case of Multimodal Person Discovery at MediaEval 2015
Johann Poignant, Hervé Bredin, Claude Barras, Mickaël Stefas, Pierrick Bruneau, Thomas Tamisier |
LREC | 2 |
| 2016 | Improving Speaker Diarization of TV Series using Talking-Face Detection and ClusteringabstractWhile successful on broadcast news, meetings or telephone conversation, state-of-the-art speaker diarization techniques tend to perform poorly on TV series or movies. In this paper, we propose to rely on state-of-the-art face clustering techniques to guide acoustic speaker diarization. Two approaches are tested and evaluated on the first season of Game Of Thrones TV series. The second (better) approach relies on a novel talking-face detection module based on bi-directional long short-term memory recurrent neural network. Both audio-visual approaches outperform the audio-only baseline. A detailed study of the behavior of these approaches is also provided and paves the way to future improvements. Hervé Bredin, Gregory Gelly |
ACM Multimedia | 1 |
| 2015 | A Visual Analytics Approach to Finding Factors Improving Automatic Speaker IdentificationsabstractClassification quality criteria such as precision, recall, and F-measure are generally the basis for evaluating contributions in automatic speaker recognition. Specifically, comparisons are carried out mostly via mean values estimated on a set of media. Whilst this approach is relevant to assess improvement w.r.t. the state-of-the-art, or ranking participants in the context of an automatic annotation challenge, it gives little insight to system designers in terms of cues for improving algorithms, hypothesis formulation, and evidence display. This paper presents a design study of a visual and interactive approach to analyze errors made by automatic annotation algorithms. A timeline-based tool emerged from prior steps of this study. A critical review, driven by user interviews, exposes caveats and refines user objectives. The next step of the study is then initiated by sketching designs combining elements of the current prototype to principles newly identified as relevant. Pierrick Bruneau, Mickaël Stefas, Hervé Bredin, Johann Poignant, Thomas Tamisier, Claude Barras |
ICMI | 3 |
| 2015 | Collaborative annotation for person identification in TV shows
Mateusz Budnik, Laurent Besacier, Johann Poignant, Hervé Bredin, Claude Barras, Mickaël Stefas, Pierrick Bruneau, Thomas Tamisier |
INTERSPEECH | 4 |
| 2015 | Structured prediction for speaker identification in TV seriesabstractInternational audience Elena Knyazeva, Guillaume Wisniewski, Hervé Bredin, François Yvon |
INTERSPEECH | 3 |
| 2015 | Lexical speaker identification in TV shows
Anindya Roy, Hervé Bredin, William Hartmann, Viet Bac Le, Claude Barras, Jean-Luc Gauvain |
Multim. Tools Appl. | 2 |
| 2014 | Collaborative Annotation of Multimedia Resources
Pierrick Bruneau, Mickaël Stefas, Mateusz Budnik, Johann Poignant, Hervé Bredin, Thomas Tamisier, Benoît Otjacques |
CDVE | 5 |
| 2014 | A Web-Based Tool for the Visual Analysis of Media AnnotationsabstractMultimedia annotation algorithms infer localized metadata in multimedia content, e.g. Speakers' voices or subjects' faces. There is a growing need of experts from this domain to perform advanced analyses, that go beyond medium-scale quality metrics. This paper describes a novel visual tool, that addresses the concerns of multimedia experts using interactive visualization principles. Multiple coordinated views, augmented by interactive inspection facilities, ease both the navigation in media annotations and the visual detection of relevant information. The usefulness of our approach is supported by experimental scenarios using a real multimedia corpus. Pierrick Bruneau, Mickaël Stefas, Hervé Bredin, Anh-Phuong Ta, Thomas Tamisier, Claude Barras |
IV | 3 |
| 2014 | TVD: A Reproducible and Multiply Aligned TV Series Dataset
Anindya Roy, Camille Guinaudeau, Hervé Bredin, Claude Barras |
LREC | 3 |
| 2014 | "Sheldon speaking, Bonjour!": Leveraging Multilingual Tracks for (Weakly) Supervised Speaker IdentificationabstractWe address the problem of speaker identification in multimedia data, and TV series in particular. While speaker identification is traditionally a supervised machine-learning task, our first contribution is to significantly reduce the need for costly preliminary manual annotations through the use of automatically aligned (and potentially noisy) fan-generated transcripts and subtitles. Hervé Bredin, Anindya Roy, Nicolas Pécheux, Alexandre Allauzen |
ACM Multimedia | 1 |
| 2013 | Integer linear programming for speaker diarization and cross-modal identification in TV broadcastabstractInternational audience Hervé Bredin, Johann Poignant |
INTERSPEECH | 1 |
| 2012 | Community-driven hierarchical fusion of numerous classifiers: Application to video semantic indexingabstractWe deal with the issue of combining dozens of classifiers into a better one. Our first contribution is the introduction of the notion of communities of classifiers. We build a complete graph with one node per classifier and edges weighted by a measure of similarity between connected classifiers. The resulting community structure is uncovered from this graph using the state-of-the-art Louvain algorithm. Our second contribution is a hierarchical fusion approach driven by these communities. First, intra-community fusion results in one classifier per community. Then, inter-community fusion takes advantage of their complementarity to achieve much better classification performance. Application to the combination of 90 classifiers in the framework of TRECVid 2010 Semantic Indexing task shows a 30% increase in performance relative to a baseline flat fusion. Hervé Bredin |
ICASSP | 1 |
| 2012 | Segmentation of TV shows into scenes using speaker diarization and speech recognitionabstractWe investigate the use of speaker diarization (SD) and automatic speech recognition (ASR) for the segmentation of audiovisual documents into scenes. We introduce multiple monomodal and multimodal approaches based on a state-of-the-art algorithm called generalized scene transition graph (GSTG). First, we extend the latter with the use of semantic information derived from both SD and ASR. Then, multimodal fusion of color histograms, SD and ASR is investigated at various point of the GSTG pipeline (early, late or intermediate fusion). Experiments driven on a few episodes of a popular TV show indicate that SD and ASR can be successfully combined with visual information and bring an additional +11% relative increase in terms of F1-measure for scene boundary detection over the state-of-the-art baseline. Hervé Bredin |
ICASSP | 1 |
| 2012 | Unsupervised Speaker Identification using Overlaid Texts in TV BroadcastabstractPoster Session: Speaker Recognition III Johann Poignant, Hervé Bredin, Viet Bac Le, Laurent Besacier, Claude Barras, Georges Quénot |
INTERSPEECH | 2 |
| 2012 | StoViz: story visualization of TV seriesabstractRecent TV series tend to have more and more complex plot. They follow the lives of numerous characters and are made of multiple intertwined stories. In this paper, we introduce StoViz, a web-based interface allowing a fast overview of this kind of episode structure, based on our plot de-interlacing system. StoViz has two main goals. First, it provides the user with a useful overview of the episode by displaying each story separately and a short abstract extracted from them. Then, it allows an efficient visual comparison of the output of any automatic plot de-interlacing algorithm with the manual annotation in terms of stories and is therefore very helpful for evaluation purposes. StoViz is available online at http://stoviz.niderb.fr. Philippe Ercolessi, Hervé Bredin, Christine Sénac |
ACM Multimedia | 2 |
| 2009 | An interactive and multi-level framework for summarising user generated videosabstractWe present an interactive and multi-level abstraction framework for user-generated video (UGV) summarisation, allowing a user the flexibility to select a summarisation criterion out of a number of methods provided by the system. First, a given raw video is segmented into shots, and each shot is further decomposed into sub-shots in line with the change in dominant camera motion. Secondly, principal component analysis (PCA) is applied to the colour representation of the collection of sub-shots, and a content map is created using the first few components. Each sub-shot is represented with a "footprint" on the content map, which reveals its content significance (coverage) and the most dynamic segment. The final stage of abstraction is devised in a user-assisted manner whereby a user is able to specify a desired summary length, with options to interactively perform abstraction at different granularity of visual comprehension. The results obtained show the potential benefit in significantly alleviating the burden of laborious user intervention associated with conventional video editing/browsing. Saman Cooray, Hervé Bredin, Li-Qun Xu, Noel E. O'Connor |
ACM Multimedia | 2 |
| 2009 | Audio-visual speech asynchrony detection using co-inertia analysis and coupled hidden markov models
Enrique Argones-Rúa, Hervé Bredin, Carmen García-Mateo, Gérard Chollet, Daniel González-Jiménez |
Pattern Anal. Appl. | 2 |
| 2008 | Making talking-face authentication robust to deliberate impostureabstractWe expose the limitations of existing frameworks designed for the evaluation of audiovisual biometric authentication algorithms. The weakness of a classical audiovisual authentication system is uncovered when confronted to realistic deliberate impostors. A client-dependent audiovisual synchrony measure is used in order to deal with deliberate impostors and three new fusion strategies and their performance against random and deliberate impostors are studied. Hervé Bredin, Gérard Chollet |
ICASSP | 1 |
| 2008 | Some results from the biosecure talking face evaluation campaignabstractThe BioSecure Network of Excellence has collected a large multi- biometric publicly available database and organized the BioSecure Multimodal Evaluation Campaigns (BMEC) in 20072. This paper reports on the Talking Faces campaign. Open source reference systems were made available to participants and four laboratories submitted executable code to the organizer who performed tests on sequestered data. Several deliberate impostures were tested. It is demonstrated that forgeries are a real threat for such systems. A technological race is ongoing between deliberate impostors and system developers. Benoit G. B. Fauve, Hervé Bredin, Walid Karam, Florian Verdet, Aurélien Mayoue, Gérard Chollet, Jean Hennebert, Richard P. Lewis 0002, John S. D. Mason, Chafic Mokbel, Dijana Petrovska-Delacrétaz |
ICASSP | 2 |
| 2007 | Audio-Visual Speech Synchrony Measure for Talking-Face Identity VerificationabstractWe investigate the use of audio-visual speech synchrony measure in the framework of identity verification based on talking faces. Two synchrony measures based on canonical correlation analysis and co-inertia analysis respectively are introduced and their performances are evaluated on the specific task of detecting synchronized and not-synchronized audio-visual speech sequences. The notion of high-effort impostor attacks is also introduced as a dangerous threat for current biometric system based on speaker verification and face recognition. A novel biometric modality based on synchrony measures is introduced in order to improve the overall performance of identity verification, and more specifically its robustness to replay attacks. Hervé Bredin, Gérard Chollet |
ICASSP (2) | 1 |
| 2006 | Detecting Replay Attacks in Audiovisual Identity VerificationabstractWe describe an algorithm that detects a lack of correspondence between speech and lip motion by detecting and monitoring the degree of synchrony between live audio and visual signals. It is simple, effective, and computationally inexpensive; providing a useful degree of robustness against basic replay attacks and against speech or image forgeries. The method is based on a cross-correlation analysis between two streams of features, one from the audio signal and the other from the image sequence. We argue that such an algorithm forms an effective first barrier against several kinds of replay attack that would defeat existing verification systems based on standard multimodal fusion techniques. In order to provide an evaluation mechanism for the new technique we have augmented the protocols that accompany the BANCA multimedia corpus by defining new scenarios. We obtain 0% equal-error rate (EER) on the simplest scenario and 35% on a more challenging one Hervé Bredin, Antonio Miguel, Ian H. Witten, Gérard Chollet |
ICASSP (1) | 1 |