VLDB 2026 Research / reviewers in the wild / expert
Julien Pinquier
dblp:05/5283
· DBLP profile ↗
43ranked-venue papers
7as first author
12since 2021 · last 2024
0000-0003-1556-1284ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 24 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Emvd Dataset: a Dataset of Extreme Vocal Distortion Techniques Used in Heavy MetalabstractIn this paper, we introduce the Extreme Metal Vocals Dataset, which comprises a collection of recordings of extreme vocal techniques performed within the realm of heavy metal music. The dataset consists of 760 audio excerpts of 1 second to 30 seconds long, totaling about 100 min of audio material, roughly composed of 60 minutes of distorted voices and 40 minutes of clear voice recordings. These vocal recordings are from 27 different singers and are provided without accompanying musical instruments or post-processing effects. The distortion taxonomy within this dataset encompasses four distinct distortion techniques and three vocal effects, all performed in different pitch ranges. Performance of a state-of-the-art deep learning model is evaluated for two different classification tasks related to vocal techniques, demonstrating the potential of this resource for the audio processing community. Modan Tailleur, Julien Pinquier, Laurent Millot, Corsin Vogel, Mathieu Lagrange |
CBMI | 2 |
| 2024 | CoNeTTE: An Efficient Audio Captioning System Leveraging Multiple Datasets With Task EmbeddingabstractAutomated Audio Captioning (AAC) involves generating natural language descriptions of audio content, using encoder-decoder architectures. An audio encoder produces audio embeddings fed to a decoder, usually a Transformer decoder, for caption generation. In this work, we describe our model, which novelty, compared to existing models, lies in the use of a ConvNeXt architecture as audio encoder, adapted from the vision domain to audio classification. This model, called CNext-trans, achieved state-of-the-art scores on the AudioCaps (AC) dataset and performed competitively on Clotho (CL), while using four to forty times fewer parameters than existing models. We examine potential biases in the AC dataset due to its origin from AudioSet by investigating unbiased encoder's impact on performance. Using the well-known PANN's CNN14, for instance, as an unbiased encoder, we observed a 0.017 absolute reduction in SPIDEr score (where higher scores indicate better performance). To improve cross-dataset performance, we conducted experiments by combining multiple AAC datasets (AC, CL, MACS, WavCaps) for training. Although this strategy enhanced overall model performance across datasets, it still fell short compared to models trained specifically on a single target dataset, indicating the absence of a one-size-fits-all model. To mitigate performance gaps between datasets, we introduced a Task Embedding (TE) token, allowing the model to identify the source dataset for each input sample. We provide insights into the impact of these TEs on both the form (words) and content (sound event types) of the generated captions. The resulting model, named CoNeTTE, an unbiased CNext-trans model enriched with dataset-specific Task Embeddings, achieved SPIDEr scores of 0.467 and 0.310 on AC and CL, respectively. Etienne Labbé, Thomas Pellegrini, Julien Pinquier |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Can We Use Speaker Embeddings On Spontaneous Speech Obtained From Medical Conversations To Predict Intelligibility?abstractThe automatic prediction of speech intelligibility is a recurrent problem in the context of pathological speech. Despite recent developments, these systems are normally applied to specific speech tasks recorded in clean conditions that do not necessarily reflect a healthcare environment. In the present paper, we intend to test the reliability of an intelligibility predictor on data obtained in clinical conditions, in the specific case of head and neck cancer. In order to do so, we present a system based on speaker embeddings trained on a multi-task methodology to simultaneous predict speech intelligibility and speech disorder severity. The results obtained on the different evaluation tasks display correlations as high as 0.891 on a hospital patient set, showing robustness to the type of speech material used in these automatic assessments. Moreover, the usage of spontaneous speech during the evaluation shed light on an understudied, but with more ecological validity, type of speech material which displayed promising results. The reliability displayed across the different tasks suggests a direct deployment of the developed systems in a hospital setting. Sebastião Quintas, Mathieu Balaguer, Julie Mauclair, Virginie Woisard, Julien Pinquier |
ASRU | 5 |
| 2023 | Towards Reducing Patient Effort for the Automatic Prediction of Speech Intelligibility in Head and Neck CancersabstractThe automatic prediction of speech intelligibility can be seen as a growing and relevant alternative to the perceptual evaluations used clinically, which are known to be biased, variant and subjective. We propose an automatic way to regress an intelligibility score based on a recurrent model with a self-attention mechanism. This approach not only presented a high correlation of 0.87 when applied to a pseudo-word task designed for head and neck cancers, but also a significant decrease in error of more than 50%, when compared to previous approaches. Moreover, we have also studied the reliability of the same system when operating with smaller amounts of data at inference time. The results suggest that we can reduce the linguistic sample size to only 30% of the full sample, without losing performance. This aspect validates the reliability of using a smaller subset of data when predicting intelligibility, which can be extremely useful to prevent patient’s fatigue by creating smaller batteries of clinical exams. Sebastião Quintas, Alberto Abad, Julie Mauclair, Virginie Woisard, Julien Pinquier |
ICASSP | 5 |
| 2023 | Audio-video fusion strategies for active speaker detection in meetings
Lionel Pibre, Francisco Madrigal, Cyrille Equoy, Frédéric Lerasle, Thomas Pellegrini, Julien Pinquier, Isabelle Ferrané |
Multim. Tools Appl. | 6 |
| 2022 | Prediction of L2 speech proficiency based on multi-level linguistic featuresabstractInternational audience Verdiana De Fino, Lionel Fontan, Julien Pinquier, Isabelle Ferrané, Sylvain Detey |
INTERSPEECH | 3 |
| 2022 | Automatic Assessment of Speech Intelligibility using Consonant Similarity for Head and Neck CancerabstractInternational audience Sebastião Quintas, Julie Mauclair, Virginie Woisard, Julien Pinquier |
INTERSPEECH | 4 |
| 2021 | Towards a content-based prediction of personalized musical preferences using transfer learningabstractMusic recommender systems attempt to provide to the users tracks in accordance with their preferences. Thus, content-based recommender systems rely on the audio signal in order to infer users' preferences. We present a preliminary study of a personalized recommendation task which is here considered as a prediction between two classes: `like' and `dislike'. First, a categorization experiment was conducted in order to generate users' preference data. Each volunteer has sorted pieces of music according to whether he liked those or not, by creating personal playlists (around 10 hours per volunteer) on a music streaming platform. Pieces of music were either coming from the free browsing of the volunteer on the platform or from a corpus of pieces that was built according to musicological criteria. We used an inductive transfer learning with music genre classification as a source task and the prediction of `like' and `dislike' classes as a target task. We pre-trained our models with two different corpora: GTZAN and FMA. Cross-validation results obtained are promising, with very satisfying results (more than 80% satisfaction) for 14 volunteers out of 20. This method could be used for a music recommendation purpose and further research will be conducted in this direction. Nicolas Dauban, Christine Sénac, Julien Pinquier, Pascal Gaillard |
CBMI | 3 |
| 2021 | Automatic macro segmentation into interaction sequence: a silence-based approach for meeting structuringabstractMeetings are a common activity in professional contexts, and it remains difficult to analyze them because they are not always structured and people cut each other off (in a debate of ideas for example). A first step, to facilitate their analysis, is to segment the meeting into homogeneous zones at interaction level. To do so, we studied the typology of the non-speech segments (pauses and silences) in order to determine the different sequences during a meeting. Indeed, information such as the frequency and lengths of the non-speech segments will be different during a presentation or a debate. In this article, we propose an original approach to segment meetings using only the non-speech segments. We apply a Voice Activity Detection (VAD) to find the non-speech segments from which a set of parameters are extracted to study the typology of silence segments. We then use a sliding window on the whole meeting and we apply an unsupervised approach on each of these windows. We have validated our approaches using purity and coverage metrics on part of the AMI corpus (38 meetings of about 28 minutes each). This approach is non-invasive and relies only on acoustic information and does not analyze speech content since moments containing speech, and potentially sensitive information, are not processed. Lionel Pibre, Sélim Mechrouh, Thomas Pellegrini, Julien Pinquier, Isabelle Ferrané |
CBMI | 4 |
| 2021 | Simulating Reading Mistakes for Child Speech Transformer-Based Phone RecognitionabstractInternational audience Lucile Gelin, Thomas Pellegrini, Julien Pinquier, Morgane Daniel |
Interspeech | 3 |
| 2021 | Improving vehicle re-identification using CNN latent spaces: Metrics comparison and track-to-track extensionabstractAbstract Herein, the problem of vehicle re‐identification using distance comparison of images in CNN latent spaces is addressed. First, the impact of the distance metrics, comparing performances obtained with different metrics is studied: the minimal Euclidean distance ( MED ), the minimal cosine distance ( MCD ) and the residue of the sparse coding reconstruction ( RSCR ). These metrics are applied using features extracted from five different CNN architectures, namely ResNet18, AlexNet, VGG16, InceptionV3 and DenseNet201. We use the specific vehicle re‐identification dataset VeRi to fine‐tune these CNNs and evaluate results. Overall, independently of the CNN used, MCD outperforms MED , commonly used in the literature. These results are confirmed on other vehicle retrieval datasets. Second, the state‐of‐the‐art image‐to‐track process (I2TP) is extended to a track‐to‐track process (T2TP). The three distance metrics are extended to measure distance between tracks, enabling T2TP. T2TP and I2TP are compared using the same CNN models. Results show that T2TP outperforms I2TP for MCD and RSCR. T2TP combining DenseNet201 and MCD ‐based metrics exhibits the best performances, outperforming the state‐of‐the‐art I2TP‐based models. Finally, experiments highlight two main results: i) the impact of metric choice in vehicle re‐identification, and ii) T2TP improves the performances compared with I2TP, especially when coupled with MCD ‐based metrics. Geoffrey Roman-Jimenez, Patrice Guyot, Thierry Malon, Sylvie Chambon, Vincent Charvillat, Alain Crouzil, André Péninou, Julien Pinquier, Florence Sèdes, Christine Sénac |
IET Comput. Vis. | 8 |
| 2021 | End-to-end acoustic modelling for phone recognition of young readers
Lucile Gelin, Morgane Daniel, Julien Pinquier, Thomas Pellegrini |
Speech Commun. | 3 |
| 2020 | Automatic Prediction of Speech Intelligibility Based on X-Vectors in the Context of Head and Neck CancerabstractInternational audience Sebastião Quintas, Julie Mauclair, Virginie Woisard, Julien Pinquier |
INTERSPEECH | 4 |
| 2020 | Subjective Evaluation of Comprehensibility in Movie InteractionsabstractVarious research works have dealt with the comprehensibility of textual, audio, or audiovisual documents, and showed that factors related to text (e.g. linguistic complexity), sound (e.g. speech intelligibility), image (e.g. presence of visual context), or even to cognition and emotion can play a major role in the ability of humans to understand the semantic and pragmatic contents of a given document. However, to date, no reference human data is available that could help investigating the role of the linguistic and extralinguistic information present at these different levels (i.e., linguistic, audio/phonetic, and visual) in multimodal documents (e.g., movies). The present work aimed at building a corpus of human annotations that would help to study further how much and in which way the human perception of comprehensibility (i.e., of the difficulty of comprehension, referred in this paper as overall difficulty) of audiovisual documents is affected (1) by lexical complexity, grammatical complexity, and speech intelligibility, and (2) by the modality/ies (text, audio, video) available to the human recipient. Estelle I. S. Randria, Lionel Fontan, Maxime Le Coz, Isabelle Ferrané, Julien Pinquier |
LREC | 5 |
| 2019 | Audiovisual Annotation Procedure for Multi-view Field Recordings
Patrice Guyot, Thierry Malon, Geoffrey Roman-Jimenez, Sylvie Chambon, Vincent Charvillat, Alain Crouzil, André Péninou, Julien Pinquier, Florence Sèdes, Christine Sénac |
MMM (1) | 8 |
| 2018 | Perceptual and Automatic Evaluations of the Intelligibility of Speech Degraded by Noise Induced Hearing Loss SimulationabstractInternational audience Imed Laaridh, Julien Tardieu, Cynthia Magnen, Pascal Gaillard, Jérôme Farinas, Julien Pinquier |
INTERSPEECH | 6 |
| 2018 | Carcinologic Speech Severity Index Project: A Database of Speech Disorder Productions to Assess Quality of Life Related to Speech After Cancer
Corine Astésano, Mathieu Balaguer, Jérôme Farinas, Corinne Fredouille, Pascal Gaillard, Alain Ghio, Imed Laaridh, Muriel Lalain, Benoît Lepage, Julie Mauclair, Olivier Nocaudie, Julien Pinquier, Oriol Pont, Gilles Pouchoulin, Michèle Puech, Danièle Robert, Etienne Sicard, Virginie Woisard |
LREC | 12 |
| 2018 | Toulouse campus surveillance dataset: scenarios, soundtracks, synchronized videos with overlapping and disjoint viewsabstractIn surveillance applications, humans and vehicles are the most important common elements studied. In consequence, detecting and matching a person or a car that appears on several videos is a key problem. Many algorithms have been introduced and nowadays, a major relative problem is to evaluate precisely and to compare these algorithms, in reference to a common ground-truth. In this paper, our goal is to introduce a new dataset for evaluating multi-view based methods. This dataset aims at paving the way for multidisciplinary approaches and applications such as 4D-scene reconstruction, object identification/tracking, audio event detection and multi-source meta-data modeling and querying. Consequently, we provide two sets of 25 synchronized videos with audio tracks, all depicting the same scene from multiple viewpoints, each set of videos following a detailed scenario consisting in comings and goings of people and cars. Every video was annotated by regularly drawing bounding boxes on every moving object with a flag indicating whether the object is fully visible or occluded, specifying its category (human or vehicle), providing visual details (for example clothes types or colors), and timestamps of its apparitions and disappearances. Audio events are also annotated by a category and timestamps. Thierry Malon, Geoffrey Roman-Jimenez, Patrice Guyot, Sylvie Chambon, Vincent Charvillat, Alain Crouzil, André Péninou, Julien Pinquier, Florence Sèdes, Christine Sénac |
MMSys | 8 |
| 2016 | A Multi-modal Perception based Architecture for a Non-intrusive Domestic Assistant RobotabstractWe present a multi-modal perception based architecture to realize a non-intrusive domestic assistant robot. The realized robot is non-intrusive in that it only starts interaction with a user when it detects the user's intention to do so automatically. All the robot's actions are based on multi-modal perceptions, which include: user detection based on RGB-D data, user's intention-for-interaction detection with RGB-D and audio data, and communication via speech recognition. The utilization of multi-modal cues in different parts of the robotic activity paves the way to successful robotic runs. Christophe Mollaret, Alhayat Ali Mekonnen, Julien Pinquier, Frédéric Lerasle, Isabelle Ferrané |
HRI | 3 |
| 2016 | Using Phonologically Weighted Levenshtein Distances for the Prediction of Microscopic IntelligibilityabstractInternational audience Lionel Fontan, Isabelle Ferrané, Jérôme Farinas, Julien Pinquier, Xavier Aumont |
INTERSPEECH | 4 |
| 2016 | CNN-Based Phone Segmentation Experiments in a Less-Represented LanguageabstractThese last years, there has been a regain of interest in unsupervised sub-lexical and lexical unit discovery. Speech segmentation into phone-like units may be a first interesting step for such a task. In this article, we report speech segmentation experiments in Xitsonga, a less-represented language spoken in South Africa. We chose to use convolutional neural networks (CNN) with FBANK static coefficients as input. The models take binary decisions whether a boundary is present or not at each signal sliding frame. We compare the use of a model trained exclusively on Xitsonga data to the use of a bootstrap model trained on a larger corpus of another language, the BUCKEYE U.S. English corpus. Using a two-convolution-layer model, a 79% F-measure was obtained on BUCKEYE, with a 20 ms error tolerance. This performance is equal to the human inter-annotator agreement rate. We then used this bootstrap model to segment Xitsonga data and compared the results when adapting it with 1 to 20 minutes of Xitsonga data. Céline Manenti, Thomas Pellegrini, Julien Pinquier |
INTERSPEECH | 3 |
| 2016 | A multi-modal perception based assistive robotic system for the elderly
Christophe Mollaret, Alhayat Ali Mekonnen, Frédéric Lerasle, Isabelle Ferrané, Julien Pinquier, B. Boudet, Pierre Rumeau |
Comput. Vis. Image Underst. | 5 |
| 2015 | Perceiving user's intention-for-interaction: A probabilistic multimodal data fusion schemeabstractUnderstanding people's intention, be it action or thought, plays a fundamental role in establishing coherent communication amongst people, especially in non-proactive robotics, where the robot has to understand explicitly when to start an interaction in a natural way. In this work, a novel approach is presented to detect people's intention-for-interaction. The proposed detector fuses multimodal cues, including estimated head pose, shoulder orientation and vocal activity detection, using a probabilistic discrete state Hidden Markov Model. The multimodal detector achieves up to 80% correct detection rates improving purely audio and RGB-D based variants. Christophe Mollaret, Alhayat Ali Mekonnen, Isabelle Ferrané, Julien Pinquier, Frédéric Lerasle |
ICME | 4 |
| 2015 | Automatic intelligibility measures applied to speech signals simulating age-related hearing lossabstractInternational audience Lionel Fontan, Jérôme Farinas, Isabelle Ferrané, Julien Pinquier, Xavier Aumont |
INTERSPEECH | 4 |
| 2014 | A particle swarm optimization inspired tracker applied to visual trackingabstractVisual tracking is dynamic optimization where time and object state simultaneously influence the problem. In this paper, we intend to show that we built a tracker from an evolutionary optimization approach, the PSO (Particle Swarm optimization) algorithm. We demonstrated that an extension of the original algorithm where system dynamics is explicitly taken into consideration, it can perform an efficient tracking. This tracker is also shown to outperform SIR (Sampling Importance Resampling) algorithm with random walk and constant velocity model, as well as a previously PSO inspired tracker, SPSO (Sequential Particle Swarm Optimization). Experiments were performed both on simulated data and real visual RGB-D information. Our PSO inspired tracker can be a very effective and robust alternative for visual tracking. Christophe Mollaret, Frédéric Lerasle, Isabelle Ferrané, Julien Pinquier |
ICIP | 4 |
| 2014 | Segmentation in singer turns with the Bayesian information criterionabstractAs part of a project on indexing ethno-musicological audio recordings, segmentation in singer turns automatically appeared to be essential. In this article, we present the problem of segmentation in singer turns of musical recordings and our first experiments in this direction by exploring a method based on the Bayesian Information Criterion (BIC), which are used in numerous works in audio segmentation, to detect singer turns. The BIC penalty coefficient was shown to vary when determining its value to achieve the best performance for each recording. In order to avoid the decision about which single value is best for all the documents, we propose to combine several segmentations obtained with different values of this parameter. This method consists of taking a posteriori decisions on which segment boundaries are to be kept. A gain of 7.1% in terms of F-measure was obtained compared to a standard coefficient. Marwa Thlithi, Thomas Pellegrini, Julien Pinquier, Régine André-Obrecht |
INTERSPEECH | 3 |
| 2014 | Hierarchical Hidden Markov Model in detecting activities of daily living in wearable videos for studies of dementia
Svebor Karaman, Jenny Benois-Pineau, Vladislavs Dovgalecs, Rémi Mégret, Julien Pinquier, Régine André-Obrecht, Yann Gaëstel, Jean-François Dartigues |
Multim. Tools Appl. | 5 |
| 2013 | Water sound recognition based on physical modelsabstractThis article describes an audio signal processing algorithm to detect water sounds, built in the context of a larger system aiming to monitor daily activities of elderly people. While previous proposals for water sound recognition relied on classical machine learning and generic audio features to characterize water sounds as a flow texture, we describe here a recognition system based on a physical model of air bubble acoustics. This system is able to recognize a wide variety of water sounds and does not require training. It is validated on a home environmental sound corpus with a classification task, in which all water sounds are correctly detected. In a free detection task on a real life recording, it outperformed the classical systems and obtained 70% of F-measure. Patrice Guyot, Julien Pinquier, Régine André-Obrecht |
ICASSP | 2 |
| 2013 | Two-step detection of water sound events for the diagnostic and monitoring of dementiaabstractA significant aging of world population is foreseen for the next decades. Thus, developing technologies to empower the independency and assist the elderly are becoming of great interest. In this framework, the IMMED project investigates tele-monitoring technologies to support doctors in the diagnostic and follow-up of dementia illnesses such as Alzheimer. Specifically, water sounds are very useful to track and identify abnormal behaviors form everyday activities (e.g. hygiene, household, cooking, etc.). In this work, we propose a double-stage system to detect this type of sound events. In the first stage, the audio stream is segmented with a simple but effective algorithm based on the Spectral Cover feature. The second stage improves the system precision by classifying the segmented streams into water/non-water sound events using Gammatone Cepstral Coefficients and Support Vector Machines. Experimental results reveal the potential of the combined system, yielding a F-measure higher than 80%. Patrice Guyot, Julien Pinquier, Xavier Valero, Francesc Alías |
ICME | 2 |
| 2013 | Superposed speech localisation using frequency trackingabstractOn this paper we present a new approach for the localisation of superposed speech areas. The system is based on the frequency tracking of speech segments following the evolution of the main amplitude frequencies and uses no learning of acoustic or prosodic models. The set of trackings of the frequencies are then grouped together using a distance based on the harmonicity, each group being the production of a single speaker. The co-occurrence of different harmonic groups is then used as a consequence of the presence of multiple speakers. Our method has been evaluated on the data of the French ANR evaluation campaign ETAPE, showing the usability of this approach. Maxime Le Coz, Julien Pinquier, Régine André-Obrecht |
INTERSPEECH | 2 |
| 2012 | Strategies for multiple feature fusion with Hierarchical HMM: Application to activity recognition from wearable audiovisual sensors
Julien Pinquier, Svebor Karaman, Laetitia Letoupin, Patrice Guyot, Rémi Mégret, Jenny Benois-Pineau, Yann Gaëstel, Jean-François Dartigues |
ICPR | 1 |
| 2012 | Detecting individual role using features extracted from speaker diarization results
Benjamin Bigot, Isabelle Ferrané, Julien Pinquier, Régine André-Obrecht |
Multim. Tools Appl. | 3 |
| 2011 | Distinguishing Monophonies From Polyphonies Using Weibull Bivariate DistributionsabstractIn the context of music indexation, it would be useful to have a precise information about the number of sources performing; a source is a solo voice or an isolated instrument which produces a single note at any time. This correspondence discusses the automatic distinction between monophonic music excerpts, where only one source is present, and polyphonic ones. Our method is based on the analysis of a “confidence indicator,” which gives the confidence (in fact its inverse) on the current estimated fundamental frequency (pitch). In a monophony, the confidence indicator is low. In a polyphony, the confidence indicator is higher and varies more. This leads us to compute the short term mean and variance of this indicator, take this 2-D vector as the observation vector and model its conditional distribution with Weibull bivariate models. This probability density function is characterized by five parameters. A method to perform their estimation is developed (in theory and practice). The decision is taken considering the maximum likelihood, computed over one second. The best configuration gives a global error rate of 6.3%, performed on a balanced corpus (18 minutes in total). Hélène Lachambre, Régine André-Obrecht, Julien Pinquier |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Looking for relevant features for speaker role recognitionabstractWhen listening to foreign radio or TV programs we are able to pick up some information from the way people are interacting with each others and easily identify the most dominant speaker or the person who is interviewed. Our work relies on the existence of clues about speaker roles in acoustic and prosodic low-level features extracted from audio files and from speaker segmentations. In this paper we describe an original language-independent method which achieves the recognition of 5 roles (Anchor, Journalist, Other, Punctual Journalist, Punctual Other) with an accuracy of 85% on a 13-hour corpus composed of 46 documents among which can be found different radio shows. A feature selection method is exploited in order to highlight the most relevant features for every speaker role. Index Terms: speaker role recognition, temporal and prosodic features, speaker segmentations, feature selection. Benjamin Bigot, Julien Pinquier, Isabelle Ferrané, Régine André-Obrecht |
INTERSPEECH | 2 |
| 2010 | The IMMED project: wearable video monitoring of people with age dementiaabstractIn this paper, we describe a new application for multimedia indexing, using a system that monitors the instrumental activities of daily living to assess the cognitive decline caused by dementia. The system is composed of a wearable camera device designed to capture audio and video data of the instrumental activities of a patient, which is leveraged with multimedia indexing techniques in order to allow medical specialists to analyze several hour long observation shots efficiently. Rémi Mégret, Vladislavs Dovgalecs, Hazem Wannous, Svebor Karaman, Jenny Benois-Pineau, Elie Khoury 0001, Julien Pinquier, Philippe Joly, Régine André-Obrecht, Yann Gaëstel, Jean-François Dartigues |
ACM Multimedia | 7 |
| 2009 | Improved speaker diarization system for meetingsabstractIn this paper, we investigate new approaches to improve speech activity detection, speaker segmentation and speaker clustering. The main idea behind them is to deal with the problem of speaker diarization for meetings where error rates are relatively high. In opposition to existing methods, a new iterative scheme is proposed considering those three tasks as only one problem. New bidirectional source segmentation is proposed based on the GLR/BIC method. The well-known BIC clustering is also reviewed and a new unsupervised post-processing is added to increase clusters purity. Those new proposals applied on meeting data show a relative improvement of about 40% compared to a standard speaker diarization system. Elie Khoury 0001, Christine Sénac, Julien Pinquier |
ICASSP | 3 |
| 2008 | Dynamic organization of audiovisual database using a user-defined similarity measure based on low-level featuresabstractIn this paper we explore the way to allow a user to interactively organize a multimedia database through a dynamic interface, creating its own "audiovisual concepts" freely. The user defines distances on a small subset of documents, using low-level audio and video off-line automatically extracted descriptors. The semi-supervised learning process, relying on support vector regression used in an early fusion context, leads to generate a behavioral model of those descriptors thanks to human interaction, creating a personal audiovisual similarity measure. Jeremy Philippeau, Julien Pinquier, Philippe Joly, Jean Carrive |
ICIP | 2 |
| 2006 | Audio indexing: primary components retrieval
Julien Pinquier, Régine André-Obrecht |
Multim. Tools Appl. | 1 |
| 2004 | Jingle detection and identification in audio documentsabstractThe paper addresses the soundtrack indexing of multimedia documents. Our purpose is to detect and locate one or many jingles to structure the audio dataflow in program broadcasts (reports). Each jingle is commonly represented by a sequence of spectral vectors, considered as its "signature". Potential candidates are extracted from the data flow by computing a Euclidean distance. They are validated with heuristic rules. The system evaluation is performed on TV and radio corpora (more than 10 hours, 3 TV channels and 3 radio channels). First results show that the system is efficient: among 132 jingles to recognize, we have detected 130 with our reference jingle table of 32 different key sounds. Julien Pinquier, Régine André-Obrecht |
ICASSP (4) | 1 |
| 2003 | A fusion study in speech/music classificationabstractWe present and merge two speech/music classification approaches that we have developed. The first one is a differentiated modeling approach based on a spectral analysis, which is implemented using GMM (Gaussian mixture model). The other one is based on three original features: entropy modulation, stationary segment duration and number of segments. They are merged with the classical 4 Hertz modulation energy. Our classification system is a fusion of the two approaches. It is divided in two classifications (speech/non-speech and music/non-music) and provides 94% of accuracy for speech detection and 90% for music detection, with one second of input signal. Beside the spectral information and GMM, classically used in speech/music discrimination, simple parameters bring complementary and efficient information. Julien Pinquier, Jean-Luc Rouas, Régine André-Obrecht |
ICASSP (2) | 1 |
| 2003 | A fusion study in speech / music classificationabstractIn this paper, we present and merge two speech / music classification approaches of that we have developed. The first one is a differentiated modeling approach based on a spectral analysis, which is implemented with GMM. The other one is based on three original features: entropy modulation, stationary segment duration and number of segments. They are merged with the classical 4 Hertz modulation energy. Our classification system is a fusion of the two approaches. It is divided in two classifications (speech/non-speech and music/non-music) and provides 94 % of accuracy for speech detection and 90 % for music detection, with one second of input signal. Beside the spectral information and GMM, classically used in speech / music discrimination, simple parameters bring complementary and efficient information. Julien Pinquier, Jean-Luc Rouas, Régine André-Obrecht |
ICME | 1 |
| 2002 | Speech and music classification in audio documentsabstractTo index efficiently the soundtrack of multimedia documents, it is necessary to extract elementary and homogeneous acoustic segments. In this paper, we explore such a prior partitioning which consists in detect the two basic components, which are speech and music components. The originality of this work is that music and speech are not considered as two classes and two classification systems are independently defined, a speech/non-speech one and a music/non-music one. This approach permits to better characterize and discriminate each component: in particular, two different feature spaces are necessary as two pairs of Gaussian mixture models. More, the acoustic signal is divided into four types of segments: speech, music, speech-music and other. The experiments are performed on the soundtracks of audio video documents (films, TV sport broadcasts). The performance proves the interest of this approach, so called the Differentiated Modeling Approach. Julien Pinquier, Christine Sénac |
ICASSP | 1 |
| 2002 | Robust speech / music classification in audio documentsabstractInternational audience Julien Pinquier, Jean-Luc Rouas, Régine André-Obrecht |
INTERSPEECH | 1 |