EDBT 2026 Demo / reviewers in the wild / expert
Ricard Marxer
dblp:52/7343
· DBLP profile ↗
41ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0001-5099-5059ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 13 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design choices for PixIT-based speaker-attributed ASR: Team ToTaTo at the NOTSOFAR-1 challenge
Joonas Kalda, Séverin Baroudi, Martin Lebourdais, Clément Pagés, Ricard Marxer, Tanel Alumäe, Hervé Bredin |
Comput. Speech Lang. | 5 |
| 2025 | On the Use of Self-Supervised Representation Learning for Speaker Diarization and SeparationabstractSelf-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in lowresource settings, over the past few years. Despite this, evaluations on tasks such as Speaker Diarization and Speech Separation remain limited. This paper investigates the quality of recent selfsupervised speech representations on these two speaker identityrelated tasks, highlighting gaps in the current literature that stem from limitations in the existing benchmarks-particularly the lack of diversity in evaluation datasets and variety in downstream systems associated to both diarization and separation. Séverin Baroudi, Hervé Bredin, Joseph Razik, Ricard Marxer |
ASRU | 4 |
| 2025 | Aligning Multimodal Representations through an Information BottleneckabstractContrastive losses have been extensively used as a tool for multimodal representation learning. However, it has been empirically observed that their use is not effective to learn an aligned representation space. In this paper, we argue that this phenomenon is caused by the presence of modality-specific information in the representation space. Although some of the most widely used contrastive losses maximize the mutual information between representations of both modalities, they are not designed to remove the modality-specific information. We give a theoretical description of this problem through the lens of the Information Bottleneck Principle. We also empirically analyze how different hyperparameters affect the emergence of this phenomenon in a controlled experimental setup. Finally, we propose a regularization term in the loss function that is derived by means of a variational approximation and aims to increase the representational alignment. We analyze in a set of controlled experiments and real-world applications the advantages of including this regularization term. Antonio Almudévar, José Miguel Hernández-Lobato, Sameer Khurana, Ricard Marxer, Alfonso Ortega Giménez |
ICML | 4 |
| 2025 | Optimizing Underwater Robot Navigation: A Study of DRL Algorithms and Multi-Modal Sensor FusionabstractAutonomous underwater navigation faces significant challenges due to the complexity of the environment, limited localization methods, and poor visibility. This paper investigates the performance of various reinforcement learning (RL) algorithms-Proximal Policy Optimization (PPO), Trust Region Policy Optimization (TRPO), Soft Actor-Critic (SAC), Twin Delayed DDPG (TD3), and Advantage Actor-Critic (A2C)-to improve navigation capabilities of low-cost underwater robots equipped with multi-modal sensors. Advanced depth estimation models such as MiDaS and Depth Anything, combined with domain randomization techniques, are employed to enhance the system's robustness and generalization across varying underwater conditions. The proposed approach integrates real-time sensor data and historical actions to enable 3D maneuvering in simulated environments, leading to significant improvements in sensor fusion, depth perception, and obstacle avoidance. Simulation results demonstrate that the combination of RL techniques with sensor fusion considerably improves mapless autonomous underwater exploration, providing a robust solution for navigating unstructured aquatic environments. The complete implementation is available in an open-source repository, https://github.com/eather0056/BlueROV_Nav_DRL. Md Ether Deowan, Md Shamin Yeasher Yousha, Tihan Mahmud Hossain, Shahriar Hassan, Ricard Marxer |
ICRA | 5 |
| 2025 | Factorized RVQ-GAN For Disentangled Speech TokenizationabstractInternational audience Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos, Juraj Novosad, Peter Gazdik, Ellen Zhang, Zili Huang, Amir Hussein, Ricard Marxer, Yoshiki Masuyama, Ryo Aihara, Chiori Hori, François G. Germain, Gordon Wichern, Jonathan Le Roux |
INTERSPEECH | 10 |
| 2024 | SUCRe: Leveraging Scene Structure for Underwater Color RestorationabstractUnderwater images are altered by the physical characteristics of the medium through which light rays pass before reaching the optical sensor. Scattering and wavelength-dependent absorption significantly modify the captured colors depending on the distance of observed elements to the image plane. In this paper, we aim to recover an image of the scene as if the water had no effect on light propagation. We introduce SUCRe, a novel method that exploits the scene’s 3D structure for underwater color restoration. By following points in multiple images and tracking their intensities at different distances to the sensor, we constrain the optimization of the parameters in an underwater image formation model and retrieve unattenuated pixel intensities. We conduct extensive quantitative and qualitative analyses of our approach in a variety of scenarios ranging from natural light to deep-sea environments using three underwater datasets acquired from real-world scenarios and one synthetic dataset. We also compare the performance of the proposed approach with that of a wide range of existing state-of-the-art methods. The results demonstrate a consistent benefit of exploiting multiple views across a spectrum of objective metrics. Our code is publicly available at github.com/clementinboittiaux/sucre. Clémentin Boittiaux, Ricard Marxer, Claire Dune, Aurélien Arnaubec, Maxime Ferrera, Vincent Hugel |
3DV | 2 |
| 2024 | Scaling Properties of Speech Language ModelsabstractSpeech Language Models (SLMs) aim to learn language from raw audio, without textual resources.Despite significant advances, our current models exhibit weak syntax and semantic abilities.However, if the scaling properties of neural language models hold for the speech modality, these abilities will improve as the amount of compute used for training increases.In this paper, we use models of this scaling behavior to estimate the scale at which our current methods will yield a SLM with the English proficiency of text-based Large Language Models (LLMs).We establish a strong correlation between pre-training loss and downstream syntactic and semantic performance in SLMs and LLMs, which results in predictable scaling of linguistic performance.We show that the linguistic performance of SLMs scales up to three orders of magnitude more slowly than that of text-based LLMs.Additionally, we study the benefits of synthetic data designed to boost semantic understanding and the effects of coarser speech tokenization. Santiago Cuervo, Ricard Marxer |
EMNLP | 2 |
| 2024 | Speech Foundation Models on Intelligibility Prediction for Hearing-Impaired ListenersabstractSpeech foundation models (SFMs) have been benchmarked on many speech processing tasks, often achieving state-of-the-art performance with minimal adaptation. However, the SFM paradigm has been significantly less explored for applications of interest to the speech perception community. In this paper we present a systematic evaluation of 10 SFMs on one such application: Speech intelligibility prediction. We focus on the non-intrusive setup of the Clarity Prediction Challenge 2 (CPC2), where the task is to predict the percentage of words correctly perceived by hearing-impaired listeners from speech-in-noise recordings. We propose a simple method that learns a lightweight specialized prediction head on top of frozen SFMs to approach the problem. Our results reveal statistically significant differences in performance across SFMs. Our method resulted in the winning submission in the CPC2, demonstrating its promise for speech perception applications. Santiago Cuervo, Ricard Marxer |
ICASSP | 2 |
| 2024 | Transfer Learning from Whisper for Microscopic Intelligibility PredictionabstractInternational audience Paul Best, Santiago Cuervo, Ricard Marxer |
INTERSPEECH | 3 |
| 2024 | Investigating self-supervised speech models' ability to classify animal vocalizations: The case of gibbon's vocal signaturesabstractInternational audience Jules Cauzinille, Benoît Favre, Ricard Marxer, Dena J. Clink, Abdul Hamid Ahmad, Arnaud Rey |
INTERSPEECH | 3 |
| 2024 | TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024abstractInternational audience Joonas Kalda, Tanel Alumäe, Martin Lebourdais, Hervé Bredin, Séverin Baroudi, Ricard Marxer |
INTERSPEECH | 6 |
| 2023 | On the Benefits of Self-supervised Learned Speech Representations for Predicting Human Phonetic MisperceptionsabstractInternational audience Santiago Cuervo, Ricard Marxer |
INTERSPEECH | 2 |
| 2023 | Progress and Prospects for Spoken Language Technology: Results from Five Sexennial SurveysabstractEvery six years (since 1997), a survey has been conducted at the IEEE workshop on Automatic Speech Recognition and Understanding (ASRU). The aim has been to gain an insight into the research community's perspective on the 'progress and prospects' for spoken language technology. These surveys have been based on a set of 'statements' describing possible scenarios, and respondents are asked to estimate the year (in the future or in the past) when each given scenario might be realised. A number of the statements have appeared in multiple surveys, hence it has been possible to track changes in opinion over time. This paper presents the combined results from five such surveys, the most recent of which was conducted at ASRU-2021. The latest results reveal a dramatic increase in optimism. Roger K. Moore, Ricard Marxer |
INTERSPEECH | 2 |
| 2022 | Contrastive Prediction Strategies for Unsupervised Segmentation and Categorization of Phonemes and WordsabstractWe identify a performance trade-off between the tasks of phoneme categorization and phoneme and word segmentation in several self-supervised learning algorithms based on Contrastive Predictive Coding (CPC). Our experiments suggest that context building networks, albeit necessary for high performance on categorization tasks, harm segmentation performance by causing a temporal shift on the learned representations. Aiming to tackle this trade-off, we take inspiration from the leading approaches on segmentation and propose multi-level Aligned CPC (mACPC). It builds on Aligned CPC (ACPC), a variant of CPC which exhibits the best performance on categorization tasks, and incorporates multi-level modeling and optimization for detection of spectral changes. Our methods improve in all tested categorization metrics and achieve state-of-the-art performance in word segmentation. Santiago Cuervo, Maciej Grabias, Jan Chorowski, Grzegorz Ciesielski, Adrian Lancucki, Pawel Rychlikowski, Ricard Marxer |
ICASSP | 7 |
| 2022 | Blind Speech Separation Through Direction of Arrival Estimation Using Deep Neural Networks with a Flexibility on the Number of SpeakersabstractThis paper presents a complete framework for blind multi-speaker separation in reverberant environments using neural networks while being flexible on the number of sound sources. Under the W-disjoint orthogonality (WDO) hypothesis of speech signals, with the proposed framework, we first estimate the direction of arrival (DoA) of each dominant speech signal in each time-frequency bin using a deep neural network. While existing state-of-the-art methods use the masks associated with the different DoAs directly in the separation, we propose to use them as input features of a second deep neural network to estimate refined separation masks. Each speaker's signal is then separated using an estimated Generalized Eigenvalue (GEV) Wiener filter. This approach reduces distortion, interference, and artifacts. We assessed our framework through numerical experiments on a simulated dataset with a comprehensive analysis and on a real dataset acquired from recorded spatial room impulse responses to check if the framework generalizes well to actual cases. Numerical experiments show that our contribution outperforms the state-of-the-art with almost$3dB$to$6dB$in terms of signal to distortion ratio (SDR) and an improvement of 22% to 26% in terms of word error rate using wav2vec2 ASR model. Mohammed Hafsati, Kamil Bentounes, Ricard Marxer |
MMSP | 3 |
| 2022 | Variable-rate hierarchical CPC leads to acoustic unit discovery in speechabstractThe success of deep learning comes from its ability to capture the hierarchical structure of data by learning high-level representations defined in terms of low-level ones. In this paper we explore self-supervised learning of hierarchical representations of speech by applying multiple levels of Contrastive Predictive Coding (CPC). We observe that simply stacking two CPC models does not yield significant improvements over single-level architectures. Inspired by the fact that speech is often described as a sequence of discrete units unevenly distributed in time, we propose a model in which the output of a low-level CPC module is non-uniformly downsampled to directly minimize the loss of a high-level CPC module. The latter is designed to also enforce a prior of separability and discreteness in its representations by enforcing dissimilarity of successive high-level representations through focused negative sampling, and by quantization of the prediction targets. Accounting for the structure of the speech signal improves upon single-level CPC features and enhances the disentanglement of the learned representations, as measured by downstream speech recognition tasks, while resulting in a meaningful segmentation of the signal that closely resembles phone boundaries. Santiago Cuervo, Adrian Lancucki, Ricard Marxer, Pawel Rychlikowski, Jan Chorowski |
NeurIPS | 3 |
| 2021 | Information Retrieval for ZeroSpeech 2021: The Submission by University of WroclawabstractWe present a number of low-resource approaches to the tasks of the Zero Resource Speech Challenge 2021. We build on the unsupervised representations of speech proposed by the organizers as a baseline, derived from CPC and clustered with the k-means algorithm. We demonstrate that simple methods of refining those representations can narrow the gap, or even improve upon the solutions which use a high computational budget. The results lead to the conclusion that the CPC-derived representations are still too noisy for training language models, but stable enough for simpler forms of pattern matching and retrieval. Jan Chorowski, Grzegorz Ciesielski, Jaroslaw Dzikowski, Adrian Lancucki, Ricard Marxer, Mateusz Opala, Piotr Pusz, Pawel Rychlikowski, Michal Stypulkowski |
Interspeech | 5 |
| 2021 | Aligned Contrastive Predictive CodingabstractWe investigate the possibility of forcing a self-supervised model trained using a contrastive predictive loss to extract slowly varying latent representations. Rather than producing individual predictions for each of the future representations, the model emits a sequence of predictions shorter than that of the upcoming representations to which they will be aligned. In this way, the prediction network solves a simpler task of predicting the next symbols, but not their exact timing, while the encoding network is trained to produce piece-wise constant latent codes. We evaluate the model on a speech coding task and demonstrate that the proposed Aligned Contrastive Predictive Coding (ACPC) leads to higher linear phone prediction accuracy and lower ABX error rates, while being slightly faster to train due to the reduced number of prediction heads. Jan Chorowski, Grzegorz Ciesielski, Jaroslaw Dzikowski, Adrian Lancucki, Ricard Marxer, Mateusz Opala, Piotr Pusz, Pawel Rychlikowski, Michal Stypulkowski |
Interspeech | 5 |
| 2020 | The "ScribbleLens" Dutch Historical Handwriting CorpusabstractHistorical handwritten documents guard an important part of human knowledge only at the reach of a few scholars and experts. Recent developments in machine learning have the potential of rendering this information accessible to a larger audience. Data-driven approaches to automatic manuscript recognition require large amounts of transcribed scans to work. To this end, we introduce a new handwritten corpus based on 400-year-old, cursive, early modern Dutch documents such as ship journals and daily logbooks. This is a 1000 page collection, segmented into lines, to facilitate fully-, weakly- and un-supervised research and with textual transcriptions on 20% of the pages. Other annotations such as handwriting slant, year of origin, complexity, and writer identity have been manually added. With over 80 writers this corpus is significantly larger and more varied than other existing historical data sets such as Spanish RODRIGO. We provide train/test splits, experimental results from an automatic transcription baseline and tools to facilitate its use in deep learning research. The manuscripts span over 150 years of significant journeys by captains and traders from the Vereenigde Oost-indische Company (VOC) such as Tasman, Brouwer and Van Neck, making this resource also valuable to historians and the paleography community. Hans J. G. A. Dolfing, Jerome R. Bellegarda, Jan Chorowski, Ricard Marxer, Antoine Laurent |
ICFHR | 4 |
| 2020 | Deep Learning and Domain Transfer for Orca Vocalization DetectionabstractIn this paper, we study the difficulties of domain transfer when training deep learning models, on a specific task that is orca vocalization detection. Deep learning appears to be an answer to many sound recognition tasks in human speech analysis as well as in bioacoustics. This method allows to learn from large amounts of data, and find the best scoring way to discriminate between classes (e.g. orca vocalization and other sounds). However, to learn the perfect data representation and discrimination boundaries, all possible data configurations need to be processed. This causes problems when those configurations are ever changing (e.g. in our experiment, a change in the recording system happened to considerably disturb our previously well performing model). We thus explore approaches to compensate on the difficulties faced with domain transfer, with two convolutional neural networks (CNN) architectures, one that works in the time-frequency domain, and one that works directly on the time domain. Paul Best, Maxence Ferrari, Marion Poupard, Sébastien Paris, Ricard Marxer, Helena Symonds, Paul Spong, Hervé Glotin |
IJCNN | 5 |
| 2020 | DOCC10: Open access dataset of marine mammal transient studies and end-to-end CNN classificationabstractClassification of transients is a difficult task. In bioacoustics, almost all studies are still done with human labeling. In passive acoustic monitoring (PAM), the data to label are made up from months of continuous recordings with multiple recording stations and the time required to label everything with human labeling is longer than the next recording session will take to produce new data, even with multiple experts. To help lay a foundation for the emergence of automatic labeling of marine mammal transients, we built a dataset using weak labels from a 3TB dataset of marine mammal transients of DCLDE 2018. The DCLDE dataset was made for a click classification challenge. The new dataset has strong labels and opened a new challenge, DOCC10, whose baseline is also described in this paper. The accuracy of 71% of the baseline is already good enough to curate the large dataset, leaving only some regions of interest still to be expertised. But this is far from perfect, and there remains space for future improvement, or challenging alternative techniques. A smaller version of DOCC10 named DOCC7 is also presented. Maxence Ferrari, Hervé Glotin, Ricard Marxer, Mark Asch |
IJCNN | 3 |
| 2020 | Robust Training of Vector Quantized Bottleneck ModelsabstractIn this paper we demonstrate methods for reliable and efficient training of discrete representation using Vector-Quantized Variational Auto-Encoder models (VQ-VAEs). Discrete latent variable models have been shown to learn nontrivial representations of speech, applicable to unsupervised voice conversion and reaching state-of-the-art performance on unit discovery tasks. For unsupervised representation learning, they became viable alternatives to continuous latent variable models such as the Variational Auto-Encoder (VAE). However, training deep discrete variable models is challenging, due to the inherent non-differentiability of the discretization operation. In this paper we focus on VQ-VAE, a state-of-the-art discrete bottleneck model shown to perform on par with its continuous counterparts. It quantizes encoder outputs with on-line k-means clustering. We show that the codebook learning can suffer from poor initialization and non-stationarity of clustered encoder outputs. We demonstrate that these can be successfully overcome by increasing the learning rate for the codebook and periodic date-dependent codeword re-initialization. As a result, we achieve more robust training across different tasks, and significantly increase the usage of latent codewords even for large codebooks. This has practical benefit, for instance, in unsupervised representation learning, where large codebooks may lead to disentanglement of latent representations. Adrian Lancucki, Jan Chorowski, Guillaume Sanchez, Ricard Marxer, Nanxin Chen, Hans J. G. A. Dolfing, Sameer Khurana, Tanel Alumäe, Antoine Laurent |
IJCNN | 4 |
| 2020 | A Convolutional Deep Markov Model for Unsupervised Speech Representation LearningabstractProbabilistic Latent Variable Models (LVMs) provide an alternative to self-supervised learning approaches for linguistic representation learning from speech. LVMs admit an intuitive probabilistic interpretation where the latent structure shapes the information extracted from the signal. Even though LVMs have recently seen a renewed interest due to the introduction of Variational Autoencoders (VAEs), their use for speech representation learning remains largely unexplored. In this work, we propose Convolutional Deep Markov Model (ConvDMM), a Gaussian state-space model with non-linear emission and transition functions modelled by deep neural networks. This unsupervised model is trained using black box variational inference. A deep convolutional neural network is used as an inference network for structured variational approximation. When trained on a large scale speech dataset (LibriSpeech), ConvDMM produces features that significantly outperform multiple self-supervised feature extracting methods on linear phone classification and recognition on the Wall Street Journal dataset. Furthermore, we found that ConvDMM complements self-supervised methods like Wav2Vec and PASE, improving on the results achieved with any of the methods alone. Lastly, we find that ConvDMM features enable learning better phone recognizers than any other features in an extreme low-resource regime with few labeled training examples. Sameer Khurana, Antoine Laurent, Wei-Ning Hsu, Jan Chorowski, Adrian Lancucki, Ricard Marxer, James R. Glass |
INTERSPEECH | 6 |
| 2019 | Real-time Passive Acoustic 3D Tracking of Deep Diving Cetacean by Small Non-uniform Mobile Surface AntennaabstractDetecting and localizing the echolocation clicks of sperm whales provides insight into their diving behavior, but existing methods are limited in range, imprecise, or costly. In this work, we demonstrate that we can obtain a high definition 3D track of deep diving cetaceans from a five-channel, small-aperture hydrophone array on a moving autonomous surface vehicle (ASV), enabled by the vessel’s hydrodynamic quality and a high recording sample rate. Real-time processing is achieved by splitting our non-uniform array into two parts for time delay of arrival estimation. Resulting 3D tracks depict the behavior of the cetacean in the abyss (−1.2 km), with one position per second. This high resolution allows us to observe a correlation between the repetition rate of the predator’s biosonar and the tortuosity of its track. Our proposed mobile observatory may offer new insights about whale behavior and its foraging success close to vessel traffic. Marion Poupard, Maxence Ferrari, Jan Schlüter, Ricard Marxer, Pascale Giraudet, Valentin Barchasz, Valentin Gies, G. Pavan, Hervé Glotin |
ICASSP | 4 |
| 2018 | DNN Driven Speaker Independent Audio-Visual Mask Estimation for Speech SeparationabstractHuman auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on target speaker while filtering out other noises. In this study, we propose a novel deep neural network (DNN) based audiovisual (AV) mask estimation model. The proposed AV mask estimation model contextually integrates the temporal dynamics of both audio and noise-immune visual features for improved mask estimation and speech separation. For optimal AV features extraction and ideal binary mask (IBM) estimation, a hybrid DNN architecture is exploited to leverages the complementary strengths of a stacked long short term memory (LSTM) and convolution LSTM network. The comparative simulation results in terms of speech quality and intelligibility demonstrate significant performance improvement of our proposed AV mask estimation model as compared to audio-only and visual-only mask estimation approaches for both speaker dependent and independent scenarios. Mandar Gogate, Ahsan Adeel, Ricard Marxer, Jon Barker, Amir Hussain 0001 |
INTERSPEECH | 3 |
| 2018 | The impact of the Lombard effect on audio and visual speech recognition systemsabstractWhen producing speech in noisy backgrounds talkers reflexively adapt their speaking style in ways that increase speech-in-noise intelligibility. This adaptation, known as the Lombard effect, is likely to have an adverse effect on the performance of automatic speech recognition systems that have not been designed to anticipate it. However, previous studies of this impact have used very small amounts of data and recognition systems that lack modern adaptation strategies. This paper aims to rectify this by using a new audio-visual Lombard corpus containing speech from 54 different speakers – significantly larger than any previously available – and modern state-of-the-art speech recognition techniques. The paper is organised as three speech-in-noise recognition studies. The first examines the case in which a system is presented with Lombard speech having been exclusively trained on normal speech. It was found that the Lombard mismatch caused a significant decrease in performance even if the level of the Lombard speech was normalised to match the level of normal speech. However, the size of the mismatch was highly speaker-dependent thus explaining conflicting results presented in previous smaller studies. The second study compares systems trained in matched conditions (i.e., training and testing with the same speaking style). Here the Lombard speech affords a large increase in recognition performance. Part of this is due to the greater energy leading to a reduction in noise masking, but performance improvements persist even after the effect of signal-to-noise level difference is compensated. An analysis across speakers shows that the Lombard speech energy is spectro-temporally distributed in a way that reduces energetic masking, and this reduction in masking is associated with an increase in recognition performance. The final study repeats the first two using a recognition system training on visual speech. In the visual domain, performance differences are not confounded by differences in noise masking. It was found that in matched-conditions Lombard speech supports better recognition performance than normal speech. The benefit was consistently present across all speakers but to a varying degree. Surprisingly, the Lombard benefit was observed to a small degree even when training on mismatched non-Lombard visual speech, i.e., the increased clarity of the Lombard speech outweighed the impact of the mismatch. The paper presents two generally applicable conclusions: i) systems that are designed to operate in noise will benefit from being trained on well-matched Lombard speech data, ii) the results of speech recognition evaluations that employ artificial speech and noise mixing need to be treated with caution: they are overly-optimistic to the extent that they ignore a significant source of mismatch but at the same time overly-pessimistic in that they do not anticipate the potential increased intelligibility of the Lombard speaking style. Ricard Marxer, Jon Barker, Najwa Alghamdi, Steve C. Maddock |
Speech Commun. | 1 |
| 2017 | Binary Mask Estimation Strategies for Constrained Imputation-Based Speech EnhancementabstractInternational audience Ricard Marxer, Jon Barker |
INTERSPEECH | 1 |
| 2017 | Multi-microphone speech recognition in everyday environments
Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001 |
Comput. Speech Lang. | 2 |
| 2017 | The third 'CHiME' speech separation and recognition challenge: Analysis and outcomes
Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001 |
Comput. Speech Lang. | 2 |
| 2017 | An analysis of environment, microphone and data simulation mismatches in robust speech recognition
Emmanuel Vincent 0001, Shinji Watanabe 0001, Aditya Arie Nugraha, Jon Barker, Ricard Marxer |
Comput. Speech Lang. | 5 |
| 2016 | CloudCAST - Remote Speech Technology for Speech ProfessionalsabstractInternational audience Phil D. Green, Ricard Marxer, Stuart P. Cunningham, Heidi Christensen, Frank Rudzicz, Maria Yancheva, André Coy, Massimiliano Malavasi, Lorenzo Desideri, Fabio Tamburini |
INTERSPEECH | 2 |
| 2016 | Language Effects in Noise-Induced Word MisperceptionsabstractInternational audience María Luisa García Lecumberri, Jon Barker, Ricard Marxer, Martin Cooke |
INTERSPEECH | 3 |
| 2016 | Progress and Prospects for Spoken Language Technology: Results from Four Sexennial SurveysabstractInternational audience Roger K. Moore, Ricard Marxer |
INTERSPEECH | 2 |
| 2016 | Unsupervised Incremental Online Learning and Prediction of Musical Audio SignalsabstractGuided by the idea that musical human-computer interaction may become more effective, intuitive, and creative when basing its computer part on cognitively more plausible learning principles, we employ unsupervised incremental online learning (i.e. clustering) to build a system that predicts the next event in a musical sequence, given as audio input. The flow of the system is as follows: 1) segmentation by onset detection, 2) timbre representation of each segment by Mel frequency cepstrum coefficients, 3) discretization by incremental clustering, yielding a tree of different sound classes (e.g. timbre categories/instruments) that can grow or shrink on the fly driven by the instantaneous sound events, resulting in a discrete symbol sequence, 4) extraction of statistical regularities of the symbol sequence, using hierarchical N-grams and the newly introduced conceptual Boltzmann machine that adapt to the dynamically changing clustering tree in 3) , and 5) prediction of the next sound event in the sequence, given the last $n$ previous events. The system's robustness is assessed with respect to complexity and noisiness of the signal. Clustering in isolation yields an adjusted Rand index (ARI) of 82.7%/85.7% for data sets of singing voice and drums. Onset detection jointly with clustering achieve an ARI of 81.3%/76.3% and the prediction of the entire system yields an ARI of 27.2%/39.2%. Ricard Marxer, Hendrik Purwins |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | The third 'CHiME' speech separation and recognition challenge: Dataset, task and baselinesabstractThe CHiME challenge series aims to advance far field speech recognition technology by promoting research at the interface of signal processing and automatic speech recognition. This paper presents the design and outcomes of the 3rd CHiME Challenge, which targets the performance of automatic speech recognition in a real-world, commercially-motivated scenario: a person talking to a tablet device that has been fitted with a six-channel microphone array. The paper describes the data collection, the task definition and the baseline systems for data simulation, enhancement and recognition. The paper then presents an overview of the 26 systems that were submitted to the challenge focusing on the strategies that proved to be most successful relative to the MVDR array processing and DNN acoustic modeling reference system. Challenge findings related to the role of simulated data in system training and evaluation are discussed. Jon Barker, Ricard Marxer, Emmanuel Vincent 0001, Shinji Watanabe 0001 |
ASRU | 2 |
| 2015 | Exploiting synchrony spectra and deep neural networks for noise-robust automatic speech recognitionabstractThis paper presents a novel system that exploits synchrony spectra and deep neural networks (DNNs) for automatic speech recognition (ASR) in challenging noisy environments. Synchrony spectra measure the extent to which each frequency channel in an auditory model is entrained to a particular pitch period, and they are used together with F0 estimates either in a DNN for time-frequency (T-M) mask estimation or to augment the input features for a DNN-based ASR system. The proposed approach was evaluated in the context of the CHiME 3 Challenge. Our experiments show that the synchrony spectra features work best when augmenting the input features to the DNN-based ASR system. Compared to the CHiME-3 baseline system, our best system provides a word error rate (WER) reduction of more than 14% absolute and achieved a WER of 18.56% on the evaluation test set. Ning Ma 0002, Ricard Marxer, Jon Barker, Guy J. Brown |
ASRU | 2 |
| 2015 | A framework for the evaluation of microscopic intelligibility modelsabstractInternational audience Ricard Marxer, Martin Cooke, Jon Barker |
INTERSPEECH | 1 |
| 2015 | Knowledge transfer between speakers for personalised dialogue managementabstractModel-free reinforcement learning has been shown to be a promising data driven approach for automatic dialogue policy optimization, but a relatively large amount of dialogue interactions is needed before the system reaches reasonable performance.Recently, Gaussian process based reinforcement learning methods have been shown to reduce the number of dialogues needed to reach optimal performance, and pre-training the policy with data gathered from different dialogue systems has further reduced this amount.Following this idea, a dialogue system designed for a single speaker can be initialised with data from other speakers, but if the dynamics of the speakers are very different the model will have a poor performance.When data gathered from different speakers is available, selecting the data from the most similar ones might improve the performance.We propose a method which automatically selects the data to transfer by defining a similarity measure between speakers, and uses this measure to weight the influence of the data from each speaker in the policy model.The methods are tested by simulating users with different severities of dysarthria interacting with a voice enabled environmental control system. Iñigo Casanueva, Thomas Hain, Heidi Christensen, Ricard Marxer, Phil D. Green |
SIGDIAL Conference | 4 |
| 2012 | Combining a harmonic-based NMF decomposition with transient analysis for instantaneous percussion separationabstractMany recent approaches on musical source separation rely on model-based inference methods that take into account the signal's harmonic structure. To address the particular case of instantaneous percussion separation, we propose a method that combines a harmonic-based decomposition using a Non-negative Matrix Factorization (NMF) algorithm, with the transient analysis of spectral peaks from a single audio frame. The signal model allows the estimation of harmonic and non-harmonic sources. Later, as shown in the evaluation, adding transient peak information improves the Signal-to-Distortion Ratio (SDR). Compared to other existing methods, this approach achieves a comparable performance, being suitable at the same time for low-latency conditions. Jordi Janer, Ricard Marxer, Keita Arimoto |
ICASSP | 2 |
| 2012 | A Tikhonov regularization method for spectrum decomposition in low latency audio source separationabstractWe present the use of a Tikhonov regularization based method, as an alternative to the Non-negative Matrix Factorization (NMF) approach, for source separation in professional audio recordings. This method is a direct and computationally less expensive solution to the problem, which makes it interesting in low latency scenarios. The technique sacrifices the non-negativity constraint that characterizes NMF in exchange for a closed-form solution to the problem of spectrum factorization. We quantitatively evaluated it in terms of reconstruction and separation quality on a dataset of excerpts of professionally recorded songs with singing voice. Results show that the the proposed approach achieves similar quality to that of NMF. Ricard Marxer, Jordi Janer |
ICASSP | 1 |
| 2009 | What/when causal expectation modelling applied to audio signalsabstractA causal system to represent a stream of music into musical events, and to generate further expected events, is presented. Starting from an auditory front-end that extracts low-level (i.e. MFCC) and mid-level features such as onsets and beats, an unsupervised clustering process builds and maintains a set of symbols aimed at representing musical stream events using both timbre and time descriptions. The time events are represented using inter-onset intervals relative to the beats. These symbols are then processed by an expectation module using Predictive Partial Match, a multiscale technique based on N-grams. To characterise the ability of the system to generate an expectation that matches both ground truth and system transcription, we introduce several measures that take into account the uncertainty associated with the unsupervised encoding of the musical sequence. The system is evaluated using a subset of the ENST-drums database of annotated drum recordings. We compare three approaches to combine timing (when) and timbre (what) expectation. In our experiments, we show that the induced representation is useful for generating expectation patterns in a causal way. Amaury Hazan, Ricard Marxer, Paul Brossier, Hendrik Purwins, Perfecto Herrera, Xavier Serra |
Connect. Sci. | 2 |