EDBT 2026 Demo / reviewers in the wild / expert
Panagiota Karanasou
dblp:90/8340 · also Penny Karanasou
· DBLP profile ↗
24ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0003-1939-4161ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 18 · 8 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Long-Short Decision Transformer: Bridging Global and Local Dependencies for Generalized Decision-MakingabstractDecision Transformers (DTs) effectively capture long-range dependencies using self-attention but struggle with fine-grained local relationships, especially the Markovian properties in many offline-RL datasets. Conversely, Decision Convformer (DC) utilizes convolutional filters for capturing local patterns but shows limitations in tasks demanding long-term dependencies, such as Maze2d. To address these limitations and leverage both strengths, we propose the Long-Short Decision Transformer (LSDT), a general-purpose architecture to effectively capture global and local dependencies across two specialized parallel branches (self-attention and convolution). We explore how these branches complement each other by modeling various ranged dependencies across different environments, and compare it against other baselines. Experimental results demonstrate our LSDT achieves state-of-the-art performance and notable gains over the standard DT in D4RL offline RL benchmark. Leveraging the parallel architecture, LSDT performs consistently on diverse datasets, including Markovian and non-Markovian. We also demonstrate the flexibility of LSDT's architecture, where its specialized branches can be replaced or integrated into models like DC to improve their performance in capturing diverse dependencies. Finally, we also highlight the role of goal states in improving decision-making for goal-reaching tasks like Antmaze. Panagiota Karanasou, Pengyuan Wei, Elia Gatti, Diego Martínez 0001, Dimitrios Kanoulas |
ICLR | 2 |
| 2023 | eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer
Ammar Abbas, Sri Karlapati, Bastian Schnell, Panagiota Karanasou, Marcel Granero Moya, Amith Nagaraj, Ayman Boustati, Nicole Peinelt, Alexis Moinet, Thomas Drugman |
INTERSPEECH | 4 |
| 2022 | CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody TransferabstractIn this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually appropriate prosody, and c) transferring prosody at fine-grained level between any pair of seen speakers.We do this by activating distinct parts of the network for different tasks.We train our model using a novel approach to two-stage training.In Stage I, the model learns speaker-independent word-level prosody representations from speech which it uses for many-to-many finegrained prosody transfer.In Stage II, we learn to predict these prosody representations using the contextual information available in text, thereby, enabling multi-speaker TTS with contextually appropriate prosody.We compare CC2 to two strong baselines, one in TTS with contextually appropriate prosody, and one in fine-grained prosody transfer.CC2 reduces the gap in naturalness between our baseline and copy-synthesised speech by 22.79%.In fine-grained prosody transfer evaluations, it obtains a relative improvement of 33.15% in target speaker similarity. Sri Karlapati, Panagiota Karanasou, Mateusz Lajszczak, Syed Ammar Abbas, Alexis Moinet, Peter Makarov, Ray Li, Arent van Korlaar, Simon Slangen, Thomas Drugman |
INTERSPEECH | 2 |
| 2022 | Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody
Peter Makarov, Syed Ammar Abbas, Mateusz Lajszczak, Arnaud Joly, Sri Karlapati, Alexis Moinet, Thomas Drugman, Panagiota Karanasou |
INTERSPEECH | 8 |
| 2022 | Cross-lingual Style Transfer with Conditional Prior VAE and Style Loss
Dino Rattcliffe, Alex Mansbridge, Panagiota Karanasou, Alexis Moinet, Marius Cotescu |
INTERSPEECH | 4 |
| 2021 | Camp: A Two-Stage Approach to Modelling Prosody in ContextabstractProsody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the acoustic signal (e.g. segmental information and background noise); (2) determining appropriate prosody without sufficient context is an ill-posed problem. In this paper, we propose solutions to both these issues. To mitigate the challenge of modelling a slow-varying signal, we learn to disentangle prosodic information using a word level representation. To alleviate the ill-posed nature of prosody modelling, we use syntactic and semantic information derived from text to learn a context-dependent prior over our prosodic space. Our context-aware model of prosody (CAMP) outperforms the state-of-the-art technique, closing the gap with natural speech by 26%. We also find that replacing attention with a jointly-trained duration model improves prosody significantly. Zack Hodari, Alexis Moinet, Sri Karlapati, Jaime Lorenzo-Trueba, Thomas Merritt, Arnaud Joly, Ammar Abbas, Panagiota Karanasou, Thomas Drugman |
ICASSP | 8 |
| 2021 | Prosodic Representation Learning and Contextual Sampling for Neural Text-to-SpeechabstractIn this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn a prosodic distribution at the sentence level from mel-spectrograms available during training. In Stage II, we propose a novel method to sample from this learnt prosodic distribution using the contextual information available in text. To do this, we use BERT on text, and graph-attention networks on parse trees extracted from text. We show a statistically significant relative improvement of 13.2% in naturalness over a strong baseline when compared to recordings. We also conduct an ablation study on variations of our sampling technique, and show a statistically significant improvement over the baseline in each case. Sri Karlapati, Ammar Abbas, Zack Hodari, Alexis Moinet, Arnaud Joly, Panagiota Karanasou, Thomas Drugman |
ICASSP | 6 |
| 2021 | A Learned Conditional Prior for the VAE Acoustic Space of a TTS SystemabstractMany factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow multiple renditions of the same sentence via sampling. The degree of prosodic variability depends heavily on the prior that is used when sampling. In this paper, we propose a novel method to compute an informative prior for the VAE latent space of a neural text-to-speech (TTS) system. By doing so, we aim to sample with more prosodic variability, while gaining controllability over the latent space's structure. By using as prior the posterior distribution of a secondary VAE, which we condition on a speaker vector, we can sample from the primary VAE taking explicitly the conditioning into account and resulting in samples from a specific region of the latent space for each condition (i.e. speaker). A formal preference test demonstrates significant preference of the proposed approach over standard Conditional VAE. We also provide visualisations of the latent space where well-separated condition-specific clusters appear, as well as ablation studies to better understand the behaviour of the system. Panagiota Karanasou, Sri Karlapati, Alexis Moinet, Arnaud Joly, Ammar Abbas, Simon Slangen, Jaime Lorenzo-Trueba, Thomas Drugman |
Interspeech | 1 |
| 2018 | Improving Interpretability and Regularization in Deep LearningabstractDeep learning approaches yield state-of-the-art performance in a range of tasks, including automatic speech recognition. However, the highly distributed representation in a deep neural network (DNN) or other network variations is difficult to analyze, making further parameter interpretation and regularization challenging. This paper presents a regularization scheme acting on the activation function output to improve the network interpretability and regularization. The proposed approach, referred to as activation regularization, encourages activation function outputs to satisfy a target pattern. By defining appropriate target patterns, different learning concepts can be imposed on the network. This method can aid network interpretability and also has the potential to reduce overfitting. The scheme is evaluated on several continuous speech recognition tasks: the Wall Street Journal continuous speech recognition task, eight conversational telephone speech tasks from the IARPA Babel program and a U.S. English broadcast news task. On all the tasks, the activation regularization achieved consistent performance gains over the standard DNN baselines. Chunyang Wu, Mark J. F. Gales, Anton Ragni, Panagiota Karanasou, Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | I-Vectors and Structured Neural Networks for Rapid Adaptation of Acoustic ModelsabstractA lot of interest has been risen in the last years on the adaptation of deep neural network (DNN) acoustic models, as the latter become the state-of-art in automatic speech recognition. This work focuses on approaches that allow for rapid and robust adaptation of such models. First, i-vectors are added to the DNN input as speaker-informed features. An informative prior is introduced to i-vector estimation to improve the robustness to limited adaptation data. I-vectors are then combined with a structured adaptive DNN, the multibasis adaptive neural network (MBANN), and the complementarity of these adaptation techniques is investigated. Moreover, i-vectors are used to predict the MBANN transforms, avoiding the initial decoding pass and alignment. These approaches are evaluated on a U.S. English Broadcast News (BN) transcription task with two distinct sets of test data. The first, from the BN task and BN-style Youtube videos, yields test data acoustically matched to the training data, while the second set is from acoustically mismatched Youtube videos of diverse context. The performance gains from these schemes are found to be sensitive to the level of mismatch between training and test sets. The MBANN system combined with i-vector input achieves best performance for BN test sets. The i-vector-based predictive MBANN scheme is proven to be more robust to acoustically mismatched conditions and outperforms the other adaptation schemes in such scenarios. Panagiota Karanasou, Chunyang Wu, Mark J. F. Gales, Philip C. Woodland |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Improved DNN-based segmentation for multi-genre broadcast audioabstractAutomatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for multi-genre broadcast audio with deep neural network (DNN) based speech/non-speech detection. A further stage of change-point detection and clustering is used to obtain homogeneous segments. Suitable DNN inputs, context window sizes and architectures are studied with a series of experiments using a large corpus of MGB television audio. For MGB transcription, the improved segmenter yields roughly half the increase in word error rate, over manual segmentation, compared to the baseline DNN segmenter supplied for the 2015 ASRU MGB challenge. Chao Zhang 0031, Philip C. Woodland, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin, Xunying Liu, Yanmin Qian |
ICASSP | 5 |
| 2016 | Combining i-vector representation and structured neural networks for rapid adaptationabstractRapid adaptation of deep neural networks (DNNs) with limited unsupervised data remains a significant challenge. This paper investigates the combination of two schemes that have been proposed to address this problem: i-vector representations and multi-basis adaptive neural networks (MBANNs). Two approaches for combining these schemes together are described. The first uses i-vectors as one of the input features to the MBANN. The purpose is to combine the speaker representation of the i-vector with the network interpolation of the MBANN scheme. The second approach aims to reduce the computational cost, and improve the robustness to hypothesis errors, of the MBANN scheme. Here i-vectors are used to predict the interpolation weights of the MBANN scheme. This removes the need for an initial decoding pass, and alignment, which was previously used. These approaches are evaluated using acoustic and language models trained on a U.S. English Broadcast News (BN) transcription task. Two distinct sets of test data are examined. The first from the BN task, yields test data acoustically matched to the training data. The second, acoustically mismatched, set is from Youtube videos. The performance gains from these schemes is found to be sensitive to the level of mismatch between training and test. Chunyang Wu, Panagiota Karanasou, Mark J. F. Gales |
ICASSP | 2 |
| 2016 | Selection of Multi-Genre Broadcast Data for the Training of Automatic Speech Recognition SystemsabstractThis paper compares schemes for the selection of multi-genre broadcast data and corresponding transcriptions for speech recognition model training. Selections of the same amount of data (700 hours) from lightly supervised alignments based on the same original subtitle transcripts are compared. Data segments were selected according to a maximum phone matched error rate between the lightly supervised decoding and the original transcript. The data selected with an improved lightly supervised system yields lower word error rates (WERs). Detailed comparisons of the data selected on carefully transcribed development data show how the selected portions match the true phone error rate for each genre. From a broader perspective, it is shown that for different genres, either the original subtitles or the lightly supervised output should be used for model training and a suitable combination yields further reductions in final WER. Pierre Lanchantin, Mark J. F. Gales, Panagiota Karanasou, Xunying Liu, Yanman Qian, Philip C. Woodland, Chao Zhang 0031 |
INTERSPEECH | 3 |
| 2016 | Stimulated Deep Neural Network for Speech RecognitionabstractDeep neural networks (DNNs) and deep learning approaches yield state-of-the-art performance in a range of tasks, including speech recognition.However, the parameters of the network are hard to analyze, making network regularization and robust adaptation challenging.Stimulated training has recently been proposed to address this problem by encouraging the node activation outputs in regions of the network to be related.This kind of information aids visualization of the network, but also has the potential to improve regularization and adaptation.This paper investigates stimulated training of DNNs for both of these options.These schemes take advantage of the smoothness constraints that stimulated training offers.The approaches are evaluated on two large vocabulary speech recognition tasks: a U.S. English broadcast news (BN) task and a Javanese conversational telephone speech task from the IARPA Babel program.Stimulated DNN training acquires consistent performance gains on both tasks over unstimulated baselines.On the BN task, the proposed smoothing approach is also applied to rapid adaptation, again outperforming the standard adaptation scheme. Chunyang Wu, Panagiota Karanasou, Mark J. F. Gales, Khe Chai Sim |
INTERSPEECH | 2 |
| 2015 | Speaker diarisation and longitudinal linking in multi-genre broadcast dataabstractThis paper presents a multi-stage speaker diarisation system with longitudinal Linking developed on BBC multi-genre data for the 2015 Multi-Genre Broadcast (MGB) challenge. The basic speaker diarisation system draws on techniques from the Cambridge March 2005 system with a new deep neural network (DNN)-based speech/non speech segmenter. A newly developed linking stage is next added to the basic diarisation output aiming at the identification of speakers across multiple episodes of the same series. The longitudinal constraint imposes an incremental processing of the episodes, where speaker labels for each episode can be obtained using only material from the episode in question, and those broadcast earlier in time. The nature of the data as well as the longitudinal linking constraint position this diarisation task as a new open-research topic, and a particularly challenging one. Different linking clustering metrics are compared and the lowest within-episode and cross-episode DER scores are achieved on the MGB challenge evaluation set. Panagiota Karanasou, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Yanmin Qian, Philip C. Woodland, Chao Zhang 0031 |
ASRU | 1 |
| 2015 | The development of the cambridge university alignment systems for the multi-genre broadcast challengeabstractWe describe the alignment systems developed both for the preparation of data for the Multi-Genre Broadcast (MGB) challenge and for our participation in the transcription and alignment tasks. Captions of varying quality are aligned with the audio of TV shows that range from few minutes long to more than six hours. Lightly supervised decoding is performed on the audio and the output text is aligned with the original text transcript. Reliable split points are found and the resulting text chunks are force-aligned with the corresponding audio segments. Confidence scores are associated with the aligned data. Multiple refinements — including audio segmentation based on deep neural networks (DNNs) and the use of DNN-based acoustic models — were used to improve the performance. The final MGB alignment system had the highest F-measure value on the evaluation data. Pierre Lanchantin, Mark J. F. Gales, Panagiota Karanasou, Xunying Liu, Yanmin Qian, Philip C. Woodland, Chao Zhang 0031 |
ASRU | 3 |
| 2015 | Cambridge university transcription systems for the multi-genre broadcast challengeabstractWe describe the development of our speech-to-text transcription systems for the 2015 Multi-Genre Broadcast (MGB) challenge. Key features of the systems are: a segmentation system based on deep neural networks (DNNs); the use of HTK 3.5 for building DNN-based hybrid and tandem acoustic models and the use of these models in a joint decoding framework; techniques for adaptation of DNN based acoustic models including parameterised activation function adaptation; alternative acoustic models built using Kaldi; and recurrent neural network language models (RNNLMs) and RNNLM adaptation. The same language models were used with both HTK and Kaldi acoustic models and various combined systems built. The final systems had the lowest error rates on the evaluation data. Philip C. Woodland, Xunying Liu, Yanmin Qian, Chao Zhang 0031, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin |
ASRU | 6 |
| 2015 | An investigation into speaker informed DNN front-end for LVCSRabstractDeep Neural Network (DNN) has become a standard method in many ASR tasks. Recently there is considerable interest in “informed training” of DNNs, where DNN input is augmented with auxiliary codes, such as i-vectors, speaker codes, speaker separation bottleneck (SSBN) features, etc. This paper compares different speaker informed DNN training methods in LVCSR task. We discuss mathematical equivalence between speaker informed DNN training and “bias adaptation” which uses speaker dependent biases, and give detailed analysis on influential factors such as dimension, discrimination and stability of auxiliary codes. The analysis is supported by experiments on a meeting recognition task using bottleneck feature based system. Results show that i-vector based adaptation is also effective in bottleneck feature based system (not just hybrid systems). However all tested methods show poor generalisation to unseen speakers. We introduce a system based on speaker classification followed by speaker adaptation of biases, which yields equivalent performance to an i-vector based system with 10.4% relative improvement over baseline on seen speakers. The new approach can serve as a fast alternative especially for short utterances. Yulan Liu, Panagiota Karanasou, Thomas Hain |
ICASSP | 2 |
| 2015 | I-vector estimation using informative priors for adaptation of deep neural networksabstractThis is the author accepted manuscript. The final version is available from ISCA via http://www.isca-speech.org/archive/interspeech_2015/i15_2872.html Supporting data for this paper is available at the http://www.repository.cam.ac.uk/handle/1810/248387 data repository. Panagiota Karanasou, Mark J. F. Gales, Philip C. Woodland |
INTERSPEECH | 1 |
| 2014 | Adaptation of deep neural network acoustic models using factorised i-vectorsabstractThe use of deep neural networks (DNNs) in a hybrid configuration is becoming increasingly popular and successful for speech recognition. One issue with these systems is how to efficiently adapt them to reflect an individual speaker or noise condition. Recently speaker i-vectors have been successfully used as an additional input feature for unsupervised speaker adaptation. In this work the use of i-vectors for adaptation is extended to incorporate acoustic factorisation. In particular, separate i-vectors are computed to represent speaker and acoustic environment. By ensuring "orthogonality" between the individual factor representations it is possible to represent a wide range of speaker and environment pairs by simply combining i-vectors from a particular speaker and a particular environment. In this paper the i-vectors are viewed as the weights of a cluster adaptive training (CAT) system, where the underlying models are GMMs rather than HMMs. This allows the factorisation approaches developed for CAT to be directly applied. Initial experiments were conducted on a noise distorted version of the WSJ corpus. Compared to standard speaker-based i-vector adaptation, factorised i-vectors showed performance gains. Panagiota Karanasou, Yongqiang Wang 0006, Mark J. F. Gales, Philip C. Woodland |
INTERSPEECH | 1 |
| 2013 | Discriminative training of a phoneme confusion model for a dynamic lexicon in ASRabstractInternational audience Panagiota Karanasou, François Yvon, Thomas Lavergne, Lori Lamel |
INTERSPEECH | 1 |
| 2012 | Discriminatively trained phoneme confusion model for keyword spottingabstractKeyword Spotting (KWS) aims at detecting speech segments that contain a given query within large amounts of audio data. Typically, a speech recognizer is involved in a first indexing step. One of the challenges of KWS is how to handle recognition errors and out-of-vocabulary (OOV) terms. This work proposes the use of discriminative training to construct a phoneme confusion model, which expands the phonemic index of a KWS system by adding phonemic variation to handle the abovementioned problems. The objective function that is optimized is the Figure of Merit (FOM), which is directly related to the KWS performance. The experiments conducted on English data sets show some improvement on the FOM and are promising for the use of such technique. Panagiota Karanasou, Lukás Burget, Dimitra Vergyri, Murat Akbacak, Arindam Mandal |
INTERSPEECH | 1 |
| 2011 | Automatic Generation of a Pronunciation Dictionary with Rich Variation Coverage Using SMT Methods
Panagiota Karanasou, Lori Lamel |
CICLing (2) | 1 |
| 2011 | Pronunciation variants generation using SMT-inspired approachesabstractEnriching a pronunciation dictionary with phonological variation is a challenging task, not yet solved despite several decades of research, in particular for speech-to-text transcription of real world data where it is important to cover different pronunciation variants. This paper proposes two alternative methods, inspired by machine translation, to derive pronunciation variants from an initial lexicon with limited variations. In the first case, an n-best pronunciation list is extracted directly from a machine translation tool, used as a grapheme-to-phoneme (g2p) converter. The second is a novel method based on a pivot approach, previously used for the paraphrase extraction task, and here applied as a post-processing step to the g2p converter. Some preliminary speech recognition experiments with the automatically generated pronunciation variants are reported using Quaero development data. Panagiota Karanasou, Lori Lamel |
ICASSP | 1 |