Adriana Cornelia Stan

dblp:43/9407 · also Adriana Stan · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0003-2894-5770ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Unmasking real-world audio deepfakes: A data-centric approach
abstract
5343
David Combei, Adriana Cornelia Stan, Dan Oneata, Nicolas M. Müller, Horia Cucu
INTERSPEECH2
2025 Replay Attacks Against Audio Deepfake Detection
abstract
2245
Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, Adriana Cornelia Stan, Aditya Tirumala Bukkapatnam, Karla Pizzi, Alexander Wagner, Philip Sperl
INTERSPEECH4
2025 TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
abstract
Deepfake detection has gained significant attention across audio, text, and image modalities, with high accuracy in distinguishing real from fake. However, identifying the exact source--such as the system or model behind a deepfake--remains a less studied problem. In this paper, we take a significant step forward in audio deepfake model attribution or source tracing by proposing a training-free, green AI approach based entirely on k-Nearest Neighbors (kNN). Leveraging a pre-trained self-supervised learning (SSL) model, we show that grouping samples from the same generator is straightforward--we obtain an 0.93 F1-score across five deepfake datasets. The method also demonstrates strong out-of-domain (OOD) detection, effectively identifying samples from unseen models at an F1-score of 0.84. We further analyse these results in a multi-dimensional approach and provide additional insights. All code and data protocols used in this work are available in our open repository: https://github.com/adrianastan/tada/.
Adriana Cornelia Stan, David Combei, Dan Oneata, Horia Cucu
INTERSPEECH1
2024 Towards generalisable and calibrated audio deepfake detection with self-supervised representations
abstract
Generalisation—the ability of a model to perform well on unseen data—is crucial for building reliable deepfake detectors. However, recent studies have shown that the current audio deepfake models fall short of this desideratum. In this work we investigate the potential of pretrained self-supervised representations in building general and calibrated audio deepfake detection models. We show that large frozen representations coupled with a simple logistic regression classifier are extremely effective in achieving strong generalisation capabilities: compared to the RawNet2 model, this approach reduces the equal error rate from 30.9% to 8.8% on a benchmark of eight deepfake datasets, while learning less than 2k parameters. Moreover, the proposed method produces considerably more reliable predictions compared to previous approaches making it more suitable for realistic use.
Octavian Pascu, Adriana Cornelia Stan, Dan Oneata, Elisabeta Oneata, Horia Cucu
INTERSPEECH2
2023 RoLEX: The development of an extended Romanian lexical dataset and its evaluation at predicting concurrent lexical information
abstract
Abstract In this article, we introduce an extended, freely available resource for the Romanian language, named RoLEX. The dataset was developed mainly for speech processing applications, yet its applicability extends beyond this domain. RoLEX includes over 330,000 curated entries with information regarding lemma, morphosyntactic description, syllabification, lexical stress and phonemic transcription. The process of selecting the list of word entries and semi-automatically annotating the complete lexical information associated with each of the entries is thoroughly described. The dataset’s inherent knowledge is then evaluated in a task of concurrent prediction of syllabification, lexical stress marking and phonemic transcription. The evaluation looked into several dataset design factors, such as the minimum viable number of entries for correct prediction, the optimisation of the minimum number of required entries through expert selection and the augmentation of the input with morphosyntactic information, as well as the influence of each task in the overall accuracy. The best results were obtained when the orthographic form of the entries was augmented with the complete morphosyntactic tags. A word error rate of 3.08% and a character error rate of 1.08% were obtained this way. We show that using a carefully selected subset of entries for training can result in a similar performance to the performance obtained by a larger set of randomly selected entries (twice as many). In terms of prediction complexity, the lexical stress marking posed most problems and accounts for around 60% of the errors in the predicted sequence.
Beáta Lorincz, Elena Irimia, Adriana Cornelia Stan, Verginica Barbu Mititelu
Nat. Lang. Eng.3
2022 The ZevoMOS entry to VoiceMOS Challenge 2022
abstract
This paper introduces the ZevoMOS entry to the main track of the VoiceMOS Challenge 2022.The ZevoMOS submission is based on a two-step finetuning of pretrained self-supervised learning (SSL) speech models.The first step uses a task of classifying natural versus synthetic speech, while the second step's task is to predict the MOS scores associated with each training sample.The results of the finetuning process are then combined with the confidence scores extracted from an automatic speech recognition model, as well as the raw embeddings of the training samples obtained from a wav2vec SSL speech model.The team id assigned to the ZevoMOS system within the VoiceMOS Challenge is T01.The submission was placed on the 14th place with respect to the system-level SRCC, and on the 9th place with respect to the utterance-level MSE.The paper also introduces additional evaluations of the intermediate results.
Adriana Cornelia Stan
INTERSPEECH1
2022 Gamification-Based Tools Embedded in the Helios Educational Platform
Cosmin Striletchi, Adriana Cornelia Stan, Eusebiu Jecan
WorldCIST (2)2
2021 An objective evaluation of the effects of recording conditions and speaker characteristics in multi-speaker deep neural speech synthesis
abstract
Multi-speaker spoken datasets enable the creation of text-to-speech synthesis (TTS) systems which can output several voice identities. The multi-speaker (MSPK) scenario also enables the use of fewer training samples per speaker. However, in the resulting acoustic model, not all speakers exhibit the same synthetic quality, and some of the voice identities cannot be used at all. In this paper we evaluate the influence of the recording conditions, speaker gender, and speaker particularities over the quality of the synthesised output of a deep neural TTS architecture, namely Tacotron2. The evaluation is possible due to the use of a large Romanian parallel spoken corpus containing over 81 hours of data. Within this setup, we also evaluate the influence of different types of text representations: orthographic, phonetic, and phonetic extended with syllable boundaries and lexical stress markings. We evaluate the results of the MSPK system using the objective measures of equal error rate (EER) and word error rate (WER), and also look into the distances between natural and synthesised t-SNE projections of the embeddings computed by an accurate speaker verification network. The results show that there is indeed a large correlation between the recording conditions and the speaker’s synthetic voice quality. The speaker gender does not influence the output, and that extending the input text representation with syllable boundaries and lexical stress information does not equally enhance the generated audio across all speaker identities. The visualisation of the t-SNE projections of the natural and synthesised speaker embeddings show that the acoustic model shifts some of the speakers’ neural representation, but not all of them. As a result, these speakers have lower performances of the output speech.
Beáta Lorincz, Adriana Cornelia Stan, Mircea Giurgiu
KES2
2021 An Evaluation of Word-Level Confidence Estimation for End-to-End Automatic Speech Recognition
abstract
Quantifying the confidence (or conversely the uncertainty) of a prediction is a highly desirable trait of an automatic system, as it improves the robustness and usefulness in downstream tasks. In this paper we investigate confidence estimation for end-to-end automatic speech recognition (ASR). Previous work has addressed confidence measures for lattice-based ASR, while current machine learning research mostly focuses on confidence measures for unstructured deep learning. However, as the ASR systems are increasingly being built upon deep end-to-end methods, there is little work that tries to develop confidence measures in this context. We fill this gap by providing an extensive benchmark of popular confidence methods on four well-known speech datasets. There are two challenges we overcome in adapting existing methods: working on structured data (sequences) and obtaining confidences at a coarser level than the predictions (words instead of tokens). Our results suggest that a strong baseline can be obtained by scaling the logits by a learnt temperature, followed by estimating the confidence as the negative entropy of the predictive distribution and, finally, sum pooling to aggregate at word level.
Dan Oneata, Alexandru Caranica, Adriana Cornelia Stan, Horia Cucu
SLT3
2020 RECOApy: Data Recording, Pre-Processing and Phonetic Transcription for End-to-End Speech-Based Applications
abstract
Deep learning enables the development of efficient end-to-end speech processing applications while bypassing the need for expert linguistic and signal processing features. Yet, recent studies show that good quality speech resources and phonetic transcription of the training data can enhance the results of these applications. In this paper, the RECOApy tool is introduced. RECOApy streamlines the steps of data recording and pre-processing required in end-to-end speech-based applications. The tool implements an easy-to-use interface for prompted speech recording, spectrogram and waveform analysis, utterance-level normalisation and silence trimming, as well grapheme-to-phoneme conversion of the prompts in eight languages: Czech, English, French, German, Italian, Polish, Romanian and Spanish. The grapheme-to-phoneme (G2P) converters are deep neural network (DNN) based architectures trained on lexicons extracted from the Wiktionary online collaborative resource. With the different degree of orthographic transparency, as well as the varying amount of phonetic entries across the languages, the DNN's hyperparameters are optimised with an evolution strategy. The phoneme and word error rates of the resulting G2P converters are presented and discussed. The tool, the processed phonetic lexicons and trained G2P models are made freely available.
Adriana Cornelia Stan
INTERSPEECH1
2019 All Together Now: The Living Audio Dataset
David A. Braude, Matthew P. Aylett, Caoimhín Laoide-Kemp, Simone Ashby, Kristen M. Scott, Brian Ó Raghallaigh, Anna Braudo, Alex Brouwer, Adriana Cornelia Stan
INTERSPEECH9
2016 Blind speech segmentation using spectrogram image-based features and Mel cepstral coefficients
abstract
This paper introduces a novel method for blind speech segmentation at a phone level based on image processing. We consider the spectrogram of the waveform of an utterance as an image and hypothesize that its striping defects, i.e. discontinuities, appear due to phone boundaries. Using a simple image destriping algorithm these discontinuities are found. To discover phone transitions which are not as salient in the image, we compute spectral changes derived from the time evolution of Mel cepstral parametrisation of speech. These socalled image-based and acoustic features are then combined to form a mixed probability function, whose values indicate the likelihood of a phone boundary being located at the corresponding time frame. The method is completely unsupervised and achieves an accuracy of 75.59% at a -3.26% over-segmentation rate, yielding an F-measure of 0.76 and an 0.80 R-value on the TIMIT dataset.
Adriana Cornelia Stan, Cassia Valentini-Botinhao, Bogdan Orza, Mircea Giurgiu
SLT1
2016 ALISA: An automatic lightly supervised speech segmentation and alignment tool
Adriana Cornelia Stan, Yoshitaka Mamiya, Junichi Yamagishi, Peter Bell 0001, Oliver Watts, Robert A. J. Clark, Simon King 0001
Comput. Speech Lang.1
2014 Neural net word representations for phrase-break prediction without a part of speech tagger
abstract
The use of shared projection neural nets of the sort used in language modelling is proposed as a way of sharing parameters between multiple text-to-speech system components. We experiment with pretraining the weights of such a shared projection on an auxiliary language modelling task and then apply the resulting word representations to the task of phrase-break prediction. Doing so allows us to build phrase-break predictors that rival conventional systems without any reliance on conventional knowledge-based resources such as part of speech taggers.
Oliver Watts, Siva Reddy Gangireddy, Junichi Yamagishi, Simon King 0001, Steve Renals, Adriana Cornelia Stan, Mircea Giurgiu
ICASSP6
2014 RSS-TOBI - A Prosodically Enhanced Romanian Speech Corpus
Tiberiu Boros, Adriana Cornelia Stan, Oliver Watts, Stefan Daniel Dumitrescu
LREC2
2013 Lightly supervised GMM VAD to use audiobook for speech synthesiser
abstract
Audiobooks have been focused on as promising data for training Text-to-Speech (TTS) systems. However, they usually do not have a correspondence between audio and text data. Moreover, they are usually divided only into chapter units. In practice, we have to make a correspondence of audio and text data before we use them for building TTS synthesisers. However aligning audio and text data is time-consuming and involves manual labor. It also requires persons skilled in speech processing. Previously, we have proposed to use graphemes for automatically aligning speech and text data. This paper further integrates a lightly supervised voice activity detection (VAD) technique to detect sentence boundaries as a pre-processing step before the grapheme approach. This lightly supervised technique requires time stamps of speech and silence only for the first fifty sentences. Combining those, we can semi-automatically build TTS systems from audiobooks with minimum manual intervention. From subjective evaluations we analyse how the grapheme-based aligner and/or the proposed VAD technique impact the quality of HMM-based speech synthesisers trained on audiobooks.
Yoshitaka Mamiya, Junichi Yamagishi, Oliver Watts, Robert A. J. Clark, Simon King 0001, Adriana Cornelia Stan
ICASSP6
2013 Lightly supervised discriminative training of grapheme models for improved sentence-level alignment of speech and text data
abstract
This paper introduces a method for lightly supervised discriminative training using MMI to improve the alignment of speech and text data for use in training HMM-based TTS systems for low-resource languages. In TTS applications, due to the use of long-span contexts, it is important to select training utterances which have wholly correct transcriptions. In a low-resource setting, when using poorly trained grapheme models, we show that the use of MMI discriminative training at the grapheme-level enables us to increase the amount of correctly aligned data by 40 while maintaining a 7% sentence error rate and 0.8% word error rate. We present the procedure for lightly supervised discriminative training with regard to the objective of minimising sentence error rate.
Adriana Cornelia Stan, Peter Bell 0001, Junichi Yamagishi, Simon King 0001
INTERSPEECH1
2013 TUNDRA: a multilingual corpus of found data for TTS research created with light supervision
abstract
Simple4All Tundra (version 1.0) is the first release of a standardised multilingual corpus designed for text-to-speech re-search with imperfect or found data. The corpus consists of approximately 60 hours of speech data from audiobooks in 14 languages, as well as utterance-level alignments obtained with a lightly-supervised process. Future versions of the corpus will include finer-grained alignment and prosodic annotation, all of which will be made freely available. This paper gives a gen-eral outline of the data collected so far, as well as a detailed description of how this has been done, emphasizing the mini-mal language-specific knowledge and manual intervention used to compile the corpus. To demonstrate its potential use, text-to-speech systems have been built for all languages using unsu-pervised or lightly supervised methods, also briefly presented in the paper. Index Terms: multilingual corpus, light supervision, imperfect data, found data, text-to-speech, audiobook data
Adriana Cornelia Stan, Oliver Watts, Yoshitaka Mamiya, Mircea Giurgiu, Robert A. J. Clark, Junichi Yamagishi, Simon King 0001
INTERSPEECH1
2012 A grapheme-based method for automatic alignment of speech and text data
abstract
This paper introduces a method for automatic alignment of speech data with unsynchronised, imperfect transcripts, for a domain where no initial acoustic models are available. Using grapheme-based acoustic models, word skip networks and orthographic speech transcripts, we are able to harvest 55% of the speech with a 93% utterance-level accuracy and 99% word accuracy for the produced transcriptions. The work is based on the assumption that there is a high degree of correspondence between the speech and text, and that a full transcription of all of the speech is not required. The method is language independent and the only prior knowledge and resources required are the speech and text transcripts, and a few minor user interventions.
Adriana Cornelia Stan, Peter Bell 0001, Simon King 0001
SLT1
2011 The Romanian speech synthesis (RSS) corpus: Building a high quality HMM-based speech synthesis system using a high sampling rate
Adriana Cornelia Stan, Junichi Yamagishi, Simon King 0001, Matthew P. Aylett
Speech Commun.1