EDBT 2026 Demo / reviewers in the wild / expert
Horia Cucu
dblp:11/11428
· DBLP profile ↗
21ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-8711-3854ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Easy, Interpretable, Effective: openSMILE for voice deepfake detectionabstractIn this paper, we demonstrate that attacks in the latest ASVspoof5 dataset—a de facto standard in the field of voice authenticity and deepfake detection—can be identified with surprising accuracy using a small subset of very simplistic features. These are derived from the openSMILE library, and are scalar-valued, easy to compute, and human interpretable. For example, attack A10’s unvoiced segments have a mean length of 0.09 ± 0.02, while bona fide instances have a mean length of 0.18 ± 0.07. Using this feature alone, a threshold classifier achieves an Equal Error Rate (EER) of 10.3% for attack A10. Similarly, across all attacks, we achieve up to 0.8% EER, with an overall EER of 15.7 ± 6.0%.We explore the generalization capabilities of these features and find that some of them transfer effectively between attacks, primarily when the attacks originate from similar Text-to-Speech (TTS) architectures. This finding may indicate that voice anti-spoofing is, in part, a problem of identifying and remembering signatures or fingerprints of individual TTS systems. This allows to better understand anti-spoofing models and their challenges in real-world application. Octavian Pascu, Dan Oneata, Horia Cucu, Nicolas M. Müller |
ICASSP | 3 |
| 2025 | Unmasking real-world audio deepfakes: A data-centric approachabstract5343 David Combei, Adriana Cornelia Stan, Dan Oneata, Nicolas M. Müller, Horia Cucu |
INTERSPEECH | 5 |
| 2025 | TADA: Training-free Attribution and Out-of-Domain Detection of Audio DeepfakesabstractDeepfake detection has gained significant attention across audio, text, and image modalities, with high accuracy in distinguishing real from fake. However, identifying the exact source--such as the system or model behind a deepfake--remains a less studied problem. In this paper, we take a significant step forward in audio deepfake model attribution or source tracing by proposing a training-free, green AI approach based entirely on k-Nearest Neighbors (kNN). Leveraging a pre-trained self-supervised learning (SSL) model, we show that grouping samples from the same generator is straightforward--we obtain an 0.93 F1-score across five deepfake datasets. The method also demonstrates strong out-of-domain (OOD) detection, effectively identifying samples from unseen models at an F1-score of 0.84. We further analyse these results in a multi-dimensional approach and provide additional insights. All code and data protocols used in this work are available in our open repository: https://github.com/adrianastan/tada/. Adriana Cornelia Stan, David Combei, Dan Oneata, Horia Cucu |
INTERSPEECH | 4 |
| 2025 | Evolutionary Bayesian Optimization for automated circuit sizingabstractAutomated circuit sizing using Artificial Intelligence is a rapidly increasing area of interest, primarily thanks to its potential to accelerate product time-to-market and enhance employee satisfaction. A host of methods, rooted in different fundamental research philosophies, have been devised for this class of problems. While some of them perform well in terms of convergence speed , robustness has generally been given less attention. In this study we propose a novel automatic circuit sizing framework called Evolutionary Bayesian Optimization (EBO). It is a hybrid method combining the strengths of evolutionary computation techniques and Bayesian Optimization. EBO takes full advantage of parallel simulation infrastructure, by inherently using large batches of simulations. Our method is especially designed for multi-objective problems. Thus, it can optimize a large variety of circuits without the need of constructing figure of merit functions. Moreover, the strong emphasis on exploring the high-dimensional space of design variables ensures that EBO is robust and reliable across varying levels of problem complexity. We compare our framework with two state-of-the-art methods having different underlying philosophies and with arguably the most promising multi-objective evolutionary algorithm for this class of problems on four circuits: two proprietary voltage regulators, an open-source voltage regulator, and an open-source operational amplifier . The results show that EBO is superior to the other considered methods with regard to convergence speed and robustness. Generally, it can save between 30% and 70% circuit simulations compared to the next best performing method. Furthermore, EBO is the only method that finds circuit configurations that meet the specifications for all the considered circuits. Catalin Visan, Mihai Boldeanu, Georgian Nicolae, Horia Cucu, Corneliu Burileanu, Andi Buzo |
Knowl. Based Syst. | 4 |
| 2024 | Towards generalisable and calibrated audio deepfake detection with self-supervised representationsabstractGeneralisation—the ability of a model to perform well on unseen data—is crucial for building reliable deepfake detectors. However, recent studies have shown that the current audio deepfake models fall short of this desideratum. In this work we investigate the potential of pretrained self-supervised representations in building general and calibrated audio deepfake detection models. We show that large frozen representations coupled with a simple logistic regression classifier are extremely effective in achieving strong generalisation capabilities: compared to the RawNet2 model, this approach reduces the equal error rate from 30.9% to 8.8% on a benchmark of eight deepfake datasets, while learning less than 2k parameters. Moreover, the proposed method produces considerably more reliable predictions compared to previous approaches making it more suitable for realistic use. Octavian Pascu, Adriana Cornelia Stan, Dan Oneata, Elisabeta Oneata, Horia Cucu |
INTERSPEECH | 5 |
| 2024 | Hybrid-Diarization System with Overlap Post-Processing for the DISPLACE 2024 Challenge
Gabriel Pirlogeanu, Octavian Pascu, Alexandru-Lucian Georgescu, Horia Cucu |
INTERSPEECH | 4 |
| 2023 | Adaptation of Whisper models to child speech recognition
Andrei Barcovschi, Mariam Yahayah Yiwere, Peter Corcoran 0001, Horia Cucu |
INTERSPEECH | 5 |
| 2023 | The SpeeD-ZevoTech submission at DISPLACE 2023
Gabriel Pirlogeanu, Dan Oneata, Alexandru-Lucian Georgescu, Horia Cucu |
INTERSPEECH | 4 |
| 2023 | Efficient Multi-Objective Optimization for PVT Variation-Aware Circuit Sizing Using Surrogate Models and Smart Corner SamplingabstractCircuit sizing for designs with many design variables and responses is a complex task that requires highly experienced and creative designers to invest precious time in trial and error, routine work. In addition, sizing the circuit while also taking into account PVT (process, voltage, temperature) variation corners increases the complexity further. To simplify such tasks, designers select the most unfavorable PVT corner in advance (leveraging their expertise), perform circuit sizing for this condition, and finally verify the resulting design in all PVT corners. This procedure might generate designs that fail the specifications in other PVT corners leading to more design-verification iterative loops. Recent years brought machine learning (ML) and optimization techniques to the field of circuit design, with evolutionary algorithms and Bayesian models showing good results for automated circuit sizing. However, these methods can still require an unfeasibly large number of simulations, especially if taking into account several PVT corners. In this context, we introduce a methodology that uses surrogate ML models to perform PVT variation-aware circuit sizing. We propose to dynamically select the worst PVT corners and take them into account when sizing the circuit. In addition, we explore the best ways to model process corners with Gaussian Processes, leading to more than 10x improvements for such surrogate models. We evaluate the proposed corner management method on two voltage regulators showing different levels of complexity and highlight that it enables finding feasible solutions 2x faster when compared to baseline algorithms which optimize in all PVT corners. In addition, the quality and diversity of the proposed solutions are significantly higher by one to three orders of magnitude in terms of population hypervolume. Octavian Pascu, Catalin Visan, Georgian Nicolae, Mihai Boldeanu, Horia Cucu, Cristian Diaconu, Andi Buzo, Georg Pelz |
ISLPED | 5 |
| 2022 | Automated circuit sizing with multi-objective optimization based on differential evolution and Bayesian inferenceabstractManual sizing of analog circuit specifications has become challenging owing to their ever-increasing complexity. Especially for innovative, large-scale circuit designs with numerous design variables, operating conditions, and conflicting objectives to optimize, analog designers must run time-consuming simulations for several weeks to find the optimum configuration. Recently, machine learning and optimization techniques have been applied in the field of analog circuit design, wherein evolutionary algorithms and Bayesian models have shown good results for circuit sizing tasks. In this context, we introduce multi-objective optimization based on differential evolution and Bayesian inference (MODEBI)—a design optimization method based on generalized differential evolution 3 (GDE3) and Gaussian processes (GPs). The proposed method can perform sizing for complex circuits that require optimization of many design variables and conflicting objectives. Although state-of-the-art methods reduce multi-objective problems to single-objective optimization and potentially induce a priori bias, the proposed method searches directly over the multi-objective space using Pareto dominance and ensures that designers are provided with diverse solutions to choose from. To reduce optimization time, we propose using GPs to model the circuit and employing this surrogate model to preselect candidates. However, this results in a more complex offspring selection process, and the diversity in population survival must be specifically addressed. This paper proposes several solutions to these problems, resulting in multiple MODEBI variations. To the best of our knowledge, this is the first method that specifically addresses solution diversity and simultaneously focuses on minimizing the number of simulations required to obtain feasible configurations. The evaluation performed on two voltage regulators with different complexity levels showed that the proposed offspring selection method and survival policy can obtain highly diverse feasible solutions considerably faster than GDE3 or Bayesian optimization-based algorithms. Catalin Visan, Octavian Pascu, Marius Stanescu, Elena-Diana Sandru, Cristian Diaconu, Andi Buzo, Georg Pelz, Horia Cucu |
Knowl. Based Syst. | 8 |
| 2021 | Data-Filtering Methods for Self-Training of Automatic Speech Recognition SystemsabstractSelf-training is a simple and efficient way of leveraging un-labeled speech data: (i) start with a seed system trained on transcribed speech; (ii) pass the unlabeled data through this seed system to automatically generate transcriptions; (iii) en-large the initial dataset with the self-labeled data and retrain the speech recognition system. However, in order not to pol-lute the augmented dataset with incorrect transcriptions, an important intermediary step is to select those parts of the self-labeled data that are accurate. Several approaches have been proposed in the community, but most of the works address only a single method. In contrast, in this paper we inspect three distinct classes of data-filtering for self-training, leveraging: (i) confidence scores, (ii) multiple ASR hypotheses and (iii) approximate transcriptions. We evaluate these approaches from two perspectives: quantity vs. quality of the selected data and improvement of the seed ASR by including this data. The proposed methodology achieves state-of-the-art results on Romanian speech, obtaining 25% relative improvement over prior work. Among the three methods, approximate transcriptions bring the highest performance gain, even if they yield the least quantity of data. Alexandru-Lucian Georgescu, Cristian Manolache, Dan Oneata, Horia Cucu, Corneliu Burileanu |
SLT | 4 |
| 2021 | An Evaluation of Word-Level Confidence Estimation for End-to-End Automatic Speech RecognitionabstractQuantifying the confidence (or conversely the uncertainty) of a prediction is a highly desirable trait of an automatic system, as it improves the robustness and usefulness in downstream tasks. In this paper we investigate confidence estimation for end-to-end automatic speech recognition (ASR). Previous work has addressed confidence measures for lattice-based ASR, while current machine learning research mostly focuses on confidence measures for unstructured deep learning. However, as the ASR systems are increasingly being built upon deep end-to-end methods, there is little work that tries to develop confidence measures in this context. We fill this gap by providing an extensive benchmark of popular confidence methods on four well-known speech datasets. There are two challenges we overcome in adapting existing methods: working on structured data (sequences) and obtaining confidences at a coarser level than the predictions (words instead of tokens). Our results suggest that a strong baseline can be obtained by scaling the logits by a learnt temperature, followed by estimating the confidence as the negative entropy of the predictive distribution and, finally, sum pooling to aggregate at word level. Dan Oneata, Alexandru Caranica, Adriana Cornelia Stan, Horia Cucu |
SLT | 4 |
| 2020 | RSC: A Romanian Read Speech Corpus for Automatic Speech RecognitionabstractAlthough many efforts have been made in the last decade to enhance the speech and language resources for Romanian, this language is still considered under-resourced. While for many other languages there are large speech corpora available for research and commercial applications, for Romanian language the largest publicly available corpus to date comprises less than 50 hours of speech. In this context, Speech and Dialogue research group releases Read Speech Corpus (RSC) – a Romanian speech corpus developed in-house, comprising 100 hours of speech recordings from 164 different speakers. The paper describes the development of the corpus and presents baseline automatic speech recognition (ASR) results using state-of-the-art ASR technology: Kaldi speech recognition toolkit. Alexandru-Lucian Georgescu, Horia Cucu, Andi Buzo, Corneliu Burileanu |
LREC | 2 |
| 2019 | Kite: Automatic Speech Recognition for Unmanned Aerial VehiclesabstractThis paper addresses the problem of building a speech recognition system attuned to the control of unmanned aerial vehicles (UAVs). Even though UAVs are becoming widespread, the task of creating voice interfaces for them is largely unaddressed. To this end, we introduce a multi-modal evaluation dataset for UAV control, consisting of spoken commands and associated images, which represent the visual context of what the UAV sees when the pilot utters the command. We provide baseline results and address two research directions: (i) how robust the language models are, given an incomplete list of commands at train time; (ii) how to incorporate visual information in the language model. We find that recurrent neural networks (RNNs) are a solution to both tasks: they can be successfully adapted using a small number of commands and they can be extended to use visual cues. Our results show that the image-based RNN outperforms its text-only counterpart even if the command-image training associations are automatically generated and inherently imperfect. The dataset and our code are available at this http URL. Dan Oneata, Horia Cucu |
INTERSPEECH | 2 |
| 2018 | Methodology for determining the influencing factors of lifetime variation for power devicesabstractThis paper proposes a method for explanation of the lifetime variation of power devices using data from different test stages. Understanding the lifetime variation is very useful in qualification, as well as in the characterization process, in order to improve the robustness of the power devices or to estimate more accurately the minimum guaranteed lifetime. Moreover, it helps design engineers better understand the root causes of the lifetime variation and use this knowledge to improve the performances of new power devices. In the proposed methodology, the variation of the lifetime is explained by the electrical parameters, measured before the stress-test. The Sensitivity Analysis presented here has the advantage of being simple and fast. It can be applied even when the number of test-runs is less than the number of factors. Moreover, it reveals not only linear correlations, but also quadratic effects and 2nd and 3rd order interactions. Eventually, the method provides the top of the most relevant electrical parameters which explain the lifetime variation. The validation of this approach has shown that 72% of the lifetime variation can be explained by the initial values of 5 electrical parameters. Ciprian V. Pop, Andi Buzo, Georg Pelz, Horia Cucu, Corneliu Burileanu |
ETS | 4 |
| 2018 | Multilingual Low-Resourced Prototype System for Voice-Controlled Intelligent Building Applications
Alexandru Caranica, Alexandru-Lucian Georgescu, Alexandru Vulpe, Horia Cucu |
WorldCIST (3) | 4 |
| 2017 | Detecting Overlapped Speech on Short Timeframes Using Deep Learning
Valentin Andrei, Horia Cucu, Corneliu Burileanu |
INTERSPEECH | 2 |
| 2015 | Counting competing speakers in a timeframe - human versus computer
Valentin Andrei, Horia Cucu, Andi Buzo, Corneliu Burileanu |
INTERSPEECH | 2 |
| 2014 | Detecting the number of competing speakers - human selective hearing versus spectrogram distance based estimator
Valentin Andrei, Horia Cucu, Andi Buzo, Corneliu Burileanu |
INTERSPEECH | 2 |
| 2014 | SMT-based ASR domain adaptation methods for under-resourced languages: Application to Romanian
Horia Cucu, Andi Buzo, Laurent Besacier, Corneliu Burileanu |
Speech Commun. | 1 |
| 2011 | Investigating the role of machine translated text in ASR domain adaptation: Unsupervised and semi-supervised methodsabstractThis study investigates the use of machine translated text for ASR domain adaptation. The proposed methodology is applicable when domain-specific data is available in language X only, whereas the goal is to develop a domain-specific system in language Y. Two semi-supervised methods are introduced and compared with a fully unsupervised approach, which represents the baseline. While both unsupervised and semi-supervised approaches allow to quickly develop an accurate domain-specific ASR system, the semi-supervised approaches overpass the unsupervised one by 10% to 29% relative, depending on the amount of human post-processed data available. An in-depth analysis, to explain how the machine translated text improves the performance of the domain-specific ASR, is also given at the end of this paper. Horia Cucu, Laurent Besacier, Corneliu Burileanu, Andi Buzo |
ASRU | 1 |