VLDB 2026 Research / reviewers in the wild / expert
Sofoklis Kakouros
dblp:63/7823
· DBLP profile ↗
22ranked-venue papers
17as first author
5since 2021 · last 2025
0000-0001-8996-0793ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 13 first-author · 5 since 2021Artificial intelligence and machine learning · 15 · 12 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating the Impact of Word Informativeness on Speech Emotion RecognitionabstractIn emotion recognition from speech, a key challenge lies in identifying speech signal segments that carry the most relevant acoustic variations for discerning specific emotions. Traditional approaches compute functionals for features such as energy and F0 over entire sentences or longer speech portions, potentially missing essential fine-grained variation in the long-form statistics. This research investigates the use of word informativeness, derived from a pre-trained language model, to identify semantically important segments. Acoustic features are then computed exclusively for these identified segments, enhancing emotion recognition accuracy. The methodology utilizes standard acoustic prosodic features, their functionals, and self-supervised representations. Results indicate a notable improvement in recognition performance when features are computed on segments selected based on word informativeness, underscoring the effectiveness of this approach. Sofoklis Kakouros |
INTERSPEECH | 1 |
| 2025 | Sounding Like a Winner? Prosodic Differences in Post-Match Interviews
Sofoklis Kakouros |
INTERSPEECH | 1 |
| 2023 | Speech-Based Emotion Recognition with Self-Supervised Models Using Attentive Channel-Wise Correlations and Label SmoothingabstractWhen recognizing emotions from speech, we encounter two common problems: how to optimally capture emotion-relevant information from the speech signal and how to best quantify or categorize the noisy subjective emotion labels. Self-supervised pre-trained representations can robustly capture information from speech enabling state-of-the-art results in many downstream tasks including emotion recognition. However, better ways of aggregating the information across time need to be considered as the relevant emotion information is likely to appear piecewise and not uniformly across the signal. For the labels, we need to take into account that there is a substantial degree of noise that comes from the subjective human annotations. In this paper, we propose a novel approach to attentive pooling based on correlations between the representations’ coefficients combined with label smoothing, a method aiming to reduce the confidence of the classifier on the training labels. We evaluate our proposed approach on the benchmark dataset IEMOCAP, and demonstrate high performance surpassing that in the literature. The code to reproduce the results is available at github.com/skakouros/s3prl_attentive_correlation. Sofoklis Kakouros, Themos Stafylakis, Ladislav Mosner, Lukás Burget |
ICASSP | 1 |
| 2023 | North Sámi Dialect Identification with Self-supervised Speech ModelsabstractThe North Sámi (NS) language encapsulates four primary dialectal variants that are related but that also have differences in their phonology, morphology, and vocabulary. The unique geopolitical location of NS speakers means that in many cases they are bilingual in Sámi as well as in the dominant state language: Norwegian, Swedish, or Finnish. This enables us to study the NS variants both with respect to the spoken state language and their acoustic characteristics. In this paper, we investigate an extensive set of acoustic features, including MFCCs and prosodic features, as well as state-of-the-art self-supervised representations, namely, XLS-R, WavLM, and HuBERT, for the automatic detection of the four NS variants. In addition, we examine how the majority state language is reflected in the dialects. Our results show that NS dialects are influenced by the state language and that the four dialects are separable, reaching high classification accuracy, especially with the XLS-R model. Sofoklis Kakouros, Katri Hiovain-Asikainen |
INTERSPEECH | 1 |
| 2022 | Extracting Speaker and Emotion Information from Self-Supervised Speech Models via Channel-Wise CorrelationsabstractSelf-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typically approached by using descriptive statistics, and in particular, using the first - and second-order statistics of representation coefficients. In this paper, we examine an alternative way of extracting speaker and emotion information from self-supervised trained models, based on the correlations between the coefficients of the representations - correlation pooling. We show improvements over mean pooling and further gains when the pooling methods are combined via fusion. The code is available at github.com/Lamomal/s3prl_correlation. Themos Stafylakis, Ladislav Mosner, Sofoklis Kakouros, Oldrich Plchot, Lukás Burget, Jan Cernocký |
SLT | 3 |
| 2019 | Prosodic Representations of Prominence Classification Neural Networks and Autoencoders Using Bottleneck FeaturesabstractProminence perception has been known to correlate with a complex interplay of the acoustic features of energy, fundamental frequency, spectral tilt, and duration. The contribution and importance of each of these features in distinguishing between prominent and non-prominent units in speech is not always easy to determine, and more so, the prosodic representations that humans and automatic classifiers learn have been difficult to interpret. This work focuses on examining the acoustic prosodic representations that binary prominence classification neural networks and autoencoders learn for prominence. We investigate the complex features learned at different layers of the network as well as the 10-dimensional bottleneck features (BNFs), for the standard acoustic prosodic correlates of prominence separately and in combination. We analyze and visualize the BNFs obtained from the prominence classification neural networks as well as their network activations. The experiments are conducted on a corpus of Dutch continuous speech with manually annotated prominence labels. Our results show that the prosodic representations obtained from the BNFs and higher-dimensional non-BNFs provide good separation of the two prominence categories, with, however, different partitioning of the BNF space for the distinct features, and the best overall separation obtained for F0. Sofoklis Kakouros, Antti Suni, Juraj Simko, Martti Vainio |
INTERSPEECH | 1 |
| 2018 | Comparison of spectral tilt measures for sentence prominence in speech - Effects of dimensionality and adverse noise conditions
Sofoklis Kakouros, Okko Johannes Räsänen, Paavo Alku |
Speech Commun. | 1 |
| 2017 | Connecting stimulus-driven attention to the properties of infant-directed speech - Is exaggerated intonation also more surprising?
Okko Johannes Räsänen, Sofoklis Kakouros, Melanie Soderstrom |
CogSci | 2 |
| 2017 | Evaluation of Spectral Tilt Measures for Sentence Prominence Under Different Noise ConditionsabstractSpectral tilt has been suggested to be a correlate of prominence in speech, although several studies have not replicated this empirically. This may be partially due to the lack of a standard method for tilt estimation from speech, rendering interpretations and comparisons between studies difficult. In addition, little is known about the performance of tilt estimators for prominence detection in the presence of noise. In this work, we investigate and compare several standard tilt measures on quantifying prominence in spoken Dutch and under different levels of additive noise. We also compare these measures with other acoustic correlates of prominence, namely, energy, F0, and duration. Our results provide further empirical support for the finding that tilt is a systematic correlate of prominence, at least in Dutch, even though energy, F0, and duration appear still to be more robust features for the task. In addition, our results show that there are notable differences between different tilt estimators in their ability to discriminate prominent words from non-prominent ones in different levels of noise. Sofoklis Kakouros, Okko Johannes Räsänen, Paavo Alku |
INTERSPEECH | 1 |
| 2016 | Statistical Learning of Prosodic Patterns and Reversal of Perceptual Cues for Sentence Prominence
Sofoklis Kakouros, Okko Johannes Räsänen |
CogSci | 1 |
| 2016 | A Cognitive Approach to Modeling Sentence Level Prominence Based on Stimulus Unpredictability
Sofoklis Kakouros, Okko Johannes Räsänen |
CogSci | 1 |
| 2016 | Analyzing the Contribution of Top-Down Lexical and Bottom-Up Acoustic Cues in the Detection of Sentence ProminenceabstractCopyright © 2016 ISCA. Recent work has suggested that prominence perception could be driven by the predictability of the acoustic prosodic features of speech. On the other hand, lexical predictability and part of speech information are also known to correlate with prominence. In this paper, we investigate how the bottom-up acoustic and top-down lexical cues contribute to sentence prominence by using both types of features in unsupervised and supervised systems for automatic prominence detection. The study is conducted using a corpus of Dutch continuous speech with manually annotated prominence labels. Our results show that unpredictability of speech patterns is a consistent and important cue for prominence at both the lexical and acoustic levels, and also that lexical predictability and part-of-speech information can be used as efficient features in supervised prominence classifiers. Sofoklis Kakouros, Joris Pelemans, Lyan Verwimp, Patrick Wambacq, Okko Johannes Räsänen |
INTERSPEECH | 1 |
| 2016 | Does the Importance of Word-Initial and Word-Final Information Differ in Native versus Non-Native Spoken-Word Recognition?abstractContains fulltext : 161987.pdf (Publisher’s version ) (Open Access) Odette Scharenborg, Juul Coumans, Sofoklis Kakouros, Roeland van Hout |
INTERSPEECH | 3 |
| 2016 | The Effect of Sentence Accent on Non-Native Speech Perception in Noiseabstract\n Contains fulltext :\n 162038.pdf (Publisher’s version ) (Open Access)\n Odette Scharenborg, Elea Kolkman, Sofoklis Kakouros, Brechtje Post |
INTERSPEECH | 3 |
| 2016 | 3PRO - An unsupervised method for the automatic detection of sentence prominence in speech
Sofoklis Kakouros, Okko Johannes Räsänen |
Speech Commun. | 1 |
| 2015 | Analyzing the Predictability of Lexeme-specific Prosodic Features as a Cue to Sentence Prominence
Sofoklis Kakouros, Okko Johannes Räsänen |
CogSci | 1 |
| 2015 | Automatic detection of sentence prominence in speech using predictability of word-level acoustic features
Sofoklis Kakouros, Okko Johannes Räsänen |
INTERSPEECH | 1 |
| 2014 | Statistical Unpredictability of F0 Trajectories as a Cue to Sentence Stress
Sofoklis Kakouros, Okko Johannes Räsänen |
CogSci | 1 |
| 2014 | Perception of sentence stress in English infant directed speechabstractVarious studies have examined the acoustic features in infant directed speech (IDS) and adult directed speech (ADS). However, there are few speech corpora with prominence annotation from multiple listeners or analysis of the acoustic properties of the stressed versus unstressed words, most studies and corpora focusing on syllabic stress. In order to fill this gap, the current study analyzes the acoustic properties of sentence stress in a corpus of English IDS. More specifically, the work is one of the first analyzing IDS as perceived by adult listeners, providing inter-annotator agreement ratings and an analysis of the acoustic correlates of sentence stress with regard to the most important prosodic features encountered in the literature: fundamental frequency, intensity, word duration, and spectral tilt. The analysis shows that all of the analyzed features correlate with the perception of stress, indicating that the sentential prominence in IDS is conveyed by similar acoustic characteristics that are known to be relevant for stress perception in ADS. Sofoklis Kakouros, Okko Johannes Räsänen |
INTERSPEECH | 1 |
| 2014 | Modeling Dependencies in Multiple Parallel Data Streams with Hyperdimensional ComputingabstractThis work presents an approach for modeling statistical dependencies in multivariate discrete sequences by using hyperdimensional random vectors. The system takes any number of parallel sequences as inputs and learns to predict the future states of these streams using the mutual dependencies between the inputs. Performance of the system is tested in an activity recognition task with data from multiple worn sensors. The results show that the approach outperforms the existing baseline results in the task and demonstrate that the system is capable to account for the varying reliability of different input streams. Okko Johannes Räsänen, Sofoklis Kakouros |
IEEE Signal Process. Lett. | 2 |
| 2013 | Attention based temporal filtering of sensory signals for data redundancy reductionabstractSince modern computational devices are required to store and process increasing amounts of data generated from various sources, efficient algorithms for identification of significant information in the data are becoming essential. Sensory recordings are one example where automatic and continuous storing and processing of large amounts of data is needed. Therefore, algorithms that can alleviate the computational load of the devices and reduce their storage requirements by removing uninformative data are important. In this work we propose a method for data reduction based on theories of human attention. The method detects temporally salient events based on the context in which they occur and retains only those sections of the input signal. The algorithm is tested as a pre-processing stage in a weakly supervised keyword learning experiment where it is shown to significantly improve the quality of the codebooks used in the pattern discovery process. Sofoklis Kakouros, Okko Johannes Räsänen, Unto K. Laine |
ICASSP | 1 |
| 2008 | Quantisation for Multiple Description Coding for voice over IPabstractThe transmission of voice over IP networks is heavily affected by packet losses. An increasingly popular method to increase the error resilience of these systems is the use of multiple description coding (MDC). However, the MDC techniques commonly used tend to add a significant amount of redundancies, which are not always easy to use optimally. In this paper, we propose a simple vector quantisation scheme to maximise MDC performance, and study several factors affecting its performance under various error conditions. The results show that it is possible to obtain good performance under packet loss conditions, while using only limited amounts of redundancy. Sofoklis Kakouros, Stephane Villette, Ahmet M. Kondoz |
ICASSP | 1 |