VLDB 2026 Research / reviewers in the wild / expert
Efthymios Georgiou
dblp:263/4759
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-6042-9584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Medusa: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions
Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou, Efthymios Georgiou, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos |
INTERSPEECH | 5 |
| 2024 | $\mathcal {P}$owMix: A Versatile Regularizer for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) leverages heterogeneous data sources to interpret the complex nature of human sentiments. Despite significant progress in multimodal architecture design, the field lacks comprehensive regularization methods. This paper introduces$\mathcal {P}$owMix, a versatile embedding space regularizer that builds upon the strengths of unimodal mixing-based regularization approaches and introduces novel algorithmic components that are specifically tailored to multimodal tasks.$\mathcal {P}$owMix is integrated before the fusion stage of multimodal architectures and facilitates intra-modal mixing, such as mixing text with text, to act as a regularizer.$\mathcal {P}$owMix consists of five components: 1) a varying number of generated mixed examples, 2) mixing factor reweighting, 3) anisotropic mixing, 4) dynamic mixing, and 5) cross-modal label mixing. Extensive experimentation across benchmark MSA datasets and a broad spectrum of diverse architectural designs demonstrate the efficacy of$\mathcal {P}$owMix, as evidenced by consistent performance improvements over baselines and existing mixing methods. An in-depth ablation study highlights the critical contribution of each$\mathcal {P}$owMix component and how they synergistically enhance performance. Furthermore, algorithmic analysis demonstrates how$\mathcal {P}$owMix behaves in different scenarios, particularly comparing early versus late fusion architectures. Notably,$\mathcal {P}$owMix enhances overall performance without sacrificing model robustness or magnifying text dominance. It also retains its strong performance in situations of limited data. Our findings position$\mathcal {P}$owMix as a promising versatile regularization strategy for MSA. Efthymios Georgiou, Yannis Avrithis, Alexandros Potamianos |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Mmlatch: Bottom-Up Top-Down Fusion For Multimodal Sentiment AnalysisabstractCurrent deep learning approaches for multimodal fusion rely on bottom-up fusion of high and mid-level latent modality representations (late/mid fusion) or low level sensory inputs (early fusion). Models of human perception highlight the importance of top-down fusion, where high-level representations affect the way sensory inputs are perceived, i.e. cognition affects perception. These top-down interactions are not captured in current deep learning models. In this work we propose a neural architecture that captures top-down cross-modal interactions, using a feedback mechanism in the forward pass during network training. The proposed mechanism extracts high-level representations for each modality and uses these representations to mask the sensory inputs, allowing the model to perform top-down feature masking. We apply the proposed model for multimodal sentiment recognition on CMU-MOSEI. Our method shows consistent improvements over the well established MulT and over our strong late fusion baseline, achieving state-of-the-art results. Georgios Paraskevopoulos, Efthymios Georgiou, Alexandros Potamianos |
ICASSP | 2 |
| 2022 | A Multi-Task BERT Model for Schema-Guided Dialogue State TrackingabstractTask-oriented dialogue systems often employ a Dialogue State Tracker (DST) to successfully complete conversations.Recent state-of-the-art DST implementations rely on schemata of diverse services to improve model robustness and handle zeroshot generalization to new domains [1], however such methods [2, 3] typically require multiple large scale transformer models and long input sequences to perform well.We propose a single multi-task BERT-based model that jointly solves the three DST tasks of intent prediction, requested slot prediction and slot filling.Moreover, we propose an efficient and parsimonious encoding of the dialogue history and service schemata that is shown to further improve performance.Evaluation on the SGD dataset shows that our approach outperforms the baseline SGP-DST by a large margin and performs well compared to the state-of-the-art, while being significantly more computationally efficient.Extensive ablation studies are performed to examine the contributing factors to the success of our model. Eleftherios Kapelonis, Efthymios Georgiou, Alexandros Potamianos |
INTERSPEECH | 2 |
| 2022 | Regotron: Regularizing the Tacotron2 Architecture Via Monotonic Alignment LossabstractDeep learning Text-to-Speech (TTS) systems have achieved impressive generated speech quality, close to human parity. However, they suffer from training stability issues and in-correct alignment between the intermediate acoustic representation and the text input. In this work, we propose Regotron, a regularized Tacotron2 version which alleviates the training issues by augmenting the objective function with an additional term, which penalizes non-monotonic alignments in the location-sensitive attention mechanism. By introducing this regularization term we demonstrate its effectiveness to stabilize the training process, produce a monotonic attention quicker (13% of the total number of epochs compared to Tacotron2) and reduce the alignment errors during inference. Moreover, Regotron has minimal additional computational overhead, reduces common TTS mistakes and at the same time achieves improved speech naturalness according to subjective mean opinion scores (MOS) collected from 50 evaluators. Efthymios Georgiou, Kosmas Kritsis, Georgios Paraskevopoulos, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos |
SLT | 1 |
| 2021 | M3: MultiModal Masking Applied to Sentiment Analysis
Efthymios Georgiou, Georgios Paraskevopoulos, Alexandros Potamianos |
Interspeech | 1 |
| 2019 | Deep Hierarchical Fusion with Application in Sentiment Analysis
Efthymios Georgiou, Charilaos Papaioannou, Alexandros Potamianos |
INTERSPEECH | 1 |