VLDB 2026 Research / reviewers in the wild / expert
Katerina Papadimitriou
dblp:172/9746
· DBLP profile ↗
9ranked-venue papers
8as first author
6since 2021 · last 2025
0009-0001-8647-4025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 6 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multi-Stream Framework Utilizing 3D Human Reconstruction for Cued Speech Recognition
Katerina Papadimitriou, Gerasimos Potamianos |
INTERSPEECH | 1 |
| 2024 | Multimodal Continuous Fingerspelling Recognition via Visual Alignment Learning
Katerina Papadimitriou, Gerasimos Potamianos |
INTERSPEECH | 1 |
| 2024 | A large corpus for the recognition of Greek Sign Language gestures
Katerina Papadimitriou, Galini Sapountzaki, Kyriaki Vasilaki, Eleni Efthimiou, Stavroula-Evita Fotinea, Gerasimos Potamianos |
Comput. Vis. Image Underst. | 1 |
| 2023 | Sign Language Recognition via Deformable 3D Convolutions and Modulated Graph Convolutional NetworksabstractAutomatic sign language recognition (SLR) remains challenging, especially when employing RGB video alone (i.e., with no depth or special glove-based input) and under a signer-independent (SI) framework, due to inter-personal signing variation. In this paper, we address SI isolated SLR from RGB video, proposing an innovative deep-learning framework that leverages multi-modal appearanceand skeleton-based information. Specifically, we propose three components for the first time in SLR: (i) a modified version of the ResNet2+1D network to capture signing appearance information, where spatial and temporal convolutions are substituted by their deformable counterparts, accomplishing both prevalent spatial modeling potential and motion-aware modeling adaptability; (ii) a novel spatio-temporal graph convolutional network (ST-GCN) that integrates a GCN variant, involving weight and affinity modulation for modeling diverse correlations between different body joints beyond the physical human skeleton structure, followed by a self-attention layer and a temporal convolution; and (iii) the “PIXIE” 3D human pose and shape regressor to generate 3D joint-rotation parameterization used for ST-GCN graph construction. Both appearance- and skeleton-based streams are ensembled in the proposed system and evaluated on two datasets of isolated signs, one in Turkish and one in Greek. Our system outperforms the state-of-the-art on the second set, yielding 53% relative error rate reduction (2.45% absolute), while it performs on par with the best reported system on the first. Katerina Papadimitriou, Gerasimos Potamianos |
ICASSP | 1 |
| 2023 | Multimodal Locally Enhanced Transformer for Continuous Sign Language Recognition
Katerina Papadimitriou, Gerasimos Potamianos |
INTERSPEECH | 1 |
| 2022 | Spatio-Temporal Graph Convolutional Networks for Continuous Sign Language RecognitionabstractWe address the challenging problem of continuous sign language recognition (CSLR) from RGB videos, proposing a novel deep-learning framework that employs spatio-temporal graph convolutional networks (ST-GCNs), which operate on multiple, appropriately fused feature streams, capturing the signer’s pose, shape, appearance, and motion information. In addition to introducing such networks to the continuous recognition problem, our model’s novelty lies on: (i) the feature streams considered and their blending into three ST-GCN modules; (ii) the combination of such modules with bi-directional long short-term memory networks, thus capturing both short-term embedded signing dynamics and long-range feature dependencies; and (iii) the fusion scheme, where the resulting modules operate in parallel, their posteriors aligned via a guiding connectionist temporal classification method, and fused for sign gloss prediction. Notably, concerning (i), in addition to traditional CSLR features, we investigate the utility of 3D human pose and shape parameterization via the "ExPose" approach, as well as 3D skeletal joint information that is regressed from detected 2D joints. We evaluate the proposed system on two well-known CSLR benchmarks, conducting extensive ablations on its modules. We achieve the new state-of-the-art on one of the two datasets, while reaching very competitive performance on the other. Maria Parelli, Katerina Papadimitriou, Gerasimos Potamianos, Georgios Pavlakos, Petros Maragos |
ICASSP | 2 |
| 2020 | Multimodal Sign Language Recognition via Temporal Deformable Convolutional Sequence LearningabstractIn this paper we address the challenging problem of sign language recognition (SLR) from videos, introducing an end-to-end deep learning approach that relies on the fusion of a number of spatio-temporal feature streams, as well as a fully convolutional encoder-decoder for prediction. Specifically, we examine the contribution of optical flow, human skeletal features, as well as appearance features of handshapes and mouthing, in conjunction with a temporal deformable convolutional attention-based encoder-decoder for SLR. To our knowledge, this is the first use in this task of a fully convolutional multi-step attention-based encoder-decoder employing temporal deformable convolutional block structures. We conduct experiments on three sign language datasets and compare our approach to existing state-of-the-art SLR methods, demonstrating its superiority. © 2020 ISCA Katerina Papadimitriou, Gerasimos Potamianos |
INTERSPEECH | 1 |
| 2019 | End-to-End Convolutional Sequence Learning for ASL Fingerspelling RecognitionabstractAlthough fingerspelling is an often overlooked component of sign languages, it has great practical value in the communication of important context words that lack dedicated signs. In this paper we consider the problem of fingerspelling recognition in videos, introducing an end-to-end lexicon-free model that consists of a deep auto-encoder image feature learner followed by an attention-based encoder-decoder for prediction. The feature extractor is a vanilla auto-encoder variant, employing a quadratic activation function. The learned features are subsequently fed into the attention-based encoder-decoder. The latter deviates from traditional recurrent neural network architectures, being a fully convolutional attention-based encoder-decoder that is equipped with a multi-step attention mechanism relying on a quadratic alignment function and gated linear units over the convolution output. The introduced model is evaluated on the TTIC/UChicago fingerspelling video dataset, where it outperforms previous approaches in letter accuracy under all three, signer-dependent, -adapted, and -independent, experimental paradigms. Copyright © 2019 ISCA Katerina Papadimitriou, Gerasimos Potamianos |
INTERSPEECH | 1 |
| 2015 | Tomographic image reconstruction withaspatially varying Gaussian mixture priorabstractA spatially varying Gaussian mixture model (SVGMM) prior is employed to ensure the preservation of region boundaries in penalized likelihood tomographic image reconstruction. Spatially varying Gaussian mixture models are characterized by the dependence of their mixing proportions on location (contextual mixing proportions) and they have been successfully used in image segmentation. The proposed model imposes a Student's t-distribution on the local differences of the contextual mixing proportions and its parameters are automatically estimated by a variational Expectation-Maximization (EM) algorithm. The tomographic reconstruction algorithm is an iterative process consisting of alternating between an optimization of the SVGMM parameters and an optimization for updating the unknown image using also the EM algorithm. Numerical experiments on various photon limited image scenarios show that the proposed model is more accurate than the widely used Gibbs prior. Katerina Papadimitriou, Christophoros Nikou |
ICIP | 1 |