EDBT 2026 Demo / reviewers in the wild / expert
Antigoni Tsiami
dblp:148/9872
· DBLP profile ↗
13ranked-venue papers
6as first author
1since 2021 · last 2023
0000-0001-9075-6343ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 3 first-authorSystems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Image and video processing · 61% Multimedia analysis and retrieval · 30% Audio and music processing · 9% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 25% Wearable and physiological sensing · 25% Haptics and multimodal interaction · 25% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Multimedia analysis and retrieval › audio-visual learning
audio-visual saliency |
0.4 | 1 | 2020 | STAViS: Spatio-Temporal AudioVisual Saliency Network · CVPR 2020 |
Image and video processing
saliency detection |
0.4 | 1 | 2020 | STAViS: Spatio-Temporal AudioVisual Saliency Network · CVPR 2020 |
Image and video processing › saliency detection
video saliency |
0.4 | 1 | 2020 | STAViS: Spatio-Temporal AudioVisual Saliency Network · CVPR 2020 |
Haptics and multimodal interaction › multimodal fusion
audiovisual fusion |
0.3 | 1 | 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots · ICRA 2018 |
Human-robot interaction
child-robot interaction |
0.3 | 1 | 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots · ICRA 2018 |
Interaction techniques and input › input sensing
gesture recognition |
0.3 | 1 | 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots · ICRA 2018 |
Wearable and physiological sensing
sensor fusion |
0.3 | 1 | 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots · ICRA 2018 |
Methods — techniques the papers use, named apart from their topics
multimodal fusion · 0.4deep neural network · 0.4speech recognition · 0.3sensor fusion · 0.3kinect · 0.3action recognition · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Convolutional Recurrent Neural Networks for the Classification of Cetacean Bioacoustic PatternsabstractIn this paper we focus on the development of a convolutional recurrent neural network (CRNN) to categorize biosignals collected in the Hellenic Trench, generated by two cetacean species, sperm whales (Physeter macrocephalus) and striped dolphins (Stenella coeruleoalba). We convert audio signals into mel-spectrograms and forward the input into a deep residual network (ResNet), designed to capture spectral patterns. Next, ResNet’s output is reshaped into a time-distributed layer and fed into recurrent network variants, Long Short-Term Memory (LSTMs) or Gated Recurrent Units (GRUs), able to recognize long-term time dependencies on extracted features. The hybrid network perfectly classifies audio signals into three categories (dolphins, sperm whales, ambient noise) while it also exhibits high learning ability on recognising intraclass representations of overlapping acoustic patterns (clicks vs whistles and clicks, both emitted by dolphins). The proposed scheme outperforms traditional Machine Learning (ML) techniques, baseline ResNet and LSTM architectures or their deep parallel combinations. Dimitris N. Makropoulos, Antigoni Tsiami, Aristides Prospathopoulos, Dimitris Kassis, Alexandros Frantzis, Emmanuel K. Skarsoulis, George Piperakis, Petros Maragos |
ICASSP | 2 |
| 2020 | STAViS: Spatio-Temporal AudioVisual Saliency NetworkabstractWe introduce STAViS, a spatio-temporal audiovisual saliency network that combines spatio-temporal visual and auditory information in order to efficiently address the problem of saliency estimation in videos. Our approach employs a single network that combines visual saliency and auditory features and learns to appropriately localize sound sources and to fuse the two saliencies in order to obtain a final saliency map. The network has been designed, trained end-to-end, and evaluated on six different databases that contain audiovisual eye-tracking data of a large variety of videos. We compare our method against 8 different state-of-the-art visual saliency models. Evaluation results across databases indicate that our STAViS model outperforms our visual only variant as well as the other state-of-the-art models in the majority of cases. Also, the consistently good performance it achieves for all databases indicates that it is appropriate for estimating saliency "in-the-wild". The code is available at https://github.com/atsiami/STAViS. Antigoni Tsiami, Petros Koutras, Petros Maragos |
CVPR | 1 |
| 2019 | Video Processing and Learning in Assistive Robotic ApplicationsabstractThe integration of visual perception to robotic systems is a key research area in recent years. Advances in modern computer vision techniques along with the development of faster and more accurate visual sensors led to the emergence of new methods for robotic visual perception [1]. One area of research is the development of robotic assistive vision for Human-Robot Interaction (HRI) systems [2], [3]. The rapid increase of people with special needs, such as the elderly population, and the simultaneous reduction of personal care staff, reinforce the need for robotic assistants [4], [5]. There are many challenges in this area including the familiarity of these users with new technologies and the domain specific datasets, which are required for training user oriented models. Nowadays, modern assistive and social human-robot interaction requires the multimodal communication with speech, gestures and human movements so as to enhance the classic interaction with only spoken commands. Petros Koutras, Georgia Chalvatzaki, Antigoni Tsiami, Alexandros Nikolakakis, Costas S. Tzafestas, Petros Maragos |
ICIP | 3 |
| 2019 | A behaviorally inspired fusion approach for computational audiovisual saliency modeling
Antigoni Tsiami, Petros Koutras, Athanasios Katsamanis, Argiro Vatakis, Petros Maragos |
Signal Process. Image Commun. | 1 |
| 2018 | Far-Field Audio-Visual Scene Perception of Multi-Party Human-Robot Interaction for Children and AdultsabstractHuman-robot interaction (HRI) is a research area of growing interest with a multitude of applications for both children and adult user groups, as, for example, in edutainment and social robotics. Crucial, however, to its wider adoption remains the robust perception of HRI scenes in natural, untethered, and multi-party interaction scenarios, across user groups. Towards this goal, we investigate three focal HRI perception modules operating on data from multiple audio-visual sensors that observe the HRI scene from the far-field, thus bypassing limitations and platform-dependency of contemporary robotic sensing. In particular, the developed modules fuse intra- and/or inter-modality data streams to perform: (i) audio-visual speaker localization; (ii) distant speech recognition; and (iii) visual recognition of hand-gestures. Emphasis is also placed on ensuring high speech and gesture recognition rates for both children and adults. Development and objective evaluation of the three modules is conducted on a corpus of both user groups, collected by our far-field multisensory setup, for an interaction scenario of a question-answering “guess-the-object” collaborative HRI game with a “Furhat” robot. In addition, evaluation of the game incorporating the three developed modules is reported. Our results demonstrate robust far-field audio-visual perception of the multi-party HRI scene. Antigoni Tsiami, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Petros Koutras, Gerasimos Potamianos, Petros Maragos |
ICASSP | 1 |
| 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple RobotsabstractChild-robot interaction is an interdisciplinary research area that has been attracting growing interest, primarily focusing on edutainment applications. A crucial factor to the successful deployment and wide adoption of such applications remains the robust perception of the child's multi-modal actions, when interacting with the robot in a natural and untethered fashion. Since robotic sensory and perception capabilities are platform-dependent and most often rather limited, we propose a multiple Kinect-based system to perceive the child-robot interaction scene that is robot-independent and suitable for indoors interaction scenarios. The audio-visual input from the Kinect sensors is fed into speech, gesture, and action recognition modules, appropriately developed in this paper to address the challenging nature of child-robot interaction. For this purpose, data from multiple children are collected and used for module training or adaptation. Further, information from the multiple sensors is fused to enhance module performance. The perception system is integrated in a modular multi-robot architecture demonstrating its flexibility and scalability with different robotic platforms. The whole system, called Multi3, is evaluated, both objectively at the module level and subjectively in its entirety, under appropriate child-robot interaction scenarios containing several carefully designed games between children and robots. Antigoni Tsiami, Petros Koutras, Niki Efthymiou, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos |
ICRA | 1 |
| 2017 | Room-localized spoken command recognition in multi-room, multi-microphone environments
Isidoros Rodomagoulakis, Athanasios Katsamanis, Gerasimos Potamianos, Panagiotis Giannoulis, Antigoni Tsiami, Petros Maragos |
Comput. Speech Lang. | 5 |
| 2016 | Multimodal human action recognition in assistive human-robot interactionabstractWithin the context of assistive robotics we develop an intelligent interface that provides multimodal sensory processing capabilities for human action recognition. Human action is considered in multimodal terms, containing inputs such as audio from microphone arrays, and visual inputs from high definition and depth cameras. Exploring state-of-the-art approaches from automatic speech recognition, and visual action recognition, we multimodally recognize actions and commands. By fusing the unimodal information streams, we obtain the optimum multimodal hypothesis which is to be further exploited by the active mobility assistance robot in the framework of the MOBOT EU research project. Evidence from recognition experiments shows that by integrating multiple sensors and modalities, we increase multimodal recognition performance in the newly acquired challenging dataset involving elderly people while interacting with the assistive robot. Isidoros Rodomagoulakis, Nikolaos Kardaris, Vassilis Pitsikalis, E. Mavroudi, Athanasios Katsamanis, Antigoni Tsiami, Petros Maragos |
ICASSP | 6 |
| 2016 | Towards a behaviorally-validated computational audiovisual saliency modelabstractComputational saliency models aim at predicting, in a bottom-up fashion, where human attention is drawn in the presented (visual, auditory or audiovisual) scene and have been proven useful in applications like robotic navigation, image compression and movie summarization. Despite the fact that well-established auditory and visual saliency models have been validated in behavioral experiments, e.g., by means of eye-tracking, there is no established computational audiovisual saliency model validated in the same way. In this work, building on biologically-inspired models of visual and auditory saliency, we present a joint audiovisual saliency model and introduce the validation approach we follow to show that it is compatible with recent findings of psychology and neuroscience regarding multimodal integration and attention. In this direction, we initially focus on the "pip and pop" effect which has been observed in behavioral experiments and indicates that visual search in sequences of cluttered images can be significantly aided by properly timed non-spatial auditory signals presented alongside the target visual stimuli. Antigoni Tsiami, Athanasios Katsamanis, Petros Maragos, Argiro Vatakis |
ICASSP | 1 |
| 2016 | A Phase-Based Time-Frequency Masking for Multi-Channel Speech Enhancement in Domestic Environments
Alessio Brutti, Antigoni Tsiami, Athanasios Katsamanis, Petros Maragos |
INTERSPEECH | 2 |
| 2015 | Multichannel speech enhancement using MEMS microphonesabstractIn this work, we investigate the efficacy of Micro Electro-Mechanical System (MEMS) microphones, a newly developed technology of very compact sensors, for multichannel speech enhancement. Experiments are conducted on real speech data collected using a MEMS microphone array. First, the effectiveness of the array geometry for noise suppression is explored, using a new corpus containing speech recorded in diffuse and localized noise fields with a MEMS microphone array configured in linear and hexagonal array geometries. Our results indicate superior performance of the hexagonal geometry. Then, MEMS microphones are compared to Electret Condenser Microphones (ECMs), using the ATHENA database, which contains speech recorded in realistic smart home noise conditions with hexagonal-type arrays of both microphone types. MEMS microphones exhibit performance similar to ECMs. Good performance, versatility in placement, small size, and low cost, make MEMS microphones attractive for multichannel speech processing. Z.-I. Skordilis, Antigoni Tsiami, Petros Maragos, Gerasimos Potamianos, Luca Spelgatti, Roberto Sannino |
ICASSP | 2 |
| 2014 | Robust far-field spoken command recognition for home automation combining adaptation and multichannel processingabstractThe paper presents our approach to speech-controlled home automation. We are focusing on the detection and recognition of spoken commands preceded by a key-phrase as recorded in a voice-enabled apartment by a set of multiple microphones installed in the rooms. For both problems we investigate robust modeling, environmental adaptation and multichannel processing to cope with a) insufficient training data and b) the far-field effects and noise in the apartment. The proposed integrated scheme is evaluated in a challenging and highly realistic corpus of simulated audio recordings and achieves F-measure close to 0.70 for key-phrase spotting and word accuracy close to 98% for the command recognition task. Athanasios Katsamanis, Isidoros Rodomagoulakis, Gerasimos Potamianos, Petros Maragos, Antigoni Tsiami |
ICASSP | 5 |
| 2014 | ATHENA: a Greek multi-sensory database for home automation control uthor: isidoros rodomagoulakis (NTUA, Greece)
Antigoni Tsiami, Isidoros Rodomagoulakis, Panagiotis Giannoulis, Athanasios Katsamanis, Gerasimos Potamianos, Petros Maragos |
INTERSPEECH | 1 |