EDBT 2026 Demo / reviewers in the wild / expert
Cigdem Beyan
dblp:126/0778
· DBLP profile ↗
36ranked-venue papers
15as first author
21since 2021 · last 2027
0000-0002-9583-0087ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Parameter-efficient vision-language adaptation with continuous metadata conditioning for animal re-identificationabstractLong-term animal re-identification (ReID) must remain robust to gradual morphological evolution and seasonal appearance shifts. Although recent vision–language models provide strong pretrained visual representations, adapting them to longitudinal ecological settings remains challenging, particularly under identity and temporal distribution shifts. We present a parameter-efficient CLIP adaptation framework for animal ReID and introduce a continuous metadata-conditioning mechanism that incorporates numerical attributes directly into the prompt representation during training. While low-rank visual adaptation, prompt-based supervision, and cross-modal alignment provide the adaptation framework, the proposed metadata-conditioning strategy constitutes the primary methodological contribution. By preserving the continuous structure of numerical metadata rather than discretizing it into textual categories, the proposed approach enables smooth modulation of the embedding space during training while maintaining a purely visual inference pipeline. Experiments on a seven-year longitudinal fish dataset and multiple wildlife benchmarks demonstrate improved performance under closed-set, open-set, and time-aware evaluation protocols. The results demonstrate that continuous metadata conditioning improves robustness to longitudinal appearance variation and temporal distribution shifts, while parameter-efficient adaptation enables a purely visual inference pipeline without requiring metadata at test time. Code and evaluation splits will be made publicly available upon acceptance. Anil Osman Tur, Tonje Knutsen Sørdalen, Kim Halvorsen, Cigdem Beyan |
Expert Syst. Appl. | 4 |
| 2026 | Geometry-Conditioned Diffusion for Occlusion-Robust In-Bed Pose Estimation
Navid Aslankhani Khameneh, Marco Carletti, Cigdem Beyan |
FG | 3 |
| 2026 | Towards Unconstrained Human-Object Interaction
Francesco Tonini, Alessandro Conti, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci 0001 |
FG | 4 |
| 2026 | HAC: Parameter-Efficient Hyperbolic Adaptation of CLIP for Zero-Shot VQA
Francesco Dibitonto, Cigdem Beyan, Vittorio Murino |
ICPR (3) | 2 |
| 2026 | Discriminator-Guided Adaptive Diffusion for Source-Free Test-Time Adaptation Under Image Corruptions
Francesco Olivato, Cigdem Beyan, Vittorio Murino |
ICPR (8) | 2 |
| 2025 | MadCLIP: Few-Shot Medical Anomaly Detection with CLIP
Mahshid Shiri, Cigdem Beyan, Vittorio Murino |
MICCAI (6) | 2 |
| 2025 | Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection aims to identify humans and objects within images and interpret their interactions. Existing HOI methods rely heavily on large datasets with manual annotations to learn interactions from visual cues. These annotations are labor-intensive to create, prone to inconsistency, and limit scalability to new domains and rare interactions. We argue that recent advances in Vision-Language Models (VLMs) offer untapped potential, particularly in enhancing interaction representation. While prior work has injected such potential and even proposed training-free methods, there remain key gaps. Consequently, we propose a novel training-free HOI detection framework for Dynamic Scoring with enhanced semantics (dysco) that effectively utilizes textual and visual interaction representations within a multimodal registry, enabling robust and nuanced interaction understanding. This registry incorporates a small set of visual cues and uses innovative interaction signatures to improve the semantic alignment of verbs, facilitating effective generalization to rare interactions. Additionally, we propose a unique multi-head attention mechanism that adaptively weights the contributions of the visual and textual features. Experimental results demonstrate that our dysco surpasses training-free state-of-the-art models and is competitive with training-based approaches, particularly excelling in rare interactions. Code is available at https://github.com/francescotonini/dysco. Francesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan, Elisa Ricci 0001 |
ACM Multimedia | 4 |
| 2024 | Diffusion-Based Unsupervised Pre-training for Automated Recognition of Vitality FormsabstractSocial communication involves interpreting nonverbal behaviors, detecting and anticipating others’ actions and intentions. Actions convey not only the goal and motor intention but also the form, i.e., variations in action execution. These variations, termed vitality forms, communicate attitudes during interactions, such as being gentle, calm, vigorous, and rude. Automatic vitality form recognition may have several applications in social robotics, social skills training, and therapy, yet it remains a rarely studied topic. This paper introduces an unsupervised pre-training approach that utilizes 2D-body key point trajectories as input and employs diffusion models to derive more effective features for representing these trajectories. The features learned from the diffusion model’s encoder are utilized to train a multilayer perceptron for vitality form recognition. Experimental analysis showcases the superior performance of the proposed method not only across various videos but also for action classes not encountered during training. Noemi Canovi, Federico Montagna, Radoslaw Niewiadomski, Alessandra Sciutti, Giuseppe Di Cesare, Cigdem Beyan |
AVI | 6 |
| 2024 | AL-GTD: Deep Active Learning for Gaze Target DetectionabstractGaze target detection aims at determining the image location where a person is looking. While existing studies have made significant progress in this area by regressing accurate gaze heatmaps, these achievements have largely relied on access to extensive labeled datasets, which demands substantial human labor. In this paper, our goal is to reduce the reliance on the size of labeled training data for gaze target detection. To achieve this, we propose AL-GTD, an innovative approach that integrates supervised and self-supervised losses within a novel sample acquisition function to perform active learning (AL). Additionally, it utilizes pseudo-labeling to mitigate distribution shifts during the training phase. AL-GTD achieves the best of all AUC results by utilizing only 40-50% of the training data, in contrast to state-of-the-art (SOTA) gaze target detectors requiring the entire training dataset to achieve the same performance. Importantly, AL-GTD quickly reaches satisfactory performance with 10-20% of the training data, showing the effectiveness of our acquisition function, which is able to acquire the most informative samples. We provide a comprehensive experimental analysis by adapting several AL methods for the task. AL-GTD outperforms AL competitors, simultaneously exhibiting superior performance compared to SOTA gaze target detectors when all are trained within a low-data regime. Code is available at: https://github.com/francescotonini/al-gtd. Francesco Tonini, Nicola Dall'Asen, Lorenzo Vaquero, Cigdem Beyan, Elisa Ricci 0001 |
ACM Multimedia | 4 |
| 2024 | Leveraging Next-Active Objects for Context-Aware Anticipation in Egocentric VideosabstractObjects are crucial for understanding human-object interactions. By identifying the relevant objects, one can also predict potential future interactions or actions that may occur with these objects. In this paper, we study the problem of Short-Term Object interaction anticipation (STA) and propose NAOGAT (Next-Active-Object Guided Anticipation Transformer), a multi-modal end-to-end transformer network, that attends to objects in observed frames in order to anticipate the next-active-object (NAO) and, eventually, to guide the model to predict context-aware future actions. The task is challenging since it requires anticipating future action along with the object with which the action occurs and the time after which the interaction will begin, a.k.a. the time to contact (TTC). Compared to existing video modeling architectures for action anticipation, NAOGAT captures the relationship between objects and the global scene context in order to predict detections for the next active object and anticipate relevant future actions given these detections, leveraging the objects’ dynamics to improve accuracy. One of the key strengths of our approach, in fact, is its ability to exploit the motion dynamics of objects within a given clip , which is often ignored by other models, and separately decoding the object-centric and motion-centric information. Through our experiments, we show that our model outperforms existing methods on two separate datasets, Ego4D and EpicKitchens-100 ("Unseen Set"), as measured by several additional metrics, such as time to contact, and next-active-object localization. The code can be found on project page : sanketsans.github.io/wacv24 Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue |
WACV | 2 |
| 2023 | Object-aware Gaze Target DetectionabstractGaze target detection aims to predict the image location where the person is looking and the probability that a gaze is out of the scene. Several works have tackled this task by regressing a gaze heatmap centered on the gaze location, however, they overlooked decoding the relationship between the people and the gazed objects. This paper proposes a Transformer-based architecture that automatically detects objects (including heads) in the scene to build associations between every head and the gazed-head/object, resulting in a comprehensive, explainable gaze analysis composed of: gaze target area, gaze pixel point, the class and the image location of the gazed-object. Upon evaluation of the in-the-wild benchmarks, our method achieves state-of-the-art results on all metrics (up to 2.91% gain in AUC, 50% reduction in gaze distance, and 9% gain in out-of-frame average precision) for gaze target detection and 11-13% improvement in average precision for the classification and the localization of the gazed-objects. The code of the proposed method is publicly available1. Francesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa Ricci 0001 |
ICCV | 3 |
| 2023 | Enhancing Next Active Object-Based Egocentric Action Anticipation with Guided AttentionabstractShort-term action anticipation (STA) in first-person videos is a challenging task that involves understanding the next active object interactions and predicting future actions. Existing action anticipation methods have primarily focused on utilizing features extracted from video clips, but often overlooked the importance of objects and their interactions. To this end, we propose a novel approach that applies a guided attention mechanism between the objects, and the spatiotemporal features extracted from video clips, enhancing the motion and contextual information, and further decoding the object-centric and motion-centric information to address the problem of STA in egocentric videos. Our method, GANO (Guided Attention for Next active Objects) is a multi-modal, end-to-end, single transformer-based network. The experimental results performed on the largest egocentric dataset demonstrate that GANO outperforms the existing state-of-the-art methods for the prediction of the next active object label, its bounding box location, the corresponding future action, and the time to contact the object. The ablation study shows the positive contribution of the guided attention mechanism compared to other fusion methods. Moreover, it is possible to improve the next active object location and class label prediction results of GANO by just appending the learnable object tokens with the region of interest embeddings. Related implementations are available at: sanketsans.github.io/guided-attention-egocentric.html Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue |
ICIP | 2 |
| 2023 | Exploring Diffusion Models for Unsupervised Video Anomaly DetectionabstractThis paper investigates the performance of diffusion models for video anomaly detection (VAD) within the most challenging but also the most operational scenario in which the data annotations are not used. As being sparse, diverse, contextual, and often ambiguous, detecting abnormal events precisely is a very ambitious task. To this end, we rely only on the information-rich spatio-temporal data, and the reconstruction power of the diffusion models such that a high reconstruction error is utilized to decide the abnormality. Experiments performed on two large-scale video anomaly detection datasets demonstrate the consistent improvement of the proposed method over the state-of-the-art generative models while in some cases our method achieves better scores than the more complex models. This is the first study using a diffusion model and examining its parameters’ influence to present guidance for VAD in surveillance scenarios. Anil Osman Tur, Nicola Dall'Asen, Cigdem Beyan, Elisa Ricci 0001 |
ICIP | 3 |
| 2023 | Modeling Multiple Temporal Scales of Full-Body Movements for Emotion ClassificationabstractThis work investigates classification of emotions from full-body movements by using a novel Convolutional Neural Network-based architecture. The model is composed of two shallow networks processing in parallel where the 8-bit RGB images obtained from time intervals of 3D-positional data are the inputs. One network performs a coarse-grained modelling in the time domain while the other one applies a fine-grained modelling. We show that combining different temporal scales into one architecture improves the classification results of a dataset composed of short excerpts of the performances of professional dancers who interpreted four affective states: anger, happiness, sadness, and insecurity. Additionally, we investigate the effect of data chunk duration, overlapping, the size of the input images and the contribution of several data augmentation strategies for our proposed method. Better recognition results were obtained when the duration of a data chunk was longer, and this was further improved by applying balanced data augmentation. Moreover, we test our method on other existing motion capture datasets and compare the results with prior art. In all of the experiments, our results surpassed the state-of-the-art approaches, showing that this method generalizes across diverse settings and contexts. Cigdem Beyan, Sukumar Karumuri, Gualtiero Volpe, Antonio Camurri, Radoslaw Niewiadomski |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Multimodal Across Domains Gaze Target DetectionabstractThis paper addresses the gaze target detection problem in single images captured from the third-person perspective. We present a multimodal deep architecture to infer where a person in a scene is looking. This spatial model is trained on the head images of the person-of-interest, scene and depth maps representing rich context information. Our model, unlike several prior art, do not require supervision of the gaze angles, do not rely on head orientation information and/or location of the eyes of person-of-interest. Extensive experiments demonstrate the stronger performance of our method on multiple benchmark datasets. We also investigated several variations of our method by altering joint-learning of multimodal data. Some variations outperform a few prior art as well. First time in this paper, we inspect domain adaptation for gaze target detection, and we empower our multimodal network to effectively handle the domain gap across datasets. The code of the proposed method is available at https://github.com/francescotonini/multimodal-across-domains-gaze-target-detection. Francesco Tonini, Cigdem Beyan, Elisa Ricci 0001 |
ICMI | 2 |
| 2022 | Multimodal Emotion Recognition with Modality-Pairwise Unsupervised Contrastive LossabstractEmotion recognition is involved in several real-world applications. With an increase in available modalities, automatic understanding of emotions is being performed more accurately. The success in Multimodal Emotion Recognition (MER), primarily relies on the supervised learning paradigm. However, data annotation is expensive, time-consuming, and as emotion expression and perception depends on several factors (e.g., age, gender, culture) obtaining labels with a high reliability is hard. Motivated by these, we focus on unsupervised feature learning for MER. We consider discrete emotions, and as modalities text, audio and vision are used. Our method, as being based on contrastive loss between pairwise modalities, is the first attempt in MER literature. Our end-to-end feature learning approach has several differences (and advantages) compared to existing MER methods: i) it is unsupervised, so the learning is lack of data labelling cost; ii) it does not require data spatial augmentation, modality alignment, large number of batch size or epochs; iii) it applies data fusion only at inference; and iv) it does not require backbones pre-trained on emotion recognition task. The experiments on benchmark datasets show that our method outperforms several baseline approaches and unsupervised learning methods applied in MER. Particularly, it even surpasses a few supervised MER state-of-the-art. Riccardo Franceschini, Enrico Fini, Cigdem Beyan, Alessandro Conti, Federica Arrigoni, Elisa Ricci 0001 |
ICPR | 3 |
| 2021 | Unsupervised Human Action Recognition with Skeletal Graph Laplacian and Self-Supervised Viewpoints Invariance
Giancarlo Paoletti, Jacopo Cavazza, Cigdem Beyan, Alessio Del Bue |
BMVC | 3 |
| 2021 | Predicting Gaze from Egocentric Social Interaction Videos and IMU DataabstractGaze prediction in egocentric videos is a fairly new research topic, which might have several applications for assistive technology (e.g., supporting blind people in their daily interactions), security (e.g., attention tracking in risky work environments), education (e.g., augmented / mixed reality training simulators, immersive games) and so forth. Egocentric gaze is typically estimated from video while few works attempt to use inertial measurement unit (IMU) data, a sensor modality often available in wearable devices (e.g., augmented reality headsets). Instead, in this paper, we examine whether joint learning of egocentric video and corresponding IMU data can improve the first-person gaze prediction compared to using these modalities separately. In this respect, we propose a multimodal network and evaluate it on several unconstrained social interaction scenarios captured by a first-person perspective. The proposed multimodal network achieves better results compared to unimodal methods as well as several (multimodal) baselines, showing that using egocentric video together with IMU data can boost the first-person gaze estimation performance. Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Alessio Del Bue |
ICMI | 2 |
| 2021 | S-VVAD: Visual Voice Activity Detection by Motion SegmentationabstractWe address the challenging Voice Activity Detection (VAD) problem, which determines "Who is Speaking and When?" in audiovisual recordings. The typical audio-based VAD systems can be ineffective in the presence of ambient noise or noise variations. Moreover, due to technical or privacy reasons, audio might not be always available. In such cases, the use of video modality to perform VAD is desirable. Almost all existing visual VAD methods rely on body part detection, e.g., face, lips, or hands. In contrast, we propose a novel visual VAD method operating directly on the entire video frame, without the explicit need of detecting a person or his/her body parts. Our method, named S-VVAD, learns body motion cues associated with speech activity within a weakly supervised segmentation framework. Therefore, it not only detects the speakers/not-speakers but simultaneously localizes the image positions of them. It is an end-to-end pipeline, person-independent and it does not require any prior knowledge nor pre-processing. S-VVAD performs well in various challenging conditions and demonstrates the state-of-the-art results on multiple datasets. Moreover, the better generalization capability of S-VVAD is confirmed for cross-dataset and person-independent scenarios. Muhammad Shahid 0002, Cigdem Beyan, Vittorio Murino |
WACV | 2 |
| 2021 | Personality Traits Classification Using Deep Visual Activity-Based Nonverbal Features of Key-Dynamic ImagesabstractThis paper addresses nonverbal behavior analysis for the classification of perceived personality traits using novel deep visual activity (VA)-based features extracted only from key-dynamic images. Dynamic images represent short-term VA. Key-dynamic images carry more discriminative information i.e., nonverbal features (NFs) extracted from them contribute to the classification more than NFs extracted from other dynamic images. Dynamic image construction, learning long-term VA with CNN+LSTM, and detecting spatio-temporal saliency are applied to determine key-dynamic images. Once VA-based NFs are extracted, they are encoded using covariance, and resulting representation is used for classification. This method was evaluated on two datasets: small group meetings and vlogs. For the first dataset, proposed method outperforms not only the state-of-the-art VA-based methods but also multi-modal approaches for all personality traits. For extraversion classification, it performs better than i) the most popular key-frames selection algorithm, ii) random and uniform dynamic image selection, and iii) NFs extracted from all dynamic images. Furthermore, the ablation study proves the superiority of proposed method. For the further dataset, it performs as well as the state-of-the-art visual-NFs on average, while showing improved performance for agreeableness classification. Proposed method can be adapted to any application based on nonverbal behavior analysis, thanks to being data-driven. Cigdem Beyan, Andrea Zunino, Muhammad Shahid 0002, Vittorio Murino |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | RealVAD: A Real-World Dataset and A Method for Voice Activity Detection by Body Motion AnalysisabstractWe present an automatic voice activity detection (VAD) method that is solely based on visual cues. Unlike traditional approaches processing audio, we show that upper body motion analysis is desirable for the VAD task. The proposed method consists of components for body motion representation, feature extraction from a Convolutional Neural Network (CNN) architecture and unsupervised domain adaptation. The body motion representations as images are used by the feature extraction component, which is generic and person-invariant, thus, can be applied to a subject who has never been seen. The endmost component handles the domain-shift problem, which appears due to the fact that the way people move/ gesticulate while speaking might vary from subject to subject, which results in disparate body motion features and consequently poorer VAD performance. The experimental analyses applied on a publicly available real-world VAD dataset show that the proposed method performs better than the state-of-the-art video-only and multimodal VAD approaches. Moreover, the proposed method has a better generalization ability as VAD results are more consistent across different subjects. As another major contribution, we present a new multimodal dataset (called RealVAD), created from a real-world (no role-plays) panel discussion. This dataset contains many actual situations/ challenges that are missing in the previous VAD datasets. We benchmarked the RealVAD dataset by applying the proposed method as well as cross-dataset analyses. Particularly, the results of cross-dataset experiments highlight the remarkable positive contribution of the unsupervised domain adaptation applied. Cigdem Beyan, Muhammad Shahid 0002, Vittorio Murino |
IEEE Trans. Multim. | 1 |
| 2020 | Analysis of Face-Touching Behavior in Large Scale Social Interaction DatasetabstractWe present the first publicly available annotations for the analysis of face-touching behavior. These annotations are for a dataset composed of audio-visual recordings of small group social interactions with a total number of 64 videos, each one lasting between 12 to 30 minutes and showing a single person while participating to four-people meetings. They were performed by in total 16 annotators with an almost perfect agreement (Cohen's Kappa=0.89) on average. In total, 74K and 2M video frames were labelled as face-touch and no-face-touch, respectively. Given the dataset and the collected annotations, we also present an extensive evaluation of several methods: rule-based, supervised learning with hand-crafted features and feature learning and inference with a Convolutional Neural Network (CNN) for Face-Touching detection. Our evaluation indicates that among all, CNN performed the best, reaching 83.76% F1-score and 0.84 Matthews Correlation Coefficient. To foster future research in this problem, code and dataset were made publicly available (github.com/IIT-PAVIS/Face-Touching-Behavior), providing all video frames, face-touch annotations, body pose estimations including face and hands key-points detection, face bounding boxes as well as the baseline methods implemented and the cross-validation splits used for training and evaluating our models. Cigdem Beyan, Matteo Bustreo, Muhammad Shahid 0002, Gian Luca Bailo, Nicolò Carissimi, Alessio Del Bue |
ICMI | 1 |
| 2020 | Subspace Clustering for Action Recognition with Covariance Representations and Temporal PruningabstractThis paper tackles the problem of human action recognition, defined as classifying which action is displayed in a trimmed sequence, from skeletal data. Albeit state-of-the-art approaches designed for this application are all supervised, in this paper we pursue a more challenging direction: solving the problem with unsupervised learning. To this end, we propose a novel subspace clustering method, which exploits covariance matrix to enhance the action's discriminability and a times-tamp pruning approach that allow us to better handle the temporal dimension of the data. Through a broad experimental validation, we show that our computational pipeline surpasses existing unsupervised approaches but also can result in favorable performances as compared to the supervised methods. The code is available here: https://github.com/IIT-PAVIS/subspace-clustering-action-recognition Giancarlo Paoletti, Jacopo Cavazza, Cigdem Beyan, Alessio Del Bue |
ICPR | 3 |
| 2019 | A Sequential Data Analysis Approach to Detect Emergent Leaders in Small GroupsabstractThis paper addresses the problem of predicting emergent leaders (ELs) in small groups, that is, meetings. This is a long-lasting research problem for social and organizational psychology and a relevant problem that recently gained momentum in social computing. Toward this goal, we propose a novel method, which analyzes the temporal dependencies of the audio-visual data by applying unsupervised deep learning generative models (feature learning). To the best of our knowledge, this is the first attempt that sequential data processing is performed for EL detection. Feature learning results in a single feature vector per a given time interval and all feature vectors representing a participant are aggregated using novel fusion techniques. Finally, the EL detection is performed using the state-of-the-art single and multiple kernel learning algorithms. The proposed method shows (significantly) improved results compared to the state-of-the-art methods and it can be adapted to analyze various small group interactions given that it is a general approach. Cigdem Beyan, Vasiliki-Maria Katsageorgiou, Vittorio Murino |
IEEE Trans. Multim. | 1 |
| 2018 | A Multi-View Learning Approach to Deception DetectionabstractRecently, automatic deception detection has gained momentum thanks to advances in computer vision, computational linguistics and machine learning research fields. The majority of the work in this area focused on written deception and analysis of verbal features. However, according to psychology, people display various nonverbal behavioral cues, in addition to verbal ones, while lying. Therefore, it is important to utilize additional modalities such as video and audio to detect deception accurately. When multi-modal data was used for deception detection, previous studies concatenated all verbal and nonverbal features into a single vector. This concatenation might not be meaningful, because different feature groups can have different statistical properties, leading to lower classification accuracy. Following this intuition, we apply, for the first time in deception detection, a multi-view learning (MVL) approach, where each view corresponds to a feature group. This results in improved classification results over the state of the art methods. Additionally, we show that the optimized parameters of the MVL algorithm can give insights into the contribution of each feature group to the final results, thus revealing the importance of each feature and eliminating the need of performing feature selection as well. Finally, we focus on analyzing face-based low level, not hand crafted features, which are extracted using various pre-trained Deep Neural Networks (DNNs), showing that face is the most important nonverbal cue for the detection of deception. Nicolò Carissimi, Cigdem Beyan, Vittorio Murino |
FG | 2 |
| 2018 | Investigation of Small Group Social Interactions Using Deep Visual Activity-Based Nonverbal FeaturesabstractUnderstanding small group face-to-face interactions is a prominent research problem for social psychology while the automatic realization of it recently became popular in social computing. This is mainly investigated in terms of nonverbal behaviors, as they are one of the main facet of communication. Among several multi-modal nonverbal cues, visual activity is an important one and its sufficiently good performance can be crucial for instance, when the audio sensors are missing. The existing visual activity-based nonverbal features, which are all hand-crafted, were able to perform well enough for some applications while did not perform well for some other problems. Given these observations, we claim that there is a need of more robust feature representations, which can be learned from data itself. To realize this, we propose a novel method, which is composed of optical flow computation, deep neural network based feature learning, feature encoding and classification. Additionally, a comprehensive analysis between different feature encoding techniques is also presented. The proposed method is tested on three research topics, which can be perceived during small group interactions i.e. meetings: i) emergent leader detection, ii) emergent leadership style prediction, and iii) high/low extraversion classification. The proposed method shows (significantly) better results not only as compared to the state of the art visual activity based-nonverbal features but also when the state of the art visual activity based-nonverbal features are combined with other audio-based and video-based nonverbal features. Cigdem Beyan, Muhammad Shahid 0002, Vittorio Murino |
ACM Multimedia | 1 |
| 2018 | Extracting statistically significant behaviour from fish tracking data with and without large dataset cleaningabstractExtracting a statistically significant result from video of natural phenomenon can be difficult for two reasons: (i) there can be considerable natural variation in the observed behaviour and (ii) computer vision algorithms applied to natural phenomena may not perform correctly on a significant number of samples. This study presents one approach to clean a large noisy visual tracking dataset to allow extracting statistically sound results from the image data. In particular, analyses of 3.6 million underwater trajectories of a fish with the water temperature at the time of acquisition are presented. Although there are many false detections and incorrect trajectory assignments, by a combination of data binning and robust estimation methods, reliable evidence for an increase in fish speed as water temperature increases are demonstrated. Then, a method for data cleaning which removes outliers arising from false detections and incorrect trajectory assignments using a deep learning‐based clustering algorithm is proposed. The corresponding results show a rise in fish speed as temperature goes up. Several statistical tests applied to both cleaned and not‐cleaned data confirm that both results are statistically significant and show an increasing trend. However, the latter approach also generates a cleaner dataset suitable for other analysis. Cigdem Beyan, Vasiliki-Maria Katsageorgiou, Robert B. Fisher |
IET Comput. Vis. | 1 |
| 2018 | Prediction of the Leadership Style of an Emergent Leader Using Audio and Visual Nonverbal FeaturesabstractThe coordination of a leader with group members is very important for an effective leadership given that this figure is the person who actually manages the team members to achieve a desired goal. Investigating the leadership and especially the leadership style is a prominent research topic in social and organizational psychology. However this is a new problem in social signal processing that can actually make valuable contributions by analyzing multimodal data in a more effective and efficient way. In this work we identify the leadership style of an emergent leader (i.e. the leader who naturally arises from a group not designated) as autocratic or democratic. The proposed method is applied to a dataset in-the-wild; in other words there is no role-playing which is novel for this problem. Multiple kernel learning (MKL) using multimodal nonverbal features is utilized to predict leadership styles that proved to achieve better predictions as compared to traditional learning methods. Thanks to MKL and a simple heuristic proposed the best performing features are also identified showing that better predictions can be reached only by using those features. Additionally correlation analysis between the extracted nonverbal features and the results of social psychology questionnaire is also performed. This shows that significantly high correlations exist for speaking activity based and prosodic nonverbal features. Cigdem Beyan, Francesca Capozzi, Cristina Becchio, Vittorio Murino |
IEEE Trans. Multim. | 1 |
| 2017 | Multi-task learning of social psychology assessments and nonverbal features for automatic leadership identificationabstractIn social psychology, the leadership investigation is performed using questionnaires which are either i) self-administered or ii) applied to group participants to evaluate other members or iii) filled by external observers. While each of these sources is informative, using them individually might not be as effective as using them jointly. This paper is the first attempt which addresses the automatic identification of leaders in small-group meetings, by learning effective models using nonverbal audio-visual features and the results of social psychology questionnaires that reflect assessments regarding leadership. Learning is based on Multi-Task Learning which is performed without using ground-truth data (GT), but using the results of questionnaires (having substantial agreement with GT), administered to external observers and the participants of the meetings, as tasks. The results show that joint learning results in better performance as compared to single task learning and other baselines. Cigdem Beyan, Francesca Capozzi, Cristina Becchio, Vittorio Murino |
ICMI | 1 |
| 2017 | Moving as a Leader: Detecting Emergent Leadership in Small Groups using Body PoseabstractDetecting leadership while understanding the underlying behavior is an important research topic particularly for social and organizational psychology, and has started to get attention from social signal processing research community as well. It is known that, visual activity is a useful cue to investigate the social interactions, even though previously applied nonverbal features based on head/body actions were not performing well enough for identification of emergent leaders (ELs) in small group meetings. Starting from these premises, in this study, we propose an effective method that uses 2D body pose based nonverbal features to represent the visual activity of a person. Our results suggest that, i) overall, the proposed nonverbal features derived from body pose perform better than existing visual activity based features, ii) it is possible to improve classification results by applying unsupervised feature learning as a preprocessing step, and iii) the proposed nonverbal features are able to advance the EL identification performances of other types of nonverbal features when they are used together. Cigdem Beyan, Vasiliki-Maria Katsageorgiou, Vittorio Murino |
ACM Multimedia | 1 |
| 2016 | Detecting emergent leader in a meeting environment using nonverbal visual features onlyabstractIn this paper, we propose an effective method for emergent leader detection in meeting environments which is based on nonverbal visual features. Identifying emergent leader is an important issue for organizations. It is also a well-investigated topic in social psychology while a relatively new problem in social signal processing (SSP). The effectiveness of nonverbal features have been shown by many previous SSP studies. In general, the nonverbal video-based features were not more effective compared to audio-based features although, their fusion generally improved the overall performance. However, in absence of audio sensors, the accurate detection of social interactions is still crucial. Motivating from that, we propose novel, automatically extracted, nonverbal features to identify the emergent leadership. The extracted nonverbal features were based on automatically estimated visual focus of attention which is based on head pose. The evaluation of the proposed method and the defined features were realized using a new dataset which is firstly introduced in this paper including its design, collection and annotation. The effectiveness of the features and the method were also compared with many state of the art features and methods. Cigdem Beyan, Nicolò Carissimi, Francesca Capozzi, Sebastiano Vascon, Matteo Bustreo, Antonio Pierro, Cristina Becchio, Vittorio Murino |
ICMI | 1 |
| 2015 | Classifying imbalanced data sets using similarity based hierarchical decomposition
Cigdem Beyan, Robert B. Fisher |
Pattern Recognit. | 1 |
| 2014 | A rule-based event detection system for real-life underwater domain
Concetto Spampinato, Emma Beauxis-Aussalet, Simone Palazzo, Cigdem Beyan, Jacco van Ossenbruggen, Jiyin He, Bas Boom |
Mach. Vis. Appl. | 4 |
| 2013 | Detection of Abnormal Fish Trajectories Using a Clustering Based Hierarchical ClassifierabstractWe address the analysis of fish trajectories in unconstrained underwater videos environmental changes which can be observed from the abnormal behaviour of fish. The fish trajectories are separated into normal and abnormal classes which indicate the common behaviour of fish and the behaviours that are rare/ unusual respectively. The proposed solution is based on a novel type of hierarchical classifier which builds the tree using clustered and labelled data based on similarity of data while using different feature sets at different levels of hierarchy. The paper presents a new method for fish trajectory analysis which has better performance compared to state-of-the-art techniques while the results are significant considering the challenges of underwater environments, low video quality, erratic movement of fish and highly imbalanced trajectory data that we used. Moreover, the proposed method is also powerful enough to classify highly imbalanced real-world datasets. Cigdem Beyan, Robert B. Fisher |
BMVC | 1 |
| 2013 | Detecting abnormal fish trajectories using clustered and labeled dataabstractWe propose an approach for the analysis of fish trajectories in unconstrained underwater videos. Trajectories are classified into two classes: normal trajectories which contain the usual behavior of fish and abnormal trajectories which indicate the behaviors that are not as common as the normal class. The paper presents two innovations: 1) a novel approach to abnormal trajectory detection and 2) improved performance on video based abnormal trajectory analysis of fish in unconstrained conditions. First we extract a set of features from trajectories and apply PCA. We then perform clustering on a subset of features. Based on the clustering, outlier detection is applied to each cluster. Improved results are obtained which is significant considering the challenges of underwater environments, low video quality, and erratic movement of fish. Cigdem Beyan, Robert B. Fisher |
ICIP | 1 |
| 2012 | A filtering mechanism for normal fish trajectories
Cigdem Beyan, Robert B. Fisher |
ICPR | 1 |