VLDB 2026 Research / reviewers in the wild / expert
Michal Muszynski
dblp:151/5516
· DBLP profile ↗
17ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0002-0763-2612ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 9 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorComputer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HRAI 2025: The 1st Workshop on Holistic and Responsible Affective IntelligenceabstractThe ICMI 2025 Workshop on Holistic and Responsible Affective Intelligence (HRAI 2025) aims to advance research in affective intelligence by fostering discussions on the holistic development of affective computing and the ethical challenges it entails. The workshop aims to strengthen interdisciplinary connections within the affective computing community, promoting better integration of methodologies and enhancing real-world applicability. By tackling both technical and ethical issues, HRAI 2025 aspires to shape the future of affective AI, ensuring it is not only powerful but also fair, safe, and socially responsible. Yuanchao Li, Dimitris Kollias, Guillaume Chanel, Marios A. Fanourakis, Michal Muszynski, Brandon M. Booth, Leimin Tian, Madhawa Perera, Catherine Lai, Huili Chen |
ICMI | 5 |
| 2023 | Toward Foundation Models for Earth Monitoring: Generalizable Deep Learning Models for Natural Hazard SegmentationabstractClimate change results in an increased probability of extreme weather events that put societies and businesses at risk on a global scale. Therefore, near real-time mapping of natural hazards is an emerging priority for the support of natural disaster relief, risk management, and informed governmental policy decisions. Current remote sensing based approaches to near real-time natural hazard mapping increasingly leverage advantage of deep learning (DL). Nevertheless, DL-based approaches are mainly designed for one specific task in a single geographic region based on specific frequency bands of satellite data. For that reason, DL models used to map specific natural hazards struggle with their generalization to other types of natural hazards in unseen regions. In this work, we propose a methodology to significantly improve the generalizability of DL natural hazards mappers based on pre-training on a suitable pre-task. Without access to any data from the target domain, we demonstrate that this methodology improved generalizability across four U-Net architectures for the segmentation of unseen natural hazards, such as flood events, landslides, and massive glacier collapses. Importantly, our method is strongly invariant to geographic differences and the type of input frequency bands of satellite data. That is confirmed by obtaining a balanced accuracy of up to 0.74 in comparison with performance of reference baselines. By leveraging characteristics of unlabeled images from the target domain that are publicly available, our approach is able to further improve the generalization behavior of DL models without fine-tuning. That is reflected in performance metrics. Thereby, our approach is one of first attempts to support the development of foundation models for earth monitoring with the objective of directly segmenting unseen natural hazards across novel geographic regions from different sources of satellite imagery. Johannes Jakubik, Michal Muszynski, Michael Vössing, Niklas Kühl 0001, Thomas Brunschwiler |
IGARSS | 2 |
| 2023 | Flood Mapping Using Sentinel-1 Images and Lightweight U-Nets Trained on Synthesized EventsabstractFloods cause loss of lives and multi-billion dollar damages every year. When these events strike, automated tools using remote sensing for quick mapping of the affected areas are critical for planning rescue activities and assessing impact. The most recent techniques to map floods are based on semantic segmentation and deep learning, which require large datasets and ground truth for training, that are difficult to get. To overcome this challenge, we propose an effective method for synthesizing patches of synthetic aperture radar images containing open-land flooded areas. With bi-temporal image acquisitions, we replace portions of land areas in the second acquisition with water pixels borrowed from permanent water bodies. Spatial patterns derived from elevation data help guide the process. With this approach, we build a large dataset with pre and post-event Sentinel-1-based VH-polarized radar intensity images to train deep neural networks to map real-life floods. In a case study, we employ an established U-Net architecture and show that a model version with less than 2% of the number of original parameters achieves almost identical flood detection accuracy. This allows for faster processing and is a clear advantage from an operational perspective. For comparison, we provide empirical segmentation results of four flood cases. F1 scores agree between 0.80 and 0.90 when compared to reference flood maps from the Copernicus service. This confirms that the approach is valuable for mapping areas affected by floods, e.g. during or immediately after catastrophic events, even in areas that were not included during the model training. Maciel Zortea, Michal Muszynski, Paolo Fraccaro |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Flood Event Detection from Sentinel 1 and Sentinel 2 Data: Does Land Use Matter for Performance of U-Net based Flood Segmenters?abstractFloods are among the most costly weather hazards for societies and businesses globally. With increasing global warming, these events have become even more frequent and more devastating. Thus, accurate flood mapping has become critical for disaster relief, risk management and mitigation. Current flood segmentation methods use either threshold-based approaches or deep-learning schemes, e.g. using the U-Net architecture, to differentiate between water-covered bodies or dry land on Earth observation images. Many schemes are exploiting imagery from synthetic aperture radar (e.g. Sentinel 1 satellites) or visual bands of satellites such as the Sentinel 2, but often restrict themselves to using one or very few modalities, i.e. spectral wavelengths, despite the availability of many more wavelengths or pre-processed indices with potential value to the challenge. In support of operationalizing flood segmentation on a global scale using deep learning, we propose semantic flood segmentation exploiting optionally many different modalities (i.e. multimodal flood segmentation), making the approach largely immune to geographic differences across the globe. Using U-Net at the core of our work, we observe very good generalisation of our segmentation model to unseen flood events in our holdout set at the level of 0.95 F1 Score (0.92 IoU) for both no water and water class, and 0.53 F1 Score (0.43 IoU) for water class, respectively. Michal Muszynski, Tobias Hölzer, Jonas R. M. Weiss, Paolo Fraccaro, Maciel Zortea, Thomas Brunschwiler |
IEEE Big Data | 1 |
| 2022 | Multimodal Affect and Aesthetic ExperienceabstractThe term “aesthetic experience” corresponds to the inner state of a person exposed to the form and content of artistic objects. Quantifying and interpreting the aesthetic experience of people in different contexts can contribute towards (a) creating context and (b) better understanding people’s affective reactions to different aesthetic stimuli. Focusing on different types of artistic content, such as movies, music, literature, urban art, ancient artwork, and modern interactive technology, the goal of this workshop is to enhance the interdisciplinary collaboration among researchers coming from the following domains: affective computing, aesthetics, human-robot/computer interaction, digital archaeology and art, culture, addictive games. Theodoros Kostoulas, Michal Muszynski, Leimin Tian, Edgar Roman-Rangel, Theodora Chaspari, Panos Amelidis |
ICMI | 2 |
| 2022 | Surface Water Mapping in Sentinel-1 Images: A Probabilistic Approach Combining Classic Detection MethodsabstractSurface water mapping in satellite images enables flood monitoring, a task with increasing importance under changing climate conditions. Current segmentation methods based on Deep Learning require large, curated datasets for training’ which are difficult to obtain. In this paper, we present a probabilistic approach to water segmentation based on established computer vision methods that requires little training and is easy to interpret. We use prior knowledge of typical backscatter intensity to locate seed pixels likely to be in water and land. A preliminary rough segmentation using thresholding guides the selection of two image patches that will be fully labeled using seeded region growing segmentation. Then, we sample small patches within the automatically labeled regions to train a fully connected neural network that, running in sliding windows, scores for the presence of water in the entire image. The approach is tested for mapping surface water during a large flood event in Aude, France. Outputs of the proposed approach are compared to the reference flood delineation map provided online by the Copernicus service. A F1 score of 0.67 suggests that performance of the proposed approach is similar or better than classic thresholding methods used as benchmark. Maciel Zortea, Paolo Fraccaro, Thomas Brunschwiler, Michal Muszynski, Jonas R. M. Weiss |
IGARSS | 4 |
| 2021 | Learning Language and Multimodal Privacy-Preserving Markers of Mood from Mobile DataabstractPaul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nick Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nicholas B. Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency |
ACL/IJCNLP (1) | 4 |
| 2021 | Workshop on Multimodal Affect and Aesthetic ExperienceabstractThe term “aesthetic experience” corresponds to inner states of individuals exposed to art. Investigating form, content, and aesthetic values of artistic objects, indoor and outdoor spaces, urban areas, and modern interactive technology is essential to improve social behaviour, quality of life, and health of humans in the long term. Quantifying and interpreting the aesthetic experience of art receivers in different contexts can contribute towards (a) creating art and (b) better understanding humans’ affective reactions to aesthetic stimuli. Focusing on different types of artistic content, such as movies, music, urban art, ancient artwork, and modern interactive technology, the goal of the Second International Workshop on Multimodal Affect and Aesthetic Experience is to enhance the interdisciplinary collaboration among researchers from the following domains: affective computing, aesthetics, human-robot interaction, and digital archaeology and art. Michal Muszynski, Edgar Roman-Rangel, Leimin Tian, Theodoros Kostoulas, Theodora Chaspari, Panos Amelidis |
ICMI | 1 |
| 2021 | Multimodal and Multitask Approach to Listener's Backchannel Prediction: Can Prediction of Turn-changing and Turn-management Willingness Improve Backchannel Modeling?abstractThe listener's backchannel has the important function of encouraging a current speaker to hold their turn and continue to speak, which enables smooth conversation. The listener monitors the speaker's turn-management (a.k.a. speaking and listening) willingness and his/her own willingness to display backchannel behavior. Many studies have focused on predicting the appropriate timing of the backchannel so that conversational agents can display backchannel behavior in response to a user who is speaking. To the best of our knowledge, none of them added the prediction of turn-changing and participants' turn-management willingness to the backchannel prediction model in dyad interactions. In this paper, we proposed a novel backchannel prediction model that can jointly predict turn-changing and turn-management willingness. We investigated the impact of modeling turn-changing and willingness to improve backchannel prediction. Our proposed model is based on trimodal inputs, that is, acoustic, linguistic, and visual cues from conversations. Our results suggest that adding turn-management willingness as a prediction task improves the performance of backchannel prediction within the multi-modal multi-task learning approach, while adding turn-changing prediction is not useful for improving the performance of backchannel prediction. Ryo Ishii, Xutong Ren, Michal Muszynski, Louis-Philippe Morency |
IVA | 3 |
| 2021 | Recognizing Induced Emotions of Movie Audiences from Multimodal InformationabstractRecognizing emotional reactions of movie audiences to affective movie content is a challenging task in affective computing. Previous research on induced emotion recognition has mainly focused on using audio-visual movie content. Nevertheless, the relationship between the perceptions of the affective movie content (perceived emotions) and the emotions evoked in the audiences (induced emotions) is unexplored. In this work, we studied the relationship between perceived and induced emotions of movie audiences. Moreover, we investigated multimodal modelling approaches to predict movie induced emotions from movie content based features, as well as physiological and behavioral reactions of movie audiences. To carry out analysis of induced and perceived emotions, we first extended an existing database for movie affect analysis by annotating perceived emotions in a crowd-sourced manner. We find that perceived and induced emotions are not always consistent with each other. In addition, we show that perceived emotions, movie dialogues, and aesthetic highlights are discriminative for movie induced emotion recognition besides spectators' physiological and behavioral reactions. Also, our experiments revealed that induced emotion recognition could benefit from including temporal information and performing multimodal fusion. Moreover, our work deeply investigated the gap between affective content analysis and induced emotion recognition by gaining insight into the relationships between aesthetic highlights, induced emotions, and perceived emotions. Michal Muszynski, Leimin Tian, Catherine Lai, Johanna D. Moore, Theodoros Kostoulas, Patrizia Lombardo, Thierry Pun, Guillaume Chanel |
IEEE Trans. Affect. Comput. | 1 |
| 2020 | Multimodal Affect and Aesthetic ExperienceabstractThe term 'aesthetic experience' corresponds to the inner state of a person exposed to form and content of artistic objects. Exploring certain aesthetic values of artistic objects, as well as interpreting the aesthetic experience of people when exposed to art can contribute towards understanding (a) art and (b) people's affective reactions to artwork. Focusing on different types of artistic content, such as movies, music, urban art and other artwork, the goal of this workshop is to enhance the interdisciplinary collaboration between affective computing and aesthetics researchers. Theodoros Kostoulas, Michal Muszynski, Theodora Chaspari, Panos Amelidis |
ICMI | 2 |
| 2020 | Depression Severity Assessment for Adolescents at High Risk of Mental DisordersabstractRecent progress in artificial intelligence has led to the development of automatic behavioral marker recognition, such as facial and vocal expressions. Those automatic tools have enormous potential to support mental health assessment, clinical decision making, and treatment planning. In this paper, we investigate nonverbal behavioral markers of depression severity assessed during semi-structured medical interviews of adolescent patients. The main goal of our research is two-fold: studying a unique population of adolescents at high risk of mental disorders and differentiating mild depression from moderate or severe depression. We aim to explore computationally inferred facial and vocal behavioral responses elicited by three segments of the semi-structured medical interviews: Distress Assessment Questions, Ubiquitous Questions, and Concept Questions. Our experimental methodology reflects best practise used for analyzing small sample size and unbalanced datasets of unique patients. Our results show a very interesting trend with strongly discriminative behavioral markers from both acoustic and visual modalities. These promising results are likely due to the unique classification task (mild depression vs. moderate and severe depression) and three types of probing questions. Michal Muszynski, Jamie Zelazny, Jeffrey M. Girard, Louis-Philippe Morency |
ICMI | 1 |
| 2020 | Can Prediction of Turn-management Willingness Improve Turn-changing Modeling?abstractFor smooth conversation, participants must carefully monitor the turn-management (a.k.a. speaking and listening) willingness of other conversational partners and adjust turn-changing behaviors accordingly. Many studies have focused on predicting the actual moments of speaker changes (a.k.a. turn-changing), but to the best of our knowledge, none of them explicitly modeled the turn-management willingness from both speakers and listeners in dyad interactions. We address the problem of building models for predicting this willingness of both. Our models are based on trimodal inputs, including acoustic, linguistic, and visual cues from conversations. We also study the impact of modeling willingness to help improve the task of turn-changing prediction. We introduce a dyadic conversation corpus with annotated scores of speaker/listener turn-management willingness. Our results show that using all of three modalities of speaker and listener is important for predicting turn-management willingness. Furthermore, explicitly adding willingness as a prediction task improves the performance of turn-changing prediction. Also, turn-management willingness prediction becomes more accurate with this multi-task learning approach. Ryo Ishii, Xutong Ren, Michal Muszynski, Louis-Philippe Morency |
IVA | 3 |
| 2018 | Aesthetic Highlight Detection in Movies Based on Synchronization of Spectators' ReactionsabstractDetection of aesthetic highlights is a challenge for understanding the affective processes taking place during movie watching. In this article, we study spectators’ responses to movie aesthetic stimuli in a social context. Moreover, we look for uncovering the emotional component of aesthetic highlights in movies. Our assumption is that synchronized spectators’ physiological and behavioral reactions occur during these highlights because: ( i ) aesthetic choices of filmmakers are made to elicit specific emotional reactions (e.g., special effects, empathy, and compassion toward a character) and ( ii ) watching a movie together causes spectators’ affective reactions to be synchronized through emotional contagion. We compare different approaches to estimation of synchronization among multiple spectators’ signals, such as pairwise, group, and overall synchronization measures to detect aesthetic highlights in movies. The results show that the unsupervised architecture relying on synchronization measures is able to capture different properties of spectators’ synchronization and detect aesthetic highlights based on both spectators’ electrodermal and acceleration signals. We discover that pairwise synchronization measures perform the most accurately independently of the category of the highlights and movie genres. Moreover, we observe that electrodermal signals have more discriminative power than acceleration signals for highlight detection. Michal Muszynski, Theodoros Kostoulas, Patrizia Lombardo, Thierry Pun, Guillaume Chanel |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2017 | Recognizing induced emotions of movie audiences: Are induced and perceived emotions the same?abstractPredicting the emotional response of movie audiences to affective movie content is a challenging task in affective computing. Previous work has focused on using audiovisual movie content to predict movie induced emotions. However, the relationship between the audience's perceptions of the affective movie content (perceived emotions) and the emotions evoked in the audience (induced emotions) remains unexplored. In this work, we address the relationship between perceived and induced emotions in movies, and identify features and modelling approaches effective for predicting movie induced emotions. First, we extend the LIRIS-ACCEDE database by annotating perceived emotions in a crowd-sourced manner, and find that perceived and induced emotions are not always consistent. Second, we show that dialogue events and aesthetic highlights are effective predictors of movie induced emotions. In addition to movie based features, we also study physiological and behavioural measurements of audiences. Our experiments show that induced emotion recognition can benefit from including temporal context and from including multimodal information. Our study bridges the gap between affective content analysis and induced emotion prediction. Leimin Tian, Michal Muszynski, Catherine Lai, Johanna D. Moore, Theodoros Kostoulas, Patrizia Lombardo, Thierry Pun, Guillaume Chanel |
ACII | 2 |
| 2016 | Synchronization among Groups of Spectators for Highlight Detection in MoviesabstractDetection of emotional and aesthetic highlights is a challenge for the affective understanding of movies. Our assumption is that synchronized spectators' physiological and behavioral reactions occur during these highlights. We propose to employ the periodicity score to capture synchronization among groups of spectators' signals. To uncover the periodicity score's capabilities, we compare it with baseline synchronization measures, such as the nonlinear interdependence and the windowed mutual information. The results show that the periodicity score and the pairwise synchronization measures are able to capture different properties of spectators' synchronization, and they indicate the presence of some types of emotional and aesthetic highlights in a movie based on spectators' electro-dermal and acceleration signals. Michal Muszynski, Theodoros Kostoulas, Patrizia Lombardo, Thierry Pun, Guillaume Chanel |
ACM Multimedia | 1 |
| 2015 | Spectators' Synchronization Detection based on Manifold Representation of Physiological Signals: Application to Movie Highlights DetectionabstractDetection of highlights in movies is a challenge for the affective understanding and implicit tagging of films. Under the hypothesis that synchronization of the reaction of spectators indicates such highlights, we define a synchronization measure between spectators that is capable of extracting movie highlights. The intuitive idea of our approach is to define (a) a parameterization of one spectator's physiological data on a manifold; (b) the synchronization measure between spectators as the Kolmogorov-Smirnov distance between local shape distributions of the underlying manifolds. We evaluate our approach using data collected in an experiment where the electro-dermal activity of spectators was recorded during the entire projection of a movie in a cinema. We compare our methodology with baseline synchronization measures, such as correlation, Spearman's rank correlation, mutual information, Kolmogorov-Smirnov distance. Results indicate that the proposed approach allows to accurately distinguish highlight from non-highlight scenes. Michal Muszynski, Theodoros Kostoulas, Guillaume Chanel, Patrizia Lombardo, Thierry Pun |
ICMI | 1 |