VLDB 2026 Research / reviewers in the wild / expert
Ronak Kosti
dblp:205/3984
· DBLP profile ↗
11ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0003-2453-7876ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Unsupervised Domain Adaptation Using Soft-Labeled Contrastive Learning with Reversed Monte Carlo Method for Cardiac Image Segmentation
Mingxuan Gu, Mareike Thies, Siyuan Mei, Fabian Wagner, Mingcheng Fan, Yipeng Sun, Zhaoya Pan, Sulaiman Vesal, Ronak Kosti, Dennis Possart, Jonas Utz, Andreas K. Maier |
MICCAI (9) | 9 |
| 2023 | SUMAC '23: 5th Workshop on the analySis, Understanding and proMotion of heritAge Contents: Advances in Machine Learning, Signal Processing, Multimodal Techniques and Human-machine InteractionabstractSUMAC 2023 is the fifth edition of the workshop on analySis, Understanding and proMotion of heritAge Contents. It is held in Ottawa, Canada on November 2, 2023 and is co-located with the 31st ACM International Conference on Multimedia. The workshop's objective is to present and discuss the latest and most significant trends, challenges and advances in the fields of machine learning, signal processing, multimodal techniques and human-machine interaction. The workshop is dedicated to the valorization of cultural heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field. The complete SUMAC'23 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3581783.3610949. Valérie Gouet-Brunet, Ronak Kosti, Li Weng |
ACM Multimedia | 2 |
| 2023 | ICC++: Explainable feature learning for art history using image compositions
Prathmesh Madhu, Tilman Marquart, Ronak Kosti, Dirk Suckow, Peter Bell 0007, Andreas K. Maier, Vincent Christlein |
Pattern Recognit. | 3 |
| 2022 | Supervised Contrastive Learning for Robust and Efficient Multi-modal Emotion and Sentiment AnalysisabstractExpression of human emotion and sentiment are often multi-modal consisting use of spoken speech, vision, and text. Combining multiple modalities allows learning-based models to benefit with the complementary information present across modalities to produce more accurate predictions. One of the bigger challenges in multi-modal affective computing is performance consistency in non-ideal scenarios. Most benchmarks fail to generalize in non-ideal scenarios where one of the modalities is missing or highly corrupted due to occlusion, sensor errors, or change of orientation. Consequently, various modality fusion approaches were proposed. However, most of these fusion approaches assume that each modality is equally useful. To address the challenge of performance consistency, in this work we propose to use supervised contrastive learning (SCL). We demonstrate through various experiments and comparison with state-of-the-art (SOTA) methods that the model robustness against corrupted and missing modalities improves when trained with SCL. Next, we use the Perceiver architecture [1] in order to efficiently combine the representations of different modalities. Its iterative attention mechanism allows to create a reduced latent representation in an efficient manner. We observe that it can accommodate a wide range of modality combinations, allowing for robust information fusion. Our approach allows reduction of model complexity and efficient fusion of different modalities, while maintaining the performance consistency and model robustness. We conduct ablation experiments to study the effect of each contribution in different scenarios, and we show that the proposed methods outperform the state-of-art, while simultaneously being robust to corrupted modalities. Our method also outperforms its counterparts and SOTA while using less numerical complexity (inference times and compute operations). Ahmed Gomaa, Andreas K. Maier, Ronak Kosti |
ICPR | 3 |
| 2022 | ODOR: The ICPR2022 ODeuropa Challenge on Olfactory Object RecognitionabstractThe Odeuropa Challenge on Olfactory Object Recognition aims to foster the development of object detection in the visual arts and to promote an olfactory perspective on digital heritage. Object detection in historical artworks is particularly challenging due to varying styles and artistic periods. Moreover, the task is complicated due to the particularity and historical variance of predefined target objects, which exhibit a large intra-class variance, and the long tail distribution of the dataset labels, with some objects having only very few training examples. These challenges should encourage participants to create innovative approaches using domain adaptation or few-shot learning. We provide a dataset of 2647 artworks annotated with 20 120 tightly fit bounding boxes that are split into a training and validation set (public). A test set containing 1140 artworks and 15 480 annotations is kept private for the challenge evaluation. Mathias Zinnen, Prathmesh Madhu, Ronak Kosti, Peter Bell 0007, Andreas K. Maier, Vincent Christlein |
ICPR | 3 |
| 2022 | SUMAC '22: 4th ACM International workshop on Structuring and Understanding of Multimedia heritAge ContentsabstractSUMAC 2022 is the fourth edition of the workshop on Structuring and Understanding of Multimedia heritAge Contents. It is held in Lisboa, Portugal on October 10th, 2022 and is co-located with the 30th ACM International Conference on Multimedia. Its objective is to present and discuss the latest and most significant trends and challenges in the analysis, structuring and understanding of multimedia contents dedicated to the valorization of heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field. The complete SUMAC'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552464 Valérie Gouet-Brunet, Ronak Kosti, Li Weng |
ACM Multimedia | 2 |
| 2021 | SUMAC'21: 3rd Workshop on Structuring and Understanding of Multimedia heritAge ContentsabstractSUMAC 2021 is the third edition of the workshop on Structuring and Understanding of Multimedia heritAge Contents. It is held in Chengdu, China on October 20th, 2021 and is co-located with the 29th ACM International Conference on Multimedia. Its objective is to present and discuss the latest and most significant trends and challenges in the analysis, structuring and understanding of multimedia contents dedicated to the valorization of heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field. Valérie Gouet-Brunet, Margarita Khokhlova, Ronak Kosti, Li Weng |
ACM Multimedia | 3 |
| 2021 | Adapt Everywhere: Unsupervised Adaptation of Point-Clouds and Entropy Minimization for Multi-Modal Cardiac Image SegmentationabstractDeep learning models are sensitive to domain shift phenomena. A model trained on images from one domain cannot generalise well when tested on images from a different domain, despite capturing similar anatomical structures. It is mainly because the data distribution between the two domains is different. Moreover, creating annotation for every new modality is a tedious and time-consuming task, which also suffers from high inter- and intra- observer variability. Unsupervised domain adaptation (UDA) methods intend to reduce the gap between source and target domains by leveraging source domain labelled data to generate labels for the target domain. However, current state-of-the-art (SOTA) UDA methods demonstrate degraded performance when there is insufficient data in source and target domains. In this paper, we present a novel UDA method for multi-modal cardiac image segmentation. The proposed method is based on adversarial learning and adapts network features between source and target domain in different spaces. The paper introduces an end-to-end framework that integrates: a) entropy minimization, b) output feature space alignment and c) a novel point-cloud shape adaptation based on the latent features learned by the segmentation model. We validated our method on two cardiac datasets by adapting from the annotated source domain, bSSFP-MRI (balanced Steady-State Free Procession-MRI), to the unannotated target domain, LGE-MRI (Late-gadolinium enhance-MRI), for the multi-sequence dataset; and from MRI (source) to CT (target) for the cross-modality dataset. The results highlighted that by enforcing adversarial learning in different parts of the network, the proposed method delivered promising performance, compared to other SOTA methods. Sulaiman Vesal, Mingxuan Gu, Ronak Kosti, Andreas K. Maier, Nishant Ravikumar |
IEEE Trans. Medical Imaging | 3 |
| 2020 | SUMAC 2020: The 2nd Workshop on Structuring and Understanding of Multimedia heritAge ContentsabstractSUMAC 2020 is the second edition of the workshop on Structuring and Understanding of Multimedia heritAge Contents. It is held in Seattle, USA on October 12th, 2020 and is co-located with the 28th ACM International Conference on Multimedia; this year, due to the sanitary crisis, it is organized virtually. Its objective is to present and discuss the latest and most significant trends and challenges in the analysis, structuring and understanding of multimedia contents dedicated to the valorization of heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field. Valérie Gouet-Brunet, Margarita Khokhlova, Ronak Kosti, Liming Chen 0002, Xu-Cheng Yin |
ACM Multimedia | 3 |
| 2020 | Context Based Emotion Recognition Using EMOTIC DatasetabstractIn our everyday lives and social interactions we often try to perceive the emotional states of people. There has been a lot of research in providing machines with a similar capacity of recognizing emotions. From a computer vision perspective, most of the previous efforts have been focusing in analyzing the facial expressions and, in some cases, also the body pose. Some of these methods work remarkably well in specific settings. However, their performance is limited in natural, unconstrained environments. Psychological studies show that the scene context, in addition to facial expression and body pose, provides important information to our perception of people's emotions. However, the processing of the context for automatic emotion recognition has not been explored in depth, partly due to the lack of proper data. In this paper we present EMOTIC, a dataset of images of people in a diverse set of natural situations, annotated with their apparent emotion. The EMOTIC dataset combines two different types of emotion representation: (1) a set of 26 discrete categories, and (2) the continuous dimensions Valence, Arousal, and Dominance. We also present a detailed statistical and algorithmic analysis of the dataset along with annotators' agreement analysis. Using the EMOTIC dataset we train different CNN models for emotion recognition, combining the information of the bounding box containing the person with the contextual information extracted from the scene. Our results show how scene context provides important information to automatically recognize emotional states and motivate further research in this direction. Ronak Kosti, José M. Álvarez 0004, Adrià Recasens, Àgata Lapedriza |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Emotion Recognition in ContextabstractUnderstanding what a person is experiencing from her frame of reference is essential in our everyday life. For this reason, one can think that machines with this type of ability would interact better with people. However, there are no current systems capable of understanding in detail peoples emotional states. Previous research on computer vision to recognize emotions has mainly focused on analyzing the facial expression, usually classifying it into the 6 basic emotions [11]. However, the context plays an important role in emotion perception, and when the context is incorporated, we can infer more emotional states. In this paper we present the Emotions in Context Database (EMCO), a dataset of images containing people in context in non-controlled environments. In these images, people are annotated with 26 emotional categories and also with the continuous dimensions valence, arousal, and dominance [21]. With the EMCO dataset, we trained a Convolutional Neural Network model that jointly analyzes the person and the whole scene to recognize rich information about emotional states. With this, we show the importance of considering the context for recognizing peoples emotions in images, and provide a benchmark in the task of emotion recognition in visual context. Ronak Kosti, José M. Álvarez 0004, Adrià Recasens, Àgata Lapedriza |
CVPR | 1 |