EDBT 2026 Demo / reviewers in the wild / expert
Rupayan Mallick
dblp:278/3028
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-3335-2753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ViSTA: Visual Storytelling using Multi-modal Adapters for Text-to-Image Diffusion ModelsabstractText-to-image diffusion models have achieved remarkable success, yet generating coherent image sequences for visual storytelling remains challenging. A key challenge is effectively leveraging all previous text-image pairs, referred to as history text-image pairs, which provide contextual information for maintaining consistency across frames. Existing auto-regressive methods condition on all past image-text pairs but require extensive training, while training-free subject-specific approaches ensure consistency but lack adaptability to narrative prompts. To address these limitations, we propose a multi-modal history adapter for text-to-image diffusion models, ViSTA. It consists of (1) a multimodal history fusion module to extract relevant history features and (2) a history adapter to condition the generation on the extracted relevant features. We also introduce a salient history selection strategy during inference, where the most salient history text-image pair is selected, improving the quality of the conditioning. Furthermore, we propose to employ a Visual Question Answering-based metric TIFA to assess text-image alignment in visual storytelling, providing a more targeted and interpretable assessment of generated images. Evaluated on the StorySalon and FlintStonesSV dataset, our proposed ViSTA model is not only consistent across different frames, but also well-aligned with the narrative text descriptions. Sibo Dong, Ismail Shaheen, Maggie Shen, Rupayan Mallick, Sarah Adel Bargal |
WACV | 4 |
| 2026 | Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTyabstractDifferent forms of customized 2D avatars are widely used in gaming applications, virtual communication, education, and content creation. However, existing approaches often fail to capture fine-grained facial expressions and struggle to preserve identity across different expressions. We propose Gen-AFFECT, a novel framework for personalized avatar generation that generates expressive and identity-consistent avatars with a diverse set of facial expressions. Our framework proposes conditioning a multimodal diffusion transformer on an extracted identity-expression representation. This enables identity preservation and representation of a wide range of facial expressions. Gen-AFFECT additionally employs consistent attention at inference for information sharing across the set of generated expressions, enabling the generation process to maintain identity consistency over the array of generated fine-grained expressions. Gen-AFFECT demonstrates superior performance compared to previous state-of-the-art methods on the basis of the accuracy of the generated expressions, the preservation of the identity and the consistency of the target identity across an array of fine-grained facial expressions. Hao Yu 0014, Rupayan Mallick, Margrit Betke, Sarah Adel Bargal |
WACV | 2 |
| 2024 | IFI: Interpreting for Improving: A Multimodal Transformer with an Interpretability Technique for Recognition of Risk Events
Rupayan Mallick, Jenny Benois-Pineau, Akka Zemmari |
MMM (4) | 1 |
| 2024 | A hybrid transformer with domain adaptation using interpretability techniques for the application to the detection of risk situations
Rupayan Mallick, Jenny Benois-Pineau, Akka Zemmari, Kamel Guerda, Boris Mansencal, Hélène Amieva, Laura Middleton |
Multim. Tools Appl. | 1 |
| 2022 | I Saw: A Self-Attention Weighted Method for Explanation of Visual TransformersabstractRecently, visual transformers have shown promising results in tasks such as image classification, segmentation, object detection, etc. The explanation of their decision remains a challenge. This paper focuses on exploiting self-attention for an explanation. We propose a generalized interpretation of the transformers i.e model agnostic but class-specific explanations. The main principle is in the use and weighting self-attention maps of a visual transformer. To evaluate it, we use the popular hypothesis that an explanation is good if it correlates with human perception of a visual scene. Thus, the method has been evaluated against the Gaze Fixation Density Maps obtained in a psycho-visual experiment on a public database. It has been compared with other popular explainers such as Grad-Cam, LRP, Rollout, and Adaptive Relevance methods. The proposed method outperforms the best baseline by 2% in a standard Pearson Correlation Coefficient (PCC) metric. Rupayan Mallick, Jenny Benois-Pineau, Akka Zemmari |
ICIP | 1 |
| 2022 | Pooling Transformer for Detection of Risk Events in In-The-Wild Video Ego DataabstractThe paper proposes a video transformer architecture for detection of risk events on frail adults with ego video monitoring data. First we introduce an extended taxonomy for risk events, and then we propose a transformer based video recognition model for detection of these risk events. The proposed transformer architecture consists of separable attention for spatial and temporal data. We also introduce a pooling operation on the temporal video data by learning of their importance. The experiments have been conducted on visual data of in-the-wild recorded BIRDS dataset and on Kinetics-400 for benchmarking. The use of the pooling operation in transformers gives an increment of 3% on BIRDS dataset. Rupayan Mallick, Jenny Benois-Pineau, Akka Zemmari, Thinhinane Yebda, Marion Pech, Hélène Amieva, Laura Middleton |
ICPR | 1 |
| 2021 | A GRU Neural Network with attention mechanism for detection of risk situations on multimodal lifelog dataabstractMultimedia today is also in multimodality. Working with heterogeneous signals we use multimedia techniques of data fusion and mining. Classification from real world datasets are often challenging. The paper is devoted to the detection of personal risk situations of fragile people from multi-modal sensing real world lifelog data named BIRDS. Using a real-world data is challenging as the risk situations are rare and last just a few seconds compared to the global volume of the dataset. In this paper we propose a GRU architecture with global attention block to recognise semantic risk situations from a limited taxonomy. Attention is also focused on data organisation and pre-processing with imputation and normalisation. The proposed method is applied to a real-world collected multimodal dataset and to the OpenSource dataset UCI-HAR for the sake of comparison with the state-of-the-art. Rupayan Mallick, Thinhinane Yebda, Jenny Benois-Pineau, Akka Zemmari, Marion Pech, Hélène Amieva |
CBMI | 1 |
| 2021 | Selective Spatio-Temporal Aggregation Based Pose Refinement System: Towards Understanding Human Activities in Real-World VideosabstractTaking advantage of human pose data for understanding human activities has attracted much attention these days. However, state-of-the-art pose estimators struggle in obtaining high-quality 2D or 3D pose data due to occlusion, truncation and low-resolution in real-world un-annotated videos. Hence, in this work, we propose 1) a Selective Spatio-Temporal Aggregation mechanism, named SST-A, that refines and smooths the keypoint locations extracted by multiple expert pose estimators, 2) an effective weakly-supervised self-training framework which leverages the aggregated poses as pseudo ground-truth in-stead of handcrafted annotations for real-world pose estimation. Extensive experiments are conducted for evaluating not only the upstream pose refinement but also the downstream action recognition performance on four datasets, Toyota Smarthome, NTU-RGB+D, Charades, and Kinetics-50. We demonstrate that the skeleton data refined by our Pose-Refinement system (SSTA-PRS) is effective at boosting various existing action recognition models, which achieves competitive or state-of-the-art performance. Di Yang 0002, Rui Dai 0001, Yaohui Wang 0001, Rupayan Mallick, Luca Minciullo, Gianpiero Francesca, François Brémond |
WACV | 4 |