VLDB 2026 Research / reviewers in the wild / expert
Mohsen Fayyaz
dblp:163/2062
· DBLP profile ↗
20ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-5254-3128ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Polymorph: Energy-Efficient Multi-Label Classification for Video Streams on Embedded DevicesabstractReal-time multi-label video classification on embedded devices is constrained by limited compute and energy budgets. Yet, video streams exhibit structural properties such as label sparsity, temporal continuity, and label co-occurrence that can be leveraged for more efficient inference. We introduce Polymorph, a context-aware framework that activates a minimal set of lightweight Low Rank Adapters (LoRA) per frame. Each adapter specializes in a subset of classes derived from co-occurrence patterns and is implemented as a LoRA weight over a shared backbone. At runtime, Polymorph dynamically selects and composes only the adapters needed to cover the active labels, avoiding fullmodel switching and weight merging. This modular strategy improves scalability while reducing latency, and energy overhead. Polymorph achieves 40% lower energy consumption and improves mAP by 9 points over strong baselines executing the TAO dataset. Saeid Ghafouri, Mohsen Fayyaz, Xiangchen Li, Chacko John Deepu, Bo Ji 0001, Dimitrios S. Nikolopoulos, Hans Vandierendonck |
WACV | 2 |
| 2025 | Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual EvidenceabstractDense retrieval models are commonly used in Information Retrieval (IR) applications, such as Retrieval-Augmented Generation (RAG). Since they often serve as the first step in these systems, their robustness is critical to avoid downstream failures. In this work, we repurpose a relation extraction dataset (e.g., Re-DocRED) to design controlled experiments that quantify the impact of heuristic biases, such as a preference for shorter documents, on retrievers like Dragon+ and Contriever. We uncover major vulnerabilities, showing retrievers favor shorter documents, early positions, repeated entities, and literal matches, all while ignoring the answer’s presence! Notably, when multiple biases combine, models exhibit catastrophic performance degradation, selecting the answer-containing document in less than 10% of cases over a synthetic biased document without the answer. Furthermore, we show that these biases have direct consequences for downstream applications like RAG, where retrieval-preferred documents can mislead LLMs, resulting in a 34% performance drop than providing no documents at all.https://huggingface.co/datasets/mohsenfayyaz/ColDeR Mohsen Fayyaz, Ali Modarressi, Hinrich Schütze, Nanyun Peng 0001 |
ACL (1) | 1 |
| 2025 | MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal ModelsabstractExisting multimodal retrieval benchmarks primarily focus on evaluating whether models can retrieve and utilize external textual knowledge for question answering. However, there are scenarios where retrieving visual information is either more beneficial or easier to access than textual data.
In this paper, we introduce a multimodal retrieval-augmented generation benchmark, MRAG-Bench, in which we systematically identify and categorize scenarios where visually augmented knowledge is better than textual knowledge, for instance, more images from varying viewpoints.
MRAG-Bench consists of 16,130 images and 1,353 human-annotated multiple-choice questions across 9 distinct scenarios. With MRAG-Bench, we conduct an evaluation of 10 open-source and 4 proprietary large vision-language models (LVLMs). Our results show that all LVLMs exhibit greater improvements when augmented with images compared to textual knowledge, confirming that MRAG-Bench is vision-centric. Additionally, we conduct extensive analysis with MRAG-Bench, which offers valuable insights into retrieval-augmented LVLMs. Notably, the top-performing model, GPT-4o, faces challenges in effectively leveraging retrieved knowledge, achieving only a 5.82\% improvement with ground-truth information, in contrast to a 33.16\% improvement observed in human participants. These findings highlight the importance of MRAG-Bench in encouraging the community to enhance LVLMs' ability to utilize retrieved visual knowledge more effectively. Wenbo Hu 0006, Jia-Chen Gu, Zi-Yi Dou, Mohsen Fayyaz, Pan Lu, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ICLR | 4 |
| 2025 | Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision EncodersabstractDespite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that current Video Large Language Model (Video-LLM) architectures have critical limitations in temporal understanding, struggling with tasks that require detailed comprehension of action sequences and temporal progression. In this work, we propose a Video-LLM architecture that introduces stacked temporal attention modules directly within the vision encoder. This design incorporates a temporal attention in vision encoder, enabling the model to better capture the progression of actions and the relationships between frames before passing visual tokens to the LLM. Our results show that this approach significantly improves temporal reasoning and outperforms existing models in video question answering tasks, specifically in action recognition. We improve on benchmarks including VITATECS, MVBench, and Video-MME by up to +5.5%. By enhancing the vision encoder with temporal structure, we address a critical gap in video understanding for Video-LLMs. Project page and code are available at: https://alirasekh.github.io/STAVEQ2/ Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk 0001, Mohsen Fayyaz |
NeurIPS | 5 |
| 2024 | Occlusion Handling in 3D Human Pose Estimation with Perturbed Positional Encoding
Niloofar Azizi, Mohsen Fayyaz, Horst Bischof |
ECCV (13) | 2 |
| 2023 | DecompX: Explaining Transformers Decisions by Propagating Token DecompositionabstractAli Modarressi, Mohsen Fayyaz, Ehsan Aghazadeh, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ali Modarressi, Mohsen Fayyaz, Ehsan Aghazadeh, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar |
ACL (1) | 2 |
| 2023 | Diffusion models in medical imaging: A comprehensive survey
Amirhossein Kazerouni, Ehsan Khodapanah Aghdam, Moein Heidari, Reza Azad, Mohsen Fayyaz, Ilker Hacihaliloglu, Dorit Merhof |
Medical Image Anal. | 5 |
| 2022 | Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and LanguagesabstractHuman languages are full of metaphorical expressions.Metaphors help people understand the world by connecting new concepts and domains to more familiar ones.Large pretrained language models (PLMs) are therefore assumed to encode metaphorical knowledge useful for NLP systems.In this paper, we investigate this hypothesis for PLMs, by probing metaphoricity information in their encodings, and by measuring the cross-lingual and crossdataset generalization of this information.We present studies in multiple metaphor detection datasets and in four languages (i.e., English, Spanish, Russian, and Farsi).Our extensive experiments suggest that contextual representations in PLMs do encode metaphorical knowledge, and mostly in their middle layers.The knowledge is transferable between languages and datasets, especially when the annotation is consistent across training and testing sets.Our findings give helpful insights for both cognitive and NLP scientists. Ehsan Aghazadeh, Mohsen Fayyaz, Yadollah Yaghoobzadeh |
ACL (1) | 2 |
| 2022 | TaylorSwiftNet: Taylor Driven Temporal Modeling for Swift Future Frame Prediction
Mohammad Saber Pourheydari, Emad Bahrami Rad, Mohsen Fayyaz, Gianpiero Francesca, Mehdi Noroozi, Juergen Gall |
BMVC | 3 |
| 2022 | Adaptive Token Sampling for Efficient Vision Transformers
Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsiavash, Juergen Gall |
ECCV (11) | 1 |
| 2022 | GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in TransformersabstractAli Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar |
NAACL-HLT | 2 |
| 2022 | Fast Weakly Supervised Action Segmentation Using Mutual ConsistencyabstractAction segmentation is the task of predicting the actions for each frame of a video. As obtaining the full annotation of videos for action segmentation is expensive, weakly supervised approaches that can learn only from transcripts are appealing. In this paper, we propose a novel end-to-end approach for weakly supervised action segmentation based on a two-branch neural network. The two branches of our network predict two redundant but different representations for action segmentation and we propose a novel mutual consistency (MuCon) loss that enforces the consistency of the two redundant representations. Using the MuCon loss together with a loss for transcript prediction, our proposed approach achieves the accuracy of state-of-the-art approaches while being 14 times faster to train and 20 times faster during inference. The MuCon loss proves beneficial even in the fully supervised setting. Yaser Souri, Mohsen Fayyaz, Luca Minciullo, Gianpiero Francesca, Juergen Gall |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | 3D CNNs With Adaptive Temporal Feature ResolutionsabstractWhile state-of-the-art 3D Convolutional Neural Networks (CNN) achieve very good results on action recognition datasets, they are computationally very expensive and require many GFLOPs. While the GFLOPs of a 3D CNN can be decreased by reducing the temporal feature resolution within the network, there is no setting that is optimal for all input clips. In this work, we therefore introduce a differentiable Similarity Guided Sampling (SGS) module, which can be plugged into any existing 3D CNN architecture. SGS empowers 3D CNNs by learning the similarity of temporal features and grouping similar features together. As a result, the temporal feature resolution is not anymore static but it varies for each input video clip. By integrating SGS as an additional layer within current 3D CNNs, we can convert them into much more efficient 3D CNNs with adaptive temporal feature resolutions (ATFR). Our evaluations show that the proposed module improves the state-of-the-art by reducing the computational cost (GFLOPs) by half while preserving or even improving the accuracy. We evaluate our module by adding it to multiple state-of-the-art 3D CNNs on various datasets such as Kinetics-600, Kinetics-400, mini-Kinetics, Something-Something V2, UCF101, and HMDB51. Mohsen Fayyaz, Emad Bahrami Rad, Ali Diba, Mehdi Noroozi, Ehsan Adeli-Mosabbeb, Luc Van Gool, Juergen Gall |
CVPR | 1 |
| 2021 | Long Short View Feature Decomposition via Contrastive Video Representation LearningabstractSelf-supervised video representation methods typically focus on the representation of temporal attributes in videos. However, the role of stationary versus non-stationary attributes is less explored: Stationary features, which remain similar throughout the video, enable the prediction of video-level action classes. Non-stationary features, which represent temporally varying attributes, are more beneficial for downstream tasks involving more fine-grained temporal understanding, such as action segmentation. We argue that a single representation to capture both types of features is sub-optimal, and propose to decompose the representation space into stationary and non-stationary features via contrastive learning from long and short views, i.e. long video sequences and their shorter sub-sequences. Stationary features are shared between the short and long views, while non-stationary features aggregate the short views to match the corresponding long view. To empirically verify our approach, we demonstrate that our stationary features work particularly well on an action recognition downstream task, while our non-stationary features perform better on action segmentation. Furthermore, we analyse the learned representations and find that stationary features capture more temporally stable, static attributes, while non-stationary features encompass more temporally varying ones. Nadine Behrmann, Mohsen Fayyaz, Juergen Gall, Mehdi Noroozi |
ICCV | 2 |
| 2020 | SCT: Set Constrained Temporal Transformer for Set Supervised Action SegmentationabstractTemporal action segmentation is a topic of increasing interest, however, annotating each frame in a video is cumbersome and costly. Weakly supervised approaches therefore aim at learning temporal action segmentation from videos that are only weakly labeled. In this work, we assume that for each training video only the list of actions is given that occur in the video, but not when, how often, and in which order they occur. In order to address this task, we propose an approach that can be trained end-to-end on such data. The approach divides the video into smaller temporal regions and predicts for each region the action label and its length. In addition, the network estimates the action labels for each frame. By measuring how consistent the frame-wise predictions are with respect to the temporal regions and the annotated action labels, the network learns to divide a video into class-consistent regions. We evaluate our approach on three datasets where the approach achieves state-of-the-art results. Mohsen Fayyaz, Juergen Gall |
CVPR | 1 |
| 2020 | Large Scale Holistic Video Understanding
Ali Diba, Mohsen Fayyaz, Vivek Sharma 0001, Manohar Paluri, Juergen Gall, Rainer Stiefelhagen, Luc Van Gool |
ECCV (5) | 2 |
| 2018 | AVID: Adversarial Visual Irregularity Detection
Mohammad Sabokrou, Masoud PourReza, Mohsen Fayyaz, Rahim Entezari, Mahmood Fathy, Juergen Gall, Ehsan Adeli-Mosabbeb |
ACCV (6) | 3 |
| 2018 | Spatio-temporal Channel Correlation Networks for Action Classification
Ali Diba, Mohsen Fayyaz, Vivek Sharma 0001, Mohammad Mahdi Arzani, Rahman Yousefzadeh, Juergen Gall, Luc Van Gool |
ECCV (4) | 2 |
| 2018 | Deep-anomaly: Fully convolutional neural network for fast anomaly detection in crowded scenes
Mohammad Sabokrou, Mohsen Fayyaz, Mahmood Fathy, Zahra Moayed, Reinhard Klette |
Comput. Vis. Image Underst. | 2 |
| 2017 | Deep-Cascade: Cascading 3D Deep Neural Networks for Fast Anomaly Detection and Localization in Crowded ScenesabstractThis paper proposes a fast and reliable method for anomaly detection and localization in video data showing crowded scenes. Time-efficient anomaly localization is an ongoing challenge and subject of this paper. We propose a cubicpatch- based method, characterised by a cascade of classifiers, which makes use of an advanced feature-learning approach. Our cascade of classifiers has two main stages. First, a light but deep 3D auto-encoder is used for early identification of "many" normal cubic patches. This deep network operates on small cubic patches as being the first stage, before carefully resizing remaining candidates of interest, and evaluating those at the second stage using a more complex and deeper 3D convolutional neural network (CNN). We divide the deep autoencoder and the CNN into multiple sub-stages which operate as cascaded classifiers. Shallow layers of the cascaded deep networks (designed as Gaussian classifiers, acting as weak single-class classifiers) detect "simple" normal patches such as background patches, and more complex normal patches are detected at deeper layers. It is shown that the proposed novel technique (a cascade of two cascaded classifiers) performs comparable to current top-performing detection and localization methods on standard benchmarks, but outperforms those in general with respect to required computation time. Mohammad Sabokrou, Mohsen Fayyaz, Mahmood Fathy, Reinhard Klette |
IEEE Trans. Image Process. | 2 |