Amir Ghodrati

dblp:118/3233 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-8320-9269ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Mobile Video Diffusion
abstract
Video diffusion models have achieved impressive realism and controllability but are limited by high computational demands, restricting their use on mobile devices. This paper introduces the first mobile-optimized video diffusion model. Starting from a spatio-temporal UNet from Stable Video Diffusion (SVD), we reduce memory and computational cost by reducing the frame resolution, incorporating multi-scale temporal representations, and introducing two novel pruning schema to reduce the number of channels and temporal blocks. Furthermore, we employ adversarial finetuning to reduce the denoising to a single step. Our model, coined as MobileVD, is 523x more efficient (1817.2 vs. 4.34 TFLOPs) with a slight quality drop (FVD 149 vs. 171), generating latents for a 14x512x256 px clip in 1.7 seconds on a Xiaomi-14 Pro. Our results are available at https://qualcomm-ai-research.github.io/mobile-video-diffusion/
Haitam Ben Yahia, Denis Korzhenkov, Ioannis Lelekas, Amir Ghodrati, AmirHossein Habibian
ICCV4
2024 Clockwork Diffusion: Efficient Generation With Model-Step Distillation
abstract
This work aims to improve the efficiency of text-to-image diffusion models. While diffusion models use computationallyexpensive UNet-based denoising operations in ev-ery generation step, we identify that not all operations are equally relevant for the final output quality. In par-ticular, we observe that UNet layers operating on high-res feature maps are relatively sensitive to small pertur-bations. In contrast, low-res feature maps influence the semantic layout of the final image and can often be per-turbed with no noticeable change in the output. Based on this observation, we propose Clockwork Diffusion, a method that periodically reuses computation from preceding denoising steps to approximate low-res feature maps at one or more subsequent steps. For multiple base-lines, and for both text-to-image generation and image editing, we demonstrate that Clockwork leads to compa-rable or improved perceptual scores with drastically re-duced computational complexity. As an example, for Sta-ble Diffusion vI.5 with 8 DPM++ steps we save 32% of FLOPs with negligible FID and CLIP change. We re-lease code at https://github.com/Qualcomm-AI-research/clockwork-diffusion
AmirHossein Habibian, Amir Ghodrati, Noor Fathima, Guillaume Sautière, Risheek Garrepalli, Fatih Porikli, Jens Petersen
CVPR2
2024 Skip-Attention: Improving Vision Transformers by Paying Less Attention
abstract
This work aims to improve the efficiency of vision transformers (ViTs). While ViTs use computationally expensive self-attention operations in every layer, we identify that these operations are highly correlated across layers -- a key redundancy that causes unnecessary computations. Based on this observation, we propose SkipAT a method to reuse self-attention computation from preceding layers to approximate attention at one or more subsequent layers. To ensure that reusing self-attention blocks across layers does not degrade the performance, we introduce a simple parametric function, which outperforms the baseline transformer's performance while running computationally faster. We show that SkipAT is agnostic to transformer architecture and is effective in image classification, semantic segmentation on ADE20K, image denoising on SIDD, and video denoising on DAVIS. We achieve improved throughput at the same-or-higher accuracy levels in all these tasks.
Shashanka Venkataramanan, Amir Ghodrati, Yuki Markus Asano, Fatih Porikli, AmirHossein Habibian
ICLR2
2022 SALISA: Saliency-Based Input Sampling for Efficient Video Object Detection
Babak Ehteshami Bejnordi, AmirHossein Habibian, Fatih Porikli, Amir Ghodrati
ECCV (10)4
2021 Efficient Video Super Resolution by Gated Local Self Attention
Davide Abati, Amir Ghodrati, AmirHossein Habibian
BMVC2
2021 Conditional Model Selection for Efficient Video Understanding
Mihir Jain, Haitam Ben Yahia, Amir Ghodrati, AmirHossein Habibian, Fatih Porikli
BMVC3
2021 FrameExit: Conditional Early Exiting for Efficient Video Recognition
abstract
In this paper, we propose a conditional early exiting framework for efficient video recognition. While existing works focus on selecting a subset of salient frames to re-duce the computation costs, we propose to use a simple sampling strategy combined with conditional early exiting to enable efficient recognition. Our model automatically learns to process fewer frames for simpler videos and more frames for complex ones. To achieve this, we employ a cascade of gating modules to automatically determine the earliest point in processing where an inference is sufficiently reliable. We generate on-the-fly supervision signals to the gates to provide a dynamic trade-off between accuracy and computational cost. Our proposed model outperforms competing methods on three large-scale video benchmarks. In particular, on ActivityNet1.3 and mini-kinetics, we outperform the state-of-the-art efficient video recognition methods with 1.3× and 2.1 less GFLOPs, respectively. Addition-ally, our method sets× a new state of the art for efficient video understanding on the HVU benchmark.
Amir Ghodrati, Babak Ehteshami Bejnordi, AmirHossein Habibian
CVPR1
2020 ActionBytes: Learning From Trimmed Videos to Localize Actions
abstract
This paper tackles the problem of localizing actions in long untrimmed videos. Different from existing works, which all use annotated untrimmed videos during training, we learn only from short trimmed videos. This enables learning from large-scale datasets originally designed for action classification. We propose a method to train an action localization network that segments a video into interpretable fragments, we call ActionBytes. Our method jointly learns to cluster ActionBytes and trains the localization network using the cluster assignments as pseudo-labels. By doing so, we train on short trimmed videos that become untrimmed for ActionBytes. In isolation, or when merged, the ActionBytes also serve as effective action proposals. Experiments demonstrate that our boundary-guided training generalizes to unknown action classes and localizes actions in long videos of Thumos14, MultiThumos, and ActivityNet1.2. Furthermore, we show the advantage of ActionBytes for zero-shot localization as well as traditional weakly supervised localization, that train on long videos, to achieve state-of-the-art results.
Mihir Jain, Amir Ghodrati, Cees Snoek
CVPR2
2018 Video Time: Properties, Encoders and Evaluation
Amir Ghodrati, Efstratios Gavves, Cees Snoek
BMVC1
2018 Actor and Action Video Segmentation From a Sentence
abstract
This paper strives for pixel-level segmentation of actors and their actions in video content. Different from existing works, which all learn to segment from a fixed vocabulary of actor and action pairs, we infer the segmentation from a natural language input sentence. This allows to distinguish between fine-grained actors in the same super-category, identify actor and action instances, and segment pairs that are outside of the actor and action vocabulary. We propose a fully-convolutional model for pixel-level actor and action segmentation using an encoder-decoder architecture optimized for video. To show the potential of actor and action video segmentation from a sentence, we extend two popular actor and action datasets with more than 7,500 natural language descriptions. Experiments demonstrate the quality of the sentence-guided segmentations, the generalization ability of our model, and its advantage for traditional actor and action segmentation compared to the state-of-the-art.
Kirill Gavrilyuk, Amir Ghodrati, Cees Snoek
CVPR2
2017 DeepProposals: Hunting Objects and Actions by Cascading Deep Convolutional Layers
abstract
In this paper, a new method for generating object and action proposals in images and videos is proposed. It builds on activations of different convolutional layers of a pretrained CNN, combining the localization accuracy of the early layers with the high informativeness (and hence recall) of the later layers. To this end, we build an inverse cascade that, going backward from the later to the earlier convolutional layers of the CNN, selects the most promising locations and refines them in a coarse-to-fine manner. The method is efficient, because (i) it re-uses the same features extracted for detection, (ii) it aggregates features using integral images, and (iii) it avoids a dense evaluation of the proposals thanks to the use of the inverse coarse-to-fine cascade. The method is also accurate. We show that DeepProposals outperform most of the previous object proposal and action proposal approaches and, when plugged into a CNN-based object detector, produce state-of-the-art detection performance.
Amir Ghodrati, Ali Diba, Marco Pedersoli, Tinne Tuytelaars, Luc Van Gool
Int. J. Comput. Vis.1
2017 Rank Pooling for Action Recognition
abstract
We propose a function-based temporal pooling method that captures the latent structure of the video sequence data - e.g., how frame-level features evolve over time in a video. We show how the parameters of a function that has been fit to the video data can serve as a robust new video representation. As a specific example, we learn a pooling function via ranking machines. By learning to rank the frame-level features of a video in chronological order, we obtain a new representation that captures the video-wide temporal dynamics of a video, suitable for action recognition. Other than ranking functions, we explore different parametric models that could also explain the temporal changes in videos. The proposed functional pooling methods, and rank pooling in particular, is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We evaluate our method on various benchmarks for generic action, fine-grained action and gesture recognition. Results show that rank pooling brings an absolute improvement of 7-10 average pooling baseline. At the same time, rank pooling is compatible with and complementary to several appearance and local motion based methods and features, such as improved trajectories and deep learning features.
Basura Fernando, Efstratios Gavves, José Oramas M., Amir Ghodrati, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Towards Automatic Image Editing: Learning to See another You
Xu Jia 0012, Amir Ghodrati, Marco Pedersoli, Tinne Tuytelaars
BMVC2
2016 Online Action Detection
Roeland De Geest, Efstratios Gavves, Amir Ghodrati, Cees Snoek, Tinne Tuytelaars
ECCV (5)3
2015 Modeling video evolution for action recognition
abstract
In this paper we present a method to capture video-wide temporal information for action recognition. We postulate that a function capable of ordering the frames of a video temporally (based on the appearance) captures well the evolution of the appearance within the video. We learn such ranking functions per video via a ranking machine and use the parameters of these as a new video representation. The proposed method is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We perform a large number of evaluations on datasets for generic action recognition (Hollywood2 and HMDB51), fine-grained actions (MPII- cooking activities) and gestures (Chalearn). Results show that the proposed method brings an absolute improvement of 7–10%, while being compatible with and complementary to further improvements in appearance and local motion based methods.
Basura Fernando, Efstratios Gavves, José Oramas M., Amir Ghodrati, Tinne Tuytelaars
CVPR4
2015 DeepProposal: Hunting Objects by Cascading Deep Convolutional Layers
abstract
In this paper we evaluate the quality of the activation layers of a convolutional neural network (CNN) for the generation of object proposals. We generate hypotheses in a sliding-window fashion over different activation layers and show that the final convolutional layers can find the object of interest with high recall but poor localization due to the coarseness of the feature maps. Instead, the first layers of the network can better localize the object of interest but with a reduced recall. Based on this observation we design a method for proposing object locations that is based on CNN features and that combines the best of both worlds. We build an inverse cascade that, going from the final to the initial convolutional layers of the CNN, selects the most promising object locations and refines their boxes in a coarse-to-fine manner. The method is efficient, because i) it uses the same features extracted for detection, ii) it aggregates features using integral images, and iii) it avoids a dense evaluation of the proposals due to the inverse coarse-to-fine cascade. The method is also accurate, it outperforms most of the previously proposed object proposals approaches and when plugged into a CNN-based detector produces state-of-the-art detection performance.
Amir Ghodrati, Ali Diba, Marco Pedersoli, Tinne Tuytelaars, Luc Van Gool
ICCV1
2015 Swap Retrieval: Retrieving Images of Cats When the Query Shows a Dog
abstract
Query-by-example remains popular in image retrieval because it can exploit contextual information encoded in the image, that is difficult to express in a traditional textual query. Textual queries, on the other hand, give more flexibility in that it's easy to reformulate and refine a text query based on initial results.
Amir Ghodrati, Xu Jia 0012, Marco Pedersoli, Tinne Tuytelaars
ICMR1
2014 Is 2D Information Enough For Viewpoint Estimation?
Amir Ghodrati, Marco Pedersoli, Tinne Tuytelaars
BMVC1
2014 Coupling video segmentation and action recognition
abstract
Recently a lot of progress has been made in the field of video segmentation. The question then arises whether and how these results can be exploited for this other video processing challenge, action recognition. In this paper we show that a good segmentation is actually very important for recognition. We propose and evaluate several ways to integrate and combine the two tasks: i) recognition using a standard, bottom-up segmentation, ii) using a top-down segmentation geared towards actions, iii) using a segmentation based on inter-video similarities (co-segmentation), and iv) tight integration of recognition and segmentation via iterative learning. Our results clearly show that, on the one hand, the two tasks are interdependent and therefore an iterative optimization of the two makes sense and gives better results. On the other hand, comparable results can also be obtained with two separate steps but mapping the feature-space with a non-linear kernel.
Amir Ghodrati, Marco Pedersoli, Tinne Tuytelaars
WACV1
2012 Human action categorization using discriminative local spatio-temporal feature weighting
abstract
New methods based on local spatio-temporal features have exhibited significant performance in action recognition. In these methods, feature selection plays an important role to achieve a superior performance. Actions are represented by local spatio-temporal features extracted from action videos. Ac tion representations are then classified by applying a classifier (such as k-nearest neighbor or SVM). In this paper, we have proposed two feature weighting methods to better discriminate similar actions. We have proposed a definition of feature discrimination power to be used in the feature selection process. Our proposed weighting schemes have greatly improved the final categorization accuracy on the well-known KTH and Weizmann datasets.
Amir Ghodrati, Shohreh Kasaei
Intell. Data Anal.1