AmirHossein Habibian

dblp:63/8420 · also Amirhossein Habibian · DBLP profile ↗
← Back
27ranked-venue papers
14as first author
13since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 12 since 2021Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author
YearPublicationVenuePosition
2025 Mobile Video Diffusion
abstract
Video diffusion models have achieved impressive realism and controllability but are limited by high computational demands, restricting their use on mobile devices. This paper introduces the first mobile-optimized video diffusion model. Starting from a spatio-temporal UNet from Stable Video Diffusion (SVD), we reduce memory and computational cost by reducing the frame resolution, incorporating multi-scale temporal representations, and introducing two novel pruning schema to reduce the number of channels and temporal blocks. Furthermore, we employ adversarial finetuning to reduce the denoising to a single step. Our model, coined as MobileVD, is 523x more efficient (1817.2 vs. 4.34 TFLOPs) with a slight quality drop (FVD 149 vs. 171), generating latents for a 14x512x256 px clip in 1.7 seconds on a Xiaomi-14 Pro. Our results are available at https://qualcomm-ai-research.github.io/mobile-video-diffusion/
Haitam Ben Yahia, Denis Korzhenkov, Ioannis Lelekas, Amir Ghodrati, AmirHossein Habibian
ICCV5
2025 Valid: Variable-Length Input Diffusion for Novel View Synthesis
abstract
Novel View Synthesis (NVS), which tries to produce a realistic image at the target view given source view images and their corresponding poses, is a fundamental problem in 3D Vision. As this task is heavily under-constrained, some recent work, like Zerol23 [18], tries to solve this problem with generative modeling, specifically using pre-trained diffusion models. Although this strategy generalizes well to new scenes, compared to neural radiance field-based methods, it offers low levels of flexibility. For example, it can only accept a single-view image as input, despite realistic applications often offering multiple input images. This is because the source-view images and corresponding poses are processed separately and injected into the model at different stages. Thus it is not trivial to generalize the model into multi-view source images, once they are available. To solve this issue, we try to process each pose image pair separately and then fuse them as a unified visual representation which will be injected into the model to guide image synthesis at the target-views. However, inconsistency and computation costs increase as the number of input source-view images increases. To solve these issues, the Multi-view Cross Former module is proposed which maps variable-length input data to fix-size output data. A two-stage training strategy is introduced to further improve the efficiency during training time. Qualitative and quantitative evaluation over multiple datasets demonstrates the effectiveness of the proposed method against previous approaches. The code will be released according to the acceptance.
Shijie Li 0006, Farhad G. Zanjani, Haitam Ben Yahia, Yuki Markus Asano, Juergen Gall, AmirHossein Habibian
WACV6
2024 Clockwork Diffusion: Efficient Generation With Model-Step Distillation
abstract
This work aims to improve the efficiency of text-to-image diffusion models. While diffusion models use computationallyexpensive UNet-based denoising operations in ev-ery generation step, we identify that not all operations are equally relevant for the final output quality. In par-ticular, we observe that UNet layers operating on high-res feature maps are relatively sensitive to small pertur-bations. In contrast, low-res feature maps influence the semantic layout of the final image and can often be per-turbed with no noticeable change in the output. Based on this observation, we propose Clockwork Diffusion, a method that periodically reuses computation from preceding denoising steps to approximate low-res feature maps at one or more subsequent steps. For multiple base-lines, and for both text-to-image generation and image editing, we demonstrate that Clockwork leads to compa-rable or improved perceptual scores with drastically re-duced computational complexity. As an example, for Sta-ble Diffusion vI.5 with 8 DPM++ steps we save 32% of FLOPs with negligible FID and CLIP change. We re-lease code at https://github.com/Qualcomm-AI-research/clockwork-diffusion
AmirHossein Habibian, Amir Ghodrati, Noor Fathima, Guillaume Sautière, Risheek Garrepalli, Fatih Porikli, Jens Petersen
CVPR1
2024 Object-Centric Diffusion for Efficient Video Editing
Kumara Kahatapitiya, Adil Karjauv, Davide Abati, Fatih Porikli, Yuki Markus Asano, AmirHossein Habibian
ECCV (57)6
2024 Skip-Attention: Improving Vision Transformers by Paying Less Attention
abstract
This work aims to improve the efficiency of vision transformers (ViTs). While ViTs use computationally expensive self-attention operations in every layer, we identify that these operations are highly correlated across layers -- a key redundancy that causes unnecessary computations. Based on this observation, we propose SkipAT a method to reuse self-attention computation from preceding layers to approximate attention at one or more subsequent layers. To ensure that reusing self-attention blocks across layers does not degrade the performance, we introduce a simple parametric function, which outperforms the baseline transformer's performance while running computationally faster. We show that SkipAT is agnostic to transformer architecture and is effective in image classification, semantic segmentation on ADE20K, image denoising on SIDD, and video denoising on DAVIS. We achieve improved throughput at the same-or-higher accuracy levels in all these tasks.
Shashanka Venkataramanan, Amir Ghodrati, Yuki Markus Asano, Fatih Porikli, AmirHossein Habibian
ICLR5
2023 ResQ: Residual Quantization for Video Perception
abstract
This paper accelerates video perception, such as segmentation and human pose estimation, by levering cross-frame redundancies. Unlike the existing approaches, which avoid redundant computations by warping the past features using optical-flow or by performing sparse convolutions on frame differences, we approach the problem from a different perspective: low-bit quantization. We observe that residuals, as the difference in network activations between two neighboring frames, exhibit properties that make them highly quantizable. Based on this observation, we propose a novel quantization scheme for video networks coined as Res idual Quantization. ResQ extends the standard, frame-by-frame, quantization scheme by incorporating temporal dependencies that lead to better performance in terms of accuracy vs. bit-width. Furthermore, we extend our model to dynamically adjust the bit-width proportionally to the amount of changes in the video. We showcase the superiority of our model, against the standard quantization and existing efficient video perception models, using various architectures on semantic segmentation, video object segmentation and human pose estimation benchmarks.
Davide Abati, Haitam Ben Yahia, Markus Nagel, AmirHossein Habibian
ICCV4
2022 Region-of-Interest Based Neural Video Compression
Yura Perugachi-Diaz, Guillaume Sautière, Davide Abati, Yang Yang 0010, AmirHossein Habibian, Taco Cohen
BMVC5
2022 SALISA: Saliency-Based Input Sampling for Efficient Video Object Detection
Babak Ehteshami Bejnordi, AmirHossein Habibian, Fatih Porikli, Amir Ghodrati
ECCV (10)2
2022 Delta Distillation for Efficient Video Processing
AmirHossein Habibian, Haitam Ben Yahia, Davide Abati, Efstratios Gavves, Fatih Porikli
ECCV (35)1
2021 Efficient Video Super Resolution by Gated Local Self Attention
Davide Abati, Amir Ghodrati, AmirHossein Habibian
BMVC3
2021 Conditional Model Selection for Efficient Video Understanding
Mihir Jain, Haitam Ben Yahia, Amir Ghodrati, AmirHossein Habibian, Fatih Porikli
BMVC4
2021 FrameExit: Conditional Early Exiting for Efficient Video Recognition
abstract
In this paper, we propose a conditional early exiting framework for efficient video recognition. While existing works focus on selecting a subset of salient frames to re-duce the computation costs, we propose to use a simple sampling strategy combined with conditional early exiting to enable efficient recognition. Our model automatically learns to process fewer frames for simpler videos and more frames for complex ones. To achieve this, we employ a cascade of gating modules to automatically determine the earliest point in processing where an inference is sufficiently reliable. We generate on-the-fly supervision signals to the gates to provide a dynamic trade-off between accuracy and computational cost. Our proposed model outperforms competing methods on three large-scale video benchmarks. In particular, on ActivityNet1.3 and mini-kinetics, we outperform the state-of-the-art efficient video recognition methods with 1.3× and 2.1 less GFLOPs, respectively. Addition-ally, our method sets× a new state of the art for efficient video understanding on the HVU benchmark.
Amir Ghodrati, Babak Ehteshami Bejnordi, AmirHossein Habibian
CVPR3
2021 Skip-Convolutions for Efficient Video Processing
abstract
We propose Skip-Convolutions to leverage the large amount of redundancies in video streams and save computations. Each video is represented as a series of changes across frames and network activations, denoted as residuals. We reformulate standard convolution to be efficiently computed on residual frames: each layer is coupled with a binary gate deciding whether a residual is important to the model prediction, e.g. foreground regions, or it can be safely skipped, e.g. background regions. These gates can either be implemented as an efficient network trained jointly with convolution kernels, or can simply skip the residuals based on their magnitude. Gating functions can also incorporate block-wise sparsity structures, as required for efficient implementation on hardware platforms. By replacing all convolutions with Skip-Convolutions in two state-of-the-art architectures, namely EfficientDet and HRNet, we reduce their computational cost consistently by a factor of 3 ∼ 4× for two different tasks, without any accuracy drop. Extensive comparisons with existing model compression, as well as image and video efficiency methods demonstrate that Skip-Convolutions set a new state-of-the-art by effectively exploiting the temporal redundancies in videos.
AmirHossein Habibian, Davide Abati, Taco Cohen, Babak Ehteshami Bejnordi
CVPR1
2019 Video Compression With Rate-Distortion Autoencoders
abstract
In this paper we present a deep generative model for lossy video compression. We employ a model that consists of a 3D autoencoder with a discrete latent space and an autoregressive prior used for entropy coding. Both autoencoder and prior are trained jointly to minimize a ratedistortion loss, which is closely related to the ELBO used in variational autoencoders. Despite its simplicity, we find that our method outperforms the state-of-the-art learned video compression networks based on motion compensation or interpolation. We systematically evaluate various design choices, such as the use offrame-based or spatio-temporal autoencoders, and the type of autoregressive prior. In addition, we present three extensions of the basic method that demonstrate the benefits over classical approaches to compression. First, we introduce semantic compression, where the model is trained to allocate more bits to objects of interest. Second, we study adaptive compression, where the model is adapted to a domain with limited variability, e.g. videos taken from an autonomous car, to achieve superior compression on that domain. Finally, we introduce multimodal compression, where we demonstrate the effectiveness of our model in joint compression of multiple modalities captured by non-standard imaging sensors, such as quad cameras. We believe that this opens up novel video compression applications, which have not been feasible with classical codecs.
AmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco Cohen
ICCV1
2017 Video2vec Embeddings Recognize Events When Examples Are Scarce
abstract
This paper aims for event recognition when video examples are scarce or even completely absent. The key in such a challenging setting is a semantic video representation. Rather than building the representation from individual attribute detectors and their annotations, we propose to learn the entire representation from freely available web videos and their descriptions using an embedding between video features and term vectors. In our proposed embedding, which we call Video2vec, the correlations between the words are utilized to learn a more effective representation by optimizing a joint objective balancing descriptiveness and predictability. We show how learning the Video2vec embedding using a multimodal predictability loss, including appearance, motion and audio features, results in a better predictable representation. We also propose an event specific variant of Video2vec to learn a more accurate representation for the words, which are indicative of the event, by introducing a term sensitive descriptiveness loss. Our experiments on three challenging collections of web videos from the NIST TRECVID Multimedia Event Detection and Columbia Consumer Videos datasets demonstrate: i) the advantages of Video2vec over representations using attributes or alternative embeddings, ii) the benefit of fusing video modalities by an embedding over common strategies, iii) the complementarity of term sensitive descriptiveness and multimodal predictability for event recognition. By its ability to improve predictability of present day audio-visual video features, while at the same time maximizing their semantic descriptiveness, Video2vec leads to state-of-the-art accuracy for both few- and zero-example recognition of events in video.
AmirHossein Habibian, Thomas Mensink, Cees Snoek
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Discovering Semantic Vocabularies for Cross-Media Retrieval
abstract
This paper proposes a data-driven approach for cross-media retrieval by automatically learning its underlying semantic vocabulary. Different from the existing semantic vocabularies, which are manually pre-defined and annotated, we automatically discover the vocabulary concepts and their annotations from multimedia collections. To this end, we apply a probabilistic topic model on the text available in the collection to extract its semantic structure. Moreover, we propose a learning to rank framework, to effectively learn the concept classifiers from the extracted annotations. We evaluate the discovered semantic vocabulary for cross-media retrieval on three datasets of image/text and video/text pairs. Our experiments demonstrate that the discovered vocabulary does not require any manual labeling to outperform three recent alternatives for cross-media retrieval.
AmirHossein Habibian, Thomas Mensink, Cees Snoek
ICMR1
2015 Encoding Concept Prototypes for Video Event Detection and Summarization
abstract
This paper proposes a new semantic video representation for few and zero example event detection and unsupervised video event summarization. Different from existing works, which obtain a semantic representation by training concepts over images or entire video clips, we propose an algorithm that learns a set of relevant frames as the concept prototypes from web video examples, without the need for frame-level annotations, and use them for representing an event video. We formulate the problem of learning the concept prototypes as seeking the frames closest to the densest region in the feature space of video frames from both positive and negative training videos of a target concept. We study the behavior of our video event representation based on concept prototypes by performing three experiments on challenging web videos from the TRECVID 2013 multimedia event detection task and the MED-summaries dataset. Our experiments establish that i) Event detection accuracy increases when mapping each video into concept prototype space. ii) Zero-example event detection increases by analyzing each frame of a video individually in concept prototype space, rather than considering the holistic videos. iii) Unsupervised video event summarization using concept prototypes is more accurate than using video-level concept detectors.
Masoud Mazloom, AmirHossein Habibian, Dong Liu 0001, Cees Snoek, Shih-Fu Chang
ICMR2
2014 Composite Concept Discovery for Zero-Shot Video Event Detection
abstract
We consider automated detection of events in video without the use of any visual training examples. A common approach is to represent videos as classification scores obtained from a vocabulary of pre-trained concept classifiers. Where others construct the vocabulary by training individual concept classifiers, we propose to train classifiers for combination of concepts composed by Boolean logic operators. We call these concept combinations composite concepts and contribute an algorithm that automatically discovers them from existing video-level concept annotations. We discover composite concepts by jointly optimizing the accuracy of concept classifiers and their effectiveness for detecting events. We demonstrate that by combining concepts into composite concepts, we can train more accurate classifiers for the concept vocabulary, which leads to improved zero-shot event detection. Moreover, we demonstrate that by using different logic operators, namely "AND", "OR", we discover different types of composite concepts, which are complementary for zero-shot event detection. We perform a search for 20 events in 41K web videos from two test sets of the challenging TRECVID Multimedia Event Detection 2013 corpus. The experiments demonstrate the superior performance of the discovered composite concepts, compared to present-day alternatives, for zero-shot event detection.
AmirHossein Habibian, Thomas Mensink, Cees Snoek
ICMR1
2014 On-the-Fly Video Event Search by Semantic Signatures
abstract
In this technical demonstration, we showcase an event search engine that facilities instant access to an archive of web video. Different from many search engines which rely on high dimensional low-level visual features to represent videos, we rely on our proposed semantic signature. We extract semantic signature as the detection scores obtained by applying a vocabulary of 1,346 concept detectors on videos. The semantic signatures are compact, semantic and effective, as we will demonstrate for on-the-fly event retrieval using only a few positive examples. In addition, we will show how the signatures provide a crude interpretation on why a certain video has been retrieved.
AmirHossein Habibian, Masoud Mazloom, Cees Snoek
ICMR1
2014 Stop-Frame Removal Improves Web Video Classification
abstract
Web videos available in sharing sites like YouTube, are becoming an alternative to manually annotated training data, which are necessary for creating video classifiers. However, when looking into web videos, we observe they contain several irrelevant frames that may randomly appear in any video, i.e., blank and over exposed frames. We call these irrelevant frames stop-frames and propose a simple algorithm to identify and exclude them during classifier training. Stop-frames might appear in any video, so it is hard to recognize their category. Therefore we identify stop-frames as those frames, which are commonly misclassified by any concept classifier. Our experiments demonstrates that using our algorithm improves classification accuracy by 60% and 24% in terms of mean average precision for an event and concept detection benchmark.
AmirHossein Habibian, Cees Snoek
ICMR1
2014 VideoStory: A New Multimedia Embedding for Few-Example Recognition and Translation of Events
abstract
This paper proposes a new video representation for few-example event recognition and translation. Different from existing representations, which rely on either low-level features, or pre-specified attributes, we propose to learn an embedding from videos and their descriptions. In our embedding, which we call VideoStory, correlated term labels are combined if their combination improves the video classifier prediction. Our proposed algorithm prevents the combination of correlated terms which are visually dissimilar by optimizing a joint-objective balancing descriptiveness and predictability. The algorithm learns from textual descriptions of video content, which we obtain for free from the web by a simple spidering procedure. We use our VideoStory representation for few-example recognition of events on more than 65K challenging web videos from the NIST TRECVID event detection task and the Columbia Consumer Video collection. Our experiments establish that i) VideoStory outperforms an embedding without joint-objective and alternatives without any embedding, ii) The varying quality of input video descriptions from the web is compensated by harvesting more data, iii) VideoStory sets a new state-of-the-art for few-example event recognition, outperforming very recent attribute and low-level motion encodings. What is more, VideoStory translates a previously unseen video to its most likely description from visual content only.
AmirHossein Habibian, Thomas Mensink, Cees Snoek
ACM Multimedia1
2014 Recommendations for recognizing video events by concept vocabularies
AmirHossein Habibian, Cees Snoek
Comput. Vis. Image Underst.1
2014 Evaluating multimedia features and fusion for example-based event detection
abstract
Multimedia event detection (MED) is a challenging problem because of the heterogeneous content and variable quality found in large collections of Internet videos. To study the value of multimedia features and fusion for representing and learning events from a set of example video clips, we created SESAME, a system for video SEarch with Speed and Accuracy for Multimedia Events. SESAME includes multiple bag-of-words event classifiers based on single data types: low-level visual, motion, and audio features; high-level semantic visual concepts; and automatic speech recognition. Event detection performance was evaluated for each event classifier. The performance of low-level visual and motion features was improved by the use of difference coding. The accuracy of the visual concepts was nearly as strong as that of the low-level visual features. Experiments with a number of fusion methods for combining the event detection scores from these classifiers revealed that simple fusion methods, such as arithmetic mean, perform as well as or better than other, more complex fusion methods. SESAME’s performance in the 2012 TRECVID MED evaluation was one of the best reported.
Gregory K. Myers, Ramesh Nallapati, Julien van Hout, Stephanie Pancoast, Ramakant Nevatia, Chen Sun 0002, AmirHossein Habibian, Dennis C. Koelma, Koen E. A. van de Sande, Arnold W. M. Smeulders, Cees Snoek
Mach. Vis. Appl.7
2013 Recommendations for video event recognition using concept vocabularies
abstract
Representing videos using vocabularies composed of concept detectors appears promising for event recognition. While many have recently shown the benefits of concept vocabularies for recognition, the important question what concepts to include in the vocabulary is ignored. In this paper, we study how to create an effective vocabulary for arbitrary-event recognition in web video. We consider four research questions related to the number, the type, the specificity and the quality of the detectors in concept vocabularies. A rigorous experimental protocol using a pool of 1,346 concept detectors trained on publicly available annotations, a dataset containing 13,274 web videos from the Multimedia Event Detection benchmark, 25 event groundtruth definitions, and a state-of-the-art event recognition pipeline allow us to analyze the performance of various concept vocabulary definitions. From the analysis we arrive at the recommendation that for effective event recognition the concept vocabulary should i) contain more than 200 concepts, ii) be diverse by covering object, action, scene, people, animal and attribute concepts,iii) include both general and specific concepts, and iv) increase the number of concepts rather than improve the quality of the individual detectors. We consider the recommendations for video event recognition using concept vocabularies the most important contribution of the paper, as they provide guidelines for future work.
AmirHossein Habibian, Koen E. A. van de Sande, Cees Snoek
ICMR1
2013 Video2Sentence and vice versa
abstract
In this technical demonstration, we showcase a multimedia search engine that retrieves a video from a sentence, or a sentence from a video. The key novelty is our machine translation capability that exploits a cross-media representation for both the visual and textual modality using concept vocabularies. We will demonstrate the translations using arbitrary web videos and sentences related to everyday events. What is more, we will provide an automatically generated explanation, in terms of concept detectors, on why a particular video or sentence has been retrieved as the most likely translation.
AmirHossein Habibian, Cees Snoek
ACM Multimedia1
2013 Querying for video events by semantic signatures from few examples
abstract
We aim to query web video for complex events using only a handful of video query examples, where the standard approach learns a ranker from hundreds of examples. We consider a semantic signature representation, consisting of off-the-shelf concept detectors, to capture the variance in semantic appearance of events. Since it is unknown what similarity metric and query fusion to use in such an event retrieval setting, we perform three experiments on unconstrained web videos from the TRECVID event detection task. It reveals that: retrieval with semantic signatures using normalized correlation as similarity metric outperforms a low-level bag-of-words alternative, multiple queries are best combined using late fusion with an average operator, and event retrieval is preferred over event classification when less than eight positive video examples are available.
Masoud Mazloom, AmirHossein Habibian, Cees Snoek
ACM Multimedia2
2011 Extracting salient lines by Visual Attention for omnidirectional image classification
abstract
Representing an image as a set of its key and interesting lines facilitates the image understanding and classification. In this paper, we propose a method to extract the significant and interesting lines of the scene, which probably are useful in image classification. The proposed method is inspired from the Visual Attention, which is a perceptual mechanism in human and other primates that direct their perceptions to the limited regions of the scene. The attended regions are usually valuable in performing the task. Since the approach of using the lines to classify the images is particularly useful for omnidirectional images, we specialize our method to deal with these kinds of images. In the experiments, we demonstrate how our proposed methods improve the image classification performance with processing only small parts of the input images.
AmirHossein Habibian, Majid Nili Ahmadabadi, Babak Nadjar Araabi
CIMSIVP1