VLDB 2026 Research / reviewers in the wild / expert
Hamed Rezazadegan Tavakoli
dblp:76/8066 · also Hamed R. Tavakoli
· DBLP profile ↗
46ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0002-9466-9148ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 9 first-author · 21 since 2021Artificial intelligence and machine learning · 16 · 6 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | P$^{2}$U: Progressive Precision Update for Efficient Model DistributionabstractEfficient model distribution is becoming increasingly critical in bandwidth-constrained environments. In this paper, we propose a simple yet effective approach called Progressive Precision Update (P2U) to address this problem. Instead of directly transmitting the original high-precision model, P2U transmits a lower-bit precision model, coupled with a model update representing the difference between the original high-precision model and the transmitted low precision version. With extensive experiments on various model architectures, ranging from small models ($1 - 6$million parameters) to a large model (more than 100 million parameters) and using three different data sets, e.g., chest X-Ray, PASCAL-VOC, and CIFAR-100, we demonstrate that P2U consistently achieves a better tradeoff between accuracy, bandwidth usage, and startup latency, i.e., the time it takes for the receiver to start inference. Moreover, we show that when bandwidth or startup time is the priority, aggressive quantization (e.g., 4-bit) can be used without severely compromising performance. These results establish P2U as a practical solution for scalable and efficient model distribution across distributed environments such as federated learning, edge computing, and IoT deployments. Given that P2U complements existing compression techniques and can be implemented alongside many compression methods, e.g., sparsification, quantization, pruning, etc., the potential for improvement is even greater. Homayun Afrabandpey, Hamed Rezazadegan Tavakoli |
IEEE Trans. Multim. | 2 |
| 2025 | Evaluating the Emerging MPEG Video Coding for Machines in Semantic SegmentationabstractEmerging MPEG Video Coding for Machines (MPEG VCM) standardization activities address the growing demand for machine-to-machine visual applications, including video surveillance, autonomous driving, etc. This paper proposes an evaluation methodology tailored to MPEG VCM, with an emphasis on semantic segmentation tasks using the Pandaset dataset. This is a challenging target as standardization works must follow several limitations, such as dataset's licensing, fixed tools and software, and compliance with existing common test conditions (CTC) for the standard's development. The proposed evaluation methodology includes a step to align Pandaset and COCO labels. Two task networks, Detectron2 and Mask2Former, are used to evaluate the Rate-Performance behavior for semantic segmentation under various coding configurations. The performance of MPEG VCM is benchmarked against traditional codecs (VVC and HEVC), with a detailed analysis of MPEG VCM's coding tools. The extensive evaluations reveal interesting observations. (1) Although VCM achieved reasonably good segmentation performance, some of its developed tools, such as temporal resampling and region-of-interest coding, were not well suited for segmentation task. (2) The Hybrid NNVVC Inner Codec outperformed the VVC Inner Codec. (3) VCM's performance varies significantly for segmented classes. (4) Despite significant differences in human vision performance, VVC and HEVC exhibit relatively similar performance in machine vision. The main contributions of this work are to (1) enable evaluation of MPEG VCM in a real-world semantic segmentation use case, which is one of VCM's targeted tasks, and (2) to provide a detailed assessment of VCM's performance in semantic segmentation. Khoa Dang Pham, Farhad Pakdaman, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Nam Le 0003, Jukka I. Ahonen, Moncef Gabbouj |
ISM | 4 |
| 2025 | On Progressive Compressed Neural Model StorageabstractStoring multiple profiles of a neural network model at different bit-precisions is costly. Compressing and coding models seeks to addresses this problem. The most common compressed storage approach consists of independent compression of models at particular precision points. This results in redundant data storage which could be costly at scale. To further reduce the storage requirement and subsequently the transmission bandwidth requirements, we exploit the redundancy between different bit precisions and propose storing the models in terms of incremental bit precisions. We study a scenario in which a model is encoded in terms of a$b$-bit base model and the update to a high precision$n$-bit model. We empirically demonstrate that storing compressed models in terms of incremental bit-precisions allows better overall compression than independent coding with less model degradation. Hamed Rezazadegan Tavakoli, Homayun Afrabadpey |
ISM | 1 |
| 2024 | Efficient NeRF Optimization - Not All Samples Remain Equally Hard
Juuso Korhonen, Rangu Goutham, Hamed Rezazadegan Tavakoli, Juho Kannala |
ECCV (36) | 3 |
| 2024 | EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement LearningabstractFrom a visual-perception perspective, modern graphical user interfaces (GUIs) comprise a complex graphics-rich two-dimensional visuospatial arrangement of text, images, and interactive objects such as buttons and menus. While existing models can accurately predict regions and objects that are likely to attract attention “on average”, no scanpath model has been capable of predicting scanpaths for an individual. To close this gap, we introduce EyeFormer, which utilizes a Transformer architecture as a policy network to guide a deep reinforcement learning algorithm that predicts gaze locations. Our model offers the unique capability of producing personalized predictions when given a few user scanpath samples. It can predict full scanpath information, including fixation positions and durations, across individuals and various stimulus types. Additionally, we demonstrate applications in GUI layout optimization driven by our model. Yue Jiang 0002, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva, Antti Oulasvirta |
UIST | 3 |
| 2023 | Momentum Adapt: Robust Unsupervised Adaptation for Improving Temporal Consistency in Video Semantic Segmentation During Test-Time
Amirhossein Hassankhani, Hamed Rezazadegan Tavakoli, Esa Rahtu |
BMVC | 2 |
| 2023 | UEyes: Understanding Visual Saliency across User Interface TypesabstractWhile user interfaces (UIs) display elements such as images and text in a grid-based layout, UI types differ significantly in the number of elements and how they are displayed. For example, webpage designs rely heavily on images and text, whereas desktop UIs tend to feature numerous small images. To examine how such differences affect the way users look at UIs, we collected and analyzed a large eye-tracking-based dataset, UEyes (62 participants and 1,980 UI screenshots), covering four major UI types: webpage, desktop UI, mobile UI, and poster. We analyze its differences in biases related to such factors as color, location, and gaze direction. We also compare state-of-the-art predictive models and propose improvements for better capturing typical tendencies across UI types. Both the dataset and the models are publicly available. Yue Jiang 0002, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel, Julia Kylmälä, Antti Oulasvirta |
CHI | 3 |
| 2023 | NN-VVC: Versatile Video Coding boosted by self-supervisedly learned image coding for machinesabstractThe recent progress in artificial intelligence has led to an ever-increasing usage of images and videos by machine analysis algorithms, mainly neural networks. Nonetheless, compression, storage and transmission of media have traditionally been designed considering human beings as the viewers of the content. Recent research on image and video coding for machine analysis has progressed mainly in two almost orthogonal directions. The first is represented by end-to-end (E2E) learned codecs which, while offering high performance on image coding, are not yet on par with state-of-the-art conventional video codecs and lack interoperability. The second direction considers using the Versatile Video Coding (VVC) standard or any other conventional video codec (CVC) together with pre- and post-processing operations targeting machine analysis. While the CVC-based methods benefit from interoperability and broad hardware and software support, the machine task performance is often lower than the desired level, particularly in low bitrates. This paper proposes a hybrid codec for machines called NN-VVC, which combines the advantages of an E2E-learned image codec and a CVC to achieve high performance in both image and video coding for machines. Our experiments show that the proposed system achieved up to -43.20% and -26.8% Bjøntegaard Delta rate reduction over VVC for image and video data, respectively, when evaluated on multiple different datasets and machine vision tasks. To the best of our knowledge, this is the first research paper showing a hybrid video codec that outperforms VVC on multiple datasets and multiple machine vision tasks. Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Antti Hallapuro, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 6 |
| 2022 | Bridging the Gap Between Image Coding for Machines and HumansabstractImage coding for machines (ICM) aims at reducing the bitrate required to represent an image while minimizing the drop in machine vision analysis accuracy. In many use cases, such as surveillance, it is also important that the visual quality is not drastically deteriorated by the compression process. Recent works on using neural network (NN) based ICM codecs have shown significant coding gains against traditional methods; however, the decompressed images, especially at low bitrates, often contain checkerboard artifacts. We propose an effective decoder finetuning scheme based on adversarial training to significantly enhance the visual quality of ICM codecs, while preserving the machine analysis accuracy, without adding extra bitcost or parameters at the inference phase. The results show complete removal of the checkerboard artifacts at the negligible cost of −1.6% relative change in task performance score. In the cases where some amount of artifacts is tolerable, such as when machine consumption is the primary target, this technique can enhance both pixel-fidelity and feature-fidelity scores without losing task performance. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ICIP | 5 |
| 2022 | Stochastic Binary-Ternary Quantization for Communication Efficient Federated ComputationabstractA stochastic binary-ternary (SBT) quantization approach is introduced for communication efficient federated computation; form of collaborative computing where locally trained models are exchanged between institutes. Communication of deep neural network models could be highly inefficient due to their large size. This motivates model compression in which quantization is an important step. Two well-known quantization algorithms are binary and ternary quantization. The first leads into good compression, sacrificing accuracy. The second provides good accuracy with less compression. To better benefit from trade-off between accuracy and compression, we propose an algorithm to stochastically switch between binary and ternary quantization. By combining with uniform quantization, we further extend the proposed algorithm to a hierarchical method which results in even better compression without sacrificing the accuracy. We tested the proposed algorithm using Neural network Compression Test Model (NCTM) provided by MPEG community. Our results demonstrate that the hierarchical variant of the proposed algorithm outperforms other quantization algorithms in term of compression, while maintaining the accuracy competitive to that provided by other methods. Rangu Goutham, Homayun Afrabandpey, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Miska M. Hannuksela, Hamed Rezazadegan Tavakoli |
ICIP | 7 |
| 2022 | Adaptive Multi-Scale Progressive Probability Model for Lossless Image CompressionabstractDomain adaptation is an efficient technique to improve the performance of a system by adapting a pre-trained model to the given input data. The adaptation technique has been generally applied in conventional video codecs. For neural network-based systems, the encoder may adapt the decoder to the input data by fine-tuning a pre-trained model present at the decoder side. The weight update is then transferred to the decoder and the updated model is used to decode the bitstream. However, due to the large number of parameters in deep neural networks, the overhead of the weight update may diminish the gain from the adaptation technique. In recent years, various methods have been proposed to reduce the overhead without significantly compromising the gain. In this paper, we propose an adaptive multi-scale progressive probability model for lossless image compression. The proposed method uses the data that has already been processed at the inference stage to fine-tune the probability model. Importantly, the decoder can apply the fine-tuning by itself resulting a small adaptation overhead to help the decoder in performing the fine-tuning. The proposed method achieves up to 0.28 bits-per-pixel (BPP) reduction on four benchmark datasets compared to the state-of-the-art method. Honglei Zhang 0001, Francesco Cricri, Nannan Zou, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela |
ICIP | 4 |
| 2022 | TMD: Transformed Mesh Decoder for Mesh AnimationabstractEasy and fast animation of 3D characters is attractive for both gaming and entertainment applications. 3D mesh must be properly rigged and skinned in order to create seamless animation. This process can be time consuming and requires deep knowledge of appropriate software based on kinematic animation. In this work, we present a fast and lightweight deep neural model to automate 3D human animation using skeletal representation from 2D image pose, i.e., joints in 2D space. We accomplish this using Transformed Mesh Decoder (TMD), which is a novel layer for convolutional neural networks. To train the network, we generate a large and diverse dataset using Skinned Multi-Person Linear (SMPL) model. Experiment shows that our method is effective when compared to both the ground truth and state-of-the art linear blend skinning that require manually painted skinning weights for accurate result. The animation process is fast and can achieve approximately 10-15fps in practice. The proposed method is simple which opens the possibility for future improvement in real-time application. Peter Fasogbon, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu |
ICPR | 4 |
| 2022 | The Lottery Ticket Adaptation for Neural Video CodingabstractRecently, learning based video compression methods have attracted increasing attention. However, most learning based video codecs are not adaptive to different video contents. Though adaptation at inference time is a solution to tackle this issue, adapting all the codec’s parameters is computationally expensive and brings heavy bitrate overhead. The recently proposed Lottery Ticket Hypothesis (LTH) states that an over-parameterized neural network contains smaller subnetworks (winning tickets) that can match the performance of the original network. In this paper, we present a novel lottery-ticket adaptation technique on decoder-side multiplicative parameters of a neural network, transferring the concept of winning lottery tickets to video compression tasks. At inference time, the winning multiplicative parameters are overfitted, compressed, and signaled together with encoded frames for decoding. We show that our approach outperforms the Versatile Video Coding (VVC) standard in the Multiscale Structural Similarity (MS-SSIM) at a low bitrate on both the UVG and JVET sequences. To the best of our knowledge, this is the first attempt to apply LTH in the video compression domain. Also, this is the first published end-to-end learned video codec working directly on YUV format, which outperforms VVC on UVG and JVET datasets in MS-SSIM. Nannan Zou, Francesco Cricri, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 4 |
| 2022 | On the Importance of Temporal Dependencies of Weight Updates in Communication Efficient Federated LearningabstractThis paper studies the effect of exploiting temporal dependency of successive weight updates on compressing communications in Federated Learning (FL). For this, we propose residual coding for FL, which utilizes temporal dependencies by communicating compressed residuals of the weight updates whenever they are beneficial to bandwidth. We further consider Temporal Context Adaptation (TCA) which compares co-located elements of consecutive weight updates to select optimal setting for compression of bitstream in DeepCABAC encoder. Following experimental settings of MPEG standard on Neural Network Compression (NNC), we demonstrate that both temporal dependency based technologies reduce communication overhead, where the maximum reduction is obtained using both technologies, simultaneously. Homayun Afrabandpey, Rangu Goutham, Honglei Zhang 0001, Francesco Criri, Emre Aksu, Hamed Rezazadegan Tavakoli |
VCIP | 6 |
| 2022 | A compact deep architecture for real-time saliency prediction
Samad Zabihi, Hamed Rezazadegan Tavakoli, Ali Borji, Eghbal G. Mansoori |
Signal Process. Image Commun. | 2 |
| 2022 | Overview of the Neural Network Compression and Representation (NNR) StandardabstractNeural Network Coding and Representation (NNR) is the first international standard for efficient compression of neural networks (NNs). The standard is designed as a toolbox of compression methods, which can be used to create coding pipelines. It can be either used as an independent coding framework (with its own bitstream format) or together with external neural network formats and frameworks. For providing the highest degree of flexibility, the network compression methods operate per parameter tensor in order to always ensure proper decoding, even if no structure information is provided. The NNR standard contains compression-efficient quantization and deep context-adaptive binary arithmetic coding (DeepCABAC) as core encoding and decoding technologies, as well as neural network parameter pre-processing methods like sparsification, pruning, low-rank decomposition, unification, local scaling and batch norm folding. NNR achieves a compression efficiency of more than 97% for transparent coding cases, i.e. without degrading classification quality, such as top-1 or top-5 accuracies. This paper provides an overview of the technical features and characteristics of NNR. Heiner Kirchhoffer, Paul Haase, Wojciech Samek, Karsten Müller 0001, Hamed Rezazadegan Tavakoli, Francesco Cricri, Emre Aksu, Miska M. Hannuksela, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001, Swayambhoo Jain, Shahab Hamidi-Rad, Fabien Racapé, Werner Bailer |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Mind The Structure: Adopting Structural Information For Deep Neural Network CompressionabstractDeep neural networks have huge number of parameters and require large number of bits for representation. This hinders their adoption in decentralized environments where model transfer among different parties is a characteristic of the environment while the communication bandwidth is limited. Parameter quantization is a compression approach to address this challenge by reducing the number of bits required to represent a model, e.g. a neural network. However, majority of existing neural network quantization methods do not exploit structural information of layers and parameters during quantization. In this paper, focusing on Convolutional Neural Networks (CNNs), we present a novel quantization approach by employing the structural information of neural network layers and their corresponding parameters. Starting from a pre-trained CNN, we categorize network parameters into different groups based on the similarity of their layers and their spatial structure. Parameters of each group are independently clustered and the centroid of each cluster is used as representative for all parameters in the cluster. Finally, the centroids and the cluster indexes of the parameters are used as a compact representation of the parameters. Experiments with two different tasks, i.e., acoustic scene classification and image compression, demonstrate the effectiveness of the proposed approach. Homayun Afrabandpey, Anton Muravev, Hamed Rezazadegan Tavakoli, Honglei Zhang 0001, Francesco Cricri, Moncef Gabbouj, Emre Aksu |
ICIP | 3 |
| 2021 | Hybrid Pruning And SparsificationabstractA hybrid approach based on the combination of saliency-based neural pruning and regularization-based sparsification is proposed. We propose using a graph diffusion process for determining the neuron importance for pruning. Then, we use a regularization loss based on weighted $L_{1}-$norm and $L_{2}-$norm during fine-tuning to recover the lost performance. This is followed by a threshold step to further impose sparsification. We demonstrate such a hybrid approach achieves significantly better performance in comparison to purely regularization-based sparsification for large neural networks. To this end, we assessed our proposed method on three tasks, including: image classification (3 network architectures), audio classification and image compression. Hamed Rezazadegan Tavakoli, Joachim Wabnig, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Iraj Saniee |
ICIP | 1 |
| 2021 | Learned Image Coding for Machines: A Content-Adaptive ApproachabstractToday, according to the Cisco Annual Internet Report (2018-2023), the fastest-growing category of Internet traffic is machine-to-machine communication. In particular, machine-to-machine communication of images and videos represents a new challenge and opens up new perspectives in the context of data compression. One possible solution approach consists of adapting current human-targeted image and video coding standards to the use case of machine consumption. Another approach consists of developing completely new compression paradigms and architectures for machine-to-machine communications. In this paper, we focus on image compression and present an inference-time content-adaptive fine-tuning scheme that optimizes the latent representation of an end-to-end learned image codec, aimed at improving the compression efficiency for machine-consumption. The conducted experiments targeting instance segmentation task network show that our online finetuning brings an average bitrate saving (BD-rate) of -3.66% with respect to our pretrained image codec. In particular, at low bitrate points, our proposed method results in a significant bitrate saving of -9.85%. Overall, our pretrained-and-then-finetuned system achieves - 30.54% BD-rate over the state-of-the-art image/video codec Versatile Video Coding (VVC) on instance segmentation. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Esa Rahtu |
ICME | 5 |
| 2021 | Learned Enhancement Filters for Image Coding for MachinesabstractMachine-To-Machine (M2M) communication applications and use cases, such as object detection and instance segmentation, are becoming mainstream nowadays. As a consequence, majority of multimedia content is likely to be consumed by machines in the coming years. This opens up new challenges on efficient compression of this type of data. Two main directions are being explored in the literature, one being based on existing traditional codecs, such as the Versatile Video Coding (VVC) standard, that are optimized for human-targeted use cases, and another based on end-to-end trained neural networks. However, traditional codecs have significant benefits in terms of interoperability, real-time decoding, and availability of hardware implementations over end-to-end learned codecs. Therefore, in this paper, we propose learned post-processing filters that are targeted for enhancing the performance of machine vision tasks for images reconstructed by the VVC codec. The proposed enhancement filters provide significant improvements on the target tasks compared to VVC coded images. The conducted experiments show that the proposed post-processing filters provide about 45% and 49% Bjøntegaard Delta Rate gains over VVC in instance segmentation and object detection tasks, respectively. Jukka I. Ahonen, Ramin Ghaznavi Youvalari, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 6 |
| 2021 | Enhancing Image Coding for Machines with Compressed Feature ResidualsabstractAs computer vision technologies have tremendously improved over the last decade, videos and images are often consumed by machines instead of humans which are the main target for traditional video codecs. In many use cases, although machines are the main consumers, human involvement is also required, or even mandatory. In this paper, we propose a novel image coding technique targeted for machines, while maintaining the capability for human consumption. Our proposed codec generates two bitstreams: one bitstream from a traditional codec, referred to as human bitstream, optimized for human consumption; the other bitstream, referred to as machine bitstream, generated from an end-to-end learned neural network-based codec and optimized for machine tasks. Instead of working on the image domain, the proposed machine bitstream is derived from feature residuals – the difference between the features extracted from the input image and the features extracted from the reconstructed image generated by the traditional codec. With the help of the machine bitstream, we can significantly improve machine task performance in the low bitrate range. Our system beats the state-of-the-art traditional codec, the Versatile Video Coding (VVC/H.266), achieving −40.5% in Bjontegaard delta bitrate reduction on average for bitrates up to 0.07 BPP. Joni Seppälä, Honglei Zhang 0001, Nam Le 0003, Ramin Ghaznavi Youvalari, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 6 |
| 2021 | Adaptation and Attention for Neural Video CodingabstractNeural image coding represents now the state-of-the-art image compression approach. However, a lot of work is still to be done in the video domain. In this work, we propose an end-to-end learned video codec that introduces several architectural novelties as well as training novelties, revolving around the concepts of adaptation and attention. Our codec is organized as an intra-frame codec paired with an inter-frame codec. As one architectural novelty, we propose to train the inter-frame codec model to adapt the motion estimation process based on the resolution of the input video. A second architectural novelty is a new neural block that combines concepts from split-attention based neural networks and from DenseNets. Finally, we propose to overfit a set of decoder-side multiplicative parameters at inference time. Through ablation studies and comparisons to prior art, we show the benefits of our proposed techniques in terms of coding gains. We compare our codec to VVC/H.266 and RLVC, which represent the state-of-the-art traditional and end-to-end learned codecs, respectively, and to the top performing end-to-end learned approach in 2021 CLIC competition, E2E_T_OL. Our codec clearly outperforms E2E_T_OL, and compare favorably to VVC and RLVC in some settings. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Jani Lainema, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 5 |
| 2021 | Learn to overfit better: finding the important parameters for learned image compressionabstractFor most machine learning systems, overfitting is an undesired behavior. However, overfitting a model to a test image or a video at inference time is a favorable and effective technique to improve the coding efficiency of learning-based image and video codecs. At the encoding stage, one or more neural networks that are part of the codec are finetuned using the input image or video to achieve a better coding performance. The encoder en-codes the input content into a content bitstream. If the finetuned neural network is part (also) of the decoder, the encoder signals the weight update of the finetuned model to the decoder along with the content bitstream. At the decoding stage, the decoder first updates its neural network model according to the received weight update, and then proceeds with decoding the content bitstream. Since a neural network contains a large number of parameters, compressing the weight update is critical to reducing bitrate overhead. In this paper, we propose learning-based methods to find the important parameters to be overfitted, in terms of rate-distortion performance. Based on simple distribution models for variables in the weight update, we derive two objective functions. By optimizing the proposed objective functions, the importance scores of the parameters can be calculated and the important parameters can be determined. Our experiments on lossless image compression codec show that the proposed method significantly outperforms a prior-art method where overfitted parameters were selected based on heuristics. Furthermore, our technique improved the compression performance of the state-of-the-art lossless image compression codec by 0.1 bit per pixel. Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, María Santamaría 0001, Yat-Hong Lam, Miska M. Hannuksela |
VCIP | 3 |
| 2021 | Deep saliency models : The quest for the loss function
Alexandre Bruckert, Hamed Rezazadegan Tavakoli, Zhi Liu 0003, Marc Christie, Olivier Le Meur |
Neurocomputing | 2 |
| 2020 | Image Captioning Through Image Transformer
Sen He 0001, Wentong Liao, Hamed Rezazadegan Tavakoli, Michael Ying Yang, Bodo Rosenhahn, Nicolas Pugeault |
ACCV (4) | 3 |
| 2020 | Lossless Image Compression Using a Multi-scale Progressive Statistical Model
Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Nannan Zou, Emre Aksu, Miska M. Hannuksela |
ACCV (3) | 3 |
| 2020 | Understanding Visual Saliency in Mobile User InterfacesabstractFor graphical user interface (UI) design, it is important to understand what attracts visual attention. While previous work on saliency has focused on desktop and web-based UIs, mobile app UIs differ from these in several respects. We present findings from a controlled study with 30 participants and 193 mobile UIs. The results speak to a role of expectations in guiding where users look at. Strong bias toward the top-left corner of the display, text, and images was evident, while bottom-up features such as color or size affected saliency less. Classic, parameter-free saliency models showed a weak fit with the data, and data-driven models improved significantly when trained specifically on this dataset (e.g., NSS rose from 0.66 to 0.84). We also release the first annotated dataset for investigating visual saliency in mobile UIs. Luis A. Leiva, Yunfei Xue, Avya Bansal, Hamed Rezazadegan Tavakoli, Tugçe Köroglu, Jingzhou Du, Niraj Ramesh Dayama, Antti Oulasvirta |
MobileHCI | 4 |
| 2020 | AI4TV 2020: 2nd International Workshop on AI for Smart TV Content Production, Access and DeliveryabstractTechnological developments in comprehensive video understanding - detecting and identifying visual elements of a scene, combined with audio understanding (music, speech), as well as aligned with textual information such as captions, subtitles, etc. and background knowledge - have been undergoing a significant revolution during recent years. The workshop brings together experts from academia and industry in order to discuss the latest progress in artificial intelligence research in topics related to multimodal information analysis, and in particular, semantic analysis of video, audio, and textual information for smart digital TV content production, access and delivery. Raphaël Troncy, Jorma Laaksonen, Hamed Rezazadegan Tavakoli, Lyndon J. B. Nixon, Vasileios Mezaris, Mohammad Hosseini 0002 |
ACM Multimedia | 3 |
| 2020 | L2C - Learning to Learn to CompressabstractIn this paper we present an end-to-end meta-learned system for image compression. Traditional machine learning based approaches to image compression train one or more neural network for generalization performance. However, at inference time, the encoder or the latent tensor output by the encoder can be optimized for each test image. This optimization can be regarded as a form of adaptation or benevolent overfitting to the input content. In order to reduce the gap between training and inference conditions, we propose a new training paradigm for learned image compression, which is based on meta-learning. In a first phase, the neural networks are trained normally. In a second phase, the Model-Agnostic Meta-learning approach is adapted to the specific case of image compression, where the inner-loop performs latent tensor overfitting, and the outer loop updates both encoder and decoder neural networks based on the overfitting performance. Furthermore, after meta-learning, we propose to overfit and cluster the bias terms of the decoder on training image patches, so that at inference time the optimal content-specific bias terms can be selected at encoder-side. Finally, we propose a new probability model for lossless compression, which combines concepts from both multi-scale and super-resolution probability model approaches. We show the benefits of all our proposed ideas via carefully designed experiments. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Jani Lainema, Miska M. Hannuksela, Emre Aksu, Esa Rahtu |
MMSP | 4 |
| 2020 | Geometric Image Correspondence Verification by Dense Pixel MatchingabstractThis paper addresses the problem of determining dense pixel correspondences between two images and its application to geometric correspondence verification in image retrieval. The main contribution is a geometric correspondence verification approach for re-ranking a shortlist of retrieved database images based on their dense pair-wise matching with the query image at a pixel level. We determine a set of cyclically consistent dense pixel matches between the pair of images and evaluate local similarity of matched pixels using neural network based image descriptors. Final re-ranking is based on a novel similarity function, which fuses the local similarity metric with a global similarity metric and a geometric consistency measure computed for the matched pixels. For dense matching our approach utilizes a modified version of a recently proposed dense geometric correspondence network (DGC-Net), which we also improve by optimizing the architecture. The proposed model and similarity metric compare favourably to the state-of-the-art image retrieval methods. In addition, we apply our method to the problem of long-term visual localization demonstrating promising results and generalization across datasets. Zakaria Laskar, Iaroslav Melekhov, Hamed Rezazadegan Tavakoli, Juha Ylioinas |
WACV | 3 |
| 2020 | A unified cycle-consistent neural model for text and image retrieval
Marcella Cornia, Lorenzo Baraldi 0001, Hamed Rezazadegan Tavakoli, Rita Cucchiara |
Multim. Tools Appl. | 3 |
| 2019 | Understanding and Visualizing Deep Visual Saliency ModelsabstractRecently, data-driven deep saliency models have achieved high performance and have outperformed classical saliency models, as demonstrated by results on datasets such as the MIT300 and SALICON. Yet, there remains a large gap between the performance of these models and the inter-human baseline. Some outstanding questions include what have these models learned, how and where they fail, and how they can be improved. This article attempts to answer these questions by analyzing the representations learned by individual neurons located at the intermediate layers of deep saliency models. To this end, we follow the steps of existing deep saliency models, that is borrowing a pre-trained model of object recognition to encode the visual features and learning a decoder to infer the saliency. We consider two cases when the encoder is used as a fixed feature extractor and when it is fine-tuned, and compare the inner representations of the network. To study how the learned representations depend on the task, we fine-tune the same network using the same image set but for two different tasks: saliency prediction versus scene classification. Our analyses reveal that: 1) some visual regions (e.g. head, text, symbol, vehicle) are already encoded within various layers of the network pre-trained for object recognition, 2) using modern datasets, we find that fine-tuning pre-trained models for saliency prediction makes them favor some categories (e.g. head) over some others (e.g. text), 3) although deep models of saliency outperform classical models on natural images, the converse is true for synthetic stimuli (e.g. pop-out search arrays), an evidence of significant difference between human and data-driven saliency models, and 4) we confirm that, after-fine tuning, the change in inner-representations is mostly due to the task and not the domain shift in the data. Sen He 0001, Hamed Rezazadegan Tavakoli, Ali Borji, Yang Mi, Nicolas Pugeault |
CVPR | 2 |
| 2019 | Human Attention in Image Captioning: Dataset and AnalysisabstractIn this work, we present a novel dataset consisting of eye movements and verbal descriptions recorded synchronously over images. Using this data, we study the differences in human attention during free-viewing and image captioning tasks. We look into the relationship between human atten- tion and language constructs during perception and sen- tence articulation. We also analyse attention deployment mechanisms in the top-down soft attention approach that is argued to mimic human attention in captioning tasks, and investigate whether visual saliency can help image caption- ing. Our study reveals that (1) human attention behaviour differs in free-viewing and image description tasks. Hu- mans tend to fixate on a greater variety of regions under the latter task, (2) there is a strong relationship between de- scribed objects and attended objects (97% of the described objects are being attended), (3) a convolutional neural net- work as feature encoder accounts for human-attended re- gions during image captioning to a great extent (around 78%), (4) soft-attention mechanism differs from human at- tention, both spatially and temporally, and there is low correlation between caption scores and attention consis- tency scores. These indicate a large gap between humans and machines in regards to top-down attention, and (5) by integrating the soft attention model with image saliency, we can significantly improve the model's performance on Flickr30k and MSCOCO benchmarks. The dataset can be found at: https://github.com/SenHe/ Human-Attention-in-Image-Captioning. Sen He 0001, Hamed Rezazadegan Tavakoli, Ali Borji, Nicolas Pugeault |
ICCV | 2 |
| 2019 | AWSD: Adaptive Weighted Spatiotemporal Distillation for Video RepresentationabstractWe propose an Adaptive Weighted Spatiotemporal Distillation (AWSD) technique for video representation by encoding the appearance and dynamics of the videos into a single RGB image map. This is obtained by adaptively dividing the videos into small segments and comparing two consecutive segments. This allows using pre-trained models on still images for video classification while successfully capturing the spatiotemporal variations in the videos. The adaptive segment selection enables effective encoding of the essential discriminative information of untrimmed videos. Based on Gaussian Scale Mixture, we compute the weights by extracting the mutual information between two consecutive segments. Unlike pooling-based methods, our AWSD gives more importance to the frames that characterize actions or events thanks to its adaptive segment length selection. We conducted extensive experimental analysis to evaluate the effectiveness of our proposed method and compared our results against those of recent state-of-the-art methods on four benchmark datatsets, including UCF101, HMDB51, ActivityNet v1.3, and Maryland. The obtained results on these benchmark datatsets showed that our method significantly outperforms earlier works and sets the new state-of-the-art performance in video classification. Code is available at the project webpage: https://mohammadt68.github.io/AWSD/. Mohammad Tavakolian, Hamed Rezazadegan Tavakoli, Abdenour Hadid |
ICCV | 2 |
| 2019 | AI4TV 2019: 1st International Workshop on AI for Smart TV Content Production, Access and DeliveryabstractTechnological developments in comprehensive video understanding - detecting and identifying visual elements of a scene, combined with audio understanding (music, speech), as well as aligned with textual information such as captions, subtitles, etc. and background knowledge - have been undergoing a significant revolution during recent years. The workshop brings together experts from academia and industry in order to discuss the latest progress in artificial intelligence research in topics related to multimodal information analysis, and in particular, semantic analysis of video, audio, and textual information for smart digital TV content production, access and delivery. Raphaël Troncy, Jorma Laaksonen, Hamed Rezazadegan Tavakoli, Lyndon J. B. Nixon, Vasileios Mezaris |
ACM Multimedia | 3 |
| 2019 | Semantic Matching by Weakly Supervised 2D Point Set RegistrationabstractIn this paper we address the problem of establishing correspondences between different instances of the same object. The problem is posed as finding the geometric transformation that aligns a given image pair. We use a convolutional neural network (CNN) to directly regress the parameters of the transformation model. The alignment problem is defined in the setting where an unordered set of semantic key-points per image are available, but, without the correspondence information. To this end we propose a novel loss function based on cyclic consistency that solves this 2D point set registration problem by inferring the optimal geometric transformation model parameters. We train and test our approach on a standard benchmark dataset Proposal-Flow (PF-PASCAL). The proposed approach achieves state-of-the-art results demonstrating the effectiveness of the method. In addition, we show our approach further benefits from additional training samples in PF-PASCAL generated by using category level information. Zakaria Laskar, Hamed Rezazadegan Tavakoli, Juho Kannala |
WACV | 2 |
| 2019 | Digging Deeper Into Egocentric Gaze PredictionabstractThis paper digs deeper into factors that influence egocentric gaze. Instead of training deep models for this purpose in a blind manner, we propose to inspect factors that contribute to gaze guidance during daily tasks. Bottom-up saliency and optical flow are assessed versus strong spatial prior baselines. Task-specific cues such as vanishing point, manipulation point, and hand regions are analyzed as representatives of top-down information. We also look into the contribution of these factors by investigating a simple recurrent neural model for ego-centric gaze prediction. First, deep features are extracted for all input video frames. Then, a gated recurrent unit is employed to integrate information over time and to predict the next fixation. We also propose an integrated model that combines the recurrent model with several top-down and bottom-up cues. Extensive experiments over multiple datasets reveal that (1) spatial biases are strong in egocentric videos, (2) bottom-up saliency models perform poorly in predicting gaze and underperform spatial biases, (3) deep features perform better compared to traditional features, (4) as opposed to hand regions, the manipulation point is a strong influential cue for gaze prediction, (5) combining the proposed recurrent model with bottom-up cues, vanishing points and, in particular, manipulation point results in the best gaze prediction accuracy over egocentric videos, (6) the knowledge transfer works best for cases where the tasks or sequences are similar, and (7) task and activity recognition can benefit from gaze prediction. Our findings suggest that (1) there should be more emphasis on hand-object interaction and (2) the egocentric vision community should consider larger datasets including diverse stimuli and more subjects. Hamed Rezazadegan Tavakoli, Esa Rahtu, Juho Kannala, Ali Borji |
WACV | 1 |
| 2018 | Bottom-Up Attention Guidance for Recurrent Image RecognitionabstractThis paper presents a recurrent neural network architecture, guided by the bottom-up attention, for the recognition task. The proposed architecture processes an input image as a sequence of selectively chosen patches. The patches are chosen from the salient regions of the input image. Using human driven saliency maps from gaze, the benefit of such a selection process is first shown. Next, the performance of computational models of bottom-up attention are assessed as alternative to human attention. Hamed Rezazadegan Tavakoli, Ali Borji, Rao Muhammad Anwer, Esa Rahtu, Juho Kannala |
ICIP | 1 |
| 2017 | Saliency Revisited: Analysis of Mouse Movements Versus FixationsabstractThis paper revisits visual saliency prediction by evaluating the recent advancements in this field such as crowd-sourced mouse tracking-based databases and contextual annotations. We pursue a critical and quantitative approach towards some of the new challenges including the quality of mouse tracking versus eye tracking for model training and evaluation. We extend quantitative evaluation of models in order to incorporate contextual information by proposing an evaluation methodology that allows accounting for contextual factors such as text, faces, and object attributes. The proposed contextual evaluation scheme facilitates detailed analysis of models and helps identify their pros and cons. Through several experiments, we find that (1) mouse tracking data has lower inter-participant visual congruency and higher dispersion, compared to the eye tracking data, (2) mouse tracking data does not totally agree with eye tracking in general and in terms of different contextual regions in specific, and (3) mouse tracking data leads to acceptable results in training current existing models, and (4) mouse tracking data is less reliable for model selection and evaluation. The contextual evaluation also reveals that, among the studied models, there is no single model that performs best on all the tested annotations. Hamed Rezazadegan Tavakoli, Fawad Ahmed, Ali Borji, Jorma Laaksonen |
CVPR | 1 |
| 2017 | Paying Attention to Descriptions Generated by Image Captioning ModelsabstractTo bridge the gap between humans and machines in image understanding and describing, we need further insight into how people describe a perceived scene. In this paper, we study the agreement between bottom-up saliency-based visual attention and object referrals in scene description constructs. We investigate the properties of human-written descriptions and machine-generated ones. We then propose a saliency-boosted image captioning model in order to investigate benefits from low-level cues in language models. We learn that (1) humans mention more salient objects earlier than less salient ones in their descriptions, (2) the better a captioning model performs, the better attention agreement it has with human descriptions, (3) the proposed saliencyboosted model, compared to its baseline form, does not improve significantly on the MS COCO database, indicating explicit bottom-up boosting does not help when the task is well learnt and tuned on a data, (4) a better generalization is, however, observed for the saliency-boosted model on unseen data. Hamed Rezazadegan Tavakoli, Rakshith Shetty, Ali Borji, Jorma Laaksonen |
ICCV | 1 |
| 2017 | Exploiting inter-image similarity and ensemble of extreme learners for fixation prediction using deep features
Hamed Rezazadegan Tavakoli, Ali Borji, Jorma Laaksonen, Esa Rahtu |
Neurocomputing | 1 |
| 2014 | Emotional Valence Recognition, Analysis of Salience and Eye MovementsabstractThis paper studies the performance of recorded eye movements and computational visual attention models (i.e. saliency models) in the recognition of emotional valence of an image. In the first part of this study, it employs eye movement data (fixation & saccade) to build image content descriptors and use them with support vector machines to classify the emotional valence. In the second part, it examines if the human saliency map can be substituted with the state-of-the-art computational visual attention models in the task of valence recognition. The results indicate that the eye movement based descriptors provide significantly better performance compared to the baselines, which apply low-level visual cues (e.g. color, texture and shape). Furthermore, it will be shown that the current computational models for visual attention are not able to capture the emotional information in similar extent as the real eye movements. Hamed Rezazadegan Tavakoli, Victoria Yanulevskaya, Esa Rahtu, Janne Heikkilä, Nicu Sebe |
ICPR | 1 |
| 2013 | Spherical Center-Surround for Video Saliency Detection Using Sparse Sampling
Hamed Rezazadegan Tavakoli, Esa Rahtu, Janne Heikkilä |
ACIVS | 1 |
| 2013 | Analysis of Scores, Datasets, and Models in Visual Saliency PredictionabstractSignificant recent progress has been made in developing high-quality saliency models. However, less effort has been undertaken on fair assessment of these models, over large standardized datasets and correctly addressing confounding factors. In this study, we pursue a critical and quantitative look at challenges (e.g., center-bias, map smoothing) in saliency modeling and the way they affect model accuracy. We quantitatively compare 32 state-of-the-art models (using the shuffled AUC score to discount center-bias) on 4 benchmark eye movement datasets, for prediction of human fixation locations and scan path sequence. We also account for the role of map smoothing. We find that, although model rankings vary, some (e.g., AWS, LG, AIM, and HouNIPS) consistently outperform other models over all datasets. Some models work well for prediction of both fixation locations and scan path sequence (e.g., Judd, GBVS). Our results show low prediction accuracy for models over emotional stimuli from the NUSEF dataset. Our last benchmark, for the first time, gauges the ability of models to decode the stimulus category from statistics of fixations, saccades, and model saliency values at fixated locations. In this test, ITTI and AIM models win over other models. Our benchmark provides a comprehensive high-level picture of the strengths and weaknesses of many popular models, and suggests future research directions in saliency modeling. Ali Borji, Hamed Rezazadegan Tavakoli, Dicky N. Sihite, Laurent Itti |
ICCV | 2 |
| 2013 | Stochastic bottom-up fixation prediction and saccade generation
Hamed Rezazadegan Tavakoli, Esa Rahtu, Janne Heikkilä |
Image Vis. Comput. | 1 |
| 2009 | Automated Center of Radial Distortion Estimation, Using Active Targets
Hamed Rezazadegan Tavakoli, Hamid Reza Pourreza |
ACCV (2) | 1 |