EDBT 2026 Demo / reviewers in the wild / expert
Farhad Pakdaman
dblp:165/8348
· DBLP profile ↗
23ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0001-6526-3811ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 15 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JNTD-DS: A Benchmark Dataset for Just Noticeable frame rate-based Temporal Difference in Perceptual Video CodingabstractThe Just Noticeable Difference (JND) is defined as the maximum change in a visual stimulus (image or video) which the Human Visual System (HVS) can tolerate without perceiving visual distortion. Previous JND research has largely focused on spatial distortions, yielding several datasets and models that predict spatial thresholds based on parameters such as Quantization Parameter (QP) or Quality Factor (QF). However, temporal thresholds, specifically the maximum frame rate reductions which viewers cannot detect, remain largely unexplored, despite their critical importance for efficient video coding. To address this gap, we introduce JNTD-DS, which, to the best of our knowledge, is the first benchmark dataset specifically designed to measure the Just Noticeable frame rate-based Temporal Difference (JNTD). The dataset comprises 50 video scenes covering various content, and the JNTD level associated with them. T e video scenes are studied through extensive subjective tests, comparing the high frame rate videos with their temporally downsampled versions. This forms 1196 opinion scores from 78 subjects. Analyzing the collected data confirms that JNTD thresholds, which are fundamentally defined by the HVS, are inherently complex and vary across content. By providing critical insights into HVS sensitivity to frame rate changes, the dataset enables content-adaptive frame rate optimization for perceptual video coding, allowing more efficient compression in video streaming and bandwidth-limited applications without compromising visual quality. We further demonstrate the practical impact of these insights by developing a JNTD prediction model and integrating it into a video compression pipeline, achieving an average bitrate reduction of 13.62% with only a marginal quality loss. The JNTD-DS is publicly available at https://github.com/sanaznami/JNTD-DS. Sanaz Nami, Farhad Pakdaman, Sahab Taali, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
MMSys | 2 |
| 2026 | JNTD: Toward Just Noticeable Frame Rate-Based Temporal Difference for Perceptual Video CodingabstractJust Noticeable Difference (JND) refers to the maximum level of distortion in an image or video sequence that remains imperceptible to the Human Visual System (HVS). Current JND-based studies predominantly rely on existing datasets, developing models predicting JND levels in terms of Quantization Parameter (QP) or Quality Factor (QF). However, these solutions primarily focus on spatial-based Perceptual Video Coding (PVC) and neglect temporal-based optimization, which highly affects the video bitrate. This paper addresses this limitation by introducing Just Noticeable frame rate-based Temporal Difference (JNTD) to determine the optimal Frame Rate (FR) based on human perception. A novel dataset comprising 50 high frame rate video sequences is collected through subjective assessments. Subsequently, an ensemble method is proposed to predict the JNTD, by leveraging deep and hand-crafted features, for robust prediction. Experimental evaluations include the integration of the proposed method into several codecs (H.264, H.265, H.266, and a new learned codec), showcasing its ability to reduce bitrate without compromising visual quality. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Hierarchical Transformer for Panoramic Image Inpainting with Comprehensive Attention Module
Li Yu 0004, Yanjun Gao, Yihang Yin, Farhad Pakdaman, Moncef Gabbouj |
ICIG (2) | 4 |
| 2025 | Evaluating the Emerging MPEG Video Coding for Machines in Semantic SegmentationabstractEmerging MPEG Video Coding for Machines (MPEG VCM) standardization activities address the growing demand for machine-to-machine visual applications, including video surveillance, autonomous driving, etc. This paper proposes an evaluation methodology tailored to MPEG VCM, with an emphasis on semantic segmentation tasks using the Pandaset dataset. This is a challenging target as standardization works must follow several limitations, such as dataset's licensing, fixed tools and software, and compliance with existing common test conditions (CTC) for the standard's development. The proposed evaluation methodology includes a step to align Pandaset and COCO labels. Two task networks, Detectron2 and Mask2Former, are used to evaluate the Rate-Performance behavior for semantic segmentation under various coding configurations. The performance of MPEG VCM is benchmarked against traditional codecs (VVC and HEVC), with a detailed analysis of MPEG VCM's coding tools. The extensive evaluations reveal interesting observations. (1) Although VCM achieved reasonably good segmentation performance, some of its developed tools, such as temporal resampling and region-of-interest coding, were not well suited for segmentation task. (2) The Hybrid NNVVC Inner Codec outperformed the VVC Inner Codec. (3) VCM's performance varies significantly for segmented classes. (4) Despite significant differences in human vision performance, VVC and HEVC exhibit relatively similar performance in machine vision. The main contributions of this work are to (1) enable evaluation of MPEG VCM in a real-world semantic segmentation use case, which is one of VCM's targeted tasks, and (2) to provide a detailed assessment of VCM's performance in semantic segmentation. Khoa Dang Pham, Farhad Pakdaman, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Nam Le 0003, Jukka I. Ahonen, Moncef Gabbouj |
ISM | 2 |
| 2025 | Randomized PCA forest for approximate k-nearest neighbor searchabstractk-Nearest Neighbors (kNN) search is the problem of finding k points which are the closest to a given query point . It is used widely in a wide range of tasks and is among the most important tools in applied machine learning . Traditional algorithms for kNN search require computing distances between a query point and all other points in the dataset, and therefore is very slow and inefficient for large data. In this paper, we propose an approximate algorithm for kNN search to find the nearest neighbors fast and efficiently. We employ a tree-based structure which offers robustness and scalability. We propose to use Principal Component Analysis (PCA) to find the best splitting direction to fit the data on the trees. Seeking solutions with low computational complexity , (1) we use a randomized Singular Value Decomposition solver, which reduces PCA complexity from being associated with the number of features to being associated with the number of required principal values; (2) we reuse PCA calculations in multiple nodes to save computation while maintaining accuracy; (3) we ensemble these trees for improved performance, and (4) finally, we propose several variants of the proposed method which target a higher accuracy or a higher efficiency. Extensive experimental results show that proposed solutions outperform existing methods in terms of accuracy, while maintaining competitive complexity. The fast implementation variant of the proposed method outperforms existing techniques in terms of complexity and shows competitive accuracy in performing k-nearest neighbors’ search. Muhammad Rajabinasab, Farhad Pakdaman, Arthur Zimek, Moncef Gabbouj |
Expert Syst. Appl. | 2 |
| 2025 | Deep-BrownConrady: Prediction of Camera Calibration and Distortion Parameters Using Deep Learning and Synthetic DataabstractThis research addresses the challenge of camera calibration and distortion parameter prediction from a single image using deep learning models. The main contributions of this work are: (1) demonstrating that a deep learning model, trained on a mix of real and synthetic images, can accurately predict camera and lens parameters from a single image, and (2) developing a comprehensive synthetic dataset using the AILiveSim simulation platform. This dataset includes variations in focal length and lens distortion parameters, providing a robust foundation for model training and testing. The training process predominantly relied on these synthetic images, complemented by a small subset of real images, to explore how well models trained on synthetic data can perform calibration tasks on real-world images. Traditional calibration methods require multiple images of a calibration object from various orientations, which is often not feasible due to the lack of such images in publicly available datasets. A deep learning network based on the ResNet architecture was trained on this synthetic dataset to predict camera calibration parameters following the Brown-Conrady lens model. The ResNet architecture, adapted for regression tasks, is capable of predicting continuous values essential for accurate camera calibration in applications such as autonomous driving, robotics, and augmented reality. Faiz Muhammad Chaudhry, Jarno Ralli, Jérôme Leudet, Fahad Sohrab, Farhad Pakdaman, Pierre Corbani, Moncef Gabbouj |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Panoramic Image Inpainting with Gated Convolution and Contextual Reconstruction LossabstractDeep learning-based methods have demonstrated encouraging results in tackling the task of panoramic image inpainting. However, it is challenging for existing methods to distinguish valid pixels from invalid pixels and find suitable references for corrupted areas, thus leading to artifacts in the inpainted results. In response to these challenges, we propose a panoramic image inpainting framework that consists of a Face Generator, a Cube Generator, a side branch, and two discriminators. We use the Cubemap Projection (CMP) format as network input. The generator employs gated convolutions to distinguish valid pixels from invalid ones, while a side branch is designed utilizing contextual reconstruction (CR) loss to guide the generators to find the most suitable reference patch for inpainting the missing region. The proposed method is compared with state-of-the-art (SOTA) methods on SUN360 Street View dataset in terms of PSNR and SSIM. Experimental results and ablation study demonstrate that the proposed method outperforms SOTA both quantitatively and qualitatively. Li Yu 0004, Yanjun Gao, Farhad Pakdaman, Moncef Gabbouj |
ICASSP | 3 |
| 2024 | Pixel-Wise Color Constancy Via Smoothness Techniques In Multi-Illuminant ScenesabstractMost scenes are illuminated by several light sources, where the traditional assumption of uniform illumination is invalid. This issue is ignored in most color constancy methods, primarily due to the complex spatial impact of multiple light sources on the image. Moreover, most existing multi-illuminant methods fail to preserve the smooth change of illumination, which stems from spatial dependencies in natural images. Motivated by this, we propose a novel multi-illuminant color constancy method, by learning pixel-wise illumination maps caused by multiple light sources. The proposed method enforces smoothness within neighboring pixels, by regularizing the training with the total variation loss. Moreover, a bilateral filter is provisioned further to enhance the natural appearance of the estimated images, while preserving the edges. Additionally, we propose a label-smoothing technique that enables the model to generalize well despite the uncertainties in ground truth. Quantitative and qualitative experiments demonstrate that the proposed method outperforms the state-of-the-art. Umut Cem Entok, Firas Laakom, Farhad Pakdaman, Moncef Gabbouj |
ICIP | 3 |
| 2024 | Perceptual Learned Image Compression via End-to-End JND-Based OptimizationabstractEmerging Learned image Compression (LC) achieves significant improvements in coding efficiency by end-to-end training of neural networks for compression. An important benefit of this approach over traditional codecs is that any optimization criteria can be directly applied to the encoder-decoder networks during training. Perceptual optimization of LC to comply with the Human Visual System (HVS) is among such criteria, which has not been fully explored yet. This paper addresses this gap by proposing a novel framework to integrate Just Noticeable Distortion (JND) principles into LC. Leveraging existing JND datasets, three perceptual optimization methods are proposed to integrate JND into the LC training process: (1) Pixel-Wise JND Loss (PWL) prioritizes pixel-by-pixel fidelity in reproducing JND characteristics, (2) Image-Wise JND Loss (IWL) emphasizes on overall imperceptible degradation levels, and (3) Feature-Wise JND Loss (FWL) aligns the reconstructed image features with perceptually significant features. Experimental evaluations demonstrate the effectiveness of JND integration, highlighting improvements in rate-distortion performance and visual quality, compared to baseline methods. The proposed methods add no extra complexity after training. Farhad Pakdaman, Sanaz Nami, Moncef Gabbouj |
ICIP | 1 |
| 2024 | Joint End-to-End Image Compression and Denoising: Leveraging Contrastive Learning and Multi-Scale Self-OnnsabstractNoisy images are a challenge to image compression algorithms due to the inherent difficulty of compressing noise. As noise cannot easily be discerned from image details, such as high-frequency signals, its presence leads to extra bits needed for compression Since the emerging learned image compression paradigm enables end-to-end optimization of codecs, recent efforts were made to integrate denoising into the compression model, relying on clean image features to guide denoising. However, these methods exhibit suboptimal performance under high noise levels, lacking the capability to generalize across diverse noise types. In this paper, we propose a novel method integrating a multi-scale denoiser comprising of Self Organizing Operational Neural Networks, for joint image compression and denoising. We employ contrastive learning to boost the network ability to differentiate noise from high frequency signal components, by emphasizing the correlation between noisy and clean counterparts. Experimental results demonstrate the effectiveness of the proposed method both in rate-distortion performance, and codec speed, outperforming the current state-of-the-art. Li Yu 0004, Farhad Pakdaman, Moncef Gabbouj |
ICIP | 3 |
| 2024 | Channel-Wise Feature Decorrelation for Enhanced Learned Image CompressionabstractThe emerging Learned Compression (LC) replaces the traditional codec modules with Deep Neural Networks (DNN), which are trained end-to-end for rate-distortion performance. This approach is considered as the future of image/video compression, and major efforts have been dedicated to improving its compression efficiency. However, most proposed works target compression efficiency by employing more complex DNNS, which contributes to higher computational complexity. Alternatively, this paper proposes to improve compression by fully exploiting the existing DNN capacity. To do so, the latent features are guided to learn a richer and more diverse set of features, which corresponds to better reconstruction. A channel-wise feature decorrelation loss is designed and is integrated into the LC optimization. Three strategies are proposed and evaluated, which optimize (1) the transformation network, (2) the context model, and (3) both networks. Experimental results on two established LC methods show that the proposed method improves the compression with a BD-Rate of up to 8.06%, with no added complexity. The proposed solution can be applied as a plug-and-play solution to optimize any similar LC method. Farhad Pakdaman, Moncef Gabbouj |
IEEE Signal Process. Lett. | 1 |
| 2024 | Lightweight Multitask Learning for Robust JND Prediction Using Latent Space and Reconstructed FramesabstractThe Just Noticeable Difference (JND) refers to the smallest distortion in an image or video that can be perceived by Human Visual System (HVS), and is widely used in optimizing image/video compression. However, accurate JND modeling is very challenging due to its content dependence, and the complex nature of the HVS. Recent solutions train deep learning based JND prediction models, mainly based on a Quantization Parameter (QP) value, representing a single JND level, and train separate models to predict each JND level. We point out that a single QP-distance is insufficient to properly train a network with millions of parameters, for a complex content-dependent task. Inspired by recent advances in learned compression and multitask learning, we propose to address this problem by (1) learning to reconstruct the JND-quality frames, jointly with the QP prediction, and (2) jointly learning several JND levels to augment the learning performance. We propose a novel solution where first, an effective feature backbone is trained by learning to reconstruct JND-quality frames from the raw frames. Second, JND prediction models are trained based on features extracted from latent space (i.e., compressed domain), or reconstructed JND-quality frames. Third, a multi-JND model is designed, which jointly learns three JND levels, further reducing the prediction error. Extensive experimental results demonstrate that our multi-JND method outperforms the state-of-the-art and achieves an average JND1prediction error of only 1.57 in QP, and 0.72 dB in PSNR. Moreover, the multitask learning approach, and compressed domain prediction facilitate light-weight inference by significantly reducing the complexity and the number of parameters. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Comprehensive Complexity Assessment of Emerging Learned Image Compression on CPU and GPUabstractLearned Compression (LC) is the emerging technology for compressing image and video content, using deep neural networks. Despite being new, LC methods have already gained a compression efficiency comparable to state-of-the-art image compression, such as HEVC or even VVC. However, the existing solutions often require a huge computational complexity, which discourages their adoption in international standards or products. This paper provides a comprehensive complexity assessment of several notable methods, that shed light on the matter, and guide the future development of this field by presenting key findings. To do so, six existing methods have been evaluated for both encoding and decoding, on CPU and GPU platforms. Various aspects of complexity such as the overall complexity, share of each coding module, number of operations, number of parameters, most demanding GPU kernels, and memory requirements have been measured and compared on Kodak dataset. The reported results (1) quantify the complexity of LC methods, (2) fairly compare different methods, and (3) a major contribution of the work is identifying and quantifying the key factors affecting the complexity. Farhad Pakdaman, Moncef Gabbouj |
ICASSP | 1 |
| 2023 | MTJND: Multi-Task Deep Learning Framework for Improved JND PredictionabstractThe limitation of the Human Visual System (HVS) in perceiving small distortions allows us to lower the bitrate required to achieve a certain visual quality. Predicting and applying the Just Noticeable Distortion (JND), which is a threshold for maximum unperceived level of distortions, is among the popular ways to do so. Recently, machine learning based methods have been able to reduce bitrate even further by improving JND prediction accuracy. However, accurate modeling of JND is very challenging, as it is highly content dependent. Furthermore, existing datasets provide little information to learn the best parameters. To remedy this issue, we propose a multi-task deep learning framework that jointly learns various complementary visual information. We design three separate methods and training strategies that jointly learn: (1) three JND levels, (2) visual attention map and a JND level, and (3) three JND levels and the visual attention map. We show that accumulating information from multiple tasks leads to a more robust prediction of JND. Experimental results confirm the superiority of our framework compared to the state-of-the-art. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
ICIP | 2 |
| 2023 | NCOD: Near-Optimum Video Compression for Object DetectionabstractWith the emergence of technologies like smart cities, Internet of things (IoT), and 5G, the amount of produced visual data at the edges and remote nodes has exploded. Since for a considerable portion of the captured video the target is a machine learning task, rather than a human audience, transmission of videos in such applications requires efficient video compression tailored for machine vision. However, existing compression solutions are optimized for human vision. This paper presents a methodology to optimize an existing video compression standard, HEVC, for a machine vision task, Object Detection (OD). To this end, (1) a dataset of compressed videos, including several compression-ratios and their corresponding OD performance is collected to enable modeling, (2) A trade-off point (knee-point) between bitrate and OD performance is defined, that finds the point after which no major improvements will be achieved, (3) a set of features were extracted and studied to model this point, via a practical machine learning method. The resulting solution can predict the knee-point with$\boldsymbol{\text{MAE}=1.28}$, resulting in a ∆Recall of only 0.012 and bitrate reduction of 86.56%, compared to OD with very high-quality video. Ardavan Elahi, Ali Falahati, Farhad Pakdaman, Mehdi Modarressi, Moncef Gabbouj |
ISCAS | 3 |
| 2023 | MAMIQA: No-Reference Image Quality Assessment Based on Multiscale Attention Mechanism With Natural Scene StatisticsabstractNo-Reference Image Quality Assessment aims to evaluate the perceptual quality of an image, according to human perception. Many recent studies use Transformers to assign different self-attention mechanisms to distinguish regions of an image, simulating the perception of the human visual system (HVS). However, the quadratic computational complexity caused by the self-attention mechanism is time-consuming and expensive. Meanwhile, the image resizing in the feature extraction stage loses the full-size image quality. To address these issues, we propose a lightweight attention mechanism using decomposed large-kernel convolutions to extract multiscale features, and a novel feature enhancement module to simulate HVS. We also propose to compensate the information loss caused by image resizing, with supplementary features from natural scene statistics. Experimental results on five standard datasets show that the proposed method surpasses the SOTA, while significantly reducing the computational costs. Li Yu 0004, Farhad Pakdaman, Miaogen Ling, Moncef Gabbouj |
IEEE Signal Process. Lett. | 3 |
| 2023 | BL-JUNIPER: A CNN-Assisted Framework for Perceptual Video Coding Leveraging Block-Level JNDabstractJust Noticeable Distortion (JND) finds the minimum distortion level perceivable by humans. This can be a natural solution for setting the compression for each video region in perceptual video coding. However, existing JND-based solutions estimate JND levels for each video frame and ignore the fact that different video regions have different perceptual importance. To address this issue, we propose a Block-Level Just Noticeable Distortion-based Perceptual (BL-JUNIPER) framework for video coding. The proposed four-stage framework combines different perceptual information to further improve the prediction accuracy. The JND mapping in the first stage derives block-level JNDs from frame-level information without the need to collect a new bock-level JND dataset. In the second stage, an efficient CNN-based model is proposed to predict JND levels for each block according to spatial and temporal characteristics. Unlike existing methods, BL-JUNIPER works on raw video frames and avoids re-encoding each frame several times, making it computationally practical. Third, the visual importance of each block is measured using a visual attention model. Finally, a proposed quantization control algorithm uses both JND levels and visual importance to adjust the Quantization Parameter (QP) for each block. The specific algorithm for each stage of the proposed framework can be changed, as long as the input and output formats of each block are followed, without the need to change other stages, based on any current or future methods, providing a flexible and robust solution. Extensive experimental results demonstrate that BL-JUNIPER achieves a mean bitrate reduction of 27.75% with a Delta Mean Opinion Score (DMOS) close to zero and BD-Rate gains of 25.44% based on MOS, compared to the baseline encoding, and also gains a better performance compared to competing methods. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi |
IEEE Trans. Multim. | 2 |
| 2021 | SVM based approach for complexity control of HEVC intra coding
Farhad Pakdaman, Li Yu 0004, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001, Moncef Gabbouj |
Signal Process. Image Commun. | 1 |
| 2020 | Complexity Analysis Of Next-Generation VVC Encoding And DecodingabstractWhile the next generation video compression standard, Versatile Video Coding (VVC), provides a superior compression efficiency, its computational complexity dramatically increases. This paper thoroughly analyzes this complexity for both encoder and decoder of VVC Test Model 6, by quantifying the complexity break-down for each coding tool and measuring the complexity and memory requirements for VVC encoding/decoding. These extensive analyses are performed for six video sequences of 720p, 1080p, and 2160p, under Low-Delay (LD), Random-Access (RA), and All-Intra (AI) conditions (a total of 320 encoding/decoding). Results indicate that the VVC encoder and decoder are 5× and 1.5× more complex compared to HEVC in LD, and 31× and 1.8× in AI, respectively. Detailed analysis of coding tools reveals that in LD on average, motion estimation tools with 53%, transformation and quantization with 22%, and entropy coding with 7% dominate the encoding complexity. In decoding, loop filters with 30%, motion compensation with 20%, and entropy decoding with 16%, are the most complex modules. Moreover, the required memory bandwidth for VVC encoding/decoding are measured through memory profiling, which are 30× and 3× of HEVC. The reported results and insights are a guide for future research and implementations of energy-efficient VVC encoder/decoder. Farhad Pakdaman, Mohammad Ali Adelimanesh, Moncef Gabbouj, Mahmoud Reza Hashemi |
ICIP | 1 |
| 2020 | A low complexity and computationally scalable fast motion estimation algorithm for HEVC
Farhad Pakdaman, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001 |
Multim. Tools Appl. | 1 |
| 2019 | A computationally scalable fast intra coding scheme for HEVC video encoder
Elahe Hosseini, Farhad Pakdaman, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Fast and efficient intra mode decision for HEVC, based on dual-tree complex wavelet
Farhad Pakdaman, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001 |
Multim. Tools Appl. | 1 |
| 2015 | Integrated circuit-packet switching NoC with efficient circuit setup mechanism
Farhad Pakdaman, Abbas Mazloumi, Mehdi Modarressi |
J. Supercomput. | 1 |