Mahmoud Reza Hashemi

dblp:90/5731 · DBLP profile ↗
← Back
44ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-3518-9195ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 JNTD-DS: A Benchmark Dataset for Just Noticeable frame rate-based Temporal Difference in Perceptual Video Coding
abstract
The Just Noticeable Difference (JND) is defined as the maximum change in a visual stimulus (image or video) which the Human Visual System (HVS) can tolerate without perceiving visual distortion. Previous JND research has largely focused on spatial distortions, yielding several datasets and models that predict spatial thresholds based on parameters such as Quantization Parameter (QP) or Quality Factor (QF). However, temporal thresholds, specifically the maximum frame rate reductions which viewers cannot detect, remain largely unexplored, despite their critical importance for efficient video coding. To address this gap, we introduce JNTD-DS, which, to the best of our knowledge, is the first benchmark dataset specifically designed to measure the Just Noticeable frame rate-based Temporal Difference (JNTD). The dataset comprises 50 video scenes covering various content, and the JNTD level associated with them. T e video scenes are studied through extensive subjective tests, comparing the high frame rate videos with their temporally downsampled versions. This forms 1196 opinion scores from 78 subjects. Analyzing the collected data confirms that JNTD thresholds, which are fundamentally defined by the HVS, are inherently complex and vary across content. By providing critical insights into HVS sensitivity to frame rate changes, the dataset enables content-adaptive frame rate optimization for perceptual video coding, allowing more efficient compression in video streaming and bandwidth-limited applications without compromising visual quality. We further demonstrate the practical impact of these insights by developing a JNTD prediction model and integrating it into a video compression pipeline, achieving an average bitrate reduction of 13.62% with only a marginal quality loss. The JNTD-DS is publicly available at https://github.com/sanaznami/JNTD-DS.
Sanaz Nami, Farhad Pakdaman, Sahab Taali, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj
MMSys4
2026 JNTD: Toward Just Noticeable Frame Rate-Based Temporal Difference for Perceptual Video Coding
abstract
Just Noticeable Difference (JND) refers to the maximum level of distortion in an image or video sequence that remains imperceptible to the Human Visual System (HVS). Current JND-based studies predominantly rely on existing datasets, developing models predicting JND levels in terms of Quantization Parameter (QP) or Quality Factor (QF). However, these solutions primarily focus on spatial-based Perceptual Video Coding (PVC) and neglect temporal-based optimization, which highly affects the video bitrate. This paper addresses this limitation by introducing Just Noticeable frame rate-based Temporal Difference (JNTD) to determine the optimal Frame Rate (FR) based on human perception. A novel dataset comprising 50 high frame rate video sequences is collected through subjective assessments. Subsequently, an ensemble method is proposed to predict the JNTD, by leveraging deep and hand-crafted features, for robust prediction. Experimental evaluations include the integration of the proposed method into several codecs (H.264, H.265, H.266, and a new learned codec), showcasing its ability to reduce bitrate without compromising visual quality.
Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.3
2024 Lightweight Multitask Learning for Robust JND Prediction Using Latent Space and Reconstructed Frames
abstract
The Just Noticeable Difference (JND) refers to the smallest distortion in an image or video that can be perceived by Human Visual System (HVS), and is widely used in optimizing image/video compression. However, accurate JND modeling is very challenging due to its content dependence, and the complex nature of the HVS. Recent solutions train deep learning based JND prediction models, mainly based on a Quantization Parameter (QP) value, representing a single JND level, and train separate models to predict each JND level. We point out that a single QP-distance is insufficient to properly train a network with millions of parameters, for a complex content-dependent task. Inspired by recent advances in learned compression and multitask learning, we propose to address this problem by (1) learning to reconstruct the JND-quality frames, jointly with the QP prediction, and (2) jointly learning several JND levels to augment the learning performance. We propose a novel solution where first, an effective feature backbone is trained by learning to reconstruct JND-quality frames from the raw frames. Second, JND prediction models are trained based on features extracted from latent space (i.e., compressed domain), or reconstructed JND-quality frames. Third, a multi-JND model is designed, which jointly learns three JND levels, further reducing the prediction error. Extensive experimental results demonstrate that our multi-JND method outperforms the state-of-the-art and achieves an average JND1prediction error of only 1.57 in QP, and 0.72 dB in PSNR. Moreover, the multitask learning approach, and compressed domain prediction facilitate light-weight inference by significantly reducing the complexity and the number of parameters.
Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.3
2023 MTJND: Multi-Task Deep Learning Framework for Improved JND Prediction
abstract
The limitation of the Human Visual System (HVS) in perceiving small distortions allows us to lower the bitrate required to achieve a certain visual quality. Predicting and applying the Just Noticeable Distortion (JND), which is a threshold for maximum unperceived level of distortions, is among the popular ways to do so. Recently, machine learning based methods have been able to reduce bitrate even further by improving JND prediction accuracy. However, accurate modeling of JND is very challenging, as it is highly content dependent. Furthermore, existing datasets provide little information to learn the best parameters. To remedy this issue, we propose a multi-task deep learning framework that jointly learns various complementary visual information. We design three separate methods and training strategies that jointly learn: (1) three JND levels, (2) visual attention map and a JND level, and (3) three JND levels and the visual attention map. We show that accumulating information from multiple tasks leads to a more robust prediction of JND. Experimental results confirm the superiority of our framework compared to the state-of-the-art.
Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj
ICIP3
2023 BL-JUNIPER: A CNN-Assisted Framework for Perceptual Video Coding Leveraging Block-Level JND
abstract
Just Noticeable Distortion (JND) finds the minimum distortion level perceivable by humans. This can be a natural solution for setting the compression for each video region in perceptual video coding. However, existing JND-based solutions estimate JND levels for each video frame and ignore the fact that different video regions have different perceptual importance. To address this issue, we propose a Block-Level Just Noticeable Distortion-based Perceptual (BL-JUNIPER) framework for video coding. The proposed four-stage framework combines different perceptual information to further improve the prediction accuracy. The JND mapping in the first stage derives block-level JNDs from frame-level information without the need to collect a new bock-level JND dataset. In the second stage, an efficient CNN-based model is proposed to predict JND levels for each block according to spatial and temporal characteristics. Unlike existing methods, BL-JUNIPER works on raw video frames and avoids re-encoding each frame several times, making it computationally practical. Third, the visual importance of each block is measured using a visual attention model. Finally, a proposed quantization control algorithm uses both JND levels and visual importance to adjust the Quantization Parameter (QP) for each block. The specific algorithm for each stage of the proposed framework can be changed, as long as the input and output formats of each block are followed, without the need to change other stages, based on any current or future methods, providing a flexible and robust solution. Extensive experimental results demonstrate that BL-JUNIPER achieves a mean bitrate reduction of 27.75% with a Delta Mean Opinion Score (DMOS) close to zero and BD-Rate gains of 25.44% based on MOS, compared to the baseline encoding, and also gains a better performance compared to competing methods.
Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi
IEEE Trans. Multim.3
2022 Game Audio Impacts on Players' Visual Attention, Model Performance for Cloud Gaming
abstract
Cloud gaming (CG) is a new approach to deliver a high-quality gaming experience to gamers anywhere, anytime, and on any device. To achieve this goal, CG requires a high bandwidth, which is still a major challenge. Many existing research pieces have focused on modeling or predicting the players’ Visual Attention Map (VAM) and allocating bitrate accordingly. Although studies indicate that both modalities of audio and video influence human perception, a few studies considered audio impacts in the cloud-based attention models. This paper demonstrates that the audio features in video games change the players’ VAMs in various game scenarios. Our findings indicated that incorporating game audio improves the accuracy of the predicted attention maps by 13% on average compared to the previous VAMs generated based on visual saliency by Game Attention Model for CG. The audio impact is more evident in video games with fewer visual components or indicators on the screen.
Morva Saaty, Mahmoud Reza Hashemi
ETRA2
2022 GAMORRA: An API-level workload model for rasterization-based graphics pipeline architecture
Iman Soltani Mohammadi, Mohammed Ghanbari 0001, Mahmoud Reza Hashemi
Comput. Graph.3
2022 An efficient six-parameter perspective motion model for VVC
Iman Soltani Mohammadi, Mohammed Ghanbari 0001, Mahmoud Reza Hashemi
J. Vis. Commun. Image Represent.3
2021 SVM based approach for complexity control of HEVC intra coding
Farhad Pakdaman, Li Yu 0004, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001, Moncef Gabbouj
Signal Process. Image Commun.3
2020 Complexity Analysis Of Next-Generation VVC Encoding And Decoding
abstract
While the next generation video compression standard, Versatile Video Coding (VVC), provides a superior compression efficiency, its computational complexity dramatically increases. This paper thoroughly analyzes this complexity for both encoder and decoder of VVC Test Model 6, by quantifying the complexity break-down for each coding tool and measuring the complexity and memory requirements for VVC encoding/decoding. These extensive analyses are performed for six video sequences of 720p, 1080p, and 2160p, under Low-Delay (LD), Random-Access (RA), and All-Intra (AI) conditions (a total of 320 encoding/decoding). Results indicate that the VVC encoder and decoder are 5× and 1.5× more complex compared to HEVC in LD, and 31× and 1.8× in AI, respectively. Detailed analysis of coding tools reveals that in LD on average, motion estimation tools with 53%, transformation and quantization with 22%, and entropy coding with 7% dominate the encoding complexity. In decoding, loop filters with 30%, motion compensation with 20%, and entropy decoding with 16%, are the most complex modules. Moreover, the required memory bandwidth for VVC encoding/decoding are measured through memory profiling, which are 30× and 3× of HEVC. The reported results and insights are a guide for future research and implementations of energy-efficient VVC encoder/decoder.
Farhad Pakdaman, Mohammad Ali Adelimanesh, Moncef Gabbouj, Mahmoud Reza Hashemi
ICIP4
2020 A low complexity and computationally scalable fast motion estimation algorithm for HEVC
Farhad Pakdaman, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001
Multim. Tools Appl.2
2019 HPCgnature: a hardware-based application-level intrusion detection system
abstract
In the past decade, commodity software applications have been deployed more than ever in almost every domain. Having the ability to differentiate the original trusted application at run‐time from its compromised, mimic or trojanised versions would mitigate a broad range of intrusion threats to these applications. This has been addressed by application‐level intrusion detection systems, however, such schemes mostly depend on the system software for either monitoring or modelling the application. This is while system software can itself get compromised by kernel‐level rootkit attacks. In this study, the authors have proposed a new hardware‐based app‐IDS, which works independent of the system software of the target system. The proposed method, referred to as HPCgnature , includes a new abstraction corresponding to the repetitious functionalities of programs. Such functionalities generate a distinguishing sequence of periods, referred to in this study as the Operational Periodicity . The method uses monitoring scheme based on external access to the hardware performance counters of CPUs. Implementing a prototype, they have shown how HPCgnature can detect intrusions in 12 complex interactive desktop applications. Evaluation results indicate this model could differentiate applications with 98% accuracy, and can detect even small run‐time code injection attacks by an accuracy of >75%
Seyyedeh Atefeh Musavi, Mahmoud Reza Hashemi
IET Inf. Secur.2
2019 A computationally scalable fast intra coding scheme for HEVC video encoder
Elahe Hosseini, Farhad Pakdaman, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001
Multim. Tools Appl.3
2019 An Ontology-Based Method for HW/SW Architecture Reconstruction
abstract
To address the vast variety of computing requirements in recent ubiquitous computing ecosystem, there is a constant need for more complex computing systems that consist of integrated hardware (HW) and software (SW) systems. Providing an architectural insight into such systems helps in achieving a more efficient usage of system resources, verifying the characteristics of a platform and provisioning of its security and trust. Architecture reconstruction (AR) has been used in software engineering to gain a deeper insight into specific software. Neither software AR nor hardware reverse engineering techniques are sufficient to extract the architecture of a system that incorporates both HW/SW, since they are unable to recover the relationships between the HW and SW components. Inspired by the Symphony software AR framework, we propose a method to reconstruct the architecture of a computing platform as a whole. In order to cover the wide variety of existing HW/SW technologies, our method uses an ontology-based approach. Due to the lack of a comprehensive ontology in literature, we developed PLATOnt, a new ontology that has been shown to be more effective by OntoQA evaluation framework. We used our AR method to reconstruct the architecture of an ARM-based trusted execution environment and a Raspberry Pi platform, widely used in embedded systems and IoT devices.
Seyyedeh Atefeh Musavi, Mahmoud Reza Hashemi
IEEE Trans. Computers2
2017 Fast and efficient intra mode decision for HEVC, based on dual-tree complex wavelet
Farhad Pakdaman, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001
Multim. Tools Appl.2
2016 A Testing Apparatus for Faster and More Accurate Subjective Assessment of Quality of Experience in Cloud Gaming
abstract
The number of cloud gaming (CG) users is constantly growing. The idea in cloud gaming is to render the game events on a cloud server and stream the resulted scenes as a video sequence to players. CG requires a high bandwidth in order to run appropriately and create a good quality of experience for players. In order to reduce the high required bandwidth, video should be compressed without any negative impact on user's quality of experience (QoE). Thus CG providers, researchers who develop new compression methods for CG, and those who are improving network protocols for CG require to evaluate user experience using subjective methods. Over the years, many researches have investigated the subjective quality of video, but all of them have one of the following two main drawbacks, which makes them unsuitable for game videos. The subjective quality assessment methods which are designed for short duration video sequences suffer from Forgiveness and Recency effects. On the other hand, the methods which are designed for long duration video sequences usually use some sort of a handset device for rating scores, and hence cannot be used for most games where both hands are busy while playing. In this paper, a novel subjective test apparatus for assessment of game videos is proposed, where players give their opinion scores using a foot pedal while playing the game. Evaluation results indicate that the proposed scheme is more accurate and less distractive than existing methods.
Saeed Shafiee Sabet, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001
ISM2
2016 GSET somi: a game-specific eye tracking dataset for somi
abstract
In this paper, we present an eye tracking dataset of computer game players who played the side-scrolling cloud game Somi. The game was streamed in the form of video from the cloud to the player. This dataset can be used for designing and testing game-specific visual attention models. The source code of the game is also available to facilitate further modifications and adjustments. For collecting this data, male and female candidates were asked to play the game in front of a remote eye-tracking device. For each player, we recorded gaze points, video frames of the gameplay, and mouse and keyboard commands. For each video frame, a list of its game objects with their locations and sizes was also recorded. This data, synchronized with eye-tracking data, allows one to calculate the amount of attention that each object or group of objects draw from each player. As a benchmark, we also show various attention patterns could be identified among players.
Hamed Ahmadi, Saman Zad Tootaghaj, Sajad Mowlaei, Mahmoud Reza Hashemi, Shervin Shirmohammadi
MMSys4
2016 Joint application-architeture design space exploration of multimedia applications on many-core platforms - an experimental analysis
Maryam Moghadas, Hossein Afshari, Mahmoud Reza Hashemi
Multim. Tools Appl.3
2016 New R-D Optimization Criterion for Fast Mode Decision Algorithms in Video Coding and Transrating
abstract
Mode decision has a significant effect on the quality and complexity of video coding. It is even more challenging when generating multiple bitstreams with different bitrates (BRs) in, for example, dynamic adaptive streaming for HTTP or transrating systems. Full search and simplified fast search mode decision (MD) methods either suffer from a high computational complexity or have a negative impact on quality. Furthermore, mode selection in conventional approaches strongly depends on the quantization parameter (QP). Hence, modes that have been selected for high BR compression may not be suitable for low BR when transrating a bitstream. In this paper, we propose a rate-distortion (R-D)-optimized criterion for fast MD algorithms. The proposed cost function, when adopted in different fast MD algorithms, not only improves the R-D performance by up to 6.6% in terms of Bjøntegaard delta rate, but also reduces the execution time of the encoder by up to 6.8%. We also show that modes selected by the proposed criterion are less sensitive to changes in BR or QP. As a result, the same modes in an encoded bitstream may be used even after transrating using requantization, resulting in a significant R-D performance improvement of up to 33.3%.
Alireza Aminlou, Mahmoud Reza Hashemi, Moncef Gabbouj, Bing Zeng 0001, Omid Fatemi
IEEE Trans. Circuits Syst. Video Technol.2
2016 A View-Level Rate Distortion Model for Multi-View/3D Video
abstract
Multi-view/3D video is currently available in games, entertainment, education, security, and surveillance applications . Since the amount of data in multi-view/3D increases proportionally with the number of cameras, and due to different bandwidth and playback capabilities of receivers, appropriate compression of multi-view/3D video to produce the correct bitrate while maintaining smooth video quality is crucial, a task that is mostly performed by the rate control module of the encoder. There are many existing rate control algorithms for single-view and multi-view video coding considering the specific features or aspects of these videos. In this paper, we introduce a novel view-level rate distortion (RD) model. We use a systematic methodology to derive this RD model by investigating the impact of multi-view/3D video characteristics on the bitrate of a compressed video. Our proposed RD model considers the concepts of intra-view and inter-view disparity as an effective feature of multi-view/3D video to estimate the overall bitrate of each view more accurately. Evaluation results indicate that our proposed view-level RD model outperforms existing linear models by a factor of 3 and can predict the rate of each view with relatively high precision and a low estimation error of 12% on average.
Hoda Roodaki, Zahra Iravani, Mahmoud Reza Hashemi, Shervin Shirmohammadi
IEEE Trans. Multim.3
2015 An Open Source Cloud Gaming Testbed Using DirectShow
abstract
Despite its challenges, cloud gaming is growing its share in the gaming market by attracting more players. This has led to an increasing number of researches trying to overcome cloud gaming's challenges, including the required high bandwidth and low latency, to make cloud gaming more practical and profitable. To perform this research, researchers need a testbed to evaluate their ideas and find the best solutions. Currently, GamingAnywhere is the only open source platform and testbed to serve this goal. However, it cannot be used to stream all video games, since it depends on hooking APIs which might be incompatible with some video games. In this paper, we introduce a new open source cloud gaming testbed. In this testbed, the screen capturing module is fundamentally a DirectShow filter and, hence, can be tuned for any DirectShow compatible video game. The testbed also facilitates the measurement of delay and quality as the video is processed through its modules.
Hamed Ahmadi, Mahmoud Reza Hashemi, Shervin Shirmohammadi
CloudCom2
2014 Rate-distortion optimization for scalable multi-view video coding
abstract
In recent years, multi-view/3D video applications, such as three-dimensional television (3DTV) and free-viewpoint television (FTV), have drawn increasing attention. Since the amount of data that has to be stored or transmitted increases proportionally with the number of cameras, efficient compression of multi-view/3D video is crucial. Scalable multi-view video coding is one of the methods to address this challenge. But, in streaming multi-view/3D video over a network to heterogeneous receivers, efficient video compression while maintaining a high quality of received video is very challenging. This paper presents a novel method for rate-distortion optimization in scalable multiview video. We apply the Karush-Kuhn-Tucker (KKT) conditions in minimizing the perceptual distortion of decoded video under the conditions that the sum of bits generated from different views is constrained within a given bit budget. Since the constraint-based optimization problem is usually computational intensive, our proposed approach considers the concept of disparity between layers and disparity between views to reduce this computational complexity. Simulation results indicate that the proposed approach is able to meet network bandwidth limitations with acceptable overall video quality.
Hoda Roodaki, Mahmoud Reza Hashemi, Shervin Shirmohammadi
ICME2
2014 A generic, comprehensive and granular decoder complexity model for the H.264/AVC standard
Mehdi Semsarzadeh, Mahmoud Reza Hashemi, Shervin Shirmohammadi
J. Vis. Commun. Image Represent.2
2014 A game attention model for efficient bit rate allocation in cloud gaming
Hamed Ahmadi, Saman Zad Tootaghaj, Mahmoud Reza Hashemi, Shervin Shirmohammadi
Multim. Syst.3
2013 A fine-grain distortion and complexity aware parameter tuning model for the H.264/AVC encoder
Mehdi Semsarzadeh, Atieh Lotfi, Mahmoud Reza Hashemi, Shervin Shirmohammadi
Signal Process. Image Commun.3
2012 A Two-Piece R-D Model for Hybrid Video Coding and Its Application in Fast Mode Decision
abstract
The mode decision process has a significant effect on the quality and complexity of a video encoder. The conventional method that fully codes each macro block for different modes results in the best quality performance, but it suffers from high computational complexity. On the other hand, some other methods ignore the residual part and use the prediction data, or adopt early mode selection approaches in order to reduce the computational cost. These approaches have a negative impact on the coding performance. In this paper, we have used a simple model for the residual coding part and proposed a two-piece R-D model for a macro block. Based on this model, we have introduced a mode decision algorithm that reduces the bit-rate by up to 11.62% at the expense of just 0.5% computational overhead.
Alireza Aminlou, Hana Fahim-Hashemi, Mahmoud Reza Hashemi, Moncef Gabbouj, Omid Fatemi
ICME3
2012 Complexity Modeling of the Motion Compensation Process of the H.264/AVC Video Coding Standard
abstract
With recent advances in computing and communication technologies, ubiquitous access to high quality multimedia content such as high definition video using smart phones, Net books, or tablets is a fact of our daily life. However, power is still a major concern for any mobile device, and requires optimization of power consumption using a power model for each multimedia application, such as a video decoder. In this paper, a generic decoding complexity model for the motion compensation (MC) process, which constitutes up to 25% of the computational complexity and hence power consumption of an H.264/AVC decoder, has been proposed. For the model to remain independent from a specific implementation or platform, it has been developed by analysing the MC algorithm as described in the standard. Simulation results indicate that the proposed model estimates MC complexity with an average accuracy of 95.63%, for a wide range of test sequences using both JM and x.264 software implementations of H.264/AVC. For a dedicated hardware implementation of the MC module the modeling accuracy is around 89.61%, according to our simulation results. It should be noted that in addition to power consumption control, the proposed model can be used for designing a receiver-aware H.264/AVC encoder, where the complexity constraints of the receiver side are taken into account during compression.
Mehdi Semsarzadeh, Mohsen Jamali Langroodi, Mahmoud Reza Hashemi, Shervin Shirmohammadi
ICME3
2012 A new methodology to derive objective quality assessment metrics for scalable multiview 3D video coding
abstract
With the growing demand for 3D video, efforts are underway to incorporate it in the next generation of broadcast and streaming applications and standards. 3D video is currently available in games, entertainment, education, security, and surveillance applications. A typical scenario for multiview 3D consists of several 3D video sequences captured simultaneously from the same scene with the help of multiple cameras from different positions and through different angles. Multiview video coding provides a compact representation of these multiple views by exploiting the large amount of inter-view statistical dependencies. One of the major challenges in this field is how to transmit the large amount of data of a multiview sequence over error prone channels to heterogeneous mobile devices with different bandwidth, resolution, and processing/battery power, while maintaining a high visual quality. Scalable Multiview 3D Video Coding (SMVC) is one of the methods to address this challenge; however, the evaluation of the overall visual quality of the resulting scaled-down video requires a new objective perceptual quality measure specifically designed for scalable multiview 3D video. Although several subjective and objective quality assessment methods have been proposed for multiview 3D sequences, no comparable attempt has been made for quality assessment of scalable multiview 3D video. In this article, we propose a new methodology to build suitable objective quality assessment metrics for different scalable modalities in multiview 3D video. Our proposed methodology considers the importance of each layer and its content as a quality of experience factor in the overall quality. Furthermore, in addition to the quality of each layer, the concept of disparity between layers (inter-layer disparity) and disparity between the units of each layer (intra-layer disparity) is considered as an effective feature to evaluate overall perceived quality more accurately. Simulation results indicate that by using this methodology, more efficient objective quality assessment metrics can be introduced for each multiview 3D video scalable modalities.
Hoda Roodaki, Mahmoud Reza Hashemi, Shervin Shirmohammadi
ACM Trans. Multim. Comput. Commun. Appl.2
2011 Rate-distortion-complexity optimization for VLSI implementation of integer motion estimation in H.264/AVC encoder
abstract
In order to accommodate the wide range of applications and the corresponding platforms where the H.264/AVC standard is currently in place, one should be able to optimize the encoder's computational complexity with a careful selection of the coding configuration parameters. Motion estimation is the most time-consuming part of the encoder which constitutes up to 75% of the computational complexity. In this paper, the optimum selection of configuration parameters, including search range, reference frame, degree of down-sampling and number of truncation bits have been analyzed for the VLSI implementation of integer motion estimation in terms of distortion-complexity performance. Furthermore, the optimum parameter sets have been presented for different video sizes and different constraints on computational power.
Alireza Aminlou, Zahra NajafiHaghi, Majid Namaki-Shoushtari, Mahmoud Reza Hashemi
ICME4
2011 Edge-oriented interpolation for fractional motion estimation in hybrid video coding
abstract
Fractional motion estimation, using the interpolation process, improves the quality of the compressed video by about 2dB in terms of PSNR over integer motion estimation. Most existing interpolation techniques in recent video coding standards, such as the symmetric 6-tap filter of the H.264/AVC standard, do not perform well around the object edges. This increases the value of residuals, which in turn results in higher bit rate. In this paper, a new interpolation method has been proposed that considers the edges of video objects. Simulation results indicate that the proposed method, when used in the H.264/AVC encoder, improves PSNR by up to 1.0 dB (0.4 dB in average) with respect to the standard interpolation technique. This is achieved at the expense of up to 6% (3% in average) increase in the computational complexity.
Ali Kokhazadeh, Alireza Aminlou, Mahmoud Reza Hashemi
ICME3
2011 A new Scalable Multi-View Video Coding configuration for mobile applications
abstract
Transmission of multi-view video content is not practical in most mobile environments due to the limited bandwidth and processing power of mobile devices. To support such environments, one can limit the number of views that are being transmitted, known as Scalable Multi-view Video Coding (SMVC). In this paper, we propose a new view selection method for view scalability in multi-view video coding in mobile environments, which uses inter and intra view dissimilarities to determine the most suitable views for the base layer corresponding to the prediction structure and user selected limited number of views. By selecting more correlated views for the base layer, the proposed method provides an improved performance, as confirmed by simulation results, even when all the enhancement layers are dropped due to network limitations.
Hoda Roodaki, Mahmoud Reza Hashemi, Shervin Shirmohammadi
ICME2
2011 Low-complexity unbalanced multiple description coding based on balanced clusters for adaptive peer-to-peer video streaming
Majid Roohollah Ardestani, Asghar Beheshti 0001, Mahmoud Reza Hashemi
Signal Process. Image Commun.3
2010 An improved low-complexity multiple description coding for peer-to-peer video streaming
abstract
Multiple description scalable coding based on T+2D wavelet decomposition structure is highly flexible for peer-to-peer (P2P) video streaming. Finding the optimal truncation point of each code block (CB) within each description is an NP-hard problem. To implement an efficient low-complexity solution, we propose a simple clustering algorithm for partitioning the CBs into a limited number of clusters, such that one can find the optimal cluster-level redundancy-rate assignment matrix using a low-complexity full search. This approach improves the decoding quality compared to the co-echelon frameworks in which a non-optimal rate assignment matrix is used. In addition, the proposed clustering approach may be analytically represented by closed-form relations for low-complexity computation of optimal encoding parameters. The simulation results demonstrate that the adaptive proposed framework outperforms the approaches in by (0.25~1 dB), the scheme of by (1.3~1.6 dB), and the non-adaptive multiple description coding by (2.3~4.3 dB). Furthermore, the proposed clustering approach requires %52-%88 less computations compared to the framework in.
Majid Roohollah Ardestani, Asghar Beheshti 0001, Mahmoud Reza Hashemi
ICME3
2009 A cost-error optimized architecture for 9/7 lifting based Discrete Wavelet Transform with balanced pipeline stages
abstract
Discrete wavelet transform (DWT) is increasingly recognized in image/video compression standards, as indicated by its use in JPEG2000. The lifting scheme algorithm is an alternative DWT implementation that has a lower computational complexity. In this paper, a new high performance lifting-based architecture is presented for the 9/7 DWT engine. The proposed architecture has a balanced pipeline and improves both the computational error and hardware complexity for any given working frequency. In the proposed architecture, the constant coefficients are modified by introducing new variables to the conventional lifting structure to minimize hardware cost and computational error, imposed by quantization of coefficients. Simulation results indicate a quality improvement of up to 15 dB when compared to an architecture using the standard coefficients that has the same hardware cost and working frequency. Similarly, the hardware cost is reduced by about 20% when both architectures deliver the same PSNR when operating at the same frequency.
Alireza Aminlou, Fatemeh Refan, Mahmoud Reza Hashemi, Omid Fatemi, Saeed Safari
ICASSP3
2009 Low complexity hardware implementation of reciprocal fractional motion estimation for H.264/AVC in mobile applications
abstract
Motion estimation, one of the most effective modules in H.264/AVC, constitutes 60%-90% of encoding time and computation. Close to %45 of this computation belongs to Fractional Motion Estimation (FME) which has to perform a time consuming half-pixel and quarter-pixel interpolation. In addition, interpolation is a major challenge for hardware implementation in real time, specially in mobile applications where processing and battery power is limited. Several modified sub-pixel accuracy search methods have been proposed in the literature in order to reduce the complexity of interpolation. The reciprocal method has been proposed to reduce the CPU encoding time for a software implementation of an H.264 encoder on personal computers. In this paper, the hardware implementation of the reciprocal method is evaluated and its PSNR performance and hardware cost are compared to that of a simplified Rate-Distortion Optimization (RDO) process. Simulation results indicate that using reciprocal FME with enabled RDO has less computational cost and better PSNR performance than using the conventional FME with disabled RDO.
Alireza Aminlou, Parviz Alvandi, Mahmoud Reza Hashemi
PCS3
2009 An improved R-D optimized motion estimation method for video coding
abstract
Motion estimation is one of the key tools to achieve a very low bit rate in video coding. The selection of optimum motion vectors (MV) has a significant impact on the quality of the resulted compressed video, in Rate-Distortion (R-D) sense. The established method uses the Lagrange multiplier to optimally select the MV for each block. However, it does not consider the effect of residual coding at the same time with the effect of MVs. In this paper, we have considered the effect of residual coding in motion vector (MV) selection with modeling motion estimation and residual coding as independent processes which results in a new optimization condition. It is based on local optimization, which is the bit allocation between motion estimation and residual coding, and global optimization, which is the bit allocation among different blocks. The proposed motion estimation algorithm results in a PSNR improvement of 0.5 - 3.0 dB when it is used in the H.264 standard with block size of 4 times 4.
Alireza Aminlou, Mojtaba Farmani, Mahmoud Reza Hashemi, Omid Fatemi
PCS3
2007 A Split Method for Optimized Cost-Quality Hardware Implementation of Lifting-Based Discrete Wavelet Transform
abstract
Discrete wavelet transform (DWT) is increasingly recognized in image/video compression standards, as indicated by its use in JPEG2000. The lifting scheme algorithm is an alternative DWT implementation that has a lower computational complexity. In this paper, a new high performance lifting-based architecture with optimized error vs. hardware complexity is presented for DWT. The proposed architecture modifies the constant coefficients by introducing new variables to the conventional lifting structure to minimize hardware cost and quantization error. In order to achieve the most efficient coefficients, an optimization process has been implemented. Simulation results indicate an average quality improvement of 7.5 dB with the same hardware complexity/cost. Similarly, for achieving the same quality as the conventional hardware implementations the proposed architecture is 20% less complex. The appropriate coefficients can be determined according to the cost and error requirements of each application.
Alireza Aminlou, Fatemeh Refan, Maryam Homayouni, Omid Fatemi, Mahmoud Reza Hashemi
ICASSP (2)5
2007 Pattern-Based Error Recovery of Low Resolution Subbands in JPEG2000
abstract
Digital image transmission is widely used in consumer products, such as digital cameras and cellular phones, where low bit rate coding is required. In any low bit rate encoder, such as the JPEG2000 standard, data truncation (during the encoding process), and data loss (during transmission) will result in lost bit-planes, which will be normally replaced by zeros. In this paper a new algorithm has been proposed, which recovers the lost/truncated lower bit-planes of coefficients in the LL subband of a wavelet transform in a JPEG2000 stream using the data available in higher bit-planes of the same coefficient and its eight neighbors. Simulation results indicate that the proposed algorithm achieves 5.40-8.77 dB improvement with respect to zero filling data recovery method.
Alireza Aminlou, Nasim Hajari, Hossein Badakhshannoory, Mahmoud Reza Hashemi, Omid Fatemi
ICIP (4)4
2007 Two Level Cost-Quality Optimization of 9-7 Lifting-Based Discrete Wavelet Transform
abstract
Implementing the discrete wavelet transform, which is being increasingly recognized in image/video compression standards, in hardware is highly area-consuming. In this paper, a new high-performance lifting-based architecture with optimized error vs. hardware cost is proposed for the 9-7 DWT. In the proposed architecture each constant coefficient multiplier of the conventional lifting structure is split into two new constant multipliers in order to minimize the hardware implementation cost and quantization error. Using an optimization process the appropriate coefficients are determined according to the hardware cost and quality requirements of each application. Simulation results indicate an average quality improvement of 13.5 dB with the same hardware resources. For achieving the same quality, it requires 40% less hardware resources, which makes it suitable for embedded systems.
Alireza Aminlou, Fatemeh Refan, Mahmoud Reza Hashemi, Omid Fatemi
ICIP (6)3
2006 An Efficient Deblocking Filter with Self-Transposing Memory Architecture For H.264/AVC
abstract
One of the main reasons behind the superior efficiency of the H.264/AVC video coding standard is the use of an in-loop deblocking filter. Since the deblocking filter is computation and data intensive, it has a profound impact on the speed degradation of both encoding and decoding processes. In this paper, we propose an efficient deblocking filter architecture that can be used as an IP core either in the dedicated or platform-based H.264/AVC codec systems. Novel self-transposing memory unit is used in this paper to alleviate switching between the horizontal and vertical filtering modes. Moreover, to reduce the processing latency, a two-stage pipelined architecture is designed for 1-D filter that produces output data after 2 clock cycles. With a clock of 100 MHz the proposed design is able to process a 1280times1024 (4:2:0) video at 25 frame/second. The proposed architecture offers 33% to 56% performance improvement compared to the existing state-of-the-art architectures
Mahdi Nazm Bojnordi, Omid Fatemi, Mahmoud Reza Hashemi
ICASSP (2)3
2006 A Non-Iterative R-D Optimization Algorithm for Rate-Constraint Problems
abstract
R-D optimization algorithm is frequently used where subband coding or vector quantization is required. All existing R-D optimization algorithms have an iterative process which results in more computational complexity and execution time. In this paper we propose a novel R-D optimization algorithm that has a non-iterative process with lower computational complexity. This algorithm is based on exponential modeling of R-D curves and can be used in rate-constraint problems. The proposed algorithm presents a good performance with non-convex curves as well as convex ones. While the execution time of the existing algorithms is O(Ncrv×Npt), the execution time of the proposed algorithm is O(Ncrv), where Ncrvis the number of curves and Nptis the average number of points in each curve. The quality degradation is 0.32 dB, in average, when it is tested in the rate control component of a JPEG2000 encoder.
Alireza Aminlou, Omid Fatemi, Maryam Homayouni, Mahmoud Reza Hashemi
ICIP4
2006 A New Multi-Layered Coding Sequence for JPEG2000 with Reduced Memory Requirement
abstract
The order and arrangement of the main components in a JPEG2000 encoder, referred to as coding sequence hereafter, plays a significant role in its performance and implementation cost. Typical JPEG2000 encoders, which may use a pre-compression or a post-compression bit allocation, require a large amount of memory to store the wavelet coefficients, compressed data and R-D curves. In this paper, we propose a novel coding sequence with a pre-compression bit-allocation method that requires just a small portion of data in order to generate the final bit-stream. The proposed method is based on multi-layer coding and can be used in distortion-constrained applications of JPEG2000. Using this coding sequence, the memory requirements of a JPEG2000 system is reduced more than 50% while the quality degradation is only 0.4 dB in average.
Alireza Aminlou, Amir Naghdinezhad, Omid Fatemi, Mahmoud Reza Hashemi
ICIP4
2006 A Novel Fade Detection Algorithm on H.264/AVC Compressed Domain
Bita Damghanian, Mahmoud Reza Hashemi, Mohammad Kazem Akbari
PSIVT2
1995 Persian cursive script recognition
abstract
The main objective of this paper is to design a Persian text recognition system. As even typed Persian scripts are cursive, our system includes a segmentation stage in order to separate the constituent characters. This stage is also useful for highly declined or italic Latin texts. A new segmentation algorithm with two consecutive steps is introduced in this paper. The first step separates isolated and non overlapped characters as well as some overlapped ones. The second step segments not connected overlapped characters. The novel segmentation method has been tested on some real world script and has shown an accuracy rate of more than 99.7%. In the recognition stage which involves a statistical approach, two types of feature sets along with different classification methods are evaluated.
Mahmoud Reza Hashemi, Omid Fatemi, Reza Safavi
ICDAR1