Alireza Aminlou

dblp:00/2422 · DBLP profile ↗
← Back
40ranked-venue papers
15as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 14 first-author · 8 since 2021Systems, architecture and hardware · 1Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Task Enhancement Tiles for Ultra Lightweight Post-processing in Visual Coding for Machines
abstract
The proliferation of automated visual analysis calls for compression methods tailored to the unique requirements of Video Coding for Machines (VCM). In this paper, we propose a computationally lightweight post-processing method that is based on a learned component referred to as a task enhancement tile (TET). A TET is spatially tiled over the reconstructed visual data and added to it element-wise. It only requires one addition per pixel in each color channel before the machine task can be applied. Our results with the VVC test model (VTM) demonstrate coding gains of up to 39.0% for object detection and 29.2% for instance segmentation on image datasets, while evaluation on a video dataset shows gains of up to 35.2% for object detection, relative to the VTM anchor. The proposed solution also offers extremely low computational cost, preservation of human-viewable content, full compliance with video coding standards, no requirement for side information transmission from encoder to decoder, and generalization across tasks, models, and encoding parameters.
Tero Partanen, Alban Marie, Rudolf Kortelahti, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
PCS8
2025 Convolutional Cross-Component Models for Chroma Prediction in Video Coding
abstract
In this paper we present two novel approaches for improving intra and inter chroma prediction in video coding. Our research demonstrates that treating the cross-component predictor as a two-dimensional convolutional model can significantly enhance chroma prediction performance. The proposed two convolutional models incorporate multiple spatial neighbors, a bias term, and a nonlinear term. For intra-coded blocks, we derive the model coefficients on the reconstructed neighborhood of the block, while for inter-coded blocks, the model coefficients are determined using prediction samples. To evaluate our methods, we implemented them on top of the ECM software that is currently under exploration by the ITU-T/ISO/IEC Joint Video Experts Team. Our intra cross-component predictor achieves BD-rate savings of {−1.47%, −2.90%, −3.02%}, {−0.92%, −2.04%, −2.32%} (Y, U, V) for the all intra and the random access configurations over ECM-5.0, respectively. Our inter cross-component predictor achieves BD-rate savings of {−0.09%, −1.25%, −1.46%}, {−0.04%, −3.42%, −3.85%} for the random access and the low-delay B configurations over ECM-9.0, respectively. Both proposed methods have been adopted into the ECM software.
Pekka Astola, Alireza Aminlou, Ramin Ghaznavi Youvalari, Jani Lainema
IEEE Trans. Circuits Syst. Video Technol.2
2024 Feasibility Study of Multi-Layer VVC Coding Scheme for Hybrid Machine-Human Consumption
abstract
The proliferation of machine vision applications necessitates developing more efficient visual data compression schemes for machine consumption. However, numerous automated use cases still require keeping humans in the loop, leading to the need for a machine-optimized video streaming with the option for human supervision. This paper investigates the feasibility of using the multi-layer coding approach of the emerging Versatile Video Coding (VVC) standard to create favorable conditions for hybrid machine-human consumption. We introduce a multi-layer coding scheme, where the base layer (BL) is optimized for machines and the enhancement layer (EL) complements the stream for human vision. Our results demonstrate that the bitrate of the proposed multi-layer stream (BL + EL) is, on average, 11% higher than that of a single-layer VVC. However, the more compact BL yields overall bandwidth savings as long as the EL is required less than 80% of the time.
Jaakko Laitinen, Tero Partanen, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
ICME7
2024 Luma Range Scaling for Enhanced VVC Efficiency in Video Coding for Machines
abstract
Recent years have shown significant growth in video data traffic for machine vision applications, catalyzing new standardization efforts in video coding for machines (VCM). These activities focus on compressing images and videos for machine vision tasks, rather than for human viewing. In this work, we propose a novel method that scales down the luma range to enhance the coding efficiency of Versatile Video Coding (VVC) for machine consumption. This method results in a lower bitrate after encoding and has only minimal adverse effects on the accuracy of machine vision tasks. In our experiments, we down-scale the luma channel of the input video using luma-scaling factors from 0.2 to 0.9 and evaluate coding results with optional back-scaling to the original range before machine vision tasks. Our results with the VVC Test Model (VTM) demonstrate that the proposed technique achieves coding gain of up to 37.9%and 46.1% for the same object detection and tracking accuracy, respectively.
Tero Partanen, Alban Marie, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
MMSP7
2023 Optimal Tile Size and Streaming Field of View for VR Streaming
abstract
Virtual reality (VR) video services require a high bitrate, and hence, viewport-adaptive streaming techniques like motion-constrained-tile-set (MCTS) have been found important to reduce streaming-rate and storage demands. The tiling scheme and streaming field of view (FOV) are among the key elements in designing an optimal VR viewport-adaptive streaming solution, in terms of rate-distortion (R-D) performance. The aim of this study is to propose an optimal configuration for the tile grid and streaming FOV, considering different VR viewing situations such as head motion speed, system delay, and head-mounted display FOV. To achieve this, a wide range of tiling schemes and streaming FOVs are examined to study the storage and streaming R-D performance of the MCTS-based technique in both viewport and non-viewport areas using a quality metric called Zonal-cubic PSNR. The findings demonstrate that for VR applications focused on preserving high viewport quality, fine tile grids lead to higher performance. In scenarios featuring small and large HMD FOV, the optimal configuration involves a small and medium streaming FOV, respectively.
Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
MMSP2
2022 Comparison of Boundary Artifact Removal Methods in Coding of Generalized Cubemap Projection Using VVC
abstract
Virtual reality applications use 360-degree videos and head mount displays with stereoscopic capabilities to provide full immersion experience. Among existing projection format, Cubemap projection provides and improved compression performance, but similar to other projects, it suffers from visual artifacts in the rendered viewport because of discontinuity of the content in different regions. In this work, we investigated the effect of different methods in Versatile Video Coding standard and other pre- and post-processing algorithms for removing boundary artifacts by introducing a new objective quality metric for systematic comparison. We further improve the current methods by aligning the GCMP’s face boundaries to coding unit boundaries. The observation is that the combination of these method with existing methods in VVC offers the best result.
Kianoush Jafari, Alireza Aminlou, Miska M. Hannuksela
ICASSP2
2021 VVC Adaptive Loop Filter Optimization for Subpicture-based Viewport-adaptive Streaming
abstract
Virtual reality (VR) systems require delivering high-fidelity 360° video content to immerse viewers to the captured scene. The viewport-adaptive streaming (VAS) methods have been developed to deliver 360° VR content efficiently. The Versatile Video Coding (VVC) standard introduces the subpicture picture partitioning tool, which creates isolated regions suitable for VAS. The usage of Adaptive Loop Filter (ALF) as a VVC in-loop filtering operation is limited in subpicture-based VAS. This paper aims at enabling usage of ALF in subpicture-base VAS through proposing a set of encoding constraints that are standard compliant. While ALF is activated, the proposed constrains guarantee that no coding coordination with respect to sharing of ALF parameters among subpictures is required. This further allows subpicture-based parallel encoding of high-resolution VR content. We study the performance of several methods targeting both single- and multi-thread encoding platforms. The experimental results indicate that VR content encoding can be parallelized at a subpicture-group level, while still preserving most of the ALF gain. The proposed method with subpicture-group encoding parallelization achieves on average -2.4%, -4.0%, and -4.3% Bjøntegaard delta rate reduction for Y, U, and V components respectively, compared to the case where ALF operation is deactivated.
Alireza Zare, Alireza Aminlou, Miska M. Hannuksela
MMSP2
2021 Regression-Based Motion Vector Field for Video Coding
abstract
In this paper, we study a method for compensating the non-translational motion behavior in video coding. The proposed method models the motion field of a prediction block based on the motion information of the neighboring blocks by using a linear regression approach. In order to provide a finer granularity of motion vectors the Regression-based Motion Vector Field (RMVF) method derives the motion field in 4×4 sub-block accuracy. Such approach generates a smooth and more realistic motion vector field inside the prediction block. The motion field generated with RMVF is then used as a new merge mode along with other merge modes in VTM-2.0 test model of the Versatile Video Coding (H.266/VVC) standard. The conducted experiments with JVET CTC sequences illustrate that the proposed RMVF method provides 0.77%, 0.19% and 0.41% bitrate reductions with random access (RA), low delay B (LDB) and low delay P (LDP) configurations, respectively. Furthermore, this method provides on average 0.65% bitrate saving for the 360° sequences in equirectangular projection format (ERP) with RA configuration.
Ramin Ghaznavi Youvalari, Alireza Aminlou, Jani Lainema
IEEE Trans. Circuits Syst. Video Technol.2
2020 On Subpicture-based Viewport-dependent 360-degree Video Streaming using VVC
abstract
Virtual reality applications create an immersive experience using 360° video with high resolution and frame rate. However, since the user only views a portion of 360° video according to his/her current viewport, streaming the whole content with high resolution causes bandwidth wastage. To address this issue, viewport-dependent approaches have been proposed such that only the part of the video which falls within user's current viewport is transmitted in high quality while the rest of the content is transmitted in lower quality. The selection of high- and low-quality parts is constantly adapted according to the user's head motion, which requires frequent intra coded frames at switching points, leading to an increment in the overall streaming bitrate. In this paper a viewport-adaptive streaming scheme is introduced, which avoids intra frames at switching points by introducing long intra period for non-changing parts of the content during head motion. This scheme has been realized taking advantage of mixed Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit feature of Versatile Video Coding (VVC) standard. This method reduces bitrate significantly, especially for the sequences with either no or only slow camera motion, which is common for 360° video capturing.
Maryam Homayouni, Alireza Aminlou, Miska M. Hannuksela
ISM2
2019 Shared Coded Picture Technique for Tile-Based Viewport-Adaptive Streaming of Omnidirectional Video
abstract
Tile-based viewport-adaptive streaming methods have been used in delivering omnidirectional video for virtual reality applications. In these methods, the 360° video is encoded in multiple quality versions by using the motion constrained tile set (MCTS) technique. A set of high-quality and low-quality tiles, corresponding to viewport and non-viewport areas, respectively, are selected and transmitted to the user. However, these methods require frequent intra random access points to ensure seamless viewport switching capability, very high decoding complexity, or a multi-layer coding scheme. The frequent intra random access points include very high bitrate in viewport switching points. The high decoding complexity and multi-layer decoder requirements are not aligned with the omnidirectional media format (OMAF) standard. Such requirements make these methods sub-optimal or impractical for streaming the omnidirectional video. This paper studies the current tile-based solutions for delivering the omnidirectional content. Moreover, the OMAF-compliant shared coded picture (SCP)-based scheme is proposed in this paper for streaming the omnidirectional video. The core concept of the SCP-based method is to manipulate the switching point pictures in a way that the frequent intra-coded pictures are no longer required for the viewport switching operations between different quality versions of the content. The experiments illustrated that the SCP-based method outperforms the MCTS-based method on average by 11% to 14% in terms of streaming bitrate reduction with only 4% extra decoding complexity.
Ramin Ghaznavi Youvalari, Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.3
2019 6K and 8K Effective Resolution with 4K HEVC Decoding Capability for 360 Video Streaming
abstract
The recent Omnidirectional MediA Format (OMAF) standard, which specifies the delivery of 360° video content, supports only equirectangular projection (ERP) and cubemap projection and their region-wise packing with a limitation on video decoding capability to the maximum resolution of 4K (e.g., 4,096 × 2,048). Streaming of 4K ERP content allows only a limited viewport resolution, which is lower than the resolution of many current head-mounted displays (HMDs). Therefore, to take full advantage of high-resolution HMDs, delivery of 360° video content beyond 4K resolution needs to be enabled. In this regard, we propose two specific mixed-resolution packing schemes of 6K (e.g., 6,144 × 3,072) and 8K (e.g., 8,192 × 4,096) ERP content and their realization in tile-based streaming, while complying with the 4K decoding constraint and the High Efficiency Video Coding standard. The proposed packing schemes offer 6K and 8K effective resolution at the viewport. Using our proposed test methodology, experimental results indicate that the proposed layouts significantly decrease streaming bitrates when compared to mixed-quality viewport-adaptive streaming of 4K ERP. Our results further indicate that 8K-effective packing outperforms 6K-effective packing especially in high-quality videos.
Alireza Zare, Maryam Homayouni, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ACM Trans. Multim. Comput. Commun. Appl.3
2018 Geometry-Based Motion Vector Scaling for Omnidirectional Video Coding
abstract
Virtual reality (VR) applications make use of 360° omnidirectional video content for creating immersive experience to the user. In order to utilize current 2D video compression standards, such content must be projected onto a 2D image plane. However, the projection from spherical to 2D domain introduces deformations in the projected content due to the different sampling characteristics of the 2D plane. Such deformations are not favorable for the motion models of the current video coding standards. Consequently, omnidirectional video is not efficiently compressible with current codecs. In this work, a geometry-based motion vector scaling method is proposed in order to compress the motion information of omnidirectional content efficiently. The proposed method applies a scaling technique, based on the location in the 360° video, to the motion information of the neighboring blocks in order to provide a uniform motion behavior in a certain part of the content. The uniform motion behavior provides optimal candidates for efficiently predicting the motion vectors of the current block. The conducted experiments illustrated that the proposed method provides up to 2.2% bitrate reduction and on average around 1% bitrate reduction for the content with high motion characteristics in the VTM test model of Versatile Video Coding (H.266/VVC) standard.
Ramin Ghaznavi Youvalari, Alireza Aminlou
ISM2
2018 Adaptive Motion Vector Prediction for Omnidirectional Video
abstract
Omnidirectional video is widely used in virtual reality applications in order to create the immersive experience to the user. Such content is projected onto a 2D image plane in order to make it suitable for compression purposes by using current standard codecs. However, the resulted projected video contains deformations mainly due to the oversampling of the projection plane. These deformations are not favorable for the motion models that are used in the recent video compression standards. Hence, omnidirectional video is not efficiently compressible with the current codecs. In this work, an adaptive motion vector prediction method is proposed for efficiently coding the motion information of such content. The proposed method adaptively models the motion vectors of the coding block based on the motion information of the neighboring blocks and calculates a more optimal motion vector predictor for coding the motion information. The experimented results showed that the proposed motion vector prediction method provides up to 2.2% bitrate reduction in the content with high motion and on average 1.1% bitrate reduction for the tested sequences.
Ramin Ghaznavi Youvalari, Alireza Aminlou
VCIP2
2017 Virtual reality content streaming: Viewport-dependent projection and tile-based techniques
abstract
Virtual reality (VR) head-mounted display (HMD) requires spherical panoramic contents with high-spatial and temporal fidelity to immerse the viewers into the captured scene. Hereby, VR contents are extremely bandwidth intensive and impose technical challenges for the design of a VR streaming system. A bandwidth-efficient VR streaming system can be achieved using the viewport-aware adaptation techniques, in which part of the sphere within the viewer's field of view is presented at higher quality. In this paper, two recently emerged viewport-adaptive streaming methods so-called tile-based method and truncated square pyramid (TSP) projection, a well-studied viewport-dependent projection, are compared using a proposed quality assessment methodology. The comparison is made in terms of storage and streaming bitrate performances. The simulation results indicate that the tile-based approach has slightly lower streaming performance, while offering a significant storage and encoding time saving at the server side, compared to TSP-based streaming.
Alireza Zare, Alireza Aminlou, Miska M. Hannuksela
ICIP2
2017 Comparison of HEVC coding schemes for tile-based viewport-adaptive streaming of omnidirectional video
abstract
Virtual reality applications make use of 360-degree panoramic or omnidirectional video with high resolution and high frame rate in order to create the immersive experience to the user. The user views only a portion of the captured 360-degree scene at each time instant, hence streaming the whole omnidirectional video in highest quality is not efficient. In order to alleviate the problem of bandwidth wastage, viewport-adaptive encoding and streaming schemes have been proposed. In these schemes, part of the captured scene that is within the viewer's field of view is delivered at highest quality while the rest of the scene in a lower quality. In this work, three tile-based viewport-adaptive methods using motion-constrained tile sets (MCTS), region-of-interest scalability and simulcast approach have been studied for streaming omnidirectional content. In the performed experiments with various tiling arrangements, MCTS-based scheme required highest bitrate compared to other methods. The scalable coding scheme provided the highest performance in terms of streaming bitrate saving on average up to 53% and 35% compared to streaming the whole omnidirectional video and MCTS-based method, respectively.
Ramin Ghaznavi Youvalari, Alireza Zare, Huameng Fang, Alireza Aminlou, Qingpeng Xie, Miska M. Hannuksela, Moncef Gabbouj
MMSP4
2016 Standard-Compliant Multiview Video Coding and Streaming for Virtual Reality Applications
abstract
Virtual reality (VR) systems employ multiview cameras or camera rigs to capture a scene from the entire 360-degree perspective. Due to computational or latency constraints, it might not be possible to stitch multiview videos into a single video sequence prior to encoding. In this paper we investigate the coding and streaming of multiview VR video content. We present a standard-compliant method where we first divide the camera views into two types: Primary views represent a subset of camera views with lower resolution and non-overlapping (minimally overlapping) content which cover the entire 360-degree field-of-view to guarantee immediate monoscopic viewing during very rapid head movements. Auxiliary views consist of remaining camera views with higher resolution which produce overlapping content with the primary views and are additionally used for stereoscopic viewing. Based on this categorization, we propose a coding arrangement in which, the primary views are independently coded in the base layer and the additional auxiliary views are coded as an enhancement layer, using inter-layer prediction from primary views. The proposed system not only meets the low latency requirements of VR systems, but also conforms to the existing multilayer extensions of the High Efficiency Video Coding standard. Simulation results show that the coding and streaming performance of the proposed scheme is significantly improved compared to earlier methods.
Kashyap Kammachi Sreedhar, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ISM2
2016 Viewport-Adaptive Encoding and Streaming of 360-Degree Video for Virtual Reality Applications
abstract
Virtual reality applications use 360-degree videos and head mount displays (HMDs) with stereoscopic capabilities to provide full immersion experience. In these applications it is also common to use 4K resolution or higher per view for 360-degree videos. Consequently, this leads to technical challenges in handling the bandwidth requirements while keeping the system latency to the minimal. When the content is viewed with a HMD, a subset of the entire 360-degree video is displayed at a single point of time. To improve the resolution and picture quality of the displayed content, viewport based coding is desirable. In this regard, we investigated various viewport dependent projection schemes including the existing variants of Pyramidal projection. In this regard we propose the multi-resolution versions of Equirectangular and Cubemap projections. Additionally, we developed a methodology for comparing the rate-distortion performance of these projections. Based on the simulation results, it was observed that multi-resolution projections of Equirectangle and Cubemap outperform other projection schemes, significantly.
Kashyap Kammachi Sreedhar, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ISM2
2016 Efficient Coding of 360-Degree Pseudo-Cylindrical Panoramic Video for Virtual Reality Applications
abstract
Pseudo-cylindrical panoramas represent the data distribution of spherical coordinates closely in two-dimensional domain due to the equidistant sampling of 360-degree scene. Therefore, unlike the cylindrical projections, they do not suffer from the over stretching in the polar areas. However, due to the non-rectangular format in effective picture area and sharp edges at its borders, the compression performance is inefficient. In this paper, we propose two methods which improve the compression performance of both intra-frame and inter-frame coding of pseudo-cylindrical panoramic content and meanwhile reduce the coding artifacts. In the intra-frame coding method, border edges are smoothed by modifying the content of the image in the non-effective picture area, which are cropped at the receiver side. In the inter-frame coding method, gaining the benefit of 360-degree property of the content, non-effective picture area of reference frames at border is filled with the content of the effective picture area from the opposite border to enhance the performance of motion compensation.
Ramin Ghaznavi Youvalari, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ISM2
2016 HEVC-compliant Tile-based Streaming of Panoramic Video for Virtual Reality Applications
abstract
Delivering wide-angle and high-resolution spherical panoramic video content entails a high streaming bitrate. This imposes challenges when panorama clips are consumed in virtual reality (VR) head-mounted displays (HMD). The reason is that the HMDs typically require high spatial and temporal fidelity contents and strict low-latency in order to guarantee the user's sense of presence while using them. In order to alleviate the problem, we propose to store two versions of the same video content at different resolutions, each divided into multiple tiles using the High Efficiency Video Coding (HEVC) standard. According to the user's present viewport, a set of tiles is transmitted in the highest captured resolution, while the remaining parts are transmitted from the low-resolution version of the same content. In order to enable randomly choosing different combinations, the tile sets are encoded to be independently decodable. We further study the trade-off in the choice of tiling scheme and its impact on compression and streaming bitrate performances. The results indicate streaming bitrate saving from 30% to 40%, depending on the selected tiling scheme, when compared to streaming the entire video content.
Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
ACM Multimedia2
2016 Analysis of regional down-sampling methods for coding of omnidirectional video
abstract
In order to compress omnidirectional video clips, a projection onto a two-dimensional image plane is necessary. The most commonly used projection format is the equirectangular panoramic projection, which results into a significant amount of redundant samples in the polar areas. The redundant samples incur extra bitrate and increase the encoding/decoding time. In this paper, we study regional down-sampling (RDS) for achieving better compression and smaller encoding/decoding time for omnidirectional content. We extend the persistent RDS method applied equally to all pictures to be applied to selected pictures only in our proposed temporal RDS method and then compare the persistent and temporal RDS methods. The simulation results indicate that both the persistent and temporal RDS improve the rate-distortion (RD) performance compared to the conventional coding of equirectangular panoramas, while the temporal RDS method has less sequence-wise RD performance variation and slightly better RD performance on average when compared to the persistent RDS technique. Alongside the coding methods, we study spherical quality measurement methods for VR images/video and analyze the coding methods with these quality metrics. Moreover, we propose a uniformly sampled spherical quality metric in order to evaluate the coding distortion of omnidirectional videos.
Ramin Ghaznavi Youvalari, Alireza Aminlou, Miska M. Hannuksela
PCS2
2016 HEVC-compliant viewport-adaptive streaming of stereoscopic panoramic video
abstract
Virtual reality (VR) provides unprecedented immersive experience using high-resolution spherical stereoscopic panoramic video. Such an experience is achieved by using head-mounted display (HMD) which has very strict latency bounds in order to respond promptly to user movements. Conventional streaming of VR video requires large bandwidth because the entire captured panorama is transmitted. However, only a limited field-of-view (FOV) is displayed by an HMD, resulting in wastage of bandwidth. To alleviate the problem, this paper proposes a High Efficiency Video Coding (HEVC) compliant approach for efficient coding and streaming of stereoscopic VR content. The proposed method is based on partitioning video pictures into tiles, where only the required tiles corresponding to the primary viewport are transmitted in high resolution, while the remaining parts are transmitted in low resolution. Furthermore, this method enables coding stereoscopic video contents using a conventional HEVC codec, while still achieving significant compression gain by means of adopting inter-view prediction only in intra random access point (IRAP) pictures. Using this method, the predicted view can be decoded independently of the main view, hence allowing simultaneous decoding instances. Experimental results demonstrate that the proposed approach is able to substantially improve compression efficiency and streaming bitrate performance.
Alireza Zare, Kashyap Kammachi Sreedhar, Vinod Kumar Malamal Vadakital, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj
PCS4
2016 New R-D Optimization Criterion for Fast Mode Decision Algorithms in Video Coding and Transrating
abstract
Mode decision has a significant effect on the quality and complexity of video coding. It is even more challenging when generating multiple bitstreams with different bitrates (BRs) in, for example, dynamic adaptive streaming for HTTP or transrating systems. Full search and simplified fast search mode decision (MD) methods either suffer from a high computational complexity or have a negative impact on quality. Furthermore, mode selection in conventional approaches strongly depends on the quantization parameter (QP). Hence, modes that have been selected for high BR compression may not be suitable for low BR when transrating a bitstream. In this paper, we propose a rate-distortion (R-D)-optimized criterion for fast MD algorithms. The proposed cost function, when adopted in different fast MD algorithms, not only improves the R-D performance by up to 6.6% in terms of Bjøntegaard delta rate, but also reduces the execution time of the encoder by up to 6.8%. We also show that modes selected by the proposed criterion are less sensitive to changes in BR or QP. As a result, the same modes in an encoded bitstream may be used even after transrating using requantization, resulting in a significant R-D performance improvement of up to 33.3%.
Alireza Aminlou, Mahmoud Reza Hashemi, Moncef Gabbouj, Bing Zeng 0001, Omid Fatemi
IEEE Trans. Circuits Syst. Video Technol.1
2014 Weighted-prediction-based color gamut scalability extension for the H.265/HEVC video codec
abstract
Color gamut scalability refers to coding a video in a layered manner where the base and enhancement layers are coded in different color gamut spaces. Color gamut scalability and its relationship with spatial scalability are currently being studied for the scalable extension of HEVC (SHVC) to enable coding of ultra-high definition content having BT.2020 color gamut with 10-bit precision as an enhancement layer and high definition content having BT.709 color gamut with 8-bit precision as the base layer. In this paper, we propose to use the weighted prediction tool of the SHVC standard to map the color gamut of the base layer to the enhancement layer. In addition, we also propose a high-precision bit-depth mapping of the base layer to the enhancement layer that jointly performs upsampling with a bit-depth increase. Simulation results show that these two schemes improve the coding efficiency of the All Intra and Random Access configurations by about 6.8% and 3.6% on average, respectively, compared to a basic scheme where the bit-depth of the base layer is increased by simple bit-shifting. These gains are achieved by imposing no changes to the SHVC standard; hence make the proposed method very useful for practical use-cases as well.
Alireza Aminlou, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj
ICASSP1
2014 Content adaptive depth map resampling scheme in multiview video plus depth
abstract
In this paper, we propose a content-adaptive pre-processing for downsampling depth map for 3D multiview plus depth video coding. The proposed scheme takes advantage of spatially varying content of depth maps and targets improved encoding performance. The key idea is to divide the video frames into several horizontal and vertical stripes and downsample each of them with an appropriate downsampling ratio. With this approach, the rectangular shape of the frames are preserved, hence the video sequence can be coded with a standard encoder. The downsampling ratios, defined for each Intra period, are determined via a Rate Distortion Optimization process. Simulation results show that the proposed method outperforms the reference in which no downsampling is used by 0.35dB of Bjontegaard delta Peak Signal-to-Noise Ratio (PSNR) and brings up to 31% and average of 18% of Bjontegaard delta bitrate reduction (dBR).
Maryam Homayouni, Alireza Aminlou, Payman Aflaki, Moncef Gabbouj
ISCAS2
2014 Improved weighted prediction based color gamut scalability in SHVC
abstract
One use case that the scalable extension (SHVC) of the state-of-the-art High Efficiency Video Coding (HEVC) standard aims for is to support Ultra High Definition (UHD) TV broadcast in a backwards compatible way with the existing High Definition (HD) TV broadcast. However, since UHD content typically has higher bit-depth and wider color gamut in addition to increased spatial resolution, the compression efficiency is highly affected by the inter-layer processing applied on the base layer picture. This paper proposes an improvement for the weighted prediction based color gamut scalability to have a better mapping between the color gamuts of the base and enhancement layers. The proposed method aims at capturing the nonlinear characteristics of the color gamut mapping using a piecewise linear model, whose parameters are signaled through weighted prediction mechanism and multiple inter-layer reference pictures. Compared to other existing methods for color gamut mapping in SHVC, such as the 3D Look Up Table (LUT) method, the proposed weighted prediction based approach is less complex, as it does not require any changes to the decoder. The simulation results show up to 3.8% Bjontegaard delta bitrate gain in luma for all intra and 3.0% for random access configurations compared to the existing weighted prediction based scalability method in SHVC.
Döne Bugdayci Sansli, Alireza Aminlou, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj
VCIP2
2014 Differential Coding Using Enhanced Inter-Layer Reference Picture for the Scalable Extension of H.265/HEVC Video Codec
abstract
Differential coding methods improve coding efficiency of scalable video codecs by adding the high-frequency component present in the previously coded enhancement layer (EL) pictures to the base layer (BL) picture. This paper proposes a method to enable differential coding in a scalable codec design without affecting the core coding tools, thus allowing a practical implementation to reuse single-layer hardware or software components. This is achieved by creating an additional reference picture called enhanced inter-layer reference (EILR) and inserting it to the EL decoded picture buffer and reference picture lists. An EILR picture is generated by adding differential information to the current inter-layer reference picture. The differential information is calculated using the previously decoded pictures of the BL and EL and the motion information of the BL picture. The proposed method reduces luma total bitrate on average by 2.2% and 2.8% for random access and low-delay test cases, respectively. The improvements are more significant for chroma components with the average bitrate reduction of 6.5%. The measured decoding time increase for a reference software implementation is 16% with negligible overhead on encoding time.
Alireza Aminlou, Jani Lainema, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.1
2012 A Two-Piece R-D Model for Hybrid Video Coding and Its Application in Fast Mode Decision
abstract
The mode decision process has a significant effect on the quality and complexity of a video encoder. The conventional method that fully codes each macro block for different modes results in the best quality performance, but it suffers from high computational complexity. On the other hand, some other methods ignore the residual part and use the prediction data, or adopt early mode selection approaches in order to reduce the computational cost. These approaches have a negative impact on the coding performance. In this paper, we have used a simple model for the residual coding part and proposed a two-piece R-D model for a macro block. Based on this model, we have introduced a mode decision algorithm that reduces the bit-rate by up to 11.62% at the expense of just 0.5% computational overhead.
Alireza Aminlou, Hana Fahim-Hashemi, Mahmoud Reza Hashemi, Moncef Gabbouj, Omid Fatemi
ICME1
2011 Rate-distortion-complexity optimization for VLSI implementation of integer motion estimation in H.264/AVC encoder
abstract
In order to accommodate the wide range of applications and the corresponding platforms where the H.264/AVC standard is currently in place, one should be able to optimize the encoder's computational complexity with a careful selection of the coding configuration parameters. Motion estimation is the most time-consuming part of the encoder which constitutes up to 75% of the computational complexity. In this paper, the optimum selection of configuration parameters, including search range, reference frame, degree of down-sampling and number of truncation bits have been analyzed for the VLSI implementation of integer motion estimation in terms of distortion-complexity performance. Furthermore, the optimum parameter sets have been presented for different video sizes and different constraints on computational power.
Alireza Aminlou, Zahra NajafiHaghi, Majid Namaki-Shoushtari, Mahmoud Reza Hashemi
ICME1
2011 Edge-oriented interpolation for fractional motion estimation in hybrid video coding
abstract
Fractional motion estimation, using the interpolation process, improves the quality of the compressed video by about 2dB in terms of PSNR over integer motion estimation. Most existing interpolation techniques in recent video coding standards, such as the symmetric 6-tap filter of the H.264/AVC standard, do not perform well around the object edges. This increases the value of residuals, which in turn results in higher bit rate. In this paper, a new interpolation method has been proposed that considers the edges of video objects. Simulation results indicate that the proposed method, when used in the H.264/AVC encoder, improves PSNR by up to 1.0 dB (0.4 dB in average) with respect to the standard interpolation technique. This is achieved at the expense of up to 6% (3% in average) increase in the computational complexity.
Ali Kokhazadeh, Alireza Aminlou, Mahmoud Reza Hashemi
ICME2
2009 A cost-error optimized architecture for 9/7 lifting based Discrete Wavelet Transform with balanced pipeline stages
abstract
Discrete wavelet transform (DWT) is increasingly recognized in image/video compression standards, as indicated by its use in JPEG2000. The lifting scheme algorithm is an alternative DWT implementation that has a lower computational complexity. In this paper, a new high performance lifting-based architecture is presented for the 9/7 DWT engine. The proposed architecture has a balanced pipeline and improves both the computational error and hardware complexity for any given working frequency. In the proposed architecture, the constant coefficients are modified by introducing new variables to the conventional lifting structure to minimize hardware cost and computational error, imposed by quantization of coefficients. Simulation results indicate a quality improvement of up to 15 dB when compared to an architecture using the standard coefficients that has the same hardware cost and working frequency. Similarly, the hardware cost is reduced by about 20% when both architectures deliver the same PSNR when operating at the same frequency.
Alireza Aminlou, Fatemeh Refan, Mahmoud Reza Hashemi, Omid Fatemi, Saeed Safari
ICASSP1
2009 Unequal loss-protected multiple description coding of scalable source streams using a progressive approach
abstract
An analysis-based approach for unequal loss-protected multiple description coding (packetization) of the scalable (prioritized / progressive) source code streams is proposed. For a given number of packets (descriptions) of the known size, unequal loss-protected packetization leads to segment the scalable code stream, such that the source can be reconstructed with the maximum possible fidelity at the decoder side. Here, we find an analytical relation between optimal sizes of any two consecutive segments. This idea yields a low-complexity progressive solution with a performance close to that of local search, which has been approved as an efficient method to solve the segmentation problem. Simulation results are used to confirm the efficiency of the proposed method as compared with the local search algorithm.
Majid Roohollah Ardestani, Alireza Aminlou, Asghar Beheshti 0001
ICIP2
2009 Low complexity hardware implementation of reciprocal fractional motion estimation for H.264/AVC in mobile applications
abstract
Motion estimation, one of the most effective modules in H.264/AVC, constitutes 60%-90% of encoding time and computation. Close to %45 of this computation belongs to Fractional Motion Estimation (FME) which has to perform a time consuming half-pixel and quarter-pixel interpolation. In addition, interpolation is a major challenge for hardware implementation in real time, specially in mobile applications where processing and battery power is limited. Several modified sub-pixel accuracy search methods have been proposed in the literature in order to reduce the complexity of interpolation. The reciprocal method has been proposed to reduce the CPU encoding time for a software implementation of an H.264 encoder on personal computers. In this paper, the hardware implementation of the reciprocal method is evaluated and its PSNR performance and hardware cost are compared to that of a simplified Rate-Distortion Optimization (RDO) process. Simulation results indicate that using reciprocal FME with enabled RDO has less computational cost and better PSNR performance than using the conventional FME with disabled RDO.
Alireza Aminlou, Parviz Alvandi, Mahmoud Reza Hashemi
PCS1
2009 An improved R-D optimized motion estimation method for video coding
abstract
Motion estimation is one of the key tools to achieve a very low bit rate in video coding. The selection of optimum motion vectors (MV) has a significant impact on the quality of the resulted compressed video, in Rate-Distortion (R-D) sense. The established method uses the Lagrange multiplier to optimally select the MV for each block. However, it does not consider the effect of residual coding at the same time with the effect of MVs. In this paper, we have considered the effect of residual coding in motion vector (MV) selection with modeling motion estimation and residual coding as independent processes which results in a new optimization condition. It is based on local optimization, which is the bit allocation between motion estimation and residual coding, and global optimization, which is the bit allocation among different blocks. The proposed motion estimation algorithm results in a PSNR improvement of 0.5 - 3.0 dB when it is used in the H.264 standard with block size of 4 times 4.
Alireza Aminlou, Mojtaba Farmani, Mahmoud Reza Hashemi, Omid Fatemi
PCS1
2007 A Superior Low Complexity Rate Control Algorithm
abstract
In this paper, a new low complexity Rate-Distortion optimization algorithm has been proposed. The proposed method can be employed with non-convex curves as well as convex curves. The new technique has been used in a hardware implementation of a JPEG2000 encoder. Simulation results indicate that the proposed algorithm is less sensitive to the shape of R-D curves in comparison with current algorithms. Compared to the exact method, our performance degradation is less than 0.43 dB. The low complexity of this algorithm makes it suitable for real time applications and in applications like Digital Cinema that have to process a large number of input data.
Alireza Aminlou, Maryam Homayouni, Mohammad Hossein Neishaburi, Siamak Mohammadi
AICCSA1
2007 A Split Method for Optimized Cost-Quality Hardware Implementation of Lifting-Based Discrete Wavelet Transform
abstract
Discrete wavelet transform (DWT) is increasingly recognized in image/video compression standards, as indicated by its use in JPEG2000. The lifting scheme algorithm is an alternative DWT implementation that has a lower computational complexity. In this paper, a new high performance lifting-based architecture with optimized error vs. hardware complexity is presented for DWT. The proposed architecture modifies the constant coefficients by introducing new variables to the conventional lifting structure to minimize hardware cost and quantization error. In order to achieve the most efficient coefficients, an optimization process has been implemented. Simulation results indicate an average quality improvement of 7.5 dB with the same hardware complexity/cost. Similarly, for achieving the same quality as the conventional hardware implementations the proposed architecture is 20% less complex. The appropriate coefficients can be determined according to the cost and error requirements of each application.
Alireza Aminlou, Fatemeh Refan, Maryam Homayouni, Omid Fatemi, Mahmoud Reza Hashemi
ICASSP (2)1
2007 Pattern-Based Error Recovery of Low Resolution Subbands in JPEG2000
abstract
Digital image transmission is widely used in consumer products, such as digital cameras and cellular phones, where low bit rate coding is required. In any low bit rate encoder, such as the JPEG2000 standard, data truncation (during the encoding process), and data loss (during transmission) will result in lost bit-planes, which will be normally replaced by zeros. In this paper a new algorithm has been proposed, which recovers the lost/truncated lower bit-planes of coefficients in the LL subband of a wavelet transform in a JPEG2000 stream using the data available in higher bit-planes of the same coefficient and its eight neighbors. Simulation results indicate that the proposed algorithm achieves 5.40-8.77 dB improvement with respect to zero filling data recovery method.
Alireza Aminlou, Nasim Hajari, Hossein Badakhshannoory, Mahmoud Reza Hashemi, Omid Fatemi
ICIP (4)1
2007 Two Level Cost-Quality Optimization of 9-7 Lifting-Based Discrete Wavelet Transform
abstract
Implementing the discrete wavelet transform, which is being increasingly recognized in image/video compression standards, in hardware is highly area-consuming. In this paper, a new high-performance lifting-based architecture with optimized error vs. hardware cost is proposed for the 9-7 DWT. In the proposed architecture each constant coefficient multiplier of the conventional lifting structure is split into two new constant multipliers in order to minimize the hardware implementation cost and quantization error. Using an optimization process the appropriate coefficients are determined according to the hardware cost and quality requirements of each application. Simulation results indicate an average quality improvement of 13.5 dB with the same hardware resources. For achieving the same quality, it requires 40% less hardware resources, which makes it suitable for embedded systems.
Alireza Aminlou, Fatemeh Refan, Mahmoud Reza Hashemi, Omid Fatemi
ICIP (6)1
2006 A Non-Iterative R-D Optimization Algorithm for Rate-Constraint Problems
abstract
R-D optimization algorithm is frequently used where subband coding or vector quantization is required. All existing R-D optimization algorithms have an iterative process which results in more computational complexity and execution time. In this paper we propose a novel R-D optimization algorithm that has a non-iterative process with lower computational complexity. This algorithm is based on exponential modeling of R-D curves and can be used in rate-constraint problems. The proposed algorithm presents a good performance with non-convex curves as well as convex ones. While the execution time of the existing algorithms is O(Ncrv×Npt), the execution time of the proposed algorithm is O(Ncrv), where Ncrvis the number of curves and Nptis the average number of points in each curve. The quality degradation is 0.32 dB, in average, when it is tested in the rate control component of a JPEG2000 encoder.
Alireza Aminlou, Omid Fatemi, Maryam Homayouni, Mahmoud Reza Hashemi
ICIP1
2006 A New Multi-Layered Coding Sequence for JPEG2000 with Reduced Memory Requirement
abstract
The order and arrangement of the main components in a JPEG2000 encoder, referred to as coding sequence hereafter, plays a significant role in its performance and implementation cost. Typical JPEG2000 encoders, which may use a pre-compression or a post-compression bit allocation, require a large amount of memory to store the wavelet coefficients, compressed data and R-D curves. In this paper, we propose a novel coding sequence with a pre-compression bit-allocation method that requires just a small portion of data in order to generate the final bit-stream. The proposed method is based on multi-layer coding and can be used in distortion-constrained applications of JPEG2000. Using this coding sequence, the memory requirements of a JPEG2000 system is reduced more than 50% while the quality degradation is only 0.4 dB in average.
Alireza Aminlou, Amir Naghdinezhad, Omid Fatemi, Mahmoud Reza Hashemi
ICIP1
2005 A novel efficient rate control algorithm for hardware implementation in JPEG2000
abstract
In multimedia applications where images are extremely used, the rate control method has a significant role in image encoder performance, computational complexity and hardware implementation. We propose a simple rate control algorithm suitable for the hardware implementation of a JPEG2000 encoder with less computational complexity and area. The proposed algorithm, which is based on exponential modeling of R-D curves, employs distortion instead of slope values. Simulation results show similar performance compared to the full search method, with considerable reduction in hardware resources.
Alireza Aminlou, Omid Fatemi
ICASSP (5)1