EDBT 2026 Demo / reviewers in the wild / expert
Asadollah Shahbahrami
dblp:70/4599
· DBLP profile ↗
38ranked-venue papers
11as first author
15since 2021 · last 2026
0000-0002-5195-1688ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 9 since 2021Systems, architecture and hardware · 12 · 8 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry-biased transformer for dental caries detection in panoramic X-ray imagesabstractAbstract Panoramic X-ray images are essential for dental caries detection, yet they inherently suffer from geometric distortions that complicate accurate early diagnosis. To address this, we propose a Geometry-Biased transformer that explicitly models spherical geometry. Our approach integrates equirectangular relative position embedding, distance-based attention scoring, and equirectangular-aware attention rearrangement to significantly enhance spatial feature representation. This enables the model to effectively capture both local caries lesions and global dental arch structures despite distortions. We rigorously evaluated our model primarily on a hospital-scale dataset, achieving a high accuracy of 94.90% and an AUC of 0.9603. Furthermore, extensive comparative analyses across diverse publicly available dental image datasets demonstrate the superior generalization and competitive performance of our method. Our findings highlight that geometry-aware transformers offer a robust and automated tool, revolutionizing high-precision dental caries detection and potentially other medical imaging diagnostics. Nima Esmi, Koosha Refahi, Asadollah Shahbahrami, Negar Khosravifard, Georgi Gaydadjiev |
Neural Comput. Appl. | 3 |
| 2025 | Counting vehicles types using deep learning algorithm in video surveillance systems
Ali Reza Akoushideh, Seyed Shafiullah Sadat, Asadollah Shahbahrami |
Multim. Tools Appl. | 3 |
| 2025 | GPU optimizations to accelerate an image enhancement algorithm
Amin Daemdoost, Asadollah Shahbahrami, Hesam Noorpoor, Nima Esmi, Reza Hassanpour |
J. Supercomput. | 2 |
| 2024 | Pedestrian detection in low-light conditions: A comprehensive surveyabstractPedestrian detection remains a critical problem in various domains, such as computer vision, surveillance, and autonomous driving. In particular, accurate and instant detection of pedestrians in low-light conditions and reduced visibility is of utmost importance for autonomous vehicles to prevent accidents and save lives. This paper aims to comprehensively survey various pedestrian detection approaches, baselines, and datasets that specifically target low-light conditions. The survey discusses the challenges faced in detecting pedestrians at night and explores state-of-the-art methodologies proposed in recent years to address this issue. These methodologies encompass a diverse range, including deep learning-based, feature-based, and hybrid approaches, which have shown promising results in enhancing pedestrian detection performance under challenging lighting conditions. Furthermore, the paper highlights current research directions in the field and identifies potential solutions that merit further investigation by researchers. By thoroughly examining pedestrian detection techniques in low-light conditions, this survey seeks to contribute to the advancement of safer and more reliable autonomous driving systems and other applications related to pedestrian safety. Accordingly, most of the current approaches in the field use deep learning-based image fusion methodologies (i.e., early, halfway, and late fusion) for accurate and reliable pedestrian detection. Moreover, the majority of the works in the field (approximately 48%) have been evaluated on the KAIST dataset, while the real-world video feeds recorded by authors have been used in less than six percent of the works. Bahareh Ghari, Ali Tourani, Asadollah Shahbahrami, Georgi Gaydadjiev |
Image Vis. Comput. | 3 |
| 2024 | Parallelization of license plate localization on GPU platform
Ali Reza Akoushideh, Asadollah Shahbahrami, Abdorreza Joe Afshany |
Multim. Tools Appl. | 2 |
| 2024 | A video codec based on background extraction and moving object detection
Soheib Hadi, Asadollah Shahbahrami, Hossien Azgomi |
Multim. Tools Appl. | 2 |
| 2024 | A Cost-Sensitive Machine Learning Model With Multitask Learning for Intrusion Detection in IoTabstractA problem with machine learning (ML) techniques for detecting intrusions in the Internet of Things (IoT) is that they are ineffective in the detection of low-frequency intrusions. In addition, as ML models are trained using specific attack categories, they cannot recognize unknown attacks. This article integrates strategies of cost-sensitive learning and multitask learning into a hybrid ML model to address these two challenges. The hybrid model consists of an autoencoder for feature extraction and a support vector machine (SVM) for detecting intrusions. In the cost-sensitive learning phase for the class imbalance problem, the hinge loss layer is enhanced to make a classifier strong against low-distributed intrusions. Moreover, to detect unknown attacks, we formulate the SVM as a multitask problem. Experiments on the UNSW-NB15 and BoT-IoT datasets demonstrate the superiority of our model in terms of recall, precision, and F1-score averagely 92.2%, 96.2%, and 94.3%, respectively, over other approaches. Akbar Telikani, Nima Esmi, Shiva Soleymanpour, Asadollah Shahbahrami, Jun Shen 0001, Georgi Gaydadjiev, Reza Hassanpour |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | PESTD: a large-scale Persian-English scene text dataset
Atefeh Ranjkesh Rashtehroudi, Ali Reza Akoushideh, Asadollah Shahbahrami |
Multim. Tools Appl. | 3 |
| 2022 | Text localization in digital images using a hybrid method
Ali Reza Akoushideh, Sayed Mohammad Fallah Rasoulnejad, Asadollah Shahbahrami |
Multim. Tools Appl. | 3 |
| 2022 | Multi-Metric Re-Identification for Online Multi-Person TrackingabstractMulti-person tracking plays a vital role in intelligent video surveillance systems and has attracted researchers’ growing attention in recent years. This paper proposes a tracking-by-detection method to detect and track all existing persons in video sequences. The proposed method re-identifies detected persons in the latest video frame as observed persons in previous frames and thus generates their trajectories. Re-identification of the proposed approach uses a fusion of six distance metrics. Four metrics, i.e., position, scale, distance to estimated position, and tracklet continuity, are derived from two motion-based features, and two metrics, i.e., dominant colors and histogram of oriented gradients, are derived from corresponding appearance-based features. The proposed method performs tracking in two general steps per each frame. In the first step, all persons in the video frame are detected using the state-of-the-art YOLOv3 object detector. In the second step, the re-identification algorithm generates correspondences between detected persons in the latest video frame and observed persons in previous frames, using the distance matrix built up from compound distances between alldetected–observed personpairs. Experimental results show that our simple yet effective approach achieves significant performance in multi-person tracking compared to existing state-of-the-art methods. Hamid Nodehi, Asadollah Shahbahrami |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Features' value range approach to enhance the throughput of texture classificationabstractAbstract The definition of an image's category from a database with huge texture categories needs massive computation and time cost. Existing texture classification works focus on texture representation to improve the accuracy and efficiency of classification. This research wants to reduce the categories of the main classifier to decrease the comparison time of classification. To overcome computation time, a features' value range (FR) approach to enhance the throughput of texture classification is proposed. The proposed approach decreases the number of candidate categories as a pre‐classifier in a two‐step serial classification. With the decrease in the number of candidates, the main classifier can work on a few categories to find the final category. Here, configuration parameters are defined and some criteria are proposed for evaluating the FR approach. The performance of the FR is evaluated in the presence of different levels of Gaussian noise. Finally, it is shown that using effective features (EF) and hardware implementation approaches can extend the applicability of the FR approach. Experimental results depicted that the throughput of the final decision increased up to 14.85× with considerable reliability. Ali Reza Akoushideh, Babak Mazloom-Nezhad Maybodi, Asadollah Shahbahrami |
IET Image Process. | 3 |
| 2021 | An unsupervised approach for traffic motion patterns extractionabstractAbstract Automatic analysis, understanding typical activities, and identifying vehicle behaviour in crowded traffic scenes are fundamental and challenging tasks for traffic video surveillance. Some recent researches have been using machine learning approaches to extract meaningful patterns occurring in a traffic scene, for example, intersection. In this regard, we convert visual patterns and features to visual words using dense and sparse optical flow and learning traffic motion patterns with group sparse topical coding (GSTC) algorithm. In the first step of the proposed algorithm, the input traffic video is divided into non‐overlapping clips. After that, motion vectors are extracted using dual TV‐L1 as a dense optical flow and Lucas–Kanade as a sparse optical flow and converted to flow words. For learning traffic motion patterns, the GSTC algorithm, that is, a non‐probabilistic topic model (TM) has been applied. These patterns represent priors on observable motion, which can be utilised to describe a scene and answer behaviour questions such as what are the motion patterns in a traffic scene and what is going on. The experimental results which have been obtained using a real dataset, QUML, show that the combination of the GSTC + dual TV‐L1 extracts more traffic motion patterns in comparison with the GSTC + Lucas–Kanade and previous studies. Amin Moradi, Asadollah Shahbahrami, Ali Reza Akoushideh |
IET Image Process. | 2 |
| 2021 | Facial expression recognition using a combination of enhanced local binary pattern and pyramid histogram of oriented gradients features extractionabstractAbstract Automatic facial expression recognition, which has many applications such as drivers, patients, and criminals' emotions recognition, is a challenging task. This is due to the variety of individuals and facial expression variability in different conditions, for instance, gender, race, colour and changing illumination. In addition, there are many regions in a face image such as forehead, mouth, eyes, eyebrows, nose, cheeks and chin, and extracting features of all these regions are expensive in terms of computational time. Each of the six basic emotions of anger, disgust, fear, happiness, sadness and surprise affect some regions more than the other regions. The goal of this study is to evaluate the performance of enhanced local binary pattern, pyramid histogram of oriented gradients feature‐extraction algorithms and their combination in terms of recognition accuracy, feature vector length and computational time on one, two and three combined regions of a face image. Our experimental results show that the combination of both feature‐extraction algorithms yields an average recognition accuracy of 95.33% using three regions, that is, the mouth, nose and eyes on Cohn–Kanade dataset. Besides, the mouth region is the most important part in terms of accuracy in comparison to eyes, nose and combination of both eyes and nose regions. Maede Sharifnejad, Asadollah Shahbahrami, Ali Reza Akoushideh, Reza Hassanpour |
IET Image Process. | 2 |
| 2021 | An Accurate Real-Time License Plate Detection Method Based On Deep Learning ApproachesabstractIn vision-driven Intelligent Transportation Systems (ITS) where cameras play a vital role, accurate detection and re-identification of vehicles are fundamental demands. Hence, recent approaches have employed a wide range of algorithms to provide the best possible accuracy. These methods commonly generate a vehicle detection model based on its visual appearance features such as license plate, headlights, or some other distinguishable specifications. Among different object detection approaches, Deep Neural Networks (DNNs) have the advantage of magnificent detection accuracy in case a huge amount of training data is provided. In this paper, a robust approach for license plate detection (LPD) based on YOLO v.3 is proposed which takes advantage of high detection accuracy and real-time performance. The mentioned approach can detect the license plate location of vehicles as a general representation of vehicle presence in images. To train the model, a dataset of vehicle images with Iranian license plates has been generated by the authors and augmented to provide a wider range of data for test and train purposes. It should be mentioned that the proposed method can detect the license plate area as an indicator of vehicle presence with no Optical Character Recognition (OCR) algorithm to distinguish characters inside the license plate. Experimental results have shown the high performance of the system with a precision 0.979 and recall 0.972. Saeed Khazaee, Ali Tourani, Sajjad Soroori, Asadollah Shahbahrami, Ching Y. Suen |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2021 | High-performance implementation of evolutionary privacy-preserving algorithm for big data using GPU platform
Akbar Telikani, Asadollah Shahbahrami, Amir Hossein Gandomi |
Inf. Sci. | 2 |
| 2020 | Privacy-preserving in association rule mining using an improved discrete binary artificial bee colony
Akbar Telikani, Amir Hossein Gandomi, Asadollah Shahbahrami, Mohammad Naderi Dehkordi |
Expert Syst. Appl. | 3 |
| 2020 | A survey of evolutionary computation for association rule mining
Akbar Telikani, Amir Hossein Gandomi, Asadollah Shahbahrami |
Inf. Sci. | 3 |
| 2020 | SIMD programming using Intel vector extensions
Asadollah Shahbahrami |
J. Parallel Distributed Comput. | 2 |
| 2018 | Data sanitization in association rule mining: An analytical review
Akbar Telikani, Asadollah Shahbahrami |
Expert Syst. Appl. | 2 |
| 2017 | Optimizing association rule hiding using combination of border and heuristic approaches
Akbar Telikani, Asadollah Shahbahrami |
Appl. Intell. | 2 |
| 2016 | Image quality assessment using edge based features
Abdolrahman Attar, Asadollah Shahbahrami, Reza Moradi Rad |
Multim. Tools Appl. | 2 |
| 2013 | A predictive algorithm for multimedia data compression
Reza Moradi Rad, Abdolrahman Attar, Asadollah Shahbahrami |
Multim. Syst. | 3 |
| 2012 | Algorithms and architectures for 2D discrete wavelet transform
Asadollah Shahbahrami |
J. Supercomput. | 1 |
| 2012 | Parallel implementation of Gray Level Co-occurrence Matrices and Haralick texture features on cell architecture
Asadollah Shahbahrami, Koen Bertels |
J. Supercomput. | 1 |
| 2011 | Collaboration of reconfigurable processors in grid computing: Theory and application
Mahmood Ahmadi, Asadollah Shahbahrami, Stephan Wong |
Future Gener. Comput. Syst. | 2 |
| 2010 | Collaboration of Reconfigurable Processors in Grid Computing for Multimedia Kernels
Mahmood Ahmadi, Asadollah Shahbahrami, Stephan Wong |
GPC | 2 |
| 2010 | An accurate gradient-based predictive algorithm for image compressionabstractMany prediction algorithms have been proposed and implemented so far for efficient storage and transmission. In this paper, we propose an algorithm for predictive coding. This new algorithm evaluates and uses the values of the known pixels in much more different directions than previous algorithms such as Median Edge Detection (MED) and Gradient Adjusted Prediction (GAP). The proposed algorithm, MED, and GAP algorithms are implemented and tested on different image contents and sizes. The experimental results show that the prediction error of the proposed algorithm is much less than other algorithms. In other words, the quality of obtained predicted images via the proposed method is better than MED and GAP algorithms. Additionally, the accuracy of the proposed technique is independent from the image contents and image sizes. Abdolrahman Attar, Reza Moradi Rad, Asadollah Shahbahrami |
MoMM | 3 |
| 2009 | Performance Improvement of Multimedia Kernels by Alleviating Overhead Instructions on SIMD Devices
Asadollah Shahbahrami, Ben H. H. Juurlink |
APPT | 1 |
| 2009 | SIMD Architectural Enhancements to Improve the Performance of the 2D Discrete Wavelet TransformabstractThe 2D Discrete Wavelet Transform (DWT) is a time-consuming kernel in many multimedia applications such as JPEG2000 and MPEG-4. The 2D DWT consists of horizontal filtering along the rows followed by vertical filtering along the columns. The vertical filtering is easy to vectorize (assuming row-major order), but to vectorize the horizontal filtering many overhead instructions are required. In this paper we propose some SIMD architectural enhancements, such as the MAC operation, extended subwords, and the matrix register file technique, to develop high-performance implementations of the 2D DWT on SIMD architectures. The MAC operation performs four 32-bit single-precision floating-point multiplications with accumulation. The matrix register file allows to load data stored consecutively in memory to a column of the register file, where a column corresponds to corresponding subwords of different registers. These techniques avoid the need of data rearrangement instructions. In addition, in order to avoid data type conversion instructions, the extended subword technique is applied for the (5, 3) lifting transform. Extended subwords use registers that are wider than the packed format used to store the data. These techniques provide speedups of up to 2.90 and 1.32 for the (5, 3) lifting and Daub-4 transforms, respectively. Asadollah Shahbahrami, Ben H. H. Juurlink |
DSD | 1 |
| 2008 | Data Locality Optimization Based on Comprehensive Knowledge of the Cache Miss Reason: A Case Study with DWTabstractThe overall performance of a computing system increasingly depends on the efficient use of the cache memories. Traditional approaches for cache tuning deploy performance tools to help the user optimize the source program towards a better runtime data locality. Following this conventional way, we developed a set of such toolkits including data profiling, pattern analysis, and performance visualization tools. This paper demonstrates how the toolset can be used step-by-step to understand the cache access behavior of the applications and then achieve optimized program code. The Discrete Wavelet Transform, a common used algorithm for image and video compression, is applied as an example. Our initial experimental results with this sample application show an up to 3.19 speedup in execution time compared to the original implementation. Asadollah Shahbahrami |
HPCC | 2 |
| 2008 | Optimization of Content-Based Image Retrieval FunctionsabstractFeature extraction and similarity measurement are two important operations in content-based image retrieval systems. We optimize and vectorize typical feature extraction algorithms, mean and standard deviation, and some similarity measurement functions such as the sum-of-squared-differences (SSD), the sum-of-absolute differences (SAD), and histogram intersection on a general-purpose processor enhanced with SIMD extensions. In the straightforward implementation of the mean and standard deviation, there are two passes, one to compute the mean and one to compute the standard deviation.We use a single-loop approach that computes both the mean and the standard deviation in a single pass. This technique yields a speedup of up to 1.85 over the double-loop implementation. We vectorize the single-loop implementation using the MMX and SSE2 extensions. The vectorized versions improve performance by a factor of up to 14.49. In addition,we vectorize the SSD, SAD, and histogram intersection similarity measurements using SSE. The vectorized versions provide a maximum speedup of 1.45, 2.33, and 5.24 for the SSD, the SAD, and histogram intersection, respectively,over the optimized scalar implementations. Asadollah Shahbahrami, Ben H. H. Juurlink |
ISM | 1 |
| 2008 | Versatility of extended subwords and the matrix register fileabstractExtended subwords and the matrix register file (MRF) are two micro architectural techniques that address some of the limitations of existing SIMD architectures. Extended subwords are wider than the data stored in memory. Specifically, for every byte of data stored in memory, there are four extra bits in the media register file. This avoids the need for data-type conversion instructions. The MRF is a register file organization that provides both conventional row-wise, as well as column-wise, access to the register file. In other words, it allows to view the register file as a matrix in which corresponding subwords in different registers corresponds to a column of the matrix. It was introduced to accelerate matrix transposition which is a very common operation in multimedia applications. In this paper, we show that the MRF is very versatile, since it can also be used for other permutations than matrix transposition. Specifically, it is shown how it can be used to provide efficient access to strided data, as is needed in, e.g., color space conversion. Furthermore, it is shown that special-purpose instructions (SPIs), such as the sum-of-absolute differences (SAD) instruction, have limited usefulness when extended subwords and a few general SIMD instructions that we propose are supported, for the following reasons. First, when extended subwords are supported, the SAD instruction provides only a relatively small performance improvement. Second, the SAD instruction processes 8-bit subwords only, which is not sufficient for quarter-pixel resolution nor for cost functions used in image and video retrieval. Results obtained by extending the SimpleScalar toolset show that the proposed techniques provide a speedup of up to 3.00 over the MMX architecture. The results also show that using, at most, 13 extra media registers yields an additional performance improvement ranging from 1.38 to 1.57. Asadollah Shahbahrami, Ben H. H. Juurlink, Stamatis Vassiliadis |
ACM Trans. Archit. Code Optim. | 1 |
| 2008 | Implementing the 2-D Wavelet Transform on SIMD-Enhanced General-Purpose ProcessorsabstractThe 2-D Discrete Wavelet Transform (DWT) consumes up to 68% of the JPEG2000 encoding time. In this paper, we develop efficient implementations of this important kernel on general-purpose processors (GPPs), in particular the Pentium 4 (P4). Efficient implementations of the 2-D DWT on the P4 must address three issues. First, the P4 suffers from a problem known as 64K aliasing, which can degrade performance by an order of magnitude. We propose two techniques to avoid 64K aliasing which improve performance by a factor of up to 4.20. Second, a straightforward implementation of vertical filtering incurs many cache misses. Cache performance can be improved by applying loop interchange, but there will still be many conflict misses if the filter length exceeds the cache associativity. Two methods are proposed to reduce the number of conflict misses which provide an additional performance improvement of up to 1.24. To show that these methods are general, results for the P3 and Opteron are also provided. Third, efficient implementations of the 2-D DWT must exploit the SIMD instructions supported by most GPPs, including the P4, and we present MMX and SSE implementations of horizontal and vertical filtering which provide a maximum speedup of 3.39 and 6.72, respectively. Asadollah Shahbahrami, Ben H. H. Juurlink, Stamatis Vassiliadis |
IEEE Trans. Multim. | 1 |
| 2007 | SIMD Vectorization of Histogram FunctionsabstractExisting SIMD extensions cannot efficiently vectorize the histogram function due to memory collisions. We propose two techniques to avoid this problem. In the first, a hierarchical structure of three levels is proposed. In order to provide n-way parallelism, auxiliary arrays that have n and n/2 subarrays are used in the first and second level, respectively. The last level has the primary histogram array. Indirect SIMD load and store instructions are designed in order to access different elements of different subarrays. The different subarrays in the lower levels are merged and finally at the end, the calculated results are stored in the primary histogram array. In the second method, parallel comparators are used in order to count the number of subwords within a media register that are the same. Thereafter, these numbers are added to the values of the histogram array simultaneously. Experimental results obtained by extending the SimpleScalar toolset show that proposed techniques improve the performance compared to the fastest scalar version by a factor of 7.37 and 5.52, respectively. Asadollah Shahbahrami, Ben H. H. Juurlink, Stamatis Vassiliadis |
ASAP | 1 |
| 2007 | Optimizing Cache Performance of the Discrete Wavelet Transform Using a Visualization ToolabstractThe 2D DWT consists of two 1D DWT in both directions: horizontal filtering processes the rows followed by vertical filtering processes the columns. It is well known that a straightforward implementation of the vertical filtering shows quite different performance with various working set sizes. The only reasonable explanation for this has to be the access behavior of the cache memory. As known, vertical filtering has mapping conflicts in the cache with a working set size that is power of two. However, it is not clear how this conflict forms and whether cache problems exist with other data sizes. Such knowledge is the base for efficient code optimization. In order to acquire this knowledge and to achieve more accurate optimization potentials, we apply a cache visualization tool to examine the runtime cache activities of the vertical implementation. We find that besides mapping conflicts, vertical filtering also shows a large number of capacity misses. More specifically, the visualization tool allows us to detect the parameters related to the strategies. This guarantees the feasibility of the optimization. Our initial experimental results on several different architectures show an up to 215% gain in execution time compared to an already optimized baseline implementation. Jie Tao 0001, Asadollah Shahbahrami, Ben H. H. Juurlink, Rainer Buchty, Wolfgang Karl, Stamatis Vassiliadis |
ISM | 2 |
| 2006 | Limitations of special-purpose instructions for similarity measurements in media SIMD extensionsabstractMicroprocessor vendors have provided special-purpose instructions such as psadbw and pdist to accelerate the sum-of-absolute differences (SAD) similarity measurement. The usefulness of these special-purpose instructions is limited except for the motion estimation kernel. This has several drawbacks. First, if the SAD becomes obsolete because a different similarity metric is going to be employed, then those special-purpose instructions are no longer useful. Second, these special instructions process 8-bit subwords only. This precision is not su cient for some kernels such as motion estimation in the transform domain. In addition, when employing other n-way parallel SIMD instructions to implement the SAD and sum-of-squared differences (SSD),the obtained speedup is much less than n. This is because there is a mismatch between the storage and the computational format. In this paper, we design and evaluate a variety of SIMD instructions for different data types. We synthesize special-purpose instructions using a few general-purpose SIMD instructions. In addition, we employ the extended subwords technique to avoid conversion overhead and to increase parallelism. In this technique there are four extra bits for every byte of register. The results show that using different SIMD instructions and extended subwords achieve a speedup ranging from 10.39 to 14.57 over C performance for SAD, SSD with interpolation, and SSD functions in the motion estimation kernel. While, MMX achieves a speedup ranging from 4.61 to 7.42. Additionally,the proposed SIMD instructions improve the performance of similarity measurement for image histograms by a factor ranging from 8.69 (1-way)to 11.70 (4-way) over C.While for MMX speedup is between 2.90 (1-way) and 4.33 (4-way). Asadollah Shahbahrami, Ben H. H. Juurlink, Stamatis Vassiliadis |
CASES | 1 |
| 2006 | Accelerating Color Space Conversion Using Extended Subwords and the Matrix Register FileabstractColor space conversion is an important kernel in multimedia codecs such as JPEG and MPEG. When implemented using SIMD instructions, however, the performance improvement is often limited due to two reasons. First, corresponding color space components are stored at non-unit strides and, second, intermediate results can be larger than 8 bits. In this paper we show that extended subwords and the matrix register file (MRF) can be employed to mitigate these limitations. These techniques avoid rearrangement instructions and increase the number of subwords that are processed in parallel. Experimental results have been obtained by extending the SimpleScalar toolset. The results show that extended subwords and the MRF yield a speedup of up to 2.45x and 1.78x over MMX for the RGB-to-YCbCr and YCbCr-to-RGB kernels, respectively. Compared to C implementations, speedups of up to 10.09x and 6.74x, respectively, are obtained. Additionally, the results show that the speedup over MMX is higher for low issue rates. This means that extended subwords and the MRF are suitable techniques for embedded multimedia systems where high issue rates and out-of-order execution are too expensive. The results also show that using more registers improves performance substantially Asadollah Shahbahrami, Ben H. H. Juurlink, Stamatis Vassiliadis |
ISM | 1 |
| 2005 | Performance Comparison of SIMD Implementations of the Discrete Wavelet TransformabstractThis paper focuses on SIMD implementations of the 2D discrete wavelet transform (DWT). The transforms considered are Daubechies' real-to-real method of four coefficients (Daub-4) and the integer-to-integer (5, 3) lifting scheme. Daub-4 is implemented using SSE and the lifting scheme using MMX, and their performance is compared to C implementations on a Pentium 4 processor. The MMX implementation of the lifting scheme is up to 4.0/spl times/ faster than the corresponding C program for a 1-level 2D DWT, while the SSE implementation of Daub-4 is up to 2.6/spl times/ faster than the C version. It is shown that for some image sizes, the performance is significantly hampered by the so called 64K aliasing problem, which occurs in the Pentium 4 when two data blocks are accessed that are a multiple of 64K apart. It is also shown that for the (5, 3) lifting scheme, a 12-bit word size is sufficient for a 5-level decomposition of the 2D DWT for images of up to 10 bits per pixel. Asadollah Shahbahrami, Ben H. H. Juurlink, Stamatis Vassiliadis |
ASAP | 1 |