Ali Akbari 0003

dblp:04/6898-3 · DBLP profile ↗
← Back
25ranked-venue papers
14as first author
12since 2021 · last 2025
0000-0003-1457-2977ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 2 since 2021Computer networks · 2Security and privacy · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 SLICE: Synthetic Caption-Trained Lightweight Image Captioner for Edge Devices
abstract
Image captioning has made significant steps forward in recent years. Existing state-of-the-art image captioning models are heavily parameterized and a long way from being deployable on edge devices. Providing image captioning inference on edge devices with current state-of-the-art models requires expensive server-side infrastructure and a stable network connection. In this paper, we propose Synthetic caption-trained Lightweight Image Captioning for Edge devices (SLICE). It provides state-of-the-art caption quality at a slice of the compute of existing large models. We achieve this by proposing a lightweight pipeline, redesigning the baseline architecture by making multiple modifications. The model is first pre-trained using our own generated synthetic captions and then fine-tuned by adopting a novel knowledge distillation approach to improve the caption quality. Our small-SLICE achieves 8.9x reduction in the number of model parameters, 7.9x reduction in peak memory consumption and 10.5x latency improvement, while surpassing model quality compared to the baseline model using a smaller input size achieving 138.9 CIDEr. Our large-SLICE achieves 147.3 CIDEr outperforming existing image captioning models with only 35.8M parameters on COCO Karpathy test split. On desktop CPU, our proposed small-SLICE achieves a low latency of 36ms.
Shayan Joya, Soroush Fatemifar, Ali Akbari 0003, Andrew Sheppard, Christopher Alder
ICIP3
2025 Pure anomaly detection via self-supervised deep metric learning with adaptive margin
abstract
We address the problem of anomaly detection (AD) by a deep network pretrained using self-supervised learning for an auxiliary geometric transformation (GT) classification task. Our key contribution is a novel loss function that augments the standard cross-entropy by an additional term that plays a significant role in the later stages of self-supervised learning. The proposed enabling innovation is a triplet centre loss with an adaptive margin and a learnable metric, which relentlessly drives the GT classes to exhibit continuously improving compactness and inter-class separation. The pretrained network is finetuned for the downstream task using non-anomalous data only, and a GT model for the data is constructed. Anomalies are detected by fusing the output of several decision functions defined using the learnt GT class model. In contrast to the majority of existing methods, our approach strictly adheres to the pure AD design philosophy, which relies on the use of purely non-anomalous data for the design. Extensive experiments on four publicly available AD datasets demonstrate the effectiveness of the proposed contributions and lead to significant performance gains compared to the state-of-the-art (1.8% on F-MNIST, 1.0% on CIFAR-10, 1.2% on CIFAR-100, and 1.7% on CatVsDog). https://github.com/12sf12/Deep-Anomaly-Detection • A three-stage pure anomaly detection framework using self-supervised learning. • A novel loss addressing margin and metric selection issues in triplet-based losses. • We propose a method to compute adaptive margins per sample in each mini-batch. • Experiments show the proposed method outperforms SOTA across all datasets.
Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler
Neurocomputing3
2024 RAgE: Robust Age Estimation Through Subject Anchoring With Consistency Regularisation
abstract
Modern facial age estimation systems can achieve high accuracy when training and test datasets are identically distributed and captured under similar conditions. However, domain shifts in data, encountered in practice, lead to a sharp drop in accuracy of most existing age estimation algorithms. In this article, we propose a novel method, namely RAgE, to improve the robustness and reduce the uncertainty of age estimates by leveraging unlabelled data through a subject anchoring strategy and a novel consistency regularisation term. First, we propose an similarity-preserving pseudo-labelling algorithm by which the model generates pseudo-labels for a cohort of unlabelled images belonging to the same subject, while taking into account the similarity among age labels. In order to improve the robustness of the system, a consistency regularisation term is then used to simultaneously encourage the model to produce invariant outputs for the images in the cohort with respect to an anchor image. We propose a novel consistency regularisation term the noise-tolerant property of which effectively mitigates the so-called confirmation bias caused by incorrect pseudo-labels. Experiments on multiple benchmark ageing datasets demonstrate substantial improvements over the state-of-the-art methods and robustness to confounding external factors, including subject's head pose, illumination variation and appearance of expression in the face image.
Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Syed Safwan Khalid, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Deep Order-Preserving Learning With Adaptive Optimal Transport Distance
abstract
We consider a framework for taking into consideration the relative importance (ordinality) of object labels in the process of learning a label predictor function. The commonly used loss functions are not well matched to this problem, as they exhibit deficiencies in capturing natural correlations of the labels and the corresponding data. We propose to incorporate such correlations into our learning algorithm using an optimal transport formulation. Our approach is to learn the ground metric, which is partly involved in forming the optimal transport distance, by leveraging ordinality as a general form of side information in its formulation. Based on this idea, we then develop a novel loss function for training deep neural networks. A highly efficient alternating learning method is then devised to alternatively optimise the ground metric and the deep model in an end-to-end learning manner. This scheme allows us to adaptively adjust the shape of the ground metric, and consequently the shape of the loss function for each application. We back up our approach by theoretical analysis and verify the performance of our proposed scheme by applying it to two learning tasks, i.e. chronological age estimation from the face and image aesthetic assessment. The numerical results on several benchmark datasets demonstrate the superiority of the proposed algorithm.
Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 NPT-Loss: Demystifying Face Recognition Losses With Nearest Proxies Triplet
abstract
Face recognition (FR) using deep convolutional neural networks (DCNNs) has seen remarkable success in recent years. One key ingredient of DCNN-based FR is the design of a loss function that ensures discrimination between various identities. The state-of-the-art (SOTA) solutions utilise normalised Softmax loss with additive and/or multiplicative margins. Despite being popular and effective, these losses are justified only intuitively with little theoretical explanations. In this work, we show that under the LogSumExp (LSE) approximation, the SOTA Softmax losses become equivalent to a proxy-triplet loss that focuses on nearest-neighbour negative proxies only. This motivates us to propose a variant of the proxy-triplet loss, entitled Nearest Proxies Triplet (NPT) loss, which unlike SOTA solutions, converges for a wider range of hyper-parameters and offers flexibility in proxy selection and thus outperforms SOTA techniques. We generalise many SOTA losses into a single framework and give theoretical justifications for the assertion that minimising the proposed loss ensures a minimum separability between all identities. We also show that the proposed loss has an implicit mechanism of hard-sample mining. We conduct extensive experiments using various DCNN architectures on a number of FR benchmarks to demonstrate the efficacy of the proposed scheme over SOTA methods.
Syed Safwan Khalid, Muhammad Awais 0001, Zhenhua Feng 0001, Chi-Ho Chan, Ammarah Farooq, Ali Akbari 0003, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 A Theoretical Insight Into the Effect of Loss Function for Deep Semantic-Preserving Learning
abstract
Good generalization performance is the fundamental goal of any machine learning algorithm. Using the uniform stability concept, this article theoretically proves that the choice of loss function impacts the generalization performance of a trained deep neural network (DNN). The adopted stability-based framework provides an effective tool for comparing the generalization error bound with respect to the utilized loss function. The main result of our analysis is that using an effective loss function makes stochastic gradient descent more stable which consequently leads to the tighter generalization error bound, and so better generalization performance. To validate our analysis, we study learning problems in which the classes are semantically correlated. To capture this semantic similarity of neighboring classes, we adopt the well-known semantics-preserving learning framework, namely label distribution learning (LDL). We propose two novel loss functions for the LDL framework and theoretically show that they provide stronger stability than the other widely used loss functions adopted for training DNNs. The experimental results on three applications with semantically correlated classes, including facial age estimation, head pose estimation, and image esthetic assessment, validate the theoretical insights gained by our analysis and demonstrate the usefulness of the proposed loss functions in practical applications.
Ali Akbari 0003, Muhammad Awais 0001, Manijeh Bashar, Josef Kittler
IEEE Trans. Neural Networks Learn. Syst.1
2022 Distribution Cognisant Loss for Cross-Database Facial Age Estimation With Sensitivity Analysis
abstract
Existing facial age estimation studies have mostly focused on intra-database protocols that assume training and test images are captured under similar conditions. This is rarely valid in practical applications, where we typically encounter training and test sets with different characteristics. In this article, we deal with such situations, namely subjective-exclusive cross-database age estimation. We formulate the age estimation problem as the distribution learning framework, where the age labels are encoded as a probability distribution. To improve the cross-database age estimation performance, we propose a new loss function which provides a more robust measure of the difference between ground-truth and predicted distributions. The desirable properties of the proposed loss function are theoretically analysed and compared with the state-of-the-art approaches. In addition, we compile a new balanced large-scale age estimation database. Last, we introduce a novel evaluation protocol, called subject-exclusive cross-database age estimation protocol, which provides meaningful information of a method in terms of the generalisation capability. The experimental results demonstrate that the proposed approach outperforms the state-of-the-art age estimation methods under both intra-database and subject-exclusive cross-database evaluation protocols. In addition, in this article, we provide a comparative sensitivity analysis of various algorithms to identify trends and issues inherent to their performance. This analysis introduces some open problems to the community which might be considered when designing a robust age estimation system.
Ali Akbari 0003, Muhammad Awais 0001, Zhenhua Feng 0001, Ammarah Farooq, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Developing a generic framework for anomaly detection
Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler
Pattern Recognit.3
2022 Face spoofing detection ensemble via multistage optimisation and pruning
Soroush Fatemifar, Shahrokh Asadi, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler
Pattern Recognit. Lett.4
2022 A Novel Ground Metric for Optimal Transport-Based Chronological Age Estimation
abstract
Label distribution learning (LDL) is the state-of-the-art approach to dealing with a number of real-world applications, such as chronological age estimation from a face image, where there is an inherent similarity among adjacent age labels. LDL takes into account the semantic similarity by assigning a label distribution to each instance. The well-known Kullback-Leibler (KL) divergence is the widely used loss function for the LDL framework. However, the KL divergence does not fully and effectively capture the semantic similarity among age labels, thus leading to suboptimal performance. In this article, we propose a novel loss function based on the optimal transport theory for the LDL-based age estimation. A ground metric function plays an important role in the optimal transport formulation. It should be carefully determined based on the underlying geometric structure of the label space of the application in-hand. The label space in the age estimation problem has a specific geometric structure, that is, closer ages have more inherent semantic relationships. Inspired by this, we devise a novel ground metric function, which enables the loss function to increase the influence of highly correlated ages; thus exploiting the semantic similarity among ages more effectively than the existing loss functions. We then use the proposed loss function, namely, γ -Wasserstein loss, for training a deep neural network (DNN). This leads to a notoriously computationally expensive and nonconvex optimization problem. Following the standard methodology, we formulate the optimization function as a convex problem and then use an efficient iterative algorithm to update the parameters of the DNN. Extensive experiments in age estimation on different benchmark datasets validate the effectiveness of the proposed method, which consistently outperforms state-of-the-art approaches.
Ali Akbari 0003, Muhammad Awais 0001, Soroush Fatemifar, Syed Safwan Khalid, Josef Kittler
IEEE Trans. Cybern.1
2021 Particle Swarm And Pattern Search Optimisation Of An Ensemble Of Face Anomaly Detectors
abstract
While the remarkable advances in face matching render face biometric technology more widely applicable, its successful deployment may be compromised by face spoofing. Recent studies have shown that anomaly-based face spoofing detectors offer an interesting alternative to the multiclass counterparts by generalising better to unseen types of attack. In this work, we investigate the merits of fusing multiple anomaly spoofing detectors in the unseen attack scenario via a Weighted Averaging (WA) and client-specific design. We propose to optimise the parameters of WA by a two-stage optimisation method consisting of Particle Swarm Optimisation (PSO) and the Pattern Search (PS) algorithms to avoid the local minimum problem. Besides, we propose a novel scoring normalisation method which could be effectively applied in extreme cases such as heavy-tailed distributions. We evaluate the capability of the proposed system on publicly available face anti-spoofing databases including Replay-Attack, Replay-Mobile and Rose-Youtu. The experimental results demonstrate that the proposed fusion system outperforms the majority of anomaly-based and state-of-the-art multiclass approaches.
Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler
ICIP3
2021 How Does Loss Function Affect Generalization Performance of Deep Learning? Application to Human Age Estimation
abstract
Good generalization performance across a wide variety of domains caused by many external and internal factors is the fundamental goal of any machine learning algorithm. This paper theoretically proves that the choice of loss function matters for improving the generalization performance of deep learning-based systems. By deriving the generalization error bound for deep neural models trained by stochastic gradient descent, we pinpoint the characteristics of the loss function that is linked to the generalization error and can therefore be used for guiding the loss function selection process. In summary, our main statement in this paper is: choose a stable loss function, generalize better. Focusing on human age estimation from the face which is a challenging topic in computer vision, we then propose a novel loss function for this learning problem. We theoretically prove that the proposed loss function achieves stronger stability, and consequently a tighter generalization error bound, compared to the other common loss functions for this problem. We have supported our findings theoretically, and demonstrated the merits of the guidance process experimentally, achieving significant improvements.
Ali Akbari 0003, Muhammad Awais 0001, Manijeh Bashar, Josef Kittler
ICML1
2020 Sensitivity of Age Estimation Systems to Demographic Factors and Image Quality: Achievements and Challenges
abstract
Recently, impressively growing efforts have been devoted to the challenging task of facial age estimation. The improvements in performance achieved by new algorithms are measured on several benchmarking test databases with different characteristics to check on consistency. While this is a valuable methodology in itself, a significant issue in the most age estimation related studies is that the reported results lack an assessment of intrinsic system uncertainty. Hence, a more in-depth view is required to examine the robustness of age estimation systems in different scenarios. The purpose of this paper is to conduct an evaluative and comparative analysis of different age estimation systems to identify trends, as well as the points of their critical vulnerability. In particular, we investigate four age estimation systems, including the online Microsoft service, two best state-of-the-art approaches advocated in the literature, as well as a novel age estimation algorithm. We analyse the effect of different internal and external factors, including gender, ethnicity, expression, makeup, illumination conditions, quality and resolution of the face images, on the performance of these age estimation systems. The goal of this sensitivity analysis is to provide the biometrics community with the insight and understanding of the critical subject-, camera- and environmental-based factors that affect the overall performance of the age estimation system under study.
Ali Akbari 0003, Muhammad Awais 0001, Josef Kittler
IJCB1
2020 Cross Modal Person Re-identification with Visual-Textual Queries
abstract
Classical person re-identification approaches assume that a person of interest has appeared across different cameras and can be queried by one of the existing images. However, in real-world surveillance scenarios, frequently no visual information will be available about the queried person. In such scenarios, a natural language description of the person by a witness will provide the only source of information for retrieval. In this work, person re-identification using both vision and language information is addressed under all possible gallery and query scenarios. A two stream deep convolutional neural network framework supervised by identity based cross entropy loss is presented. Canonical Correlation Analysis is performed to enhance the correlation between the two modalities in a joint latent embedding space. To investigate the benefits of the proposed approach, a new testing protocol under a multi modal ReID setting is proposed for the test split of the CUHK-PEDES and CUHK-SYSU benchmarks. The experimental results verify that the learnt visual representations are more robust and perform 20% better during retrieval as compared to a single modality system.
Ammarah Farooq, Muhammad Awais 0001, Josef Kittler, Ali Akbari 0003, Syed Safwan Khalid
IJCB4
2020 Deep Learning-Aided Finite-Capacity Fronthaul Cell-Free Massive MIMO with Zero Forcing
abstract
We consider a cell-free massive multiple-input multiple-output (MIMO) system where the channel estimates and the received signals are quantized at the access points (APs) and forwarded to a central processing unit (CPU). Zero-forcing technique is used at the CPU to detect the signals transmitted from all users. To solve the non-convex sum rate maximization problem, a heuristic sub-optimal scheme is proposed to convert the problem into a geometric programme (GP). Exploiting a deep convolutional neural network (DCNN) allows us to determine both a mapping from the large-scale fading (LSF) coefficients and the optimal power by solving the optimization problem using the quantized channel. Depending on how the optimization problem is solved, different power control schemes are investigated; i) small-scale fading (SSF)-based power control; ii) LSF use-and-then-forget (UatF)-based power control; and iii) LSF deep learning (DL)-based power control. The SSF-based power control scheme needs to be solved for each coherence interval of the SSF, which is practically impossible in real time systems. Numerical results reveal that the proposed LSF-DL-based scheme significantly increases the performance compared to the practical and well-known LSF-UatF-based power control.
Manijeh Bashar, Ali Akbari 0003, K. Cumanan, Hien Quoc Ngo, Alister Burr, Pei Xiao 0001, Mérouane Debbah
ICC2
2020 A Stacking Ensemble for Anomaly Based Client-Specific Face Spoofing Detection
abstract
To counteract spoofing attacks, the majority of recent approaches to face spoofing attack detection formulate the problem as a binary classification task in which real data and attack-accesses are both used to train spoofing detectors. Although the classical training framework has been demonstrated to deliver satisfactory results, its robustness to unseen attacks is debatable. Inspired by the recent success of anomaly detection models in face spoofing detection, we propose an ensemble of one-class classifiers fused by a Stacking ensemble method to reduce the generalisation error in the more realistic unseen attack scenario. To be consistent with this scenario, anomalous samples are considered neither for training the component anomaly classifiers nor for the design of the Stacking ensemble. To achieve better face-anti spoofing results, we adopt client-specific information to build both constituent classifiers as well as the Stacking combiner. Besides, we propose a novel 2-stage Genetic Algorithm to further improve the generalisation performance of Stacking ensemble. We evaluate the effectiveness of the proposed systems on publicly available face anti-spoofing databases including Replay-Attack, Replay-Mobile and Rose-Youtu. The experimental results following the unseen attack evaluation protocol confirm the merits of the proposed model.
Soroush Fatemifar, Muhammad Awais 0001, Ali Akbari 0003, Josef Kittler
ICIP3
2020 A Flatter Loss for Bias Mitigation in Cross-dataset Facial Age Estimation
abstract
The most existing studies in the facial age estimation assume training and test images are captured under similar shooting conditions. However, this is rarely valid in real-worlds applications, where training and test sets usually have different characteristics. In this paper, we advocate a cross-dataset protocol for age estimation benchmarking. In order to improve the cross-dataset age estimation performance, we mitigate the inherent bias caused by the learning algorithm itself. To this end, we propose a novel loss function that is more effective for neural network training. The relative smoothness of the proposed loss function is its advantage with regards to the optimisation process performed by stochastic gradient descent (SGD). Compared with existing loss functions, the lower gradient of the proposed loss function leads to the convergence of SGD to a better optimum point, and consequently a better generalisation. The cross-dataset experimental results demonstrate the superiority of the proposed method over the state-of-the-art algorithms in terms of accuracy and generalisation capability.
Ali Akbari 0003, Muhammad Awais 0001, Zhenhua Feng 0001, Ammarah Farooq, Josef Kittler
ICPR1
2020 Exploiting Deep Learning in Limited-Fronthaul Cell-Free Massive MIMO Uplink
abstract
A cell-free massive multiple-input multiple-output (MIMO) uplink is considered, where quantize-and-forward (QF) refers to the case where both the channel estimates and the received signals are quantized at the access points (APs) and forwarded to a central processing unit (CPU) whereas in combine-quantize-and-forward (CQF), the APs send the quantized version of the combined signal to the CPU. To solve the non-convex sum rate maximization problem, a heuristic sub-optimal scheme is exploited to convert the power allocation problem into a standard geometric programme (GP). We exploit the knowledge of the channel statistics to design the power elements. Employing large-scale-fading (LSF) with a deep convolutional neural network (DCNN) enables us to determine a mapping from the LSF coefficients and the optimal power through solving the sum rate maximization problem using the quantized channel. Four possible power control schemes are studied, which we refer to as i) small-scale fading (SSF)-based QF; ii) LSF-based CQF; iii) LSF use-and-then-forget (UatF)-based QF; and iv) LSF deep learning (DL)-based QF, according to where channel estimation is performed and exploited and how the optimization problem is solved. Numerical results show that for the same fronthaul rate, the throughput significantly increases thanks to the mapping obtained using DCNN.
Manijeh Bashar, Ali Akbari 0003, K. Cumanan, Hien Quoc Ngo, Alister Burr, Pei Xiao 0001, Mérouane Debbah, Josef Kittler
IEEE J. Sel. Areas Commun.2
2020 Joint Sparse Learning With Nonlocal and Local Image Priors for Image Error Concealment
abstract
Joint sparse representation (JSR) model has recently emerged as a powerful technique with wide variety of applications. In this paper, the JSR model is extended to error concealment (EC) application, being effective to recover the original image from its corrupted version. This model is based on jointly learning a dictionary pair and two mapping matrices that are trained offline from external training images. Given the trained dictionaries and mappings, the restoration is done by transferring the recovery problem into the sparse representation domain with respect to the trained dictionaries, which is further transformed into a common space using the respective mapping matrices. Then, the reconstructed image is obtained by back projection into the spatial domain. In order to improve the accuracy and stability of the proposed JSR-based EC algorithm and avoid unexpected artifacts, the local and non-local priors are seamlessly integrated into the JSR model. The non-local prior is based on the self-similarity within natural images and helps to find an accurate sparse representation by taking a weighted average of similar areas throughout the image. The local prior is based on learning the local structural regularity of the natural images and helps to regularize the sparse representation, exploiting the strong correlation in the small local areas within the image. Compared with the state-of-the-art EC algorithms, the results show that the proposed method has better reconstruction performance in terms of objective and subjective evaluations.
Ali Akbari 0003, Maria Trocan, Saeid Sanei, Bertrand Granado
IEEE Trans. Circuits Syst. Video Technol.1
2020 Compressive Imaging Using RIP-Compliant CMOS Imager Architecture and Landweber Reconstruction
abstract
In this paper, we present a new image sensor architecture for fast and accurate compressive sensing (CS) of natural images. Measurement matrices usually employed in CS CMOS image sensors are recursive pseudo-random binary matrices. We have proved that the restricted isometry property of these matrices is limited by a low sparsity constant. The quality of these matrices is also affected by the non-idealities of pseudo-random number generators (PRNG). To overcome these limitations, we propose a hardware-friendly pseudo-random ternary measurement matrix generated on-chip by means of class III elementary cellular automata (ECA). These ECA present a chaotic behavior that emulates random CS measurement matrices better than other PRNG. We have combined this new architecture with a block-based CS smoothed-projected Landweber reconstruction algorithm. By means of single value decomposition, we have adapted this algorithm to perform fast and precise reconstruction while operating with binary and ternary matrices. Simulations are provided to qualify the approach.
Marco Trevisi, Ali Akbari 0003, Maria Trocan, Ángel Rodríguez-Vázquez, Ricardo Carmona-Galán
IEEE Trans. Circuits Syst. Video Technol.2
2018 Robust Image Reconstruction for Block-Based Compressed Sensing Using a Binary Measurement Matrix
abstract
Nowadays, there are still difficulties in the implementation of Compressed Sensing (CS) sensors due to the nature of the measurement matrix. A binary measurement matrix can simplify the CS procedure significantly. However, due to the singularity of this class of measurement matrices, the convergence of some of existing CS reconstruction algorithms, such as the well-known block-based CS with smoothed-projected Landweber reconstruction (BCS-SPL) algorithm, is not guaranteed and can lead to an inaccurate recovery. In this paper we propose a simple, fast and efficient CS recovery algorithm that is able to recover the original image from compressed samples which are obtained using a binary measurement. Singular value decomposition (SVD) is coupled with the BCS-SPL algorithm in order to improve its recovery capability when a binary matrix is employed. The experimental results show that the proposed recovery algorithm has a better performance in terms of reconstruction quality when compared with existing reconstruction algorithm and yields images with quality that matches or exceeds those produced by the BCS-SPL algorithm. Additionally, the proposed algorithm is the most efficient in terms of recovery time, especially at high subrates.
Ali Akbari 0003, Maria Trocan
ICIP1
2018 Downsampling Based Image Coding Using Dual Dictionary Learning and Sparse Representations
abstract
Downsampling based image compression scheme achieves better quality at low bit rates. This paper presents a new scheme in such a paradigm based on adaptive sparse representations with respect to two trained overcomplete dictionaries. The original image is downsampled at the encoder side and an upscaling technique is employed to restore the downsampled image to its original resolution at the decoder side. Due to the downsampling, the high frequency details are removed; therefore, the bit budget of low frequency information is increased, leading to better coding performance at the low bitrates. In order to further improve the coding efficiency, we also propose to encode the residual image as side information. This residual image is obtained by difference between the original image and upscaled image. The low resolution image and the residual image are represented over two dictionaries trained by a bilevel dictionary learning algorithm. Furthermore, the visual salient information is considered into the rate allocation process to improve the rate-distortion performance. The enhanced scheme achieves improvement of the quality at a variety of bitrates at the expense of increasing the system complexity, when compared to the conventional codecs.
Ali Akbari 0003, Maria Trocan
MMSP1
2017 Sparse Recovery-Based Error Concealment
abstract
Image and video transmission over heterogeneous networks may encounter packet loss due to the channel impairments, leading to quality degradation of the received image. In this paper, a novel robust image transmission system is proposed by casting the error concealment challenge into a sparse recovery framework. To this purpose, a robust encoder is carefully designed in order to mitigate the negative effects of the packet loss. After wavelet decomposition, a quadtree structure of the wavelet coefficients is used to rearrange them into independent partitions. Random linear combinations of coefficients for each partition are then adopted to provide a high error recovery capability. This linear process coupled with a simple packetization process introduces more robustness and error resilience into the transmission system. At the receiver side, the tree-sparse structure of the wavelet coefficients is explicitly exploited in order to model the error recovery problem as a sparse recovery framework. This is achieved by adaptation of a well-known iterative sparse reconstruction algorithm to the defined tree structure built in the wavelet domain. Compared with the state-of-the-art error concealment algorithms, experimental results show that the proposed method has better reconstruction performance in terms of objective and subjective evaluations over a range of packet loss rates, and ensure that a high-quality image can be recovered for the high packet loss scenarios.
Ali Akbari 0003, Maria Trocan, Bertrand Granado
IEEE Trans. Multim.1
2016 Image compression using adaptive sparse representations over trained dictionaries
abstract
Sparse representation is a common approach for reducing the spatial redundancy by modelling an image as a linear combination of few atoms taken from an analytic or trained dictionary. This paper introduces a new image codec based on adaptive sparse representations wherein the visual salient information is considered into the rate allocation process. Firstly, the regions of the image that are more conspicuous to the human visual system are extracted using a classical graph-based method. Further, block-based sparse representation over a trained dictionary coupled with an adaptive sparse representation is proposed, such that the adaptivity is achieved by appropriately assigning more atoms of the dictionary to the blocks belonging to the salient regions. Experimental results show that the proposed method outperforms the existing image coding standards, such as JPEG and JPEG2000, which use an analytic dictionary, as well as the state-of-the-art codecs based on trained dictionaries.
Ali Akbari 0003, Maria Trocan, Bertrand Granado
MMSP1
2016 Image error concealment using sparse representations over a trained dictionary
abstract
This paper introduces a novel image error concealment technique, wherein the correlation among the correctly received pixels is implicitly exploited to recover the missing areas. This correlation is modeled as sparse representations of the correctly received surrounding areas of a lost regions, hence as a linear combination of very few atoms chosen from an over-complete dictionary. Under mild conditions, the sparse representation coefficients of a given zone, including both known and unknown pixels, can be correctly recovered from the sparse representation coefficients of its neighboring area. This linear process coupled with a simple smoothing process introduces a high quality concealed image. Compared with the state-of-the-art error concealment algorithms, experimental results show that the proposed method has better reconstruction performance in terms of objective and subjective evaluations.
Ali Akbari 0003, Maria Trocan, Bertrand Granado
PCS1