Mohsen Ali

dblp:02/10964 · DBLP profile ↗
← Back
32ranked-venue papers
3as first author
18since 2021 · last 2025
0000-0003-4809-8679ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2025 Leveraging sparse annotations for leukemia diagnosis on the large leukemia dataset
Abdul Rehman 0012, Talha Meraj, Aiman Mahmood Minhas, Ayisha Imran, Mohsen Ali, Waqas Sultani, Mubarak Shah
Medical Image Anal.5
2024 Improving Single Domain-Generalized Object Detection: A Focus on Diversification and Alignment
abstract
In this work, we tackle the problem of domain generalization for object detection, specifically focusing on the scenario where only a single source domain is available. We propose an effective approach that involves two key steps: diversifying the source domain and aligning detections based on class prediction confidence and localization. Firstly, we demonstrate that by carefully selecting a set of augmentations, a base detector can outperform existing methods for single domain generalization by a good margin. This highlights the importance of domain diversification in improving the performance of object detectors. Secondly, we introduce a method to align detections from multiple views, considering both classification and localization outputs. This alignment procedure leads to better generalized and well-calibrated object detector models, which are crucial for accurate decision-making in safety-critical applications. Our approach is detector-agnostic and can be seamlessly applied to both single-stage and two-stage detectors. To validate the effectiveness of our proposed methods, we conduct extensive experiments and ablations on challenging domain-shift scenarios. The results consistently demonstrate the superiority of our approach compared to existing methods. Our code and models are available at: https://github.com/msohaildanishIDivAlign.
Muhammad Sohail Danish, Muhammad Haris Khan, Muhammad Akhtar Munir, M. Saquib Sarfraz, Mohsen Ali
CVPR5
2024 Few-Shot Domain Adaptive Object Detection for Microscopic Images
Sumayya Inayat, Nimra Dilawar, Waqas Sultani, Mohsen Ali
MICCAI (12)4
2024 A Large-Scale Multi Domain Leukemia Dataset for the White Blood Cells Detection with Morphological Attributes for Explainability
Abdul Rehman 0012, Talha Meraj, Aiman Mahmood Minhas, Ayisha Imran, Mohsen Ali, Waqas Sultani
MICCAI (3)5
2024 Detection and Localization of Firearm Carriers in Complex Scenes for Improved Safety Measures
abstract
Detecting firearms and accurately localizing individuals carrying them in images or videos is of paramount importance in security, surveillance, and content customization. However, this task presents significant challenges in complex environments due to clutter and the diverse shapes of firearms. To address this problem, we propose a novel approach that leverages human–firearm interaction information, which provides valuable clues for localizing firearm carriers. Our approach incorporates an attention mechanism that effectively distinguishes humans and firearms from the background by focusing on relevant areas. Additionally, we introduce a saliency-driven locality-preserving constraint to learn essential features while preserving foreground information in the input image. By combining these components, our approach achieves exceptional results on a newly proposed dataset. To handle inputs of varying sizes, we pass paired human–firearm instances with attention masks as channels through a deep network for feature computation, utilizing an adaptive average pooling (AAP) layer. We extensively evaluate our approach against existing methods in human–object interaction (HOI) detection and achieve significant results (AP = 77.8%) compared to the baseline approach (AP = 63.1%). This demonstrates the effectiveness of leveraging attention mechanisms and saliency-driven locality preservation for accurate human–firearm interaction detection. Our findings contribute to advancing the fields of security and surveillance, enabling more efficient firearm localization and identification in diverse scenarios.
Arif Mahmood, Abdul Basit 0019, Muhammad Akhtar Munir, Mohsen Ali
IEEE Trans. Comput. Soc. Syst.4
2023 Towards Automated Monitoring Of Glacial Lakes In Hindu Kush And Himalayas Using Deep Learning
abstract
A glacial lake outburst flood (GLOF) is typically a natural phenomenon caused by rapid discharge of water from a glacier, leading to a flood. The frequency of GLOFs has increased significantly in the northern areas of Pakistan, which demands identification and continuous monitoring of potentially dangerous glacial lakes. In this paper, an up-to-date inventory of glacial lakes in this region is presented. This inventory (HKH-PK-2020) has been prepared using high resolution PlanetScope imagery acquired in 2020 over northern Pakistan. It contains a total of 8808 lakes. We compare our database with the High Mountain Asia (HMA) glacial lakes inventory over northern Pakistan, prepared in 2018 using Landsat imagery. The new inventory contains 6537 more glacial lakes than the HMA inventory. Furthermore, we have prepared an annotated dataset containing 3525 images (of high resolution PlanetScope imagery over a selected number of lakes from the inventory). Each image comprises 4 bands, namely red, green, blue, and near infrared. The annotations are binary: lake or background. Finally, we have performed an ablation study with two encoder-decoder based convolutional neural networks (CNNs) trained on this dataset for pixel-based classification. Our results show an intersection over union (IoU) score of 72.81% for the lake class, which is a promising first result indicating a use of deep learning for automated inventory updates in future.
Muhammad Adnan Siddique, Abdul Basit 0019, Nida Qayyum, Ehtasham Naseer, Muhammad Khurram Bhatti, Brent Minchew, Mohsen Ali, Cristian Silva-Perez, Armando Marino
IGARSS7
2023 Cal-DETR: Calibrated Detection Transformer
abstract
Albeit revealing impressive predictive performance for several computer vision tasks, deep neural networks (DNNs) are prone to making overconfident predictions. This limits the adoption and wider utilization of DNNs in many safety-critical applications. There have been recent efforts toward calibrating DNNs, however, almost all of them focus on the classification task. Surprisingly, very little attention has been devoted to calibrating modern DNN-based object detectors, especially detection transformers, which have recently demonstrated promising detection performance and are influential in many decision-making systems. In this work, we address the problem by proposing a mechanism for calibrated detection transformers (Cal-DETR), particularly for Deformable-DETR, UP-DETR, and DINO. We pursue the train-time calibration route and make the following contributions. First, we propose a simple yet effective approach for quantifying uncertainty in transformer-based object detectors. Second, we develop an uncertainty-guided logit modulation mechanism that leverages the uncertainty to modulate the class logits. Third, we develop a logit mixing approach that acts as a regularizer with detection-specific losses and is also complementary to the uncertainty-guided logit modulation technique to further improve the calibration performance. Lastly, we conduct extensive experiments across three in-domain and four out-domain scenarios. Results corroborate the effectiveness of Cal-DETR against the competing train-time methods in calibrating both in-domain and out-domain detections while maintaining or even improving the detection performance. Our codebase and pre-trained models can be accessed at \url{https://github.com/akhtarvision/cal-detr}.
Muhammad Akhtar Munir, Salman Khan 0001, Muhammad Haris Khan, Mohsen Ali, Fahad Shahbaz Khan
NeurIPS4
2023 Cross-region building counting in satellite imagery using counting consistency
Muaaz Zakria, Hamza Rawal, Waqas Sultani, Mohsen Ali
Neural Comput. Appl.4
2023 Domain Adaptive Object Detection via Balancing Between Self-Training and Adversarial Learning
abstract
Deep learning based object detectors struggle generalizing to a new target domain bearing significant variations in object and background. Most current methods align domains by using image or instance-level adversarial feature alignment. This often suffers due to unwanted background and lacks class-specific alignment. A straightforward approach to promote class-level alignment is to use high confidence predictions on unlabeled domain as pseudo-labels. These predictions are often noisy since model is poorly calibrated under domain shift. In this paper, we propose to leverage model's predictive uncertainty to strike the right balance between adversarial feature alignment and class-level alignment. We develop a technique to quantify predictive uncertainty on class assignments and bounding-box predictions. Model predictions with low uncertainty are used to generate pseudo-labels for self-training, whereas the ones with higher uncertainty are used to generate tiles for adversarial feature alignment. This synergy between tiling around uncertain object regions and generating pseudo-labels from highly certain object regions allows capturing both image and instance-level context during the model adaptation. We report thorough ablation study to reveal the impact of different components in our approach. Results on five diverse and challenging adaptation scenarios show that our approach outperforms existing state-of-the-art methods with noticeable margins.
Muhammad Akhtar Munir, Muhammad Haris Khan, M. Saquib Sarfraz, Mohsen Ali
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Towards Low-Cost and Efficient Malaria Detection
abstract
Malaria, a fatal but curable disease claims hundreds of thousands of lives every year. Early and correct diagnosis is vital to avoid health complexities, however, it depends upon the availability of costly microscopes and trained experts to analyze blood-smear slides. Deep learning-based methods have the potential to not only decrease the burden of experts but also improve diagnostic accuracy on low-cost microscopes. However, this is hampered by the absence of a reasonable size dataset. One of the most challenging aspects is the reluctance of the experts to annotate the dataset at low magnification on low-cost microscopes. We present a dataset to further the research on malaria microscopy over low-cost microscopes at low magnification. Our large-scale dataset consists of images of blood-smear slides from several malaria-infected patients, collected through micro-scopes at two different cost spectrums and multiple magnifications. Malarial cells are annotated for the localization and life-stage classification task on the images collected through the high-cost microscope at high magnification. We design a mechanism to transfer these annotations from the high-cost microscope at high magnification to the low-cost microscope, at multiple magnifications. Multiple object detectors and domain adaptation methods are presented as the baselines. Furthermore, a partially supervised domain adaptation method is introduced to adapt the object-detector to work on the images collected from the low-cost microscope. The dataset is available here: http://im.itu.edu.pk/m5-malaria-dataset/
Waqas Sultani, Wajahat Nawaz, Syed Javed, Muhammad Sohail Danish, Asma Saadia, Mohsen Ali
CVPR6
2022 Deep Learning for Monitoring Glacial Lakes Formation using Sentinel 2 Multispectral Data
abstract
Glacial lake outburst floods (GLOFs) are a major threat to the local communities and important infrastructures in the high mountain regions. This paper focuses on the development of a benchmark dataset for glacial lakes classification in Sentinel 2 multi-spectral data and subsequent detection of glacial lakes prior to a glacial lake outburst flood (GLOF). Towards this end, we collected Sentinel 2 true color scenes of High-Mountain Asia (HMA) region using glacial lakes inventory of this region. It covers an area of 2080.12 km 2 with nearly 30,121 glacial lakes. After data collection, we retained 1200 cloud free true color images and manually generated their ground truth masks. The dataset covers lakes with different shapes, sizes and radiometric signatures. For detection of glacial lakes, we used an encoder-decoder based convolutional neural network (CNN). The model is trained on the labelled dataset of glacial lakes for semantic segmentation of true color images into two relevant classes: lake and no lake. The performance of the proposed model is evaluated using intersection over union (IoU) score. It classifies glacial lakes correctly with an IoU score of 79.90%, which is quite good as far as complexity of the problem is concerned.
Abdul Basit 0019, Muhammad Khurram Bhatti, Mohsen Ali, Tooba Fatima, Brent Minchew, Muhammad Adnan Siddique
IGARSS3
2022 Towards Improving Calibration in Object Detection Under Domain Shift
abstract
With deep neural network based solution more readily being incorporated in real-world applications, it has been pressing requirement that predictions by such models, especially in safety-critical environments, be highly accurate and well-calibrated. Although some techniques addressing DNN calibration have been proposed, they are only limited to visual classification applications and in-domain predictions. Unfortunately, very little to no attention is paid towards addressing calibration of DNN-based visual object detectors, that occupy similar space and importance in many decision making systems as their visual classification counterparts. In this work, we study the calibration of DNN-based object detection models, particularly under domain shift. To this end, we first propose a new, plug-and-play, train-time calibration loss for object detection (coined as TCD). It can be used with various application-specific loss functions as an auxiliary loss function to improve detection calibration. Second, we devise a new implicit technique for improving calibration in self-training based domain adaptive detectors, featuring a new uncertainty quantification mechanism for object detection. We demonstrate TCD is capable of enhancing calibration with notable margins (1) across different DNN-based object detection paradigms both in in-domain and out-of-domain predictions, and (2) in different domain-adaptive detectors across challenging adaptation scenarios. Finally, we empirically show that our implicit calibration technique can be used in tandem with TCD during adaptation to further boost calibration in diverse domain shift scenarios.
Muhammad Akhtar Munir, Muhammad Haris Khan, M. Saquib Sarfraz, Mohsen Ali
NeurIPS4
2022 FogAdapt: Self-supervised domain adaptation for semantic segmentation of foggy images
Javed Iqbal 0007, Rehan Hafiz, Mohsen Ali
Neurocomputing3
2022 Distribution regularized self-supervised learning for domain adaptation of semantic segmentation
Javed Iqbal 0007, Hamza Rawal, Rehan Hafiz, Yu-Tseh Chi, Mohsen Ali
Image Vis. Comput.5
2022 Mapping Temporary Slums From Satellite Imagery Using a Semi-Supervised Approach
abstract
One billion people worldwide are estimated to be living in slums, and documenting and analyzing these regions is a challenging task. As compared to regular slums; the small, scattered and temporary nature of temporary slums makes data collection and labeling tedious and time-consuming. To tackle this challenging problem of temporary slums detection, we present a semi-supervised deep learning segmentation-based approach; with the strategy to detect initial seed images in the zero-labeled data settings. A small set of seed samples (32 in our case) are automatically discovered by analyzing the temporal changes, which are manually labeled to train a segmentation and representation learning module. The segmentation module gathers high dimensional image representations, and the representation learning module transforms image representations into embedding vectors. After that, a scoring module uses the embedding vectors to sample images from a large pool of unlabeled images and generates pseudo-labels for the sampled images. These sampled images with their pseudo-labels are added to the training set to update the segmentation and representation learning modules iteratively. To analyze the effectiveness of our technique, we construct a large geographically marked dataset of temporary slums. This dataset constitutes more than 200 potential temporary slum locations (2.28 square kilometers) found by sieving sixty-eight thousand images from 12 metropolitan cities of Pakistan covering 8000 square kilometers. Furthermore, our proposed method outperforms several competitive semi-supervised semantic segmentation baselines on a similar setting. The code and the dataset will be made publicly available.
M. Fasi ur Rehman, Izza Aftab, Waqas Sultani, Mohsen Ali
IEEE Geosci. Remote. Sens. Lett.4
2022 A dataset and benchmark for malaria life-cycle classification in thin blood smear images
Qazi Ammar Arshad, Mohsen Ali, Saeed-Ul Hassan, Chen Chen 0001, Ayisha Imran, Ghulam Rasul, Waqas Sultani
Neural Comput. Appl.2
2021 SSAL: Synergizing between Self-Training and Adversarial Learning for Domain Adaptive Object Detection
abstract
We study adapting trained object detectors to unseen domains manifesting significant variations of object appearance, viewpoints and backgrounds. Most current methods align domains by either using image or instance-level feature alignment in an adversarial fashion. This often suffers due to the presence of unwanted background and as such lacks class-specific alignment. A common remedy to promote class-level alignment is to use high confidence predictions on the unlabelled domain as pseudo labels. These high confidence predictions are often fallacious since the model is poorly calibrated under domain shift. In this paper, we propose to leverage model’s predictive uncertainty to strike the right balance between adversarial feature alignment and class-level alignment. Specifically, we measure predictive uncertainty on class assignments and the bounding box predictions. Model predictions with low uncertainty are used to generate pseudo-labels for self-supervision, whereas the ones with higher uncertainty are used to generate tiles for an adversarial feature alignment stage. This synergy between tiling around the uncertain object regions and generating pseudo-labels from highly certain object regions allows us to capture both the image and instance level context during the model adaptation stage. We perform extensive experiments covering various domain shift scenarios. Our approach improves upon existing state-of-the-art methods with visible margins.
Muhammad Akhtar Munir, Muhammad Haris Khan, M. Saquib Sarfraz, Mohsen Ali
NeurIPS4
2021 Leveraging orientation for weakly supervised object detection with application to firearm localization
Javed Iqbal 0007, Muhammad Akhtar Munir, Arif Mahmood, Afsheen Rafaqat Ali, Mohsen Ali
Neurocomputing5
2020 Learning from Scale-Invariant Examples for Domain Adaptation in Semantic Segmentation
M. Naseer Subhani, Mohsen Ali
ECCV (22)2
2020 Localizing Firearm Carriers By Identifying Human-Object Pairs
abstract
Visual identification of gunmen in a crowd is a challenging problem, that requires resolving the association of a person with an object (firearm). We present a novel approach to address this problem, by defining human-object interaction (and non-interaction) bounding boxes. In a given image, human and firearms are separately detected. Each detected human is paired with each detected firearm, allowing us to create a paired bounding box that contains both object and the human. A network is trained to classify these paired-bounding-boxes into human carrying the identified firearm or not. Extensive experiments were performed to evaluate the effectiveness of the algorithm, including exploiting full pose of the human, hand-keypoints, and their association with the firearm. The knowledge of spatially localized features is key to the success of our method by using multi-size proposals with adaptive average pooling. We have also extended a previously existing firearm detection dataset, by adding more images and tagging in the extended dataset the human-firearm pairs (including bounding boxes for firearms and gunmen). The experimental results $({78.5 AP}_{hold})$ demonstrate effectiveness of the proposed method.
Abdul Basit 0019, Muhammad Akhtar Munir, Mohsen Ali, Naoufel Werghi, Arif Mahmood
ICIP3
2020 EpO-Net: Exploiting Geometric Constraints on Dense Trajectories for Motion Saliency
abstract
The existing approaches for salient motion segmentation are unable to explicitly learn geometric cues and often give false detections on prominent static objects. We exploit multiview geometric constraints to avoid such shortcomings. To handle the nonrigid background like a sea, we also propose a robust fusion mechanism between motion and appearance-based features. We find dense trajectories, covering every pixel in the video, and propose trajectory-based epipolar distances to distinguish between background and foreground regions. Trajectory epipolar distances are dataindependent and can be readily computed given a few features' correspondences between the images. We show that by combining epipolar distances with optical flow, a powerful motion network can be learned. Enabling the network to leverage both of these features, we propose a simple mechanism, we call input-dropout. Comparing the motion-only networks, we outperform the previous state of the art on DAVIS-2016 dataset by 5.2% in the mean IoU score. By robustly fusing our motion network with an appearance network using the input-dropout mechanism, we also outperform the previous methods on DAVIS-2016, 2017 and Segtrackv2 dataset.
Muhammad Faisal 0003, Ijaz Akhter, Mohsen Ali, Richard I. Hartley
WACV3
2020 MLSL: Multi-Level Self-Supervised Learning for Domain Adaptation with Spatially Independent and Semantically Consistent Labeling
abstract
Most of the recent Deep Semantic Segmentation algorithms suffer from large generalization errors, even when powerful hierarchical representation models, based on convolutional neural networks, have been employed. This could be attributed to limited training data and large distribution gap in train and test domain datasets. In this paper, we propose a multi-level self-supervised learning model for domain adaptation of semantic segmentation. Exploiting the idea that an object (and most of the stuff given context) should be labeled consistently regardless of its location, we generate spatially independent and semantically consistent (SISC) pseudo-labels by segmenting multiple sub-images using base model and designing an aggregation strategy. Image level pseudo weak-labels, PWL, are computed to guide domain adaptation by capturing global context similarity in source and target domain at latent space level. Thus helping latent space learn the representation even when there are very few pixels belonging to the domain category (small object for example) compared to rest of the image. Our multi-level Self-supervised learning (MLSL) outperforms existing state-of-art (self or adversarial learning) algorithms. Specifically, keeping all setting similar and employing MLSL we obtain an mIoUgain of 5.1% on GTA-V to Cityscapes adaptation and 4.3% on SYNTHIA to Cityscapes adaptation compared to existing state-of-art method.
Javed Iqbal 0007, Mohsen Ali
WACV2
2019 Virtual learning environment to predict withdrawal by leveraging deep learning
abstract
The current evolution in multidisciplinary learning analytics research poses significant challenges for the exploitation of behavior analysis by fusing data streams toward advanced decision-making. The identification of students that are at risk of withdrawals in higher education is connected to numerous educational policies, to enhance their competencies and skills through timely interventions by academia. Predicting student performance is a vital decision-making problem including data from various environment modules that can be fused into a homogenous vector to ascertain decision-making. This research study exploits a temporal sequential classification problem to predict early withdrawal of students, by tapping the power of actionable smart data in the form of students' interactional activities with the online educational system, using the freely available Open University Learning Analytics data set by employing deep long short-term memory (LSTM) model. The deployed LSTM model outperforms baseline logistic regression and artificial neural networks by 10.31% and 6.48% respectively with 97.25% learning accuracy, 92.79% precision, and 85.92% recall.
Saeed-Ul Hassan, Hajra Waheed, Naif R. Aljohani, Mohsen Ali, Sebastián Ventura, Francisco Herrera
Int. J. Intell. Syst.4
2019 Layered convolutional dictionary learning for sparse coding itemsets
Sameen Mansha, Hoang Thanh Lam, Hongzhi Yin, Faisal Kamiran, Mohsen Ali
World Wide Web5
2017 Automatic Image transformation for inducing affect
Mohsen Ali, Afsheen Rafaqat Ali
BMVC1
2017 More for less: Insights into convolutional nets for 3D point cloud recognition
abstract
With the recent breakthrough in commodity 3D imaging solutions such as depth sensing, photogrammetry, stereoscopic vision and structured light, 3D shape recognition is becoming an increasingly important problem. A longstanding question is what should be the format of the 3D shape (such as voxel, mesh, point-cloud etc.) and what could be a good generic feature representation for shape recognition. This question is particularly important in the context of convolutional neural network (CNN) whose efficacy and complexity depends upon the choice of input shape format and the design of network. It has been seen that both 3D voxel representation as well as collection of rendered views on 2D images have produced competing results. Similarly, it have been seen that networks with few million parameters and networks with several hundred million parameters have similar performance. In this work we compare these solutions and provide an analysis on the factors resulting in increase in the parameters without significantly improving accuracy. On the basis of the above analysis we propose a representation method (point cloud to 2D grid) and architecture that results in much less parameters for the CNN but has competing accuracy.
Usama Shafiq, Murtaza Taj, Mohsen Ali
ICIP3
2017 High-Level Concepts for Affective Understanding of Images
abstract
This paper aims to bridge the affective gap between image content and the emotional response of the viewer it elicits by using High-Level Concepts (HLCs). In contrast to previous work that relied solely on low-level features or used convolutional neural network (CNN) as a blackbox, we use HLCs generated by pretrained CNNs in an explicit way to investigate the relations/associations between these HLCs and a (small) set of Ekman's emotional classes. As a proof-of-concept, we first propose a linear admixture model for modeling these relations, and the resulting computational framework allows us to determine the associations between each emotion class and certain HLCs (objects and places). This linear model is further extended to a nonlinear model using support vector regression (SVR) that aims to predict the viewer's emotional response using both low-level image features and HLCs extracted from images. These class-specific regressors are then assembled into a regressor ensemble that provide a flexible and effective predictor for predicting viewer's emotional responses from images. Experimental results have demonstrated that our results are comparable to existing methods, with a clear view of the association between HLCs and emotional classes that is ostensibly missing in most existing work.
Afsheen Rafaqat Ali, Usman Shahid, Mohsen Ali, Jeffrey Ho
WACV3
2014 Deconstructing Binary Classifiers in Computer Vision
Mohsen Ali, Jeffrey Ho
ACCV (3)1
2014 Deconstructing Kernel Machines
Mohsen Ali, Muhammad Ali Rushdi 0001, Jeffrey Ho
ECML/PKDD (1)1
2013 Block and Group Regularized Sparse Modeling for Dictionary Learning
abstract
This paper proposes a dictionary learning framework that combines the proposed block/group (BGSC) or reconstructed block/group (R-BGSC) sparse coding schemes with the novel Intra-block Coherence Suppression Dictionary Learning algorithm. An important and distinguishing feature of the proposed framework is that all dictionary blocks are trained simultaneously with respect to each data group while the intra-block coherence being explicitly minimized as an important objective. We provide both empirical evidence and heuristic support for this feature that can be considered as a direct consequence of incorporating both the group structure for the input data and the block structure for the dictionary in the learning process. The optimization problems for both the dictionary learning and sparse coding can be solved efficiently using block-gradient descent, and the details of the optimization algorithms are presented. We evaluate the proposed methods using well-known datasets, and favorable comparisons with state-of-the-art dictionary learning methods demonstrate the viability and validity of the proposed framework.
Yu-Tseh Chi, Mohsen Ali, Jeffrey Ho
CVPR2
2013 Affine-Constrained Group Sparse Coding and Its Application to Image-Based Classifications
abstract
This paper proposes a novel approach for sparse coding that further improves upon the sparse representation-based classification (SRC) framework. The proposed framework, Affine-Constrained Group Sparse Coding (ACGSC), extends the current SRC framework to classification problems with multiple input samples. Geometrically, the affineconstrained group sparse coding essentially searches for the vector in the convex hull spanned by the input vectors that can best be sparse coded using the given dictionary. The resulting objective function is still convex and can be efficiently optimized using iterative block-coordinate descent scheme that is guaranteed to converge. Furthermore, we provide a form of sparse recovery result that guarantees, at least theoretically, that the classification performance of the constrained group sparse coding should be at least as good as the group sparse coding. We have evaluated the proposed approach using three different recognition experiments that involve illumination variation of faces and textures, and face recognition under occlusions. Preliminary experiments have demonstrated the effectiveness of the proposed approach, and in particular, the results from the recognition/occlusion experiment are surprisingly accurate and robust.
Yu-Tseh Chi, Mohsen Ali, Muhammad Ali Rushdi 0001, Jeffrey Ho
ICCV2
2013 Color de-rendering using coupled dictionary learning
abstract
Consumer-level digital cameras typically post-process raw captured image data to produce enhanced visually appealing output RGB images. Post-processing operations include color gamut compression, tone mapping and other non-linear color corrections. However, raw image data is needed for many computer vision applications such as photometric stereo, shape from shading, and color constancy. Recovering raw image data from RGB images is complicated by the high non-linearity of the post-processing operations. In this paper, we propose a coupled dictionary scheme to model the relationship between the raw and RGB color image spaces of consumer cameras. Dictionary learning is regularized by sparsity constraints on feature representation. As well, we explore a more elaborate variant of coupled dictionary schemes that models the feature coupling more accurately. We test the proposed dictionary learning schemes on many commercial camera datasets. Our experimental results show accurate recovery of raw image data that looks visually indistinguishable from the ground truth.
Muhammad Ali Rushdi 0001, Mohsen Ali, Jeffrey Ho
ICIP2