VLDB 2026 Research / reviewers in the wild / expert
Naimul Mefraz Khan
dblp:82/8212 · also Naimul Khan 0001
· DBLP profile ↗
32ranked-venue papers
9as first author
14since 2021 · last 2025
0000-0002-8229-0747ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CCAFF: Object Tracking Under Heavy OcclusionabstractDeep features have become a standard in many object-tracking frameworks, replacing traditional handcrafted methods for object representation. However, recent studies have shown that deep features do not outperform handcrafted features in matching under occlusion or re-identification. Many trackers are trained using standard benchmarks under ideal conditions, but they degrade significantly in real-world settings because features deteriorate over time. These features affect the similarity distance and can lead to identity switches. Thus, we propose two novel tracking approaches using only handcrafted features and an extended variant, Contextual Cross Attention Feature Fusion (CCAFF). Both methods use a class-level codebook to capture keypoint cues for feature representation. We evaluate identity preservation for both feature sets using objects under severe occlusion. The CCAFF feature embedding demonstrates an improvement across all metrics on the HOOT dataset with a 4.63% increase in IDF1 score while decreasing identity switches by 7.78% compared to the baseline model. Abdul Bhutta, Naimul Mefraz Khan, Ling Guan |
ISM | 2 |
| 2025 | DynaGuide: A generalizable dynamic guidance framework for zero-shot guided unsupervised semantic segmentationabstractZero-shot guided unsupervised image segmentation enables dense scene understanding without relying on target-domain annotations, making it particularly valuable in domains where labeled data is scarce. However, most existing approaches struggle to reconcile global semantic coherence with fine-grained boundary precision. This paper introduces DynaGuide, an adaptive segmentation framework that addresses this challenge through a novel dual-guidance strategy and dynamic loss optimization. Building on our prior work, DynaSeg, DynaGuide integrates global pseudo-labels with local boundary refinement via a lightweight CNN trained from scratch. Crucially, the global pseudo-labels can originate either from a fully unsupervised source, such as DiffSeg, or from a supervised-pretrained model such as SegFormer. In both cases, these models act only as frozen priors on unseen data, ensuring that DynaGuide itself trains entirely without ground-truth labels in the target domain. Training is driven by a multi-component loss that dynamically balances feature similarity, Huber-smoothed spatial continuity (including diagonal relationships), and semantic alignment with the global pseudo-labels. Extensive experiments on BSD500, PASCAL VOC2012, and COCO demonstrate that DynaGuide achieves state-of-the-art performance, improving mIoU by 17.5% on BSD500, 3.1% on PASCAL VOC2012, and 11.66% on COCO. With its modular design, strong generalization, and minimal computational footprint, DynaGuide offers a scalable and practical solution for zero-shot guided unsupervised segmentation in real-world settings. • Proposes DynaGuide: a dual-guidance framework for zero-shot unsupervised segmentation. • Combines static global pseudo-labels with dynamic local CNN refinement. • Introduces adaptive multi-loss: feature similarity, diagonal Huber continuity, and global guidance. • Trains fully label-free using DiffSeg or SegFormer pseudo-labels without fine-tuning. • Outperforms recent SOTA on BSD500, PASCAL VOC2012, and COCO with fewer parameters and FLOPs. Boujemaa Guermazi, Riadh Ksantini, Naimul Mefraz Khan |
Image Vis. Comput. | 3 |
| 2024 | Attentional Feature Fusion for Few-Shot LearningabstractFew-shot Learning (FSL) approaches aim to develop generalizable models that can classify novel data points with a small set of labeled training data for each class. FSL approaches have the potential to narrow down the performance gap between machines and humans, but it is challenging because humans can quickly adapt to new activities and make decisions based on organized and reusable concepts. However, existing FSL approaches learn complex feature representations ignoring the conceptual information. We propose an Attentional Feature Fusion for Few-shot Learning (AF3), a semi-supervised approach that combines features of multiple scales and utilizes prior knowledge to learn better human interpretable concepts. Attentional feature fusion involves merging features from various layers and branches through an attention mechanism that prioritizes different features through attentional weights. AF3 extracts more discriminative features by generating attention maps for both query and support images. We evaluated our AF3 model in FSL settings on three benchmark datasets, including a fine-grained image classification. Extensive experiments show that our AF3 model outperformed the state-of-the-art in the most challenging 5-way-5-shot learning tasks. Muhammad Rehman Zafar, Naimul Mefraz Khan |
IJCNN | 2 |
| 2024 | Sliding Window Check: Repairing Object IdentitiesabstractReal-time Multiple Object Tracking (MOT) solutions are popular due to their potential in various tasks. Most real-time oriented solutions use a single model for detection and tracking, and perform pairwise comparisons using past and present features to track objects in time. We find that using a single frame of data is not robust. Some algorithms learn appearance data by relying on additional networks, leading to structures unsuited for real-time performance. By storing multiple past features with a temporal sliding window and performing multiple optimized pairwise comparisons using the window, we are able to improve tracking performance at a small computational overhead. We do so by introducing a novel repair step to the association algorithm. We propose Sliding Window Check (SWC), a tracking algorithm which can be applied to many state-of-the-art one-shot trackers with consistent improvements to HOTA and IDF1 on MOT17 and MOT20. We also address the lack of track-repair algorithms, and concerns about MOTA penalizing track-repairing algorithms. Comparisons against 4 recent works show the efficacy of SWC. Geerthan Srikantharajah, Naimul Mefraz Khan |
ISM | 2 |
| 2024 | Cross-database and cross-channel electrocardiogram arrhythmia heartbeat classification based on unsupervised domain adaptation
Md. Niaz Imtiaz, Naimul Mefraz Khan |
Expert Syst. Appl. | 2 |
| 2024 | DynaSeg: A deep dynamic fusion method for unsupervised image segmentation incorporating feature similarity and spatial continuityabstractOur work tackles the fundamental challenge of image segmentation in computer vision, which is crucial for diverse applications. While supervised methods demonstrate proficiency, their reliance on extensive pixel-level annotations limits scalability. We introduce DynaSeg, an innovative unsupervised image segmentation approach that overcomes the challenge of balancing feature similarity and spatial continuity without relying on extensive hyperparameter tuning. Unlike traditional methods, DynaSeg employs a dynamic weighting scheme that automates parameter tuning, adapts flexibly to image characteristics, and facilitates easy integration with other segmentation networks. By incorporating a Silhouette Score Phase, DynaSeg prevents undersegmentation failures where the number of predicted clusters might converge to one. DynaSeg uses CNN-based and pre-trained ResNet feature extraction, making it computationally efficient and more straightforward than other complex models. Experimental results showcase state-of-the-art performance, achieving a 12.2% and 14.12% mIOU improvement over current unsupervised segmentation approaches on COCO-All and COCO-Stuff datasets, respectively. We provide qualitative and quantitative results on five benchmark datasets, demonstrating the efficacy of the proposed approach. Code available at \url{https://github.com/RyersonMultimediaLab/DynaSeg} Boujemaa Guermazi, Riadh Ksantini, Naimul Mefraz Khan |
Image Vis. Comput. | 3 |
| 2023 | Towards Efficient Multi-view Representation LearningabstractThe proliferation of deep neural networks (DNNs) has drawn unprecedented interest in the study of various contents such as image, audio, video, to name a few. However, due to the data-driven nature, the high computational requirement and slow running time are considered as Achilles’ heels of DNN-based algorithms, limiting the progress of DNNs in time-sensitive applications. Recently, distinct discriminant canonical correlation analysis network (DDCCANet), a multi-view neural network, has shown great generalizability across multiple application domains, both analytically and experimentally. However, Although the computational requirement and running time by DDCCANet are more manageable than DNN-based algorithms, they can be substantially further improved. This paper proposes two new algorithms for multi-view feature representation learning, namely incremental DDCCANet (IDDCCANet) with substantial save in computational memory and GPU-accelerated DDCCANet (GADDCCANet) with drastically accelerated running time, forming a practically significant platform for multi-view feature representation learning. To validate the power of the proposed algorithms, experiments are conducted on several data sets with different types of inputs (e.g., raw image pixels, classical and DNN-based features). Experimental results clearly show that the proposed algorithms provide promising solutions to address the two longstanding challenges. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Ling Guan |
ISM | 4 |
| 2022 | Pan-Tompkins++: A Robust Approach to Detect R-peaks in ECG SignalsabstractR-peak detection is crucial in electrocardiogram (ECG) signal processing as it is the basis of heart rate variability analysis. The Pan-Tompkins algorithm is the most widely used QRS complex detector for the monitoring of many cardiac diseases including arrhythmia detection. However, the performance of the Pan-Tompkins algorithm in detecting the QRS complexes degrades in low-quality and noisy signals. This article introduces Pan-Tompkins++, an improved Pan-Tompkins algorithm. A bandpass filter with a passband of 5–18 Hz followed by an N-point moving average filter has been applied to remove the noise without discarding the significant signal components. Pan-Tompkins++ uses three thresholds to distinguish between R-peaks and noise peaks. Rather than using a generalized equation, different rules are applied to adjust the thresholds based on the pattern of the signal for the accurate detection of R-peaks under significant changes in signal pattern. The proposed algorithm reduces the False Positive and False Negative detections, and hence improves the robustness and performance of Pan-Tompkins algorithm. Pan-Tompkins++ has been tested on four open source datasets. The experimental results show noticeable improvement for both R-peak detection and execution time. We achieve 2.8% and 1.8% reduction in FP and FN, respectively, and 2.2% increase in F-score on average across four datasets, with 33% reduction in execution time. We show specific examples to demonstrate that in situations w here the Pan-Tompkins algorithm fails to identify R-peaks, the proposed algorithm is found to be effective. The results have also been contrasted with other well-known R-peak detection algorithms. Md. Niaz Imtiaz, Naimul Mefraz Khan |
BIBM | 2 |
| 2022 | GRIHA: synthesizing 2-dimensional building layouts from images captured using a smart phone
Shreya Goyal, Naimul Mefraz Khan, Chiranjoy Chattopadhyay, Gaurav Bhatnagar |
Multim. Tools Appl. | 2 |
| 2022 | Interpretable Artificial Intelligence through Locality Guided Neural Networks
Randy Tan, Lei Gao 0001, Naimul Mefraz Khan, Ling Guan |
Neural Networks | 3 |
| 2021 | ECG Heart-Beat Classification Using Multimodal Image FusionabstractIn this paper, we present a novel Image Fusion Model (IFM) for ECG heart-beat classification to overcome the weaknesses of existing machine learning techniques that rely either on manual feature extraction or direct utilization of 1D raw ECG signal. At the input of IFM, we first convert the heart-beats of ECG into three different images using Gramian Angular Field (GAF), Recurrence Plot (RP) and Markov Transition Field (MTF) and then fuse these images to create a single imaging modality. We use AlexNet for feature ex-traction and classification and thus employ end-to-end deep learning. We perform experiments on PhysioNet’s MIT-BIH dataset for five different arrhythmias in accordance with the AAMI EC57 standard and on PTB diagnostics dataset for myocardial infarction (MI) classification. We achieved an state-of-an-art results in terms of prediction accuracy, precision and recall. Anika Tabassum, Ling Guan, Naimul Mefraz Khan |
ICASSP | 4 |
| 2021 | A two-stream heterogeneous network for action recognition based on skeleton and RGB modalitiesabstractRecent years, skeleton based action recognition with graph convolutional network (GCN) has achieved great success. However, since skeleton data only includes human body joints coordinates, other key information on actions is missing such as the subtle motion of hands, the objects the human is interacting, leading to an unsatisfactory performance. In this respect, the RGB data can offer help to recognize actions that skeleton-based methods have limitations on. In this work, we propose a novel two-stream heterogeneous network consisting of GCN and CNN networks for action recognition. Specifically, the GCN network takes the skeletal sequence as input to exploit skeleton information. For the RGB video, the CNN model, ResNet (2+1)D, is adapted to exploit RGB information. Afterwards, the discriminant canonical correlation analysis (DCCA) method is utilized to integrate the output feature maps from the skeleton and RGB streams, resulting in improved performance. Experimental results on the large-scale dataset NTU RGB+D show that the proposed model outperforms state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISM | 3 |
| 2021 | Integrating vertex and edge features with Graph Convolutional Networks for skeleton-based action recognition
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
Neurocomputing | 3 |
| 2021 | A Multi-Stream Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action RecognitionabstractRecently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN with HCRF to retain the human skeleton structure information even during the classification stage. Our model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework which takes the relative coordinate of the joints and bone direction as two static feature streams, and the temporal displacements between two consecutive frames as the dynamic feature stream. Experimental results on three challenging benchmarks (NTU RGB+D, N-UCLA, SYSU) show the superior performance of the proposed model over state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
IEEE Trans. Multim. | 3 |
| 2020 | Locality Guided Neural Networks for Explainable Artificial IntelligenceabstractIn current deep network architectures, deeper layers in networks tend to contain hundreds of independent neurons which makes it hard for humans to understand how they interact with each other. By organizing the neurons by correlation, humans can observe how clusters of neighbouring neurons interact with each other. In this paper, we propose a novel algorithm for back propagation, called Locality Guided Neural Network (LGNN) for training networks that preserves locality between neighbouring neurons within each layer of a deep network. Heavily motivated by Self-Organizing Map (SOM), the goal is to enforce a local topology on each layer of a deep network such that neighbouring neurons are highly correlated with each other. This method contributes to the domain of Explainable Artificial Intelligence (XAI), which aims to alleviate the black-box nature of current AI methods and make them understandable by humans. Our method aims to achieve XAI in deep learning without changing the structure of current models nor requiring any post processing. This paper focuses on Convolutional Neural Networks (CNNs), but can theoretically be applied to any type of deep learning architecture. In our experiments, we train various VGG and Wide ResNet (WRN) networks for image classification on CIFAR100. In depth analyses presenting both qualitative and quantitative results demonstrate that our method is capable of enforcing a topology on each layer while achieving a small increase in classification accuracy. Randy Tan, Naimul Mefraz Khan, Ling Guan |
IJCNN | 2 |
| 2020 | A Vertex-Edge Graph Convolutional Network for Skeleton-Based Action RecognitionabstractThe Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability of exploiting the joint information from the graph structure of the skeleton data. Recently, as a strong and complementary modality for action recognition, the bone information from skeleton data has attracted more attention. However, most existing GCN-based methods extract the bone and joint features with two separate GCN networks, ignoring the dependencies between joints and bones. In this work, a Vertex-Edge Graph Convolutional Network(VE-GCN) is proposed to reveal the information across joints, bones and their relationships simultaneously. In addition, we learn the additional connections among joints and bones for various action samples besides the natural connections of the skeleton. Then we conduct the convolution operation on joints and their neighbors based on these additional connections. Moreover, the conditional random field (CRF) is utilized as the loss function to achieve improved performance. Experimental results on two large-scale datasets NTU RGB+D and NTU RGB+D 120 show that the proposed model outperforms state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISCAS | 3 |
| 2020 | Deep clustering with a Dynamic Autoencoder: From reconstruction towards centroids construction
Nairouz Mrabah, Naimul Mefraz Khan, Riadh Ksantini, Zied Lachiri |
Neural Networks | 2 |
| 2019 | Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action RecognitionabstractRecently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN and HCRF to retain the human skeleton structure information during the classification stage. The proposed model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework that takes the relative coordinates of the joints and bone direction as two static feature streams and the temporal displacements as the dynamic feature stream. Experimental results on two challenging benchmarks (NTU RGB+D, N-UCLA) show the superior performance of the proposed model over state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISM | 3 |
| 2018 | Brain MRI Segmentation using efficient 3D Fully Convolutional Neural Networks
Ghazala Khan, Naimul Mefraz Khan |
BIBM | 2 |
| 2018 | Towards Improved Human Action Recognition Using Convolutional Neural Networks and Multimodal Fusion of Depth and Inertial Sensor DataabstractThis paper attempts at improving the accuracy of Human Action Recognition (HAR) by fusion of depth and inertial sensor data. Firstly, we transform the depth data into Sequential Front view Images(SFI) and fine-tune the pre-trained AlexNet on these images. Then, inertial data is converted into Signal Images (SI) and another convolutional neural network (CNN) is trained on these images. Finally, learned features are extracted from both CNN, fused together to make a shared feature layer, and these features are fed to the classifier. We experiment with two classifiers, namely Support Vector Machines (SVM) and softmax classifier and compare their performances. The recognition accuracies of each modality, depth data alone and sensor data alone are also calculated and compared with fusion based accuracies to highlight the fact that fusion of modalities yields better results than individual modalities. Experimental results on UTD-MHAD and Kinect 2D datasets show that proposed method achieves state of the art results when compared to other recently proposed visual-inertial action recognition methods. Naimul Mefraz Khan |
ISM | 2 |
| 2018 | Deep Reinforcement Learning with Parameterized Action Space for Object DetectionabstractObject detection is a fundamental task in computer vision. With the remarkable progress made in big visual data analytics and deep learning, Reinforcement Learning (RL) is becoming a promising framework to model the object detection problem since the detection procedure can be cast as a Markov decision process (MDP). We propose a Reinforcement Learning system with parameterized action space for image object detection. The proposed system uses an active agent exploring in a scene to identify the location of a target object, and learns a policy to refine the geometry of the agent by taking simple actions in parameterized space, which integrates the discrete actions and its corresponding continuous parameters. We then optimize the representation of the generated region proposals with the discriminative multiple canonical correlation analysis (DMCCA) [11] in preparation for classification with Fast R-CNN. Experiments on PASCAL VOC 2007 and 2012 datasets show the effectiveness of the proposed method. Naimul Mefraz Khan, Lei Gao 0001, Ling Guan |
ISM | 2 |
| 2018 | A Novel Image-Centric Approach Toward Direct Volume RenderingabstractTransfer function (TF) generation is a fundamental problem in direct volume rendering (DVR). A TF maps voxels to color and opacity values to reveal inner structures. Existing TF tools are complex and unintuitive for the users who are more likely to be medical professionals than computer scientists. In this article, we propose a novel image-centric method for TF generation where instead of complex tools, the user directly manipulates volume data to generate DVR. The user’s work is further simplified by presenting only the most informative volume slices for selection. Based on the selected parts, the voxels are classified using our novel sparse nonparametric support vector machine classifier, which combines both local and near-global distributional information of the training data. The voxel classes are mapped to aesthetically pleasing and distinguishable color and opacity values using harmonic colors. Experimental results on several benchmark datasets and a detailed user survey show the effectiveness of the proposed method. Naimul Mefraz Khan, Riadh Ksantini, Ling Guan |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2017 | Towards Alzheimer's disease classification through transfer learningabstractDetection of Alzheimer's Disease (AD) from neuroimaging data such as MRI through machine learning have been a subject of intense research in recent years. Recent success of deep learning in computer vision have progressed such research further. However, common limitations with such algorithms are reliance on a large number of training images, and requirement of careful optimization of the architecture of deep networks. In this paper, we attempt solving these issues with transfer learning, where state-of-the-art architectures such as VGG and Inception are initialized with pre-trained weights from large benchmark datasets consisting of natural images, and the fully-connected layer is re-trained with only a small number of MRI images. We employ image entropy to select the most informative slices for training. Through experimentation on the OASIS MRI dataset, we show that with training size almost 10 times smaller than the state-of-the-art, we reach comparable or even better performance than current deep-learning based methods. Marcia Hon, Naimul Mefraz Khan |
BIBM | 2 |
| 2017 | Real-Time System for Human Activity AnalysisabstractWe propose a real-time human activity analysis system, where a user's activity can be quantitatively evaluated with respect to a ground truth recording. We use two Kinects to solve the problem of self-occlusion through extracting optimal joint positions using Singular Value Decomposition (SVD) and Sequential Quadratic Programming (SQP). Incremental Dynamic Time Warping (IDTW) is used to compare the user and expert (ground truth) to quantitatively score the user's performance. Furthermore, the user's performance is displayed through a visual feedback system, where colors on the skeleton represent the user's score. Our experiments use a motion capture suit as ground truth to compare our dual Kinect setup to a single Kinect. We also show that with our visual feedback method, users gain a statistically significant boost to learning as opposed to watching a simple video. Randy Tan, Naimul Mefraz Khan, Ling Guan |
ISM | 2 |
| 2015 | Intuitive volume exploration through spherical self-organizing map and color harmonization
Naimul Mefraz Khan, Matthew J. Kyan, Ling Guan |
Neurocomputing | 1 |
| 2014 | Volume visualization using sparse nonparametric support vector machines and harmoniccolorsabstractIn Direct Volume Rendering (DVR), the Transfer Function (TF) to map voxel values to color and opacity values is difficult to obtain. Existing TF design tools are complex and non-intuitive for the end user, who is more likely to be a medical professional than an expert in image processing. In this paper, we propose a volume visualization method where the user directly works on the volume data to simply select the parts he/she would like to visualize. The user's work is further simplified by presenting only the most informative volume slices for selection. Based on the selected parts, all the voxels are classified using our Sparse Nonparametric Support Vector Machine (SN-SVM) classifier, which combines both local and near-global distributional information of the training data to obtain accurate results. The voxel classes are then mapped to color and opacity values using the concept of harmonic colors, which provides easily distinguishable and aesthetically pleasing results. Experimental results on several benchmark datasets show the effectiveness of the proposed method. Naimul Mefraz Khan, Riadh Ksantini, Ling Guan |
ICASSP | 1 |
| 2014 | A Visual Evaluation Framework for In-Home Physical RehabilitationabstractWe propose a novel method for in-home physical rehabilitation, where a user can visually evaluate his/her performance compared to that of an expert. Normalized joint coordinates extracted from the Kinect skeleton are used as features. A novel Incremental Dynamic Time Warping (IDTW) algorithm is used to align the user and expert sequences. IDTW extends the classic DTW by providing accurate comparison between incomplete (the user's) and complete (the expert's) sequences while significantly reducing the computational time. Instead of providing a single measurement, the proposed method maps the IDTW measurements to a color-coded skeleton frame. Different colors on the limbs provide the user with an easy-to-interpret evaluation of how he or she is performing. Preliminary analysis involving different users and exercises and comparisons against the classic DTW algorithm show the effectiveness of the proposed method. Naimul Mefraz Khan, Stephen Lin 0001, Ling Guan, Baining Guo |
ISM | 1 |
| 2014 | Covariance-guided One-Class Support Vector Machine
Naimul Mefraz Khan, Riadh Ksantini, Imran Ahmad 0001, Ling Guan |
Pattern Recognit. | 1 |
| 2013 | Incorporating covariance information in one class support vector classificationabstractUnlike multi-class problems, the low variance directions in the training data are important for one-class classification. However, projecting in these directions before classification will result in loss of important data properties. This paper introduces a Covariance-guided One-Class Support Vector Machine (COSVM) classification method which emphasizes the low variance projectional directions of the training data without compromising any important characteristics. COSVM combines the global information from the covariance matrix of the training data with the local information of Support Vectors. Our proposed method is a convex optimization problem resulting in one global solution, which can be found efficiently with the help of existing numerical methods. The method also keeps the principal structure of the OSVM method intact, and can be implemented easily with the existing OSVM applications. Comparative experimental results with contemporary one-class classifiers on numerous benchmark datasets verify that our method results in significantly better performance. Naimul Mefraz Khan, Riadh Ksantini, Imran Ahmad 0001, Ling Guan |
ICASSP | 1 |
| 2012 | A Sparse Support Vector Machine Classifier with Nonparametric Discriminants
Naimul Mefraz Khan, Riadh Ksantini, Imran Ahmad 0001, Ling Guan |
ICANN (2) | 1 |
| 2012 | A novel SVM+NDA model for classification with an application to face recognition
Naimul Mefraz Khan, Riadh Ksantini, Imran Ahmad 0001, Boubakeur Boufama |
Pattern Recognit. | 1 |
| 2009 | A new signature for quadtree-based image matchingabstractA hierarchical two-level indexing scheme for retrieval of spatially similar images has been proposed in [1]. In this scheme, the first level of indexing for identification of potentially relevant images involves matching of image signatures and the second level of indexing is based on quadtree matching and involves more intensive computations. This method provides a significant performance improvement over the other contemporary methods. However, since the number of comparisons in the second level is based on the results of signature matching, a well devised signature representation scheme can result in significant improvement in the overall efficiency of the entire system. At the same time, a signature representation scheme is required to be such that it does not result in any false negative. In this paper we propose a new image signature representation scheme that is entirely based on the quadtree representation of an image. We also formally prove that the proposed signature representation scheme not only results in fewer number of matching signatures but also does not result in any false negative. Naimul Mefraz Khan, Imran Ahmad 0001 |
MoMM | 1 |