Raghvendra Mall

dblp:15/8321 · DBLP profile ↗
← Back
31ranked-venue papers
16as first author
4since 2021 · last 2024
0000-0003-1779-3150ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 11 first-authorApplied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 FEATURE-pHLA: Physico-chemical features efficiently predict peptide-HLA binding affinity
abstract
Human leukocyte antigen or HLA plays a crucial role in the recognition of antigenic peptides as this binding is responsible for subsequent immune response by eliciting T-cell activation. Accurate prediction of peptide-HLA binding affinity is imperative for facilitating vaccine development and immunotherapies. Recent advancements in transformer-based models and protein language models in predicting peptide-HLA interactions have shown significant improvements. Current methodologies rely on deep learning methods and GPU-intensive computations. We propose a simple and computationally cheaper method that demonstrates efficacy. Our tree-based model, named FEATUREPHLA, utilizes the physico-chemical fingerprints obtained from peptides and HLA sequences and is highly interpretable.Our goal was to estimate the predictive efficacy of these physico-chemical features for the task of peptide-HLA binding prediction. Our proposed method outperforms other methods on experimentally verified peptide-HLA binders from the HPV vaccine data securing the highest number of true positives and the lowest number of false negatives, thereby, showcasing its predictive power on real-world scenarios. Our study reveals the relevance of biology-inspired features for the calculation of molecular interactions and lays the groundwork towards developing more accurate biology-informed predictive models.
Hamda Alhosani, Raghvendra Mall, Ankita Singh, Filippo Castiglione
BIBM2
2024 VISH-Pred: an ensemble of fine-tuned ESM models for protein toxicity prediction
abstract
Peptide- and protein-based therapeutics are becoming a promising treatment regimen for myriad diseases. Toxicity of proteins is the primary hurdle for protein-based therapies. Thus, there is an urgent need for accurate in silico methods for determining toxic proteins to filter the pool of potential candidates. At the same time, it is imperative to precisely identify non-toxic proteins to expand the possibilities for protein-based biologics. To address this challenge, we proposed an ensemble framework, called VISH-Pred, comprising models built by fine-tuning ESM2 transformer models on a large, experimentally validated, curated dataset of protein and peptide toxicities. The primary steps in the VISH-Pred framework are to efficiently estimate protein toxicities taking just the protein sequence as input, employing an under sampling technique to handle the humongous class-imbalance in the data and learning representations from fine-tuned ESM2 protein language models which are then fed to machine learning techniques such as Lightgbm and XGBoost. The VISH-Pred framework is able to correctly identify both peptides/proteins with potential toxicity and non-toxic proteins, achieving a Matthews correlation coefficient of 0.737, 0.716 and 0.322 and F1-score of 0.759, 0.696 and 0.713 on three non-redundant blind tests, respectively, outperforming other methods by over $10\%$ on these quality metrics. Moreover, VISH-Pred achieved the best accuracy and area under receiver operating curve scores on these independent test sets, highlighting the robustness and generalization capability of the framework. By making VISH-Pred available as an easy-to-use web server, we expect it to serve as a valuable asset for future endeavors aimed at discerning the toxicity of peptides and enabling efficient protein-based therapeutics.
Raghvendra Mall, Ankita Singh, Chirag N. Patel, Gregory Guirimand, Filippo Castiglione
Briefings Bioinform.1
2021 Network-based identification of key master regulators associated with an immune-silent cancer phenotype
abstract
A cancer immune phenotype characterized by an active T-helper 1 (Th1)/cytotoxic response is associated with responsiveness to immunotherapy and favorable prognosis across different tumors. However, in some cancers, such an intratumoral immune activation does not confer protection from progression or relapse. Defining mechanisms associated with immune evasion is imperative to refine stratification algorithms, to guide treatment decisions and to identify candidates for immune-targeted therapy. Molecular alterations governing mechanisms for immune exclusion are still largely unknown. The availability of large genomic datasets offers an opportunity to ascertain key determinants of differential intratumoral immune response. We follow a network-based protocol to identify transcription regulators (TRs) associated with poor immunologic antitumor activity. We use a consensus of four different pipelines consisting of two state-of-the-art gene regulatory network inference techniques, regularized gradient boosting machines and ARACNE to determine TR regulons, and three separate enrichment techniques, including fast gene set enrichment analysis, gene set variation analysis and virtual inference of protein activity by enriched regulon analysis to identify the most important TRs affecting immunologic antitumor activity. These TRs, referred to as master regulators (MRs), are unique to immune-silent and immune-active tumors, respectively. We validated the MRs coherently associated with the immune-silent phenotype across cancers in The Cancer Genome Atlas and a series of additional datasets in the Prediction of Clinical Outcomes from Genomic Profiles repository. A downstream analysis of MRs specific to the immune-silent phenotype resulted in the identification of several enriched candidate pathways, including NOTCH1, TGF-$\beta $, Interleukin-1 and TNF-$\alpha $ signaling pathways. TGFB1I1 emerged as one of the main negative immune modulators preventing the favorable effects of a Th1/cytotoxic response.
Raghvendra Mall, Mohamad Saad 0001, Jessica Roelands, Darawan Rinchai, Khalid Kunji, Hossam Almeer, Wouter Hendrickx, Francesco M. Marincola, Michele Ceccarelli, Davide Bedognetti
Briefings Bioinform.1
2021 A modeling framework for embedding-based predictions for compound-viral protein activity
abstract
MOTIVATION: A global effort is underway to identify compounds for the treatment of COVID-19. Since de novo compound design is an extremely long, time-consuming and expensive process, efforts are underway to discover existing compounds that can be repurposed for COVID-19 and new viral diseases.We propose a machine learning representation framework that uses deep learning induced vector embeddings of compounds and viral proteins as features to predict compound-viral protein activity. The prediction model in-turn uses a consensus framework to rank approved compounds against viral proteins of interest. RESULTS: Our consensus framework achieves a high mean Pearson correlation of 0.916, mean R2 of 0.840 and a low mean squared error of 0.313 for the task of compound-viral protein activity prediction on an independent test set. As a use case, we identify a ranked list of 47 compounds common to three main proteins of SARS-COV-2 virus (PL-PRO, 3CL-PRO and Spike protein) as potential targets including 21 antivirals, 15 anticancer, 5 antibiotics and 6 other investigational human compounds. We perform additional molecular docking simulations to demonstrate that majority of these compounds have low binding energies and thus high binding affinity with the potential to be effective against the SARS-COV-2 virus. AVAILABILITY AND IMPLEMENTATION: All the source code and data is available at: https://github.com/raghvendra5688/Drug-Repurposing and https://dx.doi.org/10.17632/8rrwnbcgmx.3. We also implemented a web-server at: https://machinelearning-protein.qcri.org/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Raghvendra Mall, Abdurrahman Elbasir, Hossam Almeer, Zeyaul Islam, Prasanna R. Kolatkar, Sanjay Chawla, Ehsan Ullah
Bioinform.1
2020 BCrystal: an interpretable sequence-based protein crystallization predictor
abstract
MOTIVATION: X-ray crystallography has facilitated the majority of protein structures determined to date. Sequence-based predictors that can accurately estimate protein crystallization propensities would be highly beneficial to overcome the high expenditure, large attrition rate, and to reduce the trial-and-error settings required for crystallization. RESULTS: In this study, we present a novel model, BCrystal, which uses an optimized gradient boosting machine (XGBoost) on sequence, structural and physio-chemical features extracted from the proteins of interest. BCrystal also provides explanations, highlighting the most important features for the predicted crystallization propensity of an individual protein using the SHAP algorithm. On three independent test sets, BCrystal outperforms state-of-the-art sequence-based methods by more than 12.5% in accuracy, 18% in recall and 0.253 in Matthew's correlation coefficient, with an average accuracy of 93.7%, recall of 96.63% and Matthew's correlation coefficient of 0.868. For relative solvent accessibility of exposed residues, we observed higher values to associate positively with protein crystallizability and the number of disordered regions, fraction of coils and tripeptide stretches that contain multiple histidines associate negatively with crystallizability. The higher accuracy of BCrystal enables it to accurately screen for sequence variants with enhanced crystallizability. AVAILABILITY AND IMPLEMENTATION: Our BCrystal webserver is at https://machinelearning-protein.qcri.org/ and source code is available at https://github.com/raghvendra5688/BCrystal. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Abdurrahman Elbasir, Raghvendra Mall, Khalid Kunji, Reda Rawi, Zeyaul Islam, Gwo-Yu Chuang, Prasanna R. Kolatkar, Halima Bensmail
Bioinform.2
2019 DeepCrystal: a deep learning framework for sequence-based protein crystallization prediction
abstract
MOTIVATION: Protein structure determination has primarily been performed using X-ray crystallography. To overcome the expensive cost, high attrition rate and series of trial-and-error settings, many in-silico methods have been developed to predict crystallization propensities of proteins based on their sequences. However, the majority of these methods build their predictors by extracting features from protein sequences, which is computationally expensive and can explode the feature space. We propose DeepCrystal, a deep learning framework for sequence-based protein crystallization prediction. It uses deep learning to identify proteins which can produce diffraction-quality crystals without the need to manually engineer additional biochemical and structural features from sequence. Our model is based on convolutional neural networks, which can exploit frequently occurring k-mers and sets of k-mers from the protein sequences to distinguish proteins that will result in diffraction-quality crystals from those that will not. RESULTS: Our model surpasses previous sequence-based protein crystallization predictors in terms of recall, F-score, accuracy and Matthew's correlation coefficient (MCC) on three independent test sets. DeepCrystal achieves an average improvement of 1.4, 12.1% in recall, when compared to its closest competitors, Crysalis II and Crysf, respectively. In addition, DeepCrystal attains an average improvement of 2.1, 6.0% for F-score, 1.9, 3.9% for accuracy and 3.8, 7.0% for MCC w.r.t. Crysalis II and Crysf on independent test sets. AVAILABILITY AND IMPLEMENTATION: The standalone source code and models are available at https://github.com/elbasir/DeepCrystal and a web-server is also available at https://deeplearning-protein.qcri.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Abdurrahman Elbasir, Balasubramanian Moovarkumudalvan, Khalid Kunji, Prasanna R. Kolatkar, Raghvendra Mall, Halima Bensmail
Bioinform.5
2018 DeepCrystal: A Deep Learning Framework for Sequence-based Protein Crystallization Prediction
Abdurrahman Elbasir, Balasubramanian Moovarkumudalvan, Khalid Kunji, Prasanna R. Kolatkar, Halima Bensmail, Raghvendra Mall
BIBM6
2018 DeepSol: a deep learning framework for sequence-based protein solubility prediction
abstract
Motivation: Protein solubility plays a vital role in pharmaceutical research and production yield. For a given protein, the extent of its solubility can represent the quality of its function, and is ultimately defined by its sequence. Thus, it is imperative to develop novel, highly accurate in silico sequence-based protein solubility predictors. In this work we propose, DeepSol, a novel Deep Learning-based protein solubility predictor. The backbone of our framework is a convolutional neural network that exploits k-mer structure and additional sequence and structural features extracted from the protein sequence. Results: DeepSol outperformed all known sequence-based state-of-the-art solubility prediction methods and attained an accuracy of 0.77 and Matthew's correlation coefficient of 0.55. The superior prediction accuracy of DeepSol allows to screen for sequences with enhanced production capacity and can more reliably predict solubility of novel proteins. Availability and implementation: DeepSol's best performing models and results are publicly deposited at https://doi.org/10.5281/zenodo.1162886 (Khurana and Mall, 2018). Supplementary information: Supplementary data are available at Bioinformatics online.
Sameer Khurana, Reda Rawi, Khalid Kunji, Gwo-Yu Chuang, Halima Bensmail, Raghvendra Mall
Bioinform.6
2018 PaRSnIP: sequence-based protein solubility prediction using gradient boosting machine
abstract
Motivation: Protein solubility can be a decisive factor in both research and production efficiency, and in silico sequence-based predictors that can accurately estimate solubility outcomes are highly sought. Results: In this study, we present a novel approach termed PRotein SolubIlity Predictor (PaRSnIP), which uses a gradient boosting machine algorithm as well as an approximation of sequence and structural features of the protein of interest. Based on an independent test set, PaRSnIP outperformed other state-of-the-art sequence-based methods by more than 9% in accuracy and 0.17 in Matthew's correlation coefficient, with an overall accuracy of 74% and Matthew's correlation coefficient of 0.48. Additionally, PaRSnIP provides importance scores for all features used in training. We observed higher fractions of exposed residues to associate positively with protein solubility and tripeptide stretches with multiple histidines to associate negatively with solubility. The improved prediction accuracy of PaRSnIP should enable it to predict protein solubility with greater reliability and to screen for sequence variants with enhanced manufacturability. Availability and implementation: PaRSnIP software is available for download under GitHub (https://github.com/RedaRawi/PaRSnIP). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Reda Rawi, Raghvendra Mall, Khalid Kunji, Chen-Hsiang Shen, Peter D. Kwong, Gwo-Yu Chuang
Bioinform.2
2017 An adaptive refinement for community detection methods for disease module identification in biological networks using novel metric based on connectivity, conductance & modularity
abstract
Disease processes are usually driven by several genes interacting in molecular modules or pathways leading to the disease. The identification of such modules in gene or protein networks is at the core of several analysis methods in biomedical research. However, there is still a need to develop a generic framework to uncover biologically relevant modules for different types of networks. With this pretext in mind, the Disease Module Identification DREAM Challenge was initiated as an effort to systematically assess module identification methods on a panel of 6 diverse state-of-the-art genomic networks. Methods: In this paper, we propose a generic refinement method based on ideas of merging and splitting the hierarchical tree obtained from any community detection technique for constrained disease module identification in biological networks. The only constraint for a module to be considered as a candidate disease module was size of the community to be: 3 ≤ community size ≤ 100. Here, we propose a novel quality metric, called F-score, computed from several unsupervised quality metrics like modularity, conductance and connectivity to determine the quality of a graph partition at a given level of hierarchy. We also propose a quality metric, namely Inverse Confidence, which ranks and prune insignificant modules to obtain a curated list of candidate disease modules for a given biological network. The predicted modules are then evaluated on the basis of the total number of unique candidate modules that are associated with complex traits and diseases from over 200 genome-wide association study (GWAS) datasets. Results: We stood 9thout of a total of 42 teams in the competition at the offical FDR cut-off of 0.05 for identifying statistically significant disease associated modules in the 6 benchmark networks. Our proposed approach detected a total of 44 disease modules in the 6 benchmark networks in comparison to 60 for the winner of the DREAM Challenge. For several benchmark networks we were better or competitive with the winner.
Raghvendra Mall, Ehsan Ullah, Khalid Kunji, Halima Bensmail, Michele Ceccarelli
BIBM1
2017 Identification of cancer drug sensitivity biomarkers
abstract
Anti-cancer therapies have different responses to different patients. There is a need to identify biomarkers for effectiveness of drugs beside the biomarkers for the diseases like cancer to advance the field personalized medicine. We have used a panel of cancer cell lines from Genomics of Drug Sensitivity in Cancer (GDSC) to capture the sensitivity of drugs. By combining the genetic information such as mutations, copy number variations and gene expression levels from Catalogue of Somatic Mutations in Cancer (COSMIC) and The Cancer Genome ATLAS (TCGA) for the afore mentioned cell lines, we are able to identify genomic features associated with the sensitivity of the drugs.
Ehsan Ullah, Raghvendra Mall, Halima Bensmail, Reda Rawi, Saila Shama, Nooral Al Muftah, Ian Richard Thmpson
BIBM2
2016 Fast in-memory spectral clustering using a fixed-size approach
Rocco Langone, Raghvendra Mall, Vilen Jumutc, Johan A. K. Suykens
ESANN2
2016 Denoised Kernel Spectral data Clustering
abstract
Kernel Spectral Clustering (KSC) solves a weighted kernel principal component analysis problem in a primal-dual optimization framework. It builds an unsupervised model on a small subset of data using the dual solution of the optimization problem. This allows KSC to have a powerful out-of-sample extension property leading to good cluster generalization w.r.t. unseen data points. However, in the presence of noise that causes overlapping data, the technique often fails to provide good generalization capability. In this paper, we propose a two-step process for clustering noisy data. We first denoise the data using kernel principal component analysis (KPCA) with a recently proposed Model selection criterion based on point-wise Distance Distributions (MDD) to obtain the underlying information in the data. We then use the KSC technique on this denoised data to obtain good quality clusters. One advantage of model based techniques is that we can use the same training and validation set for denoising and for clustering. We discovered that using the same kernel bandwidth parameter obtained from MDD for KPCA works efficiently with KSC in combination with the optimal number of clusters k to produce good quality clusters. We compare the proposed approach with normal KSC and KSC with KPCA using a heuristic method based on reconstruction error for several synthetic and real-world datasets to showcase the effectiveness of the proposed approach.
Raghvendra Mall, Halima Bensmail, Rocco Langone, Carolina Varon, Johan A. K. Suykens
IJCNN1
2016 COUSCOus: improved protein contact prediction using an empirical Bayes covariance estimator
abstract
BACKGROUND: The post-genomic era with its wealth of sequences gave rise to a broad range of protein residue-residue contact detecting methods. Although various coevolution methods such as PSICOV, DCA and plmDCA provide correct contact predictions, they do not completely overlap. Hence, new approaches and improvements of existing methods are needed to motivate further development and progress in the field. We present a new contact detecting method, COUSCOus, by combining the best shrinkage approach, the empirical Bayes covariance estimator and GLasso. RESULTS: Using the original PSICOV benchmark dataset, COUSCOus achieves mean accuracies of 0.74, 0.62 and 0.55 for the top L/10 predicted long, medium and short range contacts, respectively. In addition, COUSCOus attains mean areas under the precision-recall curves of 0.25, 0.29 and 0.30 for long, medium and short contacts and outperforms PSICOV. We also observed that COUSCOus outperforms PSICOV w.r.t. Matthew's correlation coefficient criterion on full list of residue contacts. Furthermore, COUSCOus achieves on average 10% more gain in prediction accuracy compared to PSICOV on an independent test set composed of CASP11 protein targets. Finally, we showed that when using a simple random forest meta-classifier, by combining contact detecting techniques and sequence derived features, PSICOV predictions should be replaced by the more accurate COUSCOus predictions. CONCLUSION: We conclude that the consideration of superior covariance shrinkage approaches will boost several research fields that apply the GLasso procedure, amongst the presented one of residue-residue contact prediction as well as fields such as gene network reconstruction.
Reda Rawi, Raghvendra Mall, Khalid Kunji, Mohammed El Anbari, Michaël Aupetit 0001, Ehsan Ullah, Halima Bensmail
BMC Bioinform.2
2015 Ranking Overlap and Outlier Points in Data using Soft Kernel Spectral Clustering
Raghvendra Mall, Rocco Langone, Johan A. K. Suykens
ESANN1
2015 Kernel spectral document clustering using unsupervised precision-recall metrics
abstract
Kernel Spectral Clustering (KSC) solves a weighted kernel principal component analysis problem in a primal-dual optimization framework. The KSC model is built on a small subset of data using a proper training, model selection and a test phase. The clustering model is obtained using the dual solution of the problem and has a powerful out-of-sample extensions property which allows cluster affiliation for previously unseen data points. In the model selection phase, we estimate the appropriate number of clusters using a metric that evaluates the quality of the clusters. Traditional quality indices like inertia, Davies-Bouldin (DB) index and silhouette (SIL) are known to be method-dependent and not perform well in case of complex heterogeneous data like textual data. In this paper, we utilize the quality evaluation techniques based on an unsupervised version of Precision, Recall and F-measure proposed in [1] to come up with a new kernel spectral document clustering (KSDC) model which generates homogeneous clusters of documents. We compare the quality of the clusters obtained by the proposed KSDC technique with k-means and neural gas algorithm, which are more oriented towards these metrics, on several real world textual data.
Raghvendra Mall, Johan A. K. Suykens
IJCNN1
2015 Hierarchical semi-supervised clustering using KSC based model
abstract
This paper introduces a methodology to incorporate the label information in discovering the underlying clusters in a hierarchical setting using multi-class semi-supervised clustering algorithm. The method aims at revealing the relationship between clusters given few labels associated to some of the clusters. The problem is formulated as a regularized kernel spectral clustering algorithm in the primal-dual setting. The available labels are incorporated in different levels of hierarchy from top to bottom. As we advance towards the lowers levels in the tree all the previously added labels are used in the generation of the new levels of hierarchy. The model is trained on a subset of the data and then applied to the rest of the data in a learning framework. Thanks to the previously learned model, the out-of-sample extension property of the model allows then to predict the memberships of a new point. A combination of an internal clustering quality index and classification accuracy is used for model selection. Experiments are conducted on synthetic data and real image segmentation problems to show the applicability of the proposed approach.
Siamak Mehrkanoon, Oscar Mauricio Agudelo, Raghvendra Mall, Johan A. K. Suykens
IJCNN3
2015 Identifying intervals for hierarchical clustering using the Gershgorin circle theorem
Raghvendra Mall, Siamak Mehrkanoon, Johan A. K. Suykens
Pattern Recognit. Lett.1
2015 Very Sparse LSSVM Reductions for Large-Scale Data
abstract
Least squares support vector machines (LSSVMs) have been widely applied for classification and regression with comparable performance with SVMs. The LSSVM model lacks sparsity and is unable to handle large-scale data due to computational and memory constraints. A primal fixed-size LSSVM (PFS-LSSVM) introduce sparsity using Nyström approximation with a set of prototype vectors (PVs). The PFS-LSSVM model solves an overdetermined system of linear equations in the primal. However, this solution is not the sparsest. We investigate the sparsity-error tradeoff by introducing a second level of sparsity. This is done by means of L0 -norm-based reductions by iteratively sparsifying LSSVM and PFS-LSSVM models. The exact choice of the cardinality for the initial PV set is not important then as the final model is highly sparse. The proposed method overcomes the problem of memory constraints and high computational costs resulting in highly sparse reductions to LSSVM models. The approximations of the two models allow to scale the models to large-scale datasets. Experiments on real-world classification and regression data sets from the UCI repository illustrate that these approaches achieve sparse models without a significant tradeoff in errors.
Raghvendra Mall, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.1
2015 Multiclass Semisupervised Learning Based Upon Kernel Spectral Clustering
abstract
This paper proposes a multiclass semisupervised learning algorithm by using kernel spectral clustering (KSC) as a core model. A regularized KSC is formulated to estimate the class memberships of data points in a semisupervised setting using the one-versus-all strategy while both labeled and unlabeled data points are present in the learning process. The propagation of the labels to a large amount of unlabeled data points is achieved by adding the regularization terms to the cost function of the KSC formulation. In other words, imposing the regularization term enforces certain desired memberships. The model is then obtained by solving a linear system in the dual. Furthermore, the optimal embedding dimension is designed for semisupervised clustering. This plays a key role when one deals with a large number of clusters.
Siamak Mehrkanoon, Carlos Alzate, Raghvendra Mall, Rocco Langone, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.3
2014 Representative subsets for big data learning using k-NN graphs
abstract
In this paper we propose a deterministic method to obtain subsets from big data which are a good representative of the inherent structure in the data. We first convert the large scale dataset into a sparse undirected k-NN graph using a distributed network generation framework that we propose in this paper. After obtaining the k-NN graph we exploit the fast and unique representative subset (FURS) selection method [1], [2] to deterministically obtain a subset for this big data network. The FURS selection technique selects nodes from different dense regions in the graph retaining the natural community structure. We then locate the points in the original big data corresponding to the selected nodes and compare the obtained subset with subsets acquired from state-of-the-art subset selection techniques. We evaluate the quality of the selected subset on several synthetic and real-life datasets for different learning tasks including big data classification and big data clustering.
Raghvendra Mall, Vilen Jumutc, Rocco Langone, Johan A. K. Suykens
IEEE BigData1
2014 Clustering data over time using kernel spectral clustering with memory
abstract
This paper discusses the problem of clustering data changing over time, a research domain that is attracting increasing attention due to the increased availability of streaming data in the Web 2.0 era. In the analysis conducted throughout the paper we make use of the kernel spectral clustering with memory (MKSC) algorithm, which is developed in a constrained optimization setting. Since the objective function of the MKSC model is designed to explicitly incorporate temporal smoothness, the algorithm belongs to the family of evolutionary clustering methods. Experiments over a number of real and synthetic datasets provide very interesting insights in the dynamics of the clusters evolution. Specifically, MKSC is able to handle objects leaving and entering over time, and recognize events like continuing, shrinking, growing, splitting, merging, dissolving and forming of clusters. Moreover, we discover how one of the regularization constants of the MKSC model, referred as the smoothness parameter, can be used as a change indicator measure. Finally, some possible visualizations of the cluster dynamics are proposed.
Rocco Langone, Raghvendra Mall, Johan A. K. Suykens
CIDM2
2014 Agglomerative hierarchical kernel spectral data clustering
abstract
In this paper we extend the agglomerative hierarchical kernel spectral clustering (AH-KSC [1]) technique from networks to datasets and images. The kernel spectral clustering (KSC) technique builds a clustering model in a primal-dual optimization framework. The dual solution leads to an eigen-decomposition. The clustering model consists of kernel evaluations, projections onto the eigenvectors and a powerful out-of-sample extension property. We first estimate the optimal model parameters using the balanced angular fitting (BAF) [2] criterion. We then exploit the eigen-projections corresponding to these parameters to automatically identify a set of increasing distance thresholds. These distance thresholds provide the clusters at different levels of hierarchy in the dataset which are merged in an agglomerative fashion as shown in [1], [4]. We showcase the effectiveness of the AH-KSC method on several datasets and real world images. We compare the AH-KSC method with several agglomerative hierarchical clustering techniques and overcome the issues of hierarchical KSC technique proposed in [5].
Raghvendra Mall, Rocco Langone, Johan A. K. Suykens
CIDM1
2014 Agglomerative hierarchical kernel spectral clustering for large scale networks
Raghvendra Mall, Rocco Langone, Johan A. K. Suykens
ESANN1
2014 Optimal reduced sets for sparse kernel spectral clustering
abstract
Kernel spectral clustering (KSC) solves a weighted kernel principal component analysis problem in a primal-dual optimization framework. It results in a clustering model using the dual solution of the problem. It has a powerful out-of-sample extension property leading to good clustering generalization w.r.t. the unseen data points. The out-of-sample extension property allows to build a sparse model on a small training set and introduces the first level of sparsity. The clustering dual model is expressed in terms of non-sparse kernel expansions where every point in the training set contributes. The goal is to find reduced set of training points which can best approximate the original solution. In this paper a second level of sparsity is introduced in order to reduce the time complexity of the computationally expensive out-of-sample extension. In this paper we investigate various penalty based reduced set techniques including the Group Lasso, L0, L1+ L0penalization and compare the amount of sparsity gained w.r.t. a previous L1penalization technique. We observe that the optimal results in terms of sparsity corresponds to the Group Lasso penalization technique in majority of the cases. We showcase the effectiveness of the proposed approaches on several real world datasets and an image segmentation dataset.
Raghvendra Mall, Siamak Mehrkanoon, Rocco Langone, Johan A. K. Suykens
IJCNN1
2013 Self-tuned kernel spectral clustering for large scale networks
abstract
We propose a parameter-free kernel spectral clustering model for large scale complex networks. The kernel spectral clustering (KSC) method works by creating a model on a subgraph of the complex network. The model requires a kernel function which can have parameters and the number of communities k has be detected in the large scale network. We exploit the structure of the projections in the eigenspace to automatically identify the number of clusters. We use the concept of entropy and balanced clusters for this purpose. We show the effectiveness of the proposed approach by comparing the cluster memberships w.r.t. several large scale community detection techniques like Louvain, Infomap and Bigclam methods. We conducted experiments on several synthetic networks of varying size and mixing parameter along with large scale real world experiments to show the efficiency of the proposed approach.
Raghvendra Mall, Rocco Langone, Johan A. K. Suykens
IEEE BigData1
2013 Soft kernel spectral clustering
abstract
In this paper we propose an algorithm for soft (or fuzzy) clustering. In soft clustering each point is not assigned to a single cluster (like in hard clustering), but it can belong to every cluster with a different degree of membership. Generally speaking, this property is desirable in order to improve the interpretability of the results. Our starting point is a state-of-the art technique called kernel spectral clustering (KSC). Instead of using the hard assignment method present therein, we suggest a fuzzy assignment based on the cosine distance from the cluster prototypes. We then call the new method soft kernel spectral clustering (SKSC). We also introduce a related model selection technique, called average membership strength criterion, which solves the drawbacks of the previously proposed method (namely balanced linefit). We apply the new algorithm to synthetic and real datasets, for image segmentation and community detection on networks. We show that in many cases SKSC outperforms KSC.
Rocco Langone, Raghvendra Mall, Johan A. K. Suykens
IJCNN2
2013 Sparse Reductions for Fixed-Size Least Squares Support Vector Machines on Large Scale Data
Raghvendra Mall, Johan A. K. Suykens
PAKDD (1)1
2011 Comparative Behaviour of Recent Incremental and Non-incremental Clustering Methods on Text: An Extended Study
Jean-Charles Lamirel, Raghvendra Mall, Mumtaz Ahmad 0002
IEA/AIE (1)2
2011 Variations to incremental growing neural gas algorithm based on label maximization
abstract
Neural clustering algorithms show high performance in the general context of the analysis of homogeneous textual dataset. This is especially true for the recent adaptive versions of these algorithms, like the incremental growing neural gas algorithm (IGNG) and the labeling maximization based incremental growing neural gas algorithm (IGNG-F). In this paper we highlight that there is a drastic decrease of performance of these algorithms, as well as the one of more classical algorithms, when a heterogeneous textual dataset is considered as an input. Specific quality measures and cluster labeling techniques that are independent of the clustering method are used for the precise performance evaluation. We provide new variations to incremental growing neural gas algorithm exploiting in an incremental way knowledge from clusters about their current labeling along with cluster distance measure data. This solution leads to significant gain in performance for all types of datasets, especially for the clustering of complex heterogeneous textual data.
Jean-Charles Lamirel, Raghvendra Mall, Pascal Cuxac, Ghada Safi
IJCNN2
2010 PERFICT: Perturbed Frequent Itemset Based Classification Technique
abstract
This paper presents Perturbed Frequent Itemset based Classification Technique (PERFICT), a novel associative classification approach based on perturbed frequent itemsets. Most of the existing associative classifiers work well on transactional data where each record contains a set of boolean items. They are not very effective in general for relational data that typically contains real valued attributes. In PERFICT, we handle real attributes by treating items as (attribute, value) pairs, where the value is not the original one, but is perturbed by a small amount and is a range based value. We also propose our own similarity measure which captures the nature of real valued attributes and provide effective weights for the itemsets. The probabilistic contributions of different itemsets is taken into considerations during classification. Some of the applications where such a technique is useful are in signal classification, medical diagnosis and handwriting recognition. Experiments conducted on the UCI Repository datasets show that PERFICT is highly competitive in terms of accuracy in comparison with popular associative classification methods.
Raghvendra Mall, Prakhar Jain, Vikram Pudi
ICTAI (1)1