EDBT 2026 Demo / reviewers in the wild / expert
Alok Sharma
dblp:90/2375
· DBLP profile ↗
50ranked-venue papers
23as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 20 · 9 first-author · 3 since 2021Systems, architecture and hardware · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | scHDeepInsight: a hierarchical deep learning framework for precise immune cell annotation in single-cell RNA-seq dataabstractAccurate classification of immune cells is crucial for elucidating their diverse roles in health and disease. However, this task remains very challenging in single-cell RNA sequencing (scRNA-seq) data due to the complex and hierarchical relationships of immune cell types. To address this, we introduce scHDeepInsight, a deep learning framework that extends our previous scDeepInsight model by integrating a biologically-informed classification architecture with an adaptive hierarchical focal loss (AHFL). The framework builds on our established method of converting gene expression data into two-dimensional structured images, enabling convolutional neural networks to effectively capture both global and fine-grained transcriptomic features. This design utilizes hierarchical relationships among immune cell types to enhance the classification ability beyond the flat classification approaches. scHDeepInsight dynamically adjusts loss contributions to balance performance across the hierarchy levels. Comprehensive benchmarking across seven diverse tissue datasets shows scHDeepInsight achieves an average accuracy of 93.2%, surpassing contemporary methods by 5.1 percentage points. The model successfully distinguishes 50 distinct immune cell subtypes with high accuracy, demonstrating proficiency for identifying rare and closely related cell subtypes. Additionally, SHAP-based interpretability quantifies individual gene contributions to reveal the biological basis of classification decisions. These qualities make scHDeepInsight a robust tool for high-resolution cell subtype characterization, well-suited for detailed profiling in immunological studies and extensible to nonimmune cell types. Shangru Jia, Artem Lysenko, Keith A. Boroevich, Alok Sharma, Tatsuhiko Tsunoda |
Briefings Bioinform. | 4 |
| 2024 | multi-GAT: Integrative Analysis of scRNA-seq and scATAC-seq Data Using Graph Attention Networks for Cell Annotation
Shangru Jia, Tatsuhiko Tsunoda, Alok Sharma |
PRICAI (1) | 3 |
| 2023 | scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learningabstractAnnotation of cell-types is a critical step in the analysis of single-cell RNA sequencing (scRNA-seq) data that allows the study of heterogeneity across multiple cell populations. Currently, this is most commonly done using unsupervised clustering algorithms, which project single-cell expression data into a lower dimensional space and then cluster cells based on their distances from each other. However, as these methods do not use reference datasets, they can only achieve a rough classification of cell-types, and it is difficult to improve the recognition accuracy further. To effectively solve this issue, we propose a novel supervised annotation method, scDeepInsight. The scDeepInsight method is capable of performing manifold assignments. It is competent in executing data integration through batch normalization, performing supervised training on the reference dataset, doing outlier detection and annotating cell-types on query datasets. Moreover, it can help identify active genes or marker genes related to cell-types. The training of the scDeepInsight model is performed in a unique way. Tabular scRNA-seq data are first converted to corresponding images through the DeepInsight methodology. DeepInsight can create a trainable image transformer to convert non-image RNA data to images by comprehensively comparing interrelationships among multiple genes. Subsequently, the converted images are fed into convolutional neural networks such as EfficientNet-b3. This enables automatic feature extraction to identify the cell-types of scRNA-seq samples. We benchmarked scDeepInsight with six other mainstream cell annotation methods. The average accuracy rate of scDeepInsight reached 87.5%, which is more than 7% higher compared with the state-of-the-art methods. Shangru Jia, Artem Lysenko, Keith A. Boroevich, Alok Sharma, Tatsuhiko Tsunoda |
Briefings Bioinform. | 4 |
| 2023 | Memory capacity of recurrent neural networks with matrix representationabstractIt is well known that canonical recurrent neural networks (RNNs) faced limitations in learning long-term dependencies which has been addressed by memory structures in long short-term memory (LSTM) networks. Neural Turing machines (NTMs) are novel RNNs that implement the notion of programmable computers with neural network controllers which can learn simple algorithmic tasks. Matrix neural networks feature matrix representation which inherently preserves the spatial structure of data when compared to canonical neural networks that use vector-based representation. The matrix-representation of neural networks also have the potential to provide better memory capacity. In this paper, we define and study a probabilistic notion of memory capacity based on Fisher information for matrix-based RNNs. We find bounds on memory capacity for such networks under various hypotheses and compare them with their vector counterparts. In particular, we show that the memory capacity of such networks is bounded by N2 for N×N state matrix which generalizes the one known for vector networks. We also show and analyze the increase in memory capacity for such networks which is introduced when one exhibits an external state memory, such as Neural Turing Machines (NTMs). This motivates us to construct NTMs with RNN controllers with matrix-based representation of external memory, leading us to introduce Matrix NTMs. We demonstrate the performance of this class of memory networks under certain algorithmic learning tasks such as copying and recall and compare it with Matrix RNNs. We find an improvement in the performance of Matrix NTMs by the addition of external memory. Animesh Renanse, Alok Sharma, Rohitash Chandra |
Neurocomputing | 2 |
| 2022 | A Regression and Neural Network-Based Methodology to Improve Vertical Resolution of Matched Indian Ground Radar Reflectivity ObservationsabstractIn this work, a direct comparison between Indian ground radar (GR) and precipitation radar (PR) onboard the Tropical Rainfall Measuring Mission (TRMM) is first performed. Further, the matched GR observations are converted to a resolution (5 km X 0.25 km) same as that of space radar (SR) observations using regression methodology. Furthermore, GR observations are corrected against SR observations using an artificial neural network (ANN). Alok Sharma, Srinivasa Ramanujam Kannan |
IGARSS | 1 |
| 2021 | Demo: A WhatsApp Bot for Citizen Journalism in Rural IndiaabstractIncreasing penetration of Internet-enabled smartphones in low-resource areas makes them an attractive platform for engaging emerging users. In this paper, we demonstrate how a voice forum for citizen journalism in rural India– previously accessible via an Interactive Voice Response (IVR) system– can be naturally supported and enriched using a chatbot. Implemented using the WhatsApp Business API, the bot enables submission of both audio (with or without image) and video stories. Following review by moderators, stories are published on a website and social media sites, and can also be browsed interactively using the WhatsApp bot. This multi-way, intermediated model of communication expands the scope and functionality of typical WhatsApp groups while offering significant cost savings relative to IVR systems. In the first 9 weeks of a long-term deployment, the bot demonstrated high usability and acceptance and resulted in 218 published stories from 27 users. Ananya Saxena, Alok Sharma, Bill Thies, Devansh Mehta |
COMPASS | 3 |
| 2021 | Upscaling IMD Ground Radar Vertical Reflectivity Using TRMM PR Observations and Artificial Neural NetworkabstractIn the present work, a one-to-one comparison of the Indian Meteorological Department's (IMD's) Ground Radar (GR) at Delhi against Precipitation Radar (PR) onboard Tropical Rainfall Measuring Mission (TRMM) using alignment methodology is presented. The GR and PR reflectivity data are aligned to a common volume. The GR reflectivity is then upscaled to PR resolution using a weighting methodology. Further, an artificial neural network (ANN) based technique is applied to correct the bias in upscaled GR reflectivity against the PR reflectivity. Alok Sharma, Srinivasa Ramanujam Kannan |
IGARSS | 1 |
| 2021 | DeepFeature: feature selection in nonimage data using convolutional neural networkabstractArtificial intelligence methods offer exciting new capabilities for the discovery of biological mechanisms from raw data because they are able to detect vastly more complex patterns of association that cannot be captured by classical statistical tests. Among these methods, deep neural networks are currently among the most advanced approaches and, in particular, convolutional neural networks (CNNs) have been shown to perform excellently for a variety of difficult tasks. Despite that applications of this type of networks to high-dimensional omics data and, most importantly, meaningful interpretation of the results returned from such models in a biomedical context remains an open problem. Here we present, an approach applying a CNN to nonimage data for feature selection. Our pipeline, DeepFeature, can both successfully transform omics data into a form that is optimal for fitting a CNN model and can also return sets of the most important genes used internally for computing predictions. Within the framework, the Snowfall compression algorithm is introduced to enable more elements in the fixed pixel framework, and region accumulation and element decoder is developed to find elements or genes from the class activation maps. In comparative tests for cancer type prediction task, DeepFeature simultaneously achieved superior predictive performance and better ability to discover key pathways and biological processes meaningful for this context. Capabilities offered by the proposed framework can enable the effective use of powerful deep learning methods to facilitate the discovery of causal mechanisms in high-dimensional biomedical data. Alok Sharma, Artem Lysenko, Keith A. Boroevich, Edwin Vans, Tatsuhiko Tsunoda |
Briefings Bioinform. | 1 |
| 2021 | FEATS: feature selection-based clustering of single-cell RNA-seq dataabstractMOTIVATION: Advances in next-generation sequencing have made it possible to carry out transcriptomic studies at single-cell resolution and generate vast amounts of single-cell RNA sequencing (RNA-seq) data rapidly. Thus, tools to analyze this data need to evolve as well as to improve accuracy and efficiency. RESULTS: We present FEATS, a Python software package, that performs clustering on single-cell RNA-seq data. FEATS is capable of performing multiple tasks such as estimating the number of clusters, conducting outlier detection and integrating data from various experiments. We develop a univariate feature selection-based approach for clustering, which involves the selection of top informative features to improve clustering performance. This is motivated by the fact that cell types are often manually determined using the expression of only a few known marker genes. On a variety of single-cell RNA-seq datasets, FEATS gives superior performance compared with the current tools, in terms of adjusted Rand index and estimating the number of clusters. It achieves a 22% improvement in clustering and more accurately estimates the number of clusters when compared with other tools. In addition to cluster estimation, FEATS also performs outlier detection and data integration while giving an excellent computational performance. Thus, FEATS is a comprehensive clustering tool capable of addressing the challenges during the clustering of single-cell RNA-seq data. AVAILABILITY: The installation instructions and documentation of FEATS is available at https://edwinv87.github.io/feats/. SUPPLEMENTARY DATA: Supplementary data are available online at https://academic.oup.com/bib. Edwin Vans, Ashwini Patil, Alok Sharma |
Briefings Bioinform. | 3 |
| 2021 | Forecasting the spread of COVID-19 using LSTM networkabstractBACKGROUND: The novel coronavirus (COVID-19) is caused by severe acute respiratory syndrome coronavirus 2, and within a few months, it has become a global pandemic. This forced many affected countries to take stringent measures such as complete lockdown, shutting down businesses and trade, as well as travel restrictions, which has had a tremendous economic impact. Therefore, having knowledge and foresight about how a country might be able to contain the spread of COVID-19 will be of paramount importance to the government, policy makers, business partners and entrepreneurs. To help social and administrative decision making, a model that will be able to forecast when a country might be able to contain the spread of COVID-19 is needed. RESULTS: The results obtained using our long short-term memory (LSTM) network-based model are promising as we validate our prediction model using New Zealand's data since they have been able to contain the spread of COVID-19 and bring the daily new cases tally to zero. Our proposed forecasting model was able to correctly predict the dates within which New Zealand was able to contain the spread of COVID-19. Similarly, the proposed model has been used to forecast the dates when other countries would be able to contain the spread of COVID-19. CONCLUSION: The forecasted dates are only a prediction based on the existing situation. However, these forecasted dates can be used to guide actions and make informed decisions that will be practically beneficial in influencing the real future. The current forecasting trend shows that more stringent actions/restrictions need to be implemented for most of the countries as the forecasting model shows they will take over three months before they can possibly contain the spread of COVID-19. Shiu Kumar, Ronesh Sharma, Tatsuhiko Tsunoda, Thirumananseri Kumarevel, Alok Sharma |
BMC Bioinform. | 5 |
| 2021 | SPECTRA: a tool for enhanced brain wave signal recognitionabstractBACKGROUND: Brain wave signal recognition has gained increased attention in neuro-rehabilitation applications. This has driven the development of brain-computer interface (BCI) systems. Brain wave signals are acquired using electroencephalography (EEG) sensors, processed and decoded to identify the category to which the signal belongs. Once the signal category is determined, it can be used to control external devices. However, the success of such a system essentially relies on significant feature extraction and classification algorithms. One of the commonly used feature extraction technique for BCI systems is common spatial pattern (CSP). RESULTS: The performance of the proposed spatial-frequency-temporal feature extraction (SPECTRA) predictor is analysed using three public benchmark datasets. Our proposed predictor outperformed other competing methods achieving lowest average error rates of 8.55%, 17.90% and 20.26%, and highest average kappa coefficient values of 0.829, 0.643 and 0.595 for BCI Competition III dataset IVa, BCI Competition IV dataset I and BCI Competition IV dataset IIb, respectively. CONCLUSIONS: Our proposed SPECTRA predictor effectively finds features that are more separable and shows improvement in brain wave signal recognition that can be instrumental in developing improved real-time BCI systems that are computationally efficient. Shiu Kumar, Tatsuhiko Tsunoda, Alok Sharma |
BMC Bioinform. | 3 |
| 2020 | Using Mobile Airtime Credits to Incentivize Learning, Sharing and Survey Response: Experiences from the FieldabstractIn the Global South, mobile airtime payment has emerged as a popular way to incentivize different research studies, including ones on survey completion or disseminating information to people. Building on this literature, we report deployment experiences from three different studies in India that used airtime incentives. The first was used to promote awareness about HIV/AIDS, the second for promoting awareness and surveying preparedness for an upcoming election, and the third to measure learning and encourage people to vote in a conflict-hit region for a different election. Unlike past work, we found that a delivery mechanism that focuses on asking questions first, rather than presenting a tutorial and then asking questions, worked well in practice. In addition, we found multiple challenges in adoption of the technology and tried different ways to incentivize peer sharing of our system. Between the three deployments, we also addressed other technical and human-centered challenges such as delayed airtime payments and people using the system on behalf of someone else. We hope that our experiences and insights can be helpful to others seeking to deploy applications that utilize mobile airtime payments for learning, sharing, and survey response. Devansh Mehta, Ramaravind Kommiya Mothilal, Alok Sharma, William Thies, Amit Sharma 0007 |
COMPASS | 3 |
| 2020 | Learnings from Technological Interventions in a Low Resource Language: A Case-Study on GondiabstractThe primary obstacle to developing technologies for low-resource languages is the lack of usable data. In this paper, we report the adaption and deployment of 4 technology-driven methods of data collection for Gondi, a low-resource vulnerable language spoken by around 2.3 million tribal people in south and central India. In the process of data collection, we also help in its revival by expanding access to information in Gondi through the creation of linguistic resources that can be used by the community, such as a dictionary, children’s stories, an app with Gondi content from multiple sources and an Interactive Voice Response (IVR) based mass awareness platform. At the end of these interventions, we collected a little less than 12,000 translated words and/or sentences and identified more than 650 community members whose help can be solicited for future translation efforts. The larger goal of the project is collecting enough data in Gondi to build and deploy viable language technologies like machine translation and speech to text systems that can help take the language onto the internet. Devansh Mehta, Sebastin Santy, Ramaravind Kommiya Mothilal, Brij Mohan Lal Srivastava, Alok Sharma, Anurag Shukla, Vishnu Prasad, U. Venkanna 0001, Amit Sharma 0007, Kalika Bali |
LREC | 5 |
| 2019 | Subject-Specific-Frequency-Band for Motor Imagery EEG Signal Recognition Based on Common Spatial Spectral Pattern
Shiu Kumar, Alok Sharma, Tatsuhiko Tsunoda |
PRICAI (2) | 2 |
| 2019 | Computational Prediction of Lysine Pupylation Sites in Prokaryotic Proteins Using Position Specific Scoring Matrix into Bigram for Feature Extraction
Alok Sharma, Abel Avitesh Chandra, Abdollah Dehzangi, Daichi Shigemizu, Tatsuhiko Tsunoda |
PRICAI (3) | 2 |
| 2019 | Clustering of Small-Sample Single-Cell RNA-Seq Data via Feature Clustering and Selection
Edwin Vans, Alok Sharma, Ashwini Patil, Daichi Shigemizu, Tatsuhiko Tsunoda |
PRICAI (3) | 2 |
| 2019 | PyFeat: a Python-based effective feature generation tool for DNA, RNA and protein sequencesabstractMOTIVATION: Extracting useful feature set which contains significant discriminatory information is a critical step in effectively presenting sequence data to predict structural, functional, interaction and expression of proteins, DNAs and RNAs. Also, being able to filter features with significant information and avoid sparsity in the extracted features require the employment of efficient feature selection techniques. Here we present PyFeat as a practical and easy to use toolkit implemented in Python for extracting various features from proteins, DNAs and RNAs. To build PyFeat we mainly focused on extracting features that capture information about the interaction of neighboring residues to be able to provide more local information. We then employ AdaBoost technique to select features with maximum discriminatory information. In this way, we can significantly reduce the number of extracted features and enable PyFeat to represent the combination of effective features from large neighboring residues. As a result, PyFeat is able to extract features from 13 different techniques and represent context free combination of effective features. The source code for PyFeat standalone toolkit and employed benchmarks with a comprehensive user manual explaining its system and workflow in a step by step manner are publicly available. RESULTS: https://github.com/mrzResearchArena/PyFeat/blob/master/RESULTS.md. AVAILABILITY AND IMPLEMENTATION: Toolkit, source code and manual to use PyFeat: https://github.com/mrzResearchArena/PyFeat/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rafsanjani Muhammod, Sajid Ahmed, Dewan Md. Farid 0001, Swakkhar Shatabda, Alok Sharma, Abdollah Dehzangi |
Bioinform. | 5 |
| 2019 | GlyStruct: glycation prediction using structural properties of amino acid residuesabstractBACKGROUND: Glycation is a one of the post-translational modifications (PTM) where sugar molecules and residues in protein sequences are covalently bonded. It has become one of the clinically important PTM in recent times attributed to many chronic and age related complications. Being a non-enzymatic reaction, it is a great challenge when it comes to its prediction due to the lack of significant bias in the sequence motifs. RESULTS: We developed a classifier, GlyStruct based on support vector machine, to predict glycated and non-glycated lysine residues using structural properties of amino acid residues. The features used were secondary structure, accessible surface area and the local backbone torsion angles. For this work, a benchmark dataset was extracted containing 235 glycated and 303 non-glycated lysine residues. GlyStruct demonstrated improved performance of approximately 10% in comparison to benchmark method of Gly-PseAAC. The performance for GlyStruct on the metrics, sensitivity, specificity, accuracy and Mathew's correlation coefficient were 0.7013, 0.7989, 0.7562, and 0.5065, respectively for 10-fold cross-validation. CONCLUSION: Glycation has emerged to be one of the clinically important PTM of proteins in recent times. Therefore, the development of computational tools become necessary to predict glycation, which could help medical professionals administer drugs and manage patients more effectively. The proposed predictor manages to classify glycated and non-glycated lysine residues with promising results consistently on various cross-validation schemes and outperforms other state of the art methods. Hamendra Manhar Reddy, Alok Sharma, Abdollah Dehzangi, Daichi Shigemizu, Abel Avitesh Chandra, Tatsuhiko Tsunoda |
BMC Bioinform. | 2 |
| 2019 | Discovering MoRFs by trisecting intrinsically disordered protein sequence into terminals and middle regionsabstractBACKGROUND: Molecular Recognition Features (MoRFs) are short protein regions present in intrinsically disordered protein (IDPs) sequences. MoRFs interact with structured partner protein and upon interaction, they undergo a disorder-to-order transition to perform various biological functions. Analyses of MoRFs are important towards understanding their function. RESULTS: Performance is reported using the MoRF dataset that has been previously used to compare the other existing MoRF predictors. The performance obtained in this study is equivalent to the benchmarked OPAL predictor, i.e., OPAL achieved AUC of 0.815, whereas the model in this study achieved AUC of 0.819 using TEST set. CONCLUSION: Achieving comparable performance, the proposed method can be used as an alternative approach for MoRF prediction. Ronesh Sharma, Alok Sharma, Ashwini Patil, Tatsuhiko Tsunoda |
BMC Bioinform. | 2 |
| 2018 | OPAL: prediction of MoRF regions in intrinsically disordered protein sequencesabstractMotivation: Intrinsically disordered proteins lack stable 3-dimensional structure and play a crucial role in performing various biological functions. Key to their biological function are the molecular recognition features (MoRFs) located within long disordered regions. Computationally identifying these MoRFs from disordered protein sequences is a challenging task. In this study, we present a new MoRF predictor, OPAL, to identify MoRFs in disordered protein sequences. OPAL utilizes two independent sources of information computed using different component predictors. The scores are processed and combined using common averaging method. The first score is computed using a component MoRF predictor which utilizes composition and sequence similarity of MoRF and non-MoRF regions to detect MoRFs. The second score is calculated using half-sphere exposure (HSE), solvent accessible surface area (ASA) and backbone angle information of the disordered protein sequence, using information from the amino acid properties of flanks surrounding the MoRFs to distinguish MoRF and non-MoRF residues. Results: OPAL is evaluated using test sets that were previously used to evaluate MoRF predictors, MoRFpred, MoRFchibi and MoRFchibi-web. The results demonstrate that OPAL outperforms all the available MoRF predictors and is the most accurate predictor available for MoRF prediction. It is available at http://www.alok-ai-lab.com/tools/opal/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Ronesh Sharma, Gaurav Raicar, Tatsuhiko Tsunoda, Ashwini Patil, Alok Sharma |
Bioinform. | 5 |
| 2017 | An improved discriminative filter bank selection approach for motor imagery EEG signal classification using mutual informationabstractBACKGROUND: Common spatial pattern (CSP) has been an effective technique for feature extraction in electroencephalography (EEG) based brain computer interfaces (BCIs). However, motor imagery EEG signal feature extraction using CSP generally depends on the selection of the frequency bands to a great extent. METHODS: In this study, we propose a mutual information based frequency band selection approach. The idea of the proposed method is to utilize the information from all the available channels for effectively selecting the most discriminative filter banks. CSP features are extracted from multiple overlapping sub-bands. An additional sub-band has been introduced that cover the wide frequency band (7-30 Hz) and two different types of features are extracted using CSP and common spatio-spectral pattern techniques, respectively. Mutual information is then computed from the extracted features of each of these bands and the top filter banks are selected for further processing. Linear discriminant analysis is applied to the features extracted from each of the filter banks. The scores are fused together, and classification is done using support vector machine. RESULTS: The proposed method is evaluated using BCI Competition III dataset IVa, BCI Competition IV dataset I and BCI Competition IV dataset IIb, and it outperformed all other competing methods achieving the lowest misclassification rate and the highest kappa coefficient on all three datasets. CONCLUSIONS: Introducing a wide sub-band and using mutual information for selecting the most discriminative sub-bands, the proposed method shows improvement in motor imagery EEG signal classification. Shiu Kumar, Alok Sharma, Tatsuhiko Tsunoda |
BMC Bioinform. | 2 |
| 2017 | 2D-EM clustering approach for high-dimensional data through folding feature vectorsabstractBACKGROUND: Clustering methods are becoming widely utilized in biomedical research where the volume and complexity of data is rapidly increasing. Unsupervised clustering of patient information can reveal distinct phenotype groups with different underlying mechanism, risk prognosis and treatment response. However, biological datasets are usually characterized by a combination of low sample number and very high dimensionality, something that is not adequately addressed by current algorithms. While the performance of the methods is satisfactory for low dimensional data, increasing number of features results in either deterioration of accuracy or inability to cluster. To tackle these challenges, new methodologies designed specifically for such data are needed. RESULTS: We present 2D-EM, a clustering algorithm approach designed for small sample size and high-dimensional datasets. To employ information corresponding to data distribution and facilitate visualization, the sample is folded into its two-dimension (2D) matrix form (or feature matrix). The maximum likelihood estimate is then estimated using a modified expectation-maximization (EM) algorithm. The 2D-EM methodology was benchmarked against several existing clustering methods using 6 medically-relevant transcriptome datasets. The percentage improvement of Rand score and adjusted Rand index compared to the best performing alternative method is up to 21.9% and 155.6%, respectively. To present the general utility of the 2D-EM method we also employed 2 methylome datasets, again showing superior performance relative to established methods. CONCLUSIONS: The 2D-EM algorithm was able to reproduce the groups in transcriptome and methylome data with high accuracy. This build confidence in the methods ability to uncover novel disease subtypes in new datasets. The design of 2D-EM algorithm enables it to handle a diverse set of challenging biomedical dataset and cluster with higher accuracy than established methods. MATLAB implementation of the tool can be freely accessed online ( http://www.riken.jp/en/research/labs/ims/med_sci_math or http://www.alok-ai-lab.com /). Alok Sharma, Piotr J. Kamola, Tatsuhiko Tsunoda |
BMC Bioinform. | 1 |
| 2017 | Divisive hierarchical maximum likelihood clusteringabstractBACKGROUND: Biological data comprises various topologies or a mixture of forms, which makes its analysis extremely complicated. With this data increasing in a daily basis, the design and development of efficient and accurate statistical methods has become absolutely necessary. Specific analyses, such as those related to genome-wide association studies and multi-omics information, are often aimed at clustering sub-conditions of cancers and other diseases. Hierarchical clustering methods, which can be categorized into agglomerative and divisive, have been widely used in such situations. However, unlike agglomerative methods divisive clustering approaches have consistently proved to be computationally expensive. RESULTS: The proposed clustering algorithm (DRAGON) was verified on mutation and microarray data, and was gauged against standard clustering methods in the literature. Its validation included synthetic and significant biological data. When validated on mixed-lineage leukemia data, DRAGON achieved the highest clustering accuracy with data of four different dimensions. Consequently, DRAGON outperformed previous methods with 3-,4- and 5-dimensional acute leukemia data. When tested on mutation data, DRAGON achieved the best performance with 2-dimensional information. CONCLUSIONS: This work proposes a computationally efficient divisive hierarchical clustering method, which can compete equally with agglomerative approaches. The proposed method turned out to correctly cluster data with distinct topologies. A MATLAB implementation can be extraced from http://www.riken.jp/en/research/labs/ims/med_sci_math/ or http://www.alok-ai-lab.com. Alok Sharma, Yosvany López, Tatsuhiko Tsunoda |
BMC Bioinform. | 1 |
| 2016 | Decimation filter with Common Spatial Pattern and Fishers Discriminant Analysis for motor imagery classificationabstractBrain Computer Interface (BCI) system converts thoughts into commands for driving external device with Electroencephalography (EEG). This paper presents the use of decimation filters for filtering the EEG signal. Common Spatial Pattern (CSP) technique is used to transform the filtered signal to a new time series in order to have optimal variance for the discrimination of different tasks. Fishers Discriminant Analysis (FDA) is applied to the CSP features and the FDA scores are fed to a Support Vector Machine (SVM) classifier. The method is evaluated on BCI Competition III Dataset IVa and compared with other related state-of-the-art approaches. The results show that our method outperforms all other approaches in terms of average classification error rate. Compared to best performing method that uses only CSP features, the results obtained in this research offer on average a reduction of 1.07% in the classification error rate. Shiu Kumar, Ronesh Sharma, Alok Sharma, Tatsuhiko Tsunoda |
IJCNN | 3 |
| 2016 | Highly accurate sequence-based prediction of half-sphere exposures of amino acid residues in proteinsabstractMOTIVATION: Solvent exposure of amino acid residues of proteins plays an important role in understanding and predicting protein structure, function and interactions. Solvent exposure can be characterized by several measures including solvent accessible surface area (ASA), residue depth (RD) and contact numbers (CN). More recently, an orientation-dependent contact number called half-sphere exposure (HSE) was introduced by separating the contacts within upper and down half spheres defined according to the Cα-Cβ (HSEβ) vector or neighboring Cα-Cα vectors (HSEα). HSEα calculated from protein structures was found to better describe the solvent exposure over ASA, CN and RD in many applications. Thus, a sequence-based prediction is desirable, as most proteins do not have experimentally determined structures. To our best knowledge, there is no method to predict HSEα and only one method to predict HSEβ. RESULTS: This study developed a novel method for predicting both HSEα and HSEβ (SPIDER-HSE) that achieved a consistent performance for 10-fold cross validation and two independent tests. The correlation coefficients between predicted and measured HSEβ (0.73 for upper sphere, 0.69 for down sphere and 0.76 for contact numbers) for the independent test set of 1199 proteins are significantly higher than existing methods. Moreover, predicted HSEα has a higher correlation coefficient (0.46) to the stability change by residue mutants than predicted HSEβ (0.37) and ASA (0.43). The results, together with its easy Cα-atom-based calculation, highlight the potential usefulness of predicted HSEα for protein structure prediction and refinement as well as function prediction. AVAILABILITY AND IMPLEMENTATION: The method is available at http://sparks-lab.org CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rhys Heffernan, Abdollah Dehzangi, James G. Lyons, Kuldip K. Paliwal, Alok Sharma, Jihua Wang, Abdul Sattar 0001, Yaoqi Zhou, Yuedong Yang |
Bioinform. | 5 |
| 2016 | Predicting MoRFs in protein sequences using HMM profilesabstractBACKGROUND: Intrinsically Disordered Proteins (IDPs) lack an ordered three-dimensional structure and are enriched in various biological processes. The Molecular Recognition Features (MoRFs) are functional regions within IDPs that undergo a disorder-to-order transition on binding to a partner protein. Identifying MoRFs in IDPs using computational methods is a challenging task. METHODS: In this study, we introduce hidden Markov model (HMM) profiles to accurately identify the location of MoRFs in disordered protein sequences. Using windowing technique, HMM profiles are utilised to extract features from protein sequences and support vector machines (SVM) are used to calculate a propensity score for each residue. Two different SVM kernels with high noise tolerance are evaluated with a varying window size and the scores of the SVM models are combined to generate the final propensity score to predict MoRF residues. The SVM models are designed to extract maximal information between MoRF residues, its neighboring regions (Flanks) and the remainder of the sequence (Others). RESULTS: To evaluate the proposed method, its performance was compared to that of other MoRF predictors; MoRFpred and ANCHOR. The results show that the proposed method outperforms these two predictors. CONCLUSIONS: Using HMM profile as a source of feature extraction, the proposed method indicates improvement in predicting MoRFs in disordered protein sequences. Ronesh Sharma, Shiu Kumar, Tatsuhiko Tsunoda, Ashwini Patil, Alok Sharma |
BMC Bioinform. | 5 |
| 2016 | Stepwise iterative maximum likelihood clustering approachabstractBACKGROUND: Biological/genetic data is a complex mix of various forms or topologies which makes it quite difficult to analyze. An abundance of such data in this modern era requires the development of sophisticated statistical methods to analyze it in a reasonable amount of time. In many biological/genetic analyses, such as genome-wide association study (GWAS) analysis or multi-omics data analysis, it is required to cluster the plethora of data into sub-categories to understand the subtypes of populations, cancers or any other diseases. Traditionally, the k-means clustering algorithm is a dominant clustering method. This is due to its simplicity and reasonable level of accuracy. Many other clustering methods, including support vector clustering, have been developed in the past, but do not perform well with the biological data, either due to computational reasons or failure to identify clusters. RESULTS: The proposed SIML clustering algorithm has been tested on microarray datasets and SNP datasets. It has been compared with a number of clustering algorithms. On MLL datasets, SIML achieved highest clustering accuracy and rand score on 4/9 cases; similarly on SRBCT dataset, it got for 3/5 cases; on ALL subtype it got highest clustering accuracy for 5/7 cases and highest rand score for 4/7 cases. In addition, SIML overall clustering accuracy on a 3 cluster problem using SNP data were 97.3, 94.7 and 100 %, respectively, for each of the clusters. CONCLUSIONS: In this paper, considering the nature of biological data, we proposed a maximum likelihood clustering approach using a stepwise iterative procedure. The advantage of this proposed method is that it not only uses the distance information, but also incorporate variance information for clustering. This method is able to cluster when data appeared in overlapping and complex forms. The experimental results illustrate its performance and usefulness over other clustering methods. A Matlab package of this method (SIML) is provided at the web-link http://www.riken.jp/en/research/labs/ims/med_sci_math/ . Alok Sharma, Daichi Shigemizu, Keith A. Boroevich, Yosvany López, Yoichiro Kamatani, Michiaki Kubo, Tatsuhiko Tsunoda |
BMC Bioinform. | 1 |
| 2015 | Gram-positive and gram-negative subcellular localization using rotation forest and physicochemical-based featuresabstractBACKGROUND: The functioning of a protein relies on its location in the cell. Therefore, predicting protein subcellular localization is an important step towards protein function prediction. Recent studies have shown that relying on Gene Ontology (GO) for feature extraction can improve the prediction performance. However, for newly sequenced proteins, the GO is not available. Therefore, for these cases, the prediction performance of GO based methods degrade significantly. RESULTS: In this study, we develop a method to effectively employ physicochemical and evolutionary-based information in the protein sequence. To do this, we propose segmentation based feature extraction method to explore potential discriminatory information based on physicochemical properties of the amino acids to tackle Gram-positive and Gram-negative subcellular localization. We explore our proposed feature extraction techniques using 10 attributes that have been experimentally selected among a wide range of physicochemical attributes. Finally by applying the Rotation Forest classification technique to our extracted features, we enhance Gram-positive and Gram-negative subcellular localization accuracies up to 3.4% better than previous studies which used GO for feature extraction. CONCLUSION: By proposing segmentation based feature extraction method to explore potential discriminatory information based on physicochemical properties of the amino acids as well as using Rotation Forest classification technique, we are able to enhance the Gram-positive and Gram-negative subcellular localization prediction accuracies, significantly. Abdollah Dehzangi, Sohrab Sohrabi, Rhys Heffernan, Alok Sharma, James G. Lyons, Kuldip K. Paliwal, Abdul Sattar 0001 |
BMC Bioinform. | 4 |
| 2015 | A deterministic approach to regularized linear discriminant analysis
Alok Sharma, Kuldip K. Paliwal |
Neurocomputing | 1 |
| 2014 | Improving protein fold recognition using the amalgamation of evolutionary-based and structural based informationabstractDeciphering three dimensional structure of a protein sequence is a challenging task in biological science. Protein fold recognition and protein secondary structure prediction are transitional steps in identifying the three dimensional structure of a protein. For protein fold recognition, evolutionary-based information of amino acid sequences from the position specific scoring matrix (PSSM) has been recently applied with improved results. On the other hand, the SPINE-X predictor has been developed and applied for protein secondary structure prediction. Several reported methods for protein fold recognition have only limited accuracy. In this paper, we have developed a strategy of combining evolutionary-based information (from PSSM) and predicted secondary structure using SPINE-X to improve protein fold recognition. The strategy is based on finding the probabilities of amino acid pairs (AAP). The proposed method has been tested on several protein benchmark datasets and an improvement of 8.9% recognition accuracy has been achieved. We have achieved, for the first time over 90% and 75% prediction accuracies for sequence similarity values below 40% and 25%, respectively. We also obtain 90.6% and 77.0% prediction accuracies, respectively, for the Extended Ding and Dubchak and Taguchi and Gromiha benchmark protein fold recognition datasets widely used for in the literature. Kuldip K. Paliwal, Alok Sharma, James G. Lyons, Abdollah Dehzangi |
BMC Bioinform. | 2 |
| 2014 | A feature selection method using improved regularized linear discriminant analysis
Alok Sharma, Kuldip K. Paliwal, Seiya Imoto, Satoru Miyano |
Mach. Vis. Appl. | 1 |
| 2014 | A Segmentation-Based Method to Extract Structural and Evolutionary Features for Protein Fold RecognitionabstractProtein fold recognition (PFR) is considered as an important step towards the protein structure prediction problem. Despite all the efforts that have been made so far, finding an accurate and fast computational approach to solve the PFR still remains a challenging problem for bioinformatics and computational biology. In this study, we propose the concept of segmented-based feature extraction technique to provide local evolutionary information embedded in position specific scoring matrix (PSSM) and structural information embedded in the predicted secondary structure of proteins using SPINE-X. We also employ the concept of occurrence feature to extract global discriminatory information from PSSM and SPINE-X. By applying a support vector machine (SVM) to our extracted features, we enhance the protein fold prediction accuracy for 7.4 percent over the best results reported in the literature. We also report 73.8 percent prediction accuracy for a data set consisting of proteins with less than 25 percent sequence similarity rates and 80.7 percent prediction accuracy for a data set with proteins belonging to 110 folds with less than 40 percent sequence similarity rates. We also investigate the relation between the number of folds and the number of features being used and show that the number of features should be increased to get better protein fold prediction results when the number of folds is relatively large. Abdollah Dehzangi, Kuldip K. Paliwal, James G. Lyons, Alok Sharma, Abdul Sattar 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2013 | A strategy to select suitable physicochemical attributes of amino acids for protein fold recognitionabstractBACKGROUND: Assigning a protein into one of its folds is a transitional step for discovering three dimensional protein structure, which is a challenging task in bimolecular (biological) science. The present research focuses on: 1) the development of classifiers, and 2) the development of feature extraction techniques based on syntactic and/or physicochemical properties. RESULTS: Apart from the above two main categories of research, we have shown that the selection of physicochemical attributes of the amino acids is an important step in protein fold recognition and has not been explored adequately. We have presented a multi-dimensional successive feature selection (MD-SFS) approach to systematically select attributes. The proposed method is applied on protein sequence data and an improvement of around 24% in fold recognition has been noted when selecting attributes appropriately. CONCLUSION: The MD-SFS has been applied successfully in selecting physicochemical attributes of the amino acids. The selected attributes show improved protein fold recognition performance. Alok Sharma, Kuldip K. Paliwal, Abdollah Dehzangi, James G. Lyons, Seiya Imoto, Satoru Miyano |
BMC Bioinform. | 1 |
| 2012 | Improved Pseudoinverse Linear Discriminant Analysis Method for Dimensionality ReductionabstractPseudoinverse linear discriminant analysis (PLDA) is a classical method for solving small sample size problem. However, its performance is limited. In this paper, we propose an improved PLDA method which is faster and produces better classification accuracy when experimented on several datasets. Kuldip K. Paliwal, Alok Sharma |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2012 | A new perspective to null linear discriminant analysis method and its fast implementation using random matrix multiplication with scatter matrices
Alok Sharma, Kuldip K. Paliwal |
Pattern Recognit. | 1 |
| 2012 | A two-stage linear discriminant analysis for face-recognition
Alok Sharma, Kuldip K. Paliwal |
Pattern Recognit. Lett. | 1 |
| 2012 | A Top-r Feature Selection Algorithm for Microarray Gene Expression DataabstractMost of the conventional feature selection algorithms have a drawback whereby a weakly ranked gene that could perform well in terms of classification accuracy with an appropriate subset of genes will be left out of the selection. Considering this shortcoming, we propose a feature selection algorithm in gene expression data analysis of sample classifications. The proposed algorithm first divides genes into subsets, the sizes of which are relatively small (roughly of size h), then selects informative smaller subsets of genes (of size r < h) from a subset and merges the chosen genes with another gene subset (of size r) to update the gene subset. We repeat this process until all subsets are merged into one informative subset. We illustrate the effectiveness of the proposed algorithm by analyzing three distinct gene expression data sets. Our method shows promising classification accuracy for all the test data sets. We also show the relevance of the selected genes in terms of their biological functions. Alok Sharma, Seiya Imoto, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2010 | Improved direct LDA and its application to DNA microarray gene expression data
Kuldip K. Paliwal, Alok Sharma |
Pattern Recognit. Lett. | 2 |
| 2008 | Cancer classification by gradient LDA technique using microarray gene expression data
Alok Sharma, Kuldip K. Paliwal |
Data Knowl. Eng. | 1 |
| 2008 | A Gradient Linear Discriminant Analysis for Small Sample Sized Problem
Alok Sharma, Kuldip K. Paliwal |
Neural Process. Lett. | 1 |
| 2008 | Rotational Linear Discriminant Analysis Technique for Dimensionality ReductionabstractThe linear discriminant analysis (LDA) technique is very popular in pattern recognition for dimensionality reduction. It is a supervised learning technique that finds a linear transformation such that the overlap between the classes is minimum for the projected feature vectors in the reduced feature space. This overlap, if present, adversely affects the classification performance. In this paper, we introduce prior to dimensionality-reduction transformation an additional rotational transform that rotates the feature vectors in the original feature space around their respective class centroids in such a way that the overlap between the classes in the reduced feature space is further minimized. As a result, the classification performance significantly improves, which is demonstrated using several data corpuses. Alok Sharma, Kuldip K. Paliwal |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2007 | Intrusion detection using text processing techniques with a kernel based similarity measure
Alok Sharma, Arun K. Pujari, Kuldip K. Paliwal |
Comput. Secur. | 1 |
| 2007 | Fast principal component analysis using fixed-point algorithm
Alok Sharma, Kuldip K. Paliwal |
Pattern Recognit. Lett. | 1 |
| 2006 | Subspace independent component analysis using vector kurtosis
Alok Sharma, Kuldip K. Paliwal |
Pattern Recognit. | 1 |
| 2006 | Class-dependent PCA, MDC and LDA: A combined classifier for pattern classification
Alok Sharma, Kuldip K. Paliwal, Godfrey C. Onwubolu |
Pattern Recognit. | 1 |
| 1994 | Register Estimation from Behavioral SpecificationsabstractProvides answers to the following problems: (1) Given a data flow graph and a performance constraint, determine a lower-bound on the storage area required for executing the data flow graph while satisfying the performance constraint. (2) Determine a lower-bound on performance for executing a data flow graph under fixed storage area constraints. The results demonstrate that our approach produces solutions which are very close to the optimal.> Alok Sharma, Rajiv Jain |
ICCD | 1 |
| 1993 | InSyn: Integrated Scheduling for DSP ApplicationsabstractIn this paper, we present the l~SV~, an integrated allocation and schedulingapproachfor high-levelsynthesisapplications.The scheduler considers functional units, busses and registers while performing time-step assignment.The results show that incorporating all these features during scheduling can produce very good designs. Alok Sharma, Rajiv Jain |
DAC | 1 |
| 1993 | Estimating Architectural Resources and Performance for High-Level Synthesis ApplicationsabstractIn this paper we present a solution to the following problems related to architectural synthesis. Given an input specification and a perfomance constraint, determine a lower bound number of resources (active and interconnect) required to execute the data flow graph while satisfying the performance constraint. Conversely, determine a lower bound performance for executing an input specification for a given number of resources (active and interconnect). The generated bounds are close to the actual designs synthesized by several existing systems. Alok Sharma, Rajiv Jain |
DAC | 1 |
| 1993 | Estimating architectural resources and performance for high-level synthesis applicationsabstractThe authors present a solution to the following problems related to architectural synthesis. (1) Given an input specification and a performance constraint, determine a lower bound number of resources (active and interconnect) required to execute the data flow graph while satisfying the performance constraint. (2) Determine a lower bound performance for executing an input specification for a given number of resources (active and interconnect). These bounds are close to the actual designs synthesized by several existing systems.> Alok Sharma, Rajiv Jain |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1991 | Empirical Evaluation of Some High-Level Synthesis Scheduling HeuristicsabstractOver the past few years a large number of heuristics for performing scheduling Rajiv Jain, Ashutosh Mujumdar, Alok Sharma, Hueymin Wang |
DAC | 3 |