VLDB 2026 Research / reviewers in the wild / expert
Halima Bensmail
dblp:60/2303
· DBLP profile ↗
31ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-6700-5752ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 1 first-authorDatabases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient and interpretable DNA/RNA representation using Komlós-Hadamard transformsabstractThis study introduces a novel encoding scheme for DNA/RNA sequences, integrating Komlós and Hadamard transforms. Unlike traditional One-Hot encoding, this approach offers a more informative representation of omics data while significantly reducing computational complexity. However, it is important to note that the Komlós transform component provides fewer features and does not utilize sparse codes. By leveraging the inherent properties of these transforms, our method effectively captures complex patterns within the data, leading to improved model accuracy and reduced training times. When combined with an image transformation, this encoding scheme demonstrates particularly efficient results, achieving superior performance across various predictive tasks with significantly lower computational resource demands compared to One-Hot encoding. Our findings suggest that this novel encoding scheme, particularly when integrated with Hilbert Curve mapping or sequence to image analysis, holds significant promise for advancing DNA/RNA data analysis by offering a more efficient and effective approach to feature representation. Kareem Kabbani, Samir Brahim Belhaouari, Michaël Aupetit 0001, Aisha Al-Qahtani, Ahmad Halabi, Sophia L. Haoudi, Halima Bensmail |
BMC Bioinform. | 7 |
| 2025 | Tisslet tissues-based learning estimation for transcriptomicsabstractIn the context of multi-omics data analytics for various diseases, transcriptome-wide association studies leveraging genetically predicted gene expression hold promise for identifying novel regions linked to complex traits. However, existing methods for multi-tissue gene expression prediction often fail to account for tissue-tissue expression interactions, limiting their accuracy and effectiveness. This research addresses the challenge of predicting gene expression across multiple tissues by incorporating tissue-tissue expression correlations based on a nonlinear multivariate model. Our findings demonstrate that this model excels in estimating tissue-tissue interactions and accurately predicting missing data. These results have significant implications for multi-omics data analytics and transcriptome-wide association studies, suggesting a novel approach for identifying regions associated with complex traits. Ahmed Miloudi, Aisha Al-Qahtani, Thamanna Hashir, Mohamed Chikri, Halima Bensmail |
BMC Bioinform. | 5 |
| 2023 | Progressive Fourier Transform (PFT): Enhancing Time-Frequency Representation of EEG signals for Stress and Seizure DetectionabstractThe detection and classification of neurological and psychological phenomena heavily rely on Electroencephalography (EEG). This study investigates the effectiveness of various feature extraction techniques and machine learning classifiers in EEG-based classification tasks. Stress detection using the Bird et al. dataset, which encompasses multiple emotional states, and seizure detection using the CHB-MIT dataset, known for its challenges in distinguishing seizure from non-seizure patterns, are specifically explored.The results highlight the crucial role of feature extraction methods in EEG-based classification. Among the techniques tested, our Progressive Fourier Transform (PFT) method consistently outperforms others, emerging as the superior choice.In stress detection, our proposed PFT achieves an outstanding accuracy of 98.41% on the Bird et al. dataset, surpassing existing methods based on statistical features. For seizure detection, our model attains a competitive accuracy of 96.88% on the CHBMIT dataset, showcasing efficiency even with a reduced number of channels.This study demonstrates the potential of EEG-based classification techniques in practical applications such as stress monitoring and seizure prediction. Furthermore, it emphasizes the significance of advanced feature extraction methods in achieving accurate results. Future research may involve refining these techniques further and expanding their applicability to diverse EEG datasets and other neurological and psychological disorders. Nisreen Said Amer, Samir Brahim Belhaouari, Halima Bensmail |
BIBM | 3 |
| 2023 | OutSingle: a novel method of detecting and injecting outliers in RNA-Seq count data using the optimal hard threshold for singular valuesabstractMOTIVATION: Finding outliers in RNA-sequencing (RNA-Seq) gene expression (GE) can help in identifying genes that are aberrant and cause Mendelian disorders. Recently developed models for this task rely on modeling RNA-Seq GE data using the negative binomial distribution (NBD). However, some of those models either rely on procedures for inferring NBD's parameters in a nonbiased way that are computationally demanding and thus make confounder control challenging, while others rely on less computationally demanding but biased procedures and convoluted confounder control approaches that hinder interpretability. RESULTS: In this article, we present OutSingle (Outlier detection using Singular Value Decomposition), an almost instantaneous way of detecting outliers in RNA-Seq GE data. It uses a simple log-normal approach for count modeling. For confounder control, it uses the recently discovered optimal hard threshold (OHT) method for noise detection, which itself is based on singular value decomposition (SVD). Due to its SVD/OHT utilization, OutSingle's model is straightforward to understand and interpret. We then show that our novel method, when used on RNA-Seq GE data with real biological outliers masked by confounders, outcompetes the previous state-of-the-art model based on an ad hoc denoising autoencoder. Additionally, OutSingle can be used to inject artificial outliers masked by confounders, which is difficult to achieve with previous approaches. We describe a way of using OutSingle for outlier injection and proceed to show how OutSingle outperforms its competition on 16 out of 18 datasets that were generated from three real datasets using OutSingle's injection procedure with different outlier types and magnitudes. Our methods are applicable to other types of similar problems involving finding outliers in matrices under the presence of confounders. AVAILABILITY AND IMPLEMENTATION: The code for OutSingle is available at https://github.com/esalkovic/outsingle. Edin Salkovic, Mohammad Amin Sadeghi, Abdelkader Baggag, Ahmed Gamal Rashed Salem, Halima Bensmail |
Bioinform. | 5 |
| 2021 | Computational prediction and interpretation of both general and specific types of promoters in Escherichia coli by exploiting a stacked ensemble-learning frameworkabstractPromoters are short consensus sequences of DNA, which are responsible for transcription activation or the repression of all genes. There are many types of promoters in bacteria with important roles in initiating gene transcription. Therefore, solving promoter-identification problems has important implications for improving the understanding of their functions. To this end, computational methods targeting promoter classification have been established; however, their performance remains unsatisfactory. In this study, we present a novel stacked-ensemble approach (termed SELECTOR) for identifying both promoters and their respective classification. SELECTOR combined the composition of k-spaced nucleic acid pairs, parallel correlation pseudo-dinucleotide composition, position-specific trinucleotide propensity based on single-strand, and DNA strand features and using five popular tree-based ensemble learning algorithms to build a stacked model. Both 5-fold cross-validation tests using benchmark datasets and independent tests using the newly collected independent test dataset showed that SELECTOR outperformed state-of-the-art methods in both general and specific types of promoter prediction in Escherichia coli. Furthermore, this novel framework provides essential interpretations that aid understanding of model success by leveraging the powerful Shapley Additive exPlanation algorithm, thereby highlighting the most important features relevant for predicting both general and specific types of promoters and overcoming the limitations of existing 'Black-box' approaches that are unable to reveal causal relationships from large amounts of initially encoded features. Fuyi Li, ZongYuan Ge, Yanwei Yue, Morihiro Hayashida, Abdelkader Baggag, Halima Bensmail, Jiangning Song |
Briefings Bioinform. | 8 |
| 2020 | Comprehensive review and assessment of computational methods for predicting RNA post-transcriptional modification sites from RNA sequencesabstractRNA post-transcriptional modifications play a crucial role in a myriad of biological processes and cellular functions. To date, more than 160 RNA modifications have been discovered; therefore, accurate identification of RNA-modification sites is fundamental for a better understanding of RNA-mediated biological functions and mechanisms. However, due to limitations in experimental methods, systematic identification of different types of RNA-modification sites remains a major challenge. Recently, more than 20 computational methods have been developed to identify RNA-modification sites in tandem with high-throughput experimental methods, with most of these capable of predicting only single types of RNA-modification sites. These methods show high diversity in their dataset size, data quality, core algorithms, features extracted and feature selection techniques and evaluation strategies. Therefore, there is an urgent need to revisit these methods and summarize their methodologies, in order to improve and further develop computational techniques to identify and characterize RNA-modification sites from the large amounts of sequence data. With this goal in mind, first, we provide a comprehensive survey on a large collection of 27 state-of-the-art approaches for predicting N1-methyladenosine and N6-methyladenosine sites. We cover a variety of important aspects that are crucial for the development of successful predictors, including the dataset quality, operating algorithms, sequence and genomic features, feature selection, model performance evaluation and software utility. In addition, we also provide our thoughts on potential strategies to improve the model performance. Second, we propose a computational approach called DeepPromise based on deep learning techniques for simultaneous prediction of N1-methyladenosine and N6-methyladenosine. To extract the sequence context surrounding the modification sites, three feature encodings, including enhanced nucleic acid composition, one-hot encoding, and RNA embedding, were used as the input to seven consecutive layers of convolutional neural networks (CNNs), respectively. Moreover, DeepPromise further combined the prediction score of the CNN-based models and achieved around 43% higher area under receiver-operating curve (AUROC) for m1A site prediction and 2-6% higher AUROC for m6A site prediction, respectively, when compared with several existing state-of-the-art approaches on the independent test. In-depth analyses of characteristic sequence motifs identified from the convolution-layer filters indicated that nucleotide presentation at proximal positions surrounding the modification sites contributed most to the classification, whereas those at distal positions also affected classification but to different extents. To maximize user convenience, a web server was developed as an implementation of DeepPromise and made publicly available at http://DeepPromise.erc.monash.edu/, with the server accepting both RNA sequences and genomic sequences to allow prediction of two types of putative RNA-modification sites. Zhen Chen 0009, Fuyi Li, Yanan Wang 0003, Alexander Ian Smith, Geoffrey I. Webb, Tatsuya Akutsu, Abdelkader Baggag, Halima Bensmail, Jiangning Song |
Briefings Bioinform. | 9 |
| 2020 | BCrystal: an interpretable sequence-based protein crystallization predictorabstractMOTIVATION: X-ray crystallography has facilitated the majority of protein structures determined to date. Sequence-based predictors that can accurately estimate protein crystallization propensities would be highly beneficial to overcome the high expenditure, large attrition rate, and to reduce the trial-and-error settings required for crystallization. RESULTS: In this study, we present a novel model, BCrystal, which uses an optimized gradient boosting machine (XGBoost) on sequence, structural and physio-chemical features extracted from the proteins of interest. BCrystal also provides explanations, highlighting the most important features for the predicted crystallization propensity of an individual protein using the SHAP algorithm. On three independent test sets, BCrystal outperforms state-of-the-art sequence-based methods by more than 12.5% in accuracy, 18% in recall and 0.253 in Matthew's correlation coefficient, with an average accuracy of 93.7%, recall of 96.63% and Matthew's correlation coefficient of 0.868. For relative solvent accessibility of exposed residues, we observed higher values to associate positively with protein crystallizability and the number of disordered regions, fraction of coils and tripeptide stretches that contain multiple histidines associate negatively with crystallizability. The higher accuracy of BCrystal enables it to accurately screen for sequence variants with enhanced crystallizability. AVAILABILITY AND IMPLEMENTATION: Our BCrystal webserver is at https://machinelearning-protein.qcri.org/ and source code is available at https://github.com/raghvendra5688/BCrystal. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Abdurrahman Elbasir, Raghvendra Mall, Khalid Kunji, Reda Rawi, Zeyaul Islam, Gwo-Yu Chuang, Prasanna R. Kolatkar, Halima Bensmail |
Bioinform. | 8 |
| 2019 | DeepCrystal: a deep learning framework for sequence-based protein crystallization predictionabstractMOTIVATION: Protein structure determination has primarily been performed using X-ray crystallography. To overcome the expensive cost, high attrition rate and series of trial-and-error settings, many in-silico methods have been developed to predict crystallization propensities of proteins based on their sequences. However, the majority of these methods build their predictors by extracting features from protein sequences, which is computationally expensive and can explode the feature space. We propose DeepCrystal, a deep learning framework for sequence-based protein crystallization prediction. It uses deep learning to identify proteins which can produce diffraction-quality crystals without the need to manually engineer additional biochemical and structural features from sequence. Our model is based on convolutional neural networks, which can exploit frequently occurring k-mers and sets of k-mers from the protein sequences to distinguish proteins that will result in diffraction-quality crystals from those that will not. RESULTS: Our model surpasses previous sequence-based protein crystallization predictors in terms of recall, F-score, accuracy and Matthew's correlation coefficient (MCC) on three independent test sets. DeepCrystal achieves an average improvement of 1.4, 12.1% in recall, when compared to its closest competitors, Crysalis II and Crysf, respectively. In addition, DeepCrystal attains an average improvement of 2.1, 6.0% for F-score, 1.9, 3.9% for accuracy and 3.8, 7.0% for MCC w.r.t. Crysalis II and Crysf on independent test sets. AVAILABILITY AND IMPLEMENTATION: The standalone source code and models are available at https://github.com/elbasir/DeepCrystal and a web-server is also available at https://deeplearning-protein.qcri.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Abdurrahman Elbasir, Balasubramanian Moovarkumudalvan, Khalid Kunji, Prasanna R. Kolatkar, Raghvendra Mall, Halima Bensmail |
Bioinform. | 6 |
| 2019 | Prediction of protein group function by iterative classification on functional relevance networkabstractMOTIVATION: Biological experiments including proteomics and transcriptomics approaches often reveal sets of proteins that are most likely to be involved in a disease/disorder. To understand the functional nature of a set of proteins, it is important to capture the function of the proteins as a group, even in cases where function of individual proteins is not known. In this work, we propose a model that takes groups of proteins found to work together in a certain biological context, integrates them into functional relevance networks, and subsequently employs an iterative inference on graphical models to identify group functions of the proteins, which are then extended to predict function of individual proteins. RESULTS: The proposed algorithm, iterative group function prediction (iGFP), depicts proteins as a graph that represents functional relevance of proteins considering their known functional, proteomics and transcriptional features. Proteins in the graph will be clustered into groups by their mutual functional relevance, which is iteratively updated using a probabilistic graphical model, the conditional random field. iGFP showed robust accuracy even when substantial amount of GO annotations were missing. The perspective of 'group' function annotation opens up novel approaches for understanding functional nature of proteins in biological systems.Availability and implementation: http://kiharalab.org/iGFP/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ishita K. Khan, Aashish Jain, Reda Rawi, Halima Bensmail, Daisuke Kihara |
Bioinform. | 4 |
| 2019 | KinVis: a visualization tool to detect cryptic relatedness in genetic datasetsabstractMOTIVATION: It is important to characterize individual relatedness in terms of familial relationships and underlying population structure in genome-wide association studies for correct downstream analysis. The characterization of individual relatedness becomes vital if the cohort is to be used as reference panel in other studies for association tests and for identifying ethnic diversities. In this paper, we propose a kinship visualization tool to detect cryptic relatedness between subjects. We utilize multi-dimensional scaling, bar charts, heat maps and node-link visualizations to enable analysis of relatedness information. AVAILABILITY AND IMPLEMENTATION: Available online as well as can be downloaded at http://shiny-vis.qcri.org/public/kinvis/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ehsan Ullah, Michaël Aupetit 0001, Arun Das 0004, Abhishek Patil, Nooral Al Muftah, Reda Rawi, Mohamad Saad 0001, Halima Bensmail |
Bioinform. | 8 |
| 2019 | ClustMe: A Visual Quality Measure for Ranking Monochrome Scatterplots based on Cluster PatternsabstractAbstract We propose ClustMe, a new visual quality measure to rank monochrome scatterplots based on cluster patterns. ClustMe is based on data collected from a human‐subjects study, in which34participants judged synthetically generated cluster patterns in1000scatterplots. We generated these patterns by carefully varying the free parameters of a simple Gaussian Mixture Model with two components, and asked the participants to count the number of clusters they could see (1 or more than 1). Based on the results, we form ClustMe by selecting the model that best predicts these human judgments among7different state‐of‐the‐art merging techniques (Demp). To quantitatively evaluate ClustMe, we conducted a second study, in which31human subjects ranked 435 pairs of scatterplots of real and synthetic data in terms of cluster patterns complexity. We use this data to compare ClustMe's performance to 4 other state‐of‐the‐art clustering measures, including the well‐known Clumpiness scagnostics. We found that of all measures, ClustMe is in strongest agreement with the human rankings. Mostafa M. Abbas, Michaël Aupetit 0001, Michael Sedlmair, Halima Bensmail |
Comput. Graph. Forum | 4 |
| 2018 | DeepCrystal: A Deep Learning Framework for Sequence-based Protein Crystallization Prediction
Abdurrahman Elbasir, Balasubramanian Moovarkumudalvan, Khalid Kunji, Prasanna R. Kolatkar, Halima Bensmail, Raghvendra Mall |
BIBM | 5 |
| 2018 | DeepSol: a deep learning framework for sequence-based protein solubility predictionabstractMotivation: Protein solubility plays a vital role in pharmaceutical research and production yield. For a given protein, the extent of its solubility can represent the quality of its function, and is ultimately defined by its sequence. Thus, it is imperative to develop novel, highly accurate in silico sequence-based protein solubility predictors. In this work we propose, DeepSol, a novel Deep Learning-based protein solubility predictor. The backbone of our framework is a convolutional neural network that exploits k-mer structure and additional sequence and structural features extracted from the protein sequence. Results: DeepSol outperformed all known sequence-based state-of-the-art solubility prediction methods and attained an accuracy of 0.77 and Matthew's correlation coefficient of 0.55. The superior prediction accuracy of DeepSol allows to screen for sequences with enhanced production capacity and can more reliably predict solubility of novel proteins. Availability and implementation: DeepSol's best performing models and results are publicly deposited at https://doi.org/10.5281/zenodo.1162886 (Khurana and Mall, 2018). Supplementary information: Supplementary data are available at Bioinformatics online. Sameer Khurana, Reda Rawi, Khalid Kunji, Gwo-Yu Chuang, Halima Bensmail, Raghvendra Mall |
Bioinform. | 5 |
| 2017 | An adaptive refinement for community detection methods for disease module identification in biological networks using novel metric based on connectivity, conductance & modularityabstractDisease processes are usually driven by several genes interacting in molecular modules or pathways leading to the disease. The identification of such modules in gene or protein networks is at the core of several analysis methods in biomedical research. However, there is still a need to develop a generic framework to uncover biologically relevant modules for different types of networks. With this pretext in mind, the Disease Module Identification DREAM Challenge was initiated as an effort to systematically assess module identification methods on a panel of 6 diverse state-of-the-art genomic networks. Methods: In this paper, we propose a generic refinement method based on ideas of merging and splitting the hierarchical tree obtained from any community detection technique for constrained disease module identification in biological networks. The only constraint for a module to be considered as a candidate disease module was size of the community to be: 3 ≤ community size ≤ 100. Here, we propose a novel quality metric, called F-score, computed from several unsupervised quality metrics like modularity, conductance and connectivity to determine the quality of a graph partition at a given level of hierarchy. We also propose a quality metric, namely Inverse Confidence, which ranks and prune insignificant modules to obtain a curated list of candidate disease modules for a given biological network. The predicted modules are then evaluated on the basis of the total number of unique candidate modules that are associated with complex traits and diseases from over 200 genome-wide association study (GWAS) datasets. Results: We stood 9thout of a total of 42 teams in the competition at the offical FDR cut-off of 0.05 for identifying statistically significant disease associated modules in the 6 benchmark networks. Our proposed approach detected a total of 44 disease modules in the 6 benchmark networks in comparison to 60 for the winner of the DREAM Challenge. For several benchmark networks we were better or competitive with the winner. Raghvendra Mall, Ehsan Ullah, Khalid Kunji, Halima Bensmail, Michele Ceccarelli |
BIBM | 4 |
| 2017 | Identification of cancer drug sensitivity biomarkersabstractAnti-cancer therapies have different responses to different patients. There is a need to identify biomarkers for effectiveness of drugs beside the biomarkers for the diseases like cancer to advance the field personalized medicine. We have used a panel of cancer cell lines from Genomics of Drug Sensitivity in Cancer (GDSC) to capture the sensitivity of drugs. By combining the genetic information such as mutations, copy number variations and gene expression levels from Catalogue of Somatic Mutations in Cancer (COSMIC) and The Cancer Genome ATLAS (TCGA) for the afore mentioned cell lines, we are able to identify genomic features associated with the sensitivity of the drugs. Ehsan Ullah, Raghvendra Mall, Halima Bensmail, Reda Rawi, Saila Shama, Nooral Al Muftah, Ian Richard Thmpson |
BIBM | 3 |
| 2017 | Advanced Computation of Sparse Precision Matrices for Big Data
Abdelkader Baggag, Halima Bensmail, Jaideep Srivastava |
PAKDD (2) | 2 |
| 2016 | A design study to identify inconsistencies in kinship information: The case of the 1000 Genomes projectabstractGenome Wide Association Studies (GWAS) examine genetic variants in different individuals to detect variants associated to specific diseases. The 1000 Genomes project is such a collaborative research effort to sequence the genomes of at least 1000 participants of 26 different ethnicities, to establish a detailed summary of human genetic variation. The kinship information is a measure of individuals ancestor relationships within the considered populations. We study the design of kinship data visualizations allowing the experts to discover anomalies in GWAS data. The visual analysis of the 1000 Genomes Project kinship data reveals inconsistencies which call for a deeper analysis of the data quality within this project. Michaël Aupetit 0001, Ehsan Ullah, Reda Rawi, Halima Bensmail |
PacificVis | 4 |
| 2016 | Denoised Kernel Spectral data ClusteringabstractKernel Spectral Clustering (KSC) solves a weighted kernel principal component analysis problem in a primal-dual optimization framework. It builds an unsupervised model on a small subset of data using the dual solution of the optimization problem. This allows KSC to have a powerful out-of-sample extension property leading to good cluster generalization w.r.t. unseen data points. However, in the presence of noise that causes overlapping data, the technique often fails to provide good generalization capability. In this paper, we propose a two-step process for clustering noisy data. We first denoise the data using kernel principal component analysis (KPCA) with a recently proposed Model selection criterion based on point-wise Distance Distributions (MDD) to obtain the underlying information in the data. We then use the KSC technique on this denoised data to obtain good quality clusters. One advantage of model based techniques is that we can use the same training and validation set for denoising and for clustering. We discovered that using the same kernel bandwidth parameter obtained from MDD for KPCA works efficiently with KSC in combination with the optimal number of clusters k to produce good quality clusters. We compare the proposed approach with normal KSC and KSC with KPCA using a heuristic method based on reconstruction error for several synthetic and real-world datasets to showcase the effectiveness of the proposed approach. Raghvendra Mall, Halima Bensmail, Rocco Langone, Carolina Varon, Johan A. K. Suykens |
IJCNN | 2 |
| 2016 | COUSCOus: improved protein contact prediction using an empirical Bayes covariance estimatorabstractBACKGROUND: The post-genomic era with its wealth of sequences gave rise to a broad range of protein residue-residue contact detecting methods. Although various coevolution methods such as PSICOV, DCA and plmDCA provide correct contact predictions, they do not completely overlap. Hence, new approaches and improvements of existing methods are needed to motivate further development and progress in the field. We present a new contact detecting method, COUSCOus, by combining the best shrinkage approach, the empirical Bayes covariance estimator and GLasso. RESULTS: Using the original PSICOV benchmark dataset, COUSCOus achieves mean accuracies of 0.74, 0.62 and 0.55 for the top L/10 predicted long, medium and short range contacts, respectively. In addition, COUSCOus attains mean areas under the precision-recall curves of 0.25, 0.29 and 0.30 for long, medium and short contacts and outperforms PSICOV. We also observed that COUSCOus outperforms PSICOV w.r.t. Matthew's correlation coefficient criterion on full list of residue contacts. Furthermore, COUSCOus achieves on average 10% more gain in prediction accuracy compared to PSICOV on an independent test set composed of CASP11 protein targets. Finally, we showed that when using a simple random forest meta-classifier, by combining contact detecting techniques and sequence derived features, PSICOV predictions should be replaced by the more accurate COUSCOus predictions. CONCLUSION: We conclude that the consideration of superior covariance shrinkage approaches will boost several research fields that apply the GLasso procedure, amongst the presented one of residue-residue contact prediction as well as fields such as gene network reconstruction. Reda Rawi, Raghvendra Mall, Khalid Kunji, Mohammed El Anbari, Michaël Aupetit 0001, Ehsan Ullah, Halima Bensmail |
BMC Bioinform. | 7 |
| 2015 | Residue-residue contact prediction in the HIV-1 envelope glycoprotein complexabstractHIV-1 Env glycoprotein complex is the key protein that mediates binding and entry of HIV-1 into human host cells. The complex entry process involves three main steps. First, the attachment, the interaction of gp120 and CD4. Second, the coreceptor binding, where gp120 binds the chemokine receptor CCR5 or CXCR4, and finally, the fusion of viral and host cell membranes. Despite the fact that several coordinate structures of HIV-1 Env in unliganded state exist (as well as in complex with CD4, CD4 mimics, or various antibodies, and of gp41 in intermediate and post-fusion state), a comprehensive understanding of structural arrangements and communication within gp120 and gp41 domains during the entry is far from complete. In this study, we applied a direct amino acid interaction detecting method to analyse the function and structure of HIV-1 Env protein sequences representing all group M subtypes. We identified more than 400 coevolving residue pairs within Env, of which the majority are real contacts and proximal in the available coordinate structures, or have functional implications such as receptor binding, variable loop, gp120-gp41, and interdomain interactions. This work provides a new dimension of information in HIV research, important in assisting protein coordinate structure prediction and in designing new and effective entry inhibitors. Reda Rawi, Khalid Kunji, Abdelali Haoudi, Halima Bensmail |
BIBM | 4 |
| 2015 | Supervised Cross-Modal Factor Analysis for Multiple Modal Data ClassificationabstractIn this paper we study the problem of learning from multiple modal data for purpose of document classification. In this problem, each document is composed two different modals of data, i.e., An image and a text. Cross-modal factor analysis (CFA) has been proposed to project the two different modals of data to a shared data space, so that the classification of a image or a text can be performed directly in this space. A disadvantage of CFA is that it has ignored the supervision information. In this paper, we improve CFA by incorporating the supervision information to represent and classify both image and text modals of documents. We project both image and text data to a shared data space by factor analysis, and then train a class label predictor in the shared space to use the class label information. The factor analysis parameter and the predictor parameter are learned jointly by solving one single objective function. With this objective function, we minimize the distance between the projections of image and text of the same document, and the classification error of the projection measured by hinge loss function. The objective function is optimized by an alternate optimization strategy in an iterative algorithm. Experiments in two different multiple modal document data sets show the advantage of the proposed algorithm over other CFA methods. Jingbin Wang, Kanghong Duan, Jim Jing-Yan Wang, Halima Bensmail |
SMC | 5 |
| 2014 | Domain transfer nonnegative matrix factorizationabstractDomain transfer learning aims to learn an effective classifier for a target domain, where only a few labeled samples are available, with the help of many labeled samples from a source domain. The source and target domain samples usually share the same features and class label space, but have significantly different In these experiments error of the classifier distributions. Nonnegative Matrix Factorization (NMF) has been studied and applied widely as a powerful data representation method. However, NMF is limited to single domain learning problem. It can not be directly used in domain transfer learning problem due to the significant differences between the distributions of the source and target domains. In this paper, we extend the NMF method to domain transfer learning problem. The Maximum Mean Discrepancy (MMD) criteria is employed to reduce the mismatch of source and target domain distributions in the coding vector space. Moreover, we also learn a classifier in the coding vector space to directly utilize the class labels from both the two domains. We construct an unified objective function for the learning of both NMF parameters and classifier parameters, which is optimized alternately in an iterative algorithm. The proposed algorithm is evaluated on two challenging domain transfer tasks, and the encouraging experimental results show its advantage over state-of-the-art domain transfer learning algorithms. Jim Jing-Yan Wang, Yijun Sun, Halima Bensmail |
IJCNN | 3 |
| 2014 | Feature selection and multi-kernel learning for sparse representation on a manifold
Jim Jing-Yan Wang, Halima Bensmail, Xin Gao 0001 |
Neural Networks | 2 |
| 2014 | Unified framework for representing and ranking
Jim Jing-Yan Wang, Halima Bensmail |
Pattern Recognit. | 2 |
| 2013 | Cross-domain sparse codingabstractSparse coding has shown its power as an effective data representation method. However, up to now, all the sparse coding approaches are limited within the single domain learning problem. In this paper, we extend the sparse coding to cross domain learning problem, which tries to learn from a source domain to a target domain with significant different distribution. We impose the Maximum Mean Discrepancy (MMD) criterion to reduce the cross-domain distribution difference of sparse codes, and also regularize the sparse codes by the class labels of the samples from both domains to increase the discriminative ability. The encouraging experiment results of the proposed cross-domain sparse coding algorithm on two challenging tasks --- image classification of photograph and oil painting domains, and multiple user spam detection --- show the advantage of the proposed method over other cross-domain data representation methods. Jim Jing-Yan Wang, Halima Bensmail |
CIKM | 2 |
| 2013 | Discriminative sparse coding on multi-manifoldsabstractSparse coding has been popularly used as an effective data representation method in various applications, such as computer vision, medical imaging and bioinformatics. However, the conventional sparse coding algorithms and their manifold-regularized variants (graph sparse coding and Laplacian sparse coding), learn codebooks and codes in an unsupervised manner and neglect class information that is available in the training set. To address this problem, we propose a novel discriminative sparse coding method based on multi-manifolds, that learns discriminative class-conditioned codebooks and sparse codes from both data feature spaces and class labels. First, the entire training set is partitioned into multiple manifolds according to the class labels. Then, we formulate the sparse coding as a manifold–manifold matching problem and learn class-conditioned codebooks and codes to maximize the manifold margins of different classes. Lastly, we present a data sample-manifold matching-based strategy to classify the unlabeled data samples. Experimental results on somatic mutations identification and breast tumor classification based on ultrasonic images demonstrate the efficacy of the proposed data representation and classification approach. Jim Jing-Yan Wang, Halima Bensmail, Nan Yao, Xin Gao 0001 |
Knowl. Based Syst. | 2 |
| 2013 | Multiple graph regularized nonnegative matrix factorization
Jim Jing-Yan Wang, Halima Bensmail, Xin Gao 0001 |
Pattern Recognit. | 2 |
| 2013 | Joint learning and weighting of visual vocabulary for bag-of-feature based tissue classification
Jim Jing-Yan Wang, Halima Bensmail, Xin Gao 0001 |
Pattern Recognit. | 2 |
| 2012 | RAFNI: Robust Analysis of Functional NeuroImages with Non-normal α-Stable Error
Halima Bensmail, Samreen Anjum, Othmane Bouhali, Mohammed El Anbari |
ICONIP (1) | 1 |
| 2012 | Multiple graph regularized protein domain rankingabstractBACKGROUND: Protein domain ranking is a fundamental task in structural biology. Most protein domain ranking methods rely on the pairwise comparison of protein domains while neglecting the global manifold structure of the protein domain database. Recently, graph regularized ranking that exploits the global structure of the graph defined by the pairwise similarities has been proposed. However, the existing graph regularized ranking methods are very sensitive to the choice of the graph model and parameters, and this remains a difficult problem for most of the protein domain ranking methods. RESULTS: To tackle this problem, we have developed the Multiple Graph regularized Ranking algorithm, MultiG-Rank. Instead of using a single graph to regularize the ranking scores, MultiG-Rank approximates the intrinsic manifold of protein domain distribution by combining multiple initial graphs for the regularization. Graph weights are learned with ranking scores jointly and automatically, by alternately minimizing an objective function in an iterative algorithm. Experimental results on a subset of the ASTRAL SCOP protein domain database demonstrate that MultiG-Rank achieves a better ranking performance than single graph regularized ranking methods and pairwise similarity based ranking methods. CONCLUSION: The problem of graph model and parameter selection in graph regularized protein domain ranking can be solved effectively by combining multiple graphs. This aspect of generalization introduces a new frontier in applying multiple graphs to solving protein domain ranking applications. Halima Bensmail, Xin Gao 0001 |
BMC Bioinform. | 2 |
| 2005 | A novel approach for clustering proteomics data using Bayesian fast Fourier transformabstractMOTIVATION: Bioinformatics clustering tools are useful at all levels of proteomic data analysis. Proteomics studies can provide a wealth of information and rapidly generate large quantities of data from the analysis of biological specimens. The high dimensionality of data generated from these studies requires the development of improved bioinformatics tools for efficient and accurate data analyses. For proteome profiling of a particular system or organism, a number of specialized software tools are needed. Indeed, significant advances in the informatics and software tools necessary to support the analysis and management of these massive amounts of data are needed. Clustering algorithms based on probabilistic and Bayesian models provide an alternative to heuristic algorithms. The number of clusters (diseased and non-diseased groups) is reduced to the choice of the number of components of a mixture of underlying probability. The Bayesian approach is a tool for including information from the data to the analysis. It offers an estimation of the uncertainties of the data and the parameters involved. RESULTS: We present novel algorithms that can organize, cluster and derive meaningful patterns of expression from large-scaled proteomics experiments. We processed raw data using a graphical-based algorithm by transforming it from a real space data-expression to a complex space data-expression using discrete Fourier transformation; then we used a thresholding approach to denoise and reduce the length of each spectrum. Bayesian clustering was applied to the reconstructed data. In comparison with several other algorithms used in this study including K-means, (Kohonen self-organizing map (SOM), and linear discriminant analysis, the Bayesian-Fourier model-based approach displayed superior performances consistently, in selecting the correct model and the number of clusters, thus providing a novel approach for accurate diagnosis of the disease. Using this approach, we were able to successfully denoise proteomic spectra and reach up to a 99% total reduction of the number of peaks compared to the original data. In addition, the Bayesian-based approach generated a better classification rate in comparison with other classification algorithms. This new finding will allow us to apply the Fourier transformation for the selection of the protein profile for each sample, and to develop a novel bioinformatic strategy based on Bayesian clustering for biomarker discovery and optimal diagnosis. Halima Bensmail, Jennifer Golek, Michelle M. Moody, O. John Semmes, Abdelali Haoudi |
Bioinform. | 1 |