EDBT 2026 Demo / reviewers in the wild / expert
Luis Rueda 0001
dblp:r/LuisRueda · also Luís G. Rueda
· DBLP profile ↗
71ranked-venue papers
24as first author
13since 2021 · last 2025
0000-0001-7988-2058ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 25 · 15 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Delaunay Triangulations: A New Avenue for Classification of Biomedical Images Using Graph Neural NetworksabstractWe present a geometry-driven graph learning framework for histopathology that converts Tissue Microarray (TMA) images into Delaunay triangulation graphs of superpixels. Unlike convolutional-based approaches that rely on pre-trained feature extractors, our Voronoi Graph Convolutional Network (VGCN) uses compact node descriptors—spatial coordinates and LAB statistics—preserving structural context while reducing computation. Two variants are evaluated on the Harvard Dataverse prostate TMA dataset: DTGNN-Class (benign vs. cancerous) and DTGNN-Grade (multi-label Gleason component classification), exhibiting 90.6% and 74.6% accuracy respectively. Mustafa Mohammadi Gharasuie, Luis Rueda 0001 |
BIBE | 2 |
| 2025 | Heterophily-Aware Hypergraph Neural Networks for Cell Type Prediction Using Ligand-Receptor-Informed Single-Cell RNA-Seq DataabstractAccurately predicting cell types from single-cell RNA sequencing (scRNA-seq) data requires modeling complex cellular interactions that extend beyond pairwise transcriptional similarity. Ligand—receptor-mediated signaling is a key driver of such interactions, often spanning diverse cell types and exhibiting both homophilic and heterophilic patterns. In this work, we introduce a biologically informed framework for cell type prediction based on heterophily-aware hypergraph neural networks (HGNNs), where hyperedges represent multi-cell communication events derived from curated ligand—receptor pairs. This construction enables higher-order modeling of intercellular signaling and captures the combinatorial nature of ligand—receptor communication. We evaluate nine state-of-the-art hypergraph-based models—including HGNN, HyperGCN, UniGCNII, HyperND, AllDeepSets, AllSetTransformer, ED-HNN, SheafHyperGNN, and HyperUFG—that encompass diverse message passing paradigms such as spectral convolutions, diffusion dynamics, permutationinvariant set operations, and sheaf-theoretic encoding. Experiments on six benchmark scRNA-seq datasets reveal that architectures tailored to heterophilic structure substantially outperform their homophily-oriented counterparts. Our results underscore the importance of both biologically grounded hypergraph design and heterophily-aware learning in advancing automated cell type annotation for complex tissue systems. Mahshad Hashemi, Sharjeel Mustafa, Alioune Ngom, Luis Rueda 0001 |
BIBM | 4 |
| 2025 | HeteroGraphNet: A Ligand-Receptor Informed, Heterophily-Adapted Graph Neural Network for Cell Type Prediction in scRNA-Seq DataabstractGraph Neural Networks (GNNs) have emerged as powerful tools for modeling complex relational data, yet most existing architectures assume homophily-where connected nodes share similar features-an assumption that does not hold in many biological systems. In single-cell RNA sequencing (scRNA-seq) data, intercellular communication networks often exhibit heterophily, with meaningful interactions occurring between dissimilar cell types. Moreover, conventional graph construction in this domain frequently relies on arbitrary similarity thresholds, overlooking biologically validated interaction pathways. We address these limitations with HeteroGraphNet, a heterophily-adapted GNN that incorporates ligand-receptor ($\mathbf{L}-\mathbf{R}$) interactions inferred from scRNA-seq data using LIANA to construct biologically grounded cell-cell graphs. Our model combines a bi-kernel aggregation mechanism-capable of capturing both homophilic and heterophilic signals-with cosine similarity-guided adaptive random walks that dynamically update neighborhoods during training. We further mitigate class imbalance through weighted loss functions, ensuring robust performance across underrepresented cell types. Across six scRNA-seq datasets, HeteroGraphNet consistently outperforms a multi-layer perceptron (MLP), four standard GNNs (GCN, GAT, GraphSAGE, MixHop), and two heterophily-specific GNNs (H2GCN, GBK-GNN), with especially strong gains on low-homophily graphs. These results demonstrate that incorporating ligand-receptor-informed connectivity with adaptive neighborhood exploration enables more accurate and biologically meaningful cell type prediction in heterogeneous single-cell interaction networks, offering a scalable framework for broader biological network analysis. Mahshad Hashemi, Sharjeel Mustafa, Alioune Ngom, Luis Rueda 0001 |
BIBM | 4 |
| 2024 | Harnessing the Power of Graph Propagation in Lung Nodule Detection
Sudipta Modak, Yash Trivedi, Esam Abdel-Raheem, Luis Rueda 0001 |
AIME (2) | 4 |
| 2024 | Improving Out-of-Distribution Data Handling and Corruption Resistance via Modern Hopfield Networks
Saleh Sargolzaei, Luis Rueda 0001 |
ICPR (26) | 2 |
| 2024 | SEGCECO: Subgraph Embedding of Gene expression matrix for prediction of CEll-cell COmmunicationabstractRecent advances in single-cell RNA sequencing technology have eased analyses of signaling networks of cells. Recently, cell-cell interaction has been studied based on various link prediction approaches on graph-structured data. These approaches have assumptions about the likelihood of node interaction, thus showing high performance for only some specific networks. Subgraph-based methods have solved this problem and outperformed other approaches by extracting local subgraphs from a given network. In this work, we present a novel method, called Subgraph Embedding of Gene expression matrix for prediction of CEll-cell COmmunication (SEGCECO), which uses an attributed graph convolutional neural network to predict cell-cell communication from single-cell RNA-seq data. SEGCECO captures the latent and explicit attributes of undirected, attributed graphs constructed from the gene expression profile of individual cells. High-dimensional and sparse single-cell RNA-seq data make converting the data into a graphical format a daunting task. We successfully overcome this limitation by applying SoptSC, a similarity-based optimization method in which the cell-cell communication network is built using a cell-cell similarity matrix which is learned from gene expression data. We performed experiments on six datasets extracted from the human and mouse pancreas tissue. Our comparative analysis shows that SEGCECO outperforms latent feature-based approaches, and the state-of-the-art method for link prediction, WLNM, with 0.99 ROC and 99% prediction accuracy. The datasets can be found at https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE84133 and the code is publicly available at Github https://github.com/sheenahora/SEGCECO and Code Ocean https://codeocean.com/capsule/8244724/tree. Akram Vasighizaker, Sheena Hora, Raymond Zeng, Luis Rueda 0001 |
Briefings Bioinform. | 4 |
| 2023 | Lung Nodule Segmentation on CT Scan Images Using Patchwise Iterative Graph ClusteringabstractOne of the most important steps in lung nodule diagnosis is the automatic segmentation of nodules irrespective of their position and size in the lung parenchyma. In this paper, we propose a new way of applying graph clustering to nodule segmentation. Firstly, the image is preprocessed to extract the lung parenchyma from the CT scan image and identify the region of interest. This is followed by the application of Patchwise Iterative Graph Clustering to spilt the patches and generate superpixels. Next, a region adjacency graph is generated, and agglomerative hierarchical clustering is used to merge the superpixels into different structures such as nodules, and blood vessels. A thresholding algorithm is then used to extract the nodules from the clusters. The proposed method has shown good performance in segmentation with an average dice score of 0.88, an intersection over union score of 0.81, and a high average sensitivity of 89.32 %. Furthermore, the proposed method has been compared to several state-of-the-art methods in the field and has shown an increase in performance in terms of the evaluation metrics. Sudipta Modak, Esam Abdel-Raheem, Luis Rueda 0001 |
ISCAS | 3 |
| 2023 | Robust Emotion Recognition in EEG Signals Based on a Combination of Multiple Domain Adaptation TechniquesabstractConventional classification approaches for EEG- based emotion recognition cannot often adapt to different domains, such as cross-subject or cross-dataset scenarios, leading to poor performance. To handle this challenge, we introduce a novel fusion method using a combination of multiple domain adaptation techniques to improve the emotional states in EEG datasets via classification accuracy. For this aim, Our proposed approach exploits domain adaptation approaches such as Transfer Component Analysis (TCA), Correlation Alignment (CORAL), Transfer Joint Matching (TJM), Geodesic Flow Kernel (GFK), and Joint Distribution Adaptation (JDA), to enhance the overall classification performance. Later, a new fusion approach called Multiple Domain Adaptation based on a Neuro-Fuzzy Inference System (MDA-NF) is applied to combine the classifiers using proper fuzzy membership functions and deliver maximum separation between classes. The main contribution is by applying the fusion approach using MDA- NF technique, adaptability is sufficiently enhanced. Another advantage is to employ multiple adaptation techniques that improve separation between classes. In experimental test results conducted with cross-subject and cross-dataset scenarios, the MDA-NF approach demonstrates superior performance in terms of accuracy for both the valence and arousal aspects, as observed in two public DEAP and DREAMER datasets. Alireza Mirzaee, Mojtaba Kordestani, Luis Rueda 0001, Mehrdad Saif |
SMC | 3 |
| 2022 | DeePSLiM: A Deep Learning Approach to Identify Predictive Short-linear Motifs for Protein Sequence ClassificationabstractSLiMs (Short Linear Motifs) are patterns of three to 20 amino acids within proteins that are sufficient to fulfill certain functions. SLiMs play a critical role in many biological processes. Hence, with the increasing quantity of biological data, it is important to develop algorithms that can quickly find patterns in large databases of DNA, RNA and protein sequences. Previous research has been very successful at applying deep learning methods to the problems of motif detection as well as classification of biological sequences. There are, however, limitations to these approaches. Most are limited to finding motifs of a single length. In addition, most research has focused on DNA and RNA, both of which use a four-letter alphabet. A few of these have attempted to apply deep learning methods on the larger, twenty letter, alphabet of proteins. We present an enhanced deep learning model, called DeePSLiM, capable of detecting predictive SLiMs in protein sequences. The model is a shallow network that can be trained quickly on large amounts of data. The SLiMs are predictive because they can be used to classify the sequences into their respective families. In this study, first, we propose a new deep learning approach for finding predictive SLiMs in protein sequences. Then, we use these predictive SLiMs for the classification task of protein sequences to evaluate our proposed method. The model was able to reach scores of 94.5% on accuracy, precision, recall, F1-Score and Matthews-correlation coefficient, as well as 99.9% area under the receiver operator characteristic curve (AUROC). Availability: The source code, sample data, and supplementary material are available via a Github project at https://github.com/sshaghayeghs/DeePSLiM. Alexandru Filip, Seyedeh Shaghayegh Sadeghi, Alioune Ngom, Luis Rueda 0001 |
CIBCB | 4 |
| 2022 | Computationally repurposing drugs for breast cancer subtypes using a network-based approachabstract'De novo' drug discovery is costly, slow, and with high risk. Repurposing known drugs for treatment of other diseases offers a fast, low-cost/risk and highly-efficient method toward development of efficacious treatments. The emergence of large-scale heterogeneous biomolecular networks, molecular, chemical and bioactivity data, and genomic and phenotypic data of pharmacological compounds is enabling the development of new area of drug repurposing called 'in silico' drug repurposing, i.e., computational drug repurposing (CDR). The aim of CDR is to discover new indications for an existing drug (drug-centric) or to identify effective drugs for a disease (disease-centric). Both drug-centric and disease-centric approaches have the common challenge of either assessing the similarity or connections between drugs and diseases. However, traditional CDR is fraught with many challenges due to the underlying complex pharmacology and biology of diseases, genes, and drugs, as well as the complexity of their associations. As such, capturing highly non-linear associations among drugs, genes, diseases by most existing CDR methods has been challenging. We propose a network-based integration approach that can best capture knowledge (and complex relationships) contained within and between drugs, genes and disease data. A network-based machine learning approach is applied thereafter by using the extracted knowledge and relationships in order to identify single and pair of approved or experimental drugs with potential therapeutic effects on different breast cancer subtypes. Indeed, further clinical analysis is needed to confirm the therapeutic effects of identified drugs on each breast cancer subtype. Forough Firoozbakht, Iman Rezaeian, Luis Rueda 0001, Alioune Ngom |
BMC Bioinform. | 3 |
| 2022 | Guest editorial: Deep neural networks for precision medicine
Fang-Xiang Wu, Min Li 0007, Lukasz A. Kurgan, Luis Rueda 0001 |
Neurocomputing | 4 |
| 2022 | Identification of Enriched Regions in ChIP-Seq Data via a Linear-Time Multi-Level Thresholding AlgorithmabstractChromatin immunoprecipitation (ChIP-Seq) has emerged as a superior alternative to microarray technology as it provides higher resolution, less noise, greater coverage and wider dynamic range. While ChIP-Seq enables probing of DNA-protein interaction over the entire genome, it requires the use of sophisticated tools to recognize hidden patterns and extract meaningful data. Over the years, various attempts have resulted in several algorithms making use of different heuristics to accurately determine individual peaks corresponding to unique DNA-protein. However, finding all the significant peaks with high accuracy in a reasonable time is still a challenge. In this work, we propose the use of Multi-level thresholding algorithm, which we call LinMLTBS, used to identify the enriched regions on ChIP-Seq data. Although various suboptimal heuristics have been proposed for multi-level thresholding, we emphasize on the use of an algorithm capable of obtaining an optimal solution, while maintaining linear-time complexity. Testing various algorithm on various ENCODE project datasets shows that our approach attains higher accuracy relative to previously proposed peak finders while retaining a reasonable processing speed. Musab Naik, Luis Rueda 0001, Akram Vasighizaker |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Spam review detection using self-organizing maps and convolutional neural networks
Ashraf Neisari, Luis Rueda 0001, Sherif Saad |
Comput. Secur. | 2 |
| 2020 | Unsupervised Identification of SARS-CoV-2 Target Cell Groups via Nonlinear Dimensionality Reduction on Single-cell RNA-Seq DataabstractRecent emergence of a new coronavirus, SARS-CoV2, has caused the disease COVID-19 and has been declared a worldwide pandemic. Identification of relevant modules such as target cells is a significant step for characterizing diseases and consequently leads to better diagnosis, treatment and prognosis. High-throughput single-cell RNA-Seq (scRNA-seq) technologies have advanced in recent years, enabling researchers to investigate cells individually and understand their biological mechanisms. Computational techniques such as data clustering, which are categorized via unsupervised learning methods, are the more suitable for the pre-processing step in scRNA-seq data analysis. They can be used to identify a group of genes that belong to a specific cell type based on similar gene expression patterns. However, due to the sparsity and high-dimensional nature of this type of data, classical clustering methods are not efficient. Therefore, the use of nonlinear dimensionality reduction techniques to improve clustering results is crucial. In this work, we aim to find representative clusters of SARS-CoV-2 target cell lung by combining dimensionality reduction and clustering techniques. We first perform upstream analysis on data, including normalization and filtering using quality control metrics. We then assess the impact of different dimensionality reduction techniques on the clustering results. Our results show that modified Locally Linear Embedding combined with Independent Component Analysis have a very positive impact on clustering large-scale COVID19 scRNA-seq data. To validate our findings, we identified target cell types involved in immune system functionality and a list of overlapping marker genes among COVID-19, Influenza A and HSV-1 infection. Saiteja Danda, Akram Vasighizaker, Luis Rueda 0001 |
BIBM | 3 |
| 2020 | iSOM-GSN: an integrative approach for transforming multi-omic data into gene similarity networks via self-organizing mapsabstractMOTIVATION: One of the main challenges in applying graph convolutional neural networks (CNNs) on gene-interaction data is the lack of understanding of the vector space to which they belong, and also the inherent difficulties involved in representing those interactions on a significantly lower dimension, viz Euclidean spaces. The challenge becomes more prevalent when dealing with various types of heterogeneous data. We introduce a systematic, generalized method, called iSOM-GSN, used to transform 'multi-omic' data with higher dimensions onto a 2D grid. Afterwards, we apply a CNN to predict disease states of various types. Based on the idea of Kohonen's self-organizing map, we generate a 2D grid for each sample for a given set of genes that represent a gene similarity network. RESULTS: We have tested the model to predict breast and prostate cancer using gene expression, DNA methylation and copy number alteration. Prediction accuracies in the 94-98% range were obtained for tumor stages of breast cancer and calculated Gleason scores of prostate cancer with just 14 input genes for both cases. The scheme not only outputs nearly perfect classification accuracy, but also provides an enhanced scheme for representation learning, visualization, dimensionality reduction and interpretation of multi-omic data. AVAILABILITY AND IMPLEMENTATION: The source code and sample data are available via a Github project at https://github.com/NaziaFatima/iSOM_GSN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nazia Fatima, Luis Rueda 0001 |
Bioinform. | 2 |
| 2018 | Identifying suutype specific network-Uiomarkers of breast cancer survivauilityabstractCurrent studies of breast cancer find small subsets of gene biomarkers able to accurately predict the survivability of patients. In these studies, the selected genes are not necessarily functionally related, and hence, they may not correctly indicate the molecular mechanism behind breast cancer survivability. Also, several studies have shown there is a very low overlap between the different respective biomarkers subsets for the same cancer disease. To improve the robustness of classification performance and stability of detected biomarkers, recent methods take existing knowledge on relations between genes into account in the classifier, by aggregating functionality related genes to produce discriminative gene subnetworks called network-biomarkers. In this paper, given a breast cancer dataset of patients with different subtypes, we devise a novel network-based approach by integrating protein-protein interaction network (PPI) with gene expression data (1) to identify the network-biomarkers (metagene) of breast cancer survivability and (2) to predict the survivability of breast cancer patients based on subtypes. Our method uses the concept of seed gene for identification of network-biomarkers, ADASYN to solve class imbalance and random forest to predict survivability of patients. We obtained best classification performance with distance 3 from seed gene protein where the gmean, fl-measure and accuracy are respectively 0.900, 0.800 and 90.34%. The maximum size of a network biomarkers with distance 3 is 9. Maximum 34 genes are needed to predict survivability of patients. Sheikh Jubair, Luis Rueda 0001, Alioune Ngom |
IJCNN | 2 |
| 2018 | The predictive performance of short-linear motif features in the prediction of calmodulin-binding proteinsabstractBACKGROUND: The prediction of calmodulin-binding (CaM-binding) proteins plays a very important role in the fields of biology and biochemistry, because the calmodulin protein binds and regulates a multitude of protein targets affecting different cellular processes. Computational methods that can accurately identify CaM-binding proteins and CaM-binding domains would accelerate research in calcium signaling and calmodulin function. Short-linear motifs (SLiMs), on the other hand, have been effectively used as features for analyzing protein-protein interactions, though their properties have not been utilized in the prediction of CaM-binding proteins. RESULTS: We propose a new method for the prediction of CaM-binding proteins based on both the total and average scores of known and new SLiMs in protein sequences using a new scoring method called sliding window scoring (SWS) as features for the prediction module. A dataset of 194 manually curated human CaM-binding proteins and 193 mitochondrial proteins have been obtained and used for testing the proposed model. The motif generation tool, Multiple EM for Motif Elucidation (MEME), has been used to obtain new motifs from each of the positive and negative datasets individually (the SM approach) and from the combined negative and positive datasets (the CM approach). Moreover, the wrapper criterion with random forest for feature selection (FS) has been applied followed by classification using different algorithms such as k-nearest neighbors (k-NN), support vector machines (SVM), naive Bayes (NB) and random forest (RF). CONCLUSIONS: Our proposed method shows very good prediction results and demonstrates how information contained in SLiMs is highly relevant in predicting CaM-binding proteins. Further, three new CaM-binding motifs have been computationally selected and biologically validated in this study, and which can be used for predicting CaM-binding proteins. Yixun Li, Mina Maleki, Nicholas Carruthers, Paul M. Stemmer, Alioune Ngom, Luis Rueda 0001 |
BMC Bioinform. | 6 |
| 2017 | A Hybrid Scheme for Fault Diagnosis with Partially Labeled Sets of ObservationsabstractMachine learning techniques are widely used for diagnosing faults to guarantee the safe and reliable operation of the systems. Among various techniques, semi-supervised learning can help in diagnosing faulty states and decision making in partially labeled data, where only a few number of labeled observations along with a large number of unlabeled observations are collected from the process. Thus, it is crucial to conduct a critical study on the use of semi-supervised techniques for both dimensionality reduction and fault classification. In this work, three state-of-the- art semi-supervised dimensionality reduction techniques are used to produce informative features for semi-supervised fault classifiers. This study aims to achieve the best pair of the semisupervised dimensionality reduction and classification techniques that can be integrated into the diagnostic scheme for decision making under partially labeled sets of observations. Roozbeh Razavi-Far, Ehsan Hallaji, Mehrdad Saif, Luis Rueda 0001 |
ICMLA | 4 |
| 2016 | Identification of discriminative genes for predicting breast cancer subtypesabstractBreast cancer is a widespread cancer type in females and accounts for lots of cancer cases and cancer deaths in the world. Identifying the type of breast cancer plays a crucial role in selecting the best treatment. In this paper an optimized hierarchical model is proposed to predict the breast cancer subtype. Suitable filter feature selection methods and new hybrid feature selection methods are utilized in our model to find discriminative genes. The multi-class problem is handled using a proper classifier at each step in the hierarchical model to separate a subtype from the others. The parameters of each classifier are optimized to achieve a better performance. Our proposed model achieves 100% of accuracy for predicting the breast cancer subtypes using the same or even less number of genes. Roohollah Etemadi, Abedalrhman Alkhateeb, Iman Rezaeian, Luis Rueda 0001 |
BIBM | 4 |
| 2016 | A new feature selection approach for optimizing prediction models, applied to breast cancer subtype classificationabstractFeature selection is a useful technique in classification (and regression) problems to find the most informative features for predicting but still preserves the data generality. However, some feature subset searching methods are too exhaustive while others are too greedy. On the other hand, parameter searching is another factor to improve the prediction performance. But, if it is conducted separately after feature selection stage the classification model might not be as optimal as it should. In this study, we propose a new method, called Apriori-like Feature Selection that can overcome those drawbacks. Given a classifier and a dataset, it searches for the optimal parameters and the optimal feature subset in the combined space of features and parameters. Moreover, its greedy search behavior is controllable by running options. When applying this approach on a breast cancer dataset of five subtypes, it yielded the overall classification accuracy of more than 99% but requires only about 12 genes; a significant improvement as compared to another study. Huy Quang Pham, Alioune Ngom, Luis Rueda 0001 |
BIBM | 3 |
| 2016 | Efficient feature extraction of vibration signals for diagnosing bearing defects in induction motorsabstractThis paper presents a model to extract and select a proper set of features for diagnosing bearing defects in induction motors. An efficient pre-processing of the vibration signals is of paramount importance to provide informative features for the fault classification module. The vibration signals are firstly analyzed by the wavelet packet transform to extract informative frequency domain features. The dimension of the set of extracted features is reduced by resorting to linear discriminant analysis to provide a small-size set of informative features for decision making. The fault classification module contains different classifiers that can learn the features-faults relations and classify multiple bearing defects including ball, inner race and outer race defects of different diameters. Experimental results verify the effectiveness of the proposed technique for diagnosing multiple bearing defects in induction motors. Maryam Farajzadeh-Zanjani, Roozbeh Razavi-Far, Mehrdad Saif, Luis Rueda 0001 |
IJCNN | 4 |
| 2016 | A novel model used to detect differential splice junctions as biomarkers in prostate cancer from RNA-Seq data
Iman Rezaeian, Ahmad Tavakoli, Dora Cavallo-Medved, Lisa A. Porter, Luis Rueda 0001 |
J. Biomed. Informatics | 5 |
| 2015 | ZSeq 2.0: A fully automatic preprocessing method for next generation sequencing dataabstractPreprocessing is a critical step in next generation sequencing (NGS) data analysis, since any error or artifact in library preparation and the sequencing process can affect subsequent steps, leading to possibly erroneous biological conclusions. In this work, we propose ZSeq 2.0, a fully automatic NGS preprocessing method, which combines the strength of the original ZSeq method with a free of parameters scheme that automatically detects and filters out low complexity and highly biased regions, without any need for parameter adjustment. We estimate parameters by applying dynamic penalty rates to high and low GC-content sequences. We also use a labeling rule method to detect outlier sequences that have very low NUS. Some other preprocessing features have been added to ZSeq2.0, including adapter detection and low-quality nucleotides trimming at each side of the sequence. ZSeq2.0 is publicly available and can be downloaded from http://sourceforge.net/p/ZSeq/wiki/Home/. Abed Alkhateeb, Iman Rezaeian, Luis Rueda 0001 |
BIBM | 3 |
| 2015 | Obtaining biomarkers in cancer progression from outliers of time-series clustersabstractStudying the expression of transcripts throughout the various stages of prostate cancer may provide insight into the factors that influence the progression of the disease. Moreover, it may also reveal outlier transcripts, which have different trends than the majority of the transcripts. In this study, we use a time-series profile hierarchical clustering method to separate dissimilar groups of aligned transcripts that have maximum distance with the other group expression patterns throughout the various stages/sub-stages of prostate cancer progression. The isolated outliers can serve as biomarkers in analyzing different stages/sub-stages. This paper suggests that the combination of proper clustering, distance function and index validation for clusters are suitable model to find a pattern of trending for transcript abundance throughout different prostate cancer stages/sub-stages. The stages/sub-stages represent the time points, and the growth of the transcript abundance throughout those time points are cubic spline interpolated. The trending throughout those stages can lead to understanding the relationships among the transcripts and provide a better analysis of prostate cancer development through stages. Abed Alkhateeb, Iman Rezaeian, Siva Singireddy, Luis Rueda 0001 |
BIBM | 4 |
| 2015 | A new compact set of biomarkers for distinguishing among ten breast cancer subtypesabstractWorld-wide, one in nine women are diagnosed with breast cancer in their lifetime and breast cancer is the second leading cause of death among women. Accurate diagnosis of the specific subtypes of this disease is vital to ensure that the patients will have the best possible response to therapy. Using the newly proposed ten subtypes of breast cancer we hypothesized that machine learning techniques would offer many benefits for selecting the most informative biomarkers. Unlike existing gene selection approaches, we use a hierarchical classification approach that selects genes and builds the classifier concurrently. Our results support that this modified approach to gene selection yields a small subset of 82 genes that can predict each of these ten subtypes with accuracies ranging from 92% to 99%. Forough Firoozbakht, Iman Rezaeian, Alioune Ngom, Luis Rueda 0001 |
BIBM | 4 |
| 2015 | A novel approach for finding informative genes in ten subtypes of breast cancerabstractWorld wide, one in nine women are diagnosed with breast cancer in their lifetime and breast cancer is the second leading cause of death among women. Accurate diagnosis of the specific subtypes of this disease is vital to ensure that patients will have the best possible response to therapy. One way to discriminate subtypes of breast cancer is to study those genes that differentially express across different subtypes. In this study, we use different machine learning techniques to select the most informative genes corresponding to ten subtypes of breast cancer. In particular, we propose a new bottom-up hierarchical classification approach to select the most informative genes for different subtypes, while we identify the similarity level between these subtypes. Our results support that this new approach to gene selection yields a small subset of genes that can predict each of these ten subtypes with very high accuracy. Moreover, the proposed model provides an insightful structure for further analysis of these subtypes. Forough Firoozbakht, Iman Rezaeian, Alioune Ngom, Luis Rueda 0001, Lisa A. Porter |
CIBCB | 4 |
| 2015 | Prediction of high-throughput protein-protein interactions based on protein sequence informationabstractPrediction of protein-protein interaction (PPI) is one of the most challenging problems in biology. Although great progress has been devoted to the development of methodology for predicting PPIs and PIN using machine learning methods, the problem is still far from being solved since the application of most existing methods is limited. In this study, we propose a method for PPI prediction based on amino acids differences between pairs of protein sequences. 10-fold cross-validation tests based on human PPI datasets with balanced positive-to-negative ratios indicate that it performs comparably well. Therefore, our finding suggests that amino acids differences of interacting protein pairs are relevant to the prediction of PPIs and hence provide important information on sequence-based encoding schemes. Yixun Li, Behzad Rezaei, Alioune Ngom, Luis Rueda 0001 |
CIBCB | 4 |
| 2015 | Classification via correlation-based feature groupingabstractEmploying the most relevant and discriminating features is very important to achieve a successful classification with low computational cost. Although, different feature selection methods have been recently developed for this purpose, feature grouping can deal with high dimensional sparse feature vectors more effectively, yielding better interpretation of the data. In this paper, a correlation-based feature grouping (CFG) method is proposed. First, the features are grouped based on the variety of their correlation scores, and then, a new representative feature vector is generated for each group by combining its features. To investigate the strength of CFG method, two filter methods of χ2and correlation are employed for feature selection, while classification is performed using a support vector machine (SVM) and k-Nearest Neighbor (k-NN). The empirical study on two datasets of protein-protein interactions (PPIs) and breast cancer verifies that the idea of employing feature grouping is more efficient than employing feature selection in identifying a set of features that exhibit high classification accuracy. In addition, a CFG diagram is introduced in this paper, which is used to visualize the groups and their corresponding features found by the proposed method. Mina Maleki, Luis Rueda 0001 |
CIBCB | 2 |
| 2015 | Identifying differentially expressed transcripts associated with prostate cancer progression using RNA-Seq and machine learning techniquesabstractBackground: Prostate cancer is complicated by a high level of unexplained variability in the aggressiveness of newly diagnosed disease. Given that this is one of the most prevalent cancers worldwide, finding biomarkers to effectively stratify high risk patient populations is a vital next step in improving survival rates and quality of life after treatment. Materials and Methods: In this study, we selected a dataset consisting of 106 prostate cancer samples, which represent various stages of prostate cancer and developed by RNA-Seq technology. Our objective is to identify differentially expressed transcripts associated with prostate cancer progression using pair-wise stage comparisons. Results: Using machine learning techniques, we identified 44 transcripts that are correlated to different stages of progression. Expression of an identified transcript, USP13, is reduced in stage T3 in comparison with stage T2c, a pattern also observed in breast cancer tumourigenesis. We also identified another differentially expressed transcript, PTGFR, which has also been reported to be involved in prostate cancer progression and has also been linked to breast, ovarian and renal cancers. Conclusions: The results support the use of RNA-Seq along with machine learning techniques as an essential tool in identifying potential biomarkers for prostate cancer progression. Further studies elucidating the biochemical role of identified transcripts in vitro are crucial in validating the use of these biomarkers in the prediction of disease progression and development of effective therapeutic strategies. Siva Singireddy, Abed Alkhateeb, Iman Rezaeian, Luis Rueda 0001, Dora Cavallo-Medved, Lisa A. Porter |
CIBCB | 4 |
| 2015 | Pattern classification using a new border identification paradigm: The nearest border technique
Yifeng Li 0001, B. John Oommen, Alioune Ngom, Luis Rueda 0001 |
Neurocomputing | 4 |
| 2014 | A model based on minimotifs for classification of stable protein-protein complexesabstractPrediction of protein-protein interactions (PPIs) is an important problem in biology, since interactions play key role in most biological processes and functions in living cells. PPIs have been studied from many perspectives. Of these, an important problem is prediction of different complex types such as obligate vs. non obligate and transient vs. permanent, among others. We focus on prediction of obligate protein complexes, which are more stable and perform a specific function, as opposed to transient and non-obligate complexes which last for a short period of time. We have modeled the prediction problem using minimotifs, aka short-linear motifs, to extract information contained in the protein sequences to distinguish between obligate and non-obligate PPIs. Incorporating different classifiers such as the k-nearest neighbor (k-NN), the support vector machine (SVM) and linear dimensionality reduction (LDR) yields a very powerful scheme for prediction. On two well-known datasets, the model delivers classification accuracies as high as 99%. Analysis and cross-dataset validation show that the information contained in the training sequences is crucial for prediction and determination of stability in PPIs. Luis Rueda 0001, Manish Pandit |
CIBCB | 1 |
| 2012 | Finding genomic features from enriched regions in ChlP-Seq dataabstractFinding genomic features in ChlP-Seq data has become an attractive research topic lately, because of the power, resolution and low-noise of next generation sequencing, making it a much better alternative to traditional microarrays such as ChlP-chip and other related methods. However, handling ChlP-Seq data is not straightforward, mainly because of the large amounts of data produced by next generation sequencing. ChlP-Seq has widespread over a range of applications in finding biomarkers, especially those associated with important genomic features in epigenomics and transcriptomics, including binding sites, promoters, exons/introns, transcription sites, among others. Efficient algorithms for finding relevant regions in ChlP-Seq data have been proposed, which capture the most significant peaks from the sequence reads. Among these, multilevel thresholding algorithms have been applied successfully for transcriptomics and genomics data analysis, in particular for detecting significant regions based on next generation sequencing data. We show that the Optimal Multilevel Thresholding algorithm (OMT) achieves higher accuracy in detecting enriched regions and genomic features of detected regions on FoxAl data. OMT finds more gene-related regions (gene, exon, promoter) in comparison with other methods. Using a small number of parameters is another advantage of the proposed method. Iman Rezaeian, Luis Rueda 0001 |
BIBM | 2 |
| 2012 | A model to predict and analyze protein-protein interaction types using electrostatic energiesabstractIdentification and analysis of types of protein-protein interactions (PPI) is an important problem in molecular biology because of their key role in many biological processes in living cells. We propose a model to predict and analyze protein interaction types using electrostatic energies as properties to distinguish between obligate and non-obligate interactions. Our prediction approach uses electrostatic energies for pairs of atoms and amino acids present in interfaces where the interaction occurs. Our results confirm that electrostatic energy is an important property to predict obligate and non obligate protein interaction types achieving accuracy of over 96% on two well known datasets. The classifiers used are support vector machines and linear dimensionality reduction. Gokul Vasudev, Luis Rueda 0001 |
BIBM | 2 |
| 2012 | Prediction of crystal packing and biological protein-protein interactionsabstractPrediction of protein-protein interactions are important to understand any biological processes. The structural models of the complexes resulting from these interactions are necessary to understand those processes at the molecular level. X-ray crystallography is the most popular method to determine the three dimensional structures of protein complexes. However, some of the observed interactions in the structures of protein complexes determined by X-ray crystallography are crystal packing contacts and are not biologically relevant. Thus, it is important to discriminate between biologically relevant interactions and crystal packing contacts. We propose a classification approach to predict these two types of complexes. Our approach has two main features. Firstly, we have calculated various interface property features from the quaternary structures of these interactions. Various features are extracted for each complex, namely number-based and area-based amino acid compositions. Secondly, these features are treated as the input features of the classifiers. The classification is performed with support vector machines (SVM) and linear dimensionality reduction (LDR) coupled with Bayesian classifiers. The results on a standard benchmark dataset of crystal packing and biological protein complexes show increasing prediction accuracy when compared. Sridip Banerjee, Luis Rueda 0001, Mina Maleki |
CIBCB | 2 |
| 2012 | Using structural domains to predict obligate and non-obligate protein-protein interactionsabstractThe identification and prediction of particular types of protein-protein interactions (PPIs) based on knowledge of their interacting domains is a problem that has drawn the attention of researchers in the past few years. We focus on the prediction and analysis of obligate and non-obligate complexes by using structural domains from the CATH database. Our proposed prediction model uses desolvation energies of domain-domain interactions (DDIs) present in the interfaces of such complexes. The prediction is performed via linear dimensionality reduction (LDR) and support vector machines (SVMs). Our results on two well-known datasets show that DDI features of the first three levels of CATH, especially level 2, are more powerful and discriminative than features of other levels in predicting these types of complexes. Furthermore, a detailed analysis shows that different DDIs are present in obligate and non-obligate complexes, and that homo-DDIs are more likely to be present in obligate interactions. Mina Maleki, Michael Hall, Luis Rueda 0001 |
CIBCB | 3 |
| 2011 | A Novel Recursive Feature Subset Selection AlgorithmabstractUnivariate filter methods, which rank single genes according to how well they each separate the classes, are widely used for gene ranking in the field of microarray analysis of gene expression datasets. These methods rank all of the genes by considering all of the samples; however some of these samples may never be classified correctly by adding new genes and these methods keep adding redundant genes covering only some parts of the space and finally the returned subset of genes may never cover the space perfectly. In this paper we introduce a new gene subset selection approach which aims to add genes covering the space which has not been covered by already selected genes in a recursive fashion. Our approach leads to significant improvement on many different benchmark datasets. Amirali Jafarian, Alioune Ngom, Luis Rueda 0001 |
BIBE | 3 |
| 2011 | Applications of Multilevel Thresholding Algorithms to Transcriptomics Data
Luis Rueda 0001, Iman Rezaeian |
CIARP | 1 |
| 2011 | A Fully Automatic Gridding Method for cDNA Microarray ImagesabstractBACKGROUND: Processing cDNA microarray images is a crucial step in gene expression analysis, since any errors in early stages affect subsequent steps, leading to possibly erroneous biological conclusions. When processing the underlying images, accurately separating the sub-grids and spots is extremely important for subsequent steps that include segmentation, quantification, normalization and clustering. RESULTS: We propose a parameterless and fully automatic approach that first detects the sub-grids given the entire microarray image, and then detects the locations of the spots in each sub-grid. The approach, first, detects and corrects rotations in the images by applying an affine transformation, followed by a polynomial-time optimal multi-level thresholding algorithm used to find the positions of the sub-grids in the image and the positions of the spots in each sub-grid. Additionally, a new validity index is proposed in order to find the correct number of sub-grids in the image, and the correct number of spots in each sub-grid. Moreover, a refinement procedure is used to correct possible misalignments and increase the accuracy of the method. CONCLUSIONS: Extensive experiments on real-life microarray images and a comparison to other methods show that the proposed method performs these tasks fully automatically and with a very high degree of accuracy. Moreover, unlike previous methods, the proposed approach can be used in various type of microarray images with different resolutions and spot sizes and does not need any parameter to be adjusted. Luis Rueda 0001, Iman Rezaeian |
BMC Bioinform. | 1 |
| 2010 | A parameterless automatic spot detection method for cDNA microarray imagesabstractGridding cDNA microarray images is a critical step in gene expression analysis, since any errors in this stage are propagated in future steps in the analysis. We propose a fully automatic approach to detect the locations of the spots. The approach first detects and corrects rotations in the sub-grids by an affine transformation, followed by a polynomial-time optimal multi-level thresholding algorithm that finds the positions of the spots. Additionally, a new validity index is proposed in order to find the correct number of spots in each sub-grid, followed by a refinement procedure used to improve the performance of the method. Extensive experiments on real-life microarray images show that the proposed method performs these tasks automatically and with very high accuracy. Iman Rezaeian, Luis Rueda 0001 |
BIBM | 2 |
| 2010 | Protein-protein interaction prediction using desolvation energies and interface propertiesabstractAn important aspect in understanding and classifying protein-protein interactions (PPI) is to analyze their interfaces in order to distinguish between transient and obligate complexes. We propose a classification approach to discriminate between these two types of complexes. Our approach has two important aspects. First, we have used desolvation energies - amino acid and atom type - of the residues present in the interface, which are the input features of the classifiers. Principal components of the data were found and then the classification is performed via linear dimensionality reduction (LDR) methods. Second, we have investigated various interface properties of these interactions. From the analysis of protein quaternary structures, physicochemical properties are treated as the input features of the classifiers. Various features are extracted from each complex, and the classification is performed via different linear dimensionality reduction (LDR) methods. The results on standard benchmarks of transient and obligate protein complexes show that (i) desolvation energies are better discriminants than solvent accessibility and conservation properties, among others, and (ii) the proposed approach outperforms previous solvent accessible area based approaches using support vector machines. Luis Rueda 0001, Sridip Banerjee, Md. Mominul Aziz, Mohammad Raza |
BIBM | 1 |
| 2010 | Alignment versus variation methods for clustering microarray time-series dataabstractIn the past few years, it has been shown that traditional clustering methods do not necessarily perform well on time-series data because of the temporal relationships involved in such data - this makes it a particularly difficult problem. In this paper, we compare two clustering methods that have been introduced recently, especially for gene expression time-series data, namely, multiple-alignment (MA) clustering and variation-based co-expression detection (VCD) clustering approaches. Both approaches are based on a transformation of the data that takes into account the temporal relationships, and have been shown to effectively detect groups of co-expressed genes. We investigate the performances of the MA and VCD approaches on two microarray time-series data sets and discuss their strengths and weaknesses. Our experiments show the superior accuracy of MA over VCD when finding groups of co-expressed genes. Numanul Subhani, Yifeng Li 0001, Alioune Ngom, Luis Rueda 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2010 | Missing value imputation methods for gene-sample-time microarray data analysisabstractWith the recent advances in microarray technology, the expression levels of genes with respect to the samples can be monitored synchronically over a series of time points. Such three-dimensional microarray data, termed gene-sample-time microarray data or GST data for short, may contain missing values. Current microarray analysis methods require complete data sets, and thus, either each row, column or tube containing missing values must be removed from the original GST data, or these missing values must be estimated before analysis. Imputation of missing values is, however, more recommended than removal of data in order to increase the effectiveness of analysis algorithms. In this paper, we extend automated imputation methods, devised for two-dimensional microarray data, to GST data. We implemented imputation methods for GST data based on Singular Value Decomposition (3SVDimpute), K-Nearest Neighbor (3KNNimpute), and gene and sample average methods (3Aimpute), and show that methods based on KNN yield the best results with the lowest normalized root mean squared error. Yifeng Li 0001, Alioune Ngom, Luis Rueda 0001 |
CIBCB | 3 |
| 2010 | New approaches to clustering microarray time-series data using multiple expression profile alignmentabstractAn important process in functional genomic studies is clustering microarray time-series data, where genes with similar expression profiles are expected to be functionally related. Clustering microarray time-series data via pairwise alignment of piecewise linear profiles has been recently introduced. In this paper, we propose a clustering approach based on a multiple profile alignment of natural cubic spline and piecewise linear representations of gene expression profiles. We combine these multiple alignment approaches with k-means. We ran our methods on a well-known data set of pre-clustered Saccharomyces cerevisiae gene expression profiles and a data set of 3315 Pseudomonas aeruginosa expression profiles. We assessed the validity of the resulting clusters and applied a c-nearest neighbor classifier for evaluating the performance of our approaches, obtaining accuracies of 89.51% and 86.12% respectively, on Saccharomyces cerevisiae data, and 90.90% and 93.71% accuracies for cubic spline and piecewise linear respectively on Pseudomonas aeruginosa data. Numanul Subhani, Luis Rueda 0001, Alioune Ngom, Conrad J. Burden |
CIBCB | 2 |
| 2010 | Multiple gene expression profile alignment for microarray time-series data clusteringabstractMOTIVATION: Clustering gene expression data given in terms of time-series is a challenging problem that imposes its own particular constraints. Traditional clustering methods based on conventional similarity measures are not always suitable for clustering time-series data. A few methods have been proposed recently for clustering microarray time-series, which take the temporal dimension of the data into account. The inherent principle behind these methods is to either define a similarity measure appropriate for temporal expression data, or pre-process the data in such a way that the temporal relationships between and within the time-series are considered during the subsequent clustering phase. RESULTS: We introduce pairwise gene expression profile alignment, which vertically shifts two profiles in such a way that the area between their corresponding curves is minimal. Based on the pairwise alignment operation, we define a new distance function that is appropriate for time-series profiles. We also introduce a new clustering method that involves multiple expression profile alignment, which generalizes pairwise alignment to a set of profiles. Extensive experiments on well-known datasets yield encouraging results of at least 80% classification accuracy. Numanul Subhani, Luis Rueda 0001, Alioune Ngom, Conrad J. Burden |
Bioinform. | 2 |
| 2010 | Multi-class pairwise linear dimensionality reduction using heteroscedastic schemes
Luis Rueda 0001, B. John Oommen, Claudio Henríquez |
Pattern Recognit. | 1 |
| 2010 | Selection based heuristics for the non-unique oligonucleotide probe selection problem in microarray design
Alioune Ngom, Luis Rueda 0001, Robin Gras |
Pattern Recognit. Lett. | 2 |
| 2009 | Biofilm Image Segmentation Using Optimal Multi-level ThresholdingabstractA microbial biofilm is structured mainly by a protective sticky matrix of extracellular polymeric substances. Quantifying such structures is useful for microbiologists and a correct image segmentation process helps substantially reduce errors in quantification. This paper proposes an approach to segmentation of biofilm images using optimal multilevel thresholding and indices of clustering validity. A direct comparison through Rand index and a quantification process is performed in a laboratory, obtaining results similar to the quantification and segmentation done by an expert. Darío Rojas, Luis Rueda 0001, Alioune Ngom, Homero Urrutia, Gerardo Carcamo |
BIBM | 2 |
| 2008 | Chernoff-Based Multi-class Pairwise Linear Dimensionality Reduction
Luis Rueda 0001, Claudio Henríquez, B. John Oommen |
CIARP | 1 |
| 2008 | Evolution strategy with greedy probe selection heuristics for the non-unique oligonucleotide probe selection problemabstractIn order to accurately measure the gene expression levels in microarray experiments, it is crucial to design unique, highly specific and highly sensitive oligonucleotide probes for the identification of biological agents such as genes in a sample. Unique probes are difficult to obtain for closely related genes such as the known strains of HIV genes. The non-unique probe selection problem is to find a smallest probe set that is able to uniquely identify targets in a biological sample. This is an NP-hard problem. We present two approaches for finding near-minimal non-unique probe sets. Each approach combines of a deterministic greedy probe selection heuristic that selects good probes, with an evolution strategy that optimizes the selected probe sets. The heuristics, guided by selection functions defined over a probe set, decide at each moment which probes are the best to be included in, or excluded from, a candidate solution. Our methods produce results that are very close to, and in many cases better than, those of the current state-of-the-art approaches for the non-unique probe selection problem, namely integer linear programming, optimal cutting-plane and genetic algorithm approaches. Alioune Ngom, Robin Gras, Luis Rueda 0001 |
CIBCB | 4 |
| 2008 | Linear dimensionality reduction by maximizing the Chernoff distance in the transformed space
Luis Rueda 0001, Myriam Herrera |
Pattern Recognit. | 1 |
| 2008 | A theoretical comparison of two-class Fisher's and heteroscedastic linear dimensionality reduction schemes
Luis Rueda 0001, Myriam Herrera |
Pattern Recognit. Lett. | 1 |
| 2007 | Sub-grid Detection in DNA Microarray Images
Luis Rueda 0001 |
PSIVT | 1 |
| 2006 | A Theoretical Comparison of Two Linear Dimensionality Reduction Techniques
Luis Rueda 0001, Myriam Herrera |
CIARP | 1 |
| 2006 | A New Approach to Multi-class Linear Dimensionality Reduction
Luis Rueda 0001, Myriam Herrera |
CIARP | 1 |
| 2006 | A fast and efficient nearly-optimal adaptive Fano coding scheme
Luis Rueda 0001, B. John Oommen |
Inf. Sci. | 1 |
| 2006 | Stochastic learning-based weak estimation of multinomial random variables and its applications to pattern recognition in non-stationary environments
B. John Oommen, Luis Rueda 0001 |
Pattern Recognit. | 2 |
| 2006 | Geometric visualization of clusters obtained from fuzzy clustering algorithms
Luis Rueda 0001, Yuanquan Zhang |
Pattern Recognit. | 1 |
| 2006 | A Hill-Climbing Approach for Automatic Gridding of cDNA Microarray ImagesabstractImage and statistical analysis are two important stages of cDNA microarrays. Of these, gridding is necessary to accurately identify the location of each spot while extracting spot intensities from the microarray images and automating this procedure permits high-throughput analysis. Due to the deficiencies of the equipment used to print the arrays, rotations, misalignments, high contamination with noise and artifacts, and the enormous amount of data generated, solving the gridding problem by means of an automatic system is not trivial. Existing techniques to solve the automatic grid segmentation problem cover only limited aspects of this challenging problem and require the user to specify the size of the spots, the number of rows and columns in the grid, and boundary conditions. In this paper, a hill-climbing automatic gridding and spot quantification technique is proposed which takes a microarray image (or a subgrid) as input and makes no assumptions about the size of the spots, rows, and columns in the grid. The proposed method is based on a hill-climbing approach that utilizes different objective functions. The method has been found to effectively detect the grids on microarray images drawn from databases from GEO and the Stanford genomic laboratories. Luis Rueda 0001, Vidya Vidyadharan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2006 | Stochastic Automata-Based Estimators for Adaptively Compressing Files With Nonstationary DistributionsabstractThis correspondence shows that learning automata techniques, which have been useful in developing weak estimators, can be applied to data compression applications in which the data distributions are nonstationary. The adaptive coding scheme utilizes stochastic learning-based weak estimation techniques to adaptively update the probabilities of the source symbols, and this is done without resorting to either maximum likelihood, Bayesian, or sliding-window methods. The authors have incorporated the estimator in the adaptive Fano coding scheme and in an adaptive entropy-based scheme that "resembles" the well-known arithmetic coding. The empirical results obtained for both of these adaptive methods are obtained on real-life files that possess a fair degree of nonstationarity. From these results, it can be seen that the proposed schemes compress nearly 10% more than their respective adaptive methods that use maximum-likelihood estimator-based estimates. Luis Rueda 0001, B. John Oommen |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2005 | A formal analysis of why heuristic functions work
B. John Oommen, Luis Rueda 0001 |
Artif. Intell. | 2 |
| 2005 | A one-dimensional analysis for the probability of error of linear classifiers for normally distributed classes
Luis Rueda 0001 |
Pattern Recognit. | 1 |
| 2004 | New Bounds and Approximations for the Error of Linear Classifiers
Luis Rueda 0001 |
CIARP | 1 |
| 2004 | A nearly-optimal Fano-based coding algorithm
Luis Rueda 0001, B. John Oommen |
Inf. Process. Manag. | 1 |
| 2004 | An efficient approach to compute the threshold for multi-dimensional linear classifiers
Luis Rueda 0001 |
Pattern Recognit. | 1 |
| 2004 | Selecting the best hyperplane in the framework of optimal pairwise linear classifiers
Luis Rueda 0001 |
Pattern Recognit. Lett. | 1 |
| 2003 | A New Approach That Selects a Single Hyperplane from the Optimal Pairwise Linear Classifier
Luis Rueda 0001 |
CIARP | 1 |
| 2003 | On optimal pairwise linear classifiers for normal distributions: the d-dimensional case
Luis Rueda 0001, B. John Oommen |
Pattern Recognit. | 1 |
| 2002 | The Efficiency of Histogram-like Techniques for Database Query OptimizationabstractOne of the most difficult tasks in modern day database management systems is information retrieval. Basically, this task involves a user query, written in a high-level language such as the Structured Query Language, and some internal operations, which are transparent to the user. The internal operations are carried out through very complex modules that decompose, optimize and execute the different operations. We consider the problem of Query Optimization which consists of the system choosing, among many different query evaluation plans (QEPs), the most economical one. Since the number of QEPs increases exponentially as the number of relations involving the query increases, query optimization is a very complex problem. Many estimation techniques have been developed in order to approximate the cost of a QEP. Histogram-based techniques are the most used methods in this context. In this paper, we discuss the efficiency of some of these methods: Equi-width, Equi-depth, the Rectangular Attribute Cardinality Map (R-ACM) and the Trapezoidal Attribute Cardinality Map (T-ACM). These methods are used to estimate the cost of the different QEP, whence they attempt to determine the optimal one. It has been shown that the errors of the estimates from R-ACM and T-ACM are significantly less than the corresponding errors obtained from Equi-width and Equi-depth. This fact has been formally demonstrated using reasonable statistical distributions for the cost of a QEP, the doubly exponential distribution and the normal distribution. For the empirical analysis, we have developed a formal, rigorous prototype model used to analyze these methods on random databases. Our empirical results demonstrate that R-ACM chooses a superior QEP more than two times as often as Equi-width and Equi-depth. Similar results have been obtained for T-ACM when compared to the traditional methods. Indeed, in the most general scenario, we analytically prove that under certain models the better the accuracy of an estimation technique, the greater the probability of choosing the most efficient QEP. B. John Oommen, Luis Rueda 0001 |
Comput. J. | 2 |
| 2002 | On Optimal Pairwise Linear Classifiers for Normal Distributions: The Two-Dimensional CaseabstractOptimal Bayesian linear classifiers have been studied in the literature for many decades. We demonstrate that all the known results consider only the scenario when the quadratic polynomial has coincident roots. Indeed, we present a complete analysis of the case when the optimal classifier between two normally distributed classes is pairwise and linear. We focus on some special cases of the normal distribution with nonequal covariance matrices. We determine the conditions that the mean vectors and covariance matrices have to satisfy in order to obtain the optimal pairwise linear classifier. As opposed to the state of the art, in all the cases discussed here, the linear classifier is given by a pair of straight lines, which is a particular case of the general equation of second degree. We also provide some empirical results, using synthetic data for the Minsky's paradox case, and demonstrated that the linear classifier achieves very good performance. Finally, we have tested our approach on real life data obtained from the UCI machine learning repository. The empirical results that we obtained show the superiority of our scheme over the traditional Fisher's discriminant classifier. Luis Rueda 0001, B. John Oommen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Histogram Methods in Query Optimization: The Relation between Accuracy and OptimalityabstractWe have solved the following problem using pattern classification techniques (PCT): given two histogram methods, M/sub 1/ and M/sub 2/, used in query optimization, if the estimation accuracy of M/sub 1/ is greater than that of M/sub 2/, then M/sub 1/ has a higher probability of leading to the optimal query evaluation plan (QEP) than M/sub 2/. To the best of our knowledge, this problem has been open for at least two decades, the difficulty of the problem partially being due to the hurdles involved in the formulation itself. By formulating the problem from a pattern recognition perspective, we use PCT to present a rigorous mathematical proof of this fact, and show some uniqueness results. We also report empirical results demonstrating the power of these theoretical results on well-known histogram estimation methods. B. John Oommen, Luis Rueda 0001 |
DASFAA | 2 |
| 2001 | Enhanced static Fano codingabstractStatistical coding techniques have been used for a long time in lossless data compression, using methods such as Huffman's algorithm, arithmetic coding, Shannon's method, Fano's method, etc. Most of these methods can be implemented either statically or adaptively. Canonical codes, in which the code words are arranged in a lexicographical order, are advantageous because they can be decoded extremely expediently. Although Huffman's algorithm is optimal, the generation of a canonical Huffman code is not straightforward. Conversely, while the Fano coding is sub-optimal, it can lead to canonical codes. In this paper, we resolve the dilemma by focusing on the static implementation of Fano's method. By taking advantage of the properties of the encoding schemes generated by this method, and the concept of "code word arrangement", we present an enhanced version of the static Fano's method, namely Fano/sup +/. We formally analyze Fanol by presenting some properties of Fano trees, and the theory of list rearrangements. Our enhanced algorithm achieves compression ratios arbitrarily close to those of Huffman's algorithm. Empirical results on files of the Canterbury corpus corroborate the almost-optimal efficiency of our enhanced algorithm and its canonical nature. We believe that the compression efficiency of Fano+ can be made to attain the compression ratios of the best known schemes if a structure model of the data is also incorporated. Luis Rueda 0001, B. John Oommen |
SMC | 1 |