VLDB 2026 Research / reviewers in the wild / expert
Hesham Ali 0001
dblp:a/HHAli · also Hesham H. Ali
· DBLP profile ↗
42ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0002-8016-6144ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 3 since 2021Systems, architecture and hardware · 5Computer networks · 3Human-computer interaction and ubiquitous computing · 2Theory of computation · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On the Robustness of Correlation Network Models in Predicting the Safety of Bridges
Prasad Chetti, Hesham Ali 0001 |
COMPLEXIS | 2 |
| 2023 | A Network Analysis Approach for the Classification of Psychiatric Disorders Using Multi-Modal DataabstractPsychiatric disorders span a broad range of conditions, exhibiting diverse clinical and behavioral traits. Traditional diagnostic methods primarily reliant on subjective behavioral assessments present substantial challenges in achieving accurate identification and characterization. To overcome this challenge, we propose an innovative data-driven integrated classification approach that combines mobility data from wearable sensors with clinical data. Initially, we construct network representations of mobility patterns, enabling the identification of distinct communities within the population. Subsequently, we classify each community by exploring its unique clinical parameters. Our results reveal the presence of at least five distinct communities within the studied population, each characterized by a unique combination of mobility and clinical data. These groups are identified as follows: ADHD (Attention-Deficit / Hyperactivity Disorder) with ADD (Attention-Deficit Disorder), Bipolar with Anxiety, Unipolar with Anxiety, and ADHD with Bipolar and Anxiety. This classification is achieved by leveraging both mobility and clinical data present in the dataset, allowing for a more comprehensive and nuanced understanding of these psychiatric conditions. The resulting findings provide valuable insights into the heterogeneity associated with different disorders, thereby facilitating the development of personalized interventions and treatment strategies. The salient aspect of this study is that our results are neither predicated solely on class labels nor driven solely by mobility data; instead, they derive strength from the harmonious integration of data-driven analysis and crucial clinical insights, culminating in a robust classification model. Rama Krishna Thelagathoti, Hesham Ali 0001 |
BIBM | 2 |
| 2022 | A Data-Driven Approach for the Analysis of Behavioral Disorders With a Focus on Classification and Severity EstimationabstractIt has been well-established that behavioral disorders such as depression and schizophrenia are hard to diagnose. Classifying the type of the disorder is a challenging task because of the diversity of the symptoms. Furthermore, estimation of the severity level of such disorders is another challenging factor in the diagnosis process. Current clinical diagnostic procedures tend to be mostly dependent on the self-reporting of patients or the limited observational evaluations of clinicians. Hence, there is a need for objective analytical methods that leverage the current informatics advancements in studying behavior disorders. In this study, we proposed two major contributions. First, we introduce a data-driven approach that takes advantage of the available public databases to differentiate between multiple behavioral disorders. Second, we developed an index (Behavioral Health Score (BHS)) to assess the severity level. The proposed approach is based on the established correlations between mobility and health by analyzing the mobility data collected from 72 participants suffering from multiple behavioral disorders using wearable sensors. Obtained results demonstrate that the proposed correlation network model can distinguish between different types of disorders and BHS index can be utilized to estimate the severity levels of the disorder. Moreover, the proposed model can add a useful information to healthcare providers to better diagnose behavioral conditions and develop data-driven treatment options. Rama Krishna Thelagathoti, Hesham Ali 0001 |
BIBM | 2 |
| 2021 | Parallel Planar Approximation for Large NetworksabstractPlanarity is an extensively studied topic within the graph theory domain. Planarity describes the potential for a graph to be embedded on a Euclidean plane without having any of its edges cross, and planarity is consequently a subject of particular interest in the context of graphs with nodes representing entities that are subject to physical constraints. While graph planarity has been extensively studied as a theoretical topic, it has also been successfully practically applied to architectural design [1] and circuit board design [2]. Our prior work found planarity to be a topic of interest in biological networks, specifically with regard to protein-protein interaction and domain-domain interaction networks. This work presents a novel parallel algorithm that produces planar approximations of large networks, which may be used to assess the relative planarity of large networks, for further analyses or ensemble methods, and visualization. William Gasper, Kathryn M. Cooper, Nathan Cornelius, Hesham Ali 0001 |
BIBM | 4 |
| 2020 | A Dynamic Approach for Managing Heterogeneous Wireless Networks in Smart Environments
Yara Mahfood Haddad, Hesham Ali 0001 |
ISCC | 2 |
| 2019 | Identifying Structural Changes in Correlation Networks Models of Cancer Gene Expression by StageabstractGene expression analysis using correlation network modeling can help to identify systems-level cellular changes and cooperation among genes. Network modeling is a relatively novel method for comparing changes across stages of cancer, or between primary tumor tissue and the metastatic tissue. In this study, we develop a pipeline to identify the dynamic changes of cancer in gene expression level through time-dependent network analysis of cancer data from The Cancer Genome Atlas (TCGA). A total of 16 correlation networks were built from four (4) stages in four (4) different types of cancers: Thyroid Carcinoma, Colon Adenocarcinoma and Rectum Adenocarcinoma, Stomach Adenocarcinoma, and Kidney Renal Clear Cell Carcinoma. To identify the basic changes in network structure, we performed Jaccard similarity comparison of structurally relevant nodes. We employed mutation analysis to measure and present the time-based changes in mutation rate of genes that are specific to each cancer type. Finally, we present a case study to identify the gene expression changes among primary tumor tissue and metastatic tissue in skin cutaneous melanoma (SKCM). Qianran Li, Dario Ghersi, Ishwor Thapa, Hesham Ali 0001, Kathryn M. Cooper |
BIBM | 5 |
| 2019 | Uncovering and characterizing splice variants associated with survival in lung cancer patientsabstractSplice variants have been shown to play an important role in tumor initiation and progression and can serve as novel cancer biomarkers. However, the clinical importance of individual splice variants and the mechanisms by which they can perturb cellular functions are still poorly understood. To address these issues, we developed an efficient and robust computational method to: (1) identify splice variants that are associated with patient survival in a statistically significant manner; and (2) predict rewired protein-protein interactions that may result from altered patterns of expression of such variants. We applied our method to the lung adenocarcinoma dataset from TCGA and identified splice variants that are significantly associated with patient survival and can alter protein-protein interactions. Among these variants, several are implicated in DNA repair through homologous recombination. To computationally validate our findings, we characterized the mutational signatures in patients, grouped by low and high expression of a splice variant associated with patient survival and involved in DNA repair. The results of the mutational signature analysis are in agreement with the molecular mechanism suggested by our method. To the best of our knowledge, this is the first attempt to build a computational approach to systematically identify splice variants associated with patient survival that can also generate experimentally testable, mechanistic hypotheses. Code for identifying survival-significant splice variants using the Null Empirically Estimated P-value method can be found at https://github.com/thecodingdoc/neep. Code for construction of Multi-Granularity Graphs to discover potential rewired protein interactions can be found at https://github.com/scwest/SINBAD. Sean West, Surinder K. Batra, Hesham Ali 0001, Dario Ghersi |
PLoS Comput. Biol. | 4 |
| 2018 | A Graph-Theoretic Approach for Identifying Bacterial Inter-correlations and Functional Pathways in Microbiome Data
Su Yeon Kim, Ishwor Thapa, Hesham Ali 0001 |
BIBM | 3 |
| 2017 | A network-based approach to mine temporal genes exhibiting significant expression variation in Caenorhabditis elegansabstractIt is critical to be able to identify longitudinally changing genes in temporal data so that studies can be focused on how gene expression changes in a dynamic way. While biological networks continue to play a significant role in modeling and characterizing complex relationships in biological systems, most network modeling studies in biomedical research focus on snapshot or “static” network-based analysis to identify genes of interest. In this study, we use a temporal non-sampling network-based approach to identify and rank genes that exhibit significant co-expression variation over time. We use in the C. elegans gene correlation network obtained from mRNA expression profiles to illustrate the value of the proposed approach. We compare the results of this method to results obtained from traditional statistical analysis that focuses on identifying simple differentially expressed genes. We show that rank-based temporal network analysis can identify genes that contribute to changes in the network structure and consequently contribute to changes in the genetic regulatory machine. Kathryn M. Cooper, Wail Hassan, Hesham Ali 0001 |
BIBM | 3 |
| 2017 | On the integration of assembly and non-assembly approaches for comparing biological sequencesabstractAs Next Generation Sequencing (NGS) technologies continue to expand rapidly, the need to assemble and manipulate NGS data, available in the form of short genomic reads, remains the primary source of biological data in many Bioinformatics applications. As a result, many assemblers have been developed to assemble NSG short reads into long genomic sequences or contigs ready for advanced analysis such as Whole Genome Wide Studies (GWAS). However, the lack of high levels of robustness and reproducibility continue to limit the impact of Bioinformatics research and many biomedical researchers remain skeptical of results obtained from bioinformatics applications. In this study, we conduct a comparative study of various widely used assemblers and compare their performances using several NGS datasets associated with various organisms. We highlight the advantages and disadvantage of each assembler and explore the factors that impact the performance of each approach. In addition, we survey the assembly-free compression approach recently developed to process NGS short reads to analyze their performance in comparing genomic sequences represented by sets of short reads. We use phylogeny trees obtained from simulated and real datasets to evaluate the accuracy of each assembly-free approach. We test the hypothesis that non-assembly approaches could potentially overcome the limitations and inaccuracies of assembly approaches in comparing sequences, especially for large read sizes. Moreover, we proposed a hybrid approach by integrating both assembly and non-assembly approach for classifying genomic sequences. The proposed approach incorporates results obtained from partially assembling short reads as input for assembly-free methods to complete the NGS manipulation process. Preliminary superior results show that the hybrid approach is potential in comparing genomic sequences. Vi Dam, Hesham Ali 0001 |
BIBM | 2 |
| 2017 | A systems biology approach for modeling microbiomes using split graphsabstractWith the recent advances in sequencing technology, researchers now have opportunities to study microbiomes associated with various environments. Recent studies have shown that the composition of microbiomes in our bodies and our environments play a significant role in our health. For example, 90% of human DNA is composed of bacterial microbiomes. In this study, we propose a systems biology approach using split graphs to analyze the composition of microbiomes and the impact of such composition on the health and growth of organisms living in associated environments. We focus on a case study related to the composition of microbiomes in fish guts and its impact on various growth parameters for three types of fish. The proposed model explores features in the aquatic ecosystem including correlations among its microorganisms and their abundance levels. The results of the study show that single or groups of bacteria are significantly associated with multiple growth phenotypes in different gut portions of the fish. We also identify bacterial clusters that provide new insight to functional relevance of these bacteria and their contribution to the fish gut microbial ecosystem. Su Yeon Kim, Ishwor Thapa, Guoqing Lu, Lifeng Zhu, Hesham Ali 0001 |
BIBM | 5 |
| 2017 | Using gait parameters to recognize various stages of Parkinson's diseaseabstractMonitoring gait patterns seamlessly and continuously over time provides valuable information that could help physicians diagnose diseases in the early stages. Currently, traditional gait measurement approaches do not support continuous monitoring of gait and focus on collecting limited data points in controlled lab environments. However, with advancements in wireless technology, movement patterns can be recorded using small portable wearable devices. Parkinson's disease (PD) is a progressively disabling neurodegenerative disorder that is affecting gait and posture and consequently leads to higher risk of falling. Several research studies have looked into changes in the gait parameters of PD patients compared to healthy adults. However, there are only few studies with the focus on gait assessment of PD patients in the early stages as compared to patterns associated with patients at advanced stages. In addition, the number of gait-related studies in this domain using accelerometers on ankle is very limited. Knowing which body location could serve as a target place for accelerometers to provide accurate information is a necessary step toward the health assessment of PD patients. The purpose of this study was to evaluate the gait parameters of patients with mild or moderate PD using accelerometers on ankles. A number of gait parameters, including average stride time, stride time variability, stride time symmetry, and oscillation of acceleration in the mediolateral (ML) direction were calculated and compared between PD patients and healthy elderlies. Preliminary results indicate that features extracted from accelerometers on ankles can be effective in differentiating between healthy elderlies and PD patients at mid-stages of disease but less so at earlier stages of disease. Elham Rastegari, Vivien Marmelat, Lotfollah Najjar, Dhundy Bastola, Hesham Ali 0001 |
BIBM | 5 |
| 2017 | Analysis of clustering algorithms in biological networksabstractBiological data is often represented as networks, as in the case of protein-protein interactions and metabolic pathways. Modeling, analyzing, and visualizing networks can help make sense of large volumes of data generated by high-throughput experiments. However, due to their size and complex structure, biological networks can be difficult to interpret without further processing. Cluster analysis is a widely-used approach to extract meaningful information from biological networks. In this work, we provide a study that surveys some of the widely used clustering algorithms used for clustering biological data. We identify the advantages and disadvantages of each algorithm and attempt to identify features associated with datasets that align well with each approach. We also propose a new clustering method based on graph matching and node merging techniques in an attempt to fill the gap left by the current clustering approaches. Asuda Sharma, Hesham Ali 0001 |
BIBM | 2 |
| 2017 | Evaluation of the oral microbiome as a biomarker for early detection of human oral carcinomasabstractWith the advent of precision medicine, biomarkers have recently come into focus as a promising tool for early cancer detection and treatment individualization. In particular, much interest has been shown in the oral microbiome as a promising potential cancer biomarker, especially for head and neck cancers. The American Cancer Society estimates that there will be nearly 50,000 new cases and roughly 10,000 deaths from oral or oropharyngeal cancer in 2017. The five-year survival rate for cancers of the oral cavity and pharynx is 66% for Caucasian individuals and 47% for African American individuals. However, when caught early while the cancer is at a local stage, the 5-year survival rates rise to 83% and 79%, respectively. The oral carcinoma cancer is typically diagnosed by an oral health care provider by visual screening. However, many of these cancers are discovered when they have already progressed to a later stage. The goal of this research is to evaluate oral microbiome based biomarkers for early oral carcinoma detection. The outcome of this research is a machine-learning based framework for microbiome-based early cancer detection. The ability to identify at risk patients using minimally invasive biomarkers will allow for more rapid treatment plan development and improved outcome. The early diagnosis and treatment of cancer is essential for increasing patient survival odds and mitigating patient suffering. Julia Warnke-Sommer, Hesham Ali 0001 |
BIBM | 2 |
| 2017 | Granularity-aware fusion of biological networks for information extractionabstractCurrent biomedical research approaches capture multiple data sources, fusing and integrating them to extract accurate information. While the purpose of fusion in the biomedical domain is to create a more holistic picture of biological reality, latent sensitivities of the fusion process often result in information loss. It has been shown that biomedical data fusion is sensitive to granularity dimensions, scales of semantic relationships across specificity and mereology. Low granularity data casts a wide net, while high granularity data has focus. When data fusion occurs between low and high granularities, these benefits counteract each other. In this study, we reexamine the union function as a basis for biomedical data fusions, via comparison to a granularity-aware function. This granularity-aware approach uses domain knowledge (low granularity) as a filter for expression relationships (high granularity). We use pancreatic cancer expression data and domain knowledge networks to test the fusion approaches. We support previous findings that granularity-unaware fusion allows domain knowledge to eclipse condition-specific data. In addition, we find that the granularity-aware approach tends to outperform both the union and non-fusion networks, resulting in higher information extraction scores. Further, the granularity-aware approach increases the network information extraction effect size between disease and normal networks, allowing for a more distinctive delineation between the two conditions. Sean West, Hesham Ali 0001 |
BIBM | 2 |
| 2015 | On the comparison of state- and transition-based analysis of biological relevance in gene co-expression networksabstractTraditional correlation network analysis typically involves creating a network using gene expression data and then identifying biologically relevant clusters from that network by enrichment with Gene Ontology or pathway information. When one wants to examine these networks in a dynamic way - such as between controls versus treatment or over time - a “snapshot” approach is taken by comparing network structures at each time point. The biological relevance of these structures are then reported and compared. In this research, we examine the same “snapshot” networks but focus on the enrichment of changes in structure to determine if these results give any more insight into the mechanisms behind observed phenotypes. Our main hypothesis is that more information, particularly related to potential dynamic changes, can be obtained through transition-based analysis of biological networks. To test this hypothesis, we compare gene expression data from the mouse hippocampus at three different time points: young, middle-aged, and aged, and compare the traditional state-based approach to the dynamic transition-based enrichment approach. In this study we use a clustering approach (SPICi) designed specifically for clustering of large biological networks. The results of this study verify an inconsistency between traditional and dynamic structure identification approaches through biological enrichment. These results highlight an intriguing issue for those performing, critiquing, and using network based approaches in their research - that a black box or workflow type of approach typically used in network based research can be supplemented with a transitionbased approach to support movement from in silico to in vivo experimentation of target genes. Kathryn M. Cooper, Prasuna Vemuri, Hesham Ali 0001 |
BIBM | 3 |
| 2015 | Next generation sequence assembler mis-assembly of phage genomes with terminal redundancyabstractNext generation sequencing (NGS) has become the platform of numerous biomedical applications. The study of viral genomes using NGS technologies has led to the characterization of viral species in numerous environments including the human gut microbiome and plant hosts. Many viral genomes are circular or have terminally redundant ends. Circular or linear viral genomes with indeterminate starting and ending points pose a challenge for NGS assemblers, which may erroneously duplicate sections of these genomes. The length of an assembly, often characterized by the N50 length, is frequently used as an indication of an assembly's completeness and even quality. In this paper, we show that the longest contig produced by various assemblers is not always the best assembly for circular or terminally redundant phage genomes and may represent erroneously repeated genomic regions. Results demonstrate that assembly tools may even produce assembled genomes of different lengths for the same species, depending on content inaccurately repeated, leading to results that might be confusing to or inaccurately used by a researcher. To overcome this problem, we introduce strategies for using coverage depth to identify inaccurately repeated content in circular or terminally redundant phage genomes. We conclude the paper by providing the results of assembling two bacteriophage genomes and a bacteriophage metagenomics dataset, highlighting the impact of using the proposed strategies. Julia Warnke-Sommer, Ishwor Thapa, Hesham Ali 0001 |
BIBM | 3 |
| 2013 | A base composition analysis of natural patterns for the preprocessing of metagenome sequencesabstractBACKGROUND: On the pretext that sequence reads and contigs often exhibit the same kinds of base usage that is also observed in the sequences from which they are derived, we offer a base composition analysis tool. Our tool uses these natural patterns to determine relatedness across sequence data. We introduce spectrum sets (sets of motifs) which are permutations of bacterial restriction sites and the base composition analysis framework to measure their proportional content in sequence data. We suggest that this framework will increase the efficiency during the pre-processing stages of metagenome sequencing and assembly projects. RESULTS: Our method is able to differentiate organisms and their reads or contigs. The framework shows how to successfully determine the relatedness between these reads or contigs by comparison of base composition. In particular, we show that two types of organismal-sequence data are fundamentally different by analyzing their spectrum set motif proportions (coverage). By the application of one of the four possible spectrum sets, encompassing all known restriction sites, we provide the evidence to claim that each set has a different ability to differentiate sequence data. Furthermore, we show that the spectrum set selection having relevance to one organism, but not to the others of the data set, will greatly improve performance of sequence differentiation even if the fragment size of the read, contig or sequence is not lengthy. CONCLUSIONS: We show the proof of concept of our method by its application to ten trials of two or three freshly selected sequence fragments (reads and contigs) for each experiment across the six organisms of our set. Here we describe a novel and computationally effective pre-processing step for metagenome sequencing and assembly tasks. Furthermore, our base composition method has applications in phylogeny where it can be used to infer evolutionary distances between organisms based on the notion that related organisms often have much conserved code. Oliver Bonham-Carter, Hesham Ali 0001, Dhundy Bastola |
BMC Bioinform. | 2 |
| 2013 | An efficient and scalable graph modeling approach for capturing information at different levels in next generation sequencing readsabstractBACKGROUND: Next generation sequencing technologies have greatly advanced many research areas of the biomedical sciences through their capability to generate massive amounts of genetic information at unprecedented rates. The advent of next generation sequencing has led to the development of numerous computational tools to analyze and assemble the millions to billions of short sequencing reads produced by these technologies. While these tools filled an important gap, current approaches for storing, processing, and analyzing short read datasets generally have remained simple and lack the complexity needed to efficiently model the produced reads and assemble them correctly. RESULTS: Previously, we presented an overlap graph coarsening scheme for modeling read overlap relationships on multiple levels. Most current read assembly and analysis approaches use a single graph or set of clusters to represent the relationships among a read dataset. Instead, we use a series of graphs to represent the reads and their overlap relationships across a spectrum of information granularity. At each information level our algorithm is capable of generating clusters of reads from the reduced graph, forming an integrated graph modeling and clustering approach for read analysis and assembly. Previously we applied our algorithm to simulated and real 454 datasets to assess its ability to efficiently model and cluster next generation sequencing data. In this paper we extend our algorithm to large simulated and real Illumina datasets to demonstrate that our algorithm is practical for both sequencing technologies. CONCLUSIONS: Our overlap graph theoretic algorithm is able to model next generation sequencing reads at various levels of granularity through the process of graph coarsening. Additionally, our model allows for efficient representation of the read overlap relationships, is scalable for large datasets, and is practical for both Illumina and 454 sequencing technologies. Julia Warnke-Sommer, Hesham Ali 0001 |
BMC Bioinform. | 2 |
| 2012 | On the design of advanced filters for biological networks using graph theoretic propertiesabstractNetwork modeling of biological systems is a powerful tool for analysis of high-throughput datasets by computational systems biologists. Integration of networks to form a heterogeneous model requires that each network be as noise-free as possible while still containing relevant biological information. In earlier work, we have shown that the graph theoretic properties of gene correlation networks can be used to highlight and maintain important structures such as high degree nodes, clusters, and critical links between sparse network branches while reducing noise. In this paper, we propose the design of advanced network filters using structurally related graph theoretic properties. While spanning trees and chordal subgraphs provide filters with special advantages, we hypothesize that a hybrid subgraph sampling method will allow for the design of a more effective filter preserving key properties in biological networks. That the proposed approach allows us to optimize a number of parameters associated with the filtering process which in turn improves upon the identification of essential genes in mouse aging networks. Kathryn Dempsey, Tzu-Yi Chen, Sanjukta Bhowmick, Hesham Ali 0001 |
BIBM | 4 |
| 2012 | A Novel Multithreaded Algorithm for Extracting Maximal Chordal SubgraphsabstractChordal graphs are triangulated graphs where any cycle larger than three is bisected by a chord. Many combinatorial optimization problems such as computing the size of the maximum clique and the chromatic number are NP-hard on general graphs but have polynomial time solutions on chordal graphs. In this paper, we present a novel multithreaded algorithm to extract a maximal chordal sub graph from a general graph. We develop an iterative approach where each thread can asynchronously update a subset of edges that are dynamically assigned to it per iteration and implement our algorithm on two different multithreaded architectures - Cray XMT, a massively multithreaded platform, and AMD Magny-Cours, a shared memory multicore platform. In addition to the proof of correctness, we present the performance of our algorithm using a test set of synthetical graphs with up to half-a-billion edges and real world networks from gene correlation studies and demonstrate that our algorithm achieves high scalability for all inputs on both types of architectures. Mahantesh Halappanavar, John Feo, Kathryn Dempsey, Hesham Ali 0001, Sanjukta Bhowmick |
ICPP | 4 |
| 2009 | Community wireless networks: Emerging wireless commons for digital inclusionabstractCommunity wireless networks (CWNs) have emerged as collective actions where communities develop common telecommunication infrastructure for their use. These collective projects have gained widespread popularity as they serve as digital highways for community members to join the information society. This paper explores the role of CWNs in achieving digital inclusion. In particular, we used a survey-based instrument to measure the size and the capacity of a number of networks. We also explored other related variables such as service pricing, funding instruments and the life duration of the investigated CWNs to distinguish them from similar networks. This study provides useful theoretical accounts and analytical insights that we believe that can guide future studies and advance CWNs as a form of common projects. Also, it has the potential to help policy makers and community developers in promoting such collective projects. Abdelnasser Abdelaal, Hesham Ali 0001 |
ISTAS | 2 |
| 2009 | RFID-based information system for preventing medical errorsabstractA report by the Institute of Medicine of the National Academy of Sciences estimates that as many as 98,000 people die in U.S. hospitals each year because of medical errors. In this project, we propose an innovative IT-based approach to prevent errors in various medical processes by utilizing advance Jong-Hoon Youn, Hesham Ali 0001, Hamid Sharif, Biplav Chhetri |
MobiQuitous | 2 |
| 2008 | A parallel architecture for regulatory motif algorithm assessmentabstractComputational discovery of cis-regulatory motifs has become one of the more challenging problems in bioinformatics. In recent years, over 150 methods have been proposed as solutions, however, it remains difficult to characterize the advantages and disadvantages of these approaches because of the wide variability of approaches and datasets. Although biologists desire a set of parameters and a program most appropriate for cis-regulatory discovery in their domain of interest, compiling such a list is a great computational challenge. First, a discovery pipeline for 150+ methods must be automated and then each dataset of interest must used to grade the methods. Automation is challenging because these programs are intended to be used over a small set of sites and consequently have many manual steps intended to help the user in fine-tuning the program to specific problems or organisms. If a program is fine-tuned to parameters other than those used in the original paper, it is not guaranteed to have the same sensitivity and specificity. Consequently, there are few methods that rank motif discovery tools. This paper proposes a parallel framework for the automation and evaluation of cis-regulatory motif discovery tools. This evaluation platform can both run and benchmark motif discovery tools over a wide range of parameters and is the first method to consider both multiple binding locations within a regulatory region and regulatory regions of orthologous genes. Because of the large amount of tests required, we implemented this platform on a computing cluster to increase performance. Daniel Quest, Kathryn Dempsey, Mohammad Shafiullah, Dhundy Bastola, Hesham Ali 0001 |
IPDPS | 5 |
| 2008 | A Hidden Markov Model Approach for Prediction of Genomic Alterations from Gene Expression Profiling
Huimin Geng, Hesham Ali 0001, Wing C. Chan |
ISBRA | 2 |
| 2008 | MTAP: The Motif Tool Assessment PlatformabstractBACKGROUND: In recent years, substantial effort has been applied to de novo regulatory motif discovery. At this time, more than 150 software tools exist to detect regulatory binding sites given a set of genomic sequences. As the number of software packages increases, it becomes more important to identify the tools with the best performance characteristics for specific problem domains. Identifying the correct tool is difficult because of the great variability in motif detection software. Consequently, many labs spend considerable effort testing methods to find one that works well in their problem of interest. RESULTS: In this work, we propose a method (MTAP) that substantially reduces the effort required to assess de novo regulatory motif discovery software. MTAP differs from previous attempts at regulatory motif assessment in that it automates motif discovery tool pipelines (something that traditionally required many manual steps), automatically constructs orthologous upstream sequences, and provides automated benchmarks for many popular tools. As a proof of concept, we have run benchmarks over human, mouse, fly, yeast, E. coli and B. subtilis. CONCLUSION: MTAP presents a new approach to the challenging problem of assessing regulatory motif discovery methods. The most current version of MTAP can be downloaded from http://biobase.ist.unomaha.edu/ Daniel Quest, Kathryn Dempsey, Mohammad Shafiullah, Dhundy Bastola, Hesham Ali 0001 |
BMC Bioinform. | 5 |
| 2007 | Reducing Folding Scenario Candidates in Pseudoknots Detection Using PLMM_DPSS Algorithm Integrated With Energy FiltersabstractPseudoknots are important functional structures in RNA. Despite of numerous pseudoknot prediction tools, biologists still need a pseudoknot detection method with much higher sensitivity. We have previously developed the pseudoknot local motif model and dynamic partner sequence stacking (PLMMDPSS) algorithm which predicts short loop 2 pseudoknots with high sensitivity. Integrated with Mfold, PLMM_DPSS is capable of predicting both H-type and complicated type pseudoknots. In this study, we have developed PLMM_DPSS_SF Mfold_FF with the following modifications: the extension of the PLM model to include long loop 2 pseudoknots, the incorporation of overall folding energy calculation, recombination of non-overlapping pseudoknots within one sequence and two filters for higher specificity, The prediction results have shown that PLMM_DPSS_SF_Mfold_FF is more sensitive than other leading pseudoknot prediction tools and results in a low alternative folding number, and results also support the notion that some kinetic barriers may trap RNA folding in local minimums. Xiaolu Huang, Hesham Ali 0001 |
BIBE | 2 |
| 2007 | A Robust Scalable Cluster-Based Multi-hop Routing Protocol for Wireless Sensor Networks
Sudha Mudundi, Hesham Ali 0001 |
ISPA | 2 |
| 2007 | Applications and Performances of Extended TTDDs in Large-Scale Wireless Sensor Networks
Hong Zhou 0005, Lu Jin 0005, Hesham Ali 0001, Chulho Won |
MSN | 4 |
| 2007 | WLAN-Based Real-Time Asset Tracking System in Healthcare Environments
Jong-Hoon Youn, Hesham Ali 0001, Hamid Sharif, Jitender S. Deogun, Jason Uher, Steven H. Hinrichs |
WiMob | 2 |
| 2006 | On Effective Utilization of Wireless Networks in Collaborative ApplicationsabstractCollaborative applications allow a group of users to work together by sharing information and processes. Traditionally, these applications run in a lab environment with the support of traditional communication networks. Employing wireless technologies in supporting such applications will increase its effectiveness and usability. Wireless networks are rapidly emerging to be the network architecture of choice for many application domains due to their adaptability and scalability. However, skepticism about the level of quality of service (QoS) guaranteed by wireless networks has prevented them from replacing traditional networks in time-critical applications. Early attempts to naively replace traditional networks with wireless networks in supporting collaborative applications have produced less than acceptable results. In this paper, we show that the performance of various group decision support systems (GDSS) modules highly depends on the architecture of the employed wireless network and that selecting a proper configuration leads to significant performance improvements comparable to those obtained in wired networks Surekha Kanapuram, Hesham Ali 0001, Gert-Jan de Vreede |
CollaborateCom | 2 |
| 2005 | On Clustering Biological Data Using Unsupervised and Semi-Supervised Message PassingabstractNoticing that unsupervised clustering may produce clusters that are irrelevant to the research hypotheses and interests, we generalize traditional unsupervised clustering into semi-supervised clustering based on our previously proposed message passing clustering (MPC). In the semi-supervised MPC, prior knowledge such as instance-level and attribute-level constraints are used to guide the clustering process towards better and interpretable partitions. We applied the unsupervised MPC (null background) to phylogenetic analysis of Mycobacterium and the semi-supervised MPC to colon cancer microarray data analysis. The results show that MPC is superior to the widely accepted neighbor-joining and hierarchical clustering methods, and the semi-supervised MPC is even more powerful in biological data analysis such as gene selection and cancer diagnosis using microarray. Huimin Geng, Xutao Deng, Dhundy Bastola, Hesham Ali 0001 |
BIBE | 4 |
| 2005 | Protein Motif Searching Through Similar Enriched Parikh Vector IdentificationabstractBiological researches have shown that some protein regions sharing similar functions or structures have inversed-ordered or highly dispersed sequence similarities and some intra-sequence similarities such as palindrome repeats also play important roles in protein folding. The current protein analysis tools cannot detect these "nontraditional" similarities. Although some tools can be modified for searching intra-sequence inversed or forward ordered similarities, their maximally optimal path processes will miss many suboptimal biologically meaningful similarities. The similar enriched Parikh vector searching (SRPVS) algorithm searches similarities by separating the subsequence composition and order information. The SRPVS first breaks sequences into groups of predefined-sized subsequences, each represented by an enriched Parikh vector (RPV); then similar RPV pairs (SRPV) are searched in each nonoverlapping RPV pair based on various order restrictions - forward, inversed, or shuffled. In this study, SRPVS has been applied to the protein ligand motif finding and the intra-sequence protein inversed repeats finding. Xiaolu Huang, Hesham Ali 0001, Anguraj Sadanandam, Rakesh K. Singh |
BIBE | 2 |
| 2005 | A new clustering algorithm using message passing and its applications in analyzing microarray dataabstractIn this paper, we proposed a new clustering algorithm that employs the concept of message passing to describe parallel and spontaneous biological processes. Inspired by real-life situations in which people in large gatherings form groups by exchanging messages, message passing clustering (MPC) allows data objects to communicate with each other and produces clusters in parallel, thereby making the clustering process intrinsic and improving the clustering performance. We have proved that MPC shares similarity with hierarchical clustering but offers significantly improved performance because it takes into account both local and global structure. MPC can be easily implemented in a parallel computing platform for the purpose of speed-up. To validate the MPC method, we applied MPC to microarray data from the Stanford yeast cell-cycle database. The results show that MPC gave better clustering solutions in terms of homogeneity and separation values than other clustering methods. Huimin Geng, Xutao Deng, Hesham Ali 0001 |
ICMLA | 3 |
| 2005 | HYDRA: A New Approach for Integrating Various Wireless EnvironmentsabstractIn this paper we present the design of Hydra, an experimental Linux platform for integrating various wireless environments. Currently, mobile devices are often equipped with many network interfaces, which may be of different access technologies, like wireless, cellular and wired. Applications have different requirements, which results in different network preferences. Also, the network preferences of some applications change over time and they would like to use multiple access technologies to satisfy their needs best. Hydra is an attempt to provide users control such that they may manage their own and available devices in a more flexible way than the existing networks are offering. With Hydra, there is a way for applications to provide the operating system, features of the environment they are interested in. Also, there is a mechanism that enables applications to track their environment. The biggest feature is the ability to integrate multiple wireless technologies. Karthik Ramachandra 0001, Hesham Ali 0001 |
ISCC | 2 |
| 2005 | A method of precise mRNA/DNA homology-based gene structure predictionabstractBACKGROUND: Accurate and automatic gene finding and structural prediction is a common problem in bioinformatics, and applications need to be capable of handling non-canonical splice sites, micro-exons and partial gene structure predictions that span across several genomic clones. RESULTS: We present a mRNA/DNA homology based gene structure prediction tool, GIGOgene. We use a new affine gap penalty splice-enhanced global alignment algorithm running in linear memory for a high quality annotation of splice sites. Our tool includes a novel algorithm to assemble partial gene structure predictions using interval graphs. GIGOgene exhibited a sensitivity of 99.08% and a specificity of 99.98% on the Genie learning set, and demonstrated a higher quality of gene structural prediction when compared to Sim4, est2genome, Spidey, Galahad and BLAT, including when genes contained micro-exons and non-canonical splice sites. GIGOgene showed an acceptable loss of prediction quality when confronted with a noisy Genie learning set simulating ESTs. CONCLUSION: GIGOgene shows a higher quality of gene structure prediction for mRNA/DNA spliced alignment when compared to other available tools. Alexander G. Churbanov, Mark A. Pauley, Daniel Quest, Hesham Ali 0001 |
BMC Bioinform. | 4 |
| 2004 | A New Fast Fault Tolerant Scheduling Approach in Distributed Systems
Mohana Desiraju, Hesham Ali 0001 |
CAINE | 2 |
| 2004 | Adaptive Data Forwarding Techniques for Energy Conservation in Wireless Sensor Networks
Kavitha Gundappachikkenahalli, Hesham Ali 0001 |
CAINE | 2 |
| 2004 | A Common Interval Searching Algorithm for Protein Sequence Comparison
Xiaolu Huang, Hesham Ali 0001 |
CAINE | 2 |
| 2004 | Identification of Mycobacterium Species Using Curated Custom DatabasesabstractSummary form only given. Advances in molecular biology have resulted in the development of diagnostic tests for infectious diseases based on genetic profiles. While probe based assays dominate the field today, sequence based assays hold great promise for the future. However, the variability in quality of sequence information currently present in public databases limits the potential growth and use of sequence based analysis. To address this problem a standardized method for DNA sequence validation and building of custom databases was developed using mycobacterium as a development model. With this model, a computational approach to identification of infectious diseases was developed and evaluated. The Web-based application, termed BioDatabase, accomplished genetic sequence identification via the creation of curated databases containing a relatively small set of genetic data specific to a species or group. The process for creation of the custom database included multiple steps beginning with identification of highly conserved start and end sequences and intervening sequence validation parameters. The process eliminated the need for multiple sequence alignment with GenBank sequences, whose information is valuable, yet difficult to properly utilize due to its size and quality. The custom database approach maximized application performance with minimal impact on analysis response time, allowing investigation of optimal sequences for identification of all mycobacterium to the species level. In comparison to the 16S and ITS genetic regions, a curated ITS based approach proved most effective for identification of mycobacterium isolates. Dan Kuyper, Hesham Ali 0001, Amr M. Mohamed, Steven H. Hinrichs |
IPDPS | 2 |
| 1995 | An Optimal Algorithm for Scheduling Interval Ordered Tasks with Communication on N Processors
Hesham Ali 0001, Hesham El-Rewini |
J. Comput. Syst. Sci. | 1 |
| 1995 | Static Scheduling of Conditional Branches in Parallel Programs
Hesham El-Rewini, Hesham Ali 0001 |
J. Parallel Distributed Comput. | 2 |