EDBT 2026 Demo / reviewers in the wild / expert
Kamal Taha
dblp:51/4159
· DBLP profile ↗
46ranked-venue papers
35as first author
10since 2021 · last 2026
0000-0002-6674-4614ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 17 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10 · 10 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 2 since 2021Security and privacy · 6 · 4 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TriGAN-SiaMT: A triple-segmentor adversarial network with bounding box priors for semi-supervised brain lesion segmentation
Mohammad Alshurbaji, Maregu Assefa, Ahmad Obeid 0001, Mohamed L. Seghier, Taimur Hassan, Kamal Taha, Naoufel Werghi |
Pattern Recognit. Lett. | 6 |
| 2025 | Empirical and Experimental Insights into Data Mining Techniques for Crime Prediction: A Comprehensive SurveyabstractThis survey article presents a comprehensive analysis of crime prediction methodologies, exploring the various techniques and technologies utilized in this area. The article covers the statistical methods, machine learning algorithms, and deep learning techniques employed to analyze crime data, while also examining their effectiveness and limitations. We propose a methodological taxonomy that classifies crime prediction algorithms into specific techniques. This taxonomy is structured into four tiers, including methodology category, methodology sub-category, methodology techniques, and methodology sub-techniques. Empirical and experimental evaluations are provided to rank the different techniques. The empirical evaluation assesses the crime prediction techniques based on three criteria, while the experimental evaluation ranks the algorithms that employ the same sub-technique, the different sub-techniques that employ the same technique, the different techniques that employ the same methodology sub-category, the different methodology sub-categories within the same category, and the different methodology categories. The combination of methodological taxonomy, empirical evaluations, and experimental comparisons allows for a nuanced and comprehensive understanding of crime prediction algorithms, aiding researchers in making informed decisions. Finally, the article provides a glimpse into the future of crime prediction techniques, highlighting potential advancements and opportunities for further research in this field. Kamal Taha |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2024 | Learning a deep-feature clustering model for gait-based individual identificationabstractGait biometrics which concern with recognizing individuals by the way they walk are of a paramount importance these days. Human gait is a candidate pathway for such identification tasks since other mechanisms can be concealed. Most common methodologies rely on analyzing 2D/3D images captured by surveillance cameras. Thus, the performance of such methods depends heavily on the quality of the images and the appearance variations of individuals. In this study, we describe how gait biometrics could be used in individuals' identification using a deep feature learning and inertial measurement unit (IMU) technology. We propose a model that recognizes the biological and physical characteristics of individuals, such as gender, age, height, and weight, by examining high-level representations constructed during its learning process. The effectiveness of the proposed model has been demonstrated by a set of experiments with a new gait dataset generated using a shoe-type based on a gait analysis sensor system. The experimental results show that the proposed model can achieve better identification accuracy than existing models, while also demonstrating more stable predictive performance across different classes. This makes the proposed model a promising alternative to current image-based modeling. Kamal Taha, Paul D. Yoo, Yousof Al-Hammadi, Sami Muhaidat, Chan Yeob Yeun |
Comput. Secur. | 1 |
| 2024 | Analysis of Cancer-Associated Mutations of POLB Using Machine Learning and BioinformaticsabstractDNA damage is a critical factor in the onset and progression of cancer. When DNA is damaged, the number of genetic mutations increases, making it necessary to activate DNA repair mechanisms. A crucial factor in the base excision repair process, which helps maintain the stability of the genome, is an enzyme called DNA polymerase β (Pol β) encoded by the POLB gene. It plays a vital role in the repair of damaged DNA. Additionally, variations known as Single Nucleotide Polymorphisms (SNPs) in the POLB gene can potentially affect the ability to repair DNA. This study uses bioinformatics tools that extract important features from SNPs to construct a feature matrix, which is then used in combination with machine learning algorithms to predict the likelihood of developing cancer associated with a specific mutation. Eight different machine learning algorithms were used to investigate the relationship between POLB gene variations and their potential role in cancer onset. This study not only highlights the complex link between POLB gene SNPs and cancer, but also underscores the effectiveness of machine learning approaches in genomic studies, paving the way for advanced predictive models in genetic and cancer research. Razan Alkhanbouli, Amira Al-Aamri, Maher Maalouf, Kamal Taha, Andreas Henschel, Dirar Homouz |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Employing Machine Learning Techniques to Detect Protein Function: A Survey, Experimental, and Empirical EvaluationsabstractThis review article delves deeply into the various machine learning (ML) methods and algorithms employed in discerning protein functions. Each method discussed is assessed for its efficacy, limitations, potential improvements, and future prospects. We present an innovative hierarchical classification system that arranges algorithms into intricate categories and unique techniques. This taxonomy is based on a tri-level hierarchy, starting with the methodology category and narrowing down to specific techniques. Such a framework allows for a structured and comprehensive classification of algorithms, assisting researchers in understanding the interrelationships among diverse algorithms and techniques. The study incorporates both empirical and experimental evaluations to differentiate between the techniques. The empirical evaluation ranks the techniques based on four criteria. The experimental assessments rank: (1) individual techniques under the same methodology sub-category, (2) different sub-categories within the same category, and (3) the broad categories themselves. Integrating the innovative methodological classification, empirical findings, and experimental assessments, the article offers a well-rounded understanding of ML strategies in protein function identification. The paper also explores techniques for multi-task and multi-label detection of protein functions, in addition to focusing on single-task methods. Moreover, the paper sheds light on the future avenues of ML in protein function determination. Kamal Taha |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | Semi-supervised and un-supervised clustering: A review and experimental evaluation
Kamal Taha |
Inf. Syst. | 1 |
| 2023 | Context-Driven-Based Community DetectionabstractDetecting cross-communities constructed based on the commonalities of their adaptive social traits is crucial for solving many settings of real-world problems, such as examining the dynamics of a social network, examining the behavior patterns that influence the outbreak and spread of disease, determining a criminal organization’s influential individuals, and amplifying a business potential. Unfortunately, investigating approaches that emphasize the detection of such cross-communities has been understudied. Moreover, few current approaches may detect cross-communities but without regard for their granularities. To overcome this, we introduce a novel methodology that analyses the overlapping among social traits to detect the most granular cross-communities with multisocial traits. The methodology is implemented in a working system called CDBCD. It can detect the most granular multisocial traits cross-communities, to which an active user (i.e., a context user) belongs. Such cross-communities are detected and selected by the system if they exhibit evidence of multisocial traits’ homophily among their users. We propose novel context-driven search techniques that infer the relationships among various social traits. We evaluated CDBCD by comparing it with ten methods. The results demonstrated that CDBCD can detect granular cross-communities with marked accuracy. The improvement of CDBCD over the ten methods combined is 23% and 19% in terms of ARI and F1-score, respectively. Kamal Taha |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Identifying and Protecting Cyber-Physical Systems' Influential Devices for Sustainable CybersecurityabstractFor sustainable cyber-physical systems (CPS) security, proactive measures to cybersecurity need to be implemented instead of reactive measures. Towards this, we introduce in this paper a proactive methodology implemented in a system called IDI_CPS. It is based on the observation that CPS devices that have LAN-based network sharing (e.g., via Wi-Fi connections) need first to be clustered using some clustering criterion. Then, the influential and central devices in these clusters need to be identified to pay more attention to their file sharing. These influential devices may have network sharing with devices at the WAN level. Therefore, the influential devices at the WAN level that have network sharing with the influential devices in the clusters need also to be identified to pay more attention to their file sharing. We propose novel techniques for: (1) clustering the devices that have LAN-based network sharing using k-clique modeling, (2) employing clustering coefficient-based techniques for identifying the most influential device in each cluster, and (3) employing Independent Cascades model-based techniques for identifying the influential devices at the WAN level that have network sharing with the influential devices in the clusters. We experimentally evaluated our proposed system IDI_CPS and compared it with four comparable methods. Results showed marked improvement. Kamal Taha |
IEEE Trans. Sustain. Comput. | 1 |
| 2022 | Inferring the densest multi-profiled cross-community for a user
Kamal Taha, Paul D. Yoo, Fatima Zohra Eddinari, Siniya Nedunkulathil |
Knowl. Based Syst. | 1 |
| 2021 | Detecting Disjoint Communities in a Social Network Based on the Degrees of Association Between Edges and Influential NodesabstractDetecting communities is crucial to understanding the dynamics of their members. However, the detection of “good” communities is deemed demonstrably problematic, which is mainly due to the following two factors. First, real-world networks are complex and require optimizing multi-objective functions for capturing their community structures, whereas most current approaches optimize only one or two objective functions. Second, most current approaches detect communities in respect of the independence of how closely associated their connections are based on the global relative influences of the edges connecting them. To overcome these limitations, a clustering method needs to optimize multi-objective functions and employ global preprocessing techniques that consider the topology of the entire network. We, therefore, proposed a system called DAVE, which optimizes four objective functions that capture the community structures in most real-word network settings, and detects communities with regards of how closely associated their connections are based on the relative influences of the edges connecting them. We proposed novel formulas that capture these functions. Our method is the first to utilize the prediction of node-node associations based on global node-edge degrees of association. After ranking nodes based on their global relative influences on the network, some of the top-ranked ones will serve as core seeds for constructing communities. Then, the degrees of association between influential edges and seed nodes are computed. DAVE assigns a node to a community, only if each edge in the shortest path from this node to the community's core seed node is both influential and has significant degree of association with the core node. We evaluated DAVE by comparing it empirically and experimentally with 16 methods. Results showed a remarkable improvement. Kamal Taha |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | R2BN: An Adaptive Model for Keystroke-Dynamics-Based Educational Level ClassificationabstractOver the past decade, keystroke-based pattern recognition techniques, as a forensic tool for behavioral biometrics, have gained increasing attention. Although a number of machine learning-based approaches have been proposed, they are limited in terms of their capability to recognize and profile a set of an individual's characteristics. In addition, up to today, their focus was primarily gender and age, which seem to be more appropriate for commercial applications (such as developing commercial software), leaving out from research other characteristics, such as the educational level. Educational level is an acquired user characteristic, which can improve targeted advertising, as well as provide valuable information in a digital forensic investigation, when it is known. In this context, this paper proposes a novel machine learning model, the randomized radial basis function network, which recognizes and profiles the educational level of an individual who stands behind the keyboard. The performance of the proposed model is evaluated by using the empirical data obtained by recording volunteers' keystrokes during their daily usage of a computer. Its performance is also compared with other well-referenced machine learning models using our keystroke dynamic datasets. Although the proposed model achieves high accuracy in educational level prediction of an unknown user, it suffers from high computational cost. For this reason, we examine ways to reduce the time that is needed to build our model, including the use of a novel data condensation method, and discuss the tradeoff between an accurate and a fast prediction. To the best of our knowledge, this is the first model in the literature that predicts the educational level of an individual based on the keystroke dynamics information only. Ioannis Tsimperidis, Paul D. Yoo, Kamal Taha, Alexios Mylonas, Vasilios Katos |
IEEE Trans. Cybern. | 3 |
| 2019 | DEMISe: Interpretable Deep Extraction and Mutual Information Selection Techniques for IoT Intrusion DetectionabstractRecent studies have proposed that traditional security technology -- involving pattern-matching algorithms that check predefined pattern sets of intrusion signatures -- should be replaced with sophisticated adaptive approaches that combine machine learning and behavioural analytics. However, machine learning is performance driven, and the high computational cost is incompatible with the limited computing power, memory capacity and energy resources of portable IoT-enabled devices. The convoluted nature of deep-structured machine learning means that such models also lack transparency and interpretability. The knowledge obtained by interpretable learners is critical in security software design. We therefore propose two novel models featuring a common Deep Extraction and Mutual Information Selection (DEMISe) element which extracts features using a deep-structured stacked autoencoder, prior to feature selection based on the amount of mutual information (MI) shared between each feature and the class label. An entropy-based tree wrapper is used to optimise the feature subsets identified by the DEMISe element, yielding the DEMISe with Tree Evaluation and Regression Detection (DETEReD) model. This affords 'white box' insight, and achieves a time to build of 603 seconds, a 99.07% detection rate, and 98.04% model accuracy. When tested against AWID, the best-referenced intrusion detection dataset, the new models achieved a test error comparable to or better than state-of-the-art machine-learning models, with a lower computational cost and higher levels of transparency and interpretability. Luke R. Parker, Paul D. Yoo, A. Taufiq Asyhari, Lounis Chermak, Yoonchan Jhi, Kamal Taha |
ARES | 6 |
| 2019 | Predicting the Functions of Proteins from their Co-occurrences with Implicit and Explicit Functional Terms in TextsabstractRecent computational methods take advantage of the exponential explosion of biomedical literatures to predict protein functions. They do so by extracting information from the literatures that directly (i.e., explicitly)describe the functions of already annotated proteins. We observe that some biological terms pertaining protein functions may co-occur implicitly with proteins in biomedical texts. This has led us to believe that the methods that rely only on explicitly mentioned biomedical terms in texts may miss vital information about protein functions that is implicitly mentioned in the texts. Towards this, we propose in this paper an Information Extraction system called IPFI that employs techniques for predicting the functions of proteins from their co-occurrences in biomedical texts with both implicitly and explicitly mentioned biological terms pertaining functional categories. That is, IPFI uses a combination of explicit term extraction methods and logic-based implicit term extraction methods. It extracts explicit terms using Natural Language Processing techniques. It infers implicit terms by employing the inference rules of predicate logic. It triggers protein specification rules recursively. These rules are represented in the form of predicate logic's premises. We evaluated IPFI by comparing it experimentally with four existing methods. Results revealed marked improvement. Kamal Taha |
CIBCB | 1 |
| 2019 | Detecting Overlapping Communities of Nodes with Multiple Attributes from Heterogeneous Networks
Kamal Taha, Paul D. Yoo |
CollaborateCom | 1 |
| 2019 | Analyzing a co-occurrence gene-interaction network to identify disease-gene associationabstractBACKGROUND: Understanding the genetic networks and their role in chronic diseases (e.g., cancer) is one of the important objectives of biological researchers. In this work, we present a text mining system that constructs a gene-gene-interaction network for the entire human genome and then performs network analysis to identify disease-related genes. We recognize the interacting genes based on their co-occurrence frequency within the biomedical literature and by employing linear and non-linear rare-event classification models. We analyze the constructed network of genes by using different network centrality measures to decide on the importance of each gene. Specifically, we apply betweenness, closeness, eigenvector, and degree centrality metrics to rank the central genes of the network and to identify possible cancer-related genes. RESULTS: We evaluated the top 15 ranked genes for different cancer types (i.e., Prostate, Breast, and Lung Cancer). The average precisions for identifying breast, prostate, and lung cancer genes vary between 80-100%. On a prostate case study, the system predicted an average of 80% prostate-related genes. CONCLUSIONS: The results show that our system has the potential for improving the prediction accuracy of identifying gene-gene interaction and disease-gene associations. We also conduct a prostate cancer case study by using the threshold property in logistic regression, and we compare our approach with some of the state-of-the-art methods. Amira Al-Aamri, Kamal Taha, Yousof Al-Hammadi, Maher Maalouf, Dirar Homouz |
BMC Bioinform. | 2 |
| 2019 | Predicting protein functions by applying predicate logic to biomedical literatureabstractBACKGROUND: A large number of computational methods have been proposed for predicting protein functions. The underlying techniques adopted by most of these methods revolve around predicting the functions of an unannotated protein p from already annotated proteins that have similar characteristics as p. Recent Information Extraction methods take advantage of the huge growth of biomedical literature to predict protein functions. They extract biological molecule terms that directly describe protein functions from biomedical texts. However, they consider only explicitly mentioned terms that co-occur with proteins in texts. We observe that some important biological molecule terms pertaining functional categories may implicitly co-occur with proteins in texts. Therefore, the methods that rely solely on explicitly mentioned terms in texts may miss vital functional information implicitly mentioned in the texts. RESULTS: To overcome the limitations of methods that rely solely on explicitly mentioned terms in texts to predict protein functions, we propose in this paper an Information Extraction system called PL-PPF. The proposed system employs techniques for predicting the functions of proteins based on their co-occurrences with explicitly and implicitly mentioned biological molecule terms that pertain functional categories in biomedical literature. That is, PL-PPF employs a combination of statistical-based explicit term extraction techniques and logic-based implicit term extraction techniques. The statistical component of PL-PPF predicts some of the functions of a protein by extracting the explicitly mentioned functional terms that directly describe the functions of the protein from the biomedical texts associated with the protein. The logic-based component of PL-PPF predicts additional functions of the protein by inferring the functional terms that co-occur implicitly with the protein in the biomedical texts associated with it. First, the system employs its statistical-based component to extract the explicitly mentioned functional terms. Then, it employs its logic-based component to infer additional functions of the protein. Our hypothesis is that important biological molecule terms pertaining functional categories of proteins are likely to co-occur implicitly with the proteins in biomedical texts. We evaluated PL-PPF experimentally and compared it with five systems. Results revealed better prediction performance. CONCLUSIONS: The experimental results showed that PL-PPF outperformed the other five systems. This is an indication of the effectiveness and practical viability of PL-PPF's combination of explicit and implicit techniques. We also evaluated two versions of PL-PPF: one adopting the complete techniques (i.e., adopting both the implicit and explicit techniques) and the other adopting only the explicit terms co-occurrence extraction techniques (i.e., without the inference rules for predicate logic). The experimental results showed that the complete version outperformed significantly the other version. This is attributed to the effectiveness of the rules of predicate logic to infer functional terms that co-occur implicitly with proteins in biomedical texts. A demo application of PL-PPF can be accessed through the following link: http://ecesrvr.kustar.ac.ae:8080/plppf/. Kamal Taha, Youssef Iraqi, Amira Al-Aamri |
BMC Bioinform. | 1 |
| 2019 | Shortlisting the Influential Members of Criminal Organizations and Identifying Their Important Communication ChannelsabstractLow-level criminals, who do the legwork in a criminal organization, are the most likely to be arrested, whereas the high-level ones tend to avoid attention. But crippling the work of criminal organizations is not possible unless investigators can identify the most influential, high-level members and monitor their communication channels. Investigators often approach this task by requesting the mobile phone service records of the arrested low-level criminals to identify contacts, and then they build a network model of the organization, where each node denotes a criminal and the edges represent communications. Network analysis can be used to infer the most influential criminals and most important communication channels within the network, but screening all the nodes and links in a network is laborious and time consuming. Here, we propose a new forensic analysis system called identifying influential criminals and their communication channels (IICCC) that can effectively and efficiently infer the high-level criminals and short-list the important communication channels in a criminal organization, based on the mobile phone communications of its members. IICCC can also be used to build a network from crime incident reports. We evaluated IICCC experimentally and compared it with five other systems, confirming its superior prediction performance. Kamal Taha, Paul D. Yoo |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Inferring the Functions of Proteins from the Interrelationships between Functional CategoriesabstractThis study proposes a new method to determine the functions of an unannotated protein. The proteins and amino acid residues mentioned in biomedical texts associated with an unannotated protein can be considered as characteristics terms for , which are highly predictive of the potential functions of . Similarly, proteins and amino acid residues mentioned in biomedical texts associated with proteins annotated with a functional category can be considered as characteristics terms of . We introduce in this paper an information extraction system called IFP_IFC that predicts the functions of an unannotated protein by representing and each functional category by a vector of weights. Each weight reflects the degree of association between a characteristic term and (or a characteristic term and ). First, IFP_IFC constructs a network, whose nodes represent the different functional categories, and its edges the interrelationships between the nodes. Then, it determines the functions of by employing random walks with restarts on the mentioned network. The walker is the vector of . Finally, is assigned to the functional categories of the nodes in the network that are visited most by the walker. We evaluated the quality of IFP_IFC by comparing it experimentally with two other systems. Results showed marked improvement. Kamal Taha |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | Disjoint Community Detection in Networks Based on the Relative Association of MembersabstractWe propose in this paper a hybrid system called DCD_RAM that detects disjoint communities. It is based, in part, on the underlying techniques of the network-centric, hierarchy-centric, vertex-centric, and group-centric approaches. It adopts most of the underlying techniques of the four approaches. Most of these approaches work well in only networks with certain topologies. DCD_RAM aims at overcoming the limitations of each of the four approaches to enable it to work well in networks with all types of topologies. It does so by: 1) measuring the betweenness of each edge (u, v) in such a way that the betweenness acts as an indicator of the influences of vertices u and v over the flow of information in the entire network; 2) employing a novel logarithm-based formula that captures and further enhances the eigenvector principle in order to characterize the global influence of each vertex in the network; 3) employing a novel agglomerative-like formula that discovers natural divisions of a network; and 4) employing a novel belonging formula that helps in discovering disjoint communities. We evaluated DCD_RAM by comparing it empirically and experimentally with nine methods. Results showed marked improvement. Kamal Taha |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2017 | Using the Spanning Tree of a Criminal Network for Identifying Its LeadersabstractWe introduce a forensic analysis system called ECLfinder that identifies the influential members of a criminal organization as well as the immediate leaders of a given list of lower-level criminals. Criminal investigators usually seek to identify the influential members of criminal organizations, because eliminating them is most likely to hinder and disrupt the operations of these organizations and put them out of business. First, ECLfinder constructs a network representing a criminal organization from either mobile communication data associated with the organization or crime incident reports that include information about the organization. It then constructs a minimum spanning tree (MST) of the network. It identifies the influential members of a criminal organization by determining the important vertices in the network representing the organization, using the concept of existence dependence. Each vertex v is assigned a score, which is the number of other vertices, whose existence in MST is dependent on v. Vertices are ranked based on their scores. Criminals represented by the top ranked vertices are considered the influential members of the criminal organization represented by the network. We evaluated the quality of ECLfinder by comparing it experimentally with three other systems. Results showed marked improvement. Kamal Taha, Paul D. Yoo |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | Predicting the functions of a protein from its ability to associate with other moleculesabstractBACKGROUND: All proteins associate with other molecules. These associated molecules are highly predictive of the potential functions of proteins. The association of a protein and a molecule can be determined from their co-occurrences in biomedical abstracts. Extensive semantically related co-occurrences of a protein's name and a molecule's name in the sentences of biomedical abstracts can be considered as indicative of the association between the protein and the molecule. Dependency parsers extract textual relations from a text by determining the grammatical relations between words in a sentence. They can be used for determining the textual relations between proteins and molecules. Despite their success, they may extract textual relations with low precision. This is because they do not consider the semantic relationships between terms in a sentence (i.e., they consider only the structural relationships between the terms). Moreover, they may not be well suited for complex sentences and for long-distance textual relations. RESULTS: We introduce an information extraction system called PPFBM that predicts the functions of unannotated proteins from the molecules that associate with these proteins. PPFBM represents each protein by the other molecules that associate with it in the abstracts referenced in the protein's entries in reliable biological databases. It automatically extracts each co-occurrence of a protein-molecule pair that represents semantic relationship between the pair. Towards this, we present novel semantic rules that identify the semantic relationship between each co-occurrence of a protein-molecule pair using the syntactic structures of sentences and linguistics theories. PPFBM determines the functions of an un-annotated protein p as follows. First, it determines the set S r of annotated proteins that is semantically similar to p by matching the molecules representing p and the annotated proteins. Then, it assigns p the functional category FC if the significance of the frequency of occurrences of S r in abstracts associated with proteins annotated with FC is statistically significantly different than the significance of the frequency of occurrences of S r in abstracts associated with proteins annotated with all other functional categories. We evaluated the quality of PPFBM by comparing it experimentally with two other systems. Results showed marked improvement. CONCLUSIONS: The experimental results demonstrated that PPFBM outperforms other systems that predict protein function from the textual information found within biomedical abstracts. This is because these system do not consider the semantic relationships between terms in a sentence (i.e., they consider only the structural relationships between the terms). PPFBM's performance over these system increases steadily as the number of training protein increases. That is, PPFBM's prediction performance becomes more accurate constantly, as the size of training proteins gets larger. This is because every time a new set of test proteins is added to the current set of training proteins. A demo of PPFBM that annotates each input Yeast protein (SGD (Saccharomyces Genome Database). Available at: http://www.yeastgenome.org/download-data/curation) with the functions of Gene Ontology terms is available at: (see Appendix for more details about the demo) http://ecesrvr.kustar.ac.ae:8080/PPFBM/. Kamal Taha, Paul D. Yoo |
BMC Bioinform. | 1 |
| 2016 | Erratum to: Predicting the functions of a protein from its ability to associate with other moleculesabstractUpon publication of the original article [1] it was found that author corrections for Table 1 were not correctly implemented. You can now find the correct Table 1 below and also updated in the original article.
Table 1
The distribution of semantically related and semantically unrelated co-occurrences of molecule m i and Protein p Pair in an Abstract A j Kamal Taha, Paul D. Yoo |
BMC Bioinform. | 1 |
| 2016 | Applying Monte Carlo Simulation to Biomedical Literature to Approximate Genetic NetworkabstractBiologists often need to know the set of genes associated with a given set of genes or a given disease. We propose in this paper a classifier system called Monte Carlo for Genetic Network (MCforGN) that can construct genetic networks, identify functionally related genes, and predict gene-disease associations. MCforGN identifies functionally related genes based on their co-occurrences in the abstracts of biomedical literature. For a given gene g , the system first extracts the set of genes found within the abstracts of biomedical literature associated with g. It then ranks these genes to determine the ones with high co-occurrences with g . It overcomes the limitations of current approaches that employ analytical deterministic algorithms by applying Monte Carlo Simulation to approximate genetic networks. It does so by conducting repeated random sampling to obtain numerical results and to optimize these results. Moreover, it analyzes results to obtain the probabilities of different genes' co-occurrences using series of statistical tests. MCforGN can detect gene-disease associations by employing a combination of centrality measures (to identify the central genes in disease-specific genetic networks) and Monte Carlo Simulation. MCforGN aims at enhancing state-of-the-art biological text mining by applying novel extraction techniques. We evaluated MCforGN by comparing it experimentally with nine approaches. Results showed marked improvement. Rami Al-Dalky, Kamal Taha, Dirar Mohammad Al Homouz, Murad Qasaimeh |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2016 | Data Randomization and Cluster-Based Partitioning for Botnet Intrusion DetectionabstractBotnets, which consist of remotely controlled compromised machines called bots, provide a distributed platform for several threats against cyber world entities and enterprises. Intrusion detection system (IDS) provides an efficient countermeasure against botnets. It continually monitors and analyzes network traffic for potential vulnerabilities and possible existence of active attacks. A payload-inspection-based IDS (PI-IDS) identifies active intrusion attempts by inspecting transmission control protocol and user datagram protocol packet's payload and comparing it with previously seen attacks signatures. However, the PI-IDS abilities to detect intrusions might be incapacitated by packet encryption. Traffic-based IDS (T-IDS) alleviates the shortcomings of PI-IDS, as it does not inspect packet payload; however, it analyzes packet header to identify intrusions. As the network's traffic grows rapidly, not only the detection-rate is critical, but also the efficiency and the scalability of IDS become more significant. In this paper, we propose a state-of-the-art T-IDS built on a novel randomized data partitioned learning model (RDPLM), relying on a compact network feature set and feature selection techniques, simplified subspacing and a multiple randomized meta-learning technique. The proposed model has achieved 99.984% accuracy and 21.38 s training time on a well-known benchmark botnet dataset. Experiment results demonstrate that the proposed methodology outperforms other well-known machine-learning models used in the same detection task, namely, sequential minimal optimization, deep neural network, C4.5, reduced error pruning tree, and randomTree. Omar Y. Al-Jarrah, Omar Alhussein, Paul D. Yoo, Sami Muhaidat, Kamal Taha, Kwangjo Kim |
IEEE Trans. Cybern. | 5 |
| 2016 | SIIMCO: A Forensic Investigation Tool for Identifying the Influential Members of a Criminal OrganizationabstractMembers of a criminal organization, who hold central positions in the organization, are usually targeted by criminal investigators for removal or surveillance. This is because they play key and influential roles by acting as commanders, who issue instructions or serve as gatekeepers. Removing these central members (i.e., influential members) is most likely to disrupt the organization and put it out of business. Most often, criminal investigators are even more interested in knowing the portion of these influential members, who are the immediate leaders of lower level criminals. These lower level criminals are the ones who usually carry out the criminal works; therefore, they are easier to identify. The ultimate goal of investigators is to identify the immediate leaders of these lower level criminals in order to disrupt future crimes. We propose, in this paper, a forensic analysis system called SIIMCO that can identify the influential members of a criminal organization. Given a list of lower level criminals in a criminal organization, SIIMCO can also identify the immediate leaders of these criminals. SIIMCO first constructs a network representing a criminal organization from either mobile communication data that belongs to the organization or crime incident reports. It adopts the concept space approach to automatically construct a network from crime incident reports. In such a network, a vertex represents an individual criminal, and a link represents the relationship between two criminals. SIIMCO employs formulas that quantify the degree of influence/importance of each vertex in the network relative to all other vertices. We present these formulas through a series of refinements. All the formulas incorporate novel-weighting schemes for the edges of networks. We evaluated the quality of SIIMCO by comparing it experimentally with two other systems. Results showed marked improvement. Kamal Taha, Paul D. Yoo |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2015 | A System for Analyzing Criminal Social NetworksabstractThe influential members of a criminal organization are usually targeted by criminal investigators for removal or surveillance. Identifying and capturing these influential members will most likely to disrupt the organization. We propose in this paper a forensic analysis system called CLDRI that can identify the most influential members of a criminal organization. First, a network representing a criminal organization is built from Mobile Communication Data that belongs to the organization. In such a network, a vertex represents an individual criminal and an edge represents the communication attempts between two criminals. CLDRI employs formulas that quantify the degree of importance of each vertex in the network relative to all other vertices. We present these formulas through series of improvement refinements. All the formulas incorporate novel-weighting schemes for the edges of networks. We evaluated the quality of CLDRI by comparing it experimentally with two systems. Results showed improvement. Kamal Taha, Paul D. Yoo |
ASONAM | 1 |
| 2015 | Semantic rules for extracting proteins functions information from biomedical abstractsabstractWe present a classifier system called SRPFP that predicts the functions of un-annotated proteins. SRPFP aims at enhancing the state of the art of biological text mining. It analyzes biomedical texts in order to discover protein function information that is difficult to retrieve. It employs semantic rules for extracting proteins functions information from biomedical abstracts. It applies a novel model and linguistic computational techniques for extracting the functional relationship from different structural forms of terms in the sentences of biological abstracts. Specifically, SRPFP extracts phrases that represent functional relationships between proteins and molecules. These molecules usually bind to the proteins and are highly predictive of the functions of these proteins. The proposed semantic rules can identify the semantic relationship between each co-occurrence of a protein-molecule pair using the syntactic structures of sentences and linguistics theories. SRPFP represents each protein by the molecules that have high co-occurrences with the protein in biomedical abstracts. This is because such molecules are good characteristics and indicators of the functions of proteins. SRPFP measures the semantic similarity between the molecules representing an un-annotated protein p and the molecules representing annotated proteins and assigns p the functions of annotated proteins that are similar to p. Kamal Taha |
BIBM | 1 |
| 2015 | An information extraction system for protein function predictionabstractWe present a Natural Language Processing extraction system called IESforPFP, which can retrieve useful information from biomedical abstracts. IESforPFP aims at enhancing the state of the art of biological text mining by applying novel linguistic computational technique. By retrieving significant patterns of associations between proteins and molecules from biomedical abstracts, IESforPFP can determine the functions of un-annotated proteins. The system determines the semantic relationship between each protein-molecule pair in sentences using novel semantic rules. It applies a semantic relationship extraction model that retrieves information from different structural forms of constituents in sentences. In the framework of IESforPFP, each protein p is represented by a vector of weights. Each weight reflects the significance of a molecule m in the biomedical abstracts associated with p. That is, each weight quantifies the likelihood of the association between m and p. IESforPFP determines the set of annotated proteins that is semantically similar to p by comparing their vectors. It then annotates p with the functions of these annotated proteins. We evaluated the quality of IESforPFP by comparing it experimentally with two other systems. Results showed marked improvement. Kamal Taha, Paul D. Yoo |
CIBCB | 1 |
| 2015 | Randomized Subspace Learning for Proline Cis-Trans Isomerization PredictionabstractProline residues are common source of kinetic complications during folding. The X-Pro peptide bond is the only peptide bond for which the stability of the cis and trans conformations is comparable. The cis-trans isomerization (CTI) of X-Pro peptide bonds is a widely recognized rate-limiting factor, which can not only induces additional slow phases in protein folding but also modifies the millisecond and sub-millisecond dynamics of the protein. An accurate computational prediction of proline CTI is of great importance for the understanding of protein folding, splicing, cell signaling, and transmembrane active transport in both the human body and animals. In our earlier work, we successfully developed a biophysically motivated proline CTI predictor utilizing a novel tree-based consensus model with a powerful metalearning technique and achieved 86.58 percent Q2 accuracy and 0.74 Mcc, which is a better result than the results (70-73 percent Q2 accuracies) reported in the literature on the well-referenced benchmark dataset. In this paper, we describe experiments with novel randomized subspace learning and bootstrap seeding techniques as an extension to our earlier work, the consensus models as well as entropy-based learning methods, to obtain better accuracy through a precise and robust learning scheme for proline CTI prediction. Omar Y. Al-Jarrah, Paul D. Yoo, Kamal Taha, Sami Muhaidat, Abdallah Shami, Nazar Zaki |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2015 | iPFPi: A System for Improving Protein Function Prediction through Cumulative IterationsabstractWe propose a classifier system called iPFPi that predicts the functions of un-annotated proteins. iPFPi assigns an un-annotated protein P the functions of GO annotation terms that are semantically similar to P. An un-annotated protein P and a GO annotation term T are represented by their characteristics. The characteristics of P are GO terms found within the abstracts of biomedical literature associated with P. The characteristics of Tare GO terms found within the abstracts of biomedical literature associated with the proteins annotated with the function of T. Let F and F/ be the important (dominant) sets of characteristic terms representing T and P, respectively. iPFPi would annotate P with the function of T, if F and F/ are semantically similar. We constructed a novel semantic similarity measure that takes into consideration several factors, such as the dominance degree of each characteristic term t in set F based on its score, which is a value that reflects the dominance status of t relative to other characteristic terms, using pairwise beats and looses procedure. Every time a protein P is annotated with the function of T, iPFPi updates and optimizes the current scores of the characteristic terms for T based on the weights of the characteristic terms for P. Set F will be updated accordingly. Thus, the accuracy of predicting the function of T as the function of subsequent proteins improves. This prediction accuracy keeps improving over time iteratively through the cumulative weights of the characteristic terms representing proteins that are successively annotated with the function of T. We evaluated the quality of iPFPi by comparing it experimentally with two recent protein function prediction systems. Results showed marked improvement. Kamal Taha, Paul D. Yoo, Mohammed Alzaabi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | CISRI: A Crime Investigation System Using the Relative Importance of Information Spreaders in Networks Depicting Criminals CommunicationsabstractIn this paper, we propose a forensic analysis system called crime investigation system using the relative importance (CISRI) that helps forensic investigators determine the most influential members of a criminal group, who are related to known members of the group, for the purposes of investigation. In the CISRI framework, we describe the structural relationships between the members of a criminal group in terms of a graph. In such a graph, a node represents a member of a criminal group, an edge connecting two nodes represents the relationship between two members of the group, and the weight of an edge represents the degree of the relationship between those two members. Using this representation, we propose a method that determines the relative importance of nodes in a graph with respect to a given set of query nodes. Most current approaches that study relative importance determine the relative importance of a node under consideration by estimating the contribution of each query node individually to the importance of this node while overlooking the contribution of the query nodes collectively to the importance of the node under consideration. This may lead to results with low precision. CISRI overcomes this limitation by: 1) computing the contribution of the overall set of query nodes to the importance of a node under consideration and 2) adopting a tight constraint calculation that considers how much each query node contributes to the relative importance of a node under consideration. This leads to accurate identification of nodes in the graph that are important, in relation to the query nodes. In the framework of CISRI, a graph is constructed from mobile communication records (e.g., phone calls and messages), where a node represents a caller and the weight of an edge reflects the number of contacts between two callers. We evaluated the quality of CISRI by comparing it experimentally with three comparable methods. Our results showed marked improvement. Mohammed Alzaabi, Kamal Taha, Thomas Martin 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | Simplified Subspaced Regression Network for Identification of Defect Patterns in Semiconductor Wafer MapsabstractWafer defects, which are primarily defective chips on a wafer, are of the key challenges facing the semiconductor manufacturing companies, as they could increase the yield losses to hundreds of millions of dollars. Fortunately, these wafer defects leave unique patterns due to their spatial dependence across wafer maps. It is thus possible to identify and predict them in order to find the point of failure in the manufacturing process accurately. This paper introduces a novel simplified subspaced regression framework for the accurate and efficient identification of defect patterns in semiconductor wafer maps. It can achieve a test error comparable to or better than the state-of-the-art machine-learning (ML)-based methods, while maintaining a low computational cost when dealing with large-scale wafer data. The effectiveness and utility of the proposed approach has been demonstrated by our experiments on real wafer defect datasets, achieving detection accuracy of 99.884% and R2of 99.905%, which are far better than those of any existing methods reported in the literature. Fatima Adly, Omar Alhussein, Paul D. Yoo, Yousof Al-Hammadi, Kamal Taha, Sami Muhaidat, Youngseon Jeong 0001, Uihyoung Lee, Mohammed Ismail 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2015 | Extracting Various Classes of Data From Biological Text Using the Concept of Existence DependencyabstractOne of the key goals of biological natural language processing (NLP) is the automatic information extraction from biomedical publications. Most current constituency and dependency parsers overlook the semantic relationships between the constituents comprising a sentence and may not be well suited for capturing complex long-distance dependences. We propose in this paper a hybrid constituency-dependency parser for biological NLP information extraction called EDCC. EDCC aims at enhancing the state of the art of biological text mining by applying novel linguistic computational techniques that overcome the limitations of current constituency and dependency parsers outlined earlier, as follows: 1) it determines the semantic relationship between each pair of constituents in a sentence using novel semantic rules; and 2) it applies a semantic relationship extraction model that extracts information from different structural forms of constituents in sentences. EDCC can be used to extract different types of data from biological texts for purposes such as protein function prediction, genetic network construction, and protein-protein interaction detection. We evaluated the quality of EDCC by comparing it experimentally with six systems. Results showed marked improvement. Kamal Taha |
IEEE J. Biomed. Health Informatics | 1 |
| 2014 | Inferring the relationships among genes from weighted GO graphabstractBiologists may need to know the set of genes that are semantically related to a given set of genes. For instance, a biologist may need to know the set of genes related to another set of genes known to be involved in a specific disease. Some works use the concept of gene clustering in order to identify semantically related genes. Others propose tools that return the set of genes that are semantically related to a given set of genes. Most of these gene similarity measures determine the semantic similarities among the genes based solely on the proximity to each other of the GO terms annotating the genes, while overlook the structural dependencies among these GO terms, which may lead to low recall and precision of results. We propose in this paper a search engine called IRG, which overcomes the problems of current gene similarity measures outlined above. The search engine constructs a minimum spanning tree of GO graph based on their weights. Let S' be the set of genes that are semantically related to set S. In the framework of IRG, the set S' is annotated to the GO term located at the convergence of the subtree of the minimum spanning tree that passes through the GO terms annotating the set S. We evaluated IRG experimentally and compared it with a gene prediction tool called DynGO and with two other systems we proposed previously. Results showed marked improvement. Kamal Taha, Paul D. Yoo |
CIBCB | 1 |
| 2014 | BioHCDP: A Hybrid Constituency-Dependency Parser for Biological NLP information extractionabstractOne of the key goals of biological Natural Language Processing (NLP) is the automatic information extraction from biomedical publications. Most current constituency and dependency parsers overlook the semantic relationships between the constituents comprising a sentence and may not be well suited for capturing complex long-distance dependencies. We propose in this paper a hybrid constituency-dependency parser for biological NLP information extraction called BioHCDP. BioHCDP aims at enhancing the state of the art of biological text mining by applying novel linguistic computational techniques that overcome the limitations of current constituency and dependency parsers outlined above, as follows: (1) it determines the semantic relationship between each pair of constituents in a sentence using novel semantic rules, and (2) it applies semantic relationship extraction models that represent the relationships of different patterns of usage in different contexts. BioHCDP can be used to extract various classes of data from biological texts, including protein function assignments, genetic networks, and protein-protein interactions. We compared BioHCDP experimentally with three systems. Results showed marked improvement. Kamal Taha, Mohammed Alzaabi |
CIDM | 1 |
| 2014 | Determining Semantically Related Significant GenesabstractGO relation embodies some aspects of existence dependency. If GO term xis existence-dependent on GO term y, the presence of y implies the presence of x. Therefore, the genes annotated with the function of the GO term y are usually functionally and semantically related to the genes annotated with the function of the GO term x. A large number of gene set enrichment analysis methods have been developed in recent years for analyzing gene sets enrichment. However, most of these methods overlook the structural dependencies between GO terms in GO graph by not considering the concept of existence dependency. We propose in this paper a biological search engine called RSGSearch that identifies enriched sets of genes annotated with different functions using the concept of existence dependency. We observe that GO term xcannot be existence-dependent on GO term y, if x- and y- have the same specificity (biological characteristics). After encoding into a numeric format the contributions of GO terms annotating target genes to the semantics of their lowest common ancestors (LCAs), RSGSearch uses microarray experiment to identify the most significant LCA that annotates the result genes. We evaluated RSGSearch experimentally and compared it with five gene set enrichment systems. Results showed marked improvement. Kamal Taha |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | Intelligent Consensus Modeling for ProlineCis-Trans Isomerization PredictionabstractProline cis-trans isomerization (CTI) plays a key role in the rate-determining steps of protein folding. Accurate prediction of proline CTI is of great importance for the understanding of protein folding, splicing, cell signaling, and transmembrane active transport in both the human body and animals. Our goal is to develop a state-of-the-art proline CTI predictor based on a biophysically motivated intelligent consensus modeling through the use of sequence information only (i.e., position specific scores generated by PSI-BLAST). The current computational proline CTI predictors reach about 70-73 percent Q2 accuracies and about 0.40 Matthew correlation coefficient (Mcc) through the use of sequence-based evolutionary information as well as predicted protein secondary structure information. However, our approach that utilizes a novel decision tree-based consensus model with a powerful randomized-metal earning technique has achieved 86.58 percent Q2 accuracy and 0.74 Mcc, on the same proline CTI data set, which is a better result than those of any existing computational proline CTI predictors reported in the literature. Paul D. Yoo, Sami Muhaidat, Kamal Taha, Jamal Bentahar, Abdallah Shami |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | GRank: a middleware search engine for ranking genes by relevance to given genesabstractBACKGROUND: Biologists may need to know the set of genes that are semantically related to a given set of genes. For instance, a biologist may need to know the set of genes related to another set of genes known to be involved in a specific disease. Some works use the concept of gene clustering in order to identify semantically related genes. Others propose tools that return the set of genes that are semantically related to a given set of genes. Most of these gene similarity measures determine the semantic similarities among the genes based solely on the proximity to each other of the GO terms annotating the genes, while overlook the structural dependencies among these GO terms, which may lead to low recall and precision of results. RESULTS: We propose in this paper a search engine called GRank, which overcomes the limitations of the current gene similarity measures outlined above as follows. It employs the concept of existence dependency to determine the structural dependencies among the GO terms annotating a given set of gene. After determining the set of genes that are semantically related to input genes, GRank would use microarray experiment to rank these genes based on their degree of relativity to the input genes. We evaluated GRank experimentally and compared it with a comparable gene prediction tool called DynGO, which retrieves the genes and gene products that are relatives of input genes. Results showed marked improvement. CONCLUSIONS: The experimental results demonstrated that GRank overcomes the limitations of current gene similarity measures. We attribute this performance to GRank's use of existence dependency concept for determining the semantic relationships among gene annotations. The recall and precision values for two benchmarking datasets showed that GRank outperforms DynGO tool, which does not employ the concept of existence dependency. The demo of GRank using 11000 KEGG yeast genes and a Gene Expression Omnibus (GEO) microarray file named "GSM34635.pad" is available at: http://ecesrvr.kustar.ac.ae:8080/ (click on the link labelled Gene Ontology 2). Kamal Taha, Dirar Mohammad Al Homouz, Hassan Al-Muhairi, Zaid Al Mahmoud |
BMC Bioinform. | 1 |
| 2013 | Determining the Semantic Similarities Among Gene Ontology TermsabstractWe present in this paper novel techniques that determine the semantic relationships among GeneOntology (GO) terms. We implemented these techniques in a prototype system called GoSE, which resides between user application and GO database. Given a set S of GO terms, GoSE would return another set S' of GO terms, where each term in S' is semantically related to each term in S. Most current research is focused on determining the semantic similarities among GO ontology terms based solely on their IDs and proximity to one another in the GO graph structure, while overlooking the contexts of the terms, which may lead to erroneous results. The context of a GO term T is the set of other terms, whose existence in the GO graph structure is dependent on T. We propose novel techniques that determine the contexts of terms based on the concept of existence dependency. We present a stack-based sort-merge algorithm employing these techniques for determining the semantic similarities among GO terms.We evaluated GoSE experimentally and compared it with three existing methods. The results of measuring the semantic similarities among genes in KEGG and Pfam pathways retrieved from the DBGET and Sanger Pfam databases, respectively, have shown that our method outperforms the other three methods in recall and precision. Kamal Taha |
IEEE J. Biomed. Health Informatics | 1 |
| 2012 | Personalization with Dynamic Group ProfileabstractIn this paper, we propose an XML-based recommender system, called PDGP. It is a type of collaborative information filtering system. PDGP uses ontology-driven social networks, where nodes represent social groups. A social group is an entity that defines a group based on demographic, ethnic, cultural, religious, age, or other characteristics. In the PDGP framework, query results are filtered and ranked based on the preferences of the social groups to which the user belongs. The user's social groups are inferred implicitly by the system without involving the user. PDGP constructs the social groups and identifies their preferences dynamically on the fly. These preferences are determined from the preferences of the social groups' member users using a group modeling strategy. PDGP can be used for various practical applications, such as Internet or other businesses that market preference-driven products. We experimentally compared PDGP with an existing system. Results showed marked improvement. Kamal Taha, Ramez Elmasri |
ASONAM | 1 |
| 2012 | GOcSim: GO context-driven similarityabstractWe present in this paper novel context-driven search techniques that compute the semantic similarities among different GO ontology terms. We implemented these techniques in a middleware search engine called GOcSim, which resides between user application and GO database. Most current research is focused on determining semantic similarities between GO ontology terms based solely on their IDs and proximity to one another in GO graph structure, while overlooking the contexts of the terms, which may lead to erroneous results. The context of a term T is determined by the set of other terms, whose existence is dependent on T. We propose novel techniques that determine the contexts of terms based on the concept of existence dependency. We present a stack-based sort-merge algorithm employing these techniques for determining the semantic similarities between GO terms. We evaluated GOcSim experimentally and compared it with four other methods. The results showed marked improvement. Kamal Taha, Ramez Elmasri |
CIBCB | 1 |
| 2012 | Automatic academic advisorabstractOne of the problems that face a Distance Education academic advisor (and for lesser degree local academic advisors) is to identify courses that best suit a student’s interests and academic skills from a wide collection of elective courses. This is because an advisor needs to select courses that suit Kamal Taha |
CollaborateCom | 1 |
| 2010 | SPGProfile: Speak Group Profile
Kamal Taha, Ramez Elmasri |
Inf. Syst. | 1 |
| 2010 | BusSEngine: a business search engine
Kamal Taha, Ramez Elmasri |
Knowl. Inf. Syst. | 1 |
| 2010 | XCDSearch: An XML Context-Driven Search EngineabstractWe present in this paper, a context-driven search engine called XCDSearch for answering XML Keyword-based queries as well as Loosely Structured queries, using a stack-based sort-merge algorithm. Most current research is focused on building relationships between data elements based solely on their labels and proximity to one another, while overlooking the contexts of the elements, which may lead to erroneous results. Since a data element is generally a characteristic of its parent, its context is determined by its parent. We observe that we could treat each set of elements consisting of a parent and its children data elements as one unified entity, and then use a stack-based sort-merge algorithm employing context-driven search techniques for determining the relationships between the different unified entities. We evaluated XCDSearch experimentally and compared it with five other search engines. The results showed marked improvement. Kamal Taha, Ramez Elmasri |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2009 | OOXKSearch: A Search Engine for Answering XML Keyword and Loosely Structured Queries Using OO TechniquesabstractOOXKSearch is a semantic search engine that answers XML keyword-based queries as well as loosely structured queries using Object Oriented techniques. There has been extensive research in XML keyword-based and loosely structured querying. Some frameworks work well for certain types of XML data models while fail in others. The reason is that the proposed techniques are based solely on establishing relationships between individual elements while overlooking the context of these elements. The context of a data element is determined by its parent, because it specifies one of the characteristics of the parent. Since data elements are nothing but characteristics of their parents, we observe that we could treat each parent-children set of elements as one unified entity. We then find semantic relationships between the different unified entities. If two distinct unified entities are semantically related, their data elements are also semantically related. The search performance and quality of OOXKSearch were evaluated experimentally and compared with three recent proposed systems. The results showed marked improvement. Kamal Taha, Ramez Elmasri |
J. Database Manag. | 1 |