EDBT 2026 Demo / reviewers in the wild / expert
Reda Alhajj
dblp:78/11148
· DBLP profile ↗
99ranked-venue papers in the field
17as first author
11since 2021 · last 2025
0000-0001-6657-9738ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 54 (3 first)Database Systems & Data Management · 16 (8 first)Knowledge Engineering, Semantic Web & Information Systems · 14 (3 first)Information Retrieval & Web Search · 12 (2 first)Other / Interdisciplinary · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vendor-Specific Vulnerability Analysis: A 26-Year Study of CVE Distribution Patterns
Yasamin Akrami, Melisa Saritas, Malek Malkawi, Reda Alhajj |
ASONAM (3) | 4 |
| 2024 | Medical Report Generation from Medical Images Using Vision Transformer and Bart Deep Learning Architectures
Murat Ucan, Buket Kaya, Mehmet Kaya, Reda Alhajj |
ASONAM (4) | 4 |
| 2023 | KoExPubMed: A Tool for Effective and Customized Knowledge Extraction from PubMedabstractAn exponential growth in the literature in general and the medical literature in particular raises a need for effective intelligent analysis strategies and tools to provide valuable insights to researchers about the current evolving literature. While existing applications provide more specific approaches to the problem, such as focusing on particular genome or protein information, in this paper, the proposed application provides effective and detailed analysis of PubMed. The developed tool, named KoExPubMed, follows a more generalized and holistic way by taking into consideration different types of information such as authors, countries, genes, and the interactions between them. The developed application consists of four main components; (1) keyword search and ID extraction, (2) PubMed article information and abstract retrieval, (3) country and address extraction, and (4) gene information extraction. In addition to the fundamental components, the tool provides a variety of visualization options for showing the extracted information and the related associations, including line charts for densities and countries, chord charts for collaborations of authors, network graphs for the genes mentioned together, bubble charts for gene frequencies, etc. By addressing the need for a generalized data mining tool, we propose a comprehensive application which is capable of employing data mining and machine learning techniques to extract from PubMed knowledge valuable to researchers and practitioners who are interested in closely investigating the achievements of others. Tansel Özyer, Reda Alhajj, Jon G. Rokne, Kashfia Sailunaz, Gabriela Jurca, Deniz Bestepe, Lama Alhajj, Busra Kartay |
ASONAM | 2 |
| 2023 | Investigating The Roles of microRNAs / lncRNAs in Characterizing Breast Cancer Subtypes and PrognosisabstractMolecular subtyping is a method of separating tumor clusters in a cancer type with common features according to molecular data and classification models. Genome datasets are taken from many different people and some genetic material, more precisely genetic markers, are obtained to predict the presence of a disease. In addition, breast cancer occurs due to mutation or modification observed in cells. miRNAs and lncRNAs take participation in cell cycle, regulation, and even chromatic inhibition of cell. For example, miRNAs function in cell cycle regulation as the degradation of mRNAs. Therefore, the aim of this work is to investigate the roles of miRNAs and lncRNAs in prognosis and characterizing the subtypes of Breast Cancer. Tansel Özyer, Reyhan Zeynep Pek, Muhammed Talha Zavalsiz, Melis Serdar, Sleiman Alhajj, Lama Alhajj, Jon G. Rokne, Reda Alhajj, Kashfia Sailunaz |
ASONAM | 8 |
| 2023 | Creating a Learning Profile by Using Face and Emotion RecognitionabstractThe aim of this work is to employ face recognition for creating learning profiles of the analysed persons who are students in this study. Generating education profiles will help experts in the diagnosis of Attention Deficit Hyperactivity Disorder (ADHD), which is a serious problem in children. Children with ADHD often have the ability and potential to learn. However, it may be difficult to reveal their capabilities and skills. Accordingly, a suffering child may have a hard time succeeding in real life when he/she is ignored and expected to mix with other children. The unrealized gap and deficiency may lead to other problems and more complicated situation with unpredictable consequences. Thanks to the system developed in this study, and the like, which will help in diagnosing the ADHD disease, and hence suffering individuals will be able to recognize their deficiencies, understand their ability to learn and adapt when approached differently in a way which suits his/her situation. This personalized handling of infected students will be an excellent guide to advance their potential and integration within the society carefully and smoothly. The system analyzes the face of a student to inspire his/her emotional state. The reported test results demonstrate how the system works well and produces high accuracy under a variety of severe conditions such as skewed angle, less illumination, accessories etc. Tansel Özyer, Gözde Yurtdas, Loubaba Alhajj, Jon G. Rokne, Kashfia Sailunaz, Reda Alhajj |
ASONAM | 6 |
| 2022 | NetDriller-V3: A Powerful Social Network Analysis ToolabstractThe development in technology has led to the generation of huge amounts of data from various sources, including biological data, social networking data, etc. Accordingly, social network analysis has received considerable attention with the availability of more raw datasets which could be realized using a network structure. Most of the datasets can be represented as a social network which is a graph consisting of actors having relationships. Many tools exist for social network analysis inspired to extract knowledge from the networks. NetDriller has been developed as a social network extraction, manipulation and analysis tool to cover the lack that exists in other tools. It is capable of constructing social networks from raw data by employing a variety of data mining and machine learning techniques. In this paper, we describe an extend version of NetDriller, which has some new essential functions, including social network construction using data collection from Twitter, DBLP and IEEE. We also added (1) a new chart for viewing the network property and metrics, and (2) new graph manipulation techniques using GUI to keep the tool up to date with the huge volume of networks and the different types of raw data available on the web. Salim Afra, Tansel Özyer, Jon G. Rokne, Reda Alhajj |
ASONAM | 4 |
| 2022 | Classes versus Communities: Outlier Detection and Removal in Tabular Datasets via Social Network Analysis (ClaCO)abstractIn this research, we introduce a model to detect inconsistent & anomalous samples in tabular labeled datasets which are used in machine learning classification tasks, frequently. Our model, abbreviated as the ClaCO (Classes vs. Communities: SNA for Outlier Detection), first converts tabular data with labels into an attributed and labeled undirected network graph. Following the enrichment of the graph, it analyses the edge structure of the individual egonets, in terms of the class and community belongings, by introducing a new SNA metric named as ‘the Consistency Score of a Node - CSoN’. Through an exhaustive analysis of the ego network of a node, CSoN tries to exhibit consistency of a node by examining the similarity of its immediate neighbors in terms of shared class and/or shared community belongings. To prove the efficiency of the proposed ClaCO, we employed it as a subsidiary method for detecting anomalous samples in the train part in the traditional ML classification task. With the help of this new consistency score, the least CSoN scored set of nodes flagged as outliers and removed from the training dataset, and remaining part fed into the ML model to see the effect on classification performance with the ‘whole’ dataset through competing outlier detection methods. We have shown this outlier detection model as an efficient method since it improves classification performance both on the whole dataset and reduced datasets with competing outlier detection methods, over several known both real-life and synthetic datasets. Serkan Üçer, Tansel Özyer, Reda Alhajj |
ASONAM | 3 |
| 2021 | Predictions of drug metabolism pathways through CYP 3A4 enzyme by analysing drug-target interactions network graphabstractThe available data of drugs and their targets has increased widely in recent years. Far from the traditional way of studying the drug-target interactions, we propose a network-based computational method to identify new targets for known drugs. In this study, the Stanford Biomedical Network Dataset Collection (BIOSNAP Datasets) is used. A network graph is constructed and analyzed to study the relationship between the drugs and their targets. Different centrality and similarity measures analyses are applied and predict new potential metabolism pathways for five drugs, namely (Wortmannin, Voacamine, Vancomycin, Dactinomycin and Arundic acid) through Cytochrome P450 3A4 enzyme in the liver. The application of network theory to the analysis of this dataset reveals a new significant approach. Finally the molecular docking is performed to confirm the results. Also, the importance of the presented method in drug discovery is highlighted/pointed out. M. Taleb Albrijawi, Amrou Haj Ibrahim, Reda Alhajj |
ASONAM | 3 |
| 2021 | Detecting spam tweets using machine learning and effective preprocessingabstractNowadays, with the rapid increase in popularity of online social networks (OSNs), these platforms are realized as ideal places for spammers. Unfortunately, these spammers can easily publish malicious content, advertise phishing scams by taking advantage of OSNs. Therefore, effective identification and filtering of spam tweets will be beneficial to both OSNs and users. However, it is becoming increasingly difficult to check and eliminate spam tweets due to this great flow of posts. Motivated by these observations, in this paper we propose an approach for the detection of spam tweets using machine learning and effective preprocessing techniques. The approach proposes the advantages of the preprocessing and which of these preprocessing techniques are the most effective. To compare these techniques UtkML Twitter spam dataset is used in testing. After the most effective methods determined, the detection accuracy of the spam tweets will be better optimized by combining them. We have evaluated our solution with four different machine learning algorithms namely - Naïve Bayes Classifier, Neural Network, Logistic Regression and Support Vector Machine. With SVM Classifier, we are able to achieve an accuracy of 93.02%. Experimental results show that our approach can improve the performance of spam tweet classification effectively. Berk Kardas, Ismail Erdem Bayar, Tansel Özyer, Reda Alhajj |
ASONAM | 4 |
| 2021 | Automation of active reconnaissance phase: an automated API-based port and vulnerability scannerabstractThe unprecedented growth in technology has increased the importance of the required information security that is still hard to be reached. Recently, network and web application attacks have occurred frequently, causing confidential data to be stolen by the available vulnerabilities in the systems and the most prominent is in the form of open ports. This causes the CIA (Confidentiality Integrity and Availability) Triad Model to break. Penetration testing is one of the key techniques used in real life to accurately detect the possible threats and potential attacks against the system, and the first step for hackers to conduct attacks is information collection. In this paper, we present a useful schema for the active information-gathering phase that can be used during penetration testing and by system administrators. It will be the first feature of a security engine going to be implemented. The work involves an automated API-based IP and port scanner, service-version enumerator, and vulnerability detection system. This scheme is based on the Network Mapper (Nmap) to collect the information with high accuracy depending on the provided rules in our schema. Besides, the work has been implemented as a RESTful-API server, aiming at easy integration for real-life cases and allowing administrators to scan and secure their networks more quickly and easily. The effectiveness and efficiency of this technique has been proved by the various test cases applied considering different scenarios from the real world. The average time of scanning a server and detecting the vulnerabilities is 2.2 minutes. Regardless of the number of vulnerabilities, the increase in time for each open port is just about 12 seconds. Malek Malkawi, Tansel Özyer, Reda Alhajj |
ASONAM | 3 |
| 2021 | Hot topic detection and evaluation of multi-relation effectsabstractWith the growth of social media, Twitter has become one of the most popularly used microblogging communication platforms between people. Due to the wide preference of Twitter, popular issues in public, events like local or global news and daily life stories can immediately publish on Twitter. Thus, a substantial number of hot topics are created by Twitter users in real-time. These topics can exhibit every incident of everyday life. Therefore, detection of hot topics can be used in many applications such as observing public judgment, product recommendation, and incidence detection. In this paper, we propose a method for detecting Twitter hot topics and evaluate the effect of multi-relations such as retweets and hashtags on hot topics. The dataset was generated by fetching tweets for a certain time and location by using GetOldTweets3 API. Then using the LDA topic modeling algorithm the hot topics were identified for each multi relation. Finally, the effect of each relation is described by using the coherence scores) Nadir Emre Zirbilek, Mustafa Erakin, Tansel Özyer, Reda Alhajj |
ASONAM | 4 |
| 2020 | Recent Trends in Emotion Analysis: A Big Data Analysis PerspectiveabstractHuman action recognition has recently started to find its way into applications in different applications. Accordingly, human action recognition methods are becoming increasingly important in our daily life. They are used for different purposes such as automation, security, surveillance, health, smart home systems, and customer behaviour prediction, among others. Though have more systems with methods provides a rich pool of choices, it is important to well understand the performance of these systems and their success rates in recognizing the right activities in order to decide on the most appropriate system for the current application domain. This survey tackles this issue by analyzing and commenting on the available human action recognition systems and methods. Tansel Özyer, Ak Duygu Selin, Reda Alhajj |
ASONAM | 3 |
| 2019 | A model based on random walk with restart to predict circRNA-disease associations on heterogeneous networkabstractRecent studies show that circRNAs have critical roles in many biological processes. Knowing the associations between circRNAs and diseases may contribute to the understanding of the mechanism of circRNAs and to the diagnostic and therapeutic methods of diseases at the molecular level. A small number of computation models have been developed to estimate CircRNA-disease associations. Therefore, in this study, a computational model has been developed. Similarity matrices have been obtained for circRNA and disease respectively by applying gaussian on the data obtained from the circRNADisease database. Then, random walk with restart algorithm applied on the combined matrices. The AUC value was obtained by 5-fold cross validation is 0.861 and this demonstrates the reliability of the model. Hüseyin Vural, Mehmet Kaya, Reda Alhajj |
ASONAM | 3 |
| 2019 | Multivariate motif detection in local weather big dataabstractIn recent years, there are very frequent reports of disasters attributed to the climate change and there are several reports that these extreme phenomena will further affect people not only as weather disasters but also indirectly with the shortage of natural resources such as water or food due to the climate change. Towards this direction, there is an on-going research that studies weather phenomena by collecting data not only in the surface of the globe but also at the different levels of the atmosphere. Having such a large volume of data, traditional numerical weather prediction models may not be able to assimilate those data and extract knowledge useful for the prediction of extreme phenomena. Thus, analysis of weather data has been transformed into a big data analytics problem which may enable weather scientists to better understand the interrelations of the weather variables and use the knowledge discovered to improve their prediction models. In this context, the current paper proposes a big data analytics methodology that is able to detect all common patterns between different weather variables in neighboring or distant points in a specific time window revealing useful associations between weather variables which is not possible to detect otherwise with the traditional numerical methods. The proposed methodology is based on a data structure that is able to store the magnitude of the weather data in different dimensions and a pattern detection algorithm which is able to detect all common patterns. The experimental results using weather data from the National Oceanic and Atmospheric Administration (NOAA) revealed interesting otherwise unknown patterns in two weather variables for two specific locations that were studied. Konstantinos F. Xylogiannopoulos, Panagiotis Karampelas, Reda Alhajj |
ASONAM | 3 |
| 2019 | Text mining for malware classification using multivariate all repeated patterns detectionabstractMobile phones have become nowadays a commodity to the majority of people. Using them, people are able to access the world of Internet and connect with their friends, their colleagues at work or even unknown people with common interests. This proliferation of the mobile devices has also been seen as an opportunity for the cyber criminals to deceive smartphone users and steel their money directly or indirectly, respectively, by accessing their bank accounts through the smartphones or by blackmailing them or selling their private data such as photos, credit card data, etc. to third parties. This is usually achieved by installing malware to smartphones masking their malevolent payload as a legitimate application and advertise it to the users with the hope that mobile users will install it in their devices. Thus, any existing application can easily be modified by integrating a malware and then presented it as a legitimate one. In response to this, scientists have proposed a number of malware detection and classification methods using a variety of techniques. Even though, several of them achieve relatively high precision in malware classification, there is still space for improvement. In this paper, we propose a text mining all repeated pattern detection method which uses the decompiled files of an application in order to classify a suspicious application into one of the known malware families. Based on the experimental results using a real malware dataset, the methodology tries to correctly classify (without any misclassification) all randomly selected malware applications of 3 categories with 3 different families each. Konstantinos F. Xylogiannopoulos, Panagiotis Karampelas, Reda Alhajj |
ASONAM | 3 |
| 2018 | A Paper Recommendation System Based on User's Research InterestsabstractResearchers and scientists read articles to improve their studies. Researchers spend too much time and struggle to find the suitable article they are looking for. The purpose of article recommendation system is to reduce the time they spend and present to them the related articles they are not aware of. Classic article recommendation systems do not consider the user's information, they show the same results in the same sort for each researcher. In this study, an article recommendation system that takes into consideration the researcher's work field and the publisher's previous articles is presented. One of the most important innovations of this work is the use of TF-IDF and Cosinus similarity to make article recommendations taking user's past articles into consideration. As a result of the work, the users have been recommended articles, and the method we present has proved more successful results compared to equivalent methods according to f-Measure criterion. Betul Bulut, Buket Kaya, Reda Alhajj, Mehmet Kaya |
ASONAM | 3 |
| 2018 | Prediction of New Potential Micro RNAs-Environmental Factor Associations Based on KATZ MeasureabstractMicroRNAs (miRNAs) have effects on regulation of gene expressions and they also have various functions on biological processes, so the disruption of functions of miRNAs can cause to diseases and abnormal phenotypes. Environmental Factors (EFs) such as drugs, radiation, cigarette smoke, alcohol have important negative effects on human health because they perform interactions at molecular level in human organisms. miRNA is one of the molecules interacting with EFs, and the interactions between them have affects on diseases. In this study, we have developed a model based on KATZ measure to find new potential associations between miRNAs and EFs with using Gaussian interaction profile kernel similarity. Hüseyin Vural, Buket Kaya, Reda Alhajj, Mehmet Kaya |
ASONAM | 3 |
| 2018 | Clickstream Analytics: An Experimental Analysis of the Amazon Users' Simulated Monthly TrafficabstractOnline shopping in recent years demonstrate a constant increase and as a result the study of user behavior through clickstream has attracted again the interest of the research community. This increase though requires novel approaches to clickstream analytics since the volume of the products available online and the corresponding transactions is huge. In this paper, a sequential frequent itemsets detection methodology (SAFID) is adopted to solve a clickstream analytics problem by analyzing a composite dataset which simulates monthly traffic of Amazon U.S. online retail shop. It is shown that the methodology can perform the analysis very efficiently in a simple desktop and detect all the frequently bought together products which can provide valuable knowledge to marketers of online retail stores. The methodology used can further be improved to handle larger datasets by considering a cloud computing environment. Konstantinos F. Xylogiannopoulos, Panagiotis Karampelas, Reda Alhajj |
ASONAM | 3 |
| 2018 | Text Mining for Plagiarism Detection: Multivariate Pattern Detection for Recognition of Text SimilaritiesabstractThe problem of plagiarism the recent years has been intensified by the availability of information in digital form and the accessibility of the electronic libraries through the Internet. As a result, plagiarism detection has been transformed into a big data analytics problem since the number of digital sources is extravagant and a new document needs to be compared with millions of other existing documents. In this paper, a text mining methodology is proposed that can detect all common patterns between a document and the documents in a reference database. The technique is based on a pattern detection algorithm and the corresponding data structure that enables the algorithm to detect all common patterns. The methodology has been applied in a well-defined dataset providing very promising results identifying difficult cases of plagiarism such as technical disguise. Konstantinos F. Xylogiannopoulos, Panagiotis Karampelas, Reda Alhajj |
ASONAM | 3 |
| 2018 | Extension of neighbor-based link prediction methods for directed, weighted and temporal social networks
Ertan Butun, Mehmet Kaya, Reda Alhajj |
Inf. Sci. | 3 |
| 2018 | Integrating flexibility and fuzziness into a question driven query model
Abdullah Sarhan, Jon G. Rokne, Reda Alhajj |
Inf. Sci. | 3 |
| 2017 | A Supervised Learning Method for Prediction Citation Count of Scientists in Citation NetworksabstractPredicting the future impacts of scientists and their publications is an important task for future scientific studies. Studies predicting scientific effects in the academia are mostly based on estimating citation numbers of publications. However, there are few studies focusing on predicting the impact of scientists. Determining the potential influence of scientists is useful for future scientific studies and organizations. In this respect, we propose a new approach for predicting future impacts of scientists by considering the number of citations that each scientist individually has. A great number of studies predicting citation count of publications use content based information such as title, abstract and keywords of publication, institution of scientists, journal impact factor of publisher, etc. and they don't benefit from graph structures in citation networks. However, citations among scientific publications create more complex networks which have the potential for mining scientific impacts. So, we formulated the problem of predicting future citation counts of a given scientist, as a link prediction problem for directed, weighted and dynamic citation networks. The proposed link prediction method is able to predict links and its weights. Firstly, we introduce a dynamic metric based on increasing and decreasing trends during transitions time frames in citation networks. Then, we use the introduced dynamic proximity metric with some simple topological features in classifiers to predict citation count. Experiments are conducted on 12 sampled datasets which are extracted from two citation networks. The experimental results show that the proposed approach performs well when considering the problem of prediction links with its weights is firstly addressed in the context of link prediction problem in directed, weighted and dynamic networks. Ertan Butun, Mehmet Kaya, Reda Alhajj |
ASONAM | 3 |
| 2017 | From Social Media Analysis to Ubiquitous Event Monitoring: The case of Turkish TweetsabstractThe work described in this paper illustrates how social media is a valuable source of data which may be processed for informative knowledge discovery which may help in better decision making. We concentrate on Twitter as the source for the data to be processed. In particular, we extracted and captured tweets written in Turkish. We analyzed tweets online and real-time to determine most recent trending events, their location and time. The outcome may help predicting next hot events to be broadcasted in the news. It may also raise alert and warn people related to upcoming or ongoing disaster or an event which should be avoided, e.g., traffic jam, terror attacks, earthquake, flood, storm, fire, etc. To achieve this, a tweet may be labeled with more than one event. Named entity recognition combined with multinomial naive Bayes and stochastic gradient descent have been integrated in the process. The reported 95% success rate demonstrate the applicability and effectiveness of the proposed approach. Ahmet Enis Erdogan, Tolga Yilmaz, Onur Can Sert, Mirun Akyüz, Tansel Özyer, Reda Alhajj |
ASONAM | 6 |
| 2017 | Effectiveness of Mobile Electrocardiogram in Healthcare: From Mobile Application and Development to Community ReactionabstractChronic diseases such as heart and blood vessels are considered among the most common and serious reasons of mortality in the world. In Europe alone, over four million deaths a year (45% of all deaths) are caused by heart diseases [1]. In addition, chronic diseases are responsible for 70 % of United States deaths, and account for more than 75% of annual United States medical care cost [2]. For instance, Cardio Vascular Diseases (CVD) are considered the main cause for around 14.3% of total deaths in Denmark5, and it is the main cause for over 45% of the total death in Lebanon6. It is too costly to keep CVD patients under control locally within the vicinity of a healthcare unit. Thus, researchers recently started to realize the need for automated monitoring in health systems that are expected to reduce the overall death rate and cost associated with monitoring of patients. However, a general monitoring health system will not cover all diseases at once. Therefore, there is a pushing necessity for monitoring health systems which are dedicated to specific health cases. To contribute to the ongoing efforts, this work develops an automated system which could be customized for various chronic diseases. A mobile application based solution is proposed. Further, the work concentrates on CVD by conducting a survey in Lebanon to investigate the acceptance and awareness of ECG for remote monitoring of patients. The results are promising and reflect how specialists are aware of the need to utilize the rapid development in technology combined with the widespread usage of mobile phone which may be used as the main device to guarantee 24/7 communication link for ECG. Adopting ECG in the healthcare system will allow for capturing some valuable data which could guide the development of a recommendation system. This will issue necessary alerts to specialists and guidance to patients and their careers so that specialists could attend to the case on timely basis and patients with their careers could follow the recommendations to keep the case under control until the specialist becomes available. Finally, a secure forum based communication system will be developed to allow patients to share their experience and specialists to provide consultancy and guidance on demand. Ahmad Kassem, Umut Ozan Yildirim, Kadir Anil Turgut, Uffe Kock Wiil, Tansel Özyer, Reda Alhajj |
ASONAM | 6 |
| 2016 | A new topological metric for link prediction in directed, weighted and temporal networksabstractOne of the most interesting tasks in social network analysis is link prediction. There are a lot of studies dealing with link prediction task in the literature. In recent years, there is an increasing on link prediction methods trying to model network as more close to real networks such as heterogeneous, temporal and directed network models to gain better link prediction performance. Many of the existing link prediction methods don't take into account links directions in directed networks. In this paper we propose a new neighbor and graph pattern based topological metric considering direction of links for link prediction. The proposed metric also takes into account temporal and weighted information, which are useful to increase link prediction performance. Accuracy of the proposed metric is evaluated by comparison with multiple baseline metrics from literature in supervised learning methods. Experimental results demonstrate that the proposed metric improves remarkably the accuracy of link prediction. Ertan Butun, Mehmet Kaya, Reda Alhajj |
ASONAM | 3 |
| 2016 | Classification of HIV data by constructing a social network with frequent itemsetsabstractAcquired immune deficiency syndrome (AIDS) is the last and the most life-threatening phase of Human Immunodeficiency Virus (HIV) disease. HIV attacks and heavily affects the immune system of the body which remains unable to resist the disease. HIV uses white blood cells to replicate itself and spreads everywhere in the body. The lifecycle of HIV disease, especially the replication stage must be prominently understood in order to develop effective drugs for treatment. HIV-1 protease enzyme is in charge of cleaving an amino acid octamer into peptides which are used to create proteins by virus. It should be scrutinized properly since it is a potential target to tightly bind drugs to protease for blocking the virus action at an early stage before cell infection. It is very critical to induce a model and predict cleavage of HIV-1 protease on octamers. Several machine learning approaches have been applied for predicting and profiling cleavage rules. However, we propose a novel general approach that can also be applied on different domains. It basically utilizes social network analysis and data mining techniques for classification. This method yet presents promising results that are comparable with existing machine learning methods, besides it gives the opportunity to validate the results obtained by using other techniques from social network analysis perspective. We have used the HIV-1 protease cleavage data set from UCI machine learning repository and demonstrated the effectiveness of our proposed method by comparing it with decision tree, Naive-Bayes and k-nearest neighbor methods. Yunuscan Kocak, Tansel Özyer, Reda Alhajj |
ASONAM | 3 |
| 2016 | Time preference aware dynamic recommendation enhanced with location, social network and temporal informationabstractSocial networks and location based social networks have many active users who provide various kind of data, such as where they have been, who their friends are, which items they like more, when they go to a venue. Location, social network and temporal information provided by them can be used by recommendation systems to give more accurate suggestions. Also, recommendation systems can provide dynamic recommendations based on the users' preferences, such that they can give different recommendations for different hours of the day or different days of the week. In this paper, we propose a recommendation system which considers the users' temporal preference to give dynamic recommendation. The recommendation method uses multi-objective optimization approach and gives point of interest (POI) recommendation using several different criteria, namely past check-in locations, hometown of users, time of check-ins, friendship and influence among users. Makbule Gulcin Ozsoy, Faruk Polat, Reda Alhajj |
ASONAM | 3 |
| 2016 | Frequent and non-frequent pattern detection in big data streams: An experimental simulation in 1 trillion data pointsabstractBig data streaming analysis nowadays has become one of the most important topic in the list of data analysts since enormous amount of data are produced daily by the numerous smart devices. The analysis of such data is very important and the detection of frequent or even non-frequent patterns can be critical for many aspects of our lives. In the current paper, we propose a new methodology based on our previous work regarding the detection of all repeated patterns in a string in order to analyze a very big data stream with 1 Trillion digits, composed from 1 thousand subsequences of 1 billion digits each one. More specifically, using the novel data structure, LERP Reduced Suffix Array, and the innovative ARPaD algorithm which allows the detection of all repeated patterns in a string we managed to analyze each one of the 1 billion data points, using 10 computers with standard hardware configuration, in 33 minutes which outperforms to the best of our knowledge any other existing methodology, which is equivalent to data point generation every 2 microseconds. Konstantinos F. Xylogiannopoulos, Reda Alhajj, Panagiotis Karampelas |
ASONAM | 2 |
| 2015 | Time Frame based Link Prediction in Directed Citation NetworksabstractLink prediction is a well-known problem in field of social network analysis which intends to guess the likelihood of the occurrences of connections between nodes. By using the structure of the network up to a given time appearance of links in future can be predict. In the most of previous studies, for performing the link prediction task just according to the exploration of the state of the network at a specific moment by applying the proximity metrics like topological based metrics to non-connected nodes has been used in order to predict new links. In those studies the behavior of links along the time or directed networks didn't considered which can be point as a limitation in link prediction studies. In this study we tried to overcome the above mentioned limitation by analyzing the development of topological measures in a citation network on a specific pried of time. For achieving this aim, chosen similarity matric deployed to all non-connected pairs of nodes in different frames of time in the network. Then, time frames are built for each pair to record their values which provided by the metric. Experiments on unsupervised prediction on a directed citation network show that the proposed method finds satisfactory results and is promising. Mujtaba Jawed, Mehmet Kaya, Reda Alhajj |
ASONAM | 3 |
| 2015 | Tactics, weapons, targets and rationale behind the actions of the mostly operational terrorist groups across EuropeabstractThis paper discusses the practices employed by various terrorist groups that operated in European countries between the years of 1968 and 2009. We focus on the deployment of the terrorist operations as presented in the RAND Database of Worldwide Terrorism Activities. In this context we elaborate on the tactics of the terrorist groups with the highest frequency of actions that operated in the European countries showing the highest rate of terrorist activity. Their targets, the weapons used and the consequences suffered as a result of their actions (both fatalities and injuries intended for the original targets, as well as any kind of collateral damage caused to third parties) are analyzed, in order to evaluate their ideological and -- perhaps - ethical standing. In particular, we look at the groups' targets as well as the tactics they used to achieve them, in a bid to explore whether there is a correlation between targeting of specific people or groups of people or other types of targets with certain international events - and if so, how these events influenced the actions of the terrorists. Within this line of thought, we also provide an outline of the political and ideological framework of the groups on focus in an effort to place them within the general historical and political context during their operational years. This is of great importance, as it enables us to run a comparison between terrorist groups that operated in different countries (albeit with similar aims) both from an ideological and operational viewpoint. Ioanna K. Lekea, Panagiotis Karampelas, Konstantinos F. Xylogiannopoulos, Reda Alhajj |
ASONAM | 4 |
| 2015 | Modeling Individuals and Making Recommendations Using Multiple Social NetworksabstractWeb-based platforms, such as social networks, review web-sites, and e-commerce web-sites, commonly use recommendation systems to serve their users. The common practice is to have each platform captures and maintains data related to its own users. Later the data is analyzed to produce user specific recommendations. We argue that recommendations could be enriched by considering data consolidated from multiple sources instead of limiting the analysis to data captured from a single source. Integrating data from multiple sources is analogous to watching the behavior and preferences of each user on multiple platforms instead of a limited one platform based vision. Motivated by this, we developed a recommendation framework which utilizes user specific data collected from multiple platforms. To the best of our knowledge, this is the first work aiming to make recommendations by consulting multiple social networks to produce a rich modeling of user behavior. For this purpose, we collected and anonymized a specific dataset that contains information from BlogCatalog, Twitter and Flickr web-sites. We implemented several different types of recommendation methodologies to observe their performances while using single versus multiple features from a single source versus multiple sources. The conducted experiments showed that using multiple features from multiple social networks produces a wider perspective of user behavior and preferences leading to improved recommendation outcome. Makbule Gulcin Ozsoy, Faruk Polat, Reda Alhajj |
ASONAM | 3 |
| 2015 | Sequential All Frequent Itemsets Detection: A Method to Detect All Frequent Sequential Itemsets Using LERP-Reduced Suffix Array Data Structure and ARPaD AlgorithmabstractSequential frequent itemsets detection is one of the core problems in data mining. In the current paper we propose a new methodology based on our previous work regarding the detection of all repeated patterns in a string. By analyzing big datasets from FIMI website of up to one million transactions we were able to detect not only the most frequent sequential itemsets but any sequential itemset occurred at least twice in the transactions' database. For this purpose we have used a novel data structure the LERP Reduced Suffix Array and the innovative ARPaD algorithm which allows the detection of all repeated patterns in a string. The methodology uses a pre-statistical analysis of the transactions that allows constructing in a very efficient way smaller LERP-RSA data structures for each transaction. The integration and classification of all LERP-RSAs let ARPaD algorithm to be executed in parallel and to detect every sequential itemset that occurs at least twice in a very efficient way. Konstantinos F. Xylogiannopoulos, Panagiotis Karampelas, Reda Alhajj |
ASONAM | 3 |
| 2015 | On personalizing Web search using social network analysis
M. Omair Shafiq, Reda Alhajj, Jon G. Rokne |
Inf. Sci. | 2 |
| 2014 | Simpler is better? Lexicon-based ensemble sentiment classification beats supervised methodsabstractIt has been shown in this paper that simplistic Bag of Words (BoW) lexicon methods for sentiment polarity assignment with ensemble classifiers are much faster than a supervised approach to sentiment classification while yielding similar accuracy. BoW methods also proved to be efficient and fast across all examined datasets. Moreover, a new approach to lexicon extraction that can be successfully used for sentiment polarity assignment is presented in the paper. It has been shown that accuracy obtained from such lexicons outperforms other lexicon based approaches. Lukasz Augustyniak, Tomasz Kajdanowicz, Piotr Szymanski, Wlodzimierz Tuliglowicz, Przemyslaw Kazienko, Reda Alhajj, Boleslaw K. Szymanski |
ASONAM | 6 |
| 2014 | Handling incomplete data using semantic logging based Social Network Analysis Hexagon for effective application monitoring and managementabstractMonitoring and management of large scale applications is already a complex task because of syntactic and unstructured nature of execution data. Traditional application monitoring and management solutions focused on employing analysis techniques on unstructured and syntactic log information become limited as unstructured information cannot be well utilized to find out related events information or correlate such information with other related information from applications. Our proposed solution of semantically formalized logging fills this gap by bringing formal semantics and combining it in a meaningful way to enable automated monitoring and management of applications. Such formalized and well-structured log information helps analytical solution to maximally automate the process of monitoring and management of applications. However, while formalizing and structuring the log information, we came across several missing and incomplete data which causes hindrance in this process. In this paper, we tackle this problem and propose a social network analysis based solution to handle incomplete and missing data from application execution, possibly compute it and use it by our proposed solution of semantically formalizing and structured logs with adapted data mining techniques to enable automated and effective application monitoring and management. We demonstrate from an industrial use-case application that how historical data from application execution is stored using semantic logging and utilized with standard social-network analysis techniques to find out missing values in incomplete data and perform application monitoring and management. M. Omair Shafiq, Reda Alhajj, Jon G. Rokne |
ASONAM | 2 |
| 2014 | Development of multidimensional academic information networks with a novel data cube based modeling method
Mehmet Kaya, Reda Alhajj |
Inf. Sci. | 2 |
| 2013 | Combining information extraction and text mining for cancer biomarker detectionabstractInformation technology is advancing faster than anticipated. The amount of data captured and stored in electronic form by far exceeds the capabilities available for comprehensive analysis and effective knowledge discovery. There is always a need for new sophisticated techniques that could extract more of the knowledge hidden in the raw data collected continuously in huge repositories. Biomedicine and computational biology is one of the domains overwhelmed with huge amounts of data that should be carefully analyzed for valuable knowledge that may help uncovering many of the still unknown information related to various diseases threatening the human body. Biomarker detection is one of the areas which have received considerable attention in the research community. There are two sources of data that could be analyzed for biomarker detection, namely gene expression data and the rich literature related to the domain. Our research group has reported achievements analyzing both domains. In this paper, we concentrate on the latter domain by describing a powerful tool which is capable of extracting from the content of a repository (like PubMed) the parts related to a given specific domain like cancer, analyze the retrieved text to extract the key terms with high frequency, present the extracted terms to domain experts for selecting those most relevant to the investigated domain, retrieve from the analyzed text molecules related to the domain by considering the relevant terms, derive the network which will be analyzed to identify potential biomarkers. For the work described in this paper, we considered PubMed and extracted abstracts related to prostate and breast cancer. The reported results are promising; they demonstrate the effectiveness and applicability of the proposed approach. Khaled Dawoud, Shang Gao 0005, Ala Qabaja, Panagiotis Karampelas, Reda Alhajj |
ASONAM | 5 |
| 2013 | Analyzing the scalability of a social network of agentsabstractSocial networks are ever-growing systems by inheritance. The increase in the number nodes in these systems often brings forth the need to add additional functionalities. However due to the distributed nature of social networks, system growth can be a challenging task. Therefore scalability of the system is of vital importance in the design of social networks. This research attempts to establish a comprehensive framework for analysis and validation of requirements and design documents for software systems. In previous work, we applied this framework to analyze the requirements of a social network of agents; expressed using scenario-based specifications. Scenarios are appealing because of their expressive power and simplicity. Moreover due to the clear and concise notation of scenarios, they can be used to analyze the system requirements for general validity, lack of deadlock, and existence of emergent behavior. In this paper a methodology to analyze the scalability of social networks is presented. This methodology is devised to indicate whether or not the new requirements of the system are consistent with the current requirements in place. A larger prototype of a social network of MSA for semantic search is utilized to illustrate the developed methodology. Mohammad Moshirpour, Shimaa M. El-Sherif, Reda Alhajj, Behrouz Homayoun Far |
ASONAM | 3 |
| 2013 | Effectiveness of template detection on noise reduction and websites summarization
Derar Alassi, Reda Alhajj |
Inf. Sci. | 2 |
| 2013 | Out-of-core detection of periodicity from sequence databases
Faraz Rasheed, Muhaimenul Adnan, Reda Alhajj |
Knowl. Inf. Syst. | 3 |
| 2012 | Stock Market Investment Advice: A Social Network ApproachabstractMaking investment decision on various available stocks in the market is a challenging task. Econometric and statistical models, as well as machine learning and data mining techniques, have proposed heuristic based solutions with limited long-range success. In practice, the capabilities and intelligence of financial experts is required to build a managed portfolio of stocks. However, for non-professional investors, it is too complicated to make subjective judgments on available stocks and thus they might be interested to follow an expert's investment decision. For this purpose, it is critical to find an expert with similar investment preferences. In this work, we propose to benefit from the power of Social Network Analysis in this domain. We first build a social network of financial experts based on their publicly available portfolios. This social network is then used for further analysis to recommend an appropriate managed portfolio to non-professional investors based on their behavioral similarities to the expert investors. This approach is evaluated through a case study on real portfolios. The result shows that the proposed portfolio recommendation approach works well in terms of Sharpe ratio as the portfolio performance metric. Negar Koochakzadeh, Keivan Kianmehr, Atieh Sarraf, Reda Alhajj |
ASONAM | 4 |
| 2012 | Developing an Efficient Health Clinical Application: IIOP Distributed Objects FrameworkabstractThe Middleware is a piece of software lying between the operating system and the application layer. Distributed applications are gaining popularity with the widespread of reliable communication services. It is affordable to have data accessible 24/7 from almost anywhere, thanks to the well-developed mobile technology and handheld devices. Healthcare domain is one of the very demanding application areas that highly benefits from these developments. Actually, health clinic system is an evolving and promising area in which clients can easily log into a distant health clinic server and retrieve, add, or update the data and diagnosis of the patients using Common Object Request Broker Architecture (CORBA) Internet Inter-Orb Protocol (IIOP) middleware. In this paper, we describe the development and implementation of a middleware model that utilizes the effective connection between two different programming languages (java and .Net) to send and receive requested patients' data in a very efficient response time. This is achieved by using the reference of the object in the server. Based on our knowledge, our approach has the best response time compared to the existing works in the area. Our approach has been successfully tested and evaluated. Ayman N. Murshed, Wadhah Almansoori, Konstantinos F. Xylogiannopoulos, Mohamad Elzohbi, Reda Alhajj, Jon G. Rokne |
ASONAM | 5 |
| 2012 | AskFuzzy: Attractive Visual Fuzzy Query BuilderabstractThe user-centric query interface is very common application that allows expressing both the input and the output using fuzzy terms. This is becoming a need in the evolving internet-based era where web-based applications are very common and the number of users accessing structured databases is increasing rapidly. Restricting the user group to only experts in query coding must be avoided. The Ask Fuzzy system has been developed to address this vital issue which has social and industrial impact. It is an attractive and friendly visual user interface that facilitates expressing queries using both fuzziness and traditional methods. The fuzziness is not expressed explicitly inside the database, it is rather absorbed and effectively handled by an intermediate layer which is cleverly incorporated between the front-end visual user-interface and the back-end database. Keivan Kianmehr, Negar Koochakzadeh, Reda Alhajj |
ICDE | 3 |
| 2012 | TempoXML: Nested bitemporal relationship modeling and conversion tool for fuzzy XML
Ömer Özgün Isikman, Tansel Özyer, Omar Zarour, Reda Alhajj, Faruk Polat |
Inf. Sci. | 4 |
| 2012 | Effectiveness of NAQ-tree in handling reverse nearest-neighbor queries in high-dimensional metric space
Reda Alhajj |
Knowl. Inf. Syst. | 2 |
| 2011 | Simple and effective behavior tracking by post processing of association rules into segmentsabstractFrequent pattern mining and consequently association rule mining is a useful technique for discovering relationships between items in databases. However, as the size of the data to be analyzed increases or the values of the pruning thresholds decrease, larger number of frequent pattern and more association rules will be generated with little information about the association rules in relation to each other. This research paper discusses a method to segment rules into different sets with no internal conflicts. The goal is to establish an effective method to reduce the difficulty for businesses to review the association rules of different customer segments, and track the behaviors of market segments based on their buying behaviors. The method established in this paper has the advantage of not needing customer information, thus removing the need for businesses to obtain customer information. This removes the threat of intrusions into customer privacy. The method also generates the rule sets based on conflicting rules, and dividing rules based on customer behaviors is more accurate than customer characteristics. The proposed method has been validated by running some tests. Alan Chia-Lung Chen, Konstantinos F. Xylogiannopoulos, Tamer N. Jarada, Omar Zarour, Panagiotis Karampelas, Jon G. Rokne, Reda Alhajj |
iiWAS | 7 |
| 2011 | Semantically enhanced matchmaking of consumers and providers: a Canadian real estate case studyabstractMatchmaking services connecting consumers and providers on the internet have become phenomenally important in today's world. Whereas the matchmaking was by traditional media such as print and television in the past it is now expected to include the internet. Consumers expect these services to be readily available as this can only be accomplished via the internet Providers that do not have an internet presence are therefore severely disadvantage in the competition for customers (consumers). The decline of the traditional physical music store and the ascendancy of the virtual iTunes store is a perfect example of this. As the traditional stores go bankrupt and go out of business has also become the largest music retailer in the United States of America. Consumers simply do not want to spend extra energy to get what they want. If the music can be purchased and downloaded from the comfort of their own home, the consumers (customers) will do just that. Therefore online matchmaking services are a hot topic of discussion for many companies, as they are finding ways to provide the fastest, cheapest and most efficient methods for consumers to easily access their services. Freddy Poon, Thomas Chin, Matt Bentrovato, M. Omair Shafiq, Alan Chia-Lung Chen, Flouris Triant, Jon G. Rokne, Reda Alhajj |
iiWAS | 8 |
| 2011 | Effective monitoring by efficient fingerprint matching using a forest of NAQ-trees
Keivan Kianmehr, Reda Alhajj |
J. Intell. Inf. Syst. | 3 |
| 2011 | Efficient Periodicity Mining in Time Series Databases Using Suffix TreesabstractPeriodic pattern mining or periodicity detection has a number of applications, such as prediction, forecasting, detection of unusual activities, etc. The problem is not trivial because the data to be analyzed are mostly noisy and different periodicity types (namely symbol, sequence, and segment) are to be investigated. Accordingly, we argue that there is a need for a comprehensive approach capable of analyzing the whole time series or in a subsection of it to effectively handle different types of noise (to a certain degree) and at the same time is able to detect different types of periodic patterns; combining these under one umbrella is by itself a challenge. In this paper, we present an algorithm which can detect symbol, sequence (partial), and segment (full cycle) periodicity in time series. The algorithm uses suffix tree as the underlying data structure; this allows us to design the algorithm such that its worstcase complexity is O(k.n2), where k is the maximum length of periodic pattern and n is the length of the analyzed portion (whole or subsection) of the time series. The algorithm is noise resilient; it has been successfully demonstrated to work with replacement, insertion, deletion, or a mixture of these types of noise. We have tested the proposed algorithm on both synthetic and real data from different domains, including protein sequences. The conducted comparative study demonstrate the applicability and effectiveness of the proposed algorithm; it is generally more time-efficient and noise-resilient than existing algorithms. Faraz Rasheed, Mohammed Al-Shalalfa, Reda Alhajj |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2010 | A Global Measure for Estimating the Degree of Organization of Terrorist NetworksabstractThe motivation for the study described in this paper is realizing the fact that organizational structure of a group is a key indicator in determining its strengths and weaknesses. A general knowledge of the prevalent models of terrorist organizations leads to a better understanding of their capabilities. Knowledge of the different labels and systems of classification that have been applied to groups and individuals aid us in discarding useless or irrelevant terms, and in understanding the purposes and usefulness of different terminologies. Previous studies in network analysis have mostly dealt with legal networks with transparent structures. Terrorist networks share some features with conventional (real world) networks, but they are harder to identify because they mostly hide their illicit activities. In this paper we describe a novel approach for extracting structural patterns of terrorist networks with the help of social network analysis measures and techniques. We propose a global measure for estimating the degree of organization of social networks; the measure is global in terms of being applied to the whole network as an entity and being extracted from the major well-known SNA measures. The importance of such research comes from the fact that individuals in organized intellectual networks and especially terrorist networks tend to hide their individual rules and thus there is a need to deal with such networks as a whole, discovering the degree of organization and thus its strengths and weaknesses. Khaled Dawoud, Reda Alhajj, Jon G. Rokne |
ASONAM | 2 |
| 2010 | Community Aware Personalized Web SearchabstractSearching for the right information over the Web is not straight-forward. In the era of high speed internet, high capacity networks, and interactive Web applications, it has become even easier for the users to publish data online. A huge amount of data is published over the internet; every data is in the form of web pages, news, blogs and other material, etc. Similarly, for search engines like Google and Yahoo, it becomes rather hard to find out the right information, i.e., as per user's preferences; search results for same query differ in priority for different users. In this paper, we proposed a way to prioritize search results of search engines like Google, based on the personal interests and context of users. In order to find out personal interest and context, we follow a unique approach of (1) finding out activities of a user of his/her social-network, (2) finding out what information does the social networks (i.e., friends and community) provide to the user. Based on this information, we have developed a methodology that takes into account the information about social networks and prioritize search results from Web search engine. M. Omair Shafiq, Reda Alhajj, Jon G. Rokne |
ASONAM | 2 |
| 2010 | Mapping rules for converting from ODL to XML schemasabstractThis paper presents a comprehensive approach for the transformation of ODL Schemas into XML Schemas. The approach starts with an incomplete set of rules described in the literature to assist in the transformation process. The fact that the rules provided a solid foundation for expansion, as well as the fact that the rules only cover a small subset of ODL, was our main motivation for continuing the study of this topic. In this paper, we first analyze an existing set of nine transformation rules. After evaluating the correctness and completeness of the rules, we proceed to propose some improvements and extensions into a more complete set of rules that cover the whole transformation process. By modifying the existing rule set, we are able to handle a much wider variety of ODL. Finally, we discuss some ODL scenarios that the original rule set cannot handle. This is meant to justify the need for the proposed extension as described in this paper. The presented more complete rule set is capable of handling a larger subset of ODL (including dictionaries, global and local scope enumerations, and most importantly, inheritance). Tamer N. Jarada, Kelvin Chung, Armen Shimoon, Panagiotis Karampelas, Reda Alhajj, Jon G. Rokne |
iiWAS | 5 |
| 2010 | A novel client-based approach for signing and checking web forms by using XML against DoS attacksabstractIn parallel to rapid growth of internet technologies, security becomes more critical in various real life applications such as e-finance, e-health, and e-government. These applications strictly require data authentication mechanisms. To address this essential issue, we grasp the idea of client based authenticity for interactive web technologies. We proposed a novel client based web form signing and checking with XML data structure method. Our method specifically uses XML structure for the involvement of data exchange between web applications. Our method curbs the DoS (Denial of Service) attacks for protection of the server. In order to illustrate our ideas, we adapted our digital signature mechanism on health related forms with two commonly used web browsers. Kaziim Sarikaya, Duygu Sarikaya, Tamer N. Jarada, Shang Gao 0005, Tansel Özyer, Reda Alhajj |
iiWAS | 6 |
| 2010 | Skyline queries with constraints: Integrating skyline and traditional query operators
Reda Alhajj |
Data Knowl. Eng. | 2 |
| 2010 | XML materialized views and schema evolution in VIREX
Anthony Chiu Wa Lo, Tansel Özyer, Radwan Tahboub, Keivan Kianmehr, Jamal Jida, Reda Alhajj |
Inf. Sci. | 6 |
| 2010 | VIREX and VRXQuery: interactive approach for visual querying of relational databases to produce XML
Anthony Chiu Wa Lo, Tansel Özyer, Keivan Kianmehr, Reda Alhajj |
J. Intell. Inf. Syst. | 4 |
| 2010 | Fuzzy clustering-based discretization for gene expression classification
Keivan Kianmehr, Mohammed Al-Shalalfa, Reda Alhajj |
Knowl. Inf. Syst. | 3 |
| 2010 | Effectiveness of NAQ-tree as index structure for similarity search in high-dimensional metric space
Reda Alhajj |
Knowl. Inf. Syst. | 2 |
| 2009 | The Economic Benefts of Web MiningabstractIn this paper, we investigate and explore the process of analyzing log data of website visitor traffic in order to assist the owner of a website in understanding the behavior of its visitors. The developed approach involves the review of statistical data on the types of visitors that come to the website, as well as the steps they take to reach and satisfy the goal of their visit. The value added from the analysis of this data provides eCommerce and commercial website owners with the information needed to display targeted advertisements or messages to their customers. In the long run, this is expected to allow for the increase in sales and overall customer loyalty. David Kinzel, Micah Klettke, Paul Uppal, Naheed Visram, Keivan Kianmehr, Reda Alhajj, Jon G. Rokne |
ASONAM | 6 |
| 2009 | Text summarization techniques: SVM versus neural networksabstractAutomated text summarization is important to for humans to better manage the massive information explosion. Several machine learning approaches could be successfully used to handle the problem. This paper reports the results of our study to compare the performance between neural networks and support vector machines for text summarization. Both models have the ability to discover non-linear data and are effective model when dealing with large datasets. Keivan Kianmehr, Shang Gao 0005, Jawad Al Attari, M. Mushfiqur Rahman, Kofi Akomeah, Reda Alhajj, Jon G. Rokne, Ken Barker 0001 |
iiWAS | 6 |
| 2009 | Mining online shopping patterns and communitiesabstractThe great increase in online transactions and the thousands of online retailers has created a great demand for companies to gain competitive advantage. An easy way for a company to gain customer advantage is through the use of data mining. Due to this high demand we have developed a prototypical tool to help with the analysis of these online transactions. From the raw data generated by running these transactions we are able to find consumer trends and shopping patterns by using hierarchical clustering and association rules mining algorithm. The focus of this research is to demonstrate how this development can be useful and effective in a business situation for companies to gain competitive advantage. Keivan Kianmehr, X. Peng, Chris Luce, Justin Chung, Nam Pham, Walter Chung, Reda Alhajj, Jon G. Rokne, Ken Barker 0001 |
iiWAS | 7 |
| 2009 | Data governance strategy: a key issue in building Enterprise Data WarehouseabstractThis paper articulates data governance as one of the key issue in building Enterprise Data Warehouse. The key goals of this document are to: define the strategy for Data Governance processes and procedures; define the scope of and identify major components of the data governance processes; adhere to enterprise Data Management standards, principles and guidelines; and articulate a vision for building, managing and safeguarding enterprise data foundation. The client-centric focus of business organizations coupled with aggressive attention to the bottom line propelled initiatives such as Data Governance to the top of the list of IT and business executives. The recent financial crisis which spawned the worldwide economic meltdown has been to a great extent blamed on non-trustworthy and non-transparent data. It is becoming progressively and patently evident that data MUST be managed like other assets such as financial and human resources. It has to have defined and mandated set of controls where compliance can be objectively measured and reported. Mohammad Rifaie, Reda Alhajj, Mick J. Ridley |
iiWAS | 2 |
| 2008 | Optimal incremental multi-step nearest-neighbor searchabstractThe distance measures used to determine the dissimilarities between high-dimensional feature vectors are often expensive to compute. To reduce the number of expensive distance calculations in the search process, Korn, et al [5] proposed a multi-step algorithm, which involves two stages: filtering and refinement. This algorithm was later improved by Seidl and Kriegel [8] to produce optimal-sized candidate set in the filtering stage; the improved algorithm is said to be filtering optimal, but can not produce the result incrementally in the refinement stage. In this paper, we propose an extended version of the algorithm that can produce the nearest neighbors incrementally in an optimal way. Our algorithm is both filtering and refinement optimal, and well serves real applications. We proved the optimality of the proposed extended algorithm. Reda Alhajj, Jon G. Rokne |
GIS | 2 |
| 2008 | Flexible approach for representing object oriented databases in XML formatabstractIn this paper, we address the representation of object-oriented databases in XML format. This will allow for better and platform independent sharing of data stored in object-oriented format. We handle the process by first applying a reverse engineering approach to extract what we call an object-graph from the object-oriented database. Then, we convert the object-graph into XML schema; we avoid DTD representation because XML schema is more flexible and widely accepted as part of the standardization of XML. Deriving the object-graph empowers the proposed approach to smoothly consider both the inheritance and the nesting hierarchies during the mapping into XML schema. We provide a user-friendly interface that displays the result of each phase of the conversion process. Test results are encouraging, demonstrating the applicability and effectiveness of the proposed approach. Taher Naser, Reda Alhajj, Mick J. Ridley |
iiWAS | 2 |
| 2008 | Data warehouse architecture for GIS applicationsabstractGeo-data sets are built for use in geographical information systems (GIS). The data is modeled to suit the needs of data entry and visual representations. They are optimized for simplicity and speed of modification. These models do not lend themselves to efficiently produce enterprise reports. Hence, Geo-data sets can be very challenging to query and analyze. In this paper, we create a data warehouse (DW) for a geo-dataset to facilitate report generating processes. A data warehouse is attractive as the main repository of an organization's historical data and is optimized for reporting and analysis. For better understanding of the developed architecture, we cover the data warehouse construction process as well as the enterprise data model Mohammad Rifaie, Erwin J. Blas, Abdel Rahman M. Muhsen, Terrance Mok, Keivan Kianmehr, Reda Alhajj, Mick J. Ridley |
iiWAS | 6 |
| 2008 | Multi-objective genetic algorithms based automated clustering for fuzzy association rules mining
Reda Alhajj, Mehmet Kaya |
J. Intell. Inf. Syst. | 1 |
| 2007 | Enhanced Graph Based Genealogical Record Linkage
Cary Sweet, Tansel Özyer, Reda Alhajj |
ADMA | 3 |
| 2007 | Fuzzy Classifier Based Feature Reduction for Better Gene Selection
Mohammad Khabbaz, Keivan Kianmehr, Mohammed Al-Shalalfa, Reda Alhajj |
DaWaK | 4 |
| 2007 | A parallel multi-scale region outlier mining algorithm for meteorological dataabstractThe increase use of high dimensional, geographically distributed rich and massive meteorological data poses an increasing scientific challenge in efficient outlier mining. Properties in such meteorological data are observed to fluctuate in spatial synchrony. Capturing this spatial variation at different spatial scales requires a multi-resolution analysis. In this paper, we develop an algorithm for region outlier detection at different scales using the multi-resolution feature of wavelet analysis. Another challenge of meteorological data mining is that the data size is huge to accommodate different resolutions and number of samples varies with the spatial scales. This motivated us to design a load adaptive parallel algorithm for outlier detection which can maintain good scalability for all spatial scales. Our algorithm has been implemented on high-performance computing architecture and evaluated on real-world meteorological data. Sajib Barua, Reda Alhajj |
GIS | 2 |
| 2007 | Semantic Interoperability Between Relational Database SystemsabstractRelational database systems (RDBSs) are well-known and widely used in many organizations, however, semantic conflicts between the participating RDBSs must be resolved before data can be exchanged between them. Semantic resolution between the RDBSs is extremely difficult to address mainly because participating RDBSs are designed and built independently. Furthermore, individual RDBSs are likely to evolve over time and the changes must be reconciled dynamically. In this paper, we describe an approach to resolve the semantic conflicts between RDBSs automatically while allowing the individual RDBSs to evolve. Relational database ontology (RDBO) is created and used to ensure the semantic descriptions of the individual RDBSs are conformed to a set of vocabularies, structures, and restrictions. We show how a modified reasoning engine is used to validate and infer additional semantic relationships from the existing relationships. We also show how terms defined in different database ontologies are compared to each other semantically using semantic weights and our modified reasoning engine. As a result, RDBSs can intemperate with each other seamlessly and at the correct level of semantics defined in their ontologies. Quang Trinh, Ken Barker 0001, Reda Alhajj |
IDEAS | 3 |
| 2007 | Wrapping VRXQuery with Self-Adaptive Fuzzy CapabilitiesabstractThis paper addresses the development of a plug and run wrapper to incorporate fuzziness into VRXQuery, the querying facility of VIREX which is a user-friendly system for transforming and querying relational data as XML. Our basic argument is not to force the underlying XML data to incorporate fuzziness. Rather, fuzziness is smoothly supported in a novel plug and run manner via a wrapper. Either the user specifies the membership functions for the elements/attributes to be queried as fuzzy, or multi-objective genetic algorithm is used to automatically decide on and optimize the membership functions. The interface of VIREX has been expanded to allow specifying queries with fuzziness. Then, queries expressed in VRXQuery empowered with fuzziness are translated into corresponding XQuery code, which is run on the underlying XML and the returned result is translated into a fuzzy representation; translation into SQL is also possible. The user is given the choice to display the result either as colored- text or in graphical format. Anthony Chiu Wa Lo, Keivan Kianmehr, Mehmet Kaya, Tansel Özyer, Reda Alhajj |
Web Intelligence | 5 |
| 2006 | Extending OLAP with Fuzziness for Effective Mining of Fuzzy Multidimensional Weighted Association Rules
Mehmet Kaya, Reda Alhajj |
ADMA | 2 |
| 2006 | Support Vector Machine Approach for Fast Classification
Keivan Kianmehr, Reda Alhajj |
DaWaK | 2 |
| 2006 | VIREX: Interactive Approach for Database Querying and Integration by Re-engineering Relational Data into XMLabstractVIREX is a user-oriented novel approach for visual-querying of relational databases and converting the result into XML. VIREX works even when the catalogue of the relational database is missing; it extracts the required catalogue information by analyzing the database content. Further, VIREX allows the user to define views by specifying on the interactive diagram certain factors to be considered while converting relational data into XML. As a result, VIREX displays on the screen the XML schema that satisfies the specified characteristics and generates colored (easy to read) XML document(s) Anthony Chiu Wa Lo, Reda Alhajj, Ken Barker 0001 |
Web Intelligence | 2 |
| 2006 | Genetic algorithms based approach to database vertical partition
Reda Alhajj, Ken Barker 0001 |
J. Intell. Inf. Syst. | 2 |
| 2005 | Multiagent Association Rules Mining in Cooperative Learning Systems
Reda Alhajj, Mehmet Kaya |
ADMA | 1 |
| 2005 | Hybrid Approach to Web Content Outlier Mining Without Query Vector
Malik Agyemang, Ken Barker 0001, Reda Alhajj |
DaWaK | 3 |
| 2005 | Selection Mechanism for Locating Relevant Structured Databases in Multidatabase Environment
Mohammad A. Hassan, Reda Alhajj, Mick J. Ridley |
iiWAS | 2 |
| 2005 | Cluster Validity Analysis of Alternative Results from Multi-Objective OptimizationabstractThis paper investigates validity analysis of alternative clustering results obtained using the algorithm named Multi-objective K-Means Genetic Algorithm (MOKGA). The reported results are promising. MOKGA gives the optimal number of clusters as a solution set. The achieved clustering results are then analyzed and validated under several cluster validity techniques proposed in the literature. The optimal clusters are ranked for each validity index. The approach is tested by conducting experiments using three well-known data sets. The obtained results for each dataset are compared with those reported in the literature to demonstrate the applicability and effectiveness of the proposed approach. Reda Alhajj |
SDM | 1 |
| 2005 | Views as first-class citizens in object-oriented databases
Reda Alhajj, Faruk Polat, Cem Yílmaz |
VLDB J. | 1 |
| 2004 | Novel Clustering That Employs Genetic Algorithm with New Representation Scheme and Multiple Objectives
Emin Erkan Korkmaz, Reda Alhajj, Ken Barker 0001 |
DaWaK | 3 |
| 2004 | Flexible User Interface for Converting Relational Data into XML
Anthony Chiu Wa Lo, Reda Alhajj, Ken Barker 0001 |
FQAS | 2 |
| 2004 | ntegrating Multi-Objective Genetic Algorithms into Clustering for Fuzzy Association Rules MiningabstractIn this paper, we propose an automated method to decide on the number of fuzzy sets and for the autonomous mining of both fuzzy sets and fuzzy association rules. We compare the proposed multiobjective GA based approach with: 1) CURE based approach; 2) Chien et al. (2001) clustering approach. Experimental results on JOOK transactions extracted from the adult data of United States census in year 2000 show that the proposed method exhibits good performance over the other two approaches in terms of runtime, number of large itemsets and number of association rules. Mehmet Kaya, Reda Alhajj |
ICDM | 2 |
| 2003 | A Novel Approach to Separate Handwritten Connected DigitsabstractThis paper presents a novel approach to separateconnected digits in handwritten numerals by employing twoagents in the process. The first agent decides oncandidate cut-point as the closest feature-point to the center ofthe deepest top-valley, if any. The second agent arguescandidate cut-point as the closest feature-point to the center ofthe highest bottom-hill, if any. Then the actual cut-point isdecided by negotiation, which is influenced by a degree ofconfidence in each candidate cut-point. Experimentscon-ducted so far are promising and successful as well asjustified employing multiple agents. The obtained results arevery encouraging with a success factor of 97.8%. Reda Alhajj, Ashraf Elnagar |
ICDAR | 1 |
| 2003 | Integrating Fuzziness into OLAP for Multidimensional Fuzzy Association Rules MiningabstractWe contribute to the ongoing research on multidimensional online association rules mining by proposing a general architecture that utilizes a fuzzy data cube for knowledge discovery. Three different methods are introduced to mine fuzzy association rules in the constructed fuzzy data cube, namely single dimension, multidimensional and hybrid association rules mining. Experimental results obtained for each of the three methods on the adult data of the United States census in 2000 show their effectiveness and applicability. Reda Alhajj, Mehmet Kaya |
ICDM | 1 |
| 2003 | Facilitating Fuzzy Association Rules Mining by Using Multi-Objective Genetic Algorithms for Automated ClusteringabstractWe propose an automated clustering method based on multiobjective genetic algorithms (GA); the aim of this method is to automatically cluster values of a given quantitative attribute to obtain large number of large itemsets in low duration (time). We compare the proposed multi-objective GA-based approach with CURE-based approach. In addition to the autonomous specification of fuzzy sets, experimental results showed that the proposed automated clustering exhibits good performance over CURE-based approach in terms of runtime as well as the number of large itemsets and interesting association rules. Mehmet Kaya, Reda Alhajj |
ICDM | 2 |
| 2003 | Extracting the extended entity-relationship model from a legacy relational database
Reda Alhajj |
Inf. Syst. | 1 |
| 2002 | Efficient Automated Mining of Fuzzy Association Rules
Mehmet Kaya, Reda Alhajj, Faruk Polat, Ahmet Arslan 0001 |
DEXA | 2 |
| 2001 | Semantic information-based alternative plan generation for multiple query optimization
Faruk Polat, Ahmet Cosar, Reda Alhajj |
Inf. Sci. | 3 |
| 2000 | Deciding on the Equivalence of a Relational Schema and and Object-Oriented Schema
Reda Alhajj |
DEXA | 1 |
| 2000 | Simplifying the Formulation of a Wide Range of Object-Oriented Complex QueriesabstractWe present a model that simplifies the formulation of a wide range of complex, mainly selection-based, object-oriented queries, including linear recursive queries. They are complex because it is almost impossible for naive users to predict the formulation of their predicate expressions. Naive users are mainly decision makers who are most probably not computer professionals. Therefore, it is necessary to provide them with a straightforward and easy-to-handle approach to retrieve the information required for the decision making process. To achieve this, the definition of the selection operation and the predicate definition are adjusted to make it possible to have in the output only a subset of the objects from the actual result of a linear recursive query. Otherwise, it is infeasible to achieve the same output without an additional selection with a complicated predicate. We also define an operation that facilitates applying aggregate functions on objects. The two operations proved to be very useful and necessary to study the characteristics of trees and directed graphs. The presented model has been implemented as a part of our object-oriented database management system prototype. Reda Alhajj |
J. Database Manag. | 1 |
| 1999 | A Model for Deferred View MaintenanceabstractPresents a model for deferred maintenance of object-oriented views. Queries that define views are implemented as methods that are invoked to compute the corresponding views. A deferred update reflects to a view only those related modifications that were introduced into the database while that view was inactive. A view is updated by considering modifications performed within all classes along the inheritance and class-composition subhierarchies rooted at every class used in deriving that view. To each class, we add a modification list to keep one modification tuple per view dependent on that class. Such a tuple acts as a reference point that marks the start of the next update to the corresponding view. Reda Alhajj, Ashajj Elnagar |
IDEAS | 1 |
| 1999 | Incremental Materialization of Object-Oriented Views
Reda Alhajj, Ashraf Elnagar |
Data Knowl. Eng. | 1 |
| 1999 | Using Object-Oriented Materialized Views to Answer Selection-Based Complex Queries
Reda Alhajj, Faruk Polat |
Inf. Sci. | 1 |
| 1998 | Proper Handling of Query Results towards Maximizing Reusability in Object_oriented Databases
Reda Alhajj, Faruk Polat |
Inf. Sci. | 1 |
| 1996 | View Maintenance in Object-Oriented Databases
Reda Alhajj, Faruk Polat |
DEXA | 1 |
| 1994 | Closure Maintenance in An Object-Oriented Query ModelabstractAn object-algebra is presented as a formal query model for object-oriented data models. The algebra serves not only to access and manipulate the structure and behavior of objects, but it also supports the creation of new objects and the introduction of new relationships into the schema. It provides a more powerful and flexible tool than messages for effectively dealing with complex situations and meeting associative access requirements. Operands as well as the results of operations in the proposed algebra are formally characterized as a pair of sets—a set of objects capturing the states and a set of message expressions comprised of sequences of messages modeling the object behavior. The closure property is achieved in a natural way by letting the results of operations possess the same characteristics as the operands in an algebra expression. Some operators of the algebra resemble those of the relational algebra but with different syntax and semantics. Additional operators are introduced to complement them. A class is shown to posses the properties of an operand by defining a set of objects and deriving a set of message expressions for it. Furthermore, the result of an object algebra expression is shown to have the characteristics of a class whose superclass/subclass relationships with its operand class(es) can be established providing a mechanism to properly and persistently place it in the class lattice (schema). Reda Alhajj, Faruk Polat |
CIKM | 1 |
| 1993 | A Query Model for Object-Oriented DatabasesabstractA formal object-oriented query model is described in terms of an object algebra. Both the structure and the behavior of objects are handled. An operand and the output from a query in the object algebra are defined to have a pair of sets, i.e. a set of objects and a set of message expressions, where a message expression is a valid sequence of messages. The closure property is therefore maintained in a natural way. In addition, it is proven that the output from a query has the characteristics of a class; hence, the inheritance relationship between the operand and the output from a query is derived.> Reda Alhajj, M. Erol Arkun |
ICDE | 1 |
| 1992 | Queries in Object-Oriented Database Systems
Reda Alhajj, M. Erol Arkun |
CIKM | 1 |