EDBT 2026 Demo / reviewers in the wild / expert
Osmar R. Zaïane
dblp:z/OsmarRZaiane · also Osmar Zaïane
· DBLP profile ↗
113ranked-venue papers in the field
12as first author
17since 2021 · last 2026
0000-0002-0060-5988ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 75 (6 first)Database Systems & Data Management · 20 (5 first)Information Retrieval & Web Search · 9Other / Interdisciplinary · 5Big Data, Cloud & Distributed Data Systems · 2Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Finding Prominent Clusters in VAT Family of Algorithms
Kartik Vishal Deshpande, Osmar R. Zaïane |
DaWaK | 2 |
| 2025 | Structure-Aware Self-supervised Graph Representation Learning
Lingwen Liu, Peng Cao 0001, Guangqi Wen, Zhuolin Jia, Jinzhu Yang, Weiping Li 0002, Osmar R. Zaïane |
DASFAA (3) | 7 |
| 2025 | Enhancing Algorithms with LLMs: A Case Study
Yashar Talebirad, Amirhossein Nadiri, Osmar R. Zaïane, Christine Largeron |
iiWAS | 3 |
| 2024 | Integrating Conversational Pathways with a Chatbot Builder Platform
Varshini Prakash, Alex Lambe Foster, Jasmine Noble, Osmar R. Zaïane |
iiWAS (2) | 4 |
| 2024 | Capturing Temporal Node Evolution via Self-supervised Learning: A New Perspective on Dynamic Graph Learningabstract\beginabstract Dynamic graphs play an important role in many fields like social relationship analysis, recommender systems and medical science, as graphs evolve over time. It is fundamental to capture the evolution patterns for dynamic graphs. Existing works mostly focus on constraining the temporal smoothness between neighbor snapshots, however, fail to capture sharp shifts, which can be beneficial for graph dynamics embedding. To solve it, we assume the evolution of dynamic graph nodes can be split into temporal shift embedding and temporal consistency embedding. Thus, we propose the Self-supervised Temporal-aware Dynamic Graph representation Learning framework (STDGL) for disentangling the temporal shift embedding from temporal consistency embedding via a well-designed auxiliary task from the perspectives of both node local and global connectivity modeling in a self-supervised manner, further enhancing the learning of interpretable graph representations and improving the performance of various downstream tasks. Extensive experiments on link prediction, edge classification and node classification tasks demonstrate STDGL successfully learns the disentangled temporal shift and consistency representations. Furthermore, the results indicate significant improvements in our STDGL over the state-of-the-art methods, and appealing interpretability and transferability owing to the disentangled node representations. \endabstract Lingwen Liu, Guangqi Wen, Peng Cao 0001, Jinzhu Yang, Weiping Li 0002, Osmar R. Zaïane |
WSDM | 6 |
| 2024 | Uncovering Flat and Hierarchical Topics by Community Discovery on Word Co-occurrence NetworkabstractTopic modeling aims to discover latent themes in collections of text documents. It has various applications across fields such as sociology, opinion analysis, and media studies. In such areas, it is essential to have easily interpretable, diverse, and coherent topics. An efficient topic modeling technique should accurately identify flat and hierarchical topics, especially useful in disciplines where topics can be logically arranged into a tree format. In this paper, we propose Community Topic, a novel algorithm that exploits word co-occurrence networks to mine communities and produces topics. We also evaluate the proposed approach using several metrics and compare it with usual baselines, confirming its good performances. Community Topic enables quick identification of flat topics and topic hierarchy, facilitating the on-demand exploration of sub- and super-topics. It also obtains good results on datasets in different languages. Eric Austin, Shraddha Makwana, Amine Trabelsi, Christine Largeron, Osmar R. Zaïane |
Data Sci. Eng. | 5 |
| 2023 | csl-MTFL: Multi-task Feature Learning with Joint Correlation Structure Learning for Alzheimer's Disease Cognitive Performance Prediction
Peng Cao 0001, Xiaoli Liu 0001, Jinzhu Yang, Osmar R. Zaïane |
ADMA (3) | 6 |
| 2023 | Towards Time-Variant-Aware Link Prediction in Dynamic Graph Through Self-supervised Learning
Guangqi Wen, Peng Cao 0001, Zhiyong Jin, Ruoxian Song, Xiaoli Liu 0001, Jinzhu Yang, Osmar R. Zaïane |
ADMA (4) | 7 |
| 2023 | Label Correlation Guided Feature Selection for Multi-label Learning
Peng Cao 0001, Jinzhu Yang, Weiping Li 0002, Osmar R. Zaïane |
ADMA (4) | 6 |
| 2023 | CFAR++: Enhancing Rule Based ClassifierabstractOver the last few years, associative classifiers have shown massive success in mining patterns using association rules. These rule-based classifiers offer a level of human interpretability, addressing a common concern stemming from several deep learning models. Various associative classifiers have been proposed over the past that have shown state-of-the-art performance. However, those classifiers suffer the limitation of requiring parametric values which vary across different datasets. Furthermore, those frameworks do not consider the statistical significance of the rules. Recently, some works have addressed this limitation by proposing an associative classifier that incorporates the idea of using statistical significance to mine association classification rules. Though the recent associative classifiers show good performance, their performance is greatly affected by the dimension of the data. In this study, we explore the weakness of the recent associative classification models and experiment with using ensemble models to overcome such limitations, particularly on aggregating the ensemble models in a concise but effective predictor. We use 10 UCI datasets for evaluation of our new approach. From our study, we find the results based on the ensemble model with a delayed pruning are very competitive and can better handle large dimensional data spaces. Md Rayhan Kabir, Seeratpal Jaura, Osmar R. Zaïane |
ASONAM | 3 |
| 2023 | USIWO: A Local Community Search Algorithm for Uncertain GraphsabstractCommunity detection and community search are both critical tasks in graph mining, each serving unique purposes and presenting distinct challenges. The former aims to partition the graph vertices into densely connected subsets, while the latter adopts a more ego-centric approach, focusing on a specific node or group of nodes to identify a densely-connected subgraph that contains these query nodes. However, many real-world networks are characterized by uncertainty, leading to the notion of uncertain or probabilistic graphs. The transition from deterministic graphs to uncertain graphs introduces new challenges. We present USIWO, an efficient and practical solution for community search in unweighted uncertain graphs with edge uncertainty. In addition to being accurate, the approach utilizes an efficient data structure for storing only the relevant parts of the network in main memory, eliminating the need to store the entire graph, making it a valuable tool in finding the core of a community on very large uncertain graphs, when there is limited time and memory available. The algorithm operates through a one-node-expansion approach, based on the concepts of strong and weak links within a graph. Experimental results on several datasets demonstrate the algorithm's efficiency and performance. Yashar Talebirad, Mohammadmahdi Zafarmand, Osmar R. Zaïane, Christine Largeron |
ASONAM | 3 |
| 2023 | Exploring Dialog Act Recognition in Open Domain Conversational Agents
Maliha Sultana, Osmar R. Zaïane |
DaWaK | 2 |
| 2023 | Explaining Decisions of Black-Box Models Using BARBE
Mohammad H. Motallebi, Md. Tanvir Alam Anik, Osmar R. Zaïane |
DEXA (2) | 3 |
| 2022 | Dynamic Ensemble Associative LearningabstractAssociative classifiers have shown competitive performance with state-of-the-art methods for predicting class labels. In addition to accuracy performance, associative classifiers produce human readable rules for classification which provides an easier way to understand the decision process of the model. Early models of associative classifiers suffered from the limitation of selecting proper threshold values which are dataset specific. Recent work on associative classifiers eliminates that restriction by searching for statistically significant rules. However, a high dimensional feature vector in the training data impacts the performance of the model. Ensemble models like Random Forest are also very powerful tools for classification but the decision process of Random Forest is not easily understandable like the associative classifiers. In this study we propose Dynamic Ensemble Associative Learning (DEAL) where we use associative classifiers as base learners on feature sub-spaces. In our approach we select a subset of the feature vector to train each of the base learners. Instead of a random selection, we propose a dynamic feature sampling procedure which automatically defines the number of base learners and ensures diversity and completeness among the subset of feature vectors. We use 10 datasets from the UCI repository and evaluate the performance of the model in terms of accuracy and memory requirement. Our ensemble approach using the proposed sampling method largely decreases the memory requirement in the case of datasets having a large number of features and this without jeopardising accuracy. In fact, accuracy is also improved in most cases. Moreover, the decision process of our DEAL approach remains human interpretable by collecting and ranking the rules generated by the base learners predicting the final class label. Md. Rayhanul Kabir, Osmar R. Zaïane |
ASONAM | 2 |
| 2022 | Sentiment and Knowledge Based Algorithmic Trading with Deep Reinforcement Learning
Abhishek Nan, Anandh Perumal, Osmar R. Zaïane |
DEXA (1) | 3 |
| 2022 | Named Entity Recognition for Partially Annotated Datasets
Michael Strobl, Amine Trabelsi, Osmar R. Zaïane |
NLDB | 3 |
| 2021 | Simulated Annealing for Emotional Dialogue SystemsabstractExplicitly modeling emotions in dialogue generation has important applications, such as building empathetic personal companions. In this study, we consider the task of expressing a specific emotion for dialogue generation. Previous approaches take the emotion as a training signal, which may be ignored during inference. Here, we propose a search-based emotional dialogue system by simulated annealing (SA). Specifically, we first define a scoring function that combines contextual coherence and emotional correctness. Then, SA iteratively edits a general response, and search for a generation with a high score. In this way, we enforce the presence of the desired emotion. We evaluate our system on the NLPCC2017 dataset. The proposed method shows about 12% improvements in emotion accuracy compared with the previous state-of-the-art method, without hurting the generation quality (measured by BLEU). Chengzhang Dong, Chenyang Huang 0001, Osmar R. Zaïane, Lili Mou |
CIKM | 3 |
| 2020 | Building a Competitive Associative Classifier
Nitakshi Sood, Osmar R. Zaïane |
DaWaK | 2 |
| 2020 | Bi-Level Associative Classifier Using Automatic Learning on Rules
Nitakshi Sood, Leepakshi Bindra, Osmar R. Zaïane |
DEXA (1) | 3 |
| 2020 | Addressing the Resolution Limit and the Field of View Limit in Community MiningabstractWe introduce a novel efficient approach for community detection based on a formal definition of the notion of community. We name the links that run between communities weak links and links being inside communities strong links. We put forward a new objective function, called SIWO (Strong Inside, Weak Outside) which encourages adding strong links to the communities while avoiding weak links. This process allows us to effectively discover communities in social networks without the resolution and field of view limit problems some popular approaches suffer from. The time complexity of this new method is linear in the number of edges. We demonstrate the effectiveness of our approach on various real and artificial datasets with large and small communities. Shiva Zamani Gharaghooshi, Osmar R. Zaïane, Christine Largeron, Mohammadmahdi Zafarmand |
IDA | 2 |
| 2020 | A Late-Fusion Approach to Community Detection in Attributed NetworksabstractThe majority of research on community detection in attributed networks follows an “early fusion” approach, in which the structural and attribute information about the network are integrated together as the guide to community detection. In this paper, we propose an approach called late-fusion , which looks at this problem from a different perspective. We first exploit the network structure and node attributes separately to produce two different partitionings. Later on, we combine these two sets of communities via a fusion algorithm, where we introduce a parameter for weighting the importance given to each type of information: node connections and attribute values. Extensive experiments on various real and synthetic networks show that our late-fusion approach can improve detection accuracy from using only network structure. Moreover, our approach runs significantly faster than other attributed community detection algorithms including early fusion ones. Christine Largeron, Osmar R. Zaïane, Shiva Zamani Gharaghooshi |
IDA | 3 |
| 2020 | Model-Based Clustering with HDBSCAN
Michael Strobl, Jörg Sander 0001, Ricardo J. G. B. Campello, Osmar R. Zaïane |
ECML/PKDD (2) | 4 |
| 2020 | Framework for extreme imbalance classification: SWIM - sampling with the majority class
Colin Bellinger, Shiven Sharma, Nathalie Japkowicz, Osmar R. Zaïane |
Knowl. Inf. Syst. | 4 |
| 2019 | Detecting the Onset of Machine Failure Using Anomaly Detection Methods
Mohammad Riazi, Osmar R. Zaïane, Tomoharu Takeuchi, Anthony Maltais, Johannes Günther 0002, Micheal Lipsett |
DaWaK | 2 |
| 2019 | PhAITV: A Phrase Author Interaction Topic Viewpoint Model for the Summarization of Reasons Expressed by Polarized Stances
Amine Trabelsi, Osmar R. Zaïane |
ICWSM | 2 |
| 2019 | Neighbor-Based Link Prediction with Edge Uncertainty
Osmar R. Zaïane |
PAKDD (2) | 2 |
| 2019 | Augmenting Semantic Representation of Depressive Language: From Forums to Microblogs
Nawshad Farruque, Osmar R. Zaïane, Randy Goebel |
ECML/PKDD (3) | 2 |
| 2019 | APNEA: Intelligent Ad-Bidding Using Sentiment AnalysisabstractOnline advertising is one of the most lucrative forms of advertising, making it an important channel of advertising media. Contextual Advertising is a type of online display advertising that takes cues from the content of the triggering page and displays advertisements that are relevant to the current context. However, on several occasions, the context may have a negative connotation, and displaying advertisements that are relevant to it might prove to be detrimental to the advertiser. We refer to such a scenario as an unfortunate placement. In this work, we propose APNEA (Ad Positive NEgative Analysis), a light-weight system that uses a sentiment-oriented approach to rank the advertisers such that positively correlated brands are ranked higher than brands that are neutral or negatively correlated. Experiments show that APNEA helps avoid unfortunate placements while maintaining ad-relevance. It outperforms several baselines in terms of accuracy on human-annotated test data while having a lower run-time, which is crucial for real-time bidding systems. Samuel Suraj Bushi, Osmar R. Zaïane |
WI | 2 |
| 2018 | Detecting Local Communities in Networks with Edge UncertaintyabstractIn this work, we focus on the problem of local community detection with edge uncertainty. We use an estimator to cope with the intrinsic uncertainty of the problem. Then we illustrate with an example that periphery nodes tend to be grouped into their neighbor communities in uncertain networks, and we propose a new measure K to address this problem. Due to the very limited publicly available uncertain network datasets, we also put forward a way to generate uncertain networks. Finally, we evaluate our algorithm using existing ground truth as well as based on common metrics to show the effectiveness of our proposed approach. Osmar R. Zaïane |
ASONAM | 2 |
| 2018 | Sequence-Based Approaches to Course Recommender Systems
Ren Wang 0009, Osmar R. Zaïane |
DEXA (1) | 2 |
| 2018 | Synthetic Oversampling with the Majority Class: A New Perspective on Handling Extreme ImbalanceabstractThe class imbalance problem is a pervasive issue in many real-world domains. Oversampling methods that inflate the rare class by generating synthetic data are amongst the most popular techniques for resolving class imbalance. However, they concentrate on the characteristics of the minority class and use them to guide the oversampling process. By completely overlooking the majority class, they lose a global view on the classification problem and, while alleviating the class imbalance, may negatively impact learnability by generating borderline or overlapping instances. This becomes even more critical when facing extreme class imbalance, where the minority class is strongly underrepresented and on its own does not contain enough information to conduct the oversampling process. We propose a novel method for synthetic oversampling that uses the rich information inherent in the majority class to synthesize minority class data. This is done by generating synthetic data that is at the same Mahalanbois distance from the majority class as the known minority instances. We evaluate over 26 benchmark datasets, and show that our method offers a distinct performance improvement over the existing state-of-the-art in oversampling techniques. Shiven Sharma, Colin Bellinger, Bartosz Krawczyk, Osmar R. Zaïane, Nathalie Japkowicz |
ICDM | 4 |
| 2018 | Unsupervised Model for Topic Viewpoint Discovery in Online Debates Leveraging Author Interactions
Amine Trabelsi, Osmar R. Zaïane |
ICWSM | 2 |
| 2018 | Context Prediction in the Social Web Using Applied Machine Learning: A Study of Canadian TweetersabstractIn this ongoing work, we present the Grebe social data aggregation framework for extracting geo-fenced Twitter data for analysis of user engagement in health and wellness topics. Grebe also provides various visualization tools for analyzing temporal and geographical health trends. Grebe currently has over 18 million indexed public tweets, and is the first of its kind for Canadian researchers. The large dataset is used for analyzing three types of contexts: geographical context via prediction of user location using supervised learning, topical context via determining health-related tweets using various learning approaches, and affective context via sentiment analysis of tweets using rule-based methods. For the first, we define user location as the position from which users are posting a tweet and use standard precision metrics for evaluation with promising results for predicting provinces and cities from tweet text. For the second, we use a broader definition of health using the six dimensions of wellness model and evaluate using manually annotated documents with good results using supervised and semi-supervised machine learning. For the third, we use the indexed tweets to show current trends in emotions and opinions and demonstrate trends in polarity and emotions across various Canadian provinces. The combination of these contexts provides useful insights for digital epidemiology. Ultimately, the vision of Grebe is to provide researchers with Canada-specific social web datasets through an open source platform with an accessible RESTful API, and this paper showcases Grebe's potential and presents our progress towards achieving these goals. Hamman W. Samuel, Benyamin Noori, Sara Farazi, Osmar R. Zaïane |
WI | 4 |
| 2017 | Sentiment Analysis on Twitter to Improve Time Series Contextual Anomaly Detection for Detecting Stock Market Manipulation
Koosha Golmohammadi, Osmar R. Zaïane |
DaWaK | 2 |
| 2017 | Toward Personalized Relational LearningabstractRelational learning exploits relationships among instances manifested in a network to improve the predictive performance of many network mining tasks. Due to its empirical success, it has been widely applied in myriad domains. In many cases, individuals in a network are highly idiosyncratic. They not only connect to each other with a composite of factors but also are often described by some content information of high dimensionality specific to each individual. For example in social media, as user interests are quite diverse and personal; posts by different users could differ significantly. Moreover, social content of users is often of high dimensionality which may negatively degrade the learning performance. Therefore, it would be more appealing to tailor the prediction for each individual while alleviating the issue related to the curse of dimensionality. In this paper, we study a novel problem of Personalized Relational Learning and propose a principled framework PRL to personalize the prediction for each individual in a network. Specifically, we perform personalized feature selection and employ a small subset of discriminative features customized for each individual and some common features shared by all to build a predictive model. On this account, the proposed personalized model is more human interpretable. Experiments on real-world datasets show the superiority of the proposed PRL framework over traditional relational learning methods. Jundong Li, Liang Wu 0006, Osmar R. Zaïane, Huan Liu 0001 |
SDM | 3 |
| 2017 | DANCer: dynamic attributed networks with community structure generation
Christine Largeron, Pierre-Nicolas Mougel, Oualid Benyahia, Osmar R. Zaïane |
Knowl. Inf. Syst. | 4 |
| 2016 | Advantage of integration in big data: Feature generation in multi-relational databases for imbalanced learningabstractMost real world applications comprise databases having multiple tables. It becomes further complicated in the realm of Big Data where related information is spread over different data repositories. However, data mining techniques are usually applied on a single flat table. This work focuses on generating a mining table by aggregating information from multiple local tables and external data sources and automatically generating potentially discriminant features. It extends data aggregation techniques by navigating paths where a single table is traversed multiple times. Such paths are not considered by existing techniques, which results in the loss of several attributes. Our framework also prevents leakage of the class information by avoiding features built after the knowledge of the class label. Experiments are performed on transactional data of a U.S. consumer electronics retailer to predict causes of product returns. In addition, we augmented the dataset with Suppliers information and Reviews to show the value of data integration. The results show that our technique improves classification accuracy and generates discriminant features that mitigate the impact of class imbalance. Farrukh Ahmed, Michele Samorani, Colin Bellinger, Osmar R. Zaïane |
IEEE BigData | 4 |
| 2016 | Automatic generation of relational attributes: An application to product returnsabstractAlthough statistical and machine learning methods require the input data to be in a tabular format, in real-world applications data are often stored across several tables in a relational database. How to build a single mining table from a relational database is a critical pre-processing step of any classification method, because including the right attributes may dramatically boost the accuracy of the classifier. We propose a methodology and implement a software program, Dataconda, to automatically mine a relational database. The user selects a class attribute contained in a table of the database and the procedure builds and selects predictors by exploring the whole database and aggregating information, without any user intervention. For example, our procedure may find that the best predictor for “product return” is the proportion of products returned by the same customer in the past, even if the user has not built any such attribute. Our procedure produces more expressive attributes than existing methods. Our experiments on the ISMS Durable Goods Datasets, a publicly available data set of product returns in retailing, suggest that our method allows new knowledge to emerge. Michele Samorani, Farrukh Ahmed, Osmar R. Zaïane |
IEEE BigData | 3 |
| 2016 | DANCer: Dynamic Attributed Network with Community Structure Generator
Oualid Benyahia, Christine Largeron, Baptiste Jeudy, Osmar R. Zaïane |
ECML/PKDD (3) | 4 |
| 2016 | On discovering co-location patterns in datasets: a case study of pollutants and child cancers
Jundong Li, Aibek Adilmagambetov, Mohomed Shazan Mohomed Jabbar, Osmar R. Zaïane, Alvaro Osornio-Vargas, Osnat Wine |
GeoInformatica | 4 |
| 2016 | Mining contentious documents
Amine Trabelsi, Osmar R. Zaïane |
Knowl. Inf. Syst. | 2 |
| 2015 | Associative Classification with Statistically Significant Positive and Negative RulesabstractRule-based classifier has shown its popularity in building many decision support systems such as medical diagnosis and financial fraud detection. One major advantage is that the models are human understandable and can be edited. Associative classifiers, as an extension of rule-based classifiers, use association rules to associate attributes with class labels. A delicate issue of associative classifiers is the need for subtle thresholds: minimum support and minimum confidence. Without prior knowledge, it could be difficult to choose the proper thresholds, and the discovered rules within the support-confidence framework are not statistically significant, i.e., inclusion of noisy rules and exclusion of valuable rules. Besides, most associative classifiers proposed so far, are built with only positive association rules. Negative rules, however, are also able to provide valuable information to discriminate between classes. To solve the above mentioned problems, we propose a novel associative classifier which is built upon both positive and negative classification association rules that show statistically significant dependencies. Experimental results on real-world datasets show that our method achieves competitive or even better performance than well-known rule-based and associative classifiers in terms of both classification accuracy and computational efficiency. Jundong Li, Osmar R. Zaïane |
CIKM | 2 |
| 2015 | Time series contextual anomaly detection for detecting market manipulation in stock marketabstractAnomaly detection in time series is one of the fundamental issues in data mining that addresses various problems in different domains such as intrusion detection in computer networks, irregularity detection in healthcare sensory data and fraud detection in insurance or securities. Although, there has been extensive work on anomaly detection, majority of the techniques look for individual objects that are different from normal objects but do not take the temporal aspect of data into consideration. We are particularly interested in contextual outlier detection methods for time series that are applicable to fraud detection in securities. This has significant impacts on national and international securities markets. In this paper, we propose a prediction-based Contextual Anomaly Detection (CAD) method for complex time series that are not described through deterministic models. The proposed method improves the recall from 7% to 33% compared to kNN and Random Walk without compromising the precision. Koosha Golmohammadi, Osmar R. Zaïane |
DSAA | 2 |
| 2015 | Generalization of clustering agreements and distances for overlapping clusters and network communities
Reihaneh Rabbany, Osmar R. Zaïane |
Data Min. Knowl. Discov. | 2 |
| 2015 | Extraction and clustering of arguing expressions in contentious text
Amine Trabelsi, Osmar R. Zaïane |
Data Knowl. Eng. | 2 |
| 2015 | Recognition of Patient-Related Named Entities in Noisy Tele-Health TextsabstractWe explore methods for effectively extracting information from clinical narratives that are captured in a public health consulting phone service called HealthLink. Our research investigates the application of state-of-the-art natural language processing and machine learning to clinical narratives to extract information of interest. The currently available data consist of dialogues constructed by nurses while consulting patients by phone. Since the data are interviews transcribed by nurses during phone conversations, they include a significant volume and variety of noise. When we extract the patient-related information from the noisy data, we have to remove or correct at least two kinds of noise: explicit noise , which includes spelling errors, unfinished sentences, omission of sentence delimiters, and variants of terms, and implicit noise , which includes non-patient information and patient's untrustworthy information. To filter explicit noise, we propose our own biomedical term detection/normalization method: it resolves misspelling, term variations, and arbitrary abbreviation of terms by nurses. In detecting temporal terms, temperature, and other types of named entities (which show patients’ personal information such as age and sex), we propose a bootstrapping-based pattern learning process to detect a variety of arbitrary variations of named entities. To address implicit noise, we propose a dependency path-based filtering method. The result of our denoising is the extraction of normalized patient information, and we visualize the named entities by constructing a graph that shows the relations between named entities. The objective of this knowledge discovery task is to identify associations between biomedical terms and to clearly expose the trends of patients’ symptoms and concern; the experimental results show that we achieve reasonable performance with our noise reduction methods. Mi-Young Kim, Ying Xu 0003, Osmar R. Zaïane, Randy Goebel |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2014 | Community Dynamics: Event and Role Analysis in Social Network Analysis
Justin Fagnan, Reihaneh Rabbany, Mansoureh Takaffoli, Eric Verbeek 0002, Osmar R. Zaïane |
ADMA | 5 |
| 2014 | SSRM: Structural social role mining for dynamic social networksabstractA social role is a special position an individual possesses within a network, which indicates his or her behaviours, expectations, and responsibilities. Identifying the roles that individuals play in a social network has various direct applications, such as detecting influential members, trustworthy people, idea innovators, etc. Roles can also be used for further analyses of the network, e.g. community detection, temporal event prediction, and summarization. In this paper, we propose a structural social role mining framework (SSRM), which is built to identify roles, study their changes, and analyze their impacts on the underlying social network. We define fundamental roles in a social network (namely leader, outermost, mediator, and outsider), and then propose methodologies to identify them, and track their changes. To identify these roles, we leverage the traditional social network analyses and metrics, as well as proposing new measures, including community-based variants for the Betweenness centrality. Our results indicate how the changes in the structural roles, in combination with the changes in the community structure of a network, can provide additional clues into the dynamics of networks. Afra Abnar, Mansoureh Takaffoli, Reihaneh Rabbany, Osmar R. Zaïane |
ASONAM | 4 |
| 2014 | Using triads to identify local community structure in social networksabstractWe present our novel community mining algorithm that uses only local information to accurately identify communities, outliers, and hubs in social networks. The main component of our algorithm is the T metric, which evaluates the relative quality of a community by considering the number of internal and external triads (3-node cliques) it contains. Furthermore we propose an intuitive statistical method based on our T metric, which correctly identifies outlier and hub nodes within each discovered community. Finally, we evaluate our approach on a series of ground-truth networks and show that our method outperforms the state-of-the-art in community mining algorithms. Justin Fagnan, Osmar R. Zaïane, Denilson Barbosa 0001 |
ASONAM | 2 |
| 2014 | Community evolution prediction in dynamic social networksabstractFinding patterns of interaction and predicting the future structure of networks has many important applications, such as recommendation systems and customer targeting. Community structure of social networks may undergo different temporal events and transitions. In this paper, we propose a framework to predict the occurrence of different events and transition for communities in dynamic social networks. Our framework incorporates key features related to a community - its structure, history, and influential members, and automatically detects the most predictive features for each event and transition. Our experiments on real world datasets confirms that the evolution of communities can be predicted with a very high accuracy, while we further observe that the most significant features vary for the predictability of each event and transition. Mansoureh Takaffoli, Reihaneh Rabbany, Osmar R. Zaïane |
ASONAM | 3 |
| 2014 | Discovering Statistically Significant Co-location Rules in Datasets with Extended Spatial Objects
Jundong Li, Osmar R. Zaïane, Alvaro Osornio-Vargas |
DaWaK | 2 |
| 2014 | Detecting stock market manipulation using supervised learning algorithmsabstractMarket manipulation remains the biggest concern of investors in today's securities market, despite fast and strict responses from regulators and exchanges to market participants that pursue such practices. The existing methods in the industry for detecting fraudulent activities in securities market rely heavily on a set of rules based on expert knowledge. The securities market has deviated from its traditional form due to new technologies and changing investment strategies in the past few years. The current securities market demands scalable machine learning algorithms supporting identification of market manipulation activities. In this paper we use supervised learning algorithms to identify suspicious transactions in relation to market manipulation in stock market. We use a case study of manipulated stocks during 2003. We adopt CART, conditional inference trees, C5.0, Random Forest, Naïve Bayes, Neural Networks, SVM and kNN for classification of manipulated samples. Empirical results show that Naïve Bayes outperform other learning methods achieving F2measure of 53% (sensitivity and specificity are 89% and 83% respectively). Koosha Golmohammadi, Osmar R. Zaïane, David Diaz |
DSAA | 2 |
| 2014 | Mining Contentious Documents Using an Unsupervised Topic Model Based ApproachabstractThis work proposes an unsupervised method intended to enhance the quality of opinion mining in contentious text. It presents a Joint Topic Viewpoint (JTV) probabilistic model to analyse the underlying divergent arguing expressions that may be present in a collection of contentious documents. It extends the original Latent Dirichlet Allocation (LDA), which makes it domain and thesaurus-independent, e.g., does not rely on Word Net coverage. The conceived JTV has the potential of automatically carrying the tasks of extracting associated terms denoting an arguing expression, according to the hidden topics it discusses and the embedded viewpoint it voices. Furthermore, JTV's structure enables the unsupervised grouping of obtained arguing expressions according to their viewpoints, using a constrained clustering approach. Experiments are conducted on three types of contentious documents: polls, online debates and editorials. The qualitative and quantitative analysis of the experimental results show the effectiveness of our model to handle six different contentious issues when compared to a state-of-the-art method. Moreover, the ability to automatically generate distinctive and informative patterns of arguing expressions is demonstrated. Amine Trabelsi, Osmar R. Zaïane |
ICDM | 2 |
| 2014 | A Joint Topic Viewpoint Model for Contention Analysis
Amine Trabelsi, Osmar R. Zaïane |
NLDB | 2 |
| 2013 | Utility Enhancement for Privacy Preserving Health Data Publishing
Lengdong Wu, Osmar R. Zaïane |
ADMA (2) | 3 |
| 2013 | Biomedical text disambiguation using UMLSabstractInterest in extracting information from biomedical documents has increased significantly in recent years but has always been challenged by the ambiguity of natural language. An important source of ambiguity is the usage of polysemous words: words with multiple meanings. Word sense disambiguation algorithms attempt to solve this problem by finding the correct meaning of a polysemous word in a given context, but very few algorithms were designed to disambiguate biomedical text. In this study we propose a word sense disambiguation algorithm focused on biomedical text. The proposed algorithm does not need to be trained and uses a relatively small knowledge base. Wessam Gad El Rab, Osmar R. Zaïane |
ASONAM | 2 |
| 2013 | Incremental local community identification in dynamic social networksabstractSocial networks are usually drawn from the interactions between individuals, and therefore are temporal and dynamic in essence. Examining how the structure of these networks changes over time provides insights into their evolution patterns, factors that trigger the changes, and ultimately predict the future structure of these networks. One of the key structural characteristics of networks is their community structure --groups of densely interconnected nodes. Communities in a dynamic social network span over periods of time and are affected by changes in the underlying population, i.e. they have fluctuating members and can grow and shrink over time. In this paper, we introduce a new incremental community mining approach, in which communities in the current time are obtained based on the communities from the past time frame. Compared to previous independent approaches, this incremental approach is more effective at detecting stable communities over time. Extensive experimental studies on real datasets, demonstrate the applicability, effectiveness, and soundness of our proposed framework. Mansoureh Takaffoli, Reihaneh Rabbany, Osmar R. Zaïane |
ASONAM | 3 |
| 2013 | Discovering Co-location Patterns in Datasets with Extended Spatial Objects
Aibek Adilmagambetov, Osmar R. Zaïane, Alvaro Osornio-Vargas |
DaWaK | 2 |
| 2013 | An Optimized Cost-Sensitive SVM for Imbalanced Data Learning
Peng Cao 0001, Dazhe Zhao, Osmar R. Zaïane |
PAKDD (2) | 3 |
| 2013 | Text Document Topical Recursive Clustering and Automatic Labeling of a Hierarchy of Document Clusters
Jiyang Chen, Osmar R. Zaïane |
PAKDD (2) | 3 |
| 2012 | Relative Validity Criteria for Community Mining AlgorithmsabstractGrouping data points is one of the fundamental tasks in data mining, which is commonly known as clustering if data points are described by attributes. When dealing with interrelated data that does not have any attributes and is represented in the form of nodes and their relationships, this task is also referred to as community mining. There has been a considerable number of approaches proposed in recent years for mining communities in a given network. But little work has been done on how to evaluate community mining results. The common practice is to use an agreement measure to compare the mining result against a ground truth, however, the ground truth is not known in most of the real world applications. In this paper, we investigate relative clustering quality measures defined for evaluation of clustering data points with attributes and propose proper adaptations to make them applicable in the context of social networks. Not only these relative criteria could be used as metrics for evaluating quality of the groupings but also they could be used as objectives for designing new community mining algorithms. Reihaneh Rabbany, Mansoureh Takaffoli, Justin Fagnan, Osmar R. Zaïane, Ricardo J. G. B. Campello |
ASONAM | 4 |
| 2012 | Classifying Websites into Non-topical Categories
Chaman Thapa, Osmar R. Zaïane, Davood Rafiei, Arya M. Sharma |
DaWaK | 2 |
| 2012 | An Associative Classifier for Uncertain Datasets
Metanat HooshSadat, Osmar R. Zaïane |
PAKDD (1) | 2 |
| 2012 | Guest editorial: special issue on a decade of mining the Web
Myra Spiliopoulou, Bamshad Mobasher, Olfa Nasraoui, Osmar R. Zaïane |
Data Min. Knowl. Discov. | 4 |
| 2011 | Learning Actions in Complex Software Systems
Koosha Golmohammadi, Michael Smit, Osmar R. Zaïane |
DaWaK | 3 |
| 2011 | MODEC - Modeling and Detecting Evolutions of Communities
Mansoureh Takaffoli, Farzad Sangi, Justin Fagnan, Osmar R. Zaïane |
ICWSM | 4 |
| 2011 | Class separation through variance: a new application of outlier detection
Andrew Foss, Osmar R. Zaïane |
Knowl. Inf. Syst. | 2 |
| 2011 | On Pruning for Top-K Ranking in Uncertain DatabasesabstractTop-k ranking for an uncertain database is to rank tuples in it so that the best k of them can be determined. The problem has been formalized under the unified approach based on parameterized ranking functions (PRFs) and the possible world semantics. Given a PRF, one can always compute the ranking function values of all the tuples to determine the top-k tuples, which is a formidable task for large databases. In this paper, we present a general approach to pruning for the framework based on PRFs. We show a mathematical manipulation of possible worlds which reveals key insights in the part of computation that may be pruned and how to achieve it in a systematic fashion. This leads to concrete pruning methods for a wide range of ranking functions. We show experimentally the effectiveness of our approach. Chonghai Wang, Li-Yan Yuan, Jia-Huai You, Osmar R. Zaïane, Jian Pei 0001 |
Proc. VLDB Endow. | 4 |
| 2010 | An Occurrence Based Approach to Mine Emerging Sequences
Kang Deng, Osmar R. Zaïane |
DaWak | 2 |
| 2009 | Local Community Identification in Social NetworksabstractThere has been much recent research on identifying global community structure in networks. However, most existing approaches require complete information of the graph in question, which is impractical for some networks, e.g. the World Wide Web (WWW). Algorithms for local community detection have been proposed but their results usually contain many outliers. In this paper, we propose a new measure of local community structure, coupled with a two-phase algorithm that extracts all possible candidates first, and then optimizes the community hierarchy. We compare our results with previous methods on real world networks such as the co-purchase network from Amazon. Experimental results verify the feasibility and effectiveness of our approach. Jiyang Chen, Osmar R. Zaïane, Randy Goebel |
ASONAM | 2 |
| 2009 | A Visual Data Mining Approach to Find Overlapping Communities in NetworksabstractCommunities in social networks may overlap, with some hub nodes belonging to multiple communities. They may also have outliers, which are nodes that belong to no community. The criterion to locate hubs or outliers is network dependent. Previous methods usually require this information as input parameters, e.g., an expected number of communities, with no intuition or assistance. Here we present a visual data mining approach, which first helps the user to make appropriate parameter selections by observing initial data visualizations, and then finds and extracts overlapping community structures from the network. Experimental results verify the scalability and accuracy of our approach on real network data and show its advantages over previous methods. Jiyang Chen, Osmar R. Zaïane, Randy Goebel |
ASONAM | 2 |
| 2009 | Unsupervised Class Separation of Multivariate Data through Cumulative Variance-Based RankingabstractThis paper introduces a new extension of outlier detection approaches and a new concept, class separation through variance. We show that accumulating information about the outlierness of points in multiple subspaces leads to a ranking in which classes with differing variance naturally tend to separate. Exploiting this leads to a highly effective and efficient unsupervised class separation approach, especially useful in the difficult case of heavily overlapping distributions. Unlike typical outlier detection algorithms, this method can be applied beyond the `rare classes' case with great success. Two novel algorithms that implement this approach are provided. Additionally, experiments show that the novel methods typically outperform other state-of-the-art outlier detection methods on high dimensional data such as Feature Bagging, SOE1, LOF, ORCA and Robust Mahalanobis Distance and competes even with the leading supervised classification methods. Andrew Foss, Osmar R. Zaïane, Sandra Zilles |
ICDM | 2 |
| 2009 | Detecting Communities in Social Networks Using Max-Min ModularityabstractMany datasets can be described in the form of graphs or networks where nodes in the graph represent entities and edges represent relationships between pairs of entities. A common property of these networks is their community structure, considered as clusters of densely connected groups of vertices, with only sparser connections between groups. The identification of such communities relies on some notion of clustering or density measure. which defines the communities that can be found. However, previous community detection methods usually apply the same structural measure on all kinds of networks, despite their distinct dissimilar features. In this paper, we present a new community mining measure, Max-Min Modularity, which considers both connected pairs and criteria defined by domain experts in finding communities, and then specify a hierarchical clustering algorithm to detect communities in networks. When applied to real world networks for which the community structures are already known, our method shows improvement over previous algorithms. In addition, when applied to randomly generated networks for which we only have approximate information about communities, it gives promising results which shows the algorithm's robustness against noise. Jiyang Chen, Osmar R. Zaïane, Randy Goebel |
SDM | 2 |
| 2009 | Resolution-based outlier factor: detecting the top- n most outlying data points in engineering data
Hongqin Fan, Osmar R. Zaïane, Andrew Foss |
Knowl. Inf. Syst. | 2 |
| 2009 | Clustering and Sequential Pattern Mining of Online Collaborative Learning DataabstractGroup work is widespread in education. The growing use of online tools supporting group work generates huge amounts of data. We aim to exploit this data to support mirroring: presenting useful high-level views of information about the group, together with desired patterns characterizing the behavior of strong groups. The goal is to enable the groups and their facilitators to see relevant aspects of the group's operation and provide feedback if these are more likely to be associated with positive or negative outcomes and indicate where the problems are. We explore how useful mirror information can be extracted via a theory-driven approach and a range of clustering and sequential pattern mining. The context is a senior software development project where students use the collaboration tool TRAC. We extract patterns distinguishing the better from the weaker groups and get insights in the success factors. The results point to the importance of leadership and group interaction, and give promising indications if they are occurring. Patterns indicating good individual practices were also identified. We found that some key measures can be mined from early data. The results are promising for advising groups at the start and early identification of effective and poor practices, in time for remediation. Dilhan Perera, Judy Kay, Irena Koprinska, Kalina Yacef, Osmar R. Zaïane |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2008 | An Unsupervised Approach to Cluster Web Search Results Based on Word Sense CommunitiesabstractEffectively organizing web search results into clusters is important to facilitate quick user navigation to relevant documents. Previous methods may rely on a training process and do not provide a measure for whether page clustering is actually required. In this paper, we reformalize the clustering problem as a word sense discovery problem. Given a query and a list of result pages, our unsupervised method detects word sense communities in the extracted keyword network. The documents are assigned to several refined word sense communities to form clusters. We use the modularity score of the discovered keyword community structure to measure page clustering necessity. Experimental results verify our method's feasibility and effectiveness. Jiyang Chen, Osmar R. Zaïane, Randy Goebel |
Web Intelligence | 2 |
| 2008 | Clustering high dimensional data: A graph-based relaxed optimization approach
Osmar R. Zaïane, Ho-Hyun Park, Jiayuan Huang, Russell Greiner |
Inf. Sci. | 2 |
| 2007 | Contrasting the Contrast Sets: An Alternative ApproachabstractThe need to identify significant differences between contrasting groups or classes is ubiquitous and thus was the focus of many statisticians and data miners. Contrast sets, conjunctions of attribute-value pairs significantly more frequent in one group than another, were proposed to describe such differences, which lead to the introduction of a new data mining technique - contrast-set mining. A number of attempts have been made in this regard by various authors; however, no clear picture seems to have emerged. In this paper, we try to address the problem of finding meaningful contrast sets by using association rule based analysis. We present the results for our experiments for interesting contrast sets and compare these results with those obtained from the well-known algorithm for contrast sets-STUCCO. Amit Satsangi, Osmar R. Zaïane |
IDEAS | 2 |
| 2007 | Feature Space Enrichment by Incorporation of Implicit Features for Effective ClassificationabstractFeature space conversion for classifiers is the process by which the data that is to be fed into the classifier is transformed from one form to another. The motivation behind doing this is to enhance the "discriminative power" of the data together with preserving its "information content". In this paper, a new method of feature space conversion is explored, wherein "enrichment" of the feature space is carried out by the augmentation of the existing features with new "implicit" features. The modus operandi involves generation of association rules in one case and closed frequent patterns in another and the extraction of the new features from these. This new feature space is first made use of independently to feed the classifier and then it is used in unison with the original feature space. The effectiveness of these methods is subsequently verified experimentally and expressed in terms of the classification accuracy achieved by the classifier. Abhishek Srivastava 0001, Osmar R. Zaïane, Maria-Luiza Antonie |
IDEAS | 2 |
| 2006 | Learning to Use a Learned Model: A Two-Stage Approach to ClassificationabstractAssociation rule-based classifiers have recently emerged as competitive classification systems. However, there are still deficiencies that hinder their performance. One deficiency is the use of rules in the classification stage. Current systems assign classes to new objects based on the best rule applied or on some predefined scoring of multiple rules. In this paper we propose a new technique where the system automatically learns how to use the rules. We achieve this by developing a two-stage classification model. First, we use association rule mining to discover classification rules. Second, we employ another learning algorithm to learn how to use these rules in the prediction process. Our two-stage approach outperforms C4.5 and RIPPER on the UCI datasets in our study, and outperforms other rule- learning methods on more than half the datasets. The versatility of our method is also demonstrated by applying it to text classification, where it equals the performance of the best known systems for this task, SVMs. Maria-Luiza Antonie, Osmar R. Zaïane, Robert C. Holte |
ICDM | 2 |
| 2006 | An Efficient Reference-Based Approach to Outlier Detection in Large DatasetsabstractA bottleneck to detecting distance and density based outliers is that a nearest-neighbor search is required for each of the data points, resulting in a quadratic number of pairwise distance evaluations. In this paper, we propose a new method that uses the relative degree of density with respect to a fixed set of reference points to approximate the degree of density defined in terms of nearest neighbors of a data point. The running time of our algorithm based on this approximation is 0(Rn log n) where n is the size of dataset and R is the number of reference points. Candidate outliers are ranked based on the outlier score assigned to each data point. Theoretical analysis and empirical studies show that our method is effective, efficient, and highly scalable to very large datasets. Yaling Pei, Osmar R. Zaïane |
ICDM | 2 |
| 2006 | A Nonparametric Outlier Detection for Effectively Discovering Top-N Outliers from Engineering Data
Hongqin Fan, Osmar R. Zaïane, Andrew Foss |
PAKDD | 2 |
| 2006 | Efficient Spatial Classification Using Decoupled Conditional Random Fields
Russell Greiner, Osmar R. Zaïane |
PKDD | 3 |
| 2006 | Parallel Bifold: Large-scale parallel pattern mining with constraints
Osmar R. Zaïane |
Distributed Parallel Databases | 2 |
| 2005 | Finding All Frequent Patterns Starting from the Closure
Osmar R. Zaïane |
ADMA | 2 |
| 2005 | Relevance of Counting in Data Mining Tasks
Osmar R. Zaïane |
ADMA | 1 |
| 2005 | Scrutinizing Frequent Pattern Discovery PerformanceabstractBenchmarking technical solutions is as important as the solutions themselves. Yet many fields still lack any type of rigorous evaluation. Performance benchmarking has always been an important issue in databases and has played a significant role in the development, deployment and adoption of technologies. To help assessing the myriad algorithms for frequent itemset mining, we built an open framework and testbed to analytically study the performance of different algorithms and their implementations, and contrast their achievements given different data characteristics, different conditions, and different types of patterns to discover and their constraints. This facilitates reporting consistent and reproducible performance results using known conditions. Osmar R. Zaïane, Stella Luk |
ICDE | 1 |
| 2005 | Bifold Constraint-Based Mining by Simultaneous Monotone and Anti-Monotone CheckingabstractMining for frequent item sets can generate an overwhelming number of patterns, often exceeding the size of the original transactional database. One way to deal with this issue is to set filters and interestingness measures. Others advocate the use of constraints to apply to the patterns, either on the form of the patterns or on descriptors of the items in the patterns. However, typically the filtering of patterns based on these constraints is done as a post-processing phase. Filtering the patterns post-mining adds a significant overhead, still suffers from the sheer size of the pattern set and loses the opportunity to exploit those constraints. In this paper we propose an approach that allows the efficient mining of frequent item sets patterns, while pushing simultaneously both monotone and anti-monotone constraints during and at different strategic stages of the mining process. Our implementation shows a significant improvement when considering the constraints early and a better performance over Dualminer which also considers both types of constraints. Osmar R. Zaïane, Paul Nalos |
ICDM | 2 |
| 2005 | Pattern lattice traversal by selective jumpsabstractRegardless of the frequent patterns to discover, either the full frequent patterns or the condensed ones, either closed or maximal, the strategy always includes the traversal of the lattice of candidate patterns. We study the existing depth versus breadth traversal approaches for generating candidate patterns and propose in this paper a new traversal approach that jumps in the search space among only promising nodes. Our leaping approach avoids nodes that would not participate in the answer set and reduce drastically the number of candidate patterns. We use this approach to efficiently pinpoint maximal patterns at the border of the frequent patterns in the lattice and collect enough information in the process to generate all subsequent patterns. Osmar R. Zaïane |
KDD | 1 |
| 2005 | Considering Re-occurring Features in Associative Classifiers
Rafal Rak, Wojciech Stach, Osmar R. Zaïane, Maria-Luiza Antonie |
PAKDD | 3 |
| 2004 | Secure Association Rule Sharing
Stanley R. M. Oliveira, Osmar R. Zaïane, Yücel Saygin |
PAKDD | 2 |
| 2004 | Mining Positive and Negative Association Rules: An Approach for Confined Rules
Maria-Luiza Antonie, Osmar R. Zaïane |
PKDD | 2 |
| 2004 | Visualizing and Discovering Web Navigational PatternsabstractWeb site structures are complex to analyze. Cross-referencing the web structure with navigational behaviour adds to the complexity of the analysis. However, this convoluted analysis is necessary to discover useful patterns and understand the navigational behaviour of web site visitors, whether to improve web site structures, provide intelligent on-line tools or offer support to human decision makers. Moreover, interactive investigation of web access logs is often desired since it allows ad hoc discovery and examination of patterns not a priori known. Various visualization tools have been provided for this task but they often lack the functionality to conveniently generate new patterns. In this paper we propose a visualization tool to visualize web graphs, representations of web structure overlaid with information and pattern tiers. We also propose a web graph algebra to manipulate and combine web graphs and their layers in order to discover new patterns in an ad hoc manner. Jiyang Chen, Lisheng Sun, Osmar R. Zaïane, Randy Goebel |
WebDB | 3 |
| 2003 | Non-recursive Generation of Frequent K-itemsets from Frequent Pattern Tree Representations
Osmar R. Zaïane |
DaWaK | 2 |
| 2003 | Protecting Sensitive Knowledge By Data SanitizationabstractWe address the problem of protecting some sensitive knowledge in transactional databases. The challenge is on protecting actionable knowledge for strategic decisions, but at the same time not losing the great benefit of association rule mining. To accomplish that, we introduce a new, efficient one-scan algorithm that meets privacy protection and accuracy in association rule mining, without putting at risk the effectiveness of the data mining per se. Stanley R. M. Oliveira, Osmar R. Zaïane |
ICDM | 2 |
| 2003 | Incremental Mining of Frequent Patterns without Candidate Generation or Support ConstraintabstractIn this paper, we propose a novel data structure called CATS Tree. CATS Tree extends the idea of FPTree to improve storage compression and allow frequent pattern mining without generation of candidate item sets. The proposed algorithms enable frequent pattern mining with different supports without rebuilding the tree structure. Furthermore, the algorithms allow mining with a single pass over the database as well as efficient insertion or deletion of transactions at any time. Osmar R. Zaïane |
IDEAS | 2 |
| 2003 | Algorithms for Balancing Privacy and Knowledge Discovery in Association Rule MiningabstractThe discovery of association rules from large databases has proven beneficial for companies since such rules can be very effective in revealing actionable knowledge that leads to strategic decisions. In tandem with this benefit, association rule mining can also pose a threat to privacy protection. The main problem is that from non-sensitive information or unclassified data, one is able to infer sensitive information, including personal information, facts, or even patterns that are not supposed to be disclosed. This scenario reveals a pressing need for techniques that ensure privacy protection, while facilitating proper information accuracy and mining. In this paper, we introduce new algorithms for balancing privacy and knowledge discovery in association rule mining. We show that our algorithms require only two scans, regardless of the database size and the number of restrictive association rules that must be protected. Our performance study compares the effectiveness and scalability of the proposed algorithms and analyzes the fraction of association rules, which are preserved after sanitizing a database. We also report the main results of our performance evaluation and discuss some open research issues. Stanley R. M. Oliveira, Osmar R. Zaïane |
IDEAS | 2 |
| 2003 | Inverted matrix: efficient discovery of frequent items in large datasets in the context of interactive miningabstractExisting association rule mining algorithms suffer from many problems when mining massive transactional datasets. One major problem is the high memory dependency: either the gigantic data structure built is assumed to fit in main memory, or the recursive mining process is too voracious in memory resources. Another major impediment is the repetitive and interactive nature of any knowledge discovery process. To tune parameters, many runs of the same algorithms are necessary leading to the building of these huge data structures time and again. This paper proposes a new disk-based association rule mining algorithm called Inverted Matrix, which achieves its efficiency by applying three new ideas. First, transactional data is converted into a new database layout called Inverted Matrix that prevents multiple scanning of the database during the mining phase, in which finding frequent patterns could be achieved in less than a full scan with random access. Second, for each frequent item, a relatively small independent tree is built summarizing co-occurrences. Finally, a simple and non-recursive mining process reduces the memory requirements as minimum candidacy generation and counting is needed. Experimental studies reveal that our Inverted Matrix approach outperform FP-Tree especially in mining very large transactional databases with a very large number of unique items. Our random access disk-based approach is particularly advantageous in a repetitive and interactive setting. Osmar R. Zaïane |
KDD | 2 |
| 2002 | Text Document Categorization by Term AssociationabstractA good text classifier is a classifier that efficiently categorizes large sets of text documents in a reasonable time frame and with an acceptable accuracy, and that provides classification rules that are human readable for possible fine-tuning. If the training of the classifier is also quick, this could become in some application domains a good asset for the classifier. Many techniques and algorithms for automatic text categorization have been devised. According to published literature, some are more accurate than others, and some provide more interpretable classification models than others. However, none can combine all the beneficial properties enumerated above. In this paper we present a novel approach for automatic text categorization that borrows from market basket analysis techniques using association rule mining in the data-mining field. We focus on two major problems: (1) finding the best term association rules in a textual database by generating and pruning; and (2) using the rules to build a text classifier. Our text categorization method proves to be efficient and effective, and experiments on well-known collections show that the classifier performs well. In addition, training as well as classification are both fast and the generated rules are human readable. Maria-Luiza Antonie, Osmar R. Zaïane |
ICDM | 2 |
| 2002 | A Parameterless Method for Efficiently Discovering Clusters of Arbitrary Shape in Large DatasetsabstractClustering is the problem of grouping data based on similarity and consists of maximizing the intra-group similarity while minimizing the inter-group similarity. The problem Of clustering data sets is also known as unsupervised classification, since no class labels are given. However, all existing clustering algorithms require some parameters to steer the clustering process, such as the famous k for the number of expected clusters, which constitutes a supervision of a sort. We present in this paper a new, efficient, fast and scalable clustering algorithm that clusters over a range of resolutions and finds a potential optimum clustering without requiring any parameter input. Our experiments show that our algorithm outperforms most existing clustering algorithms in quality and speed for large data sets. Andrew Foss, Osmar R. Zaïane |
ICDM | 2 |
| 2002 | Clustering Spatial Data when Facing Physical ConstraintsabstractClustering spatial data is a well-known problem that has been extensively studied to find hidden patterns or meaningful sub-groups and has many applications such as satellite imagery, geographic information systems, medical image analysis, etc. Although many methods have been proposed in the literature, very few have considered constraints such that physical obstacles and bridges linking clusters may have significant consequences on the effectiveness of the clustering. Taking into account these constraints during the clustering process is costly, and the effective modeling of the constraints is of paramount importance for good performance. In this paper we define the clustering problem in the presence of constraints - obstacles and crossings - and investigate its efficiency and effectiveness for large databases. In addition, we introduce a new approach to model these constraints to prune the search space and reduce the number of polygons to test during clustering. The algorithm DBCluC we present detects clusters of arbitrary shape and is insensitive to noise and the input order Its average running complexity is O(NlogN) where N is the number of data objects. Osmar R. Zaïane |
ICDM | 1 |
| 2002 | Clustering Spatial Data in the Presence of Obstacles: a Density-Based ApproachabstractClustering spatial data is a well-known problem that has been extensively studied. Grouping similar data in large 2-dimensional spaces to find hidden patterns or meaningful sub-groups has many applications such as satellite imagery, geographic information systems, medical image analysis, marketing, computer visions, etc. Although many methods have been proposed in the literature, very few have considered physical obstacles that may have significant consequences on the effectiveness of the clustering. Taking into account these constraints during the clustering process is costly and the modeling of the constraints is paramount for good performance. In this paper, we investigate the problem of clustering in the presence of constraints such as physical obstacles and introduce a new approach to model these constraints using polygons. We also propose a strategy to prune the search space and reduce the number of polygons to test during clustering. We devise a density-based clustering algorithm, DBCluC, which takes advantage of our constraint modeling to efficiently cluster data objects while considering all physical constraints. The algorithm can detect clusters of arbitrary shape and is insensitive to noise, the input order and the difficulty of constraints. Its average running complexity is O(NlogN) where N is the number of data points. Osmar R. Zaïane |
IDEAS | 1 |
| 2002 | On Data Clustering Analysis: Scalability, Constraints, and Validation
Osmar R. Zaïane, Andrew Foss |
PAKDD | 1 |
| 2002 | Guest Editor's Introduction - Special Issue on Multimedia Data Mining
Osmar R. Zaïane |
J. Intell. Inf. Syst. | 1 |
| 2001 | Towards a Novel OLAP Interface for Distributed Data Warehouses
Ayman Ammoura, Osmar R. Zaïane, Randy Goebel |
DaWaK | 2 |
| 2001 | Fast Parallel Association Rule Mining without Candidacy GenerationabstractIn this paper we introduce a new parallel algorithm MLFPT (multiple local frequent pattern tree) for parallel mining of frequent patterns, based on FP-growth mining, that uses only two full I/O scans of the database, eliminating the need for generating candidate items, and distributing the work fairly among processors. We have devised partitioning strategies at different stages of the mining process to achieve near optimal balancing between processors. We have successfully tested our algorithm on datasets larger than 50 million transactions. Osmar R. Zaïane, Paul Lu |
ICDM | 1 |
| 2001 | Building virtual web views
Osmar R. Zaïane |
Data Knowl. Eng. | 1 |
| 2000 | Mining Recurrent Items in Multimedia with Progressive Resolution RefinementabstractDespite the overwhelming amounts of multimedia data recently generated and the significance of such data, very few people have systematically investigated multimedia data mining. With our previous studies on content-based retrieval of visual artifacts, we study in this paper the methods for mining content-based associations with recurrent items and with spatial relationships from large visual data repositories. A progressive resolution refinement approach is proposed in which frequent item-sets at rough resolution levels are mined, and progressively, finer resolutions are mined only on the candidate frequent items-sets derived from mining rough resolution levels. Such a multi-resolution mining strategy substantially reduces the overall data mining cost without loss of the quality and completeness of the results. Osmar R. Zaïane, Jiawei Han 0001 |
ICDE | 1 |
| 2000 | Multimedia data mining (workshop session - title only)abstractNo abstract available. Simeon J. Simoff, Osmar R. Zaïane |
KDD | 2 |
| 1998 | MultiMediaMiner: A System Prototype for Multimedia Data MiningabstractMultimedia data mining is the mining of high-level multimedia information and knowledge from large multimedia databases. A multimedia data mining system prototype, MultiMediaMiner, has been designed and developed. It includes the construction of a multimedia data cube which facilitates multiple dimensional analysis of multimedia data, primarily based on visual content, and the mining of multiple kinds of knowledge, including summarization, comparison, classification, association, and clustering. Osmar R. Zaïane, Jiawei Han 0001, Ze-Nian Li, Sonny Han Seng Chee, Jenny Chiang |
SIGMOD Conference | 1 |
| 1996 | DBMiner: A System for Mining Knowledge in Large Relational Databases
Jiawei Han 0001, Yongjian Fu 0001, Wei Wang 0009, Jenny Chiang, Wan Gong, Krzysztof Koperski, Deyi Li, Amynmohamed Rajan, Nebojsa Stefanovic, Betty Xia, Osmar R. Zaïane |
KDD | 12 |
| 1996 | DBMiner: Interactive Mining of Multiple-Level Knowledge in Relational DatabasesabstractBased on our years-of-research, a data mining system, DB-Miner, has been developed for interactive mining of multiple-level knowledge in large relational databases. The system implements a wide spectrum of data mining functions, including generalization, characterization, association, classification, and prediction. By incorporation of several interesting data mining techniques, including attribute-oriented induction, progressive deepening for mining multiple-level rules, and meta-rule guided knowledge mining, the system provides a user-friendly, interactive data mining environment with good performance. Jiawei Han 0001, Yongjian Fu 0001, Wei Wang 0009, Jenny Chiang, Osmar R. Zaïane, Krzysztof Koperski |
SIGMOD Conference | 5 |
| 1995 | Resource and Knowledge Discovery in Global Information Systems: A Preliminary Design and Experiment
Osmar R. Zaïane, Jiawei Han 0001 |
KDD | 1 |