Aijun An

dblp:99/3459 · DBLP profile ↗
← Back
116ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0003-1765-5751ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 71 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 60 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 5 since 2021Theory of computation · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Adapt-As-You-Walk Through the Clouds: Training-Free Online Test-Time Adaptation of 3D Vision-Language Foundation Models
abstract
3D Vision-Language Foundation Models (VLFMs) have demonstrated strong generalization and zero-shot recognition capabilities in open-world point cloud processing tasks. However, their performance often degrades in practical scenarios where data are noisy, incomplete, or drawn from distributions that differ from the training data. To address this challenge, we propose Uni-Adapter, a novel training-free online test-time adaptation (TTA) strategy for 3D VLFMs based on dynamic prototype learning. Uni-Adapter maintains a 3D cache that stores class-specific cluster centers as prototypes, which are continuously updated to capture intra-class variability under heterogeneous data distributions. These dynamic prototypes serve as anchors for cache-based logit computation through similarity scoring. In parallel, a graph-based label smoothing module models inter-prototype similarities to enforce label consistency among related prototypes. Finally, predictions from the original 3D VLFM and the refined 3D cache are unified through entropy-weighted aggregation to ensure reliable adaptation. Without retraining, Uni-Adapter effectively mitigates distribution shifts and achieves state-of-the-art performance across diverse 3D benchmarks and multiple 3D VLFMs, improving performance on ModelNet-40C by 10.55%, ScanObjectNN-C by 8.26%, and ShapeNet-C by 4.49% over the source 3D VLFMs.
Mehran Tamjidi, Hamidreza Dastmalchi, Mohammadreza Alimoradijazi, Ali Cheraghian, Aijun An, Morteza Saberi
AAAI5
2025 Canada AI Day
abstract
Artificial Intelligence (AI) research in Canada is driving a vibrant ecosystem of startups and industry innovation, thanks significantly to the leadership of Canada CIFAR AI Chairs. Canada AI Day at KDD 2025 will showcase cutting-edge research by some of these leading experts, with a focus on building ethical, interpretable, and accessible AI systems. The event will feature a series of invited talks followed by a panel discussion bringing together academic, government, and industry researchers to address the challenges and opportunities in developing effective and responsible AI.
Filippo Maria Sposini, Aijun An
KDD (2)2
2025 Test-Time Adaptation of 3D Point Clouds via Denoising Diffusion Models
abstract
Test-time adaptation (TTA) of 3D point clouds is crucial for mitigating discrepancies between training and testing samples in real-world scenarios, particularly when handling corrupted point clouds. LiDAR data, for instance, can be affected by sensor failures or environmental factors, causing domain gaps. Adapting models to these distribution shifts online is crucial, as training for every possible variation is impractical. Existing methods often focus on fine-tuning pre-trained models based on self-supervised learning or pseudo-labeling, which can lead to forgetting valuable source domain knowledge over time and reduce generalization on future tests. In this paper, we introduce a novel 3D test-time adaptation method, termed 3DD-TTA, which stands for 3D Denoising Diffusion Test-Time Adaptation. This method uses a diffusion strategy that adapts input point cloud samples to the source domain while keeping the source model parameters intact. The approach uses a Variational Autoencoder (VAE) to encode the corrupted point cloud into a shape latent and latent points. These latent points are corrupted with Gaussian noise and subjected to a denoising diffusion process. During this process, both the shape latent and latent points are updated to preserve fidelity, guiding the denoising toward generating consistent samples that align more closely with the source domain. We conduct extensive experiments on the ShapeNet dataset and investigate its generalizability on ModelNet40 and ScanObjectNN, achieving state-of-the-art results. The code has been released at https://github.com/hamidreza-dastmalchi/3DD-Tta.
Hamidreza Dastmalchi, Aijun An, Ali Cheraghian, Shafin Rahman, Sameera Ramasinghe
WACV2
2025 Disease Outbreak Detection and Forecasting: A Review of Methods and Data Sources
abstract
Infectious diseases occur when pathogens from other individuals or animals infect a person, causing harm to both individuals and society. Outbreaks of such diseases can pose a significant threat to human health. However, early detection and tracking of these outbreaks have the potential to reduce mortality rates. To address these threats, public health authorities have endeavored to establish comprehensive mechanisms for collecting disease data. Many countries have implemented infectious disease surveillance systems, with epidemic detection as a primary objective. The clinical healthcare system, local/state health agencies, federal agencies, academic/professional groups, and collaborating governmental entities all play pivotal roles within this system. Moreover, search engines and social media platforms can serve as valuable tools for monitoring disease trends. The Internet and social media have become significant platforms where users share information about their preferences and relationships. This real-time information can be harnessed to gauge the influence of ideas and societal opinions, proving highly useful across various domains and research areas, such as marketing campaigns, financial predictions, and public health. This article provides a review of the existing standard methods developed by researchers for detecting outbreaks using time series data. These methods leverage various data sources, including conventional data sources and social media data or Internet data sources. The review particularly concentrates on works published within the timeframe of 2015 to 2022.
Ghazaleh Babanejad Dehaki, Aijun An, Manos Papagelis
ACM Trans. Comput. Heal.2
2025 Topology-aware GPU job scheduling with deep reinforcement learning and heuristics
Hajer Ayadi, Aijun An, Hossein Pourmedheji, Junjie Deng, Jimmy Huang 0001, Michael Feiman
J. Parallel Distributed Comput.2
2025 Optimizing domain-generalizable ReID through non-parametric normalization
Amran Bhuiyan, Aijun An, Jimmy Huang 0001, Jialie Shen 0001
Pattern Recognit.2
2024 DAMOCRO: A Data Migration Framework Using Online Classification and Reordering
abstract
This paper introduces DAMOCRO, a data migration framework using online classification and tuple reordering to improve throughput and decrease the costs of data migration. The DAMOCRO workflow consists of four main steps. First, it classifies records into subgroups to maximize the similarity within each group. Next, it reorders tuples within these groups, ensuring that similar tuples are adjacent. Subsequently, column-wise compression is applied to each group. Finally, the compressed data is transferred from the source to the target machine. The initial two steps enhance the compression ratio, thereby boosting throughput and reducing costs. Our evaluations on five real-world datasets and two benchmark datasets, show that the online classification process in DAMOCRO improves throughput by more than 24% and reduces costs by over 19% compared to baselines. Besides, implementing reordering based on functional dependencies brings an additional cost reduction ranging from 10% to 60%, while also enhancing throughput.
Zhongxin Hu, Xingjian Mao, Jingfeng Pan, Yunfei Peng, Aijun An, Xiaohui Yu 0001, Dariusz Jania
CIKM6
2024 IGMG: Instance-guided multi-granularity for domain generalizable person re-identification
Amran Bhuiyan, Jimmy Huang 0001, Aijun An
Comput. Vis. Image Underst.3
2024 The Role of Preprocessing for Word Representation Learning in Affective Tasks
abstract
Affective tasks, including sentiment analysis, emotion classification, and sarcasm detection have drawn a lot of attention in recent years due to a broad range of useful applications in various domains. The main goal of affect detection tasks is to recognizestatessuch as mood, sentiment, and emotions from textual data (e.g., news articles or product reviews). Despite the importance of utilizingpreprocessingsteps in different stages (i.e., word representation learning and building a classification model) of affect detection tasks, this topic has not been studied well. To that end, we explore whether applying various preprocessing methods (stemming, lemmatization, stopword removal, punctuation removal and so on) and their combinations in different stages of the affect detection pipeline can improve the model performance. The are many preprocessing approaches that can be utilized in affect detection tasks. However, their influence on the final performance depends on the type of preprocessing and the stages that they are applied. Moreover, the preprocessing impacts vary across different affective tasks. Our analysis provides thorough insights into how preprocessing steps can be applied in building an effect detection pipeline and their respective influence on performance.
Nastaran Babanejad, Heidar Davoudi, Ameeta Agrawal, Aijun An, Manos Papagelis
IEEE Trans. Affect. Comput.4
2024 A Survey on Graph Representation Learning Methods
abstract
Graph representation learning has been a very active research area in recent years. The goal of graph representation learning is to generate graph representation vectors that capture the structure and features of large graphs accurately. This is especially important because the quality of the graph representation vectors will affect the performance of these vectors in downstream tasks such as node classification, link prediction and anomaly detection. Many techniques have been proposed for generating effective graph representation vectors, which generally fall into two categories: traditional graph embedding methods and graph neural network (GNN)–based methods. These methods can be applied to both static and dynamic graphs. A static graph is a single fixed graph, whereas a dynamic graph evolves over time and its nodes and edges can be added or deleted from the graph. In this survey, we review the graph-embedding methods in both traditional and GNN-based categories for both static and dynamic graphs and include the recent papers published until the time of submission. In addition, we summarize a number of limitations of GNNs and the proposed solutions to these limitations. Such a summary has not been provided in previous surveys. Finally, we explore some open and ongoing research directions for future work.
Shima Khoshraftar, Aijun An
ACM Trans. Intell. Syst. Technol.2
2023 A Hybrid Outbreak Detection using Ontology-based Data Collection from Social Media
abstract
Rapidly spreading diseases pose a significant threat, leading to substantial loss of life and economic devastation, as seen in the global COVID-19 outbreaks. Developing disease prediction models is crucial for preemptive pandemic control and minimizing their impact. As internet accessibility grows through computers and mobile devices, social media platforms provide a direct conduit to disseminate vital health information to the public. Unlike traditional methods that rely on bureaucratic channels, these platforms offer accurate and timely information distribution. We propose a framework that employs ontology to identify these symptoms and gather relevant tweets. Subsequently, the XGBoost-BiLSTM hybrid model harnesses this data to predict the count of infected cases. This hybrid model capitalizes on XGBoost’s prowess in handling limited dataset sizes, a prevalent challenge during outbreaks with insufficient time series data. Moreover, it enriches data for BiLSTM, amplifying its efficacy in predicting and monitoring outbreaks. To construct our dataset, we extracted tweets discussing symptoms from six distinct infectious disease outbreaks (Ebola, Zika, MERS, H1N1, Chikungunya, COVID-19) spanning from 2012 to 2021. Our results demonstrate that the proposed hybrid model outperforms nine cutting-edge and baseline models. This advancement can significantly assist health authorities in minimizing fatalities and preparing preemptively for potential outbreaks.
Ghazaleh Babanejad Dehaki, Aijun An, Heidar Davoudi
BIBM2
2023 A Case Study on ChatGPT Question Generation
abstract
The advent of transformers and the subsequent development of Large Language Models (LLMs) based on these technologies has revolutionized the field of Natural Language Processing (NLP). These models are able to understand and generate coherent natural language and hold conversations with humans continuously. Meanwhile, ChatGPT has become famous among many LLMs for its general-purpose characteristics and versatility. With that in mind, we investigate the capabilities of ChatGPT, which is very successful in many downstream NLP tasks on the task of Question Generation (QG). In particular, our experiments show that appropriate context through our designed prompts makes ChatGPT an appropriate tool for accurately performing the QG task. We compare ChatGPT’s question generation results with the state-of-the-art models, particularly on the SQuAD and car manual datasets. The results show that ChatGPT is able to compete with or even outperform some of the baseline models. Furthermore, we illustrate that we may improve ChatGPT through additional fine-tuning of the prompts. Finally, we also investigate the use of ChatGPT to evaluate QG models. While the use of ChatGPT for such purposes is still in its early stages, our results demonstrate that ChatGPT can potentially be a strong QG accuracy evaluator comparable to human evaluators.
Winston Chan, Aijun An, Heidar Davoudi
IEEE Big Data2
2023 Question Generation Using Sequence-to-Sequence Model with Semantic Role Labels
abstract
Automatic generation of questions from text has gained increasing attention due to its useful applications.We propose a novel question generation method that combines the benefits of rule-based and neural sequence-to-sequence (Seq2Seq) models.The proposed method can automatically generate multiple questions from an input sentence covering different views of the sentence as in rule-based methods, while more complicated "rules" can be learned via the Seq2Seq model.The method utilizes semantic role labeling to convert training examples into their semantic representations, and then trains a Seq2Seq model over the semantic representations.Our extensive experiments on three realworld data sets show that the proposed method significantly improves the state-of-the-art neural question generation approaches.
Alireza Naeiji, Aijun An, Heidar Davoudi, Marjan Delpisheh, Muath Alzghool
EACL2
2022 Temporal Graph Representation Learning via Maximal Cliques
abstract
Graph Neural Networks (GNNs) have been proposed to learn graph representations for various graph mining tasks such as link prediction and node classification. These methods aggregate information from neighbors of a node to generate the node representation vector. Temporal GNN models consider the temporal and neighborhood information of nodes. However, few temporal GNN methods consider network substructures such as triads and cliques. In this paper, we present a temporal GNN-based method that generates node embeddings by aggregating neighbors of a node that exist in the maximal cliques of the graph containing the node. The reason for considering neighbors that form a maximal clique with the target node is that nodes in a maximal clique are highly connected to each other and most likely share similar characteristics. In addition, we consider the time dependency of nodes by generating temporal walks on the cliques such that in these walks the time order of the nodes is respected. The node embedding is based on the aggregation of the node’s temporal walks. Our experiments on seven datasets show the effectiveness of our method in both link prediction and node classification tasks. Furthermore, our method is faster than other baselines we evaluate.
Shima Khoshraftar, Aijun An, Nastaran Babanejad
IEEE Big Data2
2021 Centrality-based Interpretability Measures for Graph Embeddings
abstract
Many real-world data are considered as graphs, such as computer networks, social networks and protein-protein interaction networks. Graph embedding methods are powerful tools for representing large graphs in various domains. A graph embedding method projects the components of a graph, such as its nodes or edges, into a vector space with a lower dimensionality than the adjacency matrix of the graph, and aims to preserve the characteristics of the graph. The generated embedding vectors have been utilized in various graph mining applications such as node classification, link prediction and anomaly detection. Despite the wide success of the graph embedding methods, little study has been done to facilitate a better understanding of the graph embeddings. In this paper, inspired by advancements in interpreting word embeddings, we propose two interpretability measures to quantify the interpretability of graph embeddings by leveraging useful network centrality properties and perform comparisons of different graph embedding methods. Using these scores, we can provide insights into the representational power of graph embedding methods.
Shima Khoshraftar, Sedigheh Mahdavi, Aijun An
DSAA3
2021 ZipLine: An Optimized Algorithm for the Elastic Bulk Synchronous Parallel Model
abstract
The bulk synchronous parallel (BSP) is a celebrated synchronization model for distributed training of deep learning models. A shortcoming of the BSP is that it requires workers to wait for the straggler at every iteration. Therefore, employing BSP increases the waiting time of the faster workers of a cluster and results in an overall prolonged training time. To ameliorate this shortcoming of BSP, we proposed ElasticBSP [1], a model that aims to relax its strict synchronization requirement with an elastic synchronization by allowing delayed synchronization to minimize the waiting time. ELASTICBSP is realized by the algorithm named ZipLine. In this work, we show the theoretical proof of ZipLine and further propose algorithmic and implementation optimizations of ZipLine, namely ZipLineOpt and Ziplineoptbs, which reduce the time complexity of ZipLine to linearithmic time. The experiments show that ZipLineOpt and ZipLineOptBs enable the scalability of ElasticBSP. Further experimental evaluation on large deep neural networks on large ImageNet dataset demonstrate that our proposed Elas-ticbspmodel, materialized by the proposed optimized ZipLine variants, converges faster and to a higher accuracy than the predominant BSP.
Xing Zhao 0004, Manos Papagelis, Aijun An, Bao Xin Chen, Junfeng Liu 0005, Yonggang Hu
DSAA3
2021 ZipLine: an optimized algorithm for the elastic bulk synchronous parallel model
Xing Zhao 0004, Manos Papagelis, Aijun An, Bao Xin Chen, Junfeng Liu 0005, Yonggang Hu
Mach. Learn.3
2021 A Survey of Self-Sovereign Identity Ecosystem
abstract
Self-sovereign identity is the next evolution of identity management models. This survey takes a journey through the origin of identity, defining digital identity and progressive iterations of digital identity models leading up to self-sovereign identity. It then states the relevant research initiatives, platforms, projects, and regulatory frameworks, as well as the building blocks including decentralized identifiers, verifiable credentials, distributed ledger, and various privacy engineering protocols. Finally, the survey provides an overview of the key challenges and research opportunities around self-sovereign identity.
Uyen Trang Nguyen, Aijun An
Secur. Commun. Networks3
2021 Extending Isolation Forest for Anomaly Detection in Big Data via K-Means
abstract
Industrial Information Technology infrastructures are often vulnerable to cyberattacks. To ensure security to the computer systems in an industrial environment, it is required to build effective intrusion detection systems to monitor the cyber-physical systems (e.g., computer networks) in the industry for malicious activities. This article aims to build such intrusion detection systems to protect the computer networks from cyberattacks. More specifically, we propose a novel unsupervised machine learning approach that combines the K-Means algorithm with the Isolation Forest for anomaly detection in industrial big data scenarios. Since our objective is to build the intrusion detection system for the big data scenario in the industrial domain, we utilize the Apache Spark framework to implement our proposed model that was trained in large network traffic data (about 123 million instances of network traffic) stored in Elasticsearch. Moreover, we evaluate our proposed model on the live streaming data and find that our proposed system can be used for real-time anomaly detection in the industrial setup. In addition, we address different challenges that we face while training our model on large datasets and explicitly describe how these issues were resolved. Based on our empirical evaluation in different use cases for anomaly detection in real-world network traffic data, we observe that our proposed system is effective to detect anomalies in big data scenarios. Finally, we evaluate our proposed model on several academic datasets to compare with other models and find that it provides comparable performance with other state-of-the-art approaches.
Md. Tahmid Rahman Laskar, Jimmy Huang 0001, Vladan Smetana, Chris Stewart, Kees Pouw, Aijun An, Stephen Chan, Lei Liu 0047
ACM Trans. Cyber Phys. Syst.6
2021 Paywall Policy Learning in Digital News Media
abstract
Subscription-based online newspapers usually offer non-subscribed users a certain number of free articles in a period of time, and then directs them to a page (called paywall) asking for subscription. This approach (also known as metered or fixed paywall) does not consider the user's reading history nor the articles that the user may read in the future, and consequently, it may disengage many potential subscribers. To that end, we propose adaptive paywall mechanisms to make optimal paywall decisions (i.e., showing the article or the paywall) by balancing the benefit of showing the article against that of presenting the paywall. We define the notions of utility and cost which are used to define an objective function for the optimal paywall decision problem. We propose the Lookahead policy (LAP) and QPaywall policy (QP) as two data-driven approaches to solve the adaptive paywall problem. While the LAP method makes paywall decisions on the fly by simulating trajectories of article requests using Monte Carlo sampling, the QP approach is based on reinforcement learning and learns a neural network-based action-value (Q) function for this purpose. We compare advantages of the proposed approaches and discuss the practical considerations of using them in a real environment. Empirical studies on a real dataset from a major newspaper in Canada show that the proposed methods outperform several baseline approaches in terms of various business objectives.
Heidar Davoudi, Zana Rashidi, Aijun An, Morteza Zihayat, Gordon Edall
IEEE Trans. Knowl. Data Eng.3
2020 A Comprehensive Analysis of Preprocessing for Word Representation Learning in Affective Tasks
abstract
Affective tasks such as sentiment analysis, emotion classification and sarcasm detection have been popular in recent years due to abundance of user-generated data, accurate computational linguistic models, and broad range of relevant applications in various domains.At the same time, many studies have highlighted the importance of text preprocessing, as an integral step to any natural language processing prediction model and downstream task.While preprocessing in affective systems is well-studied, preprocessing in word vector based models applied to affective systems, is not.To address this limitation, we conduct a comprehensive analysis of the role of preprocessing techniques in affective analysis based on word vector models.Our analysis is the first of its kind and provides useful insights of the importance of each preprocessing technique when applied at the training phase, commonly ignored in pretrained word vector models, and/or at the downstream task phase.
Nastaran Babanejad, Ameeta Agrawal, Aijun An, Manos Papagelis
ACL3
2020 Affective and Contextual Embedding for Sarcasm Detection
abstract
Automatic sarcasm detection from text is an important classification task that can help identify the actual sentiment in user-generated data, such as reviews or tweets.Despite its usefulness, sarcasm detection remains a challenging task, due to a lack of any vocal intonation or facial gestures in textual data.To date, most of the approaches to addressing the problem have relied on hand-crafted affect features, or pre-trained models of non-contextual word embeddings, such as Word2vec.However, these models inherit limitations that render them inadequate for the task of sarcasm detection.In this paper, we propose two novel deep neural network models for sarcasm detection, namely ACE 1 and ACE 2. Given as input a text passage, the models predict whether it is sarcastic (or not).Our models extend the architecture of BERT by incorporating both affective and contextual features.To the best of our knowledge, this is the first attempt to directly alter BERT's architecture and train it from scratch to build a sarcasm classifier.Extensive experiments on different datasets demonstrate that the proposed models outperform state-of-the-art models for sarcasm detection with significant margins.
Nastaran Babanejad, Heidar Davoudi, Aijun An, Manos Papagelis
COLING3
2020 Learning to Determine the Quality of News Headlines
abstract
Today, most newsreaders read the online version of news articles rather than traditional paper-based newspapers. Also, news media publishers rely heavily on the income generated from subscriptions and website visits made by newsreaders. Thus, online user engagement is a very important issue for online newspapers. Much effort has been spent on writing interesting headlines to catch the attention of online users. On the other hand, headlines should not be misleading (e.g., clickbaits); otherwise, readers would be disappointed when reading the content. In this paper, we propose four indicators to determine the quality of published news headlines based on their click count and dwell time, which are obtained by website log analysis. Then, we use soft target distribution of the calculated quality indicators to train our proposed deep learning model which can predict the quality of unpublished news headlines. The proposed model not only processes the latent features of both headline and body of the article to predict its headline quality but also considers the semantic relation between headline and body as well. To evaluate our model, we use a real dataset from a major Canadian newspaper. Results show our proposed model outperforms other state-of-the-art NLP models.
Amin Omidvar, Hossein Pourmodheji, Aijun An, Gordon Edall
ICAART (1)3
2020 Decentralized and Privacy-Preserving Key Management Model
abstract
Data centralization and the growing rate of security breaches and identity fraud have led us to seek privacy-preserving and decentralized identity management solutions. The success of decentralized models, such as Self-Sovereign Identity, hinges on the positive-sum combination of usability and security. To support this goal, we propose a decentralized system capable of performing key management operations including key generation, key backup, and key recovery. In addition we propose our preliminary solution for a decentralized identity verification protocol. To this end, we design a digital wallet that relies on Shamir's Secret sharing scheme and blockchain technology, and we present a number of the security parameters affecting our model.
Uyen Trang Nguyen, Aijun An
ISNCC3
2020 Adaptive Momentum Coefficient for Neural Network Optimization
Zana Rashidi, Kasra Ahmadi K. A., Aijun An, Xiaogang Wang 0007
ECML/PKDD (2)3
2020 Leveraging Transitions of Emotions for Sarcasm Detection
abstract
One popular thread of research in computational sarcasm detection involves modeling sarcasm as a contrast between positive and negative sentiment polarities or exploring more fine-grained categories of emotions such as happiness, sadness, surprise, and so on. Most current models, however, treat these affective features independently, without regard for the sequential information encoded among the affective states. In order to explore the role of transitions in affective states, we formulate the task of sarcasm detection as a sequence classification problem by leveraging the natural shifts in various emotions over the course of a piece of text. Experiments conducted on datasets from two different genres suggest that our proposed approach particularly benefits datasets with limited labeled data and longer instances of text.
Ameeta Agrawal, Aijun An, Manos Papagelis
SIGIR2
2019 Dynamic Graph Embedding via LSTM History Tracking
abstract
Many real world networks are very large and constantly change over time. These dynamic networks exist in various domains such as social networks, traffic networks and biological interactions. To handle large dynamic networks in downstream applications such as link prediction and anomaly detection, it is essential for such networks to be transferred into a low dimensional space. Recently, network embedding, a technique that converts a large graph into a low-dimensional representation, has become increasingly popular due to its strength in preserving the structure of a network. Efficient dynamic network embedding, however, has not yet been fully explored. In this paper, we present a dynamic network embedding method that integrates the history of nodes over time into the current state of nodes. The key contribution of our work is 1) generating dynamic network embedding by combining both dynamic and static node information 2) tracking history of neighbors of nodes using LSTM 3) significantly decreasing the time and memory by training an autoencoder LSTM model using temporal walks rather than adjacency matrices of graphs which are the common practice. We evaluate our method in multiple applications such as anomaly detection, link prediction and node classification in datasets from various domains.
Shima Khoshraftar, Sedigheh Mahdavi, Aijun An, Yonggang Hu, Junfeng Liu 0005
DSAA3
2019 Dynamic Stale Synchronous Parallel Distributed Training for Deep Learning
abstract
Deep learning is a popular machine learning technique and has been applied to many real-world problems, ranging from computer vision to natural language processing. However, training a deep neural network is very time-consuming, especially on big data. It has become difficult for a single machine to train a large model over large datasets. A popular solution is to distribute and parallelize the training process across multiple machines using the parameter server framework. In this paper, we present a distributed paradigm on the parameter server framework called Dynamic Stale Synchronous Parallel (DSSP) which improves the state-of-the-art Stale Synchronous Parallel (SSP) paradigm by dynamically determining the staleness threshold at the run time. Conventionally to run distributed training in SSP, the user needs to specify a particular stalenes threshold as a hyper-parameter. However, a user does not usually know how to set the threshold and thus often finds a threshold value through trial and error, which is time-consuming. Based on workers' recent processing time, our approach DSSP adaptively adjusts the threshold per iteration at running time to reduce the waiting time of faster workers for synchronization of the globally shared parameters (the weights of the model), and consequently increases the frequency of parameters updates (increases iteration through-put), which speedups the convergence rate. We compare DSSP with other paradigms such as Bulk Synchronous Parallel (BSP), Asynchronous Parallel (ASP), and SSP by running deep neural networks (DNN) models over GPU clusters in both homogeneous and heterogeneous environments. The results show that in a heterogeneous environment where the cluster consists of mixed models of GPUs, DSSP converges to a higher accuracy much earlier than SSP and BSP and performs similarly to ASP. In a homogeneous distributed cluster, DSSP has more stable and slightly better performance than SSP and ASP, and converges much faster than BSP.
Xing Zhao 0004, Aijun An, Junfeng Liu 0005, Bao Xin Chen
ICDCS2
2019 Elastic Bulk Synchronous Parallel Model for Distributed Deep Learning
abstract
The bulk synchronous parallel (BSP) is a celebrated synchronization model for general-purpose parallel computing that has successfully been employed for distributed training of machine learning models. A prevalent shortcoming of the BSP is that it requires workers to wait for the straggler at every iteration. To ameliorate this shortcoming of classic BSP, we propose ELASTICBSP a model that aims to relax its strict synchronization requirement. The proposed model offers more flexibility and adaptability during the training phase, without sacrificing on the accuracy of the trained model. We also propose an efficient method that materializes the model, named ZIPLINE. The algorithm is tunable and can effectively balance the trade-off between quality of convergence and iteration throughput, in order to accommodate different environments or applications. A thorough experimental evaluation demonstrates that our proposed ELASTICBSP model converges faster and to a higher accuracy than the classic BSP. It also achieves comparable (if not higher) accuracy than the other sensible synchronization models.
Xing Zhao 0004, Manos Papagelis, Aijun An, Bao Xin Chen, Junfeng Liu 0005, Yonggang Hu
ICDM3
2019 A utility-based news recommendation system
Morteza Zihayat, Anteneh Ayanso, Xing Zhao 0004, Heidar Davoudi, Aijun An
Decis. Support Syst.5
2018 dynnode2vec: Scalable Dynamic Network Embedding
abstract
Network representation learning in low dimensional vector space has attracted considerable attention in both academic and industrial domains. Most real-world networks are dynamic with addition/deletion of nodes and edges. The existing graph embedding methods are designed for static networks and they cannot capture evolving patterns in a large dynamic network. In this paper, we propose a dynamic embedding method, dynnode2vec, based on the well-known graph embedding method node2vec. Node2vec is a random walk based embedding method for static networks. Applying static network embedding in dynamic settings has two crucial problems: 1) Generating random walks for every time step is time consuming 2) Embedding vector spaces in each timestamp are different. In order to tackle these challenges, dynnode2vec uses evolving random walks and initializes the current graph embedding with previous embedding vectors. We demonstrate the advantages of the proposed dynamic network embedding by conducting empirical evaluations on several large dynamic network datasets.
Sedigheh Mahdavi, Shima Khoshraftar, Aijun An
IEEE BigData3
2018 Learning Emotion-enriched Word Representations
abstract
Most word representation learning methods are based on the distributional hypothesis in linguistics, according to which words that are used and occur in the same contexts tend to possess similar meanings. As a consequence, emotionally dissimilar words, such as “happy” and “sad” occurring in similar contexts would purport more similar meaning than emotionally similar words, such as “happy” and “joy”. This complication leads to rather undesirable outcome in predictive tasks that relate to affect (emotional state), such as emotion classification and emotion similarity. In order to address this limitation, we propose a novel method of obtaining emotion-enriched word representations, which projects emotionally similar words into neighboring spaces and emotionally dissimilar ones far apart. The proposed approach leverages distant supervision to automatically obtain a large training dataset of text documents and two recurrent neural network architectures for learning the emotion-enriched representations. Through extensive evaluation on two tasks, including emotion classification and emotion similarity, we demonstrate that the proposed representations outperform several competitive general-purpose and affective word representations.
Ameeta Agrawal, Aijun An, Manos Papagelis
COLING2
2018 Decoupling the Layers in Residual Networks
Ricky Fok, Aijun An, Zana Rashidi, Xiaogang Wang 0007
ICLR (Poster)2
2018 Adaptive Paywall Mechanism for Digital News Media
abstract
Many online news agencies utilize the paywall mechanism to increase reader subscriptions. This method offers a non-subscribed reader a fixed number of free articles in a period of time (e.g., a month), and then directs the user to the subscription page for further reading. We argue that there is no direct relationship between the number of paywalls presented to readers and the number of subscriptions, and that this artificial barrier, if not used well, may disengage potential subscribers and thus may not well serve its purpose of increasing revenue. Moreover, the current paywall mechanism neither considers the user browsing history nor the potential articles which the user may visit in the future. Thus, it treats all readers equally and does not consider the potential of a reader in becoming a subscriber. In this paper, we propose an adaptive paywall mechanism to balance the benefit of showing an article against that of displaying the paywall (i.e., terminating the session). We first define the notion of cost and utility that are used to define an objective function for optimal paywall decision making. Then, we model the problem as a stochastic sequential decision process. Finally, we propose an efficient policy function for paywall decision making. The experimental results on a real dataset from a major newspaper in Canada show that the proposed model outperforms the traditional paywall mechanism as well as the other baselines.
Heidar Davoudi, Aijun An, Morteza Zihayat, Gordon Edall
KDD2
2018 Improving Real-Time Pedestrian Detection Using Adaptive Confidence Thresholding and Inter-Frame Correlation
abstract
The pedestrian detection algorithms form a key component in the multiple pedestrian tracking (MPT) systems. Despite efforts to detect a pedestrian accurately, it is still a challenging task. We propose a novel and efficient online method to improve the performance of the multiple person/pedestrian detector by introducing novel post-processing steps. These steps use an adaptive approach to determine both area and confidence score constraints for the output of any given multiple pedestrian detector. In this paper, we focus on pedestrian detection in video surveillance applications that require an automated, accurate and precise pedestrian detection algorithm. We demonstrate that the new steps make the multiple pedestrian detector more accurate, precise and tolerant to false positive detections. This is illustrated by evaluating the performance of the proposed method in test video sequences taken from the Pedestrian Detection Challenge, Multiple Object Tracking Benchmark (MOT Challenge 2017).
Mufleh Al-Shatnawi, Vida Movahedi, Amir Asif, Aijun An
MMSP4
2018 Affective Representations for Sarcasm Detection
abstract
Sarcasm detection from text has gained increasing attention. While one thread of research has emphasized the importance of affective content in sarcasm detection, another avenue of research has explored the effectiveness of word representations. In this paper, we introduce a novel model for automated sarcasm detection in text, called Affective Word Embeddings for Sarcasm (AWES), which incorporates affective information into word representations. Extensive evaluation on sarcasm detection on six datasets across three domains of text (tweets, reviews and forum posts) demonstrates the effectiveness of the proposed model. The experimental results indicate that while sentiment affective representations yield best results on datasets comprising of short length text such as tweets, richer representations derived from fine-grained emotions are more suitable for detecting sarcasm from longer length documents such as product reviews and discussion forum posts.
Ameeta Agrawal, Aijun An
SIGIR2
2018 Machine learning and BIM visualization for maintenance issue classification and enhanced data collection
J. J. McArthur, Nima Shahbazi, Ricky Fok, Christopher Raghubar, Brandon Bortoluzzi, Aijun An
Adv. Eng. Informatics6
2017 Authority-based Team Discovery in Social Networks
Morteza Zihayat, Aijun An, Lukasz Golab, Mehdi Kargar, Jarek Szlichta
EDBT2
2017 Contrast Pattern Based Collaborative Behavior Recommendation for Life Improvement
Yan Chen 0021, Margot Lisa-Jing Yann, Heidar Davoudi, Joy Choi, Aijun An
PAKDD (2)5
2017 Time-Aware Subscription Prediction Model for User Acquisition in Digital News Media
abstract
User acquisition is one of the most challenging problems for online news providers. In fact, due to availability of different news media, users have a lot of choices in selecting the news source. To date, most of digital news portals have tried to approach the solution indirectly by targeting the user satisfaction through the recommendation systems. In contrast, we address the problem directly by identifying valuable visitors who are likely potential subscribers in the future. First, we suggest that the decision for subscription is not a sudden, instantaneous action, but is the informed decision based on positive experience with digital medium. As such, we propose effective engagement measures and show that they are effective in building the predictive model for subscription. We design a model that not only predicts the potential subscribers but also answers queries about the subscription occurrence time. The proposed model can be used to predict the subscription time and recommend accurately the “potential users” to the current marketing campaign. We evaluate the proposed model using a real dataset from The Globe and Mail which is a major newspaper in Canada. The experimental results show that the proposed model outperforms the traditional state-of-the-art approaches significantly.
Heidar Davoudi, Morteza Zihayat, Aijun An
SDM3
2017 Mining Evolving Data Streams with Particle Filters
abstract
We propose a particle filter‐based learning method, PF‐LR, for learning logistic regression models from evolving data streams. The method inherently handles concept drifts in a data stream and is able to learn an ensemble of logistic regression models with particle filtering. A key feature of PF‐LR is that in its resampling, step particles are sampled from the ones that maximize the classification accuracy on the current data batch. Our experiments show that PF‐LR gives good performance, even with relatively small batch sizes. It reacts to concept drifts quicker than conventional particle filters while being robust to noise. In addition, PF‐LR learns more accurate models and is more computationally efficient than the gradient descent method for learning logistic regression models. Furthermore, we evaluate PF‐LR on both synthetic and real data sets and find that PF‐LR outperforms some other state‐of‐the‐art streaming mining algorithms on most of the data sets tested.
Ricky Fok, Aijun An, Xiaogang Wang 0007
Comput. Intell.2
2017 Efficiently mining high utility sequential patterns in static and streaming data
abstract
High utility sequential pattern (HUSP) mining has emerged as a novel topic in data mining. Although some preliminary works have been conducted on this topic, they incur the problem of producing a large search space for high utility sequential patterns. In addition, they mainly focus on mining HUSPs in static databases and do not take streaming data into account, where unbounded data come continuously and often at a high speed. To efficiently deal with both problems, we propose a novel framework for mining high utility sequential patterns over static and streaming databases. In this regard, two efficient data structures named ItemUtilLists (Item Utility Lists) and HUSP-Tree (High Utility Sequential Pattern Tree) are proposed to maintain essential information for mining HUSPs in both offline and online fashions. In addition, a novel utility model called Sequence-Suffix Utility is proposed for effectively pruning the search space in HUSP mining. We propose an algorithm named HUSP-Miner (High Utility Sequential Pattern Miner) to find HUSPs in static databases efficiently. Then, a one-pass algorithm named HUSP-Stream (High Utility Sequential Pattern mining over Data Streams) is proposed to incrementally update ItemUtilLists and HUSP-Tree online and find HUSPs over data streams. To the best of our knowledge, HUSP-Stream is the first method to find HUSPs over data streams. Experimental results on both real and synthetic datasets show that HUSP-Miner outperforms the compared algorithms substantially in terms of execution time, memory usage and number of generated candidates. The experiments also demonstrate impressive performance of HUSP-Stream to update the data structures and discover HUSPs over data streams.
Morteza Zihayat, Cheng-Wei Wu, Aijun An, Vincent S. Tseng, Chien Lin
Intell. Data Anal.3
2017 Geodesic and contour optimization using conformal mapping
Ricky Fok, Aijun An, Xiaogang Wang 0007
J. Glob. Optim.2
2017 Memory-adaptive high utility sequential pattern mining over data streams
Morteza Zihayat, Yan Chen 0021, Aijun An
Mach. Learn.3
2016 Top-k utility-based gene regulation sequential pattern discovery
abstract
Sequential pattern mining has been used in bioinformatics to discover frequent gene regulation sequential patterns based on time course microarray datasets. While mining frequent sequences are important in biological studies for disease treatment, to date, most of the approaches do not consider the importance of the genes with respect to a disease being studied when identifying gene regulation sequential patterns. In addition, they focus on the more general up/down effects of genes in a microarray dataset and do not take into account the various degrees of expression during the mining process. As a result, the current techniques return too many sequences which may not be informative enough for biologists to explore relationships between the disease and underlying causes encoded in gene regulation sequences. In this paper, we propose a utility model by considering both the importance of genes with respect to a disease and their degrees of expression levels under a biological investigation. Then, we design a new method, called TU-SEQ, for identifying top-k high utility gene regulation sequential patterns from a time-course microarray dataset. The evaluation results show that our approach can effectively and efficiently discover key patterns representing meaningful gene regulation sequential patterns in a time course microarray dataset.
Morteza Zihayat, Heidar Davoudi, Aijun An
BIBM3
2016 Deep parallelization of parallel FP-growth using parent-child MapReduce
abstract
MapReduce is an important programming model for processing in distributed environments. Compared to other distributed programming models, MapReduce reduces communication overheads between computers and improves fault tolerance. However, the MapReduce model does not allow for automatic synchronization between jobs. A large number of data analytics algorithms use a recursive divide-and-conquer approach, which inherently allows for parallelism at each level of recursion. However, it is often difficult to parallelize such algorithms using the traditional MapReduce model if the process requires synchronization. In this paper we introduce Parent-Child MapReduce, a version of the MapReduce programming model that allows for MapReduce tasks to be created dynamically and synchronized in a hierarchical parent-child fashion. Using the Parallel FP-Growth (PFP) algorithm for mining frequent patterns as a reference, we show that Parent-Child MapReduce can be used to parallelize recursive divide-and-conquer algorithms using the MapReduce model and that this can lead to significant speed ups in the computational speed of such algorithms. Our evaluation shows that we can achieve 68% (or 3 times) performance gain when used with PFP.
Adetokunbo Makanju, Zahra Farzanyar, Aijun An, Nick Cercone, Zane Zhenhua Hu, Yonggang Hu
IEEE BigData3
2016 Distributed and parallel high utility sequential pattern mining
abstract
The problem of mining high utility sequential patterns (HUSP) has been studied recently. Existing solutions are mostly memory-based, which assume that data can fit into the main memory of a computer. However, with advent of big data, such an assumption does not hold any longer. Hence, existing algorithms are not applicable to the big data environments, where data are often distributed and too large to be dealt with by a single machine. In this paper, we propose a new framework for mining HUSPs in big data. A distributed and parallel algorithm called BigHUSP is proposed to discover HUSPs efficiently. At its heart, BigHUSP uses multiple MapReduce-like steps to process data in parallel. We also propose a number of pruning strategies to minimize search space in a distributed environment, and thus decrease computational and communication costs, while still maintaining correctness. Our experiments with real life and large synthetic datasets validate the effectiveness of BigHUSP for mining HUSPs from large sequence datasets.
Morteza Zihayat, Zane Zhenhua Hu, Aijun An, Yonggang Hu
IEEE BigData3
2016 Selective Co-occurrences for Word-Emotion Association
abstract
Emotion classification from text typically requires some degree of word-emotion association, either gathered from pre-existing emotion lexicons or calculated using some measure of semantic relatedness. Most emotion lexicons contain a fixed number of emotion categories and provide a rather limited coverage. Current measures of computing semantic relatedness, on the other hand, do not adapt well to the specific task of word-emotion association and therefore, yield average results. In this work, we propose an unsupervised method of learning word-emotion association from large text corpora, called Selective Co-occurrences (SECO), by leveraging the property of mutual exclusivity generally exhibited by emotions. Extensive evaluation, using just one seed word per emotion category, indicates the effectiveness of the proposed approach over three emotion lexicons and two state-of-the-art models of word embeddings on three datasets from different domains.
Ameeta Agrawal, Aijun An
COLING2
2016 Ranking Documents Through Stochastic Sampling on Bayesian Network-based Models: A Pilot Study
abstract
Using approximate inference techniques, we investigate in this paper the applicability of Bayesian Networks to the problem of ranking a large set of documents. Topology of the network is a bipartite. Network parameters (conditional probability distributions) are determined through an adoption of the weighting scheme tf-idf. Rank of a document with respect to a given query is defined as the corresponding posterior probability, which is estimated through performing Rejection Sampling. Experimental results suggest that performance of the model is at least comparable to the baseline ones such as BM25. The framework of this model potentially offers new and novel ways in weighting documents. Integrating the model with other ranking algorithms, meanwhile, is expected to bring in performance improvement in document ranking.
Xing Tan 0002, Jimmy Huang 0001, Aijun An
SIGIR3
2016 Detecting the Magnitude of Events from News Articles
abstract
Forced migration is increasingly becoming a global issue of concern. In this paper, we present an effective model of targeted event detection, as an essential step towards the forced migration detection problem. To date, most of the the approaches deal with the event detection in a general setting with the main objective of detecting the presence or onset of an event. However, we focus on analyzing the magnitude of a given event from a collection of text documents such as news articles from multiple sources. We use violence as an illustration as it is one of the most critical factors of forced migration. The recent advancements in semantic similarity measures are adopted to obtain relevant violence scores for each word in the vocabulary of news articles in an unsupervised manner. The resulting scores are then used to compute the average daily violence scores over a period of three months. Evaluation of the proposed model against a manually annotated data set yields a Pearson's correlation of 0.8. We also include a case study exploring the relationship between violence and key events.
Ameeta Agrawal, Raghavender Sahdev, Heidar Davoudi, Forouq Khonsari, Aijun An, Susan McGrath
WI5
2015 Meaningful keyword search in relational databases with large and complex schema
abstract
Keyword search over relational databases offers an alternative way to SQL to query and explore databases that is effective for lay users who may not be well versed in SQL or the database schema. This becomes more pertinent for databases with large and complex schemas. An answer in this context is a join tree spanning tuples containing the query's keywords. As there are potentially many answers to the query, and the user is often only interested in seeing the top-k answers, how to rank the answers based on their relevance is of paramount importance. We focus on the relevance of join as the fundamental means to rank answers. We devise means to measure relevance of relations and foreign keys in the schema over the information content of the database. This can be done offline with no need for external models. We compare the proposed measures against a gold standard we derive from a real workload over TPC-E and evaluate the effectiveness of our methods. Finally, we test the performance of our measures against existing techniques to demonstrate a marked improvement, and perform a user study to establish naturalness of the ranking of the answers.
Mehdi Kargar, Aijun An, Nick Cercone, Parke Godfrey, Jarek Szlichta, Xiaohui Yu 0001
ICDE2
2015 Ontology-Based Topic Labeling and Quality Prediction
Heidar Davoudi, Aijun An
ISMIS2
2015 A Bayesian model for canonical circuits in the neocortex for parallelized and incremental learning of symbol representations
Martin Dimkovski, Aijun An
Neurocomputing2
2015 Finding top-k, r-cliques for keyword search from graphs in polynomial delay
Mehdi Kargar, Aijun An
Knowl. Inf. Syst.2
2014 Topic Modeling Using Collapsed Typed Dependency Relations
Elnaz Delpisheh, Aijun An
PAKDD (2)2
2014 MeanKS: meaningful keyword search in relational databases with complex schema
abstract
Keyword search in relational databases was introduced in the last decade to assist users who are not familiar with a query language, the schema of the database, or the content of the data. An answer is a join tree of tuples that contains the query keywords. When searching a database with a complex schema, there are potentially many answers to the query. Therefore, ranking answers based on their relevance is crucial in this context. Prior work has addressed relevance based on the size of the answer or the IR scores of the tuples. However, this is not sufficient when searching a complex schema. We demonstrate MeanKS, a new system for meaningful keyword search over relational databases. The system first captures the user's interest by determining the roles of the keywords. Then, it uses schema-based ranking to rank join trees that cover the keyword roles. This uses the relevance of relations and foreign-key relationships in the schema over the information content of the database. In the demonstration, attendees can execute queries against the TPC-E warehouse and compare the proposed measures against a gold standard derived from a real workload over TPC-E to test the effectiveness of our methods.
Mehdi Kargar, Aijun An, Nick Cercone, Parke Godfrey, Jarek Szlichta, Xiaohui Yu 0001
SIGMOD Conference2
2014 Mining top-k high utility patterns over data streams
Morteza Zihayat, Aijun An
Inf. Sci.2
2014 Efficient Duplication Free and Minimal Keyword Search in Graphs
abstract
Keyword search over a graph searches for a subgraph that contains a set of query keywords. A problem with most existing keyword search methods is that they may produce duplicate answers that contain the same set of content nodes (i.e., nodes containing a query keyword) although these nodes may be connected differently in different answers. Thus, users may be presented with many similar answers with trivial differences. In addition, some of the nodes in an answer may contain query keywords that are all covered by other nodes in the answer. Removing these nodes does not change the coverage of the answer but can make the answer more compact. The answers in which each content node contains at least one unique query keyword are called minimal answers in this paper. We define the problem of finding duplication-free and minimal answers, and propose algorithms for finding such answers efficiently. Extensive performance studies using two large real data sets confirm the efficiency and effectiveness of the proposed methods.
Mehdi Kargar, Aijun An, Xiaohui Yu 0001
IEEE Trans. Knowl. Data Eng.2
2013 Finding Affordable and Collaborative Teams from a Network of Experts
abstract
Given an expert network, we tackle the problem of finding a team of experts that covers a set of required skills and also minimizes the communication cost as well as the personnel cost of the team. Since two costs need to be minimized, this is a bicriteria optimization problem. We show that the problem of minimizing these objectives is NP-hard. We use two approaches to solve this bicriteria optimization problem. In the first approach, we propose several (α, β)-approximation algorithms that receive a budget on one objective and minimizes the other objective within the budget with guaranteed performance bounds. In the second approach, an approximation algorithm is proposed to find a set of Pareto-optimal teams, in which each team is not dominated by other feasible teams in terms of the personnel and communication costs. The proposed approximation algorithms have provable performance bounds. Extensive experiments on real datasets demonstrate the effectiveness and scalability of the proposed algorithms.
Aijun An, Mehdi Kargar, Morteza Zihayat
SDM1
2013 Riding the tide of sentiment change: sentiment analysis with evolving online reviews
Yang Liu 0008, Xiaohui Yu 0001, Aijun An, Jimmy Huang 0001
World Wide Web3
2012 Efficient Top-k Keyword Search in Graphs with Polynomial Delay
abstract
A system for efficient keyword search in graphs is demonstrated. The system has two components, a search through only the nodes containing the input keywords for a set of nodes that are close to each other and together cover the input keywords and an exploration for finding how these nodes are related to each other. The system generates all or top-k answers in polynomial delay. Answers are presented to the user according to a ranking criterion so that the answers with nodes closer to each other are presented before the ones with nodes farther away from each other. In addition, the set of answers produced by our system is duplication free. The system uses two methods for presenting the final answer to the user. The presentation methods reveal relationships among the nodes in an answer through a tree or a multi-center graph. We will show that each method has its own advantages and disadvantages. The system is demonstrated using two challenging datasets, very large DBLP and highly cyclic Mondial. Challenges and difficulties in implementing an efficient keyword search system are also demonstrated.
Mehdi Kargar, Aijun An
ICDE2
2012 Efficient Bi-objective Team Formation in Social Networks
Mehdi Kargar, Aijun An, Morteza Zihayat
ECML/PKDD (2)2
2012 Unsupervised Emotion Detection from Text Using Semantic and Syntactic Relations
abstract
Emotion detection from text is a relatively new classification task. This paper proposes a novel unsupervised context-based approach to detecting emotion from text at the sentence level. The proposed methodology does not depend on any existing manually crafted affect lexicons such as Word Net-Affect, thereby rendering our model flexible enough to classify sentences beyond Ekman's model of six basic emotions. Our method computes an emotion vector for each potential affect bearing word based on the semantic relatedness between words and various emotion concepts. The scores are then fine tuned using the syntactic dependencies within the sentence structure. Extensive evaluation on various data sets shows that our framework is a more generic and practical solution to the emotion classification problem and yields significantly more accurate results than recent unsupervised approaches.
Ameeta Agrawal, Aijun An
Web Intelligence2
2012 Feature evaluation for web crawler detection with data mining techniques
Dusan Stevanovic, Aijun An, Natalija Vlajic
Expert Syst. Appl.2
2012 Mining Online Reviews for Predicting Sales Performance: A Case Study in the Movie Domain
abstract
Posting reviews online has become an increasingly popular way for people to express opinions and sentiments toward the products bought or services received. Analyzing the large volume of online reviews available would produce useful actionable knowledge that could be of economic values to vendors and other interested parties. In this paper, we conduct a case study in the movie domain, and tackle the problem of mining reviews for predicting product sales performance. Our analysis shows that both the sentiments expressed in the reviews and the quality of the reviews have a significant impact on the future sales performance of products in question. For the sentiment factor, we propose Sentiment PLSA (S-PLSA), in which a review is considered as a document generated by a number of hidden sentiment factors, in order to capture the complex nature of sentiments. Training an S-PLSA model enables us to obtain a succinct summary of the sentiment information embedded in the reviews. Based on S-PLSFA, we propose ARSA, an Autoregressive Sentiment-Aware model for sales prediction. We then seek to further improve the accuracy of prediction by considering the quality factor, with a focus on predicting the quality of a review in the absence of user-supplied indicators, and present ARSQA, an Autoregressive Sentiment and Quality Aware model, to utilize sentiments and quality for predicting product sales performance. Extensive experiments conducted on a large movie data set confirm the effectiveness of the proposed approach.
Xiaohui Yu 0001, Yang Liu 0008, Jimmy Huang 0001, Aijun An
IEEE Trans. Knowl. Data Eng.4
2011 Discovering top-k teams of experts with/without a leader in social networks
abstract
We study the problem of discovering a team of experts from a social network. Given a project whose completion requires a set of skills, our goal is to find a set of experts that together have all of the required skills and also have the minimal communication cost among them. We propose two communication cost functions designed for two types of communication structures. We show that the problem of finding the team of experts that minimizes one of the proposed cost functions is NP-hard. Thus, an approximation algorithm with an approximation ratio of two is designed. We introduce the problem of finding a team of experts with a leader. The leader is responsible for monitoring and coordinating the project, and thus a different communication cost function is used in this problem. To solve this problem, an exact polynomial algorithm is proposed. We show that the total number of teams may be exponential with respect to the number of required skills. Thus, two procedures that produce top-k teams of experts with or without a leader in polynomial delay are proposed. Extensive experiments on real datasets demonstrate the effectiveness and scalability of the proposed methods.
Mehdi Kargar, Aijun An
CIKM2
2011 Towards Automatic Acquisition of a Fully Sense Tagged Corpus for Persian
Bahareh Sarrafzadeh, Nikolay Yakovets, Nick Cercone, Aijun An
ISMIS4
2011 Detecting Web Crawlers from Web Server Access Logs with Data Mining Classifiers
Dusan Stevanovic, Aijun An, Natalija Vlajic
ISMIS2
2011 Combining integrated sampling with SVM ensembles for learning from imbalanced datasets
Yang Liu 0008, Xiaohui Yu 0001, Jimmy Huang 0001, Aijun An
Inf. Process. Manag.4
2011 Finding best evidence for evidence-based best practice recommendations in health care: the initial decision support system design
Nick Cercone, Xiangdong An 0001, Jiye Li, Zhenmei Gu, Aijun An
Knowl. Inf. Syst.5
2011 Keyword Search in Graphs: Finding r-cliques
abstract
Keyword search over a graph finds a substructure of the graph containing all or some of the input keywords. Most of previous methods in this area find connected minimal trees that cover all the query keywords. Recently, it has been shown that finding subgraphs rather than trees can be more useful and informative for the users. However, the current tree or graph based methods may produce answers in which some content nodes (i.e., nodes that contain input keywords) are not very close to each other. In addition, when searching for answers, these methods may explore the whole graph rather than only the content nodes. This may lead to poor performance in execution time. To address the above problems, we propose the problem of finding r -cliques in graphs. An r -clique is a group of content nodes that cover all the input keywords and the distance between each two nodes is less than or equal to r . An exact algorithm is proposed that finds all r -cliques in the input graph. In addition, an approximation algorithm that produces r -cliques with 2-approximation in polynomial delay is proposed. Extensive performance studies using two large real data sets confirm the efficiency and accuracy of finding r -cliques in graphs.
Mehdi Kargar, Aijun An
Proc. VLDB Endow.2
2010 Partial drift detection using a rule induction framework
abstract
The major challenge in mining data streams is the issue of concept drift, the tendency of the underlying data generation process to change over time. In this paper, we propose a general rule learning framework that can efficiently handle concept-drifting data streams and maintain a highly accurate classification model. The main idea is to focus on partial drifts by allowing individual rules to monitor the stream and detect if there is a drift in the regions they cover. A rule quality measure then decides whether the affected rules are inconsistent with the concept drift. The model is accordingly updated to only include rules that are consistent with the newly arrived concept. A dynamically maintained set of instances deemed relevant to the most recent concept is also kept at memory. Learning a new concept from a larger set of instances reduces the variance of data distribution and allows for a more accurate, stable classification model. Our experiments show that this approach not only handles the drift efficiently, but it also can provide higher classification accuracy compared to other competitive approaches on a variety of real and synthetic data sets.
Damon Sotoudeh, Aijun An
CIKM2
2010 Medical search and classification tools for recommendation
abstract
their patients' records from paper to computer, enormous amounts of electronic medical records (EMR) have become available for medical research. Some of the EMR data are well-structured, for which traditional database management systems can provide effective retrieval and management functions. However, most of the EMR data (such as progress notes and consultation letters) are in free text formats. How to effectively and efficiently retrieve and discover useful information from the vast amount of such semi-structured data is a challenge faced by medical professionals. Without proper tools, the rich information and knowledge buried in the medical health records are unavailable for clinical research and decision-making. The objective of our research is to develop text analytics tools that are capable of parsing clinical medical data so that predefined search subjects that correspond to a list of medical diagnoses can be extracted. In addition to this particular core functionality, it is also desired that several important assets should be present within the text-analytics tools in order to improve its overall ability to be used as recommendation tools. In this research, we work with research scientists at the Institute for Clinical Evaluative Sciences (ICES) in Toronto and examine a number of techniques for structuring and processing free text documents in order to effectively and efficiently search and analyze vast amount of medical records. We implement several powerful medical text analytics tools for clinical data searching and classification. For data classification, our tools sort through a great amount of patientrecords to identify the likelihood of a patient having myocardial infarction (MI) or hypertension (HTN), and classify the patients accordingly. Our tools can also identify the likelihood of a patient being a smoker, previous smoker or non-smoker based on the text data of medical records.
Jimmy Huang 0001, Aijun An, Qinmin Hu
SIGIR2
2010 S-PLASA+: adaptive sentiment analysis with application to sales performance prediction
abstract
Analyzing the large volume of online reviews would produce useful knowledge that could be of economic values to vendors and other interested parties. In particular, the sentiments expressed in the online reviews have been shown to be strongly correlated with the sales performance of products. In this paper, we present an adaptive sentiment analysis model called S-PLSA+, which aims to capture the hidden sentiment factors in the reviews with the capability to be incrementally updated as more data become available. We show how S-PLSA+ can be applied to sales performance prediction using an ARSA model developed in previous literature. A case study is conducted in the movie domain, and results from preliminary experiments confirm the effectiveness of the proposed model.
Yang Liu 0008, Xiaohui Yu 0001, Jimmy Huang 0001, Aijun An
SIGIR4
2010 An Adaptive Model for Probabilistic Sentiment Analysis
abstract
Online reviews, which are getting increasingly prevalent with the rapid growth of Web 2.0, have been shown to be second only to "word-of-mouth" in terms of influencing purchase decisions. It is therefore imperative to analyze them and distill useful knowledge that could be of economic values to vendors and other interested parties. Previous studies have confirmed that the sentiments expressed in the online reviews are strongly correlated with the sales performance of products. In particular, a model called ARSA has been proposed for predicting sales performance using a model called S-PLSA. In this paper, we build upon that work, and present an adaptive sentiment analysis model called S-PLSA+, which not only can capture the hidden sentiment factors in the reviews, but has the capability to be incrementally updated as more data become available. We show how the proposed S-PLSA+model can be applied to sales performance prediction using the ARSA model. A case study is conducted in the movie domain, and results from preliminary experiments confirm the effectiveness of the proposed model.
Xiaohui Yu 0001, Yang Liu 0008, Aijun An
Web Intelligence3
2010 A quality-aware model for sales prediction using reviews
abstract
Writing and publishing reviews online has become an increasingly popular way for people to express opinions and sentiments. Analyzing the large volume of online reviews available can produce useful knowledge that are of interest to vendors and other parties. Prior studies in the literature have shown that online reviews have a significant correlation with the sales of products, and therefore mining the reviews could help predict the sales performance of relevant products. However, those studies fail to consider one important factor that may significantly affect the accuracy of the prediction, i.e., the quality of the reviews. In this paper, we propose a regression model that explicitly takes into account the quality factor, and discusses how this quality information can be predicted when it is not readily available. Experimental results on a movie review dataset confirm the effectiveness of the proposed model.
Xiaohui Yu 0001, Yang Liu 0008, Jimmy Huang 0001, Aijun An
WWW4
2010 Introduction to the special issue on advanced information retrieval and databases
Aijun An, Dominik Slezak
J. Intell. Inf. Syst.1
2009 Discovering Temporal Associations among Significant Changes in Gene Expression
abstract
One of the most demanding problems in mining temporal data is to identify how multivariate change associations might be discovered and used to better understand data interactions and dependencies. This paper introduces a framework to mine associations among significant changes in multivariate time-series data. Building on statistical methods, we detect significant changes in time-series data and use marginal change rates to qualify the direction of change at significant change points. Furthermore, a propositional confirmation-guided rule discovery method is used to discover associations among these significant changes. We apply our approach to gene expression data measured in yeast cell cycles and demonstrate that our method can learn novel and high-quality significant change associations among different genes. Such associations can be used to cluster genes and build gene interaction networks.
Hashmat Rohian, Aijun An, Jiashu Zhao, Jimmy Huang 0001
BIBM2
2009 Diverging patterns: discovering significant frequency change dissimilarities in large databases
abstract
In this paper, we present a framework for mining diverging patterns, a new type of contrast patterns whose frequency changes significantly differently in two data sets, e.g., it changes from a relatively low to a relatively high value in one dataset, but from high to low in the other. In this framework, a measure called diverging ratio is defined and used to discover diverging patterns. We use a four-dimensional vector to represent a pattern, and define the pattern's diverging ratio based on the angular difference between its vectors in two datasets. An algorithm is proposed to mine diverging patterns from a pair of datasets, which makes use of a standard frequent pattern mining algorithm to compute vector components efficiently. We demonstrate the effectiveness of our approach on real-world datasets, showing that the method can reveal novel knowledge from large databases.
Aijun An, Qian Wan 0001, Jiashu Zhao, Jimmy Huang 0001
CIKM1
2009 A roadmap of clustering algorithms: finding a match for a biomedical application
abstract
Clustering is ubiquitously applied in bioinformatics with hierarchical clustering and k-means partitioning being the most popular methods. Numerous improvements of these two clustering methods have been introduced, as well as completely different approaches such as grid-based, density-based and model-based clustering. For improved bioinformatics analysis of data, it is important to match clusterings to the requirements of a biomedical application. In this article, we present a set of desirable clustering features that are used as evaluation criteria for clustering algorithms. We review 40 different clustering algorithms of all approaches and datatypes. We compare algorithms on the basis of desirable clustering features, and outline algorithms' benefits and drawbacks as a basis for matching them to biomedical applications.
Bill Andreopoulos, Aijun An, Xiaogang Wang 0007, Michael Schroeder 0001
Briefings Bioinform.2
2009 Efficient layered density-based clustering of categorical data
Bill Andreopoulos, Aijun An, Xiaogang Wang 0007, Dirk Labudde
J. Biomed. Informatics2
2009 Discovering Transitional Patterns and Their Significant Milestones in Transaction Databases
abstract
A transaction database usually consists of a set of time-stamped transactions. Mining frequent patterns in transaction databases has been studied extensively in data mining research. However, most of the existing frequent pattern mining algorithms (such as Apriori and FP-growth) do not consider the time stamps associated with the transactions. In this paper, we extend the existing frequent pattern mining framework to take into account the time stamp of each transaction and discover patterns whose frequency dramatically changes over time. We define a new type of patterns, called transitional patterns, to capture the dynamic behavior of frequent patterns in a transaction database. Transitional patterns include both positive and negative transitional patterns. Their frequencies increase/decrease dramatically at some time points of a transaction database. We introduce the concept of significant milestones for a transitional pattern, which are time points at which the frequency of the pattern changes most significantly. Moreover, we develop an algorithm to mine from a transaction database the set of transitional patterns along with their significant milestones. Our experimental studies on real-world databases illustrate that mining positive and negative transitional patterns is highly promising as a practical and useful approach for discovering novel and interesting knowledge from large databases.
Qian Wan 0001, Aijun An
IEEE Trans. Knowl. Data Eng.2
2008 Modeling and Predicting the Helpfulness of Online Reviews
abstract
Online reviews provide a valuable resource for potential customers to make purchase decisions. However, the sheer volume of available reviews as well as the large variations in the review quality present a big impediment to the effective use of the reviews, as the most helpful reviews may be buried in the large amount of low quality reviews. The goal of this paper is to develop models and algorithms for predicting the helpfulness of reviews, which provides the basis for discovering the most helpful reviews for given products. We first show that the helpfulness of a review depends on three important factors: the reviewerpsilas expertise, the writing style of the review, and the timeliness of the review. Based on the analysis of those factors, we present a nonlinear regression model for helpfulness prediction. Our empirical study on the IMDB movie reviews dataset demonstrates that the proposed approach is highly effective.
Yang Liu 0008, Jimmy Huang 0001, Aijun An, Xiaohui Yu 0001
ICDM3
2008 HelpMeter: A Nonlinear Model for Predicting the Helpfulness of Online Reviews
abstract
With the flourish of the Internet, online review mining has attracted a lot of attention from the research community. However, compared to various well-studied sentiment analysis and opinion summarization problems, less effort has been made to analyze the quality of online reviews. The objective of this paper is to fill in this gap by automatically evaluating the "helpfulness" of reviews and consequently developing novel models to identify the most helpful reviews for a particular product. In particular, based on a thorough analysis of various factors that may affect the review quality, we propose HelpMeter, a nonlinear regression model for helpfulness prediction. Some preliminary experiments were conducted on a movie review data set, and the performance results confirm the superiority of the proposed method.
Yang Liu 0008, Jimmy Huang 0001, Aijun An, Xiaohui Yu 0001
Web Intelligence3
2007 Transitional Patterns and Their Significant Milestones
abstract
Mining frequent patterns in transaction databases has been studied extensively in data mining research. However, most of the existing frequent pattern mining algorithms do not consider the time stamps associated with the transactions. In this paper, we extend the existing frequent pattern mining framework to take into account the time stamp of each transaction and discover patterns whose frequency dramatically changes over time. We define a new type of patterns, called transitional patterns, to capture the dynamic behavior of frequent patterns in a transaction database. Transitional patterns include both positive and negative transitional patterns. Their frequencies increase/decrease dramatically at some time points of a transaction database. We introduce the concept of significant milestones for a transitional pattern, which are time points at which the frequency of the pattern changes most significantly. Moreover, we develop an algorithm to mine from a transaction database the set of transitional patterns along with their significant milestones. Our experimental studies on real-world databases illustrate that mining positive and negative transitional patterns is highly promising as a practical and useful approach to discovering novel and interesting knowledge from large databases.
Qian Wan 0001, Aijun An
ICDM2
2007 Hierarchical Density-Based Clustering of Categorical Data and a Simplification
Bill Andreopoulos, Aijun An, Xiaogang Wang 0007
PAKDD2
2007 Fast Counting with AV-Space for Efficient Rule Induction
abstract
We present AV-space, a new data structure for caching data set statistics for efficiently learning classification rules from large data sets. The AV-space is designed to work with sequential-covering rule induction algorithms. It is used to accelerate queries about the count of the examples in a data set that satisfy a conjunction of attribute-value pairs. With an AV-space, the learning algorithm does not have to access the training data to obtain the statistics about the data. We present the structure of an AV-space, algorithms for building and querying an AV-space, and procedures for dynamically updating the AV-space during the rule induction process. We present an experimental evaluation that compares the AV-space with a commonly-used data structure that simply loads the (encoded) training examples into memory. We show that the use of AV-space significantly improves the speed of rule induction and that it consumes less memory on large data sets.
Linyan Wang, Aijun An
SDM2
2007 ARSA: a sentiment-aware model for predicting sales performance using blogs
abstract
Due to its high popularity, Weblogs (or blogs in short) present a wealth of information that can be very helpful in assessing the general public's sentiments and opinions. In this paper, we study the problem of mining sentiment information from blogs and investigate ways to use such information for predicting product sales performance. Based on an analysis of the complex nature of sentiments, we propose Sentiment PLSA (S-PLSA), in which a blog entry is viewed as a document generated by a number of hidden sentiment factors. Training an S-PLSA model on the blog data enables us to obtain a succinct summary of the sentiment information embedded in the blogs. We then present ARSA, an autoregressive sentiment-aware model, to utilize the sentiment information captured by S-PLSA for predicting product sales performance. Extensive experiments were conducted on a movie data set. We compare ARSA with alternative models that do not take into account the sentiment information, as well as a model with a different feature selection method. Experiments confirm the effectiveness and superiority of the proposed approach.
Yang Liu 0008, Jimmy Huang 0001, Aijun An, Xiaohui Yu 0001
SIGIR3
2007 Clustering by common friends finds locally significant proteins mediating modules
abstract
MOTIVATION: Much research has been dedicated to large-scale protein interaction networks including the analysis of scale-free topologies, network modules and the relation of domain-domain to protein-protein interaction networks. Identifying locally significant proteins that mediate the function of modules is still an open problem. METHOD: We use a layered clustering algorithm for interaction networks, which groups proteins by the similarity of their direct neighborhoods. We identify locally significant proteins, called mediators, which link different clusters. We apply the algorithm to a yeast network. RESULTS: Clusters and mediators are organized in hierarchies, where clusters are mediated by and act as mediators for other clusters. We compare the clusters and mediators to known yeast complexes and find agreement with precision of 71% and recall of 61%. We analyzed the functions, processes and locations of mediators and clusters. We found that 55% of mediators to a cluster are enriched with a set of diverse processes and locations, often related to translocation of biomolecules. Additionally, 82% of clusters are enriched with one or more functions. The important role of mediators is further corroborated by a comparatively higher degree of conservation across genomes. We illustrate the above findings with an example of membrane protein translocation from the cytoplasm to the inner nuclear membrane. AVAILABILITY: All software is freely available under Supplementary information.
Bill Andreopoulos, Aijun An, Xiaogang Wang 0007, Michalis Faloutsos, Michael Schroeder 0001
Bioinform.2
2007 Clustering large software systems at multiple layers
Bill Andreopoulos, Aijun An, Vassilios Tzerpos, Xiaogang Wang 0007
Inf. Softw. Technol.2
2007 Personalized recommendation with adaptive mixture of markov models
abstract
Abstract With more and more information available on the Internet, the task of making personalized recommendations to assist the user's navigation has become increasingly important. Considering there might be millions of users with different backgrounds accessing a Web site everyday, it is infeasible to build a separate recommendation system for each user. To address this problem, clustering techniques can first be employed to discover user groups. Then, user navigation patterns for each group can be discovered, to allow the adaptation of a Web site to the interest of each individual group. In this paper, we propose to model user access sequences as stochastic processes, and a mixture of Markov models based approach is taken to cluster users and to capture the sequential relationships inherent in user access histories. Several important issues that arise in constructing the Markov models are also addressed. The first issue lies in the complexity of the mixture of Markov models. To improve the efficiency of building/maintaining the mixture of Markov models, we develop a lightweight adapt‐ive algorithm to update the model parameters without recomputing model parameters from scratch. The second issue concerns the proper selection of training data for building the mixture of Markov models. We investigate two different training data selection strategies and perform extensive experiments to compare their effectiveness on a real dataset that is generated by a Web‐based knowledge management system, Livelink.
Yang Liu 0008, Jimmy Huang 0001, Aijun An
J. Assoc. Inf. Sci. Technol.3
2006 Applying Data Mining to Pseudo-Relevance Feedback for High Performance Text Retrieval
abstract
In this paper, we investigate the use of data mining, in particular the text classification and co-training techniques, to identify more relevant passages based on a small set of labeled passages obtained from the blind feedback of a retrieval system. The data mining results are used to expand query terms and to re-estimate some of the parameters used in a probabilistic weighting function. We evaluate the data mining based feedback method on the TREC HARD data set. The results show that data mining can be successfully applied to improve the text retrieval performance. We report our experimental findings in detail.
Jimmy Huang 0001, Yan Rui Huang, Miao Wen, Aijun An, Yang Liu 0008, Josiah Poon
ICDM4
2006 Mining and Modeling Database User Access Patterns
Qingsong Yao, Aijun An, Jimmy Huang 0001
ISMIS2
2006 Boosting Prediction Accuracy on Imbalanced Datasets with SVM Ensembles
Yang Liu 0008, Aijun An, Jimmy Huang 0001
PAKDD2
2006 A platform for Okapi-based contextual information retrieval
abstract
We present an extensible java-based platform for contextual retrieval based on the probabilistic information retrieval model. Modules for dual indexes, relevance feedback with blind or machine learning approaches and query expansion with context are integrated into the Okapi system to deal with the contextual information. This platform allows easy extension to include other types of contextual information.
Jimmy Huang 0001, Miao Wen, Aijun An, Yan Rui Huang
SIGIR3
2006 An efficient approach to mining indirect associations
Qian Wan 0001, Aijun An
J. Intell. Inf. Syst.2
2006 Applying language modeling to session identification from database trace logs
Jimmy Huang 0001, Qingsong Yao, Aijun An
Knowl. Inf. Syst.3
2005 Finding and Analyzing Database User Sessions
Qingsong Yao, Aijun An, Jimmy Huang 0001
DASFAA2
2005 A Machine Learning Approach to Identifying Database Sessions Using Unlabeled Data
Qingsong Yao, Jimmy Huang 0001, Aijun An
DaWaK3
2005 A Distance-Based Algorithm for Clustering Database User Sessions
Qingsong Yao, Aijun An, Jimmy Huang 0001
ISMIS2
2004 Characterizing Database User's Access Patterns
Qingsong Yao, Aijun An
DEXA2
2004 Clustering Web Surfers with Probabilistic Models in a Real Application
abstract
The problem of modeling and predicting a Web surfer's browsing patterns has gained increasing attention in recent years. In this paper we present our experience in clustering Web surfers using a mixture of Markov models with a real application of Livelink log data. We propose different techniques to improve the clustering performance, and evaluate the techniques through experiments.
Yang Liu 0008, Jimmy Huang 0001, Aijun An, Gary Promhouse
Web Intelligence3
2004 Dynamic Web log session identification with statistical language models
abstract
Abstract We present a novel session identification method based on statistical language modeling. Unlike standard timeout methods, which use fixed time thresholds for session identification, we use an information theoretic approach that yields more robust results for identifying session boundaries. We evaluate our new approach by learning interesting association rules from the segmented session files. We then compare the performance of our approach to three standard session identification methods—the standard timeout method, the reference length method, and the maximal forward reference method—and find that our statistical language modeling approach generally yields superior results. However, as with every method, the performance of our technique varies with changing parameter settings. Therefore, we also analyze the influence of the two key factors in our language‐modeling–based approach: the choice of smoothing technique and the language model order. We find that all standard smoothing techniques, save one, perform well, and that performance is robust to language model order.
Jimmy Huang 0001, Fuchun Peng, Aijun An, Dale Schuurmans
J. Assoc. Inf. Sci. Technol.3
2003 Using User Access Patterns for Semantic Query Caching
Qingsong Yao, Aijun An
DEXA2
2003 Objective and Subjective Algorithms for Grouping Association Rules
abstract
We propose two algorithms for grouping and summarizing association rules. The first algorithm recursively groups rules according to the structure of the rules and generates a tree of clusters as a result. The second algorithm groups the rules according to the semantic distance between the rules by making use of an automatically tagged semantic tree-structured network of items. We provide a case study in which the proposed algorithms are evaluated. The results show that our grouping methods are effective and produce good grouping results.
Aijun An, Shakil M. Khan 0001, Jimmy Huang 0001
ICDM1
2003 SQL-Relay: An Event-Driven Rule-Based Database Gateway
Qingsong Yao, Aijun An
WAIM2
2002 Comparison of interestingness functions for learning web usage patterns
abstract
Livelink is a collaborative intranet, extranet and e-business application that enables employees and business partners of an organization to capture, share and reuse business information and knowledge. The usage of the Livelink software has been recorded by the Livelink Web server in its log files. We present an application of data mining techniques to the Livelink Web usage data. In particular, we focus on how to find interesting association rules and sequential patterns from the Livelink log files. A number of interestingness measures are used in our application to identify interesting rules and patterns. We present a comparison of these measures based on the feedback from domain experts. Some of the interestingness measures are found to be better than others.
Jimmy Huang 0001, Nick Cercone, Aijun An
CIKM3
2002 Discovery of Interesting Association Rules from Livelink Web Log Data
abstract
We present our experience in mining web usage patterns from a large collection of Livelink log data. Livelink is a web-based product of Open Text, which provides automatic management and retrieval of different types of information objects over an intranet or extranet. We report our experience in preprocessing raw log data and post-processing the mining results for finding interesting rules. In particular we compare and evaluate a number of rule interestingness measures and find that two of the measures that have not been used in association rule learning work very well.
Jimmy Huang 0001, Aijun An, Nick Cercone, Gary Promhouse
ICDM2
2001 Comparisons of Classification Methods for Screening Potential Compounds
abstract
We compare a number of data mining and statistical methods on the drug design problem of modeling molecular structure-activity relationships. The relationships can be used to identify active compounds based on their chemical structures from a large inventory of chemical compounds. The data set of this application has a highly skewed class distribution, in which only 2% of the compounds are considered active. We apply a number of classification methods to this extremely imbalanced data set and propose to use different performance measures to evaluate these methods. We report our findings on the characteristics of the performance measures, the effect of using pruning techniques in this application and a comparison of local learning methods with global techniques. We also investigate whether reducing the imbalance in the training data by up-sampling or down-sampling would improve the predictive performance.
Aijun An
ICDM1
2001 Rule Quality Measures for Rule Induction Systems: Description and Evaluation
abstract
A rule quality measure is important to a rule induction system for determining when to stop generalization or specialization. Such measures are also important to a rule‐based classification procedure for resolving conflicts among rules. We describe a number of statistical and empirical rule quality formulas and present an experimental comparison of these formulas on a number of standard machine learning datasets. We also present a meta‐learning method for generating a set of formula‐behavior rules from the experimental results which show the relationships between a formula's performance and the characteristics of a dataset. These formula‐behavior rules are combined into formula‐selection rules that can be used in a rule induction system to select a rule quality formula before rule induction. We will report the experimental results showing the effects of formula‐selection on the predictive performance of a rule induction system.
Aijun An, Nick Cercone
Comput. Intell.1
2000 Rule Quality Measures Improve the Accuracy of Rule Induction: An Experimental Approach
Aijun An, Nick Cercone
ISMIS1
2000 Probability-Based Chinese Text Processing and Retrieval
abstract
We discuss the use of probability‐based natural language processing for Chinese text retrieval. We focus on comparing different text extraction methods and probabilistic weighting methods. Several document processing methods and probabilistic weighting functions are presented. A number of experiments have been conducted on large standard text collections. We present the experimental results that compare a word‐based text processing method with a character‐based method. The experimental results also compare a number of term‐weighting functions including both single‐unit weighting and compound‐unit weighting functions.
Jimmy Huang 0001, Stephen E. Robertson, Nick Cercone, Aijun An
Comput. Intell.4
1999 ORTES: The Design of a Real-Time Control Expert System
Aijun An, Nick Cercone, Christine W. Chan
ISMIS1
1999 Discretization of Continuous Attributes for Learning Classification Rules
Aijun An, Nick Cercone
PAKDD1
1999 Rule-Induction and Case-Based Reasoning: Hybrid Architectures Appear Advantageous
abstract
Researchers have embraced a variety of machine learning (ML) techniques in their efforts to improve the quality of learning programs. The recent evolution of hybrid architectures for machine learning systems has resulted in several approaches that combine rule induction methods with case-based reasoning techniques to engender performance improvements over more traditional single-representation architectures. We briefly survey several major rule-induction and case-based reasoning ML systems. We then examine some interesting hybrid combinations of these systems and explain their strengths and weaknesses as learning systems. We present a balanced approach to constructing a hybrid architecture, along with arguments in favor of this balance and mechanisms for achieving a proper balance. Finally, we present some initial empirical results from testing our ideas and draw some conclusions based on those results.
Nick Cercone, Aijun An, Christine W. Chan
IEEE Trans. Knowl. Data Eng.2
1997 Integrating Rule Induction and Case-Based Reasoning to Enhance Problem Solving
Aijun An, Nick Cercone, Christine W. Chan
ICCBR1