Shengrui Wang

dblp:w/ShengruiWang · DBLP profile ↗
← Back
47ranked-venue papers in the field
1as first author
10since 2021 · last 2024
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 33Database Systems & Data Management · 9 (1 first)Information Retrieval & Web Search · 4Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2024 FR3LS: A Forecasting Model with Robust and Reduced Redundancy Latent Series
Abdallah Aaraba, Shengrui Wang, Jean-Marc Patenaude
PAKDD (6)2
2024 Kernel Representation Learning with Dynamic Regime Discovery for Time Series Forecasting
Kunpeng Xu 0002, Lifei Chen, Jean-Marc Patenaude, Shengrui Wang
PAKDD (6)4
2024 RHINE: A Regime-Switching Model with Nonlinear Representation for Discovering and Forecasting Regimes in Financial Markets
abstract
We investigate the problem of discovering and forecasting regular regime switches in a financial ecosystem comprising multiple time series. Such regime switches, indicative of varying market behaviors across distinct time intervals, are pivotal for a nuanced understanding of market dynamics, which in turn allows informed model selection for forecasting and enhanced interpretability of predictive outcomes. Despite strides in this domain, prevailing methodologies often falter due to: (1) an inability to effectively model the temporal behaviors inherent in financial series; and (2) neglecting the interdependencies among series when discovering regimes. In this paper, we propose RHINE, a Regime-switcHIng model with Nonlinear rEpresentation. RHINE stands out with its kernel-based representation, adept at capturing the dynamic shifts in market regimes. This representation encapsulates the nonlinear interplay across multiple financial time series. By leveraging the kernel representation, we introduce an eigengap thresholding measure, designed to automatically discern the optimal number of financial market regimes, enhancing the model's adaptability to market fluctuations. Empirical assessments on both synthetic and real-world stock market datasets underscore RHINE's prowess. The findings illuminate that the inherent structures governing financial market behaviors are dynamic, and harnessing these dynamics via RHINE leads to a regime-based model that outperforms both conventional and state-of-the-art neural network models in predictive capabilities.
Kunpeng Xu 0002, Lifei Chen, Jean-Marc Patenaude, Shengrui Wang
SDM4
2023 Modeling of Repeated Measures for Time-to-event Prediction
Jianfei Zhang 0002, Lifei Chen, Shengrui Wang
ADMA (1)3
2023 Toward Healthy Aging: Temporal Regression for Disability Prediction and Warning Decision-Making
Jianfei Zhang 0002, Lifei Chen, Shengrui Wang
DEXA (2)3
2023 Rethinking Temporal Dependencies in Multiple Time Series: A Use Case in Financial Data
abstract
These days, complex systems yield copious time series data, necessitating understanding co-generation, often assessed through pairwise comparisons. However, this method lacks scalability and temporal dynamics handling. In this paper, we advocate using a temporal graph to capture contiguous effects among multiple time series efficiently. Our two-step approach identifies patterns and temporal influences with low execution time, showcasing its potential in financial system incident prediction.
Patrick Owusu, Etienne Gael Tajeuna, Jean-Marc Patenaude, Armelle Brun, Shengrui Wang
ICDM5
2023 Modeling Regime Shifts in Multiple Time Series
abstract
We investigate the problem of discovering and modeling regime shifts in an ecosystem comprising multiple time series known as co-evolving time series. Regime shifts refer to the changing behaviors exhibited by series at different time intervals. Learning these changing behaviors is a key step toward time series forecasting. While advances have been made, existing methods suffer from one or more of the following shortcomings: (1) failure to take relationships between time series into consideration for discovering regimes in multiple time series; (2) lack of an effective approach that models time-dependent behaviors exhibited by series; (3) difficulties in handling data discontinuities which may be informative. Most of the existing methods are unable to handle all of these three issues in a unified framework. This, therefore, motivates our effort to devise a principled approach for modeling interactions and time-dependency in co-evolving time series. Specifically, we model an ecosystem of multiple time series by summarizing the heavy ensemble of time series into a lighter and more meaningful structure called a mapping grid . By using the mapping grid, our model first learns time series behavioral dependencies through a dynamic network representation, then learns the regime transition mechanism via a full time-dependent Cox regression model. The originality of our approach lies in modeling interactions between time series in regime identification and in modeling time-dependent regime transition probabilities, usually assumed to be static in existing work.
Etienne Gael Tajeuna, Mohamed Bouguessa, Shengrui Wang
ACM Trans. Knowl. Discov. Data3
2022 Discovering Affinity Relationships between Personality Types
abstract
Psychology research findings suggest that personality is related to differences in friendship characteristics and that some personality traits correlate with linguistic behavior. In this paper, we investigate the influence that personality may have on affinity formation. To this end, we derive affinity relationships from social media interactions, examine personality based on language use to discover the emotional stability of affinity relationships, and measure semantic similarity at the personality type level to understand the logic behind the development of affinity. Specifically, we conduct extensive experiments using a publicly available dataset containing information on individuals who self-identified with a Myers-Briggs personality type. Our results identify certain influential personality types that weigh more heavily on affinity relationships and show that personality can be predicted from spontaneous language with an F-1 score superior to 0.76. Future research avenues are proposed.
Jean Marie Tshimula, Belkacem Chikhaoui, Shengrui Wang
ASONAM3
2022 Clustering-Based Cross-Sectional Regime Identification for Financial Market Forecasting
Rongbo Chen, Kunpeng Xu 0002, Jean-Marc Patenaude, Shengrui Wang
DEXA (2)5
2021 Mining Customers' Changeable Electricity Consumption for Effective Load Forecasting
abstract
Most existing approaches for electricity load forecasting perform the task based on overall electricity consumption. However, using such a global methodology can affect load forecasting accuracy, as it does not consider the possibility that customers’ consumption behavior may change at any time. Predicting customers’ electricity consumption in the presence of unstable behaviors poses challenges to existing models. In this article, we propose a principled approach capable of handling customers’ changeable electricity consumption. We devise a network-based method that first builds and tracks clusters of customer consumption patterns over time. Then, on the evolving clusters, we develop a framework that exploits long short-term memory recurrent neural network and survival analysis techniques to forecast electricity consumption. Our experiments on real electricity consumption datasets illustrate the suitability of the proposed approach.
Etienne Gael Tajeuna, Mohamed Bouguessa, Shengrui Wang
ACM Trans. Intell. Syst. Technol.3
2020 On Predicting Behavioral Deterioration in Online Discussion Forums
abstract
Early detection of behavioral deterioration can be of great importance in preventing individuals' misbehavior from escalating in severity. This paper addresses the problem of behavioral deterioration in the context of online discussion forums. We propose a novel method that builds behavioral sequences from temporal information to gain a better understanding of behaviors exhibited by forum members, and then explores n-gram features to predict behavioral deterioration from consecutive combinations of sequential patterns corresponding to misbehavior. We conduct extensive experiments using real-world datasets and demonstrate the ability of our method to predict behavioral deterioration with a high degree of accuracy, as evaluated by F-1 scores. Our quantitative analysis of the model's performance yields F-1 scores of over 0.7. Specifically, we find that the best-performing model is linear SVM, with an average F-1 score of 0.74. Some future research avenues are proposed.
Jean Marie Tshimula, Belkacem Chikhaoui, Shengrui Wang
ASONAM3
2020 A Pre-training Approach for Stance Classification in Online Forums
abstract
Stance detection is the task of automatically determining whether the author of a piece of text is in favor of, against, or neutral towards a target such as a topic, entity, or claim. In this paper, we propose a method based on RoBERTa to classify stances by capturing the context of the discussion through the examination of pairs of stances and relational structures of debates specific to each topic within the defined window of each forum participant's interventions. Furthermore, we examine the degree of disagreement and neutrality in various debate topics to measure divergence of opinion in the course of the debate and estimate the emotional state manifested in different debate topics. We conduct extensive experiments using two publicly available datasets and demonstrate that our method considers more stance classes, provides better results and yields statistical improvements over existing techniques. Our quantitative analysis of model performance yields F-1 scores of over 0.745. Interestingly, we obtained the highest F-1 score, 0.814, on a stance class which was not taken into consideration in prior work. We report that none of the metrics utilized to measure divergence of opinion yield values exceeding 50 % and the correlations between the same topics over 10-fold cross-validation are statistically significant for the majority of them (p <; 0.005). Several future research avenues are proposed.
Jean Marie Tshimula, Belkacem Chikhaoui, Shengrui Wang
ASONAM3
2020 Survival neural networks for time-to-event prediction in longitudinal study
Jianfei Zhang 0002, Lifei Chen, Yanfang Ye 0001, Gongde Guo, Rongbo Chen, Alain Vanasse, Shengrui Wang
Knowl. Inf. Syst.7
2019 HAR-search: a method to discover hidden affinity relationships in online communities
abstract
This paper addresses the problem of discovering hidden affinity relationships in online communities. Online discussions assemble people to talk about various types of topics and to share information. People progressively develop the affinity, and they get closer as frequently as they mention themselves in messages and they send positive messages to one another. We propose an algorithm, named HAR-search, for discovering hidden affinity relationships between individuals. Based on Markov Chain Models, we derive the affinity scores amongst individuals in an online community. We show that our method allows to track the evolution of the affinity over time and to predict affinity relationships arisen from the influence of certain community members. The comparison with the state-of-the-art method shows that our method results in robust discovery and considers minute details.
Jean Marie Tshimula, Belkacem Chikhaoui, Shengrui Wang
ASONAM3
2019 Time-Dependent Survival Neural Network for Remaining Useful Life Prediction
Jianfei Zhang 0002, Shengrui Wang, Lifei Chen, Gongde Guo, Rongbo Chen, Alain Vanasse
PAKDD (1)2
2019 Modeling and Predicting Community Structure Changes in Time-Evolving Social Networks
abstract
As time evolves, communities in a social network may undergo various changes known as critical events. For instance, a community can either split into several other communities, expand into a larger community, shrink to a smaller community, remain stable or merge into another community. Prediction of critical events has attracted increasing attention in the recent literature. Learning the evolution of communities over time is a key step towards predicting the critical events the communities may undergo. This is an important and difficult issue in the study of social networks. In the work to date, there is a lack of formal approaches for modeling and predicting critical events over time. This motivates our effort to design a new statistical method for event prediction in order to make better use of histories of past changes. To this end, this paper proposes a sliding window analysis from which we develop a model that simultaneously exploits an autoregressive model and survival analysis techniques. The autoregressive model is employed here to simulate the evolution of the community structure, whereas the survival analysis techniques allow the prediction of future changes the community may undergo.
Etienne Gael Tajeuna, Mohamed Bouguessa, Shengrui Wang
IEEE Trans. Knowl. Data Eng.3
2018 A Variable-Order Regime Switching Model to Identify Significant Patterns in Financial Markets
abstract
The identification and prediction of complex behaviors in time series are fundamental problems of interest in the field of financial data analysis. Autoregressive (AR) model and Regime switching (RS) models have been used successfully to study the behaviors of financial time series. However, conventional RS models evaluate regimes by using a fixed-order Markov chain and underlying patterns in the data are not considered in their design. In this paper, we propose a novel RS model to identify and predict regimes based on a weighted conditional probability distribution (WCPD) framework capable of discovering and exploiting the significant underlying patterns in time series. Experimental results on stock market data, with 200 stocks, suggest that the structures underlying the financial market behaviors exhibit different dynamics and can be leveraged to better define regimes with superior prediction capabilities than traditional models.
Philippe Chatigny, Rongbo Chen, Jean-Marc Patenaude, Shengrui Wang
ICDM4
2017 A Comparative Study of Different Approaches for Tracking Communities in Evolving Social Networks
abstract
In real-world social networks, there is an increasing interest in tracking the evolution of groups of users and detecting the various changes they are liable to undergo. Several approaches have been proposed for this. In studying these approaches, we observed that most of them use a two-stage process. In the first stage, they run an algorithm to identify groups of users at each timestamp. In the second stage, a pairwise comparison based on a similarity measure is employed to track groups of users and detect changes they may undergo. While the majority of existing approaches use a two-stage process, they all run different algorithms to identify communities and rely on different similarity measures to track groups of users over time. Noting that the different approaches may perform differently depending on the dynamic social network under investigation, we decided to make a high level survey of some existing tracking approaches and then do a comparative analysis of some of them. In our analysis, we compared the algorithms in two main situations: (1) when groups of users do not overlap and (2) when the groups are overlapping. The study was done on three different testbeds extracted from the DBLP, Autonomous System (AS) and Yelp datasets.
Ziwei He, Etienne Gael Tajeuna, Shengrui Wang, Mohamed Bouguessa
DSAA3
2017 Multiple Bayesian discriminant functions for high-dimensional massive data classification
Jianfei Zhang 0002, Shengrui Wang, Lifei Chen, Patrick Gallinari
Data Min. Knowl. Discov.2
2017 Detecting Communities of Authority and Analyzing Their Influence in Dynamic Social Networks
abstract
Users in real-world social networks are organized into communities that differ from each other in terms of influence, authority, interest, size, etc. This article addresses the problems of detecting communities of authority and of estimating the influence of such communities in dynamic social networks. These are new issues that have not yet been addressed in the literature, and they are important in applications such as marketing and recommender systems. To facilitate the identification of communities of authority, our approach first detects communities sharing common interests, which we call “meta-communities,” by incorporating topic modeling based on users’ community memberships. Then, communities of authority are extracted with respect to each meta-community, using a new measure based on the betweenness centrality. To assess the influence between communities over time, we propose a new model based on the Granger causality method. Through extensive experiments on a variety of social network datasets, we empirically demonstrate the suitability of our approach for community-of-authority detection and assessment of the influence between communities over time.
Belkacem Chikhaoui, Mauricio Chiazzaro, Shengrui Wang, Martin Sotir
ACM Trans. Intell. Syst. Technol.3
2016 Predicting COPD Failure by Modeling Hazard in Longitudinal Clinical Data
abstract
Chronic obstructive pulmonary disease (COPD) accounts for the highest rate of hospital readmissions and is the third leading cause of death in Canada, the United States and worldwide. Predicting COPD failure provides a prognostic warning of death or readmission, and is crucial to early intervention and decision-making. The aim of this study is to perform COPD failure prediction on longitudinal data. To address the inappropriate estimation of Cox hazard in current approaches, we propose a new representation of hazard to capture the relationship between survival probability and time-varying risk factors in a concise but effective way. To optimize model parameters, we design and maximize a new joint likelihood that comprises two components used to estimate survival status separately for failure and censored patients. A regularized optimization is performed on the joint likelihood to prevent overfitting arising from model learning. Our approach is applied to a real-life COPD data set and outperforms the current state-of-the-art prediction models in terms of the survival AUC, concordance index and Birer score metrics, this reveals that the great promise of our approach for clinical prediction.
Jianfei Zhang 0002, Shengrui Wang, Josiane Courteau, Lifei Chen, Aurélien Bach, Alain Vanasse
ICDM2
2015 Discovering and tracking influencer-influencee relationships between online communities
abstract
This paper addresses a new problem concerning the discovery and tracking of influencer-influencee relationships between communities in dynamic social networks. A weighted temporal multigraph is employed to represent the dynamics of the social networks. To discover and track influencer-influencee relationships over time, communities sharing common interests are first grouped together in meta-communities using a topic modeling approach. Then, influencer-influencee relationships are discovered and tracked using the transfer entropy causality method. Through extensive experiments on the DBLP research publication dataset, we empirically demonstrate the suitability of our model for the discovery of influencer-influencee relationships between communities and the tracking of such relationships over time.
Belkacem Chikhaoui, Mauricio Chiazzaro, Shengrui Wang
DSAA3
2015 Tracking the evolution of community structures in time-evolving social networks
abstract
In real-world social networks, there is increasing interest in tracking the evolution of groups of users. Existing approaches track evolving communities, in a time-sequential way, by comparing communities in terms of nodes using a similarity measure such as the Jaccard or a modified Jaccard measure. The measure allows the use of a one-to-one comparison in order to match communities. However, tracking a given community based on this measure alone may, at the end of its lifespan yield a community that does not share any node with the community initially observed. In this paper we present a novel approach for modeling and detecting the evolution of communities. In our model, we first build a matrix that counts the number of nodes shared between two communities. The individual rows of the obtained matrix are then used to represent nodes shared by a community with all other communities over time. This effectively captures the trace of the communities that should be compared over the period of observation. We then propose a new similarity measure, named mutual transition, for tracking the communities and rules for capturing significant transition events a community can undergo. The proposed approach is general in the sense that it can be applied to different social networks. To demonstrate the suitability of the proposed method, we conducted experiments on real data extracted from the DBLP, Autonomous System and YELP.
Etienne Gael Tajeuna, Mohamed Bouguessa, Shengrui Wang
DSAA3
2014 Pattern-based causal relationships discovery from event sequences for modeling behavioral user profile in ubiquitous environments
Belkacem Chikhaoui, Shengrui Wang, Tengke Xiong, Hélène Pigot
Inf. Sci.2
2014 A Novel Variable-order Markov Model for Clustering Categorical Sequences
abstract
Clustering categorical sequences is an important and difficult data mining task. Despite recent efforts, the challenge remains, due to the lack of an inherently meaningful measure of pairwise similarity. In this paper, we propose a novel variable-order Markov framework, named weighted conditional probability distribution (WCPD), to model clusters of categorical sequences. We propose an efficient and effective approach to solve the challenging problem of model initialization. To initialize the WCPD model, we propose to use a first-order Markov model built on a weighted fuzzy indicator vector representation of categorical sequences, which we call the WFI Markov model. Based on a cascade optimization framework that combines the WCPD and WFI models, we design a new divisive hierarchical clustering algorithm for clustering categorical sequences. Experimental results on data sets from three different domains demonstrate the promising performance of our models and clustering algorithm.
Tengke Xiong, Shengrui Wang, Qingshan Jiang, Joshua Zhexue Huang
IEEE Trans. Knowl. Data Eng.2
2013 Efficient Cluster Labeling for Support Vector Clustering
abstract
We propose a new efficient algorithm for solving the cluster labeling problem in support vector clustering (SVC). The proposed algorithm analyzes the topology of the function describing the SVC cluster contours and explores interconnection paths between critical points separating distinct cluster contours. This process allows distinguishing disjoint clusters and associating each point to its respective one. The proposed algorithm implements a new fast method for detecting and classifying critical points while analyzing the interconnection patterns between them. Experiments indicate that the proposed algorithm significantly improves the accuracy of the SVC labeling process in the presence of clusters of complex shape, while reducing the processing time required by existing SVC labeling algorithms by orders of magnitude.
V. D'Orangeville, André Mayers, Ernest Monga, Shengrui Wang
IEEE Trans. Knowl. Data Eng.4
2013 Information-Theoretic Outlier Detection for Large-Scale Categorical Data
abstract
Outlier detection can usually be considered as a pre-processing step for locating, in a data set, those objects that do not conform to well-defined notions of expected behavior. It is very important in data mining for discovering novel or rare events, anomalies, vicious actions, exceptional phenomena, etc. We are investigating outlier detection for categorical data sets. This problem is especially challenging because of the difficulty of defining a meaningful similarity measure for categorical data. In this paper, we propose a formal definition of outliers and an optimization model of outlier detection, via a new concept of holoentropy that takes both entropy and total correlation into consideration. Based on this model, we define a function for the outlier factor of an object which is solely determined by the object itself and can be updated efficiently. We propose two practical 1-parameter outlier detection methods, named ITB-SS and ITB-SP, which require no user-defined parameters for deciding whether an object is an outlier. Users need only provide the number of outliers they want to detect. Experimental results show that ITB-SS and ITB-SP are more effective and efficient than mainstream methods and can be used to deal with both large and high-dimensional data sets where existing algorithms fail.
Shengrui Wang
IEEE Trans. Knowl. Data Eng.2
2012 Semi-naive Bayesian Classification by Weighted Kernel Density Estimation
Lifei Chen, Shengrui Wang
ADMA2
2012 Automated feature weighting in naive bayes for high-dimensional data classification
abstract
Naive Bayes (NB for short) is one of the popular methods for supervised classification in a knowledge management system. Currently, in many real-world applications, high-dimensional data pose a major challenge to conventional NB classifiers, due to noisy or redundant features and local relevance of these features to classes. In this paper, an automated feature weighting solution is proposed to result in a NB method effective in dealing with high-dimensional data. We first propose a locally weighted probability model, for Bayesian modeling in high-dimensional spaces, to implement a soft feature selection scheme. Then we propose an optimization algorithm to find the weights in linear time complexity, based on the Logitnormal priori distribution and the Maximum a Posteriori principle. Experimental studies show the effectiveness and suitability of the proposed model for high-dimensional data classification.
Lifei Chen, Shengrui Wang
CIKM2
2012 DHCC: Divisive hierarchical clustering of categorical data
Tengke Xiong, Shengrui Wang, André Mayers, Ernest Monga
Data Min. Knowl. Discov.2
2012 Model-Based Method for Projective Clustering
abstract
Clustering high-dimensional data is a major challenge due to the curse of dimensionality. To solve this problem, projective clustering has been defined as an extension to traditional clustering that attempts to find projected clusters in subsets of the dimensions of a data space. In this paper, a probability model is first proposed to describe projected clusters in high-dimensional data space. Then, we present a model-based algorithm for fuzzy projective clustering that discovers clusters with overlapping boundaries in various projected subspaces. The suitability of the proposal is demonstrated in an empirical study done with synthetic data set and some widely used real-world data set.
Lifei Chen, Qingshan Jiang, Shengrui Wang
IEEE Trans. Knowl. Data Eng.3
2011 A New Markov Model for Clustering Categorical Sequences
abstract
Clustering categorical sequences remains an open and challenging task due to the lack of an inherently meaningful measure of pair wise similarity between sequences. Model initialization is an unsolved problem in model-based clustering algorithms for categorical sequences. In this paper, we propose a simple and effective Markov model to approximate the conditional probability distribution (CPD) model, and use it to design a novel two-tier Markov model to represent a sequence cluster. Furthermore, we design a novel divisive hierarchical algorithm for clustering categorical sequences based on the two-tier Markov model. The experimental results on the data sets from three different domains demonstrate the promising performance of our models and clustering algorithm.
Tengke Xiong, Shengrui Wang, Qingshan Jiang, Joshua Zhexue Huang
ICDM2
2011 Semi-supervised Parameter-Free Divisive Hierarchical Clustering of Categorical Data
Tengke Xiong, Shengrui Wang, André Mayers, Ernest Monga
PAKDD (1)2
2011 Rating-based collaborative filtering combined with additional regularization
abstract
The collaborative filtering (CF) approach to recommender system has received much attention recently. However, previous work mainly focuses on improving the formula of rating prediction, e.g. by adding user and item biases, implicit feedback and time-aware factors, etc, to reach a better prediction by minimizing an objective function. However, little effort has been made on improving CF by incorporating additional regularization to the objective function. Regularization can further bound the searching range of predicted ratings. In this paper, we improve the conventional rating-based objective function by using ranking constraints as the supplementary regularization to restrict the searching of predicted ratings in smaller and more likely ranges, and develop a novel method, called RankSVD++, based on the SVD++ model. Experimental results show that RankSVD++ achieves better performance than existing main-streaming methods due to the addition of informative ranking-based regularization. The idea proposed here can also be easily incorporated to the other CF models.
Shengrui Wang
SIGIR2
2011 Measuring the component overlapping in the Gaussian mixture model
Haojun Sun, Shengrui Wang
Data Min. Knowl. Discov.2
2010 A general measure of similarity for categorical sequences
Abdellali Kelil, Shengrui Wang, Qingshan Jiang, Ryszard Brzezinski
Knowl. Inf. Syst.2
2010 Discovering Knowledge-Sharing Communities in Question-Answering Forums
abstract
In this article, we define a knowledge-sharing community in a question-answering forum as a set of askers and authoritative users such that, within each community, askers exhibit more homogeneous behavior in terms of their interactions with authoritative users than elsewhere. A procedure for discovering members of such a community is devised. As a case study, we focus on Yahoo! Answers, a large and diverse online question-answering service. Our contribution is twofold. First, we propose a method for automatic identification of authoritative actors in Yahoo! Answers. To this end, we estimate and then model the authority scores of participants as a mixture of gamma distributions. The number of components in the mixture is determined using the Bayesian Information Criterion (BIC), while the parameters of each component are estimated using the Expectation-Maximization (EM) algorithm. This method allows us to automatically discriminate between authoritative and nonauthoritative users. Second, we represent the forum environment as a type of transactional data such that each transaction summarizes the interaction of an asker with a specific set of authoritative users. Then, to group askers on the basis of their interactions with authoritative users, we propose a parameter-free transaction data clustering algorithm which is based on a novel criterion function. The identified clusters correspond to the communities that we aim to discover. To evaluate the suitability of our clustering algorithm, we conduct a series of experiments on both synthetic data and public real-life data. Finally, we put our approach to work using data from Yahoo! Answers which represent users’ activities over one full year.
Mohamed Bouguessa, Shengrui Wang, Benoît Dumoulin
ACM Trans. Knowl. Discov. Data2
2009 A New MCA-Based Divisive Hierarchical Algorithm for Clustering Categorical Data
abstract
Clustering categorical data faces two challenges, one is lacking of inherent similarity measure, and the other is that the clusters are prone to being embedded in different subspace. In this paper, we propose the first divisive hierarchical clustering algorithm for categorical data. The algorithm, which is based on multiple correspondence analysis (MCA), is systematic, efficient and effective. In our algorithm, MCA plays an important role in analyzing the data globally. The proposed algorithm has five merits. First, our algorithm yields a dendrogram representing nested groupings of patterns and similarity levels at different granularities. Second, it is parameter-free, fully automatic and, most importantly, requires no assumption regarding the number of clusters. Third, it is independent of the order in which the data are processed. Forth, it is scalable to large data sets; and finally, using the novel data representation and Chi-square distance measures makes our algorithm capable of seamlessly discovering the clusters embedded in the subspaces. Experiments on both synthetic and real data demonstrate the superior performance of our algorithm.
Tengke Xiong, Shengrui Wang, André Mayers, Ernest Monga
ICDM2
2009 Mining Projected Clusters in High-Dimensional Spaces
abstract
Clustering high-dimensional data has been a major challenge due to the inherent sparsity of the points. Most existing clustering algorithms become substantially inefficient if the required similarity measure is computed between data points in the full-dimensional space. To address this problem, a number of projected clustering algorithms have been proposed. However, most of them encounter difficulties when clusters hide in subspaces with very low dimensionality. These challenges motivate our effort to propose a robust partitional distance-based projected clustering algorithm. The algorithm consists of three phases. The first phase performs attribute relevance analysis by detecting dense and sparse regions and their location in each attribute. Starting from the results of the first phase, the goal of the second phase is to eliminate outliers, while the third phase aims to discover clusters in different subspaces. The clustering process is based on the k-means algorithm, with the computation of distance restricted to subsets of attributes where object values are dense. Our algorithm is capable of detecting projected clusters of low dimensionality embedded in a high-dimensional space and avoids the computation of the distance in the full-dimensional space. The suitability of our proposal has been demonstrated through an empirical study using synthetic and real datasets.
Mohamed Bouguessa, Shengrui Wang
IEEE Trans. Knowl. Data Eng.2
2008 A Probability Model for Projective Clustering on High Dimensional Data
abstract
Clustering high dimensional data is a big challenge in data mining due to the curse of dimensionality. To solve this problem, projective clustering has been defined as an extension of traditional clustering that seeks to find projected clusters in subsets of dimensions of a data space. In this paper, the problem of modeling projected clusters is first discussed, and an extended Gaussian model is proposed. Second, a general objective criterion used with k-means type projective clustering is presented based on the model. Finally, the expressions to learn model parameters are derived and then used in a new algorithm named FPC to perform fuzzy clustering on high dimensional data. The experimental results on document clustering show the effectiveness of the proposed clustering model.
Lifei Chen, Qingshan Jiang, Shengrui Wang
ICDM3
2008 SCS: A New Similarity Measure for Categorical Sequences
abstract
Measuring the similarity between categorical sequences is a fundamental process in many data mining applications. A key issue is to extract and make use of significant features hidden behind the chronological and structural dependencies found in these sequences. Almost all existing algorithms designed to perform this task are based on the matching of patterns in chronological order, but such sequences often have similar structural features in chronologically different positions. In this paper we propose SCS, a novel method for measuring the similarity between categorical sequences, based on an original pattern matching scheme that makes it possible to capture chronological and non-chronological dependencies. SCS captures significant patterns that represent the natural structure of sequences, and reduces the influence of those representing noise. It constitutes an effective approach for measuring the similarity of data such as biological sequences, natural language texts and financial transactions. To show its effectiveness, we have tested SCS extensively on a range of datasets, and compared the results with those obtained by various mainstream algorithms.
Abdellali Kelil, Shengrui Wang
ICDM2
2008 Identifying authoritative actors in question-answering forums: the case of Yahoo! answers
abstract
We consider the problem of identifying authoritative users in Yahoo! Answers. A common approach is to use link analysis techniques in order to provide a ranked list of users based on their degree of authority. A major problem for such an approach is determining how many users should be chosen as authoritative from a ranked list. To address this problem, we propose a method for automatic identification of authoritative actors. In our approach, we propose to model the authority scores of users as a mixture of gamma distributions. The number of components in the mixture is estimated by the Bayesian Information Criterion (BIC) while the parameters of each component are estimated using the Expectation-Maximization (EM) algorithm. This method allows us to automatically discriminate between authoritative and non-authoritative users. The suitability of our proposal is demonstrated in an empirical study using datasets from Yahoo! Answers.
Mohamed Bouguessa, Benoît Dumoulin, Shengrui Wang
KDD3
2007 PCGEN: A Practical Approach to Projected Clustering and its Application to Gene Expression Data
abstract
Clustering samples in gene expression data has always been a major challenge because of the high dimensionality of the input space (typically in the tens of thousands) and the small number of samples (typically less than a hundred). Moreover, clusters may hide in subspaces with very low dimensionalities. Most existing clustering algorithms become substantially inefficient if the required similarity measure is computed between data points in the full-dimensional space. These challenges motivate our effort to propose a new and efficient partitional distance-based projected clustering algorithm for clustering samples in gene expression data. Our algorithm is capable of detecting projected clusters of extremely low dimensionality embedded in a high-dimensional space and avoids the computation of the distance in the full-dimensional space. The suitability of our proposal has been demonstrated through an empirical study using public microarray datasets.
Mohamed Bouguessa, Shengrui Wang
CIDM2
2005 Address Extraction: Extraction of Location-Based Information from the Web
Wentao Cai 0003, Shengrui Wang, Qingshan Jiang
APWeb2
2005 An Experiment on the Matching and Reuse of XML Schemas
Jianguo Lu, Shengrui Wang, Ju Wang 0009
ICWE2
2005 Approximate Common Structures in XML Schema Matching
Shengrui Wang, Jianguo Lu, Ju Wang 0009
WAIM1
2001 An Artificial Network Simulating Cause-to-Effect Reasoning: Cancellation Interactions and Numerical Studies
Lotfi Ben Romdhane 0001, Béchir el Ayeb, Shengrui Wang
Knowl. Inf. Syst.3