VLDB 2026 Research / reviewers in the wild / expert
Masoud Makrehchi
dblp:10/5941
· DBLP profile ↗
20ranked-venue papers in the field
8as first author
7since 2021 · last 2025
0000-0003-0287-9568ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 9 (4 first)Information Retrieval & Web Search · 6 (3 first)Data Mining & Knowledge Discovery · 4 (1 first)Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LineDi2Vec: An Edge-Based Graph Embedding on Signed Social Networks
Chen Xing, Masoud Makrehchi |
ASONAM (2) | 2 |
| 2024 | Coherence Graphs: Bridging the Gap in Text Segmentation with Unsupervised Learning
Amit Maraj, Miguel Vargas Martin, Masoud Makrehchi |
NLDB (2) | 3 |
| 2023 | The 3rd International Workshop on Mining and Learning in the Legal DomainabstractThe increasing accessibility of legal corpora and databases create opportunities to develop data-driven techniques and advanced tools that can facilitate a variety of tasks in the legal domain, such as legal search and research, legal document review and summary, legal contract drafting, and legal outcome prediction. Compared with other application domains, the legal domain is characterized by the huge scale of natural language text data, the high complexity of specialist knowledge, and the critical importance of ethical considerations. The MLLD workshop aims to bring together researchers and practitioners to share the latest research findings and innovative approaches in employing data mining, machine learning, information retrieval, and knowledge management techniques to transform the legal sector. Building upon the previous successes, the third edition of the MLLD workshop will emphasize the exploration of new research opportunities brought about by recent rapid advances in Large Language Models and Generative AI. We encourage submissions that intersect computer science and law, from both academia and industry, embodying the interdisciplinary spirit of CIKM. Masoud Makrehchi, Dell Zhang, Alina Petrova, John Armour |
CIKM | 1 |
| 2023 | Uncertainty Quantification for Text Classification
Dell Zhang, Murat Sensoy, Masoud Makrehchi, Bilyana Taneva-Popova |
ECIR (3) | 3 |
| 2023 | Making a Computational AttorneyabstractThis “blue sky idea” paper outlines the opportunities and challenges in data mining and machine learning involving making a computational attorney — an intelligent software agent capable of helping human lawyers with a wide range of complex high-level legal tasks such as drafting legal briefs for the prosecution or defense in court. In particular, we discuss what a ChatGPT-like Large Legal Language Model (L3M) can and cannot do today, which will inspire researchers with promising short-term and long-term research objectives. Dell Zhang, Frank Schilder, Jack G. Conrad, Masoud Makrehchi, David von Rickenbach, Isabelle Moulinier |
SDM | 4 |
| 2023 | Uncertainty Quantification for Text ClassificationabstractThis full-day tutorial introduces modern techniques for practical uncertainty quantification specifically in the context of multi-class and multi-label text classification. First, we explain the usefulness of estimating aleatoric uncertainty and epistemic uncertainty for text classification models. Then, we describe several state-of-the-art approaches to uncertainty quantification and analyze their scalability to big text data: Virtual Ensemble in GBDT, Bayesian Deep Learning (including Deep Ensemble, Monte-Carlo Dropout, Bayes by Backprop, and their generalization Epistemic Neural Networks), Evidential Deep Learning (including Prior Networks and Posterior Networks), as well as Distance Awareness (including Spectral-normalized Neural Gaussian Process and Deep Deterministic Uncertainty). Next, we talk about the latest advances in uncertainty quantification for pre-trained language models (including asking language models to express their uncertainty, interpreting uncertainties of text classifiers built on large-scale language models, uncertainty estimation in text generation, calibration of language models, and calibration for in-context learning). After that, we discuss typical application scenarios of uncertainty quantification in text classification (including in-domain calibration, cross-domain robustness, and novel class detection). Finally, we list popular performance metrics for the evaluation of uncertainty quantification effectiveness in text classification. Practical hands-on examples/exercises are provided to the attendees for them to experiment with different uncertainty quantification methods on a few real-world text classification datasets such as CLINC150. Dell Zhang, Murat Sensoy, Masoud Makrehchi, Bilyana Taneva-Popova, Lin Gui 0003, Yulan He 0001 |
SIGIR | 3 |
| 2021 | A More Effective Sentence-Wise Text Segmentation Approach Using BERT
Amit Maraj, Miguel Vargas Martin, Masoud Makrehchi |
ICDAR (4) | 3 |
| 2019 | Basketball lineup performance prediction using network analysisabstractWinning a game in professional sports is the most significant matter for a team. All teams strive to bring their best performance to a game, and this requires considering all the possible lineups which coaches have available. Therefore, determining the lineup is more and more significant for a team in their winning endeavour. The ongoing result during a game defines the next decision coaches have to make to maintain or improve the outcome. Adaptive changes in a lineup of a team requires a complex decision making system. This system must consider the advantages, drawbacks, and previous experience about both teams' performance under similar situations. In order to analyze and predict lineups' performance, the authors create a directed, weighted, and signed network of all lineups that teams use against each other from 2007-2016 seasons in National Basketball Association (NBA) games. The proposed model uses machine learning and network analysis techniques to predict the performance of a lineup under a given situation by utilizing graph theory and Inverse Squared Metric. Mahboubeh Ahmadalinezhad, Masoud Makrehchi, Neil Seward |
ASONAM | 2 |
| 2016 | Mining Social Media Content for Crime PredictionabstractSocial media provides increasing opportunities for users to voluntarily share their thoughts and concerns in a large volume of data. While user-generated data from each individual may not provide considerable information, when combined, they include hidden variables, which may convey significant events. In this paper, we pursue the question of whether social media context can provide socio-behavior "signals" for crime prediction. The hypothesis is that crowd publicly available data in social media, in particular Twitter, may include predictive variables, which can indicate the changes in crime rates. We developed a model for crime trend prediction where the objective is to employ Twitter content to identify whether crime rates have dropped or increased for the prospective time frame. We also present a Twitter sampling model to collect historical data to avoid missing data over time. The prediction model was evaluated for different cities in the United States. The experiments revealed the correlation between features extracted from the content and crime rate directions. Overall, the study provides insight into the correlation of social content and crime trends as well as the impact of social data in providing predictive indicators. Somayyeh Aghababaei, Masoud Makrehchi |
WI | 2 |
| 2016 | We Didn't Miss You: Interpolating Missing OpinionsabstractWhen mining user streams from social media, activity gaps are inevitable, which is known as the sparsity of user data. Such sparsity can significantly degrade the performance of a predictive system that relies on time-sensitive user content. To mitigate this issue, conventional approaches generally tend to discard periods with missing data. However, this solution leads to neglecting information generated by other users which, if utilized, could potentially enhance the quality of the predictive model. So the following question arises: is it possible to alleviate the impact of absent data while preserving the available content contributed within the same timespan? Despite the fact that this problem is well-known, it has not been thoroughly studied before. The goal of this work is to find a way of interpolating missing data from user's network and his previous activities. We investigate how different types of user profiles affect overall behavior predictability. Proposed models are evaluated on a case study of a micro-blogging system for the investment community. Iuliia Chepurna, Masoud Makrehchi |
WI | 2 |
| 2016 | Social Filtering: User-Centric Approach to Social Trend PredictionabstractThe majority of techniques in socio-behavioral modeling tend to consider user-generated content in a bulk, with the assumption that this sort of aggregation would not have any negative impact on overall predictability of the system, which is not necessarily the case. We propose a novel user-centric approach designed specifically to capture most predictive hidden variables that can be discovered in a context of the specific individual. The concept of social filtering closely resembles collaborative filtering with the main difference that none of the considered users intentionally participates in the recommendation process. Its objective is to determine both the subset of best expert users able to reflect a particular social trend of interest and their transformation into feature space used for modeling. We introduce three-step selection procedure that includes activity-and relevance-based filtering and ensemble of expert users, and show that proper choice of expert individuals is critical to prediction quality. Iuliia Chepurna, Masoud Makrehchi |
WI | 2 |
| 2016 | Discovering Credible Twitter Users in Stock Market DomainabstractDespite extensive research efforts in stock market predictions using social media networks, there still exists a lack of credible sources of information in such media. This study presents a novel approach to measure the credibility of Twitter users in a domain of interest, namely the stock market. This study suggests a correlation between each user's credibility and the extracted features from each follower network: number of followers, number of stock market-related followers, extracted by a $Cashtag-based approach, ratio of stock market-related followers to the total number of followers, and the number of seed user tweets. The results support the initial hypothesis of this study. Mehran Kamkarhaghighi, Iuliia Chepurna, Somayyeh Aghababaei, Masoud Makrehchi |
WI | 4 |
| 2016 | Hierarchical Agglomerative Clustering Using Common Neighbours SimilarityabstractHierarchical clustering has been well-studied in the community of machine learning. Hierarchical clustering algorithms are deterministic, stable, and do not need a pre-determined number of clusters as input. However, they are not scalable for very large data due to their non-linear complexity. In this paper, a new approach is proposed to reduce the complexity of Hierarchical Clustering, improve the purity of the clustering algorithm, and reduce the chaining factor. The proposed method has the following components: (i) A new combination similarity based on common-neighbours of graph theory is proposed, (ii) In every iteration, instead of calculating the centroids for new clusters, new centroids are estimated from centroids in previous iteration, and (iii) In each iteration, instead of merging only one pair of objects, multiple pairs are merged at the same time. In addition to the proposed combination similarity, four well-known methods including centroid-based, group-based, complete-link, and single-link, have been also implemented. All five methods are tested and evaluated using two metrics: purity and imbalance or chaining factor. We show that our proposed algorithm outperforms other classic methods. Masoud Makrehchi |
WI | 1 |
| 2013 | Stock Prediction Using Event-Based Sentiment AnalysisabstractWe propose a novel approach to label social media text using significant stock market events (big losses or gains). Since stock events are easily quantifiable using returns from indices or individual stocks, they provide meaningful and automated labels. We extract significant stock movements and collect appropriate pre, post and contemporaneous text from social media sources (for example, tweets from twitter). Subsequently, we assign the respective label (positive or negative) for each tweet. We train a model on this collected set and make predictions for labels of future tweets. We aggregate the net sentiment per each day (amongst other metrics) and show that it holds significant predictive power for subsequent stock market movement. We create successful trading strategies based on this system and find significant returns over other baseline methods. Masoud Makrehchi, Sameena Shah, Wenhui Liao |
Web Intelligence | 1 |
| 2011 | Social link recommendation by learning hidden topicsabstractIn this paper, a new approach to predicting the structure of a social network without any prior knowledge from the social links is proposed. In absence of links among nodes, we assume there are other information resources associated with the nodes which are called node profiles. The task of link prediction and recommendation from text data is to learn similarities between the nodes and then translate pair-wise similarities into social links. In other words, the process is to convert a similarity matrix into an adjacency matrix. In this paper, an alternative approach is proposed. First, hidden topics of node profiles are learned using Latent Dirichlet Allocation. Then, by mapping node-topic and topic-topic relations, a new structure called semi-bipartite graph is generated which is slightly different from regular bipartite graph. Finally, by applying topological metrics such as Katz and short path scores to the new structure, we are able to rank and recommend relevant links to each node. The proposed technique is applied to several co-authorship networks. While most link prediction methods are low precision solutions, the proposed method performs effectively and offers high precision. Masoud Makrehchi |
RecSys | 1 |
| 2010 | Filter-Based Data Partitioning for Training Multiple Classifier SystemsabstractData partitioning methods such as bagging and boosting have been extensively used in multiple classifier systems. These methods have shown a great potential for improving classification accuracy. This study is concerned with the analysis of training data distribution and its impact on the performance of multiple classifier systems. In this study, several feature-based and class-based measures are proposed. These measures can be used to estimate statistical characteristics of the training partitions. To assess the effectiveness of different types of training partitions, we generated a large number of disjoint training partitions with distinctive distributions. Then, we empirically assessed these training partitions and their impact on the performance of the system by utilizing the proposed feature-based and class-based measures. We applied the findings of this analysis and developed a new partitioning method called "Clustering, Declustering, and Selection" (CDS). This study presents a comparative analysis of several existing data partitioning methods including our proposed CDS approach. Rozita Dara 0001, Masoud Makrehchi, Mohamed S. Kamel |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2008 | Automatic Extraction of Domain-Specific Stopwords from Labeled Documents
Masoud Makrehchi, Mohamed S. Kamel |
ECIR | 1 |
| 2007 | A Text Classification Framework with a Local Feature Ranking for Learning Social NetworksabstractIn this paper, a text classifier framework with a feature ranking scheme is proposed to extract social structures from text data. It is assumed that only a small subset of relations between the individuals in a community is known. With this assumption, the social network extraction is translated into a classification problem. The relations between two individuals are represented by merging their document vectors and the given relations are used as labels of training data. By this transformation, a text classifier such as Rocchio is used for learning the unknown relations. We show that there is a link between the intrinsic sparsity of social networks and class imbalance. Furthermore, we show that feature ranking methods usually fail in problem with unbalanced data. In order to deal with this deficiency and re-balance the unbalanced social data, a local feature ranking method, which is called reverse discrimination, is proposed. Masoud Makrehchi, Mohamed S. Kamel |
ICDM | 1 |
| 2007 | Automatic Taxonomy Extraction Using Google and Term DependencyabstractAn automatic taxonomy extraction algorithm is proposed. Given a set of terms or terminology related to a subject domain, the proposed approach uses Google page count to estimate the dependency links between the terms. A taxonomic link is an asymmetric relation between two concepts. In order to extract these directed links, neither mutual information nor normalized Google distance can be employed. Using the new measure of information theoretic inclusion index, term dependency matrix, which represents the pair-wise dependencies, is obtained. Next, using a proposed algorithm, the dependency matrix is converted into an adjacency matrix, representing the taxonomy tree. In order to evaluate the performance of the proposed approach, it is applied to several domains for taxonomy extraction. Masoud Makrehchi, Mohamed S. Kamel |
Web Intelligence | 1 |
| 2006 | Learning Social Networks from Web Documents Using Support Vector ClassifiersabstractAutomatic generation of a social network requires extracting pair-wise relations of the individuals. In this research, learning social network from incomplete relationship data is proposed. It is assumed that only a small subset of relations between the individuals is known. With this assumption, the social network extraction is translated into a text classification problem. The relations between two individuals are modeled by merging their document vectors and the given relations are used as labels of training data. By this transformation, a text classifier such as SVM is used for learning the unknown relations. We show that there is a link between the intrinsic sparsity of social networks and class distribution imbalance of the training data. In order to re-balance the unbalanced training data, a minority class down-sampling strategy is employed. The proposed framework is applied to a true FOAF (friend of a friend) database and evaluated by the macro-averaged F-measure Masoud Makrehchi, Mohamed S. Kamel |
Web Intelligence | 1 |