EDBT 2026 Demo / reviewers in the wild / expert
Faisal Kamiran
dblp:07/7790
· DBLP profile ↗
24ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-1168-9451ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 16 · 5 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multiview Commonsense Reasoning Using LLMs for Understanding Crime Drama Series
Muhammad Abdullah Zia, Sameen Mansha, Faisal Kamiran |
ASONAM (1) | 3 |
| 2025 | Reproducibility and Case Sensitivity of LLMs for Anonymizing Depressed TweetsabstractA careful analysis of the Large Language Model (LLM) results, generated through anonymized representations of the original dataset, is crucial to precisely evaluate the data-sharing procedure's limitations and facilitate valuable collaborations among Internet-based cognitive behavioral therapy (ICBT) companies and third parties. This paper presents an experimental study of fine-tuning 27 LMs for a multiclass classification task to identify depression severity using 40,191 tweets labeled by human annotators. We fine-tune 14 Bidirectional Encoder Representations from Transformers (BERT), 6 Robustly Optimized BERT Pretraining Approaches (RoBerta), 3 Generative Pretraining (GPT), and 4 Text-to-Text Transfer Transformer (T5) based LMs to classify confidential and anonymized tweets. We report that T5, through conditional generation, outperforms widely adopted BERT, RoBerta, and GPT types for classifying confidential and anonymized tweets. Anonymizing personal information safeguards user privacy and often increases LM performance. Case sensitivity can potentially improve or harm the performance of domain-specific LMs for original and anonymized text. Sameen Mansha, Hamza Mahmood, Anne Håkansson, Faisal Kamiran, Vladimir Vlassov |
DSAA | 4 |
| 2025 | FairUDT: Fairness-aware Uplift Decision Trees
Anam Zahid, Abdur Rehman Ali, Shaina Raza, Rai Shahnawaz, Faisal Kamiran, Asim Karim |
Knowl. Based Syst. | 5 |
| 2024 | Detecting Cybercrimes in Accordance with Pakistani Law: Dataset and Evaluation Using PLMsabstractCybercrime is a serious and growing threat affecting millions of people worldwide. Detecting cybercrimes from text messages is challenging, as it requires understanding the linguistic and cultural nuances of different languages and regions. Roman Urdu is a widely used language in Pakistan and other South Asian countries, however, it lacks sufficient resources and tools for natural language processing and cybercrime detection. To address this problem, we make three main contributions in this paper. (1) We create and release CRU, a benchmark dataset for text-based cybercrime detection in Roman Urdu, which covers a number of cybercrimes as defined by the Prevention of Electronic Crimes Act (PECA) of Pakistan. This dataset is annotated by experts following a standardized procedure based on Pakistan’s legal framework. (2) We perform experiments on four pre-trained language models (PLMs) for cybercrime text classification in Roman Urdu. Our results show that xlm-roberta-base is the best model for this task, achieving the highest performance on all metrics. (3) We explore the utility of prompt engineering techniques, namely prefix and cloze prompts, for enhancing the performance of PLMs for low-resource languages such as Roman Urdu. We analyze the impact of different prompt shapes and k-shot settings on the performance of xlm-roberta-base and bert-base-multilingual-cased. We find that prefix prompts are more effective than cloze prompts for Roman Urdu classification tasks, as they provide more contextually relevant completions for the models. Our work provides useful insights and resources for future research on cybercrime detection and text classification in low-resource languages. Faizad Ullah, Ali Faheem, Ubaid Azam, Muhammad Sohaib Ayub, Faisal Kamiran, Asim Karim |
LREC/COLING | 5 |
| 2022 | Locality Aware Temporal FMs for Crime PredictionabstractCrime forecasting techniques can play a leading role in hindering crime occurrences, especially in areas under possible threat. In this paper, we propose Locality Aware Temporal Factorization Machines (LTFMs) for crime prediction. Its locality representation module deploys a spatial encoder to estimate the regional dependencies using Graph Convolutional Networks (GCNs). Then, the Point of Interest (POI) encoder computes the weighted attentive aggregation of location, crime, and POI latent representations. The dynamic crime representation module utilizes the transformer-based positional encodings to capture the dependencies among space, time, and crime categories. The encodings learnt from locality representation and crime category encoders, are projected into a factorization machine-based architecture via a shared feed-forward network. An extensive comparison with state-of-art techniques, using Chicago and New York's criminal records, shows the significance of LTFMs. Sameen Mansha, Shaaf Abdullah, Faisal Kamiran, Hongzhi Yin |
CIKM | 4 |
| 2022 | A clustering framework for lexical normalization of Roman UrduabstractAbstract Roman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content. It lacks standard spelling and hence poses several normalization challenges during automatic language processing. In this article, we present a feature-based clustering framework for the lexical normalization of Roman Urdu corpora, which includes a phonetic algorithm UrduPhone, a string matching component, a feature-based similarity function, and a clustering algorithm Lex-Var. UrduPhone encodes Roman Urdu strings to their pronunciation-based representations. The string matching component handles character-level variations that occur when writing Urdu using Roman script. The similarity function incorporates various phonetic-based, string-based, and contextual features of words. The Lex-Var algorithm is a variant of the k-medoids clustering algorithm that groups lexical variations of words. It contains a similarity threshold to balance the number of clusters and their maximum similarity. The framework allows feature learning and optimization in addition to the use of predefined features and weights. We evaluate our framework extensively on four real-world datasets and show an F-measure gain of up to 15% from baseline methods. We also demonstrate the superiority of UrduPhone and Lex-Var in comparison to respective alternate algorithms in our clustering framework for the lexical normalization of Roman Urdu. Abdul Rafae Khan, Asim Karim, Hassan Sajjad 0001, Faisal Kamiran, Jia Xu 0004 |
Nat. Lang. Eng. | 4 |
| 2021 | GDFM: Gene Vectors Embodied Deep Attentional Factorization Machines for Interaction predictionabstractGene Network Graphs (GNGs) are comprised of biomedical data. Deriving structural information from these graphs remains a prime area of research in the domain of biomedical and health informatics. In this paper, we propose Gene Vectors Embodied Deep Attentional Factorization Machines (GDFMs) for the gene to gene interaction prediction. We first initialize GDFM with vector embeddings learned from gene locality configuration and an expression equivalence criterion that preserves their innate similar traits. GDFM uses an attention-based mechanism that manipulates different positions, to learn the representation of sequence, before calculating the pairwise factorized interactions. We further use hidden layers, batch normalization, and dropout to stabilize the performance of our deep structured architecture. An extensive comparison with several state-of-the-art approaches, using Ecoli and Yeast datasets for gene-gene interaction prediction shows the significance of our proposed framework. Sameen Mansha, Tayyab Khalid, Faisal Kamiran, Masroor Hussain, Syed Fawad Hussain, Hongzhi Yin |
CIKM | 3 |
| 2020 | Causal inference for social discrimination reasoning
Bilal Qureshi, Faisal Kamiran, Asim Karim, Salvatore Ruggieri, Dino Pedreschi |
J. Intell. Inf. Syst. | 2 |
| 2019 | An assessment of SMS fraud in PakistanabstractSMS fraud has become a growing concern for those working toward financial inclusion, however, it is often unclear how widespread such threats are in practice. This multi-method study investigates SMS fraud in Pakistan through identification and categorization of fraudulent messages as well as the impact on those who receive such messages. We collect fraudulent SMS messages by various means, including byway of a custom-built Android smartphone application. To complement this, we interview people exposed to SMS fraud and representatives of mobile network operators. Based on our analysis, lottery type fraud schemes dominate SMS fraud in Pakistan, and these schemes have the greatest impact on vulnerable low-income, rural populations. We offer a simple heuristic for fraud detection that has a high accuracy rate and is adaptable to evolving fraud schemes, and conclude with a recommendation for a fraud mitigation strategy to target fraudster call back numbers. Fahad Pervaiz, Rai Shahnawaz, Muhammad Umer Ramzan, Maryem Zafar Usmani, Shrirang Mare, Kurtis Heimerl, Faisal Kamiran, Richard J. Anderson 0001, Lubna Razaq |
COMPASS | 7 |
| 2019 | Balancing Prediction Errors for Robust Sentiment ClassificationabstractSentiment classification is a popular text mining task in which textual content (e.g., a message) is assigned a polarity label (typically positive or negative) reflecting the sentiment expressed in it. Sentiment classification is used widely in applications like customer feedback analysis where robustness and correctness of results are critical. In this article, we highlight that prediction accuracy alone is not sufficient for assessing the performance of a sentiment classifier; it is also important that the classifier is not biased toward positive or negative polarity, thus distorting the distribution of positive and negative messages in the predictions. We propose a measure, called Polarity Bias Rate, for quantifying this bias in a sentiment classifier. Second, we present two methods for removing this bias in the predictions of unsupervised and supervised sentiment classifiers. Our first method, called Bias-Aware Thresholding (BAT), shifts the decision boundary to control the bias in the predictions. Motivated from cost-sensitive learning, BAT is easily applicable to both lexicon-based unsupervised and supervised classifiers. Our second method, called Balanced Logistic Regression (BLR) introduces a bias-remover constraint into the standard logistic regression model. BLR is an automatic bias-free supervised sentiment classifier. We evaluate our methods extensively on seven real-world datasets. The experiments involve two lexicon-based and two supervised sentiment classifiers and include evaluation on multiple train-test data sizes. The results show that bias is controlled effectively in predictions. Furthermore, prediction accuracy is also increased in many cases, thus enhancing the robustness of sentiment classification. Mohsin Iqbal, Asim Karim, Faisal Kamiran |
ACM Trans. Knowl. Discov. Data | 3 |
| 2019 | Layered convolutional dictionary learning for sparse coding itemsets
Sameen Mansha, Hoang Thanh Lam, Hongzhi Yin, Faisal Kamiran, Mohsen Ali |
World Wide Web | 4 |
| 2018 | Exploiting reject option in classification for social discrimination control
Faisal Kamiran, Sameen Mansha, Asim Karim, Xiangliang Zhang 0001 |
Inf. Sci. | 1 |
| 2017 | Modeling Temporal Behavior of Awards Effect on Viewership of Movies
Basmah Altaf, Faisal Kamiran, Xiangliang Zhang 0001 |
PAKDD (1) | 2 |
| 2016 | A Self-Organizing Map for Identifying InfluentialCommunities in Speech-based NetworksabstractLow-literate people are unable to use many mainstream social networks due to their text-based interfaces even though they constitute a major portion of the world population. Specialized speech-based networks (SBNs) are more accessible to low-literate users through their simple speech-based interfaces. While SBNs have the potential for providing value-adding services to a large segment of society they have been hampered by the need to operate in low-income segments on low budgets. The knowledge of influential users and communities in such networks can help in optimizing their operations. In this paper, we present a self-organizing map (SOM) for discovering and visualizing influential communities of users in SBNs. We demonstrate how a friendship graph is formed from call data records and present a method for estimating influences between users. Subsequently, we develop a SOM to cluster users based on their influence, thus identifying community-level influences and their roles in information propagation. We test our approach on Polly, a SBN developed for job ads dissemination among low-literate users. For comparison, we identify influential users with the benchmark greedy algorithm and relate them to the discovered communities. The results show that influential users are concentrated in influential communities and community-level information propagation provides a ready summary of influential users. Sameen Mansha, Faisal Kamiran, Asim Karim, Aizaz Anwar |
CIKM | 2 |
| 2016 | Neural Network Based Association Rule Mining from Uncertain Data
Sameen Mansha, Zaheer Babar, Faisal Kamiran, Asim Karim |
ICONIP (4) | 3 |
| 2015 | An Unsupervised Method for Discovering Lexical Variations in Roman Urdu Informal TextabstractWe present an unsupervised method to find lexical variations in Roman Urdu informal text.Our method includes a phonetic algorithm UrduPhone, a featurebased similarity function, and a clustering algorithm Lex-C.UrduPhone encodes roman Urdu strings to their phonetic equivalent representations.This produces an initial grouping of different spelling variations of a word.The similarity function incorporates word features and their context.Lex-C is a variant of k-medoids clustering algorithm that group lexical variations.It incorporates a similarity threshold to balance the number of clusters and their maximum similarity.We test our system on two datasets of SMS and blogs and show an f-measure gain of up to 12% from baseline systems. Abdul Rafae, Muhammad Moeen Uddin, Asim Karim, Hassan Sajjad 0001, Faisal Kamiran |
EMNLP | 6 |
| 2015 | Multi-query Optimization in Federated Databases Using Evolutionary AlgorithmabstractMulti Query Optimization in federated database systems is a well-studied area. Studies have shown that similar problem arises in wide range of applications, e.g., distributed stream processing systems and wireless sensor networks. In this paper, a general distributed multiquery processing problem motivated by the need to speedup data acquisition in federated databases using evolutionary algorithm is studied. We setup a simple framework in which each individual in population is evolved in terms of cost, uniform labeling of hyper edges and validity of resource constraints through a number of generations. Variations of our general problem can be shown to be NP-Hard. Our extensive empirical evaluation over five different synthetic datasets shows a significant improvement of 8 percent in results as compared to the state-of-the-art methods. Sameen Mansha, Faisal Kamiran |
ICMLA | 2 |
| 2014 | Anti-discrimination Analysis Using Privacy Attack Strategies
Salvatore Ruggieri, Sara Hajian, Faisal Kamiran, Xiangliang Zhang 0001 |
ECML/PKDD (2) | 3 |
| 2013 | Controlling Attribute Effect in Linear RegressionabstractIn data mining we often have to learn from biased data, because, for instance, data comes from different batches or there was a gender or racial bias in the collection of social data. In some applications it may be necessary to explicitly control this bias in the models we learn from the data. This paper is the first to study learning linear regression models under constraints that control the biasing effect of a given attribute such as gender or batch number. We show how propensity modeling can be used for factoring out the part of the bias that can be justified by externally provided explanatory attributes. Then we analytically derive linear models that minimize squared error while controlling the bias by imposing constraints on the mean outcome or residuals of the models. Experiments with discrimination-aware crime prediction and batch effect normalization tasks show that the proposed techniques are successful in controlling attribute effects in linear regression models. Toon Calders, Asim Karim, Faisal Kamiran, Wasif Ali, Xiangliang Zhang 0001 |
ICDM | 3 |
| 2013 | Quantifying explainable discrimination and removing illegal discrimination in automated decision making
Faisal Kamiran, Indre Zliobaite, Toon Calders |
Knowl. Inf. Syst. | 1 |
| 2012 | Decision Theory for Discrimination-Aware ClassificationabstractSocial discrimination (e.g., against females) arising from data mining techniques is a growing concern worldwide. In recent years, several methods have been proposed for making classifiers learned over discriminatory data discrimination-aware. However, these methods suffer from two major shortcomings: (1) They require either modifying the discriminatory data or tweaking a specific classification algorithm and (2) They are not flexible w.r.t. discrimination control and multiple sensitive attribute handling. In this paper, we present two solutions for discrimination-aware classification that neither require data modification nor classifier tweaking. Our first and second solutions exploit, respectively, the reject option of probabilistic classifier(s) and the disagreement region of general classifier ensembles to reduce discrimination. We relate both solutions with decision theory for better understanding of the process. Our experiments using real-world datasets demonstrate that our solutions outperform existing state-of-the-art methods, especially at low discrimination which is a significant advantage. The superior performance coupled with flexible control over discrimination and easy applicability to multiple sensitive attributes makes our solutions an important step forward in practical discrimination-aware classification. Faisal Kamiran, Asim Karim, Xiangliang Zhang 0001 |
ICDM | 1 |
| 2011 | Handling Conditional DiscriminationabstractHistorical data used for supervised learning may contain discrimination. We study how to train classifiers on such data, so that they are discrimination free with respect to a given sensitive attribute, e.g., gender. Existing techniques that deal with this problem aim at removing all discrimination and do not take into account that part of the discrimination may be explainable by other attributes, such as, e.g., education level. In this context, we introduce and analyze the issue of conditional non-discrimination in classifier design. We show that some of the differences in decisions across the sensitive groups can be explainable and hence tolerable. We observe that in such cases, the existing discrimination aware techniques will introduce a reverse discrimination, which is undesirable as well. Therefore, we develop local techniques for handling conditional discrimination when one of the attributes is considered to be explanatory. Experimental evaluation demonstrates that the new local techniques remove exactly the bad discrimination, allowing differences in decisions as long as they are explainable. Indre Zliobaite, Faisal Kamiran, Toon Calders |
ICDM | 2 |
| 2011 | Data preprocessing techniques for classification without discriminationabstractRecently, the following Discrimination-Aware Classification Problem was introduced: Suppose we are given training data that exhibit unlawful discrimination; e.g., toward sensitive attributes such as gender or ethnicity. The task is to learn a classifier that optimizes accuracy, but does not have this discrimination in its predictions on test data. This problem is relevant in many settings, such as when the data are generated by a biased decision process or when the sensitive attribute serves as a proxy for unobserved features. In this paper, we concentrate on the case with only one binary sensitive attribute and a two-class classification problem. We first study the theoretically optimal trade-off between accuracy and non-discrimination for pure classifiers. Then, we look at algorithmic solutions that preprocess the data to remove discrimination before a classifier is learned. We survey and extend our existing data preprocessing techniques, being suppression of the sensitive attribute, massaging the dataset by changing class labels, and reweighing or resampling the data to remove discrimination without relabeling instances. These preprocessing techniques have been implemented in a modified version of Weka and we present the results of experiments on real-life data. Faisal Kamiran, Toon Calders |
Knowl. Inf. Syst. | 1 |
| 2010 | Discrimination Aware Decision Tree LearningabstractRecently, the following discrimination aware classification problem was introduced: given a labeled dataset and an attribute B, find a classifier with high predictive accuracy that at the same time does not discriminate on the basis of the given attribute B. This problem is motivated by the fact that often available historic data is biased due to discrimination, e.g., when B denotes ethnicity. Using the standard learners on this data may lead to wrongfully biased classifiers, even if the attribute B is removed from training data. Existing solutions for this problem consist in “cleaning away” the discrimination from the dataset before a classifier is learned. In this paper we study an alternative approach in which the non-discrimination constraint is pushed deeply into a decision tree learner by changing its splitting criterion and pruning strategy. Experimental evaluation shows that the proposed approach advances the state-of-the-art in the sense that the learned decision trees have a lower discrimination than models provided by previous methods, with little loss in accuracy. Faisal Kamiran, Toon Calders, Mykola Pechenizkiy |
ICDM | 1 |