EDBT 2026 Demo / reviewers in the wild / expert
Tarique Anwar
dblp:51/10773
· DBLP profile ↗
28ranked-venue papers
12as first author
10since 2021 · last 2026
0000-0001-7157-0236ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 16 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Security and privacy · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalized Local Prominence for Source Detection in Real-World Rumor Networks (Extended Abstract)
Syed Shafat Ali, Ajay Rastogi, Tarique Anwar, Syed Afzal Murtaza Rizvi, Jian Yang 0001, Jia Wu 0001, Quan Z. Sheng |
ICDE | 3 |
| 2026 | Lexicon-Based Graph Attention for Severity Estimation of Eating Disorders on Social MediaabstractEating disorders (EDs) are one of the major mental health disorders in our society today. Recent works have found that people experiencing an ED often use social media platforms to connect and interact with other like-minded people. These social media data provide a rich source to understand and analyse ED-related issues experienced and discussed by the users. Within this context, a major unresolved research problem is whether ED severity can be predicted from social media posts. To this end, we conduct a thorough study on Twitter (currently$\mathbb {X}$), starting from collecting a large ED-related dataset, manually annotating them, conducting an analysis to investigate the relationship between affective features and ED severity, and developing an advanced deep learning model calledSEEDNet. This model leverages affective computing by embedding emotional and psychological signals such as valence, arousal, and sentiment into both the text representation and lexicon-based graph structure, enabling severity estimation through affect-aware learning. It is used to generate the final output of ED severity through ordinal multiclass classification. In our experiments, we compare our results with several state-of-the-art methods and significantly outperform them (by over 3% in terms of average$F_{1}$-score and accuracy). Beyond this performance gain,SEEDNetoffers additional advantages, including a novel lexicon-guided affective graph and an ordinal-aware loss function. These components not only enhance the model's interpretability but also ensure that predictions better capture the graded nature of ED severity. These findings show the importance of integrating affective computing and deep learning for analysing social media data to improve the assessment of ED severity. This approach offers new opportunities for developing more nuanced and targeted interventions. Mohammad Abuhassan, Tarique Anwar, Chengfei Liu, Hannah K. Jarman, Siân A. McLean, Susan J. Paxton, Rachel F. Rodgers, Matthew Fuller-Tyszkiewicz |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Information Disorder Amidst Crisis: A Case Study of COVID-19 in IndiaabstractThe devastation led by the COVID-19 pandemic was accompanied by a plethora of misinformation, laden with pseudoscience, hoaxes, and myths, often intertwined with hate speech. This phenomenon was particularly pronounced in India, where the intricate political and communal landscape provided fertile ground. The misinformation, with its elements of hate speech, posed a significant threat to societal cohesion. In response, this article delves into the dynamics of misinformation during the COVID-19 crisis in India, with a specific focus on differentiating general misinformation (GM) from hateful misinformation (HM). To this end, we construct an Indian COVID-19 misinformation dataset collected from various online social and mainstream media and analyze it from various perspectives. Mainly, we focus on temporal evolution, content and topics involved, and emotions and sentiment sensationalism of COVID-19 misinformation. We found the emotions of sadness and fear as key amplifiers of misinformation in general, with negative sentiments dominating HM. Through our comprehensive analysis, we found many such interesting insights and patterns. We also perform hate detection within misinformation content using various unsupervised and supervised learning techniques. Our results show that while GM is relatively easier to identify, it is challenging to detect HM. Overall, deep learning models are found to be more effective than unsupervised methods. By discovering key insights and patterns, this study serves as a foundation for developing robust strategies to combat information disorder. Mohammad Affan, Syed Shafat Ali, Tarique Anwar, Ajay Rastogi |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Generalized Local Prominence for Source Detection in Real-World Rumor NetworksabstractThe problem of infection source detection deals with localizing the infection source in a given network. While the problem has been extensively studied in the past, researchers have mainly focused on simulated infection networks which may not be the correct reflection of the dynamics of real-world infections. More significantly, the existing methods assume that a rumor source lies at the center of an infection network (source-centrality), which is not always true in sparse real-world rumor networks. Due to the randomness of infection flow in such networks, the source may lie away from the center (source-skewness). There is also a lack of real-world infection network datasets to provide a true real-world perspective. Therefore, we revisit the source detection problem and contemplate a shift from mainstream simulations to a real-world paradigm. To this end, we generate two novel rumor network datasets, Cov19-RN and Use20-RN, based on COVID-19 and US Elections 2020 misinformation trends on Twitter (currently$\mathbb {X}$). Besides, inspired by the technicalities inherent to real-world rumor networks, we propose a real-world oriented algorithm called Generalized Exoneration and Prominence based Age, GEPA, for rumor source detection. GEPA addresses the problem of source-skewness to detect rumor sources using the concept of generalized local prominence, which we introduce in this study. Our experiments show that GEPA significantly outperforms the state-of-the-art methods, producing detection rates of 73.6% against 61.5% of the closest competing method on Cov19-RN, and 61.5% against 52.6% of the closest competing method on Use20-RN. To the best of our knowledge, this study is the first such work to deal with source detection in real-world rumor networks and address the problem of source-skewness. Our complete source code, benchmark datasets and detailed results are available athttps://github.com/tesla121/GEPA. Syed Shafat Ali, Ajay Rastogi, Tarique Anwar, Syed Afzal Murtaza Rizvi, Jian Yang 0001, Jia Wu 0001, Quan Z. Sheng |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | FROST: Controlled Label Propagation for Multisource DetectionabstractWe often see rumors rapidly spreading in online social networks. These are harmful for our society in many ways. Infection source detection is the task of identifying the sources of rumors or any other such infections in social networks, so that appropriate intervention could be performed to control the harm. Researchers have studied this problem under various scenarios, where multisource detection has been of special importance. In this article, we propose a novel infection rate controlled label propagation method for multisource detection calledFROST. It leverages the connection strengths between a pair of nodes in the form of infection rate to capture the implicit information latent within an infection. Initially, labels are assigned to nodes indicating whether the nodes are infected or not. Afterward, the labels are propagated across the network in a controlled manner based on the infection rate. Once the propagation converges, the locally prominent nodes are considered as sources. We compareFROSTagainst six state-of-the-art methods and two heuristic baselines in terms of ten evaluation measures over four social networks datasets. Our results show thatFROSTgenerally outperforms the competing methods across various evaluation measures and datasets. It also estimates the number of sources closer to the actual than the competing methods.FROSTscales effectively for large infections, including when there are infection overlaps, where the competing methods generally lag. Syed Shafat Ali, Ajay Rastogi, Tarique Anwar |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | BiCapsHate: Attention to the Linguistic Context of Hate via Bidirectional Capsules and HatebaseabstractOnline social media (OSM) communications sometimes turn into hate-filled and offensive comments or arguments. It not just disrupts the social fabric online, but also leads to hate, violence, and crime, in the real physical world in worst scenarios. The existing content moderation practices of OSM platforms often fail to control the online hate. In this article, we develop a deep learning model calledBiCapsHateto detect hate speech (HS) in OSM posts. The model consists of five layers of deep neural networks. It starts with an input layer to process the input text and follows on to an embedding layer to embed the text into a numeric representation. A BiCaps layer then learns the sequential and linguistic contextual representations, a dense layer prepares the model for final classification, and lastly the output layer produces the resulting class as either hate or non-HS (NHS). The BiCaps layer, being the most important component, effectively learns the contextual information with respect to different orientations in both forward and backward directions of the input text via capsule networks. It is further aided by our rich set of hand-crafted shallow and deep auxiliary features including theHatebaselexicon, making the model well-informed. We conduct extensive experiments on five benchmark datasets to demonstrate the efficacy of the proposedBiCapsHatemodel. In the overall results, we outperform the existing state-of-the-art methods including fBERT, HateBERT, and ToxicBERT.BiCapsHateachieves up to 94% and 92% f-score on balanced and imbalanced datasets, respectively. Our complete source code is publicly available at GitHub repositoryhttps://github.com/Ashraf-Kamal/BiCapsHate. Ashraf Kamal, Tarique Anwar, Vineet Kumar Sejwal, Mohd Fazil |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | EDNet: Attention-Based Multimodal Representation for Classification of Twitter Users Related to Eating DisordersabstractSocial media platforms provide rich data sources in several domains. In mental health, individuals experiencing an Eating Disorder (ED) are often hesitant to seek help through conventional healthcare services. However, many people seek help with diet and body image issues on social media. To better distinguish at-risk users who may need help for an ED from those who are simply commenting on ED in social environments, highly sophisticated approaches are required. Assessment of ED risks in such a situation can be done in various ways, and each has its own strengths and weaknesses. Hence, there is a need for and potential benefit of a more complex multimodal approach. To this end, we collect historical tweets, user biographies, and online behaviours of relevant users from Twitter, and generate a reasonably large labelled benchmark dataset. Thereafter, we develop an advanced multimodal deep learning model called EDNet using these data to identify the different types of users with ED engagement (e.g., potential ED sufferers, healthcare professionals, or communicators) and distinguish them from those not experiencing EDs on Twitter. EDNet consists of five deep neural network layers. With the help of its embedding, representation and behaviour modeling layers, it effectively learns the multimodalities of social media. In our experiments, EDNet consistently outperforms all the baseline techniques by significant margins. It achieves an accuracy of up to 94.32% and F1 score of up to 93.91% F1 score. To the best of our knowledge, this is the first such study to propose a multimodal approach for user-level classification according to their engagement with ED content on social media. Mohammad Abuhassan, Tarique Anwar, Chengfei Liu, Hannah K. Jarman, Matthew Fuller-Tyszkiewicz |
WWW | 2 |
| 2023 | Tracking the Evolution of Clusters in Social Media StreamsabstractTracking the evolution of clusters in social media streams is becoming increasingly important for many applications, such as early detection and monitoring of natural disasters or pandemics. In contrast to clustering on a static set of data, streaming data clustering does not have a global view of the complete data. The local (or partial) view in a high-speed stream makes clustering a challenging task. In this paper, we propose a novel density peak based algorithm,TStream, for tracking the evolution of clusters and outliers in social media streams, via the evolutionary actions of cluster adjustment, emergence, disappearance, split, and merge.TStreamis based on a temporal decay model and text stream summarisation. The decay model captures the decreasing importance of textual documents over time. The stream summarisation compactly represents them with the help of cells (akamicro-clusters) in the memory. We also propose a novel efficient index calledshared dependency tree(akaSD-Tree) based on the ideas of density peak and shared dependency. It maintains the dynamic dependency relationships inTStreamand thereby improves the overall efficiency. We conduct extensive experiments on five real datasets.TStreamoutperforms the existing state-of-the-art solutions based onMStream,MStreamF,EDMStream,OSGM, andEStream, in terms of cluster mapping measure (CMM) by up to 17.8%, 18.6%, 6.9%, 16.4%, and 20.1%, respectively. It is also significantly more efficient thanMStream,MStreamF,OSGM, andEStream, in terms of response time and throughput. Tarique Anwar, Surya Nepal, Cécile Paris, Jian Yang 0001, Jia Wu 0001, Quan Z. Sheng |
IEEE Trans. Big Data | 1 |
| 2022 | EDBase: Generating a Lexicon Base for Eating Disorders Via Social MediaabstractEating disorders (EDs) are characterised by abnormal eating habits and obsessive thought about food, weight, shape, and body image. EDs are experienced by a significant portion of our population. Social media is identified as a possible source of influence for EDs, and there is growing evidence of a large amount of ED-related discussions on the Web via social media platforms, such as Twitter. With this growing trend, automatic content analysis for EDs is becoming increasingly important. To date, there does not exist any comprehensive benchmark ED lexicon to identify ED-related conversations that would, in turn, facilitate these content analysis tasks. In this paper, we propose a novel method for generating a lexicon base for ED language, calledEDBase. The method starts with collecting over 3.7 million ED-focused tweets. In order to semantically represent potential ED terminology in a vector space, an ED word embedding model (EDModel) is trained. Then we develop a novel multi-seeded hierarchical density-based algorithm with contrasting corpora for ED lexicon expansion. TheEDModelis queried by the proposed lexicon expansion algorithm to expand the seed terms to a comprehensive lexicon base. OurEDBaseconsists of a (further expandable) list of 3794 high-quality ED terms, quantified by an ED score, and linked to their parent terms. The proposed method significantly outperforms all existing alternative baseline methods and models by over 25% in terms of precision and 1500 in terms of true positives. This research is expected to be impactful in the health data science and healthcare community. Tarique Anwar, Matthew Fuller-Tyszkiewicz, Hannah K. Jarman, Mohammad Abuhassan, Adrian Shatte, Suku Sukunesan |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Depression Intensity Estimation via Social Media: A Deep Learning ApproachabstractDepression has become a big problem in our society today. It is also a major reason for suicide, especially among teenagers. In the current outbreak of coronavirus disease (COVID-19), the affected countries have recommended social distancing and lockdown measures. Resulting in interpersonal isolation, these measures have raised serious concerns for mental health and depression. Generally, clinical psychologists diagnose depressed people via face-to-face interviews following the clinical depression criteria. However, often patients tend to not consult doctors in their early stages of depression. Nowadays, people are increasingly using social media to express their moods. In this article, we aim to predict depressed users as well as estimate their depression intensity via leveraging social media (Twitter) data, in order to aid in raising an alarm. We model this problem as a supervised learning task. We start with weakly labeling the Twitter data in a self-supervised manner. A rich set of features, including emotional, topical, behavioral, user level, and depression-related$n$-gram features, are extracted to represent each user. Using these features, we train a small long short-term memory (LSTM) network using Swish as an activation function, to predict the depression intensities. We perform extensive experiments to demonstrate the efficacy of our method. We outperform the baseline models for depression intensity estimation by achieving the lowest mean squared error of 1.42 and also outperform the existing state-of-the-art binary classification method by more than 2% of accuracy. We found that the depressed users frequently use negative words such as stress and sad, mostly post during late nights, highly use personal pronouns and sometimes also share personal events. Shreya Ghosh 0001, Tarique Anwar |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2020 | dFDA-VeD: A Dynamic Future Demand Aware Vehicle Dispatching SystemabstractWith the rising demand of smart mobility, ride-hailing service is getting popular in the urban regions. These services maintain a system for serving the incoming trip requests by dispatching available vehicles to the pickup points. As the process should be socially and economically profitable, the task of vehicle dispatching is highly challenging, specially due to the time-varying travel demands and traffic conditions. Due to the uneven distribution of travel demands, many idle vehicles could be generated during the operation in different subareas. Most of the existing works on vehicle dispatching system, designed static relocation centers to relocate idle vehicles. However, as traffic conditions and demand distribution dynamically change over time, the static solution can not fit the evolving situations. In this paper, we propose a dynamic future demand aware vehicle dispatching system. It can dynamically search the relocation centers considering both travel demand and traffic conditions. We evaluate the system on real-world dataset, and compare with the existing state-of-the-art methods in our experiments in terms of several standard evaluation metrics and operation time. Through our experiments, we demonstrate that the proposed system significantly improves the serving ratio and with a very small increase in operation cost. Tarique Anwar, Jian Yang 0001, Jia Wu 0001 |
MobiQuitous | 2 |
| 2019 | EPA: Exoneration and Prominence based Age for Infection Source IdentificationabstractInfection source identification is a well-established problem, having gained a substantial scale of research attention over the years. In this paper, we study the problem by exploiting the idea of the source being the oldest node. For the same, we propose a novel algorithm called Exoneration and Prominence based Age (EPA), which calculates the age of an infected node by considering its prominence in terms of its both infected and non-infected neighbors. These non-infected neighbors hold the key in exonerating an infected node from being the infection source. We also propose a computationally inexpensive variant of EPA, called EPA-LW. Extensive experiments are performed on seven datasets, including 5 real-world and 2 synthetic, of different topologies and varying sizes to demonstrate the effectiveness of the proposed algorithms. We consistently outperform the state-of-the-art single source identification methods in terms of average error distance. To the best of our knowledge, this is the largest scale performance evaluation of the considered problem till date. We also extend EPA to identify multiple sources by developing two new algorithms - one based on K-Means, called EPA_K-Means, and another based on successive identification of sources, called EPA_SSI. Our results show that both EPA_K-Means and EPA_SSI outperform the other multi-source heuristic approaches. Syed Shafat Ali, Tarique Anwar, Ajay Rastogi, Syed Afzal Murtaza Rizvi |
CIKM | 2 |
| 2019 | Forecasting short-term traffic speed based on multiple attributes of adjacent roads
Dongjin Yu, Chengfei Liu, Yiyu Wu, Sai Liao, Tarique Anwar, Wanqing Li 0003, Chengbiao Zhou |
Knowl. Based Syst. | 5 |
| 2018 | Capturing the Spatiotemporal Evolution in Road Traffic NetworksabstractThe urban road networks undergo frequent traffic congestions during the peak hours and around the city center. Capturing the spatiotemporal evolution of the congestion scenario in real-time in an urban-scale can aid in developing smart traffic management systems, and guiding commuters in making informed decision about route choice. The congestion scenario is often represented by a set of distinguishable network partitions that have a homogeneous level of congestion inside them but are heterogeneous to others. Due to the dynamic nature of traffic, these partitions evolve with time in terms of their structure and location. In this paper, we propose a comprehensive framework to capture the evolution by incrementally updating the partitions in an efficient manner using a two-layer approach. The physical layer maintains a set of small-sized road network building blocks in a fine granularity, and performs low-level computations to incrementally update them, whereas the logical layer performs high-level computations in order to serve as an interface to query the physical layer about the congested partitions in a coarse granularity. We also propose an in-memory index calledBinthat compactly stores the historical sets of building blocks in the main memory with no information loss, and facilitates their efficient retrieval. Our experimental results show that the proposed method is much efficient than the existing re-partitioning methods without significant sacrifice in accuracy. The proposedBinconsumes a minimum space with least redundancy at different time stamps. Tarique Anwar, Chengfei Liu, Hai Le Vu 0001, Md. Saiful Islam 0003, Timos K. Sellis |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | A Framework for Clustering and Dynamic Maintenance of XML Documents
Ahmed Al-Shammari, Chengfei Liu, Mehdi Naseriparsa, Quoc Bao Vo, Tarique Anwar, Rui Zhou 0001 |
ADMA | 5 |
| 2017 | Computing Influence of a Product through Uncertain Reverse SkylineabstractUnderstanding the influence of a product is crucially important for making informed business decisions. This paper introduces a new type of skyline queries, called uncertain reverse skyline, for measuring the influence of a probabilistic product in uncertain data settings. More specifically, given a dataset of probabilistic products P and a set of customers C, an uncertain reverse skyline of a probabilistic product q retrieves all customers c ∈ C which include q as one of their preferred products. We present efficient pruning ideas and techniques for processing the uncertain reverse skyline query of a probabilistic product using R-Tree data index. We also present an efficient parallel approach to compute the uncertain reverse skyline and influence score of a probabilistic product. Our approach significantly outperforms the baseline approach derived from the existing literature. The efficiency of our approach is demonstrated by conducting experiments with both real and synthetic datasets. Md. Saiful Islam 0003, Wenny Rahayu, Chengfei Liu, Tarique Anwar, Bela Stantic |
SSDBM | 4 |
| 2017 | Discovering and Tracking Active Online Social Groups
Md Musfique Anwar, Chengfei Liu, Jianxin Li 0001, Tarique Anwar |
WISE (1) | 4 |
| 2017 | Partitioning road networks using density peak graphs: Efficiency vs. accuracy
Tarique Anwar, Chengfei Liu, Hai Le Vu 0001, Christopher Leckie |
Inf. Syst. | 1 |
| 2016 | Q+Tree: An Efficient Quad Tree based Data Indexing for Parallelizing Dynamic and Reverse SkylinesabstractSkyline queries play an important role in multi-criteria decision making applications of many areas. Given a dataset of objects, a skyline query retrieves data objects that are not dominated by any other data object in the dataset. Unlike standard skyline queries where the different aspects of data objects are compared directly, dynamic and reverse skyline queries adhere to the around-by semantics, which is realized by comparing the relative distances of the data objects w.r.t. a given query. Though, there are a number of works on parallelizing the standard skyline queries, only a few works are devoted to the parallel computation of dynamic and reverse skyline queries. This paper presents an efficient quad-tree based data indexing scheme, called Q+Tree, for parallelizing the computations of the dynamic and reverse skyline queries. We compare the performance of Q+Tree with an existing quad-tree based indexing scheme. We also present several optimization heuristics to improve the performance of both of the indexing schemes further. Experimentation with both real and synthetic datasets verifies the efficiency of the proposed indexing scheme and optimization heuristics. Md. Saiful Islam 0003, Chengfei Liu, Wenny Rahayu, Tarique Anwar |
CIKM | 4 |
| 2016 | Tracking the Evolution of Congestion in Dynamic Urban Road NetworksabstractThe congestion scenario on a road network is often represented by a set of differently congested partitions having homogeneous level of congestion inside. Due to the changing traffic, these partitions evolve with time. In this paper, we propose a two-layer method to incrementally update the differently congested partitions from those at the previous time point in an efficient manner, and thus track their evolution. The physical layer performs low-level computations to incrementally update a set of small-sized road network building blocks, and the logical layer provides an interface to query the physical layer about the congested partitions. At each time point, the unstable road segments are identified and moved to their most suitable building blocks. Our experimental results on different datasets show that the proposed method is much efficient than the existing re-partitioning methods without significant sacrifice in accuracy. Tarique Anwar, Chengfei Liu, Hai Le Vu 0001, Md. Saiful Islam 0003 |
CIKM | 1 |
| 2015 | RoadRank: Traffic Diffusion and Influence Estimation in Dynamic Urban Road NetworksabstractWith the rapidly growing population in urban areas, these days the urban road networks are expanding at a faster rate. The frequent movement of people on them leads to traffic congestions. These congestions originate from some crowded road segments, and diffuse towards other parts of the urban road networks creating further congestions. This behavior of road networks motivates the need to understand the influence of individual road segments on others in terms of congestion. In this work, we propose RoadRank, an algorithm to compute the influence scores of each road segment in an urban road network, and rank them based on their overall influence. It is an incremental algorithm that keeps on updating the influence scores with time, by feeding with the latest traffic data at each time point. The method starts with constructing a directed graph called influence graph, which is then used to iteratively compute the influence scores using probabilistic diffusion theory. We show promising preliminary experimental results on real SCATS traffic data of Melbourne. Tarique Anwar, Chengfei Liu, Hai Le Vu 0001, Md. Saiful Islam 0003 |
CIKM | 1 |
| 2015 | Ranking Radically Influential Web Forum UsersabstractThe growing popularity of online social media is leading to its widespread use among the online community for various purposes. In the recent past, it has been found that the web is also being used as a tool by radical or extremist groups and users to practice several kinds of mischievous acts with concealed agendas and promote ideologies in a sophisticated manner. Some of the web forums are predominantly being used for open discussions on critical issues influenced by radical thoughts. The influential users dominate and influence the newly joined innocent users through their radical thoughts. This paper presents an application of collocation theory to identify radically influential users in web forums. The radicalness of a user is captured by a measure based on the degree of match of the commented posts with a threat list. Eleven different collocation metrics are formulated to identify the association among users, and they are finally embedded in a customized PageRank algorithm to generate a ranked list of radically influential users. The experiments are conducted on a standard data set provided for a challenge at ISI-KDD’12 workshop to find radical and infectious threads, members, postings, ideas, and ideologies. Experimental results show that our proposed method outperforms the existing UserRank algorithm. We also found that the collocation theory is more effective to deal with such ranking problem than the textual and temporal similarity-based measures studied earlier. Tarique Anwar, Muhammad Abulaish |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | Spatial Partitioning of Large Urban Road NetworksabstractThe rapid global migration of people towards urban areas is multiplying the traffic volume on urban road networks. As a result these networks are rapidly growing in size, in which different sub-networks exhibit distinctive traffic flow patterns. In this paper, we propose a scalable framework for traffic congestion-based spatial partitioning of large urban road networks. It aims to identify different sub-networks or partitions that exhibit homogeneous traffic congestion patterns internally, but heterogenous to others externally. To this end, we develop a two-stage procedure within our framework that first transforms the large road graph into a well-structured and condensed supergraph via clustering and link aggregation based on traffic density and adjacency connectivity, respectively. We then devise a spectral theory based novel graph cut (referred as 훼-Cut) to partition the supergraph and compare its performance with that of an ex-isting method for partitioning urban networks. Our results show that the proposed method outperforms the normalized cut based existing method in all the performance evaluation metrics for small road networks and provides good results for much larger networks where other methods may face serious problems of time and space complexities. Tarique Anwar, Chengfei Liu, Hai Le Vu 0001, Christopher Leckie |
EDBT | 1 |
| 2014 | Namesake alias mining on the Web and its role towards suspect tracking
Tarique Anwar, Muhammad Abulaish |
Inf. Sci. | 1 |
| 2012 | Identifying cliques in dark web forums - An agglomerative clustering approachabstractIn this paper, we present a novel agglomerative clustering method to identify cliques in dark Web forums. Considering each post as an individual entity accompanying all the information about its thread, author, time-stamp, etc., we have defined a similarity function to identify similarity between each pair of posts as a blend of their contextual and temporal coherence. The similarity function is employed in the proposed clustering algorithm to group threads into different clusters that are finally presented as individual cliques. The identified cliques are characterized using the homogeneity of posts therein, which also establishes the homogeneity of their authors and threads as well. Tarique Anwar, Muhammad Abulaish |
ISI | 1 |
| 2012 | An MCL-Based Text Mining Approach for Namesake Disambiguation on the WebabstractIn this paper, we propose a Markov Clustering (MCL) based text mining approach for namesake disambiguation on the Web. The novelty of the proposed technique lies in modeling the collection of web pages using a weighted graph structure and applying MCL to crystalize it into different clusters, each one containing the web pages related to a particular namesake individual. The proposed method focuses on three broad and realistic aspects to cluster web pages retrieved through search engines - content overlapping, structure overlapping, and local context overlapping. The efficacy of the proposed method is demonstrated through experimental evaluations on standard datasets. Tarique Anwar, Muhammad Abulaish |
Web Intelligence | 1 |
| 2011 | A web content mining approach for tag cloud generationabstractTag cloud, also known as word cloud, are very useful for quickly perceiving the most prominent terms embedded within a text collection to determine their relative prominence. The effectiveness of tag clouds to conceptualize a text corpus is directly proportional to the quality of the keyphrases extracted from the corpus. Although, authors provide a list of about five to ten keywords in scientific publications that are used to map them into their respective domain, due to exponential growth in non-scientific documents on the World Wide Web, an automatic mechanism is sought to identify keyphrases embedded within them for tag cloud generation. In this paper, we propose a web content mining technique to extract keyphrases from web documents for tag cloud generation. Instead of using partial or full parsing, the proposed method applies n-gram technique followed by various heuristics-based refinements to identify a set of lexical and semantic features from text documents. We propose a rich set of domain-independent features to model candidate keyphrases very effectively for establishing their keyphraseness using classification models. We also propose a font-determination function to determine the relative font-size of keyphrases for tag cloud generation. The efficacy of the proposed method is established through experimentation. The proposed method outperforms the popular keyphrase extraction system KEA. Muhammad Abulaish, Tarique Anwar |
iiWAS | 2 |
| 2011 | Web content mining for alias identification: A first step towards suspect trackingabstractIn this paper, we present the design of a web content mining system to identify and extract aliases of a given entity from the Web in an automatic way. Starting with a pattern-based information extraction process, the system applies n-gram technique to extract candidate aliases. Thereafter, various statistical measures are applied to identify feasible aliases from them. The extracted aliases can be used to generate profiles of suspects and keep track of their movements on the Web using different identities. Tarique Anwar, Muhammad Abulaish, Khaled Alghathbar |
ISI | 1 |