VLDB 2026 Research / reviewers in the wild / expert
Mudassir Shabbir
dblp:78/7323
· DBLP profile ↗
18ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0002-6961-0961ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Conversations in the wild: Data collection, automatic generation and evaluation
Nimra Zaheer, Agha Ali Raza, Mudassir Shabbir |
Comput. Speech Lang. | 3 |
| 2024 | MSDGSD: A Scalable Graph Descriptor for Processing Large GraphsabstractGraph representation methods have recently become the de facto standard for downstream machine learning tasks on graph-structured data and have found numerous applications, e.g., drug discovery & development, recommendation, and forecasting. However, the existing methods are specially designed to work in a centralized environment, which limits their applicability to small or medium-sized graphs. In this work, we present a graph embedding method that extracts graph representations in a distributed environment with independent and parallel machines. The proposed method is built-upon the existing approach, distributed graph statistical distance (DGSD), to enhance the scalability on large graphs. The key innovation of our work lies in the proposition of a batching mechanism for client-server message passing, which reduces communication overhead during the computation of the distance matrix. In addition, we present a sampling approach for computing pairwise distances between the nodes to compute the desired graph embedding. Moreover, we systematically explore six distinct variations of a distributed graph embeddings and subsequently subject them to comprehensive evaluation. Our extensive evaluations on over 20 graph datasets and ten baseline methods demonstrate improved running time and comparative classification accuracy compared to state-of-the-art embedding techniques. Anwar Said, Iqra Safder, Saeed-Ul Hassan, Naif R. Aljohani, Mudassir Shabbir |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | Network Controllability Perspectives on Graph RepresentationabstractGraph representations in fixed dimensional feature space are vital in applying learning tools and data mining algorithms to perform graph analytics. Such representations must encode the graph's topological and structural information at the local and global scales without posing significant computation overhead. This paper employs a unique approach grounded in networked control system theory to obtain expressive graph representations with desired properties. We consider graphs as networked dynamical systems and study their controllability properties to explore the underlying graph structure. The controllability of a networked dynamical system profoundly depends on the underlying network topology, and we exploit this relationship to design novel graph representations using controllability Gramian and related metrics. We discuss the merits of this new approach in terms of the desired properties (for instance, permutation and scale invariance) of the proposed representations. Our evaluation of various benchmark datasets in the graph classification framework demonstrates that the proposed representations either outperform (sometimes by more than 6 results to the state-of-the-art embeddings. Anwar Said, Obaid Ullah Ahmad, Waseem Abbas 0003, Mudassir Shabbir, Xenofon Koutsoukos |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Enhanced Graph Neural Networks with Ego-Centric Spectral Subgraph Embeddings AugmentationabstractGraph Neural Networks (GNNs) have shown remarkable merit in performing various learning-based tasks in complex networks. The superior performance of GNNs often correlates with the availability and quality of node-level features in the input networks. However, for many network applications, such node-level information may be missing or unreliable, thereby limiting the applicability and efficacy of GNNs. To address this limitation, we present a novel approach denoted as Ego-centric Spectral subGraph Embedding Augmentation (ESGEA), which aims to enhance and design node features, particularly in scenarios where information is lacking. Our method leverages the topological structure of the local subgraph to create topology-aware node features. The subgraph features are generated using an efficient spectral graph embedding technique, and they serve as node features that capture the local topological organization of the network. The explicit node features, if present, are then enhanced with the subgraph embeddings in order to improve the overall performance. ESGEA is compatible with any GNN-based architecture and is effective even in the absence of node features. We evaluate the proposed method in a social network graph classification task where node attributes are unavailable, as well as in a node classification task where node features are corrupted or even absent. The evaluation results on seven datasets and eight baseline models indicate up to a 10% improvement in AUC and a 7% improvement in accuracy for graph and node classification tasks, respectively. Anwar Said, Mudassir Shabbir, Tyler Derr, Waseem Abbas 0003, Xenofon Koutsoukos |
ICMLA | 2 |
| 2023 | NeuroGraph: Benchmarks for Graph Machine Learning in Brain ConnectomicsabstractMachine learning provides a valuable tool for analyzing high-dimensional functional neuroimaging data, and is proving effective in predicting various neurological conditions, psychiatric disorders, and cognitive patterns. In functional magnetic resonance imaging (MRI) research, interactions between brain regions are commonly modeled using graph-based representations. The potency of graph machine learning methods has been established across myriad domains, marking a transformative step in data interpretation and predictive modeling. Yet, despite their promise, the transposition of these techniques to the neuroimaging domain has been challenging due to the expansive number of potential preprocessing pipelines and the large parameter search space for graph-based dataset construction. In this paper, we introduce NeuroGraph, a collection of graph-based neuroimaging datasets, and demonstrated its utility for predicting multiple categories of behavioral and cognitive traits. We delve deeply into the dataset generation search space by crafting 35 datasets that encompass static and dynamic brain connectivity, running in excess of 15 baseline methods for benchmarking. Additionally, we provide generic frameworks for learning on both static and dynamic graphs. Our extensive experiments lead to several key observations. Notably, using correlation vectors as node features, incorporating larger number of regions of interest, and employing sparser graphs lead to improved performance. To foster further advancements in graph-based data driven neuroimaging analysis, we offer a comprehensive open-source Python package that includes the benchmark datasets, baseline implementations, model training, and standard evaluation. Anwar Said, Roza G. Bayrak, Tyler Derr, Mudassir Shabbir, Daniel Moyer, Catie Chang, Xenofon Koutsoukos |
NeurIPS | 4 |
| 2023 | Circuit design completion using graph neural networks
Anwar Said, Mudassir Shabbir, Brian Broll, Waseem Abbas 0003, Péter Völgyesi, Xenofon Koutsoukos |
Neural Comput. Appl. | 2 |
| 2023 | Computing Graph Descriptors on Edge StreamsabstractFeature extraction is an essential task in graph analytics. These feature vectors, called graph descriptors, are used in downstream vector-space-based graph analysis models. This idea has proved fruitful in the past, with spectral-based graph descriptors providing state-of-the-art classification accuracy. However, known algorithms to compute meaningful descriptors do not scale to large graphs since: (1) they require storing the entire graph in memory, and (2) the end-user has no control over the algorithm’s runtime. In this article, we present streaming algorithms to approximately compute three different graph descriptors capturing the essential structure of graphs. Operating on edge streams allows us to avoid storing the entire graph in memory, and controlling the sample size enables us to keep the runtime of our algorithms within desired bounds. We demonstrate the efficacy of the proposed descriptors by analyzing the approximation error and classification accuracy. Our scalable algorithms compute descriptors of graphs with millions of edges within minutes. Moreover, these descriptors yield predictive accuracy comparable to the state-of-the-art methods but can be computed using only 25% as much memory. Zohair Raza Hassan, Sarwan Ali, Mudassir Shabbir, Waseem Abbas 0003 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2022 | Byzantine Resilient Distributed Learning in Multirobot SystemsabstractDistributed machine learning algorithms are increasingly used in multirobot systems and are prone to Byzantine attacks. In this article, we consider a distributed implementation of the stochastic gradient descent (SGD) algorithm in a cooperative network, where networked agents optimize a global loss function using SGD on the local data and aggregation of the estimates of immediate neighbors. Byzantine agents can send arbitrary estimates to their neighbors, which may disrupt the convergence of normal agents to the optimum state. We show that if every normal agent combines its neighbors’ estimates (states) such that the aggregated state is in the convex hull of its normal neighbors’ states, then the resilient convergence is guaranteed. To assure this sufficient condition, we propose a resilient aggregation rule based on the notion ofcenterpoint, which is a generalization of the median in the higher-dimensional Euclidean space. We evaluate our results using examples of target pursuit and pattern recognition in multirobot systems. The evaluation results demonstrate that distributed learning with average, coordinate-wise median, and geometric median-based aggregation rules fail to converge to the optimum state, whereas the centerpoint-based aggregation rule is resilient in the same scenario. Waseem Abbas 0003, Mudassir Shabbir, Xenofon Koutsoukos |
IEEE Trans. Robotics | 3 |
| 2021 | SEMOUR: A Scripted Emotional Speech Repository for UrduabstractDesigning reliable Speech Emotion Recognition systems is a complex task that inevitably requires sufficient data for training purposes. Such extensive datasets are currently available in only a few languages, including English, German, and Italian. In this paper, we present SEMOUR, the first scripted database of emotion-tagged speech in the Urdu language, to design an Urdu Speech Recognition System. Our gender-balanced dataset contains 15,040 unique instances recorded by eight professional actors eliciting a syntactically complex script. The dataset is phonetically balanced, and reliably exhibits a varied set of emotions as marked by the high agreement scores among human raters in experiments. We also provide various baseline speech emotion prediction scores on the database, which could be used for various applications like personalized robot assistants, diagnosis of psychological disorders, and getting feedback from a low-tech-enabled population, etc. On a random test sample, our model correctly predicts an emotion with a state-of-the-art 92% accuracy. Nimra Zaheer, Obaid Ullah Ahmad, Muhammad Shehryar Khan, Mudassir Shabbir |
CHI | 5 |
| 2021 | Seymour's Second Neighborhood Conjecture for 6-antitransitive digraphs
Zohair Raza Hassan, Imran F. Khan, Mehvish I. Poshni, Mudassir Shabbir |
Discret. Appl. Math. | 4 |
| 2021 | DGSD: Distributed graph representation via graph statistical properties
Anwar Said, Saeed-Ul Hassan, Suppawong Tuarob, Raheel Nawaz, Mudassir Shabbir |
Future Gener. Comput. Syst. | 5 |
| 2021 | NetKI: A kirchhoff index based statistical graph embedding in nearly linear time
Anwar Said, Saeed-Ul Hassan, Waseem Abbas 0003, Mudassir Shabbir |
Neurocomputing | 4 |
| 2020 | Estimating Descriptors for Large Graphs
Zohair Raza Hassan, Mudassir Shabbir, Waseem Abbas 0003 |
PAKDD (1) | 2 |
| 2020 | Combinatorial trace method for network immunization
Muhammad Ahmad 0005, Sarwan Ali, Juvaria Tariq, Mudassir Shabbir, Arif Zaman |
Inf. Sci. | 5 |
| 2020 | Interpretable multi-scale graph descriptors via structural compression
Zohair Raza Hassan, Mudassir Shabbir |
Inf. Sci. | 3 |
| 2017 | Efficient Approximation Algorithms for Strings Kernel Based Sequence ClassificationabstractSequence classification algorithms, such as SVM, require a definition of distance (similarity) measure between two sequences. A commonly used notion of similarity is the number of matches between k-mers (k-length subsequences) in the two sequences. Extending this definition, by considering two k-mers to match if their distance is at most m, yields better classification performance. This, however, makes the problem computationally much more complex. Known algorithms to compute this similarity have computational complexity that render them applicable only for small values of k and m. In this work, we develop novel techniques to efficiently and accurately estimate the pairwise similarity score, which enables us to use much larger values of k and m, and get higher predictive accuracy. This opens up a broad avenue of applying this classification approach to audio, images, and text sequences. Our algorithm achieves excellent approximation performance with theoretical guarantees. In the process we solve an open combinatorial problem, which was posed as a major hindrance to the scalability of existing solutions. We give analytical bounds on quality and runtime of our algorithm and report its empirical performance on real world biological and music sequences datasets. Juvaria Tariq, Arif Zaman, Mudassir Shabbir |
NIPS | 4 |
| 2011 | Ray-Shooting Depth: Computing Statistical Data Depth of Point Sets in the Plane
Nabil H. Mustafa, Saurabh Ray, Mudassir Shabbir |
ESA | 3 |
| 2008 | Acceleration of Smith-Waterman using Recursive Variable ExpansionabstractThe Smith-Waterman (SW) algorithm is a local sequence alignment algorithm that attempts to align two biological sequences of varying length such that the alignment score is maximum. In this paper, we propose a new approach to reduce the time needed to perform the SW algorithm. This is done by applying the concept of recursive variable expansion, which exposes more parallelism in the algorithm than any other published method. The paper estimates the speed and hardware overhead for the newly proposed approach relative to other known acceleration methods. Using the new approach, it is possible to achieve a minimum speedup of 400 times better than the serial case for a typical sequence length of 500, which is 1.6 times higher than any other published method. The paper also shows that further speedup can be achieved using extra hardware to expose even more parallelism in the algorithm. Zubair Nawaz, Zaid Al-Ars, Koen Bertels, Mudassir Shabbir |
DSD | 4 |