Muhammad Usman Tariq

dblp:234/2844 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Agentic latency control and cooperative vehicle coordination in 5G-MEC: A prescriptive and explainable AI framework
Sheikh Muhammad Saqib, Tehseen Mazhar, Muhammad Usman Tariq, Tariq Shahzad, Asem Ibrahim Alalwan, Habib Hamam
Comput. Commun.3
2025 REPTILE: Performant Tiling of Recurrences
abstract
We introduce REPTILE, a compiler that performs tiling optimizations for programs expressed as mathematical recurrence equations. REPTILE recursively decomposes a recurrence program into a set of unique tiles and then simplifies each into a different set of recurrences. Given declarative user specifications of recurrence equations, optimizations, and optional mappings of recurrence subexpressions to external libraries calls, REPTILE generates C code that composes compiler-generated loops with calls to external hand-optimized libraries. We show that for direct linear solvers expressible as recurrence equations, the generated C code matches and often exceeds the performance of standard hand-optimized libraries. We evaluate REPTILE’s generated C code against hand-optimized implementations of linear solvers in Intel MKL, as well as two nonsolver recurrences from bioinformatics: Needleman-Wunsch and Smith-Waterman. When the user provides good tiling specifications, REPTILE achieves parity with MKL, achieving between 0.79−1.27x speedup for the LU decomposition, 0.97−1.21x speedup for the Cholesky decomposition, 1.61x−2.72x for lower triangular matrix inversion, 1.01−1.14x speedup for triangular solve with multiple right-hand sides, and 1.14−1.73x speedup over handwritten implementations of the bioinformatics recurrences.
Muhammad Usman Tariq, Shiv Sundram, Fredrik Kjolstad
Proc. ACM Program. Lang.1
2024 Compiling Recurrences over Dense and Sparse Arrays
abstract
We present a framework for compiling recurrence equations into native code. In our framework, users specify a system of recurrences, the types of data structures that store inputs and outputs, and scheduling commands for optimization. Our compiler then lowers these specifications into native code that respects the dependencies in the recurrence equations. Our compiler can generate code over both sparse and dense data structures, and determines if the recurrence system is solvable with the provided scheduling primitives. We evaluate the performance and correctness of the generated code on several recurrences, from domains as diverse as dense and sparse matrix solvers, dynamic programming, graph problems, and sparse tensor algebra. We demonstrate that the generated code has competitive performance to hand-optimized implementations in libraries. However, these handwritten libraries target specific recurrences, specific data structures, and specific optimizations. Our system, on the other hand, automatically generates implementations from recurrences, data formats, and schedules, giving our system more generality than library approaches.
Shiv Sundram, Muhammad Usman Tariq, Fredrik Kjolstad
Proc. ACM Program. Lang.2
2024 Real-Time Fake News Detection Using Big Data Analytics and Deep Neural Network
abstract
In today’s fast-paced world, the Internet has become a prevalent source of information for people worldwide. With the increasing use of various applications, people can get updates in real time, making access to information more convenient than ever. However, this easy access to the Internet has also led to the rise of fake news, making it difficult for individuals to differentiate between true and false information. This is where deep learning (DL) comes into play, offering a solution to identify and combat fake news. In this era of technology, DL can be a game-changer in detecting fake news and preventing potential damage to individuals and organizations. A hybrid N-gram and long short-term memory (LSTM) model improves accuracy, recall rate, and computation time, making the fake news detection process more elegant. This proposed model utilizes a classifier to classify fake news. It is based on the parallel and distributed platform, enabling it to build the DL model using big data analytics. This platform improves the training and testing time and enhances the accuracy of the proposed model. The proposed system classifies news into two categories“, fake news” and “real news”, while quantifying the results to develop a system that can detect fake news with high accuracy and a meager mistake rate. Integrating the deep neural network (DNN) and Spark architecture of big data makes the proposed model highly efficient, as demonstrated by the results.
Muhammad Babar 0001, Awais Ahmad 0001, Muhammad Usman Tariq, Sarah Kaleem
IEEE Trans. Comput. Soc. Syst.3
2023 An Optimized IoT-Enabled Big Data Analytics Architecture for Edge-Cloud Computing
abstract
The awareness of edge computing is attaining eminence and is largely acknowledged with the rise of Internet of Things (IoT). Edge-enabled solutions offer efficient computing and control at the network edge to resolve the scalability and latency-related concerns. Though, it comes to be challenging for edge computing to tackle diverse applications of IoT as they produce massive heterogeneous data. The IoT-enabled frameworks for Big Data analytics face numerous challenges in their existing structural design, for instance, the high volume of data storage and processing, data heterogeneity, and processing time among others. Moreover, the existing proposals lack effective parallel data loading and robust mechanisms for handling communication overhead. To address these challenges, we propose an optimized IoT-enabled big data analytics architecture for edge-cloud computing using machine learning. In the proposed scheme, an edge intelligence module is introduced to process and store the big data efficiently at the edges of the network with the integration of cloud technology. The proposed scheme is composed of two layers: IoT-edge and Cloud-processing. The data injection and storage is carried out with an optimized MapReduce parallel algorithm. Optimized Yet Another Resource Negotiator (YARN) is used for efficiently managing the cluster. The proposed data design is experimentally simulated with an authentic dataset using Apache Spark. The comparative analysis is decorated with existing proposals and traditional mechanisms. The results justify the efficiency of our proposed work.
Muhammad Babar 0001, Mian Ahmad Jan, Xiangjian He, Muhammad Usman Tariq, Spyridon Mastorakis, Ryan Alturki
IEEE Internet Things J.4
2021 Graph Theoretic Approach for the Analysis of Comprehensive Mass-Spectrometry (MS/MS) Data of Dissolved Organic Matter
abstract
Dissolved organic matter (DOM) is a highly complex mixture of organic substances found in aquatic ecosystems. This mixture results from the degradation of primary producers within the ecosystem, groundwater, and the surrounding terrestrial sources. Understanding the chemical structure of DOM is crucial to assessing its impact on aquatic ecosystems. Although multiple studies have addressed the complexity of DOM, the molecular structure of this set of compounds remains unclear. In this work, we present a novel computational framework "Graph-DOM," to assess the comprehensive fragmentation data obtained from the analysis of DOM using the Data Independent Fragmentation strategy with ESI-FT-ICR MS/MS enabling better understanding of the structural complexity of DOM. Graph-DOM uses graph algorithms to dissect a compiled output file obtained from processing hundreds of ultra-high-resolution fragment spectra. Over half a million ordered fragmentation pathways were computed for 764 isolated precursor ions assuming up to seven vector segments categorized as neutral losses (CH2, CH3, O, CH4, H2O, CO, and CO2). Families of structurally related molecules were identified using pathway overlaps, and output files compatible with network visualization software (e.g., Cytoscape) were also generated. Graph-DOM is able to efficiently process all the pathways to discover families within only a few minutes with adjustable parameters for overlap length of fragmentation pathways as well as configuring low abundance CHOS, CHON, and CHONS compounds. Graph-DOM is available at https://github.com/Usman095/Graph-DOM.
Muhammad Usman Tariq, Dennys Leyvay, Francisco Alberto Fernandez Limaz, Fahad Saeed
BIBM1
2021 BP Neural Network Combination Prediction for Big Data Enterprise Energy Management System
Ryan Alturki, Ateeq Ur Rehman 0001, Muhammad Usman Tariq
Mob. Networks Appl.4
2018 Parallel Sampling-Pipeline for Indefinite Stream of Heterogeneous Graphs using OpenCL for FPGAs
abstract
In the field of data science, a huge amount of data, generally represented as graphs, needs to be processed and analyzed. It is of utmost importance that this data be processed swiftly and efficiently to save time and energy. The volume and velocity of data, along with irregular access patterns in graph data structures, pose challenges in terms of analysis and processing. Further, a big chunk of time and energy is spent on analyzing these graphs on large compute clusters and/or data-centers. Filtering and refining of data using graph sampling techniques are one of the most effective ways to speed up the analysis. Efficient accelerators, such as FPGAs, have proven to significantly lower the energy cost of running an algorithm. To this end, we present the design and implementation of a parallel graph sampling technique, for a large number of input graphs streaming into a FPGA. A parallel approach using OpenCL for FPGAs was adopted to come up with a solution that is both time- and energy-efficient. We introduce a novel graph data structure, suitable for streaming graphs on FPGAs, that allows time- and memory-efficient representation of graphs. Our experiments show that our proposed technique is 3x faster and 2x more energy efficient as compared to serial CPU version of the algorithm.
Muhammad Usman Tariq, Fahad Saeed
IEEE BigData1