Faraz Ahmed

dblp:34/7562 · DBLP profile ↗
← Back
23ranked-venue papers
12as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 6 first-author · 1 since 2021Systems, architecture and hardware · 6 · 3 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Scaling Attention Beyond GPUs for LLM Inference
abstract
Scaling inference for large language models is increasingly constrained by limited GPU memory, primarily due to the expanding intermediate states (KV caches) required for long-context generation and multi-user workloads. Once the KV cache exceeds the capacity of high-bandwidth memory, it must be offloaded to host memory and reloaded on demand, a workflow severely bottlenecked by the CPU–GPU interconnect, typically PCIe. Existing approaches exploiting offload KV caches to CPU memory and selectively reload partial segments for attention computation often underutilize CPU compute resources and suffer from accuracy degradation. We present Beyond, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference. Beyond executes dense attention over recent KV entries stored in GPU memory while performing parallel, per-head sparse attention on salient contextual KV entries residing in CPU memory. The outputs are fused efficiently through a log-sum-exp scheme. During the bandwidth-constrained decoding phase, oversized KV caches are processed cooperatively by the aggregated CPU and GPU memory bandwidth, with only minimal PCIe data movement. Experiments across diverse models and workloads demonstrate that Beyond improves scalability, supports longer sequences and larger batch sizes, and outperforms existing sparse attention baselines in both efficiency and accuracy—all on commodity GPU hardware.
Weishu Deng, Peiran Du, Lingfeng Xiang, Chen Zhong 0002, Faraz Ahmed, Lianjie Cao, Puneet Sharma 0001, Song Jiang 0001, Hui Lu 0001, Jia Rao
HPDC7
2026 Advanced explainable ensemble models for multi-class intrusion detection in heterogeneous drone and industrial networks
abstract
The popularity of drones, unmanned aerial vehicles (UAVs), and industrial-level cyber-physical systems has deepened external threats caused by network-based cyber-attacks. Such environments have a problem of dynamic traffic behavior, temporal dependencies, class imbalance and the existence of various type of attacks, such as denial of service, injection, replay attacks, scanning, and man in the middle attacks. This paper presents an effective and justifiable multi-class attack detection model in a heterogeneous environment that can be used as an intrusion detector. Three benchmark datasets, Drone IDS, UAVIDS-2025, and ICSCASD_MPLC were evaluated comprehensively with ensemble-based machine learning models (Random Forest, Extra trees, AdaBoost, XGBoost, and CatBoost) and those with deep learning architecture (ANN, CNN, RNN, LSTM, and ResNet). Within the framework of many preprocessing steps, the accuracy, macro-averaged precision, recall, F1-score, Matthews Correlation Coefficient, Cohen’s Kappa, log loss, and ROC-AUC were used to evaluate the models. According to experimental findings, Random Forest is more effective than other ensemble models, with macro F1-scores of 0.99964, 0.99844, and 0.99994 on Drone IDS, UAVIDS-2025, and ICSCASD_MPLC datasets, respectively, with nearly perfect ROC-AUC indicators. Compared to other deep learning methods, LSTM is best at learning patterns of attack over time, ANN is well-performing with minimal computing costs, and RNN is well-performing in generalizing on industrial traffic. The validity of statistical significance of results is tested with Friedman and Wilcoxon signed-rank tests with Holm correction, bootstrap confidence intervals, and McNemar test. Also, explainable tools of AI, including SHAP and LIME, provide both local and global explanations, which are both intuitive, as well as ablation testing demonstrates that a small set of flow-based and temporal features are capable of sustaining close-optimal performance. In general, the framework proposed provides real-time and safety–critical deployments with intrusion detection algorithms that are accurate, interpretable, and validated statistically.
Faraz Ahmed, Waqas Ishtiaq, Md. Samiul Islam, Muhammad Masud Tarek
J. Inf. Secur.2
2025 Can Hardware Outsmart Software in Tiered Memory Management? A CMM-H Case Study
abstract
With the advent of Compute Express Link (CXL), hardware-managed memory tiering has become a reality. In this paper, we investigate Samsung's CXL Memory Module-Hybrid (CMM-H), a CXL Type 3 device integrating DRAM and NAND flash managed by an FPGA-based controller and providing byte-addressable memory interface via the cxl.mem protocol. We perform a detailed evaluation of CMM-H and compare its performance with OS-level and block-level tiering solutions. Our results highlight the performance benefits of CMM-H for cache-hit scenarios and identify key limitations for cache-miss situations, offering insights into the trade-offs involved in adopting hardware-managed memory tiering in emerging CXL-based systems.
Lingfeng Xiang, Lianjie Cao, Faraz Ahmed, Jia Rao, Hui Lu 0001, Puneet Sharma 0001
SYSTOR5
2024 Accelerating Containerized Machine Learning Workloads
abstract
To facilitate various Machine Learning (ML) training and inference tasks, enterprises tend to build large and expensive clusters and share them among different teams for diverse ML workloads. Virtualized platforms (containers/VMs) and schedulers are typically deployed to allow such access, manage heterogeneous resources and schedule ML jobs in these clusters. However, allocating resource budgets for different ML jobs to achieve best performance and cluster resource efficiency remains a significant challenge. This work proposes Nearchus to accelerate distributed ML training while ensuring high resource efficiency by using adaptive resource allocation. Nearchus automatically identifies potential performance bottlenecks for running jobs and re-allocates resources to provide optimized run-time performance with high resource efficiency. Nearchus’s resource configuration significantly improves the training speed of individual jobs up to 71.4%–129.1% against state-of-the-art resource schedulers, and reduces job completion and queuing time by 35.6% and 67.8%, respectively.
Ali Tariq, Lianjie Cao, Faraz Ahmed, Eric Rozner, Puneet Sharma 0001
NOMS3
2023 When Caching Systems Meet Emerging Storage Devices: A Case Study
abstract
Block-layer caching systems improve the I/O performance by using hybrid storage devices; the advent of fast, byte-addressable storage enables caching systems to further leverage new storage tiers (e.g., with persistent memory as the cache device and SSD as the backend device) to achieve better caching performance. However, the new storage devices also challenge the design and implementation of existing block-based caching systems. This paper conducts a comprehensive performance study of a popular caching system, Open CAS, and identifies new, unrevealed software bottlenecks. Our observations and root cause analysis cast light on optimizing the software stack of caching systems to incorporate emerging storage technologies.
Lianjie Cao, Faraz Ahmed, Hui Lu 0001, Puneet Sharma 0001
HotStorage3
2022 UnifiedNetManagement: Unified Network Management for Heterogeneous Edge Enterprise Network
abstract
The data explosion over the past few years has made it necessary to connect edge networks to the more powerful cloud infrastructure. Given the heterogeneity, dis-aggregation, and varied functionalities of the edge, it makes it more challenging to monitor and manage the edge using a single control plane. This is especially relevant to Edge Enterprise Network, which spans across Edge Access Networks, Wireless Access Networks, Sensor networks etc. This data explosion at the edge has also necessitated the edge Enterprise Network to support varied functionality, causing it to comprise of multiple modules to provide best performance. This renders the edge-Network cumbersome to manage and very susceptible to faults. Consequently, the highly increasing need to embrace cloud-native technology for Edge computation has become significant for service and application providers to deploy high computation functionalities closer to end users. This motivates us to build an intelligent framework to maintain and manage Edge Enterprise Networks. In this paper, we present “UnifiedNetManagement”, our Hybrid Cloud solution, that integrates any edge network to the cloud using an event-driven Workflow Manager to provide monitoring, scheduling, and troubleshooting capabilities for a smart edge network management system. We observe close to 2x performance improvement over traditional disaggregated network management.
Chinlin Chen, Uyen Chau, Anu Mercian, Faraz Ahmed
IC2E4
2021 Epinoia: Intent Checker for Stateful Networks
abstract
Intent-Based Networking (IBN) has been increasingly deployed in production enterprise networks. Automated network configuration in IBN lets operators focus on intents- i.e., the end to end business objectives-rather than spelling out details of the configurations that implement these objectives. Automation brings its own concerns as the administrators cannot rely on traditional network troubleshooting tools. This situation is further exacerbated in the case of stateful Network Functions (NFs) whose packet processing behavior depends on previously observed traffic patterns. To ensure that the network configuration and state derived from network automation matches the administrator’s specified intent, we propose, Epinoia, a network intent checker for stateful networks. Epinoia relies on a unified model for NFs by leveraging the causal precedence relationships that exist between NF packet I/Os and states. Scalability of Epinoia is achieved by decomposing intents into sub-checking tasks and maintaining a causality graph between checked invariants. Epinoia checks for network-wide intent violations incrementally to reduce overhead in the event of network changes. Our evaluation results using real-world network topologies show that Epinoia can perform comprehensive checking within a few seconds per network with intent updates.
Huazhe Wang, Puneet Sharma 0001, Faraz Ahmed, Joon-Myung Kang, Chen Qian 0001, Mihalis Yannakakis
ICCCN3
2021 Mind the Semantic Gap: Policy Intent Inference from Network Metadata
abstract
Network Policy management is a tedious and laborious task because of scale and dynamic changes in the network. The advent of Softwarized Networks has led to a renewed interest in intent-based network policy management. Intent-based Networking provides a structured way of specifying the intent of policies which are automatically translated and compiled into network device configuration. While this top-down approach of policy intent to policy configuration has worked well for cloud-native infrastructures such as data centers, it has not seen much adoption in legacy networks. We believe one of the primary reasons for this is the semantic gap between policy intents and policy configurations. The problem is further exacerbated by the heterogeneity, scale-on-the-fly, fragmentation, and lack of structure in non-intent native networks. We introduce Policy Intent Inference (PII) System to bridge the semantic gap with its advanced inference layer that extracts the policy intents from policy configurations fragmented over disparate network devices. We adopt a bottom-up approach to extract all policies within network devices, abstract them into a structured data model, and with the use of clustering and information retrieval techniques, build an optimal solution to extract network-wide policy intents from the underlying network that eases policy management especially policy troubleshooting, reducing the configuration clutter and reducing the time taken to compile and resolve conflicts in policies.
Anu Mercian, Faraz Ahmed, Shaun Wackerly, Charles Clark
NetSoft2
2020 Homa: An Efficient Topology and Route Management Approach in SD-WAN Overlays
abstract
This paper presents an efficient topology and route management approach in Software-Defined Wide Area Networks (SD-WAN). Traditional WANs suffer from low utilization and lack of global view of the network. Therefore, during failures, topology/service/traffic changes, or new policy requirements, the system does not always converge to the global optimal state. Using Software Defined Networking architectures in WANs provides the opportunity to design WANs with higher fault tolerance, scalability, and manageability. We exploit the correlation matrix derived from monitoring system between the virtual links to infer the underlying route topology and propose a route update approach that minimizes the total route update cost on all flows. We formulate the problem as an integer linear programming optimization problem and provide a centralized control approach that minimizes the total cost while satisfying the quality of service (QoS) on all flows. Experimental results on real network topologies demonstrate the effectiveness of the proposed approach in terms of disruption cost and average disrupted flows.
Diman Zad Tootaghaj, Faraz Ahmed, Puneet Sharma 0001, Mihalis Yannakakis
INFOCOM2
2018 Optimizing Internet Transit Routing for Content Delivery Networks
abstract
Content delivery networks (CDNs) maintain multiple transit routes from content distribution servers to eyeball ISP networks which provide Internet connectivity to end users. Due to the dynamics of varying performance and pricing on transit routes, CDNs need to implement a transit route selection strategy to optimize performance and cost tradeoffs. In this paper, we formalize the transit routing problem using a multi-attribute objective function to simultaneously optimize end-to-end performance and cost. Our approach allows CDNs to navigate the cost and performance tradeoff in transit routing through a single control knob. We evaluate our approach using real-world measurements from CDN servers located at 19 geographically distributed Internet exchange points. Using our approach, CDNs can reduce transit costs on average by 57% without sacrificing performance.
Faraz Ahmed, Zubair Shafiq, Amir R. Khakpour, Alex X. Liu
IEEE/ACM Trans. Netw.1
2018 Noise Tolerant Localization for Sensor Networks
Fu Xiao 0001, Lei Chen 0011, Chaoheng Sha, Ruchuan Wang 0001, Alex X. Liu, Faraz Ahmed
IEEE/ACM Trans. Netw.7
2017 Monitoring quality-of-experience for operational cellular networks using machine-to-machine traffic
abstract
It is crucial for cellular data network operators to understand the service quality perceived by its customers. The state-of-art systems deployed in cellular networks mostly report service quality aggregated on cell site level, which is typically an aggregation of tens or hundreds of customers depending on the locations of the cell sites. In this paper, we propose to enhance the measurement of customer-perceived service quality by leveraging M2M devices as sensors in the field, which provide an unprecedented opportunity for cellular network operators to measure what end-users experience with better accuracy and coverage. Our approach is to identify a set of M2M devices which are stationary and communicate continuously over the cellular network over an indefinite period of time. We use these M2M devices to estimate the customer-perceived service quality during cell site outages. We implement our methodology as a system called M2MScan and evaluate M2MScan with both synthetic outages and real outages from a large-scale operational cellular network. To the best of our knowledge, this is the first work that employs M2M devices to measure the service quality perceived by customers in operational cellular networks at a large scale.
Faraz Ahmed, Jeffrey Erman, Zihui Ge, Alex X. Liu, Jia Wang 0001
INFOCOM1
2017 Detecting and Localizing End-to-End Performance Degradation for Cellular Data Services Based on TCP Loss Ratio and Round Trip Time
abstract
Providing high end-to-end (E2E) performance experienced by users is critical for cellular service providers to best serve their customers. This paper focuses on the detection and localization of E2E performance degradation (such as slow webpage page loading and unsmooth video playing) at cellular service providers. Detecting and localizing E2E performance degradation is crucial for cellular service providers, content providers, device manufactures, and application developers to jointly troubleshoot root causes. To the best of our knowledge, the detection and localization of E2E performance degradation at cellular service providers has not been previously studied. In this paper, we propose a holistic approach to detecting and localizing E2E performance degradation at cellular service providers across the four dimensions of user locations, content providers, device types, and application types. Our approach consists of three steps: modeling, detection, and localization. First, we use training data to build models that can capture the normal performance of every E2E instance, which means the flows corresponding to a specific location, content provider, device type, and application type. Second, we use our models to detect performance degradation for each E2E instance on an hourly basis. Third, after each E2E instance has been labeled as non-degrading or degrading, we use association rule mining techniques to localize the source of performance degradation. Our system detected performance degradation instances over a period of one week. In 80% of the detected degraded instances, content providers, device types, and application types were the only factors of performance degradation.
Faraz Ahmed, Jeffrey Erman, Zihui Ge, Alex X. Liu, Jia Wang 0001
IEEE/ACM Trans. Netw.1
2016 Social Graph Publishing with Privacy Guarantees
abstract
Online social network graphs provide useful insights on various social phenomena such as information dissemination and epidemiology. Unfortunately, social network providers often refuse to publish their social network graphs due to privacy concerns. Recently, differential privacy has become the widely accepted criteria for privacy preserving data publishing because it provides strongest privacy guarantees for publishing sensitive datasets. Although some work has been done on publishing matrices with differential privacy, they are computationally unpractical as they are not designed to handle large matrices such as the adjacency matrices of OSN graphs. In this paper, we propose a random matrix approach to OSN graph publishing, which achieves storage and computational efficiency by reducing the dimensions of adjacency matrices and achieves differential privacy by adding a small amount of noise. Our key idea is to first project each row of an adjacency matrix into a low dimensional space using random projection, and then perturb the projected matrix with random noise, and finally publish the perturbed and projected matrix. In this paper, we first prove that random projection plus random perturbation preserve differential privacy, and also that the random noise required to achieve differential privacy is small. We then validate the proposed approach and evaluate the utility of the published data for two different applications, namely node clustering and node ranking, using publicly available OSN graphs of Facebook, Live Journal, and Pokec.
Faraz Ahmed, Alex X. Liu, Rong Jin 0001
ICDCS1
2016 The Internet is for Porn: Measurement and Analysis of Online Adult Traffic
abstract
Adult (or pornographic) websites attract a large number of visitors and account for a substantial fraction of the global Internet traffic. However, little is known about the makeup and characteristics of online adult traffic. In this paper, we present the first large-scale measurement study of online adult traffic using HTTP logs collected from a major commercial content delivery network. Our data set contains approximately 323 terabytes worth of traffic from 80 million users, and includes traffic from several dozen major adult websites and their users in four different continents. We analyze several characteristics of online adult traffic including content and traffic composition, device type composition, temporal dynamics, content popularity, content injection, and user engagement. Our analysis reveals several unique characteristics of online adult traffic. We also analyze implications of our findings on adult content delivery. Our findings suggest several content delivery and cache performance optimizations for adult traffic, e.g., modifications to website design, content delivery, cache placement strategies, and cache storage configurations.
Faraz Ahmed, Zubair Shafiq, Alex X. Liu
ICDCS1
2016 Optimizing Internet transit routing for content delivery networks
abstract
Content Distribution Networks (CDNs) maintain multiple transit routes from content distribution servers to eyeball ISP networks which provide Internet connectivity to end users. Due to the dynamics of varying performance and pricing on transit routes, CDNs need to implement a transit route selection strategy to optimize performance and cost tradeoffs. In this paper, we formalize the transit routing problem using a multi-attribute objective function to simultaneously optimize end-to-end performance and cost. Our approach allows CDNs to navigate the cost and performance tradeoff in transit routing through a single control knob. We evaluate our approach using real-world measurements from CDN servers located at 19 geographically distributed IXPs. Using our approach, CDNs can reduce transit costs on average by 57% without sacrificing performance.
Faraz Ahmed, Zubair Shafiq, Amir R. Khakpour, Alex X. Liu
ICNP1
2016 Detecting and localizing end-to-end performance degradation for cellular data services
abstract
Providing high end-to-end (E2E) performance is critical for cellular service providers to best serve their customers. Detecting and localizing E2E performance degradation is crucial for cellular service providers, content providers, device manufactures, and application developers to jointly troubleshoot root causes. To the best of our knowledge, detection and localization of E2E performance degradation at cellular service providers has not been previously studied. In this paper, we propose a holistic approach to detecting and localizing E2E performance degradation at cellular service providers across the four dimensions of user locations, content providers, device types, and application types. First, we use training data to build models that can capture the normal performance of every E2E-instance, which means flows corresponding to a specific location, content provider, device type, and application type. Second, we use our models to detect performance degradation for each E2E-instance on an hourly basis. Third, after each E2E-instance has been labeled as non-degrading or degrading, we use association rule mining techniques to localize the source of performance degradation. Our system detected performance degradation instances over a period of one week. In 80% of the detected degraded instances, content providers, device types, and application types were the only factors of performance degradation.
Faraz Ahmed, Jeffrey Erman, Zihui Ge, Alex X. Liu, Jia Wang 0001
INFOCOM1
2015 Detecting and Localizing End-to-End Performance Degradation for Cellular Data Services
abstract
Nowadays mobile device (e.g., smartphone) users not only have a high expectation on the availability of the cellular data service, but also increasingly depend on the high end-to-end (E2E) performance of their applications. Since the E2E performance of individual application sessions may vary greatly, depending on factors such as the cellular network condition, the content provider, the type/model of the mobile devices, and the application software, detecting and localizing service performance degradations in a timely manner at large scale is of great value to cellular service providers. In this paper, we build a holistic measurement system that tracks session-level E2E performance metrics along with the service attributes for these factors. Using data collected from a major cellular service provider, we first model the expected E2E service performance with a regression based approach, detect performance degradation conditions based on the time series of fine-grained measurement data, and finally localize the service degradation using association-rule-mining techniques. Our deployment experience reveals that in 80% of the detected problem instances, performance degradation can be attributed to non-network-location specific factors, such as a common content provider, or a set of applications running on certain models of devices.
Faraz Ahmed, Jeffrey Erman, Zihui Ge, Alex X. Liu, Jia Wang 0001
SIGMETRICS1
2013 Identification of Sybil Communities Generating Context-Aware Spam on Online Social Networks
Faraz Ahmed, Muhammad Abulaish
APWeb1
2013 A generic statistical approach for spam detection in Online Social Networks
Faraz Ahmed, Muhammad Abulaish
Comput. Commun.1
2012 An MCL-Based Approach for Spam Profile Detection in Online Social Networks
abstract
Over the past few years, Online Social Networks (OSNs) have emerged as cheap and popular communication and information sharing media. Huge amount of information is being shared through popular OSN sites. This aspect of sharing information to a large number of individuals with ease has attracted social spammers to exploit the network of trust for spreading spam messages to promote personal blogs, advertisements, phishing, scam and so on. In this paper, we present a Markov Clustering (MCL) based approach for the detection of spam profiles on OSNs. Our study is based on a real dataset of Facebook profiles, which includes both benign and spam profiles. We model social network using a weighted graph in which profiles are represented as nodes and their interactions as edges. The weight of an edge, connecting a pair of user profiles, is calculated as a function of their real social interactions in terms of active friends, page likes and shared URLs within the network. MCL is applied on the weighted graph to generate different clusters containing different categories of profiles. Majority voting is applied to handle the cases in which a cluster contains both spam and normal profiles. Our experimental results show that majority voting not only reduces the number of clusters to a minimum, but also increases the performance values in terms of FPand FBmeasures from FP=0.85 and FB=0.75 to FP=0.88 and FB=0.79, respectively.
Faraz Ahmed, Muhammad Abulaish
TrustCom1
2010 Towards a Theory of Generalizing System Call Representation for In-Execution Malware Detection
abstract
The major contribution of this paper is two-folds: (1) we present our novel variable-length system call representation scheme compared to existing fixed- length sequence schemes, and (2) using this representation, we present our in-execution malware detector that can not only identify zero-day malware without any a priori knowledge but can also detect a malicious process while it is executing. Our representation scheme - a more generalized version of n-gram - can be visualized in a k-dimensional hyperspace in which processes move depending upon their sequence of system calls. The process marks its impact in space by generating hyper-grams that are later used to evaluate an unknown process according to their profile. The proposed technique is evaluated on a real world dataset extracted from a Linux System. The results of our analysis show that our in-execution malware detector with hyper- gram representation achieves low processing overheads and improved detection accuracies as compared to conventional n-grams.
Bilal Mehdi, Faraz Ahmed, Syed Ali Khayam, Muddassar Farooq
ICC2
2010 Using Computational Intelligence to Identify Performance Bottlenecks in a Computer System
Faraz Ahmed, Farrukh Shahzad 0002, Muddassar Farooq
PPSN (1)1