Raza Ahmad

dblp:187/0845 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Efficiently Reproducing Distributed Workflows in Notebook-based Systems
Talha Azaz, Raza Ahmad, Douglas Thain, Tanu Malik
CCGrid2
2026 Improving reproducibility of interactive notebooks using application virtualization
Raza Ahmad, Naga Nithin Manne, Tanu Malik
Future Gener. Comput. Syst.1
2025 Backpacks for Notebooks: Enabling Containerized Notebook Workflows in Distributed Environments
abstract
Notebooks have become widely adopted in the scientific community due to their interactive interface and ease of sharing. However, using notebooks to execute large-scale scientific workflows remains challenging. Scientific workflows are typically distributed and require resource provisioning and data management prior to execution. Because notebooks do not natively embed workflow specifications, users often resort to inserting custom configuration steps directly within notebook cells to enable provisioning. This practice undermines reproducibility, as the same notebook may not run consistently across different cluster environments. In this paper, we introduce the concept of a notebook backpack—a companion specification that captures the embedded workflow along with all relevant configuration elements. We describe how notebook tracing can be leveraged to automatically populate the backpack. We then describe an integrated tool that provisions a backpack on distributed resources. Using real-world case studies, we demonstrate that the backpack abstraction enables minimal modification of the notebook, portable execution, and cross-site reproducibility of notebook-based workflows on HPC clusters without significantly increasing notebook execution time.
Talha Azaz, Raza Ahmad, A. S. M. Shahadat Hossain, Furqan Baig, Shaowen Wang 0001, Kevin Lannon, Tanu Malik, Douglas Thain
eScience3
2025 DDM4TST: Diffusion Model for Fine-grained Text Style Transfer by Disentangled Representation
abstract
Fine-grained Text Style Transfer (FTST) aims to make targeted and precise modifications to specific stylistic components of a sentence. Existing methods typically attempt to disentangle style and content representations for FTST. However, style and content are inherently abstract, making them difficult to formalize and manipulate independently. To address this, we propose a novel framework, Disentanglement Diffusion Model for Text Style Transfer (DDM4TST) , which reformulates style transformation as either semantic or syntactic transformation. By learning disentangled representations, the model enables accurate and fine-grained style control. Specifically, we construct a feature parsing module to effectively separate semantic and syntactic representations. We then incorporate a diffusion model conditioned on these disentangled representations, allowing for fine-grained control of stylistic attributes while preserving the core content of the original sentence. This integration into the denoising process enhances the controllability and precision of the style transformation. Extensive experiments on the benchmark StylePTB dataset demonstrate that our model consistently outperforms widely adopted baselines. The results validate the effectiveness of our approach in achieving high-quality style transformation while maintaining content fidelity.
Cencen Liu, Qiugang Zhan, Dongyang Zhang 0001, Raza Ahmad
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2024 Accurate Path Prediction of Provenance Traces
abstract
Several security and workflow applications require provenance information at the operating system level for diagnostics. The resulting provenance traces are often more informative if they are efficiently mapped to execution paths within the control flow graph. However, current provenance systems do not map traces to control flow graphs for diagnostics purposes due to the computational complexity of mapping traces to graphs. We formulate the path prediction problem for provenance traces and take a machine learning approach to solve the problem. We develop a transformer-based graph convolutional network to predict paths. Our experiments demonstrate that our machine learning model achieves more than twice the accuracy on average compared to simple probabilistic models, with an increased computation time trade-off.
Raza Ahmad, Heeyoung Jung, Tanu Malik
CIKM1
2022 Reproducible Notebook Containers using Application Virtualization
abstract
Notebooks have gained wide popularity in scientific computing. A notebook is both a web-based interactive front-end to program workflows and a lightweight container for sharing code and its output. Reproducing notebooks in different target environments, however, is a challenge. Notebooks do not share the computational environment in which they are executed. Consequently, despite being shareable they are often not reproducible. The application virtualization (AV) method enables shareability and reproducibility of applications in heterogeneous environments. AV-based tools, however, encapsulate non-interactive, batch applications. In this paper, we present FLINC, a user-space method and tool for creating reproducible notebook containers. FLINC virtualizes the notebook process that enables interactive computation and creates notebook containers, which include the environment and all data dependencies accessed by the notebook file. It relies on provenance collected during virtualization to ensure the correct behavior of a notebook when run repeatedly in different environments. We demonstrate how FLINC exports notebook containers seamlessly to non-notebook environments. Our experiments show that FLINC creates lighter weight containers as compared to equivalent non-interactive, batch containers, and preserves the same interactive workflow for the user as in current notebook platforms.
Raza Ahmad, Naga Nithin Manne, Tanu Malik
e-Science1
2020 Content-defined Merkle Trees for Efficient Container Delivery
abstract
Containerization simplifies the sharing and deployment of applications when environments change in the software delivery chain. To deploy an application, container delivery methods push and pull container images. These methods operate on file and layer (set of files) granularity, and introduce redundant data within a container. Several container operations such as upgrading, installing, and maintaining become inefficient, because of copying and provisioning of redundant data. In this paper, we reestablish recent results that block-level deduplication reduces the size of individual containers, by verifying the result using content-defined chunking. Block-level deduplication, however, does not improve the efficiency of push/pull operations which must determine the specific blocks to transfer. We introduce a content-defined Merkle Tree (CDMT) over deduplicated storage in a container. CDMT indexes deduplicated blocks and determines changes to blocks in logarithmic time on the client. CDMT efficiently pushes and pulls container images from a registry, especially as containers are upgraded and (re-)provisioned on a client. We also describe how a registry can efficiently maintain the CDMT index as new image versions are pushed. We show the scalability of CDMT over Merkle Trees in terms of disk and network I/O savings using 15 container images and 233 image versions from Docker Hub.
Raza Ahmad, Tanu Malik
HiPC2
2017 Accurate Detection of Automatically Spun Content via Stylometric Analysis
abstract
Spammers use automated content spinning techniques to evade plagiarism detection by search engines. Text spinners help spammers in evading plagiarism detectors by automatically restructuring sentences and replacing words or phrases with their synonyms. Prior work on spun content detection relies on the knowledge about the dictionary used by the text spinning software. In this work, we propose an approach to detect spun content and its seed without needing the text spinner's dictionary. Our key idea is that text spinners introduce stylometric artifacts that can be leveraged for detecting spun documents. We implement and evaluate our proposed approach on a corpus of spun documents that are generated using a popular text spinning software. The results show that our approach can not only accurately detect whether a document is spun but also identify its source (or seed) document - all without needing the dictionary used by the text spinner.
Usman Shahid, Shehroze Farooqi, Raza Ahmad, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar
ICDM3
2016 Malware Slums: Measurement and Analysis of Malware on Traffic Exchanges
abstract
Auto-surf and manual-surf traffic exchanges are an increasingly popular way of artificially generating website traffic. Previous research in this area has focused on the makeup, usage, and monetization of underground traffic exchanges. In this paper, we analyze the role of traffic exchanges as a vector for malware propagation. We conduct a measurement study of nine auto-surf and manual-surf traffic exchanges over several months. We present a first of its kind analysis of the different types of malware that are propagated through these traffic exchanges. We find that more than 26% of the URLs surfed on traffic exchanges contain malicious content. We further analyze different categories of malware encountered on traffic exchanges, including blacklisted domains, malicious JavaScript, malicious Flash, and malicious shortened URLs.
Salman Yousaf, Umar Iqbal 0002, Shehroze Farooqi, Raza Ahmad, Zubair Shafiq, Fareed Zaffar
DSN4