EDBT 2026 Demo / reviewers in the wild / expert
Md. Arifuzzaman
dblp:28/10124
· DBLP profile ↗
12ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SmartFLow: A Communication-Efficient SDN Framework for Cross-Silo Federated LearningabstractCross-silo Federated Learning (FL) enables multiple institutions to collaboratively train machine learning models while preserving data privacy. In such settings, clients repeatedly exchange model weights with a central server, making the overall training time highly sensitive to network performance. However, conventional routing methods often fail to prevent congestion, leading to increased communication latency and prolonged training. Software-Defined Networking (SDN), which provides centralized and programmable control over network resources, offers a promising way to address this limitation. To this end, we propose SmartFLow, an SDN-based framework designed to enhance communication efficiency in cross-silo FL. SmartFLow dynamically adjusts routing paths in response to changing network conditions, thereby reducing congestion and improving synchronization efficiency. Experimental results show that SmartFLow decreases parameter synchronization time by up to 47% compared to shortest-path routing and 41% compared to capacity-aware routing. Furthermore, it achieves these gains with minimal computational overhead and scales effectively to networks of up to 50 clients, demonstrating its practicality for real-world FL deployments. Osama Abu Hamdan, Hao Che, Engin Arslan, Md. Arifuzzaman |
CCNC | 4 |
| 2026 | FLEET: A Federated Learning Emulation and Evaluation Testbed for Holistic ResearchabstractFederated Learning (FL) presents a robust paradigm for privacy-preserving, decentralized machine learning. However, a significant gap persists between the theoretical design of FL algorithms and their practical performance, largely because existing evaluation tools often fail to model realistic operational conditions. Many testbeds oversimplify the critical dynamics among algorithmic efficiency, client-level heterogeneity, and continuously evolving network infrastructure. To address this challenge, we introduce the Federated Learning Emulation and Evaluation Testbed (FLEET). This comprehensive platform provides a scalable and configurable environment by integrating a versatile, framework-agnostic learning component with a high-fidelity network emulator. FLEET supports diverse machine learning frameworks, customizable real-world network topologies, and dynamic background traffic generation. The testbed collects holistic metrics that correlate algorithmic outcomes with detailed network statistics. By unifying the entire experiment configuration, FLEET enables researchers to systematically investigate how network constraints, such as limited bandwidth, high latency, and packet loss, affect the convergence and efficiency of FL algorithms. This work provides the research community with a robust tool to bridge the gap between algorithmic theory and real-world network conditions, promoting the holistic and reproducible evaluation of federated learning systems. Osama Abu Hamdan, Hao Che, Engin Arslan, Md. Arifuzzaman |
CCNC | 4 |
| 2026 | QuRA: Reinforcement Learning Based Routing for Quantum NetworksabstractQuantum routing deals with identifying a set of quantum repeaters to use to create entanglement between distant endpoints. Previous approaches proposed shortest-path and linear programming methods to find a solution to this problem. While the shortest path approach results in suboptimal performance, linear programming takes too long to find a solution as the network size and constraints increase. In this paper, we apply Deep Q-Reinforcement Learning (DQRL) to optimize routing in quantum networks both in terms of execution time and performance. The proposed Quantum Routing Algorithm (QuRA) first chooses which request to schedule among all requests. It then determines which route to take for the selected request. Since the number of all possible routes can be very high, we developed a hop-by-hop decision making model to lower the complexity while still attaining high performance. Experiments show that the proposed QuRA outperforms existing solutions by up to 90% in success rate and up to 79% in execution time. These results highlight QuRA as a scalable and effective solution for intelligent routing in quantum networks. Tasdiqul Islam, Engin Arslan, Md. Arifuzzaman |
CCNC | 3 |
| 2025 | Predicting and optimizing the mechanical properties of rejuvenated asphalt mix with RAP content
Kamrul Islam 0002, Uneb Gazder, Abdullah Al Mamun 0001, Md. Arifuzzaman, Hamad Ibrahim Al-Abdul Wahhab, Muhammad Muhitur Rahman |
Neural Comput. Appl. | 4 |
| 2024 | Deploy-Efficient and Fast Network Probing with Time-Series Foundation ModelsabstractActive probing, the practice of running data transfers for a short duration to collect or observe network metrics, is widely used for monitoring network performance, detecting anomalies, and optimizing data transfer algorithms. However, accurately estimating end-to-end throughput remains challenging due to the risks of self-inflicted network congestion during extended probing and the difficulty of balancing short probing durations with measurement precision. State-of-the-art transfer algorithms designed for large-scale HPC or Cloud data transfers rely on very short yet accurate periodic probing to achieve faster convergence to optimal solutions. To address these challenges, various approaches ranging from classical time-series forecasting to deep learning have been proposed to estimate metrics based on the first few seconds of data. While deep learning-based solutions outperform traditional time-series models, they often fail to generalize across diverse network environments with varying protocols, resources, and congestion levels, requiring significant data collection and retraining for each network. Furthermore, the evolving nature of networks renders these models obsolete within a short time. Consequently, running transfers for fixed durations and averaging metrics remains the most used approach despite its poor accuracy. To overcome these issues, we introduce npLLM, a scalable framework that leverages time-series foundation models for throughput estimation across dynamic, heterogeneous networks. By fine-tuning the pre-trained IBM Granite Time Series Foundation Model (TSFM) using data from local and production clusters, combined with data augmentation techniques, npLLM achieves high accuracy with very few samples in training, minimizing the need for extensive network-specific data collection. This transfer learning approach using the base foundation model enables efficient adaptation to new networks with minimal effort, significantly reducing setup time and resource requirements. Experimental results demonstrate that npLLM reduces probing error rate up to 70% compared to state-of-the-art solutions using as little as one-tenth of training data. Rasman Mubtasim Swargo, Md. Arifuzzaman |
IEEE Big Data | 2 |
| 2023 | Use Only What You Need: Judicious Parallelism For File Transfers in High Performance NetworksabstractParallelism is key to efficiently utilizing high-speed research networks when transferring large volumes of data. However, the monolithic design of existing transfer applications requires the same level of parallelism to be used for read, write, and network operations for file transfers. This, in turn, overburdens system resources since setting the parallelism level for the slowest component results in unnecessarily high parallelism for other components. Using more than necessary parallelism lead to increased overhead on system resources and unfair resource allocation among competing transfers. In this paper, we introduce modular file transfer architecture, Marlin, to separate I/O and network operations for file transfers so that parallelism can be independently adjusted for each component. Marlin adopts online gradient descent algorithm to swiftly search the solution space and find the optimal level of parallelism for read, transfer, and write operations. Experimental results collected under various network settings show that Marlin can identify and use a minimum parallelism level for each component, improving fairness among competing transfers and CPU utilization. Finally, separating network transfers from write operations allows Marlin to outperform the state-of-the-art solutions by more than 2x when transferring small datasets. Md. Arifuzzaman, Engin Arslan |
ICS | 1 |
| 2023 | Falcon: Fair and Efficient Online File Transfer OptimizationabstractResearch networks provide high-speed wide-area network connectivity between research and education institutions to facilitate large-scale data transfers. However, scalability issues of legacy transfer applications such as scp and FTP hinder the effective utilization of these networks. Although researchers extended the legacy transfer applications to increase their performance by exploiting I/O and network parallelism, these solutions necessitate users to fine-tune parallelism level, a task that is challenging even for experienced users due to the dynamic nature of networks. In this article, we propose an online optimization algorithm,Falcon, to tune the degree of parallelism for file transfers to maximize transfer throughput while keeping system overhead at a minimum. As research networks are shared infrastructures, we introduce a game theory-inspired novel utility function to evaluate the performance of various parallelism levels such that competing transfers are guaranteed to converge to a fair and stable solution. We assessed the performance ofFalconin isolated and production high-speed networks and found that it can discover optimal transfer parallelism in as little as 20 seconds and outperform the state-of-the-art solutions by more than$2\times$. Moreover,Falconis guaranteed to converge to Nash Equilibrium when multiple transfers compete for the same resources with the help of its game theory-inspired utility function. Finally, we demonstrate thatFalconcan also be used as a central transfer scheduler to speed up convergence time, increase stability, and enforce system/user-level resource limitations in shared networks. Md. Arifuzzaman, Brian Bockelman, Jim Basney, Engin Arslan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Be SMART, Save I/O: A Probabilistic Approach to Avoid Uncorrectable Errors in Storage SystemsabstractSilent data corruption poses a significant risk to the integrity of data in storage systems. Although error correction codes (ECC) can recover the majority of such errors, a non-negligible portion of them escape ECC, referred as uncorrectable errors (UEs). Despite being rare in nature, increasing scale of storage systems and fast-growing I/O rates decreased the mean time between UEs from months to hours. Yet, unlike disk failures, UEs are hard to predict with high precision, making it difficult to adopt proactive measures. In this paper, we introduce a probabilistic approach to deploy UE mitigation strategies that can capture significant portion of UE while keeping the system overhead at a tolerable range. To achieve this, we first estimate the probability of I/O operations to be exposed to UEs and find a minimum subset of disks for which employing UE avoidance strategies can lead to significant decrease in UE exposure. We demonstrate through extensive simulations that when the proposed probabilistic model is used to implement write verification strategy to detect and recover from UEs, more than 50% of all write-triggered UEs can be avoided with 1% read overhead, and more than 70% of UEs can be mitigated with less than 3.5% read overhead. We further measure the impact of incurred read overhead on write performance in production Lustre and GPFS file systems and validate our findings that more than half of UEs can be avoided while degrading write I/O throughout by less than 0.9%. Md. Arifuzzaman, Masudul Hasan Masud Bhuiyan, Mehmet Gümüs, Engin Arslan |
CLUSTER | 1 |
| 2021 | Towards Generalizable Network Anomaly Detection ModelsabstractFinding the root causes of network performance anomalies is critical to satisfy the quality of service requirements. In this paper, we introduce machine learning (ML) models to process TCP socket statistics to pinpoint underlying reasons of performance issues such as packet loss and jitter. More importantly, we introduce a novel feature engineering method to transform network-dependent metrics (e.g., total packet count and round trip time) in training datasets into network-independent forms to be able to transfer the models to new network settings without requiring to retrain them. Experimental results in various network settings show that the proposed feature engineering approach improves the performance of the models in previously unseen network settings from around 60% to nearly 90%. We believe ability to transfer ML models across networks will pave the way for wide adoption of ML solutions in production networks where collecting labeled data is not possible. Md. Arifuzzaman, Shafkat Islam, Engin Arslan |
LCN | 1 |
| 2021 | Online optimization of file transfers in high-speed networksabstractFile transfers in high-speed networks require network and I/O parallelism to reach high speeds, however, creating arbitrarily large numbers of I/O and network threads overwhelms system resources and causes fairness issues. In this paper, we introduce Falcon that combines a novel utility function with state-of-the-art online optimization algorithms to discover the degree of I/O and network parallelism for file transfer that can maximize the throughput while keeping system overhead low and ensuring fairness among competing transfers. Our extensive evaluations in several dedicated and production high-speed networks show that Falcon can find near-optimal solution in as little as 20 seconds and outperforms existing transfer application by 2--6×. Moreover, unlike other file transfer optimization solutions that fail to ensure fair resource allocation between competing transfers, Falcon is guaranteed to converge to Nash Equilibrium when multiple Falcon agents compete for the network resources with the help of its game theory-inspired utility function. Md. Arifuzzaman, Engin Arslan |
SC | 1 |
| 2017 | Moisture damage evaluation in SBS and lime modified asphalt using AFM and artificial intelligence
Md. Arifuzzaman, Muhammad Saiful Islam 0001, Muhammad Imtiaz Hossain |
Neural Comput. Appl. | 1 |
| 2016 | Joint Routing and MAC Layer QoS-Aware Protocol for Wireless Sensor NetworksabstractIn this paper, we propose a novel joint routing and medium access control (MAC) protocol with traffic differentiation, based on quality of service (QoS) for wireless sensor networks (WSNs). This is referred to as joint routing and MAC (JRM) protocol. By leveraging the classical layered approach and combining routing and MAC layer functions, the proposed JRM protocol achieves a solution for energy efficiency in WSNs. JRM also ensures low latency for prioritized traffic. There are three major advantages of the proposed protocol. Firstly, the instantaneous network information (e.g., estimated time to destination, and node's unavailability to forward additional packet) is piggy-backed with the control packets acknowledgement and clear- to-send of the MAC frame. Based on the updated network knowledge and the objective of the required performance metrics (e.g., energy and latency), the next hop neighbor is chosen dynamically with reduced control overhead. Secondly, the JRM protocol introduces an approach for finding the constrained shortest path for forwarding packets, which results in load balancing in WSNs. Finally, routers (nodes) in JRM require very little forwarding and routing table information, which is compatible with the resource constraints in WSNs. The efficiency of the proposed protocol is shown through simulation results. Md. Arifuzzaman, Octavia A. Dobre, Mohamed Hossam Ahmed, Telex Magloire Nkouatchah Ngatched |
GLOBECOM | 1 |