Muhammad Abdullah Adnan

dblp:90/5533 · DBLP profile ↗
← Back
38ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0003-3219-9053ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 9 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Security and privacy · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Layout Anything: One Transformer for Universal Room Layout Estimation
abstract
We present Layout Anything, a transformer-based framework for indoor layout estimation that adapts the One-Former's universal segmentation architecture to geometric structure prediction. Our approach integrates One-Former's task-conditioned queries and contrastive learning with two key modules: (1) a layout degeneration strategy that augments training data while preserving Manhattan-world constraints through topology-aware transformations, and (2) differentiable geometric losses that directly enforce planar consistency and sharp boundary predictions during training. By unifying these components in an end-to-end framework, the model eliminates complex post-processing pipelines while achieving high-speed inference at 114ms. Extensive experiments demonstrate state-of-the-art performance across standard benchmarks, with pixel error (PE) of 5.43% and corner error (CE) of 4.02% on the LSUN, PE of 7.04% (CE 5.17%) on the Hedau and PE of 4.03% (CE 3.15%) on the Matterport3D-Layout datasets. The framework's combination of geometric awareness and computational efficiency makes it particularly suitable for augmented reality applications and large-scale 3D scene reconstruction tasks.
Md Sohag Mia, Muhammad Abdullah Adnan
WACV2
2025 Bridging the Last Mile: Unpacking the Rural Digital Divide in Bangladesh
abstract
Peer Reviewed
Rayhan Rashed, Muhammad Masroor Ali, Sadia Sharmin, Md Shariful Islam Bhuyan, Muhammad Abdullah Adnan, Anindya Iqbal, Md Shohrab Hossain, Mohammad Sohel Rahman, A. B. M. Alim Al Islam
COMPASS5
2025 A framework for checking and mitigating the security vulnerabilities of cloud service RESTful APIs
Md Shohel Khan, Rubaiyat Sha Fardin Siam, Muhammad Abdullah Adnan
Serv. Oriented Comput. Appl.3
2024 Creek: Leveraging Serverless for Online Machine Learning on Streaming Data
Nazmul Takbir, Tahmeed Tarek, Muhammad Abdullah Adnan
CLOSER3
2024 GuardFL: Detection and Removal of Malicious Clients Through Majority Consensus
abstract
Federated machine learning (FL) is susceptible to poisoning attacks, where malicious clients embed backdoors into the global model. Even a single attacker can launch potent attacks, challenging existing defenses. This work proposes GuardFL, a novel defense against backdoor attacks in FL. GuardFL integrates majority consensus and client feedback mechanisms to detect and isolate malicious clients. Extensive testing demonstrates the effectiveness of GuardFL in identifying and isolating attackers. We evaluate GuardFL using three benchmark image classification datasets including CIFAR-10, MNIST, and FMNIST (Fashion MNIST). We believe our techniques yield a robust defense against backdoor poisoning in FL.
Md Hasebul Hasan, Md Shariful Islam Khan, Md Shafayatul Haque, Ajmain Yasar Ahmed, Muhammad Abdullah Adnan
COMPSAC5
2024 XGDFed: Exposing Vulnerability in Byzantine Robust Federated Binary Classifiers with Novel Aggregation-Agnostic Model Poisoning Attack
abstract
Federated Learning (FL) is a machine learning method in which multiple edge devices collaborate to train an ML model without exchanging data. Instead, each device sends its local update to the central server, which then aggregates these updates using a Byzantine-robust aggregation rule to prevent Byzantine failures of edge devices. Recent studies have shown that some carefully crafted model poisoning attacks can evade Byzantine robust aggregations. However, these state-of-the-art Byzantine robust attacks rely on knowing the central server's aggregation algorithm or other benign model updates from all the edge devices, which is impractical. To address this, we have proposed a novel model poisoning attack called XGDFed, which effectively targets the decision function of binary classifiers regardless of the server's aggregation method. Our experiments indicate that XGDFed outperforms state-of-the-art attacks using the CIFAR-I0, FashionMNIST, MNIST, IJCNN1, and Acoustic datasets. With only 10% of attackers in the Fashion-MNIST dataset, XGDFed reduced the accuracy of the federated binary classifier from 83% to 51 %. This paper is the first empirical evaluation of the robustness of Byzantine-robust aggregations against state-of-the-art model poisoning attacks on binary classifiers in a federated environment. Although binary classifiers are often overlooked in the research literature, they can be relevant in different applications of federated learning. The code we used in our research is open-source.
Israt Jahan Mouri, Muhammad Ridowan, Muhammad Abdullah Adnan
COMPSAC3
2024 BTVSL: A Novel Sentence-Level Annotated Dataset for Bangla Sign Language Translation
abstract
Sign language, a vital communication tool for individuals with hearing impairments worldwide, is prominently utilized within the BangIa-speaking community. BangIa is recognized as the seventh most spoken language globally. However, research in BangIa Sign Language Recognition (SLR) - the process of translating symbols or words from images and videos - has been predominantly confined to controlled environments with limited samples and rudimentary symbol annotations, impeding its application in real-world scenarios. In contrast to previous studies, our research concentrates on SLR. It delves into the relatively unexplored territories of BangIa Sign Language Translation (SLT) and Sign Language Production (SLP), areas that have been largely overlooked due to dataset constraints. We introduce BTVSL, a comprehensive BangIa Sign Language dataset derived from the YouTube series ‘BTV Desh o Jonopoder Khobor'. This dataset, featuring 60 hours of news content in sign language delivered by professionals, represents the largest sentence-level dataset available for BangIa SLT, encompassing a broad spectrum of expressions. Leveraging BTVSL, we evaluated four distinct SLT models, achieving an average BLEU score of 20.42. This result underscores the potential of BTVSL in enhancing the accuracy of sign language translation.
Iftekhar E. Mahbub Zeeon, Mir Mahathir Mohammad, Muhammad Abdullah Adnan
FG3
2024 Exploring Attention Mechanisms in Integration of Multi-Modal Information for Sign Language Recognition and Translation
abstract
Understanding intricate and fast-paced movements of body parts is essential for the recognition and translation of sign language. The inclusion of additional information intended to identify and locate the moving body parts has been an interesting research topic recently. However, previous works on using multi-modal information raise concerns such as sub-optimal multi-modal feature merging method, or the model itself being too computationally heavy. In our work, we have addressed such issues and used a plugin module based on cross-attention to properly attend to each modality with another. Moreover, we utilized 2-stage training to remove the dependency of separate feature extractors for additional modalities in an end-to-end approach, which reduces the concern about computational complexity. Besides, our additional cross-attention plugin module is very lightweight which doesn’t add significant computational overhead on top of the original baseline. We have evaluated the performance of our approaches on the RWTH-PHOENIX-2014 dataset for sign language recognition and the RWTH-PHOENIX-2014T dataset for the sign language translation task. Our approach reduced the WER by 0.9 on the recognition task and increased the BLEU- 4 scores by 0.8 on the translation task.
Zaber Ibn Abdul Hakim, Rasman Mubtasim Swargo, Muhammad Abdullah Adnan
ICIP3
2024 EMPATH: MediaPipe-Aided Ensemble Learning with Attention-Based Transformers for Accurate Recognition of Bangla Word-Level Sign Language
Kazi Reyazul Hasan, Muhammad Abdullah Adnan
ICPR (21)2
2023 Towards Poisoning of Federated Support Vector Machines with Data Poisoning Attacks
Israt Jahan Mouri, Muhammad Ridowan, Muhammad Abdullah Adnan
CLOSER3
2023 Compressed neural architecture utilizing dimensionality reduction and quantization
Mohammad Saqib Hasan, Rukshar Alam, Muhammad Abdullah Adnan
Appl. Intell.3
2023 A Prediction Based Replica Selection Strategy for Reducing Tail Latency in Geo-Distributed Systems
abstract
The deployment of modern applications in geo-distributed systems results in performance fluctuation which is a consequence of long-tail latency. To deliver high-quality services these applications always strive to adapt to the changing situation and an appropriate replica selection strategy is one efficient way to achieve this. Several replica selection strategies have already been developed but none of them are efficient enough to reduce tail latency and adapt to the dynamic environment of the geo-distributed systems. In this paper, we present the design and implementation of a prediction based replica selection strategy for reducing tail latency in geo-distributed systems. We have meticulously designed the proposed strategy to adapt to the dynamic behavior of the distributed system. For evaluating its effectiveness in reducing tail latency and adapting to the dynamic behavior of the geo-distributed systems we perform some extensive experiments in a 15 nodes Cassandra cluster that is deployed on Amazon EC2 over 5 geographical regions. For generating test datasets and workloads we use industry-standard Yahoo Cloud Serving Benchmark (YCSB). Our experimental results show that the proposed strategy not only reduces tail latency but also increases the overall throughput of the systems.
Santa Maria Shithil, Muhammad Abdullah Adnan
IEEE Trans. Cloud Comput.2
2022 Fast Support Vector Machine Using Singular Value Decomposition
abstract
Nowadays, data is being generated very rapidly worldwide, and we need to analyze this data distributedly and efficiently. In this paper, we provide a method named SVDSVM that can compute linear support vector classification (SVM) distributedly and efficiently. SVDSVM reduces the computation and communication overhead of QRSVM (a recently proposed SVM solver) by replacing householder QR factorization with stochastic singular value decomposition. By removing additional gathering, scattering, and distributed computation, this method reduces the per iteration time of the dual ascent step and improves overall training time compared to QRSVM. Our evaluation with benchmark datasets shows that our method reduces per iteration time of dual ascent step around 5× compared to QRSVM, which results in an overall 50% time reduction of the total algorithm. However, as singular value decomposition is not entirely lossless, it results in a small accuracy drop which we found around 0.5-1.5% for different datasets.
Sabuj Laskar, Muhammad Abdullah Adnan
IEEE Big Data2
2022 RS-PKE: Ranked Searchable Public-Key Encryption for Cloud-Assisted Lightweight Platforms
abstract
Since more and more data from lightweight platforms like IoT devices or mobile apps are being outsourced to the cloud, the need to ensure privacy while retaining data usability is essential. In this paper, we design a framework where lightweight platforms like IoT devices can encrypt documents and generate document indexes using the public key before uploading the document to the cloud, and an admin can search and retrieve the top-k most relevant documents that match a specific keyword using the private key. In most existing searchable encryption that supports IoT, all the documents that match a queried keyword are returned to the admin. This is not practical as IoT devices continuously upload data. We formally name our framework asRanked Searchable Public-Key Encryption (RS-PKE). We also implemented a prototype of RS-PKE and tested it in the Amazon EC2 cloud using the RFC dataset. The comprehensive evaluation demonstrates that RS-PKE is efficient and secure for practical deployment.
Israt Jahan Mouri, Muhammad Ridowan, Muhammad Abdullah Adnan
CODASPY3
2022 MK-RS-PKE: Multi-Keyword Ranked Searchable Public-Key Encryption for Cloud-Assisted Lightweight Platforms
abstract
Since more and more data from lightweight platforms like IoT devices or mobile apps are being outsourced to the cloud, the need to ensure privacy while retaining data usability is essential. This paper designs a framework where lightweight platforms like IoT devices can encrypt documents and generate document indexes using the public key before uploading the document to the cloud. An admin can search and retrieve the top-k most relevant documents that match a multi-keyword query using the private key. Most existing searchable encryption that supports IoT returns all the documents matching queried keywords. However, IoT devices can produce massive data, which is not practical for such schemes. We formally name our framework Multi-keyword Ranked Searchable Public-Key Encryption (MK-RS-PKE).
Israt Jahan Mouri, Muhammad Ridowan, Muhammad Abdullah Adnan
CODASPY3
2022 Breaking the Curse of Class Imbalance: Bangla Text Classification
abstract
This article addresses the class imbalance issue in a low-resource language called Bengali. As a use-case, we choose one of the most fundamental NLP tasks, i.e., text classification, where we utilize three benchmark text corpora: fake-news dataset, sentiment analysis dataset, and song lyrics dataset. Each of them contains a critical class imbalance. We attempt to tackle the problem by applying several strategies that include data augmentation with synthetic samples via text and embedding generation in order to augment the proportion of the minority samples. Moreover, we apply ensembling of deep learning models by subsetting the majority samples. Additionally, we enforce the focal loss function for class-imbalanced data classification. We also apply the outlier detection technique, data resampling, and hidden feature extraction to improve the minority-f1 score. All of our experimentations are entirely focused on textual content analysis, which results in a more than90%minority f1 score for each of the three tasks. It is an excellent outcome on such highly class-imbalanced datasets.
Md. Rafi Ur Rashid, Mahim Mahbub, Muhammad Abdullah Adnan
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2021 BAND: A Benchmark Dataset forBangla News Audio Classification
abstract
Despite being the sixth most widely spoken language in the world, Bangla has barely received any attention in the domain of audio-visual news classification. In this work, we collect, annotate, and prepare a comprehensive news audio dataset in Bangla, comprising 5120 news clips, with around 820 hours of total duration. We also conduct practical experiments to obtain a human baseline for the news audio classification task. Later, we implement one of the human approaches by performing news classification directly on the audio features using various state-of-the-art classifiers and a few transfer learning models. To the best of our knowledge, this is the very first work developing a benchmark dataset for news audio classification in Bangla.
Md. Rafi Ur Rashid, Mahim Mahbub, Muhammad Abdullah Adnan
MMAsia3
2021 Neuro-Scientific Analysis of Weights in Neural Networks
abstract
Deep learning is a popular topic among machine learning researchers nowadays, with great strides being made in recent years to develop robust artificial neural networks for faster convergence to a reasonable accuracy. Network architecture and hyperparameters of the model are fundamental aspects of model convergence. One such important parameter is the initial values of weights, also known as weight initialization. In this paper, we perform two research tasks concerned with the weights of neural networks. First, we develop three novel weight initialization algorithms inspired by the neuroscientific construction of the mammalian brains and then test them on benchmark datasets against other algorithms to compare and assess their performance. We call these algorithms the lognormal weight initialization, modified lognormal weight initialization, and skewed weight initialization. We observe from our results that these initialization algorithms provide state-of-the-art results on all of the benchmark datasets. Second, we analyze the influence of training an artificial neural network on its weight distribution by measuring the correlation between the quantitative metrics of skewness and kurtosis against the model accuracy using linear regression for different weight initializations. Results indicate a positive correlation between network accuracy and skewness of the weight distribution but no affirmative relation between accuracy and kurtosis. This analysis provides further insight into understanding the inner mechanism of neural network training using the shape of weight distribution. Overall, the works in this paper are the first of their kind in incorporating neuroscientific knowledge into the domain of artificial neural network weights.
Mohammad Saqib Hasan, Rukshar Alam, Muhammad Abdullah Adnan
Int. J. Pattern Recognit. Artif. Intell.3
2021 Fast, scalable and geo-distributed PCA for big data analytics
T. M. Tariq Adnan, Md. Mehrab Tanjim, Muhammad Abdullah Adnan
Inf. Syst.3
2021 Towards Greening MapReduce Clusters Considering Both Computation Energy and Cooling Energy
abstract
Increased processing power of MapReduce clusters generally enhances performance and availability at the cost of substantial energy consumption that often incurs higher operational costs (e.g., electricity bills) and negative environmental impacts (e.g., carbon dioxide emissions). There exist a few greening methods for computing clusters in the literature that focus mainly on computational energy consumption leaving cooling energy, which occupies a significant portion of the total energy consumed by the clusters. To this extent, in this article, we propose a machine learning-based approach that reduces the total energy consumption of a MapReduce cluster considering both computational energy and cooling energy. Our approach predicts the number of machines that results in minimum total energy consumption. We perform the prediction through applying different machine learning techniques over year-long data collected from a real setup. We evaluate performance of our approach through both real test-bed experimentation and simulation. Our evaluation reveals that our approach achieves substantial reduction in total energy consumption compared to other state-of-the-art alternatives while experiencing marginal performance degradation in a few cases.
Tarik Reza Toha, A. S. M. Rizvi, Jannatun Noor 0001, Muhammad Abdullah Adnan, A. B. M. Alim Al Islam
IEEE Trans. Parallel Distributed Syst.4
2020 Energy Efficient Decentralized Geographical Load Balancing via Dynamic Deferral of Workload
abstract
Integration of on-demand Cloud services has become an orthodox approach adopted by enterprises of all scales, especially over the past decade. Thus, efficient load balancing techniques in data centers that perform intensive cloud computations become increasingly necessary to reduce energy costs. The essence of cloud computing lies within data centers being located in diverse locations, thus in different time zones. Thus, we take advantage of spatial and temporal variations of the data centers in our load balancing decisions. We demonstrate our design of a decentralized online algorithm for geographical load balancing via dynamic deferral of arriving workload from the data center. It responds with reduced energy costs by adapting to the dynamic price changes in its region via a head start of predicting prices in the real market. We compare our algorithm with the centralized approach without future price predictions and have achieved cost savings. We validate our model by experiments run on Google traces, and show that geographical load balancing with dynamic deferral can provide up to 32% reduction in costs to execute arriving load in data centers with both temporal and geographical variations.
Zeenat Islam, Muhammad Abdullah Adnan
CLOUD3
2020 A Prediction Based Replica Selection Strategy for Reducing Tail Latency in Distributed Systems
abstract
In the distributed system, applications often suffer from unexpectedly high response time fluctuation which is a consequence of long-tail latency. One efficient way of reducing tail latency is to use an appropriate replica selection strategy. In this paper, through extensive experiments, we analyze and evaluate two popular state-of-the-art replica selection strategies of key-value store systems: Cassandra dynamic snitch and C3. After analyzing all the experimental results we develop a prediction based replica selection strategy and implement it on Cassandra 3.0. Our proposed algorithm outperforms both C3 and Cassandra dynamic snitch. For evaluation, we test Cassandra dynamic snitch, C3, and our proposed algorithm on a 15 nodes Cassandra cluster. For generating test datasets and workloads we use industry-standard Yahoo Cloud Serving Benchmark (YCSB).
Santa Maria Shithil, Muhammad Abdullah Adnan
CLOUD2
2020 Truth or Lie: Pre-emptive Detection of Fake News in Different Languages Through Entropy-based Active Learning and Multi-model Neural Ensemble
abstract
In recent times, the circulation of fake news on social networks has increased exponentially with spikes in propagation seen during and after the 2016 US elections. Hence, there has been a surge in research into automated fake news detection. However, most research tends towards supervised learning which requires a significant amount of labeled data which is difficult to obtain. Thus, in this paper, we develop a semi-supervised learning method for fake news detection incorporating active learning based on entropy as a query strategy to train a multi-model neural ensemble architecture. The goal of the research is to achieve high accuracy on fake news detection while using lower amounts of data. Our experiments against other standards indicate promising results, with our model achieving high accuracy with 4% to 28% of the dataset.
Mohammad Saqib Hasan, Rukshar Alam, Muhammad Abdullah Adnan
ASONAM3
2020 CloudPush: Smart Delivery of Push Notification to Secure Multi-User Support for IoT Devices
abstract
Internet of things (IoT) with a cloud server has become popular nowadays and it's going to be used in almost every aspect of human life. All devices will be connected to the internet and can communicate with each other where cloud plays an import role in the IoT environment. However, often cloud-enabled IoT environments have potential security risks and do not have multi-user support. In this paper, we discuss an IoT push messaging framework named CloudPush framework consisting of a client application, IoT devices, and a cloud system. In this framework, IoT devices can work on an ad hoc network and send event notifications to the client applications through the cloud. We show that CloudPush framework has vulnerabilities while maintaining multiple user accounts between a client application and IoT device in the cloud. The client application can receive unintended and unauthorized notification messages due to the lack of managing multiple accounts properly in the cloud server. To ensure stability in this framework while sending push notifications through the cloud by IoT devices, we discuss potential vulnerabilities and their solutions in this paper. We demonstrate that the aggregated throughput of CloudPush framework is 12-15% better than IoTivity framework even though IoTivity does not support multi-user for an IoT resource and a client application. If IoT device's events are sent to multiple client applications i.e. events are distributed among client applications, then the throughput of CloudPush framework increases to 12-25% compared with the IoTivity framework because the CloudPush framework runs optimized searching algorithm in cloud and scales event notifications in both cloud server and cloud push service layer. For a secured multi-user support, notification message data is encrypted that makes the CloudPush system 3-5% slower but still, it performs 9-12% better than the IoTivity framework.
Md. Shamsul Arifin Mozumder, Muhammad Abdullah Adnan
IC2E2
2020 WLEC: A Not So Cold Architecture to Mitigate Cold Start Problem in Serverless Computing
abstract
As serverless computing gains popularity among developers for its low costing and elasticity, it has emerged as a promising research field in computer science. Despite its popularity, the cold start remains an issue that needs more attention. In this paper, we address the cold start problem of the serverless platform. We propose WLEC, a container management architecture to minimize the cold start time. WLEC uses a modified S2LRU structure, called S2LRU ++ with an additional third queue. We implement WLEC in OpenLambda and evaluate it in both AWS and Local VM environment with six different metrics in addition to one real-time image resizing application. Among improvements in all metrics, 50% less memory consumption compared to the all-warm method and 31% average cold start duration reduction compared to the no-warm method are the most notable ones.
Khondokar Solaiman, Muhammad Abdullah Adnan
IC2E2
2020 Bangla Voice Command Recognition in end-to-end System Using Topic Modeling based Contextual Rescoring
abstract
In this work, we perform contextual rescoring using multi-label topic modeling to improve the performance of an End-to-End Bangla voice command recognition system. We use a hybrid of Connectionist Temporal Classification (CTC) and Attention mechanism in our End-to-End architecture. We use Recurrent Neural Network (RNN) as language model and La-beled LDA (Latent Dirichlet allocation) for contextual rescoring. Our experiments show that our rescoring method reduces Word Error Rate (WER) from 16.7% to 12.8% in Bangla voice command recognition task when the relevant context is provided. The system does not lose any performance when irrelevant context is provided.
Nafis Sadeq, Shafayat Ahmed, Sudipta Saha Shubha, Md. Nahidul Islam, Muhammad Abdullah Adnan
ICASSP5
2020 WeightGrad: Geo-Distributed Data Analysis Using Quantization for Faster Convergence and Better Accuracy
abstract
High network communication cost for synchronizing weights and gradients in geo-distributed data analysis consumes the benefits of advancement in computation and optimization techniques. Many quantization methods for weight, gradient or both have been proposed in recent years where weight-quantized model suffers from error related to weight dimension and gradient-quantized method suffers from slow convergence rate by a factor related to the gradient quantization resolution and gradient dimension. All these methods have been proved to be infeasible in terms of distributed training across multiple data centers all over the world. Moreover recent studies show that communicating over WANs can significantly degrade DNN model performance by upto 53.7x because of unstable and limited WAN bandwidth. Our goal in this work is to design a geo-distributed Deep-Learning system that (1) ensures efficient and faster communication over LAN and WAN and (2) maintain accuracy and convergence for complex DNNs with billions of parameters. In this paper, we introduce WeightGrad which acknowledges the limitations of quantization and provides loss-aware weight-quantized networks with quantized gradients for local convergence and for global convergence it dynamically eliminates insignificant communication between data centers while still guaranteeing the correctness of DNN models. Our experiments on our developed prototypes of WeightGrad running across 3 Amazon EC2 global regions and on a cluster that emulates EC2 WAN bandwidth show that WeightGrad provides 1.06% gain in top-1 accuracy, 5.36x speedup over baseline and 1.4x-2.26x over the four state-of-the-art distributed ML systems.
Syeda Nahida Akter, Muhammad Abdullah Adnan
KDD2
2020 Preparation of Bangla Speech Corpus from Publicly Available Audio & Text
abstract
Automatic speech recognition systems require large annotated speech corpus. The manual annotation of a large corpus is very difficult. In this paper, we focus on the automatic preparation of a speech corpus for Bangladeshi Bangla. We have used publicly available Bangla audiobooks and TV news recordings as audio sources. We designed and implemented an iterative algorithm that takes as input a speech corpus and a huge amount of raw audio (without transcription) and outputs a much larger speech corpus with reasonable confidence. We have leveraged speaker diarization, gender detection, etc. to prepare the annotated corpus. We also have prepared a synthetic speech corpus for handling out-of-vocabulary word problems in Bangla language. Our corpus is suitable for training with Kaldi. Experimental results show that the use of our corpus in addition to the Google Speech corpus (229 hours) significantly improves the performance of the ASR system.
Shafayat Ahmed, Nafis Sadeq, Sudipta Saha Shubha, Md. Nahidul Islam, Muhammad Abdullah Adnan, Mohammad Zuberul Islam
LREC5
2019 A Scalable Algorithm for Multi-class Support Vector Machine on Geo-Distributed Datasets
abstract
Training machine learning models on large scale data to efficiently discover valuable information while maintaining the security and privacy of data remains an important research issue. Often many real-life applications such as healthcare systems or financial organizations distribute data over many data centers and these data centers may have a different privacy policy. The joint decision over this dataset while not sharing the local information of the data centers is a must and becomes a challenging problem. The State-of-the-art method mainly relies on cryptographic technique to ensure the privacy for data communication between the data centers. This technique alone is not suitable for the large-scale geo-distributed datasets as this is designed for small-scale systems. In this paper, we propose a novel approach to achieve privacy-preserving SVM algorithm where the training set is distributed and each partition can contain large-scale data. We utilize the traditional SVM model to train the dataset of local data centers and with a few parameters sent by those training models, a centralized machine calculates the final result. We show that the proposed model is secure in an adverse environment and use simulations to demonstrate its correctness and computation speed compared to other parallel SVM training models.
Tasnim Kabir, Muhammad Abdullah Adnan
IEEE BigData2
2019 Real Time Principal Component Analysis
abstract
By processing the data in motion, real-time data processing enables us to extract instantaneous results from online input data that ensures timely responsiveness to events as well as a much enhanced capacity to process large data sets. This is especially important when decision loops include querying and processing data on the web where size and latency considerations make it impossible to process raw data in real-time. This makes dimensionality reduction techniques, like principal component analysis (PCA), an important data preprocessing tool to gain insights into data. In this paper, we propose a variant of PCA, that is suited for real-time applications. In the real-time version of the PCA problem, we maintain a window over the most recent data and project every incoming row of data into lower dimensional subspace, which we generate as the output of the model. The goal is to minimize the reconstruction error of the output from the input. We use the reconstruction error as the termination criteria to update the eigenspace as new data arrives. To verify whether our proposed model can capture the essence of the changing distribution of large datasets in real-time, we have implemented the algorithm and evaluated performance against carefully designed simulations that change distributions of data sources over time in a controllable manner. Furthermore, we have demonstrated that our algorithm can capture the changing distributions of real-life datasets by running simulations on datasets from a variety of real-time applications e.g. localization, customer expenditure, etc. We propose algorithmic enhancements that rely upon spectral analysis to improve dimensionality reduction. Results show that our method can successfully capture the changing distribution of data in a real-time scenario, thus enabling real-time PCA.
Ranak Roy Chowdhury, Muhammad Abdullah Adnan, Rajesh K. Gupta 0001
ICDE2
2018 Improving Query Execution Performance in Big Data using Cuckoo Filter
abstract
Performance is a critical concern when reading and writing billions of records stored in a Big Data warehouse. To improve performance, Big Data systems utilize the Eventual Consistency model as well as a probabilistic data structure called Bloom Filter. In recent times, cost minimization and privacy concerns have led Big Data systems to regularly delete data besides reading and writing. Bloom Filter has an important limitation of not supporting deletion of elements from within it. This limitation along with constraints of the Eventual Consistency model prevents the system to delete data from disk immediately upon an execution of a DELETE command. Rather data is actually deleted during a compaction process, which is executed very infrequently due to it being resource-intensive. During the time between data deletion and compaction (typically a few days), lookup queries fetch records from disk that includes the deleted rows, then detect and remove the deleted records from the returning result set. This leads to poor query execution performance. In this paper, we propose a scheme to improve performance of lookup queries after data deletion by replacing Bloom Filter with a better probabilistic data structure called Cuckoo Filter that supports deletion of elements from within it. We evaluate our proposed scheme using a popular Big Data database (Cassandra) that uses both the Eventual Consistency model and Bloom Filter, and show that performance of lookup queries executed after data deletion can improve for up to 100%. We also show that our scheme does not degrade performance of other data manipulation queries as a side effect.
Sharafat Ibn Mollah Mosharraf, Muhammad Abdullah Adnan
IEEE BigData2
2018 sSketch: A Scalable Sketching Technique for PCA in the Cloud
abstract
Multidimensional data appear frequently in many web-related applications, e.g., product ratings, the bag-of-words representation of web pages, etc. Principal Component Analysis (PCA) has been widely used for discovering patterns in relationships among entities in multidimensional data. However, existing algorithms for PCA have limited scalability since they explicitly materialize intermediate data, whose size rapidly grows as the dimension increases. To avoid scalability issues, we propose sSketch, a scalable sketching technique for PCA that employs several optimization ideas, such as mean propagation, efficient sparse matrix operations, and effective job consolidation to minimize intermediate data. Using sSketch, we also provide two other scalable methods for deriving singular value and 2-norm of reconstruction error, both of which are used for data analysis purpose. We provide our implementation on popular Spark framework for distributed platform. We compare our method against state-of-the-art library functions available for distributed settings, namely MLlib-PCA and Mahout-PCA with real big datasets. Our experiments show that our method outperforms both of them by a wide margin. To encourage reproducibility, the source code of sSketch is made publicly available at \hrefhttps://github.com/DataMiningResearch/sSketch https://github.com/DataMiningResearch/sSketch.
Md. Mehrab Tanjim, Muhammad Abdullah Adnan
WSDM2
2016 Dynamic Weight on Static Trust for trustworthy social media networks
abstract
Something in our mind does not always reflect with our fingers on keyboard and mouse, but we respect our blood which might be the way out of the wariness in usage of social media networks today. Social media networks are already the largest network among the population across our globe and for sure, it will not stop expanding its territory. Security vulnerabilities on the other side have been mounting in parallel. A comprehensive effort to change ourselves along with reversing the process or architecture of social media network sites might be the solution of this serious but unavoidable problem. We can create a trusted social media network model that will ensure trust level of a user to connect with other users. Users can trust each other based on the trust level. A lot of research has been done on trust model for social media networks and till now none is established effectively as most of them fail to capture user's trustworthiness with full proof. In order to signpost user's trustworthiness, we model trust with two components: i) Static Trust and ii) Dynamic Weight on Trust. Static Trust is captured by anti-anonymity process based on genealogical tree which will not be changed frequently. And Dynamic Weight on Trust is an algorithmic process to determine score based on predefined scoring scale that is applied on user's genealogical tree and community members. Evaluation of the above model on selected user group is quite positive in taking decision for users when (i) connecting with new friends (ii) self-representing in trustworthy manner. Our proposed trust model fully eliminates anonymity in social media networks which will be a trusted zone for users.
Mohammad Khaliqur Rahman, Muhammad Abdullah Adnan
PST2
2014 Workload Shaping to Mitigate Variability in Renewable Power Use by Data Centers
abstract
This paper explores the opportunity for energy saving in data centers using the flexibility from the Service Level Agreements (SLAs) and proposes a novel approach for scheduling workload that incorporates use of renewable energy sources. We investigate how much renewable power to store and how much workload to delay for increasing renewable usage while meeting latency constraints. We present an LP formulation for mitigating variability in renewable generation by dynamic deferral and give two online algorithms to determine optimal balance of workload deferral and power use. We prove the feasibility of the online algorithms and show that their worst case performances are bounded by constant factors with respect to the offline formulation. We validate our algorithms by trace-driven simulation on MapReduce workload and collected and publicly available wind and solar power generation data. Results show that the algorithms give 20-30% energy-savings compared to the naive 'follow the workload' policy.
Muhammad Abdullah Adnan, Rajesh K. Gupta 0001
IEEE CLOUD1
2013 Path Consolidation for Dynamic Right-Sizing of Data Center Networks
abstract
Data center topologies typically consist of multirooted trees with many equal-cost paths between a given pair of hosts. Existing power optimization techniques do not utilize this property of data center networks for power proportionality. In this paper, we exploit this opportunity and show that significant energy savings can be achieved via path consolidation in the network. We present an offline formulation for the flow assignment in a data center network and develop an online algorithm by path consolidation for dynamic right-sizing of the network to save energy. To validate our algorithm, we build a flow level simulator for a data center network. Our simulation on flow traces generated from MapReduce workload shows 80% reduction in network energy consumption in data center networks and ~25% more energy savings compared to the existing techniques for saving energy in data center networks.
Muhammad Abdullah Adnan, Rajesh K. Gupta 0001
IEEE CLOUD1
2013 Utility-aware deferred load balancing in the cloud driven by dynamic pricing of electricity
abstract
Distributed computing resources in a cloud computing environment provides an opportunity to reduce energy and its cost by shifting loads in response to dynamically varying availability of energy. This variation in electrical power availability is represented in its dynamically changing price that can be used to drive workload deferral against performance requirements. But such deferral may cause user dissatisfaction. In this paper, we quantify the impact of deferral on user satisfaction and utilize flexibility from the service level agreements (SLAs) for deferral to adapt with dynamic price variation. We differentiate among the jobs based on their requirements for responsiveness and schedule them for energy saving while meeting deadlines and user satisfaction. Representing utility as decaying functions along with workload deferral, we make a balance between loss of user satisfaction and energy efficiency. We model delay as decaying functions and guarantee that no job violates the maximum deadline, and we minimize the overall energy cost. Our simulation on MapReduce traces show that energy consumption can be reduced by ∼15%, with such utility-aware deferred load balancing. We also found that considering utility as a decaying function gives better cost reduction than load balancing with a fixed deadline.
Muhammad Abdullah Adnan, Rajesh K. Gupta 0001
DATE1
2012 Energy Efficient Geographical Load Balancing via Dynamic Deferral of Workload
abstract
With the increasing popularity of Cloud computing and Mobile computing, individuals, enterprises and research centers have started outsourcing their IT and computational needs to on-demand cloud services. Recently geographical load balancing techniques have been suggested for data centers hosting cloud computation in order to reduce energy cost by exploiting the electricity price differences across regions. However, these algorithms do not draw distinction among diverse requirements for responsiveness across various workloads. In this paper, we use the flexibility from the Service Level Agreements (SLAs) to differentiate among workloads under bounded latency requirements and propose a novel approach for cost savings for geographical load balancing. We investigate how much workload to be executed in each data center and how much workload to be delayed and migrated to other data centers for energy saving while meeting deadlines. We present an offline formulation for geographical load balancing problem with dynamic deferral and give online algorithms to determine the assignment of workload to the data centers and the migration of workload between data centers in order to adapt with dynamic electricity price changes. We compare our algorithms with the greedy approach and show that significant cost savings can be achieved by migration of workload and dynamic deferral with future electricity price prediction. We validate our algorithms on MapReduce traces and show that geographic load balancing with dynamic deferral can provide 20-30% cost-savings.
Muhammad Abdullah Adnan, Ryo Sugihara, Rajesh K. Gupta 0001
IEEE CLOUD1
2008 Minimum Segment Drawings of Series-Parallel Graphs with the Maximum Degree Three
Md. Abul Hassan Samee, Muhammad Jawaherul Alam, Muhammad Abdullah Adnan, Md. Saidur Rahman 0001
GD3