Adeel Malik

dblp:63/9216 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 7 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 TP-ML: A Machine-Learning-Based Tool to Identify Threonine Proteases Using Sequence-Derived Optimal Features
abstract
Threonine proteases (TPs) are enzymes vital for several biological processes and diseases including Alzheimer's disease and cancer. Their potential to target and degrade proteins intracellularly makes them valuable for various therapeutic and industrial applications. However, traditional experimental methods for identifying and characterizing novel TPs are exhaustive, time-consuming, and expensive. To address this, we developed TP-ML, a support vector machine-based prediction tool that can differentiate TP from non-TP sequences. We generated a benchmark dataset and calculated the physicochemical and compositional features using primary amino acid sequences. Subsequently, a comparison was made between the two feature selection approaches to identify the optimal feature sets from the original encodings. These optimal features were then used to train five different machine-learning classifiers, each assessed independently. TP-ML was selected as the best model showing consistent performance during cross-validation and independent evaluation, and achieved an accuracy of 0.934 and 0.888, respectively. We anticipate TP-ML to be a powerful tool for identifying TPs, aiding in their experimental characterization and industrial application exploration.
Ahmad Firoz, Adeel Malik, Nitin Mahajan, Le Thi Phan, Hani S. H. Mohammed Ali, Chang-Bae Kim, Balachandran Manavalan
IEEE Trans. Comput. Biol. Bioinform.2
2024 H2Opred: a robust and efficient hybrid deep learning model for predicting 2'-O-methylation sites in human RNA
abstract
2'-O-methylation (2OM) is the most common post-transcriptional modification of RNA. It plays a crucial role in RNA splicing, RNA stability and innate immunity. Despite advances in high-throughput detection, the chemical stability of 2OM makes it difficult to detect and map in messenger RNA. Therefore, bioinformatics tools have been developed using machine learning (ML) algorithms to identify 2OM sites. These tools have made significant progress, but their performances remain unsatisfactory and need further improvement. In this study, we introduced H2Opred, a novel hybrid deep learning (HDL) model for accurately identifying 2OM sites in human RNA. Notably, this is the first application of HDL in developing four nucleotide-specific models [adenine (A2OM), cytosine (C2OM), guanine (G2OM) and uracil (U2OM)] as well as a generic model (N2OM). H2Opred incorporated both stacked 1D convolutional neural network (1D-CNN) blocks and stacked attention-based bidirectional gated recurrent unit (Bi-GRU-Att) blocks. 1D-CNN blocks learned effective feature representations from 14 conventional descriptors, while Bi-GRU-Att blocks learned feature representations from five natural language processing-based embeddings extracted from RNA sequences. H2Opred integrated these feature representations to make the final prediction. Rigorous cross-validation analysis demonstrated that H2Opred consistently outperforms conventional ML-based single-feature models on five different datasets. Moreover, the generic model of H2Opred demonstrated a remarkable performance on both training and testing datasets, significantly outperforming the existing predictor and other four nucleotide-specific H2Opred models. To enhance accessibility and usability, we have deployed a user-friendly web server for H2Opred, accessible at https://balalab-skku.org/H2Opred/. This platform will serve as an invaluable tool for accurately predicting 2OM sites within human RNA, thereby facilitating broader applications in relevant research endeavors.
Nhat Truong Pham, Rajan Rakkiyapan, Jongsun Park 0002, Adeel Malik, Balachandran Manavalan
Briefings Bioinform.4
2023 Coded Caching in Networks With Heterogeneous User Activity
abstract
This work elevates coded caching networks from their purely information-theoretic framework to a stochastic setting, by exploring the effect of random user activity and by exploiting correlations in the activity patterns of different users. In particular, the work studies the$K$-user cache-aided broadcast channel with a limited number of cache states (i.e., the content stored at the cache of a certain user), and explores the effect of cache state association strategies in the presence of arbitrary user activity levels; a combination that strikes at the very core of the coded caching problem and its crippling subpacketization bottleneck. We first present a statistical analysis of the average worst-case delay performance of such subpacketization-constrained (state-constrained) coded caching networks, and provide computationally efficient performance bounds as well as scaling laws for any arbitrary probability distribution of the user-activity levels. The achieved performance is a result of a novel user-to-cache state association algorithm that leverages the knowledge of probabilistic user-activity levels. We then follow a data-driven approach that exploits the prior history on user-activity levels and correlations, in order to predict interference patterns, and thus better design the caching algorithm. This optimized strategy is based on the principle that users that overlap more, interfere more, and thus have higher priority to secure complementary cache states. This strategy is proven here to be within a small constant factor from the optimal. Finally, the above analysis is validated numerically using synthetic data following the Pareto principle. To the best of our understanding, this is the first work that seeks to exploit user-activity levels and correlations, in order to map future interference and design optimized coded caching algorithms that better handle this interference.
Adeel Malik, Berksan Serbetci, Petros Elia
IEEE/ACM Trans. Netw.1
2022 Stochastic Coded Caching with Optimized Shared-Cache Sizes and Reduced Subpacketization
abstract
This work studies the K-user broadcast channel with Λ caches, when the association between users and caches is random, i.e., for the scenario where each user can appear within the coverage area of – and subsequently is assisted by – a specific cache based on a given probability distribution. Caches are subject to a cumulative memory constraint that is equal to t times the size of the library. We provide a scheme that consists of three phases: the storage allocation phase, the content placement phase, and the delivery phase, and show that an optimized storage allocation across the caches together with a modified uncoded cache placement and delivery strategy alleviates the adverse effect of cache-load imbalance by significantly reducing the multiplicative performance deterioration due to randomness. In a nutshell, our work provides a scheme that manages to substantially mitigate the impact of cache-load imbalance in stochastic networks, as well as – compared to the best known state-of-the-art – the well-known subpacketization bottleneck by showing its applicability in deterministic settings for which it achieves the same delivery time – which was proven to be close to optimal for bounded values of t – with an exponential reduction in the subpacketization.
Adeel Malik, Berksan Serbetci, Petros Elia
ICC1
2022 Resolving Cache-Load Imbalance Bottleneck of Stochastic Shared-Cache Networks
abstract
This work proposes a two-layered coded caching scheme to resolve the cache-load imbalance bottleneck of the coded caching in a stochastic shared-cache network where the association between users and shared caches is random, i.e., for the scenario where each user can appear within the coverage area of – and subsequently is assisted by – a specific cache-enabled helper node based on a uniform probability distribution. To insightfully capture the effectiveness of our scheme in mitigating the adverse effect of randomness in shared-cache networks, we derive the exact scaling laws of the average delivery time. In the scenario of an error-free broadcast channel of bounded capacity per unit of time where the delivery involves K users and Λ cache-enabled helper nodes, we show that empowering users with an additional layer of caching can significantly mitigate, and in certain memory regimes completely nullify the adverse effects of the cache-load imbalance bottleneck.
Adeel Malik, Berksan Serbetci, Petros Elia
WCNC1
2021 Computational prediction and interpretation of cell-specific replication origin sites from multiple eukaryotes by exploiting stacking framework
abstract
Origins of replication sites (ORIs), which refers to the initiative locations of genomic DNA replication, play essential roles in DNA replication process. Detection of ORIs' distribution in genome scale is one of key steps to in-depth understanding their regulation mechanisms. In this study, we presented a novel machine learning-based approach called Stack-ORI encompassing 10 cell-specific prediction models for identifying ORIs from four different eukaryotic species (Homo sapiens, Mus musculus, Drosophila melanogaster and Arabidopsis thaliana). For each cell-specific model, we employed 12 feature encoding schemes that cover nucleic acid composition, position-specific and physicochemical properties information. The optimal feature set was identified from each encoding individually and developed their respective baseline models using the eXtreme Gradient Boosting (XGBoost) classifier. Subsequently, the predicted scores of 12 baseline models are integrated as a novel feature vector to train XGBoost and develop the final model. Extensive experimental results show that Stack-ORI achieves significantly better performance as compared with their baseline models on both training and independent datasets. Interestingly, Stack-ORI consistently outperforms existing predictor in all cell-specific models, not only on training but also on independent test. Moreover, our novel approach provides necessary interpretations that help understanding model success by leveraging the powerful SHapley Additive exPlanation algorithm, thus underlining the most important feature encoding schemes significant for predicting cell-specific ORIs.
Leyi Wei, Adeel Malik, Ran Su, Li-Zhen Cui 0001, Balachandran Manavalan
Briefings Bioinform.3
2021 Fundamental Limits of Stochastic Shared-Cache Networks
abstract
The work establishes the exact performance limits of stochastic coded caching when users share a bounded number of cache states, and when the association between users and caches, is random. Under the premise that more balanced user-to-cache associations perform better than unbalanced ones, our work provides a statistical analysis of the average performance of such networks, identifying in closed form, the exact optimal average delivery time. To insightfully capture this delay, we derive easy-to-compute closed-form analytical bounds that prove tight in the limit of a large number Λ of cache states. In the scenario where delivery involves K users, we conclude that the multiplicative performance deterioration due to randomness - as compared to the well-known deterministic uniform case - can be unbounded and can scale as Θ([log Λ]/[log log Λ]) at K = Θ(Λ ), and that this scaling vanishes when K = Ω(Λ log Λ ). To alleviate this adverse effect of cache-load imbalance, we consider various load-balancing methods, and show that employing proximity-bounded load balancing with an ability to choose from h neighboring caches, the aforementioned scaling reduces to Θ([log(Λ/h)]/[log log(Λ/h)]) at K=Θ(Λ ), while when the proximity constraint is removed, the scaling is of a much slower order Θ(log log Λ ). The above analysis is extensively validated numerically.
Adeel Malik, Berksan Serbetci, Emanuele Parrinello, Petros Elia
IEEE Trans. Commun.1
2021 A Personalized Preference Learning Framework for Caching in Mobile Networks
abstract
This paper comprehensively studies a content-centric mobile network based on a preference learning framework, where each mobile user is equipped with a finite-size cache. We consider a practical scenario where each user requests a content file according to its own preferences, which is motivated by the existence of heterogeneity in file preferences among different users. Under our model, we consider a single-hop-based device-to-device (D2D) content delivery protocol and characterize the average hit ratio for the following two file preference cases: the personalized file preferences and the common file preferences. By assuming that the model parameters such as user activity levels, user file preferences, and file popularity are unknown and thus need to be inferred, we present a collaborative filtering (CF)-based approach to learn these parameters. Then, we reformulate the hit ratio maximization problems into a submodular function maximization and propose two computationally efficient algorithms including a greedy approach to efficiently solve the cache allocation problems. We analyze the computational complexity of each algorithm. Moreover, we analyze the corresponding level of the approximation that our greedy algorithm can achieve compared to the optimal solution. Using a real-world dataset, we demonstrate that the proposed framework employing the personalized file preferences brings substantial gains over its counterpart for various system parameters.
Adeel Malik, Joongheon Kim, Kwang Soon Kim, Won-Yong Shin
IEEE Trans. Mob. Comput.1
2020 Stochastic Analysis of Coded Multicasting for Shared Caches Networks
abstract
The work establishes the exact fundamental limits of stochastic coded caching when users share a bounded number of cache states, and when the association between users and caches, is random. This association can greatly affect performance, which improves when the association is more balanced across the caches, and which deteriorates when this association becomes less uniform. Our work provides a statistical analysis of the average performance of such networks, quantifying the effect of randomness by identifying in closed-form, the exact optimal average delivery time. To insightfully capture this delay, we derive the exact scaling laws of the optimal average delivery time. In the scenario where delivery involves K users, we conclude that the multiplicative performance deterioration due to randomness - as compared to the well-known deterministic uniform case - can be unbounded and can scale as Θ([(logΛ )/(loglogΛ )]) at K=Θ(Λ), and that as K increases, this deterioration gradually reduces, and ceases to scale when K=Ω(ΛlogΛ). The above analysis is validated numerically.
Adeel Malik, Berksan Serbetci, Emanuele Parrinello, Petros Elia
GLOBECOM1
2018 On the Effects of Subpacketization in Content-Centric Mobile Networks
abstract
A large-scale content-centric mobile ad hoc network employing subpacketization is studied in which each mobile node having finite-size cache moves according to the reshuffling mobility model and requests a content object from the library independently at random according to the Zipf popularity distribution. Instead of assuming that one content object is transferred in a single time slot, we consider a more challenging scenario where the size of each content object is considerably large and thus only a subpacket of a file can be delivered during one time slot, which is motivated by a fast mobility scenario. Under our mobility model, we consider a single-hop-based content delivery and characterize the fundamental tradeoffs between throughput and delay. The order-optimal throughput-delay tradeoff is analyzed by presenting the following two content reception strategies: the sequential reception for uncoded caching and the random reception for maximum distance separable (MDS)-coded caching. We also perform numerical evaluation to validate our analytical results. In particular, we conduct performance comparisons between the uncoded caching and the MDS-coded caching strategies by identifying the regimes in which the performance difference between the two caching strategies becomes prominent with respect to system parameters such as the Zipf exponent and the number of subpackets. In addition, we extend our study to the random walk mobility scenario and show that our main results are essentially the same as those in the reshuffling mobility model.
Adeel Malik, Sung Hoon Lim, Won-Yong Shin
IEEE J. Sel. Areas Commun.1
2017 On expected transmissions for wireless random linear coding
abstract
In this paper we revisit the problem of deriving the expected number of transmissions for multicasting random linear coded (RLC) packets on single-hop wireless channels. We show by deriving the closed form expression for an instance of the problem that previous analytical formulation does not accurately model the true expected number of transmissions, especially for smaller finite field size. Our understanding is that similar to the wireless multi-hop network, the problem of deriving an exact closed form expression for the expected number of transmissions for a single-hop RLC wireless network is a complex open problem. As it is unknown whether a scalable closed form expression for the problem exists, we then propose a computationally efficient Monte Carlo method to derive a good approximation of the expected number of transmissions.
Jalaluddin Qureshi, Adeel Malik, Chuan Heng Foh
ISCC2