VLDB 2026 Research / reviewers in the wild / expert
Shyam Sundar Das
dblp:337/4245
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2023
0009-0004-9413-1753ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Prediction of DILI in Humans: Investigating Opportunities in DILIst Data, Chemical Features, Machine Learning and Rule-Based MethodsabstractDespite recent scientific developments, combating Drug Induced Liver Injury (DILI) associated risks is still a major challenge and there is an urgent need for breakthroughs in the assessment of DILI. To address this, we have built and evaluated several Machine Learning (ML) and rule-based models. Our approach is focused on three aspects: i) selection of three new high-quality datasets, DILI-876, DILI-955, and DILI-1148 from DILIst (DILI severity and toxicity) dataset, by applying several scientific criteria. The datasets DILI-876, DILI-955, and DILI-1148 contain drugs that fall within three chemical spaces, namely, drugs with molecular weights of ⩽500, ⩽ 750, and "no molecular weight bias" respectively. ii) Evaluation of the influence of three types of data division methods and chemical features on model performance. iii) Evaluation of four ML and one rule-based method for developing reliable DILI models. These efforts resulted in high-performing DILI models that are superior to reported ML models, based on AUC, F1, specificity, and most significantly, with superior performance in the case of DILI positive and negative drugs. The test set characteristics of the best ANN DILI model developed using DILI-876 are AUC-0.908, accuracy-0.804, sensitivity-0.804, specificity-0.815, balanced accuracy-0.809, MCC-0.594, precision-0.821, and F1-0.807. The best SVM DILI models showed similar performance characteristics. The order of predictive performance of the ML models derived from the three datasets is DILI-876>DILI-955>DILI-1148, and the order of performance of the ML methods is ANN~SVM>LR>RF. The order of performance of the models built using the three chemical features is fingerprints>descriptors>fragments. From the performance analysis of our models for individual chemical and DILI classes, we infer that our models performed equally well for DILI negative drugs e.g., the success level of prediction is 81.355% for "vNo" drugs, as reflected by the excellent specificity of the models for the imbalanced DILI data. Similarly, they performed equally well in the case of DILI positives, as reflected by the excellent sensitivity of our models. It is evident from the success level of prediction of "vMost" (81.818%), "vLess" (90.163%), and "others" (73.333%), DILI concerns, that our models are excellent candidates for easy adoption in real-world applications. Advaith Nila Narayanan, Shyam Sundar Das, T. T. Mirnalinee |
BIBM | 2 |
| 2023 | Evaluation of AutoML Frameworks for Computational ADMET Screening in Drug Discovery & DevelopmentabstractThe goal of this paper is to demonstrate the power of AutoML frameworks for computational ADMET screening in drug discovery and development. For this purpose, we evaluated i) five AutoML frameworks to automate ML tasks using two ADMET datasets, drug induced liver injury (DILI) and hERG inhibition. ii) The performance of manually generated ML and AutoML models, to assess the utility of AutoML models for real world applications. iii) The advantages and disadvantages based on the speed of optimization, robustness, and application domain of the AutoML models. It is observed that i) the performance of manual and AutoML models is found to be very close, in both positive and negative classes. ii) Among the five AutoML frameworks, AutoGluon is found to be the best, in terms of time, robustness, performance etc. iii) However, for easy adoption of AutoML frameworks for real-world applications, one may require a) the selection of high-quality datasets based on scientific context and b) the use of appropriate feature reduction and selection techniques, currently unavailable in many AutoML frameworks. Advaith Nila Narayanan, Shyam Sundar Das, T. T. Mirnalinee |
BIBM | 2 |
| 2022 | Hybrid Modified League Championship Algorithm and its Applications in Compartmental Pharmacokinetic-Pharmacodynamic Data AnalysisabstractWe present herein, the application of a hybrid modified league championship (HMLCA) algorithm and its applications for the reliable estimation of Compartmental PKPD parameters, for the first time in literature. For this purpose, we have modified the original League Championship algorithm, combined it with local optimization methods such as Nelder-Mead and Gauss-Newton, such that the final compartmental PK-PD parameters are estimated in a single step, to facilitate ease of analysis and to avoid laborious procedures. In HMCLA, the initial PK-PD parameters are estimated by modified LCA and then, they are optimized by the local optimization methods. We have studied the performance of the new approach using thirty-seven datasets that comprises of real-world PK-PD datasets fitted with 19 algebraic equations PK-PD models, 18 differential equations PK-PD models. These datasets represent different studies performed during pre-clinical and Phase I clinical stages of drug discovery and development. The performance of our models measured in terms weighted residual sum of squares (WRSS) is found to be better than the original LCA. Our approach eliminates the laborious procedure of estimating initial parameters and the need for domain expertise, thus it is a potential approach for automated compartmental PK-PD analysis Shyam Sundar Das, Geervani Koneti, Narayanan Ramamurthi |
BIBM | 1 |
| 2022 | Discovering the Knowledge in Unstructured Early Drug Development Data Using NLP and Advanced AnalyticsabstractDiscovering the knowledge in unstructured archived data is an area of active interest in academia and industry, as it offers opportunities to learn, confirm and eventually, address R&D productivity challenges. Our interest in this area prompted us to investigate an Natural Language Processing based approach to extract unstructured Pharmacokinetics (PK) and Pharmacodynamics (PD) data from PK-PD study reports, and perform analytics using in-house developed compartmental and non-compartmental analytics engines. For this purpose, we have developed a dictionary of two thousand twenty-one (2321) PK-PD keywords based on published study reports. Details of our approach and its applications in discovering the knowledge in the unstructured archived early drug development reports, is the subject of this paper. Geervani Koneti, Shyam Sundar Das, Jyotsna Bahl, Pritish Ranjan, Narayanan Ramamurthi |
BIBM | 2 |
| 2022 | Parallelized Population Based Exhaustive Replacement Method for Reproducible Features Selection and Its Applications in Drug DiscoveryabstractWe present herein, the parallelized Population based Exhaustive Replacement Method (PERM) and its applications for the reproducible selection of chemical features with improved exploration of search space. The improved performance of parallelized PERM over the earlier reports can be attributed to its ability to a) minimize local optimization and randomness of the method and b) effective exploration of search space. Consequently, it functions like a deterministic method with enhanced reproducibility and reliability. The application of parallelized PERM for the generation of reproducible predictive models is demonstrated using Human Serum Albumin and EGFR Kinase inhibitors data sets. Geervani Koneti, Mastan Vali Shaik, Shyam Sundar Das, Narayanan Ramamurthi |
BIBM | 3 |