VLDB 2026 Research / reviewers in the wild / expert
Md. Nurul Haque Mollah
dblp:96/4301
· DBLP profile ↗
8ranked-venue papers
4as first author
2since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 77% Environmental and earth informatics · 23% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics
toxicogenomics |
0.9 | 1 | 2025 | ToxAssay: a hierarchical model-driven tool for advanced toxicogenomics biomarker discovery · Bioinform. 2025 |
Environmental and earth informatics
adverse outcome pathway |
0.3 | 1 | 2025 | ToxAssay: a hierarchical model-driven tool for advanced toxicogenomics biomarker discovery · Bioinform. 2025 |
Methods — techniques the papers use, named apart from their topics
maximum likelihood estimation · 0.9hierarchical linear model · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | In-silico discovery of type-2 diabetes-causing host key genes that are associated with the complexity of monkeypox and repurposing common drugsabstractMonkeypox (Mpox) is a major global human health threat after COVID-19. Its treatment becomes complicated with type-2 diabetes (T2D). It may happen due to the influence of both disease-causing common host key genes (cHKGs). Therefore, it is necessary to explore both disease-causing cHKGs to reveal their shared pathogenetic mechanisms and candidate drugs as their common treatments without adverse side effect. This study aimed to address these issues. At first, 3 transcriptomics datasets for each of Mpox and 6 T2D datasets were analyzed and found 52 common host differentially expressed genes (cHDEGs) that can separate both T2D and Mpox patients from the control samples. Then top-ranked six cHDEGs (HSP90AA1, B2M, IGF1R, ALD1HA1, ASS1, and HADHA) were detected as the T2D-causing cHKGs that are associated with the complexity of Mpox through the protein-protein interaction network analysis. Then common pathogenetic processes between T2D and Mpox were disclosed by cHKG-set enrichment analysis with biological processes, molecular functions, cellular components and Kyoto Encyclopedia of Genes and Genomes pathways, and regulatory network analysis with transcription factors and microRNAs. Finally, cHKG-guided top-ranked three drug molecules (tecovirimat, vindoline, and brincidofovir) were recommended as the repurposable common therapeutic agents for both Mpox and T2D by molecular docking. The absorption, distribution, metabolism, excretion, and toxicity and drug-likeness analysis of these drug molecules indicated their good pharmacokinetics properties. The 100-ns molecular dynamics simulation results (root mean square deviation, root mean square fluctuation, and molecular mechanics generalized born surface area) with the top-ranked three complexes ASS1-tecovirimat, ALDH1A1-vindoline, and B2M-brincidofovir exhibited good pharmacodynamics properties. Therefore, the results provided in this article might be important resources for diagnosis and therapies of Mpox patients who are also suffering from T2D. Alvira Ajadee, Sabkat Mahmud, Md. Ahad Ali, Md. Manir Hossain Mollah, Reaz Ahmmed, Md. Nurul Haque Mollah |
Briefings Bioinform. | 6 |
| 2025 | ToxAssay: a hierarchical model-driven tool for advanced toxicogenomics biomarker discoveryabstractMOTIVATION: Understanding the genetic basis of drug-induced toxicity is crucial for drug development. In-silico analysis of toxicogenomics datasets facilitates early detection of toxicity biomarkers. However, existing tools struggle with the complex interdependencies among hierarchically structured variables, leading to inaccurate biomarker identification. To address this limitation, we developed a Hierarchical Linear Model (HLM) and implemented it in the R package ToxAssay, offering extensive functionality for comprehensive toxicity assessment. RESULTS: ToxAssay outperforms existing methods by improving biomarker detection and computational efficiency. Applied to glutathione depletion-induced toxicity, it prioritized 71 key genes and identified 26 core genes with high discriminative accuracy (AUC = 0.97) and strong cross-correlation (Pearson's r = 0.88) with external datasets. Additionally, our advance outcome pathway (AOP) analysis algorithm uncovered disease outcomes linked to glutathione depletion. These findings provide precise insights into the molecular mechanisms driving drug-induced toxicity. AVAILABILITY AND IMPLEMENTATION: ToxAssay is available as an open-source R package at https://github.com/Fun-Gene/toxassay. Md Masud Rana, Md. Nurul Haque Mollah, Mohammed H. Albujja, Sibte Hadi, Fan Liu 0004 |
Bioinform. | 2 |
| 2018 | SIPMA: A Systematic Identification of Protein-Protein Interactions in Zea mays Using Autocorrelation Features in a Machine-Learning FrameworkabstractZea mays (maize) is one of the most vital crops which are grown widely in the world. To understand the molecular structures and functions of maize, the identification of protein-protein interaction (PPI) is very important. PPI identification by wet lab experiments is time-consuming, expensive and laborious. These days in silico methods that accurately predict potential PPIs based on protein sequence information are highly demanded. Research on PPI prediction in maize is currently very limited, and no dedicated bioinformatics schemes are available. In this work, we proposed a novel approach, termed SIPMA (Systematic Identification of PPI in Maize using Autocorrelation). A machine learning random forest classifier was trained with autocorrelation features to build the prediction model. The SIPMA, which was tested by the experimentally verified PPI dataset of maize, yielded a prediction accuracy of 0.899 when the specificity was 0.969 on the training set. The SIPMA achieved promising performances on the test datasets. Compared with different sequence-based encoding and statistical learning methods, the SIPMA was a powerful computational resource for identifying PPIs in maize. Mst. Shamima Khatun, Md. Mehedi Hasan 0002, Md. Nurul Haque Mollah, Hiroyuki Kurata |
BIBE | 3 |
| 2012 | β-empirical Bayes inference and model diagnosis of microarray dataabstractBACKGROUND: Microarray data enables the high-throughput survey of mRNA expression profiles at the genomic level; however, he data presents a challenging statistical problem because of the large number of transcripts with small sample sizes that are obtained. To reduce the dimensionality, various Bayesian or empirical Bayes hierarchical models have been developed. However, because of the complexity of the microarray data, no model can explain the data fully. It is generally difficult to scrutinize the irregular patterns of expression that are not expected by the usual statistical gene by gene models. RESULTS: As an extension of empirical Bayes (EB) procedures, we have developed the β-empirical Bayes (β-EB) approach based on a β-likelihood measure which can be regarded as an 'evidence-based' weighted (quasi-) likelihood inference. The weight of a transcript t is described as a power function of its likelihood, fβ(yt|θ). Genes with low likelihoods have unexpected expression patterns and low weights. By assigning low weights to outliers, the inference becomes robust. The value of β, which controls the balance between the robustness and efficiency, is selected by maximizing the predictive β₀-likelihood by cross-validation. The proposed β-EB approach identified six significant (p<10⁻⁵) contaminated transcripts as differentially expressed (DE) in normal/tumor tissues from the head and neck of cancer patients. These six genes were all confirmed to be related to cancer; they were not identified as DE genes by the classical EB approach. When applied to the eQTL analysis of Arabidopsis thaliana, the proposed β-EB approach identified some potential master regulators that were missed by the EB approach. CONCLUSIONS: The simulation data and real gene expression data showed that the proposed β-EB method was robust against outliers. The distribution of the weights was used to scrutinize the irregular patterns of expression and diagnose the model statistically. When β-weights outside the range of the predicted distribution were observed, a detailed inspection of the data was carried out. The β-weights described here can be applied to other likelihood-based statistical models for diagnosis, and may serve as a useful tool for transcriptome and proteome studies. Mohammad Hossain Mollah, Md. Nurul Haque Mollah, Hirohisa Kishino |
BMC Bioinform. | 2 |
| 2010 | Robust extraction of local structures by the minimum beta-divergence method
Md. Nurul Haque Mollah, Nayeema Sultana, Mihoko Minami, Shinto Eguchi |
Neural Networks | 1 |
| 2008 | Robust Composite Interval Mapping for QTL Analysis by Minimum beta-Divergence MethodabstractInterval mapping (IM) is currently the most popular approach for quantitative trait loci (QTL) analysis in experimental crosses.Composite interval mapping (CIM) is a generalized version of interval mapping. However, the traditional IM and CIM approaches both are sensitive to outliers. This paper discusses a new robust CIM algorithm for QTL mapping in an experimental organisms by minimizing beta-divergence using the EM algorithm. Simulation studies show that the proposed method significantly improves the performance over the traditional CIM method for QTL mapping in presence of outliers; otherwise, it keeps equal performance. Md. Nurul Haque Mollah, Shinto Eguchi |
BIBM | 1 |
| 2007 | Robust Prewhitening for ICA by Minimizing beta-Divergence and Its Application to FastICA
Md. Nurul Haque Mollah, Shinto Eguchi, Mihoko Minami |
Neural Process. Lett. | 1 |
| 2006 | Exploring Latent Structure of Mixture ICA Models by the Minimum ß-Divergence MethodabstractIndependent component analysis (ICA) attempts to extract original independent signals (source components) that are linearly mixed in a basic framework. This letter discusses a learning algorithm for the separation of different source classes in which the observed data follow a mixture of several ICA models, where each model is described by a linear combination of independent and nongaussian sources. The proposed method is based on a sequential application of the minimum β-divergence method to separate all source classes sequentially. The proposed method searches the recovering matrix of each class on the basis of a rule of sequential change of the shifting parameter. If the initial choice of the shifting parameter vector is close to the mean of a data class, then all of the hidden sources belonging to that class are recovered properly with independent and nongaussian structure considering the data in other classes as outliers. The value of the tuning parameter β is a key in the performance of the proposed method. A cross-validation technique is proposed as an adaptive selection procedure for the tuning parameter β for this algorithm, together with applications for both real and synthetic data analysis. Md. Nurul Haque Mollah, Mihoko Minami, Shinto Eguchi |
Neural Comput. | 1 |