VLDB 2026 Research / reviewers in the wild / expert
Madhu Chetty
dblp:81/4612
· DBLP profile ↗
75ranked-venue papers
1as first author
11since 2021 · last 2024
0000-0001-7052-0413ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GRAMP: A gene ranking and model prioritisation framework for building consensus genetic networks
Hasini Nakulugamuwa Gamage, Madhu Chetty, Suryani Lim, Jennifer Hallinan |
Knowl. Based Syst. | 2 |
| 2024 | PANDORA: Deep Graph Learning Based COVID-19 Infection Risk Level ForecastingabstractCoronavirus disease 2019 (COVID-19) as a global pandemic causes a massive disruption to social stability that threatens human life and the economy. An effective forecasting system is arguably important to provide an early signal of the risk of COVID-19 infection so that the authorities are ready to protect the people from the worst. However, making a good forecasting model for infection risks in different cities or regions is not an easy task, because it has a lot of influential factors that are difficult to be identified manually. To address the current limitations, we propose a deep graph learning model, called PANDORA, to predict the infection risks of COVID-19, by considering all essential factors and integrating them into a geographical network. The framework uses geographical position relationships and transportation frequency as higher order structural properties formulated by higher order network structures (i.e., network motifs). Moreover, four significant node attributes (i.e., multiple features of a particular area, including climate, medical condition, economy, and human mobility) are also considered. We propose three different aggregators to better aggregate node attributes and structural features, namely, Hadamard, Summation, and Connection. Experimental results over real data show that PANDORA outperforms the baseline methods with higher accuracy and faster convergence speed, no matter which aggregator is chosen. Shuo Yu 0001, Feng Xia 0001, Yueru Wang, Falih Febrinanto, Madhu Chetty |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2023 | A Robust Ensemble Regression Model for Reconstructing Genetic NetworksabstractGenetic networks contain important information about biological processes, including regulatory relationships and gene-gene interactions. Numerous methods, using high-dimensional gene expression data have been developed to capture these interactions. These gene expression data, generated using high-throughput technologies, are prone to noise. However, most existing network inference methods are unable to cope with noisy data, making genetic network reconstruction challenging. In this paper, we propose a novel ensemble regression model combining quantile regression and cross-validated Ridge regression, RidgeCV, to infer interactions from noisy gene expression data. The application of quantile regression to GRN inference is novel, and its design makes it appropriate for noisy data. RidgeCV also addresses other important issues, such as data overfitting and multicollinearity. First, each regression method is independently applied to gene expression data and the output of these methods, in the form of ranked gene lists, is aggregated using a novel gene score-based method by considering the gene rank and model importance. The model importance score is evaluated based on an adjusted coefficient of determination. This method implicitly includes majority voting by averaging each gene score value across all models. The proposed model was tested on the DREAM4 datasets and publicly available small-scale real-world network datasets. Experiments with noisy datasets showed that the proposed ensemble model is more accurate and efficient than other state-of-the-art methods. Hasini Nakulugamuwa Gamage, Madhu Chetty, Suryani Lim, Jennifer Hallinan |
IJCNN | 2 |
| 2023 | User authentication and access control to blockchain-based forensic log dataabstractAbstract For dispute resolution in daily life, tamper-proof data storage and retrieval of log data are important with the incorporation of trustworthy access control for the related users and devices, while giving access to confidential data to the relevant users and maintaining data persistency are two major challenges in information security. This research uses blockchain data structure to maintain data persistency. On the other hand, we propose protocols for the authentication of users (persons and devices) to edge server and edge server to main server. Our proposed framework also provides access to forensic users according to their relevant roles and privilege attributes. For the access control of forensic users, a hybrid attribute and role-based access control (ARBAC) module added with the framework. The proposed framework is composed of an immutable blockchain-based data storage with endpoint authentication and attribute role-based user access control system. We simulate authentication protocols of the framework in AVISPA. Our result analysis shows that several security issues can efficiently be dealt with by the proposed framework. Md. Ezazul Islam, Md. Rafiqul Islam 0002, Madhu Chetty, Suryani Lim, Mehmood A. Chadhar |
EURASIP J. Inf. Secur. | 3 |
| 2023 | Meaning-Sensitive Text Data Augmentation with Intelligent MaskingabstractWith the recent popularity of applying large-scale deep neural network-based models for natural language processing (NLP), attention to develop methods for text data augmentation is at its peak, since the limited size of training data tends to significantly affect the accuracy of these models. To this end, we propose a novel text data augmentation technique called Intelligent Masking with Optimal Substitutions Text Data Augmentation (IMOSA). IMOSA, developed for labelled sentences, can identify the most favourable sentences and locate the appropriate word combinations in a particular sentence to replace and generate synthetic sentences with a meaning closer to the original sentence, while also significantly increasing the diversity of the dataset. We demonstrate that the proposed technique notably improves the performance of classifiers based on attention-based transformer models through the extensive experiments for five different text classification tasks which are performed under the low data regime in a context-aware NLP setting. The analysis clearly shows that IMOSA effectively generates more sentences using favourable original examples and completely ignores undesirable examples. Furthermore, the experiments carried out confirm IMOSA’s ability to add diversity to the augmented dataset using multiple distinct masking patterns against the same original sentence, which remarkably adds variety to the training dataset. IMOSA consistently outperforms the two key masked language model-based text data augmentation techniques, and demonstrates a robust performance against the critical challenging NLP tasks. Buddhika Kasthuriarachchy, Madhu Chetty, Adrian Shatte, Darren Walls |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2022 | Ensemble Regression Modelling for Genetic Network InferenceabstractAn accurate reconstruction of Gene Regulatory Networks (GRNs) from time series gene expression data is crucial for discovering complex biological interactions. Among many different approaches for inferring GRNs, there are several methods which produce high false positive interactions, and are unstable, requiring fine tuning for many of their parameters. In this paper, we consider the GRN inference problem as a regression problem, and propose a simple ensemble regression-based feature selection model which is a combination of cross-validated Lasso and cross-validated Ridge algorithms for reconstructing GRNs. Due to the novelty of the proposed ensemble model, it is able to eliminate overfitting, multi co-linearity issues, and irrelevant genes within one computational approach. While observing the type of gene-gene regulatory interactions the regression model also identifies the direction of these interactions. A new coefficient of determination (R2)-based approach identifies the best model to fit the data among LassoCV and RidgeCV, and evaluates the model importance in term of gene-wise maximum in-degree which decides the maximum number of regulatory genes including self-regulations that can be selected from a given method. Then, an evaluated gene score-based majority voting technique aggregates the selected gene lists from each method. In our experiments, the performance of the proposed ensemble approach was evaluated using gene expression datasets from three small-scale real gene networks. Our proposed model outperformed other state-of-the-art methods, producing high true positives, reducing false positives, and obtaining high Structural Accuracy, while maintaining model stability and efficiency. Hasini Nakulugamuwa Gamage, Madhu Chetty, Adrian Shatte, Jennifer Hallinan |
CIBCB | 2 |
| 2022 | Integrating steady-state and dynamic gene expression data for improving genetic network modellingabstractReverse engineering of Gene Regulatory Networks (GRNs) from experimentally obtained high-throughput data is an active and promising area of research. Among several modelling techniques, the S-System model, a set of tightly coupled differential equations, mimics the complexities and dynamics of biochemical systems, and thus provides realistic GRN representation. While it offers mathematical flexibility and biological relevance, the high number of learning parameters can lead to a computational burden. In our earlier work, we addressed this issue by judicious use of prior knowledge. However, another major cause of computational load is the need for numerical integration of the differential equations for the estimation of S-system model parameters. In this paper, we propose a method to obtain initial model parameter values from the steady state of the system, thereby computing simpler and less complex algebraic equations compared to the regular differential equations of S-systems. These network parameters are input as prior knowledge for the optimization of the dynamic S-System using differential equations. The proposed framework includes a novel fitness evaluation for steady-state S-System models, a novel evolutionary parameter learning framework, and a technique to incorporate the candidate solutions in dynamic S-System modelling. Our proposed methodology reached optimal model parameter values quickly, requiring only one-third of the fitness function evaluations, compared to our previously reported DRNI (Dynamically regulated network initialization) method for S-System modelling. Jaskaran Gill, Madhu Chetty, Adrian Shatte, Jennifer Hallinan |
CIBCB | 2 |
| 2021 | An Efficient Boolean Modelling Approach for Genetic Network InferenceabstractThe inference of Gene Regulatory Networks (GRNs) from time series gene expression data is an effective approach for unveiling important underlying gene-gene relationships and dynamics. While various computational models exist for accurate inference of GRNs, many are computationally inefficient, and do not focus on simultaneous inference of both network topology and dynamics. In this paper, we introduce a simple, Boolean network model-based solution for efficient inference of GRNs. First, the microarray expression data are discretized using the average gene expression value as a threshold. This step permits an experimental approach of defining the maximum indegree of a network. Next, regulatory genes, including the self-regulations for each target gene, are inferred using estimated multivariate mutual information-based Min-Redundancy Max-Relevance Criterion, and further accurate inference is performed by a swapping operation. Subsequently, we introduce a new method, combining Boolean network regulation modelling and Pearson correlation coefficient to identify the interaction types (inhibition or activation) of the regulatory genes. This method is utilized for the efficient determination of the optimal regulatory rule, consisting AND, OR, and NOT operators, by defining the accurate application of the NOT operation in conjunction and disjunction Boolean functions. The proposed approach is evaluated using two real gene expression datasets for an Escherichia coli gene regulatory network and a fission yeast cell cycle network. Although the Structural Accuracy is approximately the same as existing methods (MIBNI, REVEAL, Best-Fit, BIBN, and CST), the proposed method outperforms all these methods with respect to efficiency and Dynamic Accuracy. Hasini Nakulugamuwa Gamage, Madhu Chetty, Adrian Shatte, Jennifer Hallinan |
CIBCB | 2 |
| 2021 | Dynamically Regulated Initialization for S-system Modelling of Genetic NetworksabstractReverse engineering of gene regulatory networks through temporal gene expression data is an active area of research. Among the plethora of modelling techniques under investigation is the decoupled S-system model, which attempts to capture the non-linearity of biological systems in detail. For the model, number of parameters to be estimated are significantly high even when the network is of small or medium scale. Thus, the inference process poses a significant computational burden. In this paper, we propose: (1) a novel population initialization technique, Dynamically Regulated Prediction Initialization (DRPI), which utilises prior knowledge of biological gene expression data to create a feedback loop to produce dynamically regulated high-quality individuals for initial population; (2) an adaptive fitness function; and (3) a method for the maintenance of population diversity. The aim of this work is to reduce the computational complexity of the inference algorithm, to speed up the entire process of reverse engineering. The performance of the proposed algorithm was evaluated against a benchmark dataset and compared with other methods from earlier work. The experimental results show that we succeeded in achieving higher accuracy results in lesser fitness evaluations, considerably reducing the computational burden of the inference process. Jaskaran Gill, Madhu Chetty, Adrian Shatte, Jennifer Hallinan |
CIBCB | 2 |
| 2021 | Cost Effective Annotation Framework Using Zero-Shot Text ClassificationabstractManual and high-quality annotation of social media data has enabled companies and researchers to develop improved implementations using natural language processing. However, human text-annotation is expensive and time-consuming. Crowd-sourcing platforms such as Amazon's Mechanical Turk (MTurk) can be leveraged for the creation of large training corpora for text classification tasks using social media data. Nevertheless, the quality of annotations can vary significantly, based on the interpretations and motivations of annotators completing the tasks. Further, the labelling cost of data through MTurk will increase if target messages are small and having a significant amount of noise (e.g. promotional messages on Twitter). In this work, we propose a new annotation framework to create high-quality human-annotated datasets for text classification from social media data. We present a zero-shot text classification based pre-annotation technique reducing the adverse effects arising due to the highly skewed distribution of data across target classes. The proposed framework significantly reduces the cost and time while maintaining the quality of the annotations. Being generic, it can be applied to annotating text data from any discipline. Our experiment with a Twitter data annotation using the proposed annotation framework shows a cost reduction of 80% with no compromise to quality. Buddhika Kasthuriarachchy, Madhu Chetty, Adrian Shatte, Darren Walls |
IJCNN | 2 |
| 2021 | An improved memetic approach for protein structure prediction incorporating maximal hydrophobic core estimation concept
Rumana Nazmul, Madhu Chetty, Ahsan Raja Chowdhury |
Knowl. Based Syst. | 2 |
| 2020 | Pre-trained Language Models with Limited Data for Intent ClassificationabstractIntent analysis is capturing the attention of both the industry and academia due to its commercial and noncommercial significance. The rapid growth of unstructured data of micro-blogging platforms, such as Twitter and Facebook, are amongst the important sources for intent analysis. However, the social media data are often noisy and diverse, thus making the task very challenging. Further, the intent analysis frequently suffers from lack of sufficient data because the labeled datasets are often manually annotated. Recently, BERT (Bidirectional Encoder Representation from Transformers), a state-of-the-art language representation model, has attracted attention for accurate language modelling. In this paper, we investigate the application of BERT for its suitability for intent analysis. We study the fine-tuning of the BERT model through inductive transfer learning and investigate methods to overcome the challenges due to limited data availability by proposing a novel semantic data augmentation approach. This technique generates synthetic sentences while preserving the label-compatibility using the semantic meaning of the sentences, to improve the intent classification accuracy. Thus, based on the considerations for finetuning and data augmentation, a systematic and novel step-by-step methodology is presented for applying the linguistic model BERT for intent classification with limited data available. Our results show that the pre-trained language can be effectively used with noisy social media data to achieve state-of-the-art accuracy in intent analysis under low labeled-data regime. Moreover, our results also confirm that the proposed text augmentation technique is effective in eliminating noisy synthetic sentences, thereby achieving further performance improvements. Buddhika Kasthuriarachchy, Madhu Chetty, Gour C. Karmakar, Darren Walls |
IJCNN | 2 |
| 2018 | Relevance of Frequency of Heart-Rate Peaks as Indicator of 'Biological' Stress Level
Meena Santhanagopalan, Madhu Chetty, Cameron Foale, Sunil Aryal, Britt Klein |
ICONIP (7) | 2 |
| 2017 | Special issue on multi-objective reinforcement learning
Madalina M. Drugan, Marco A. Wiering, Peter Vamplew 0001, Madhu Chetty |
Neurocomputing | 4 |
| 2016 | Exploiting Temporal Genetic Correlations for Enhancing Regulatory Network Optimization
Ahammed Sherief Kizhakkethil Youseph, Madhu Chetty, Gour C. Karmakar |
ICONIP (1) | 2 |
| 2015 | Gene regulatory network inference using Michaelis-Menten kineticsabstractA gene regulatory network (GRN) represents a collection of genes, connected via regulatory interactions. Reverse engineering GRNs is a challenging problem in systems biology. Various models have been proposed for modeling GRNs. However, many of these models lack the capability to explain the molecular mechanisms underlying the biological process. Michaelis-Menten kinetics can be used to model the biomolecular mechanisms and is a widely used non-linear approach to represent biochemical systems. However, the model in its current form is not suitable for reverse engineering biological systems. In this paper, based on Michaelis-Menten kinetics, we develop a new model to reverse engineer GRNs. The parameter estimation is formulated as an optimization problem which is solved by adapting trigonometric differential evolution (TDE), a variant of differential evolution (DE). The model is applied for reconstructing both in silico and in vivo networks. The results are promising and as the model is fully biologically relevant, it provides a new perspective for accurate GRN inference. Ahammed Sherief Kizhakkethil Youseph, Madhu Chetty, Gour C. Karmakar |
CEC | 2 |
| 2015 | Towards large scale genetic network modelingabstractReverse Engineering Gene Regulatory Networks (GRNs) is an important and challenging problem of Systems Biology. For its superiority in both structure and parameter learning, the S-system model framework is often chosen for GRN reconstruction. The biggest challenge in reconstructing GRNs is the data having large number of genes and only a small number of samples. This “curse of dimensionality”, along with the large number of model parameters to be learnt, makes it extremely difficult to reverse engineer even a small network. For a medium or large network, the complexity becomes enormous. In this paper, we propose a method for managing large scale GRN modeling. As first step, we propose an Affinity Propagation Based Clustering to identify appropriate clusters by grouping the genes based on their time expression profiles. In the second step, the largest cluster consisting of majority of the relevant genes is considered in full detail to act as the core of the network while the other remaining clusters, which are not so significant, are each represented by their single representative gene to obtain a reduced order GRN. In the third step, we optimize the entire network by initializing the model parameters of the genes of the largest cluster with the values obtained in the second step (which are near optimal) and proceed to optimize the entire network. The initial investigations are carried out using previously reported 20-gene synthetic network. The superiority of performance is evaluated not only using the standard metrics, namely, sensitivity, specificity, precision and F-score, but also by average mean error and by comparing the time responses with those of the actual network parameters. The results obtained are promising. Rubaiya Rahtin Khan, Madhu Chetty |
CIBCB | 2 |
| 2015 | Frequency Decomposition Based Gene Clustering
Mohamed Abdur Rahman 0002, Madhu Chetty, Dieter Bulach, Pramod P. Wangikar |
ICONIP (2) | 2 |
| 2015 | Decoupled Modeling of Gene Regulatory Networks Using Michaelis-Menten Kinetics
Ahammed Sherief Kizhakkethil Youseph, Madhu Chetty, Gour C. Karmakar |
ICONIP (3) | 2 |
| 2015 | Network decomposition based large-scale reverse engineering of gene regulatory network
Ahsan Raja Chowdhury, Madhu Chetty |
Neurocomputing | 2 |
| 2014 | Significance of Non-edge Priors in Gene Regulatory Network Reconstruction
Ajay Nair, Madhu Chetty, Pramod P. Wangikar |
ICONIP (1) | 2 |
| 2014 | Sib-Based Survival Selection Technique for Protein Structure Prediction in 3D-FCC Lattice Model
Rumana Nazmul, Madhu Chetty |
ICONIP (2) | 2 |
| 2013 | An Adaptive Strategy for Assortative Mating in Genetic AlgorithmabstractIn any traditional Genetic Algorithm (GA), recombination is a dominant search operator and capable of exploring the search space by sharing genetic information among the individuals in the population. However, a simple application of recombination alone is insufficient to guide convergence to an optimal solution. The selection of parents for recombination operation has a significant role in guiding the evolution towards the optimal solution and also for maintaining genetic diversity to avoid getting trapped in local minima. A non-random mating mimics the mechanism of reproduction in nature and is effective in maintaining diversity in population. This paper proposes a new strategy for selection of mating pairs based on a type of non-random mating called as assortative mating. The proposed mate selection scheme conserves the merits of both positive and negative assortative mating in a controlled manner by allowing mating between individuals having both similar and dissimilar phenotypes. For effective cross-over, it maintains genetic diversity in population by distributing the recombination among dissimilar individuals. Furthermore, it ensures the preservation and propagation of useful genetic information to the later stages of search by the selection of mates having similar phenotypes. Experimental results, using not only the five widely used benchmark functions but also twenty newly developed modified functions, are reported. The results show significant improvements in the convergence characteristics of the proposed mating strategy over existing nonrandom mating techniques. Rumana Nazmul, Madhu Chetty |
IEEE Congress on Evolutionary Computation | 2 |
| 2013 | Inferring large scale genetic networks with S-system modelabstractGene regulatory network (GRN) reconstruction from high-throughput microarray data is an important problem in systems biology. The S-System model, a differential equation based approach, is among the mainstream approaches for modeling GRNs. It has the ability to represent GRNs accurately with precise regulatory weights. However, the current applications of S-System are limited to small and medium scale network, as inferring large network requires inhibitive computational cost. In this paper, we propose a novel S-System based framework to reconstruct biologically relevant GRNs by exploiting their special topological structure. In GRNs, the complex interactions occurring amongst transcription factors (TFs) and target genes (TGs) are unidirectional, i.e., TFs to TGs, and the vice-versa is biologically irrelevant. In addition, TFs can regulate themselves while only self-regulations may exist for TGs. As such, we decompose GRN into two sub-networks representing TF-TF and TF-TG interactions. We learn the sub-networks separately by adapting the traditional S-System model, and combining the solutions to get the entire network. Our experimental studies indicate that the proposed approach can scale up to larger networks, not achievable with other current S-System based approaches, yet with higher accuracy. Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
GECCO | 2 |
| 2013 | mDBN: motif based learning of gene regulatory networks using dynamic bayesian networksabstractSolutions for deriving the most consistent Bayesian gene regulatory network model from given data sets using evolutionary algorithms typically only result in locally optimal solutions. Further, due to genetic drift, merely increasing the size of the population does not overcome this limitation. In this paper, we propose a two-stage genetic algorithm that systematically searches the whole search space using frequent subgraph mining techniques. The approach finds representative patterns present in different local optimal solutions in the first stage and then combines these frequent subgraphs (motifs) in the second stage to converge to the global optima. We apply the algorithm to both synthetic and real life networks of yeast and E.coli and show the effectiveness of our approach. Nizamul Morshed, Madhu Chetty, Xuan Vinh Nguyen, Terry Caelli |
GECCO | 2 |
| 2013 | On the Analysis of Time-Delayed Interactions in Genetic Network Using S-System Model
Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (2) | 2 |
| 2013 | Reverse Engineering Genetic Networks with Time-Delayed S-System Model and Pearson Correlation Coefficient
Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (2) | 2 |
| 2013 | A Knowledge-Based Initial Population Generation in Memetic Algorithm for Protein Structure Prediction
Rumana Nazmul, Madhu Chetty |
ICONIP (2) | 2 |
| 2013 | Protein Structure Prediction with a New Composite Measure of Diversity and Memory-Based Diversification Strategy
Rumana Nazmul, Madhu Chetty |
ICONIP (2) | 2 |
| 2013 | Incorporating time-delays in S-System model for reverse engineering genetic networksabstractBACKGROUND: In any gene regulatory network (GRN), the complex interactions occurring amongst transcription factors and target genes can be either instantaneous or time-delayed. However, many existing modeling approaches currently applied for inferring GRNs are unable to represent both these interactions simultaneously. As a result, all these approaches cannot detect important interactions of the other type. S-System model, a differential equation based approach which has been increasingly applied for modeling GRNs, also suffers from this limitation. In fact, all S-System based existing modeling approaches have been designed to capture only instantaneous interactions, and are unable to infer time-delayed interactions. RESULTS: In this paper, we propose a novel Time-Delayed S-System (TDSS) model which uses a set of delay differential equations to represent the system dynamics. The ability to incorporate time-delay parameters in the proposed S-System model enables simultaneous modeling of both instantaneous and time-delayed interactions. Furthermore, the delay parameters are not limited to just positive integer values (corresponding to time stamps in the data), but can also take fractional values. Moreover, we also propose a new criterion for model evaluation exploiting the sparse and scale-free nature of GRNs to effectively narrow down the search space, which not only reduces the computation time significantly but also improves model accuracy. The evaluation criterion systematically adapts the max-min in-degrees and also systematically balances the effect of network accuracy and complexity during optimization. CONCLUSION: The four well-known performance measures applied to the experimental studies on synthetic networks with various time-delayed regulations clearly demonstrate that the proposed method can capture both instantaneous and delayed interactions correctly with high precision. The experiments carried out on two well-known real-life networks, namely IRMA and SOS DNA repair network in Escherichia coli show a significant improvement compared with other state-of-the-art approaches for GRN modeling. Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
BMC Bioinform. | 2 |
| 2013 | A model of the circadian clock in the cyanobacterium Cyanothece sp. ATCC 51142abstractBACKGROUND: The over consumption of fossil fuels has led to growing concerns over climate change and global warming. Increasing research activities have been carried out towards alternative viable biofuel sources. Of several different biofuel platforms, cyanobacteria possess great potential, for their ability to accumulate biomass tens of times faster than traditional oilseed crops. The cyanobacterium Cyanothece sp. ATCC 51142 has recently attracted lots of research interest as a model organism for such research. Cyanothece can perform efficiently both photosynthesis and nitrogen fixation within the same cell, and has been recently shown to produce biohydrogen--a byproduct of nitrogen fixation--at very high rates of several folds higher than previously described hydrogen-producing photosynthetic microbes. Since the key enzyme for nitrogen fixation is very sensitive to oxygen produced by photosynthesis, Cyanothece employs a sophisticated temporal separation scheme, where nitrogen fixation occurs at night and photosynthesis at day. At the core of this temporal separation scheme is a robust clocking mechanism, which so far has not been thoroughly studied. Understanding how this circadian clock interacts with and harmonizes global transcription of key cellular processes is one of the keys to realize the inherent potential of this organism. RESULTS: In this paper, we employ several state of the art bioinformatics techniques for studying the core circadian clock in Cyanothece sp. ATCC 51142, and its interactions with other key cellular processes. We employ comparative genomics techniques to map the circadian clock genes and genetic interactions from another cyanobacterial species, namely Synechococcus elongatus PCC 7942, of which the circadian clock has been much more thoroughly investigated. Using time series gene expression data for Cyanothece, we employ gene regulatory network reconstruction techniques to learn this network de novo, and compare the reconstructed network against the interactions currently reported in the literature. Next, we build a computational model of the interactions between the core clock and other cellular processes, and show how this model can predict the behaviour of the system under changing environmental conditions. The constructed models significantly advance our understanding of the Cyanothece circadian clock functional mechanisms. Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Sandeep Gaudana, Pramod P. Wangikar |
BMC Bioinform. | 2 |
| 2013 | Clustered Memetic Algorithm With Local Heuristics for Ab Initio Protein Structure PredictionabstractLow-resolution protein models are often used within a hierarchical framework for structure prediction. However, even with these simplified but realistic protein models, the search for the optimal solution remains NP complete. The complexity is further compounded by the multimodal nature of the search space. In this paper, we propose a systematic design of an evolutionary search technique, namely the memetic algorithm (MA), to effectively search the vast search space by exploiting the domain-specific knowledge and taking cognizance of the multimodal nature of the search space. The proposed MA achieves this by incorporating various novel features: 1) a modified fitness function includes two additional terms to account for the hydrophobic and polar nature of the residues; 2) a systematic (rather than random) generation of population automatically prevents an occurrence of invalid conformations; 3) a generalized nonisomorphic encoding scheme implicitly eliminates generation of twins (similar conformations) in the population; 4) the identification of a meme (protein substructures) during optimization from different basins of attraction - a process that is equivalent to implicit applications of threading principles; 5) a clustering of the population corresponds to basins of attraction that allows evolution to overcome the complexity of multimodal search space, thereby avoiding search getting trapped in a local optimum; and 6) a 2-stage framework gathers domain knowledge (i.e., substructures or memes) from different basins of attraction for a combined execution in the second stage. Experiments conducted with different lattice models using known benchmark protein sequences and comparisons carried out with recently reported approaches in this journal show that the proposed algorithm has robustness, speed, accuracy, and superior performance. The approach is generic and can easily be extended for applications to other classes of problems. Md. Kamrul Islam 0001, Madhu Chetty |
IEEE Trans. Evol. Comput. | 2 |
| 2012 | Adaptive regulatory genes cardinality for reconstructing genetic networksabstractWith the advent of microarray technology, researchers are able to determine cellular dynamics for thousands of genes simultaneously, thereby enabling reverse engineering of the gene regulatory network (GRN) from high-throughput time-series gene expression data. Amongst the various currently available models for inferring GRN, the S-System formalism is often considered as an excellent compromise between accuracy and mathematical tractability. In this paper, a novel approach for inferring GRN based on the decoupled S-System model, incorporating the new concept of adaptive regulatory genes cardinality, is proposed. Parameter learning for the S-System is carried out in an evolving manner using a versatile and robust Trigonometric Evolutionary Algorithm. The applicability and efficiency of the proposed method is studied using a well-known and widely studied synthetic network with various levels of noise, and excellent performance observed. Further, investigations of a 5 gene in-vivo synthetic biological network of Saccharomyces cerevisiae called IRMA, has succeeded in detecting higher number of correct regulations compared to other approaches reported earlier. Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
IEEE Congress on Evolutionary Computation | 2 |
| 2012 | Protein structure prediction based on optimal hydrophobic core formationabstractThe prediction of a minimum energy protein structure from its amino acid sequence represents an important and challenging problem in computational biology. In this paper, we propose a novel heuristic approach for protein structure prediction (PSP) based on the concept of optimal hydrophobic core formation. Using 2D HP model, a well-known set of sub-structures analogous to the secondary structures are obtained. Some sub-conformations are appropriately classified and then incorporated as prior knowledge. Unlike most of the popular PSP approaches which are stochastic in nature, the proposed method is deterministic. The effectiveness of the proposed algorithm is evaluated by well-known benchmark as well as non-benchmark sequences commonly used with 2D HP model. Maintaining similar accuracy as other core based and population based algorithms our method is significantly faster and reduces the computation time as it avoids blind search within the hydrophobic core (H-Core). Rumana Nazmul, Madhu Chetty, Ram Samudrala, David K. Chalmers |
IEEE Congress on Evolutionary Computation | 2 |
| 2012 | Local and Global Algorithms for Learning Dynamic Bayesian NetworksabstractLearning optimal Bayesian networks (BN) from data is NP-hard in general. Nevertheless, certain BN classes with additional topological constraints, such as the dynamic BN (DBN) models, widely applied in specific fields such as systems biology, can be efficiently learned in polynomial time. Such algorithms have been developed for the Bayesian-Dirichlet (BD), Minimum Description Length (MDL), and Mutual Information Test (MIT) scoring metrics. The BD-based algorithm admits a large polynomial bound, hence it is impractical for even modestly sized networks. The MDL-and MIT-based algorithms admit much smaller bounds, but require a very restrictive assumption that all variables have the same cardinality, thus significantly limiting their applicability. In this paper, we first propose an improvement to the MDL-and MIT-based algorithms, dropping the equicardinality constraint, thus significantly enhancing their generality. We also explore local Markov blanket based algorithms for constructing BN in the context of DBN, and show an interesting result: under the faithfulness assumption, the mutual information test based local Markov blanket algorithms yield the same network as learned by the global optimization MIT-based algorithm. Experimental validation on small and large scale genetic networks demonstrates the effectiveness of our proposed approaches. Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
ICDM | 2 |
| 2012 | On the Reconstruction of Genetic Network from Partial Microarray Data
Ahsan Raja Chowdhury, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (1) | 2 |
| 2012 | FusGP: Bayesian Co-learning of Gene Regulatory Networks and Protein Interaction Networks
Nizamul Morshed, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (5) | 2 |
| 2012 | Data Discretization for Dynamic Bayesian Network Based Modeling of Genetic Networks
Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
ICONIP (2) | 2 |
| 2012 | Gene regulatory network modeling via global optimization of high-order dynamic Bayesian networkabstractBACKGROUND: Dynamic Bayesian network (DBN) is among the mainstream approaches for modeling various biological networks, including the gene regulatory network (GRN). Most current methods for learning DBN employ either local search such as hill-climbing, or a meta stochastic global optimization framework such as genetic algorithm or simulated annealing, which are only able to locate sub-optimal solutions. Further, current DBN applications have essentially been limited to small sized networks. RESULTS: To overcome the above difficulties, we introduce here a deterministic global optimization based DBN approach for reverse engineering genetic networks from time course gene expression data. For such DBN models that consist only of inter time slice arcs, we show that there exists a polynomial time algorithm for learning the globally optimal network structure. The proposed approach, named GlobalMIT+, employs the recently proposed information theoretic scoring metric named mutual information test (MIT). GlobalMIT+ is able to learn high-order time delayed genetic interactions, which are common to most biological systems. Evaluation of the approach using both synthetic and real data sets, including a 733 cyanobacterial gene expression data set, shows significantly improved performance over other techniques. CONCLUSIONS: Our studies demonstrate that deterministic global optimization approaches can infer large scale genetic networks. Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
BMC Bioinform. | 2 |
| 2011 | An improved method to infer Gene Regulatory Network using S-SystemabstractGene Regulatory Network (GRN) plays an important role in the understanding of complex biological systems. In most cases, high throughput microarray gene expression data is used for finding these regulatory relationships among genes. In this paper, we present a novel approach, based on decoupled S System model, for reverse engineering GRNs. In the proposed method, the genetic algorithm used for scoring the networks contains several useful features for accurate network inference, namely a Prediction Initialization (PI) algorithm to initialize the individuals, a Flip Operation (FO) for better mating of values and a restricted execution of Hill Climbing Local Search over few individuals. It also includes a novel refinement technique which utilizes the fit solutions of the genetic algorithm for optimizing sensitivity and specificity of the inferred network. Comparative studies and robustness analysis using standard benchmark data set show the superiority of the proposed method. Ahsan Raja Chowdhury, Madhu Chetty |
IEEE Congress on Evolutionary Computation | 2 |
| 2011 | Novel local improvement techniques in clustered memetic algorithm for protein structure predictionabstractEvolutionary algorithms (EAs) often fail to find the global optimum due to genetic drift. As the protein structure prediction problem is multimodal having several global optima, EAs empowered with combined application of local and global search e.g., memetic algorithms, can be more effective. This paper introduces two novel local improvement techniques for the clustered memetic algorithm to incorporate both problem specific and search-space specific knowledge to find one of the optimum structures of a hydrophobic-polar protein sequence on lattice models. Experimental results show the superiority of the proposed techniques against existing EAs on benchmark sequences. Md. Kamrul Islam 0001, Madhu Chetty, M. Manzur Murshed |
IEEE Congress on Evolutionary Computation | 2 |
| 2011 | Reconstructing genetic networks with concurrent representation of instantaneous and time-delayed interactionsabstractAlthough living organisms can have some genetic interactions occurring instantaneously while others with time-delay, current modeling techniques for genetic network reconstruction make simplifications and assume that the interactions can be either of these but not both. In this paper, we propose a gene regulatory network reconstruction algorithm that can model concurrent occurrence of both, instantaneous as well as time-delayed interactions, thus providing a better representation of the original biological processes. First we introduce a novel framework using the Bayesian network (BN) formalism that can model both types of interactions. A gene regulatory network reconstruction algorithm using this proposed framework is then developed that employs an evolutionary search strategy and a decomposable scoring metric based on information theoretic quantities. Investigations of our approach are performed using both, the synthetic data as well as Saccharomyces cerevisiae gene expression data. Comparisons with recent reconstruction methods show the superiority of the proposed method. Nizamul Morshed, Madhu Chetty |
IEEE Congress on Evolutionary Computation | 2 |
| 2011 | Conflict Resolution Based Global Search Operators for Long Protein Structures Prediction
Md. Kamrul Islam 0001, Madhu Chetty, M. Manzur Murshed |
ICONIP (1) | 2 |
| 2011 | A Memetic Approach to Protein Structure Prediction in Triangular Lattices
Md. Kamrul Islam 0001, Madhu Chetty, Abu Zafer M. Dayem Ullah, Kathleen Steinhöfel |
ICONIP (1) | 2 |
| 2011 | Simultaneous Learning of Instantaneous and Time-Delayed Genetic Interactions Using Novel Information Theoretic Scoring Technique
Nizamul Morshed, Madhu Chetty, Xuan Vinh Nguyen |
ICONIP (2) | 2 |
| 2011 | Multi Agent Carbon Trading Incorporating Human Traits and Game Theory
Madhu Chetty, Suryani Lim |
ICONIP (3) | 2 |
| 2011 | Dynamic Bayesian Network Modeling of Cyanobacterial Biological Processes via Gene Clustering
Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
ICONIP (1) | 2 |
| 2011 | Polynomial Time Algorithm for Learning Globally Optimal Dynamic Bayesian Network
Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
ICONIP (3) | 2 |
| 2011 | GlobalMIT: learning globally optimal dynamic bayesian network with the mutual information test criterionabstractMOTIVATION: Dynamic Bayesian networks (DBN) are widely applied in modeling various biological networks including the gene regulatory network (GRN). Due to the NP-hard nature of learning static Bayesian network structure, most methods for learning DBN also employ either local search such as hill climbing, or a meta stochastic global optimization framework such as genetic algorithm or simulated annealing. RESULTS: This article presents GlobalMIT, a toolbox for learning the globally optimal DBN structure from gene expression data. We propose using a recently introduced information theoretic-based scoring metric named mutual information test (MIT). With MIT, the task of learning the globally optimal DBN is efficiently achieved in polynomial time. AVAILABILITY: The toolbox, implemented in Matlab and C++, is available at http://code.google.com/p/globalmit. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data is available at Bioinformatics online. Xuan Vinh Nguyen, Madhu Chetty, Ross L. Coppel, Pramod P. Wangikar |
Bioinform. | 2 |
| 2011 | Twin Removal in Genetic Algorithms for Protein Structure Prediction Using Low-Resolution ModelabstractThis paper presents the impact of twins and the measures for their removal from the population of genetic algorithm (GA) when applied to effective conformational searching. It is conclusively shown that a twin removal strategy for a GA provides considerably enhanced performance when investigating solutions to complex ab initio protein structure prediction (PSP) problems in low-resolution model. Without twin removal, GA crossover and mutation operations can become ineffectual as generations lose their ability to produce significant differences, which can lead to the solution stalling. The paper relaxes the definition of chromosomal twins in the removal strategy to not only encompass identical, but also highly correlated chromosomes within the GA population, with empirical results consistently exhibiting significant improvements solving PSP problems. Tamjidul Hoque, Madhu Chetty, Andrew Lewis 0004, Abdul Sattar 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | A Markov-Blanket-Based Model for Gene Regulatory Network InferenceabstractAn efficient two-step Markov blanket method for modeling and inferring complex regulatory networks from large-scale microarray data sets is presented. The inferred gene regulatory network (GRN) is based on the time series gene expression data capturing the underlying gene interactions. For constructing a highly accurate GRN, the proposed method performs: 1) discovery of a gene's Markov Blanket (MB), 2) formulation of a flexible measure to determine the network's quality, 3) efficient searching with the aid of a guided genetic algorithm, and 4) pruning to obtain a minimal set of correct interactions. Investigations are carried out using both synthetic as well as yeast cell cycle gene expression data sets. The realistic synthetic data sets validate the robustness of the method by varying topology, sample size, time delay, noise, vertex in-degree, and the presence of hidden nodes. It is shown that the proposed approach has excellent inferential capabilities and high accuracy even in the presence of noise. The gene network inferred from yeast cell cycle data is investigated for its biological relevance using well-known interactions, sequence analysis, motif patterns, and GO data. Further, novel interactions are predicted for the unknown genes of the network and their influence on other genes is also discussed. Ramesh Ram, Madhu Chetty |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Binary-Organoid Particle Swarm optimisation for inferring genetic networksabstractA holistic understanding of genetic interactions is crucial in the analysis of complex biological systems. However, due to the dimensionality problem (less samples and large number of genes) of microarray data, obtaining an optimal gene regulatory network is not only difficult but also computationally expensive. In this paper, a Bayesian model for the genetic interactions using the Minimum Description Length as a scoring metric is proposed. For fast optimisation of the network structure, we propose a novel Swarm Intelligence algorithm called Binary-Organoid Particle Swarm (BORG-Swarm). In BORG-Swarm we introduce the concepts of probability threshold vector and particle drift to update particle positions. Experimental studies are carried out using real-life yeast cell cycle dataset. Results indicate that existing binary swarms fail to converge and suffer from long runtimes. In constrast, BORG-Swarm's fast convergence towards the global optimum becomes apparent from results of extensive simulations. Santi S. Chanthaphavong, Madhu Chetty |
IEEE Congress on Evolutionary Computation | 2 |
| 2010 | Clustered memetic algorithm for protein structure predictionabstractMemetic algorithm (MA) often perform better than other evolutionary algorithm due to their combining the local search with the process of global optimization. However, like any other evolutionary algorithm (EA), MA due to the problem of genetic drift often result in sub-optimal solutions. The problem is more aggravated when EAs are applied to search complex landscape of NP complete problem like protein structure prediction. In this paper, to help mitigate the problem of genetic drift and also to cover large search space, we propose a novel initial population generation process and a novel MA which applies clusters for seeding the initial population. Apart from reducing the impact of genetic drift, the proposed MA also avoids processing of unnecessary individuals in the population, thus significantly reducing the computational burden, especially for large protein sequences. Simulation results presented using the 2D lattice HP model show the superiority of the proposed algorithm. Md. Kamrul Islam 0001, Madhu Chetty |
IEEE Congress on Evolutionary Computation | 2 |
| 2010 | Multiclass microarray gene expression classification based on fusion of correlation features
Girija Chetty, Madhu Chetty |
FUSION | 2 |
| 2010 | Computational Intelligence in Bioinformatics
Madhu Chetty, Alioune Ngom, Elena Marchiori |
Neurocomputing | 1 |
| 2010 | DFS-generated pathways in GA crossover for protein structure prediction
Tamjidul Hoque, Madhu Chetty, Andrew Lewis 0004, Abdul Sattar 0001, Vicky M. Avery |
Neurocomputing | 2 |
| 2010 | Pattern Recognition in Bioinformatics
Shandar Ahmad, Madhu Chetty, Bertil Schmidt |
Pattern Recognit. Lett. | 2 |
| 2009 | Combining segmental semi-Markov models with neural networks for protein secondary structure prediction
Niranjan P. Bidargaddi, Madhu Chetty, Joarder Kamruzzaman |
Neurocomputing | 2 |
| 2007 | Protein folding prediction in 3D FCC HP lattice model using genetic algorithmabstractIn most of the successful real protein structure prediction (PSP) problem, lattice models have been essentially utilized to have the folding backbone sampling at the top of the hierarchical approach. A three dimensional face-centred-cube (FCC), with the provision for providing the most compact core, can map closest to the folded protein in reality. Hence, our successful hybrid genetic algorithms (HGA) proposed earlier for a square and cube lattice model is being extended in this paper for a 3D FCC model. Furthermore, twins (conformations having similarity with each other), in GA population have also been considered for removal from the search space for improving the effectiveness of GA The HGA combined with the twin removal (TR) strategy showed best performance when compared with the simple GA (SGA), SGA with TR, and HGA only versions. Experiments were carried out on the publicly available benchmark HP sequences and results are expressed based on the fitness of the corresponding applied lattice model, which will help any future novel approach to be compared. Tamjidul Hoque, Madhu Chetty, Abdul Sattar 0001 |
IEEE Congress on Evolutionary Computation | 2 |
| 2007 | A guided genetic algorithm for learning gene regulatory networksabstractIn the post-genomic era, understanding the interactions of genes plays a vital role in the analysis of complex biological systems. Recently, we developed a causal model approach for learning gene regulatory networks from microarray data. The optimization process for this learning was implemented by using genetic algorithm (GA) as a search technique to find the best candidate over the space of possible networks. In this paper, we propose a genetic algorithm which is guided by exploiting certain characteristics of diversity and high level heuristics in order to generate good networks as quickly as possible. A comparison of this algorithm to the standard genetic algorithms implemented in our earlier work is also presented in this paper. The Guided GA (GGA) is tested on both synthetic and real-world microarray data. The novel approach of GGA shows superiority of the solutions, computational efficiency along with accuracy improvement compared to standard GA. Ramesh Ram, Madhu Chetty |
IEEE Congress on Evolutionary Computation | 2 |
| 2007 | Differential prioritization in feature selection and classifier aggregation for multiclass microarray datasets
Chia Huey Ooi, Madhu Chetty, Shyh Wei Teng |
Data Min. Knowl. Discov. | 2 |
| 2006 | A Guided Genetic Algorithm for Protein Folding Prediction Using 3D Hydrophobic-Hydrophilic ModelabstractIn this paper, a Guided Genetic Algorithm (GGA) has been presented for protein folding prediction (PFP) using 3D Hydrophobic-Hydrophilic (HP) model. Effective strategies have been formulated utilizing the core formation of the globular protein, which provides the guideline for the Genetic Algorithm (GA) while predicting protein folding. Building blocks containing Hydrophobic (H) -Hydrophilic (P or Polar) covalent bond are utilized such a way that it helps form a core that maximizes the fitness. A series of operators are developed including Diagonal Move and Tilt Move to assist in implementing the building blocks in three-dimensional space. The GGA outperformed Unger's GA in 3D HP model. The overall strategy incorporates a swing function that provides a mechanism to enable the GGA to test more potential solutions and also prevent it from developing a schema that may cause it to become trapped in local minima. Further, it helps the guidelines remain non-rigid. GGA provides improved and robust performance for PFP. Tamjidul Hoque, Madhu Chetty, Laurence Dooley |
IEEE Congress on Evolutionary Computation | 2 |
| 2006 | Fuzzy Model for Gene Regulatory NetworkabstractGene regulatory networks influence development and evolution in living organism. The advent of microarray technology has challenged computer scientists to develop better algorithms for modeling the underlying regulatory relationship in between the genes. Recently, a fuzzy logic model has been proposed to search microarray datasets for activator/repressor regulatory relationship. We improve this model for searching regulatory triplets by means of predicting changes in expression level of the target over interval time points based on input expression level, and comparing them with actual changes. This method eliminates possible false predictions from the classical fuzzy model thereby allowing a wider search space for inferring regulatory relationship. We also introduce a novel pre-processing technique using fuzzy logic that can group genes having similar changes in expression profile over all available intervals in the microarray data. This technique eliminates redundant computation performed by the proposed model. Saccharomyces cerevisiae data was applied to the model and 548 activator/repressor regulatory triplets were inferred from the data. These improvements will increase feasibility of using fuzzy logic for understanding the relationship between genes using microarray technology. Ramesh Ram, Madhu Chetty, Trevor I. Dix |
IEEE Congress on Evolutionary Computation | 2 |
| 2006 | Bayesian Segmentation using Residue Proximity for Secondary Structure and Contact PredictionabstractSecondary structure, residue contacts and contact numbers play an important role in tertiary structure determination of proteins. In the recent past, mainly due to non local interactions, the Bayesian segmentation approach has been successfully used for secondary structure prediction. In this paper, the performance of the Bayesian segmentation approach has been enhanced by taking residue contacts into account. The three state prediction accuracy increased by 2% when residue contacts were taken into account. Due to the inherent flexibility the Bayesian segmentation approach has been extended to infer residue contacts and contact numbers with the same segmentations. The proposed method achieved CorRvalues greater than 0.70 for protein sequence 1a62 and 1aba Niranjan P. Bidargaddi, Madhu Chetty, Joarder Kamruzzaman |
CIBCB | 2 |
| 2006 | Non-Isomorphic Coding in Lattice Model and its Impact for Protein Folding Prediction Using Genetic AlgorithmabstractTraditional encodings for hydrophobic(H)-hydrophilic(P) model or HP lattice models is isomorphic, which adds unwanted variations for the same solution, thereby slowing convergence. In this paper a novel non-isomorphic encoding scheme is presented for HP lattice model, which constrains the search space. In addition, similarity comparisons are made easier and more consistent and it will be shown that non-deterministic search approach such as genetic algorithm (GA) converges faster when non-isomorphic encoding is employed Tamjidul Hoque, Madhu Chetty, Laurence Dooley |
CIBCB | 2 |
| 2006 | Causal Modeling of Gene Regulatory NetworkabstractThe analysis of high-throughput experimental data, such as microarray gene expression data, is currently seen as a promising way of finding regulatory relationships between genes. Network inference algorithms are powerful computational tools for identifying putative causal interactions among variables from observed data. In this paper, we propose a network reconstruction technique to predict not only the structure but also the direction and sign of regulation using a genetic algorithm (GA). The networks consisting of nodes (genes), directed edges (gene-gene interactions) and dynamics of regulation are assigned scores using the presented causal model based on partial correlation. The highest scoring network best fits the expression data. As GAs are stochastic, the algorithm is repeated several times and the final network is reconstructed by combining the most significant connections identified from the high scoring networks. The presented technique is applied to the well known Saccharomyces cerevisiae microarray dataset and the reconstructed network is observed to be consistent with the results found in literature Ramesh Ram, Madhu Chetty, Trevor I. Dix |
CIBCB | 2 |
| 2006 | Differential prioritization between relevance and redundancy in correlation-based feature selection techniques for multiclass gene expression dataabstractBACKGROUND: Due to the large number of genes in a typical microarray dataset, feature selection looks set to play an important role in reducing noise and computational cost in gene expression-based tissue classification while improving accuracy at the same time. Surprisingly, this does not appear to be the case for all multiclass microarray datasets. The reason is that many feature selection techniques applied on microarray datasets are either rank-based and hence do not take into account correlations between genes, or are wrapper-based, which require high computational cost, and often yield difficult-to-reproduce results. In studies where correlations between genes are considered, attempts to establish the merit of the proposed techniques are hampered by evaluation procedures which are less than meticulous, resulting in overly optimistic estimates of accuracy. RESULTS: We present two realistically evaluated correlation-based feature selection techniques which incorporate, in addition to the two existing criteria involved in forming a predictor set (relevance and redundancy), a third criterion called the degree of differential prioritization (DDP). DDP functions as a parameter to strike the balance between relevance and redundancy, providing our techniques with the novel ability to differentially prioritize the optimization of relevance against redundancy (and vice versa). This ability proves useful in producing optimal classification accuracy while using reasonably small predictor set sizes for nine well-known multiclass microarray datasets. CONCLUSION: For multiclass microarray datasets, especially the GCM and NCI60 datasets, DDP enables our filter-based techniques to produce accuracies better than those reported in previous studies which employed similarly realistic evaluation procedures. Chia Huey Ooi, Madhu Chetty, Shyh Wei Teng |
BMC Bioinform. | 2 |
| 2005 | A new guided genetic algorithm for 2D hydrophobic-hydrophilic model to predict protein foldingabstractThis paper presents a novel guided genetic algorithm (GGA) for protein folding prediction (PFP) in 2D hydrophobic-hydrophilic (HP) by exploring the protein core formation concept. A proof of the shape for an optimal core is provided and a set of highly probable sub-conformations are defined which help to establish the guidelines to form the core boundary. A series of new operators including diagonal move and tilt move are defined to assist in implementing the guidelines. The underlying reasons for the failure in the folding prediction of relatively long sequences using Unger's genetic algorithm (GA) in 2D HP model are analysed and the new GGA is shown to overcome these limitations. The overall strategy incorporates a swing function that provides a mechanism to enable the GGA to test more potential solutions and also prevent it from developing a schema that may cause it to become trapped in local minima. While the guidelines do not force particular conformations, the result is a number of conformations for particular putative ground energy and superior prediction accuracy, endorsing the improved performance compared with other well established nondeterministic search approaches Tamjidul Hoque, Madhu Chetty, Laurence Dooley |
Congress on Evolutionary Computation | 2 |
| 2005 | Fuzzy Profile Hidden Markov Models for Protein Sequence Analysis
Niranjan P. Bidargaddi, Madhu Chetty, Joarder Kamruzzaman |
CIBCB | 2 |
| 2005 | An Architecture Combining Bayesian segmentation and Neural Network Ensembles for Protein Secondary Structure Prediction
Niranjan P. Bidargaddi, Madhu Chetty, Joarder Kamruzzaman |
CIBCB | 2 |
| 2005 | A Comparative Study of Two Novel Predictor Set Scoring Methods
Chia Huey Ooi, Madhu Chetty |
IDEAL | 2 |
| 2005 | Increasing Classification Accuracy by Combining Adaptive Sampling and Convex Pseudo-Data
Chia Huey Ooi, Madhu Chetty |
PAKDD | 2 |
| 2004 | An Efficient Algorithm for Computing the Fitness Function of a Hydrophobic-Hydrophilic ModelabstractThe protein folding problem is a minimization problem in which the energy function is often regarded as the fitness function. There are several models for protein folding prediction including the hydrophobic-hydrophilic (HP) model. Though this model is an elementary one, it is widely used as a test-bed for faster execution of new algorithms. Fitness computation is one of the major computational parts of the HP model. This paper proposes an efficient search (ES) approach for computing the fitness value requiring only O(n) complexity in contrast to the full search (FS) approach that requires O(n/sup 2/) complexity. The efficiency of the proposed ES approach results due to its utilization of some inherent properties of the HP model. The ES approach represents residues in a Cartesian coordinate framework and then uses relative distance and coordinate polarity to reduce complexity. Tamjidul Hoque, Madhu Chetty, Laurence Dooley |
HIS | 2 |
| 2004 | Partially Computed Fitness Function Based Genetic Algorithm for Hydrophobic-Hydrophilic ModelabstractFitness computation after each crossover or mutation operation in genetic algorithm (GA) requires computational time that increases with the increasing length of the chromosome. In this paper, an efficient GA is proposed for protein folding prediction based on the hydrophobic-hydrophilic (HP) model. The partial fitness of the parent computed from one end of sequence till crossover or mutation point is utilized for the computation of the fitness of the child. The calculated value of the partial fitness is stored with the corresponding chromosome. Although the approach requires additional memory for each hydrophobic residue of each chromosome, the computation time is reduced significantly, which is more important than the memory overhead. Tamjidul Hoque, Madhu Chetty, Laurence Dooley |
HIS | 2 |
| 2003 | An Incremental Constructive Layer Algorithm for Controller Design
Niranjan P. Bidargaddi, Madhu Chetty |
HIS | 2 |