Mahesan Niranjan

dblp:15/6640 · DBLP profile ↗
← Back
94ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0001-7021-140XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Balancing misclassification errors in image-based inference using problem domain semantics and a nested cascade architecture
abstract
Pattern recognition models, particularly neural networks, often focus on maximising classification accuracy. However, in practice, the types of errors made (misclassification between different classes) can have varying associated costs. Current methods overlook varying misclassification error types. Misclassification labels can either be available from expert knowledge or derived from semantics of textual descriptions of class labels. Exploiting such misclassification costs can have significant implications when deploying machine learning systems. Here, using five examples from image and tabular domains, we show how a deep neural architecture trained in a nested layer-wise fashion (cascade learning) in which early layers solve easier problems than later ones could exploit such hierarchical aspects of class labels. We employ a measure of performance called "severity" of errors and show how emphasis could be placed on classes that are deeper in the hierarchy, ignoring errors that arise between semantic neighbours. Supplementary Information: The online version contains supplementary material available at 10.1007/s00521-025-11613-8.
Rajesh Jena, Katayoun Farrahi, Mahesan Niranjan
Neural Comput. Appl.4
2025 A variational autoencoder for probabilistic non-negative matrix factorisation
abstract
Abstract We introduce and demonstrate the variational autoencoder (VAE) for probabilistic non-negative matrix factorisation (PAE-NMF). We design a network which can perform non-negative matrix factorisation (NMF) and add in aspects of a VAE to make the coefficients of the latent space probabilistic. By restricting the weights in the final layer of the network to be non-negative and using the non-negative Weibull distribution we produce a probabilistic form of NMF which allows us to generate new data and find a probability distribution that effectively links the latent and input variables. Our approach uses a minimum description length methodology to provide a method for achieving automatic regularisation; as it is designed using neural networks it can leverage deep learning frameworks for automatic differentiation, fast gradient descent algorithms and GPU support. We demonstrate the effectiveness of PAE-NMF on three heterogeneous datasets: images, financial time series and genomic.
Steven Squires, Adam Prügel-Bennett, Mahesan Niranjan
Pattern Anal. Appl.3
2024 3D Semantic Scene Completion From A Depth Map With Unsupervised Learning For Semantics Prioritisation
abstract
The Semantic Scene Completion (SSC) problem entails generating a comprehensive 3D voxel representation of a scene from a partial view, while simultaneously predicting volumetric occupancy and object category. A significant challenge in SSC is evaluating occluded regions in 3D space and accurately predicting object categories within an imbalanced setting. In addressing this challenge, our study explores SSC literature and introduces a simple, innovative class-balancing re-weighting technique rooted in an unsupervised clustering, leading to balanced learning and generalised representation. This method modulates the penalty on dataset classes during the CNN learning, emphasizing infrequent classes while moderately de-prioritizing the dominant ones, combining the strengths of both re-sampling and cost-sensitive learning enhancing the performance for both scene completion and scene semantics tasks. Our design, which relies on a single depth input without any RGB information, has shown to significantly outperform comparable baseline models. Our results are also competitively matched with other multi-input methods.
Mona Alawadh, Mahesan Niranjan, Hansung Kim 0001
ICIP2
2024 Protein language models meet reduced amino acid alphabets
abstract
MOTIVATION: Protein language models (PLMs), which borrowed ideas for modelling and inference from natural language processing, have demonstrated the ability to extract meaningful representations in an unsupervised way. This led to significant performance improvement in several downstream tasks. Clustering amino acids based on their physical-chemical properties to achieve reduced alphabets has been of interest in past research, but their application to PLMs or folding models is unexplored. RESULTS: Here, we investigate the efficacy of PLMs trained on reduced amino acid alphabets in capturing evolutionary information, and we explore how the loss of protein sequence information impacts learned representations and downstream task performance. Our empirical work shows that PLMs trained on the full alphabet and a large number of sequences capture fine details that are lost in alphabet reduction methods. We further show the ability of a structure prediction model(ESMFold) to fold CASP14 protein sequences translated using a reduced alphabet. For 10 proteins out of the 50 targets, reduced alphabets improve structural predictions with LDDT-Cα differences of up to 19%. AVAILABILITY AND IMPLEMENTATION: Trained models and code are available at github.com/Ieremie/reduced-alph-PLM.
Ioan Ieremie, Rob M. Ewing, Mahesan Niranjan
Bioinform.3
2024 Non-negative subspace feature representation for few-shot learning in medical imaging
Keqiang Fan, Xiaohao Cai, Mahesan Niranjan
Image Vis. Comput.3
2023 Depth Estimation for a Single Omnidirectional Image with Reversed-Gradient Warming-up Thresholds Discriminator
abstract
Depth estimation for single image using deep learning requires a large labelled depth dataset with various scenes for training. However, currently published omnidirectional depth datasets cover limited types of scenes and are not suitable for depth estimation for various real-world scenes. With the challenge of labelled real-world datasets generation and stability of the performance, we propose an architecture with the Reverse-gradient Warming-up Threshold Discriminator (RWTD) to estimate real-world depth maps from the synthetic ground truth. It takes labelled synthetic scenes of a source domain and unlabelled real-world scenes of a target domain as inputs to predict the corresponding depth maps. Compared with state-of-the-art encoder-decoder models, the proposed architecture shows an 11% points improvement on the testing dataset for depth accuracy.
Yihong Wu 0004, Yuwen Heng, Mahesan Niranjan, Hansung Kim 0001
ICASSP3
2023 IIHT: Medical Report Generation with Image-to-Indicator Hierarchical Transformer
Keqiang Fan, Xiaohao Cai, Mahesan Niranjan
ICONIP (6)3
2023 Quantum annealing-based clustering of single cell RNA-seq data
abstract
Cluster analysis is a crucial stage in the analysis and interpretation of single-cell gene expression (scRNA-seq) data. It is an inherently ill-posed problem whose solutions depend heavily on hyper-parameter and algorithmic choice. The popular approach of K-means clustering, for example, depends heavily on the choice of K and the convergence of the expectation-maximization algorithm to local minima of the objective. Exhaustive search of the space for multiple good quality solutions is known to be a complex problem. Here, we show that quantum computing offers a solution to exploring the cost function of clustering by quantum annealing, implemented on a quantum computing facility offered by D-Wave [1]. Out formulation extracts minimum vertex cover of an affinity graph to sub-sample the cell population and quantum annealing to optimise the cost function. A distribution of low-energy solutions can thus be extracted, offering alternate hypotheses about how genes group together in their space of expressions.
Michal Kubacki, Mahesan Niranjan
Briefings Bioinform.2
2022 Robust 3D rotation invariant local binary pattern for volumetric texture classification
abstract
3D local binary pattern (LBP) shows significant performance in many domains such as solid textures analysis, face recognition and tumor detection. In recent years, rotation invariant 3D LBP texture descriptors have received increasing attention and several variants have been proposed. However, they are sensitive to the noise present in the image. In this paper, we propose an efficient rotation invariant texture descriptor known as robust extended 3D LBP (RELBP) for volumetric texture classification. Unlike the current 3D LBP framework, our descriptor uses the information of neighboring voxels to reduce noise. First, the 3D weighted average filter is employed to process each voxel in the image, in which the center voxel is replaced by the average local gray level based on weights. Besides, equidistant points on a sphere are sampled to construct a set of rotation invariant features. Our experiments demonstrate that the RELBP proposed here shows superior classification performance in texture classification tasks and our method is highly robust to image noise on benchmark datasets.
Shengyu Lu, Sasan Mahmoodi, Mahesan Niranjan
ICPR3
2022 TransformerGO: predicting protein-protein interactions by modelling the attention between sets of gene ontology terms
abstract
MOTIVATION: Protein-protein interactions (PPIs) play a key role in diverse biological processes but only a small subset of the interactions has been experimentally identified. Additionally, high-throughput experimental techniques that detect PPIs are known to suffer various limitations, such as exaggerated false positives and negatives rates. The semantic similarity derived from the Gene Ontology (GO) annotation is regarded as one of the most powerful indicators for protein interactions. However, while computational approaches for prediction of PPIs have gained popularity in recent years, most methods fail to capture the specificity of GO terms. RESULTS: We propose TransformerGO, a model that is capable of capturing the semantic similarity between GO sets dynamically using an attention mechanism. We generate dense graph embeddings for GO terms using an algorithmic framework for learning continuous representations of nodes in networks called node2vec. TransformerGO learns deep semantic relations between annotated terms and can distinguish between negative and positive interactions with high accuracy. TransformerGO outperforms classic semantic similarity measures on gold standard PPI datasets and state-of-the-art machine-learning-based approaches on large datasets from Saccharomyces cerevisiae and Homo sapiens. We show how the neural attention mechanism embedded in the transformer architecture detects relevant functional terms when predicting interactions. AVAILABILITY AND IMPLEMENTATION: https://github.com/Ieremie/TransformerGO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ioan Ieremie, Rob M. Ewing, Mahesan Niranjan
Bioinform.3
2022 Convex Multi-View Clustering Via Robust Low Rank Approximation With Application to Multi-Omic Data
abstract
Recent advances in high throughput technologies have made large amounts of biomedical omics data accessible to the scientific community. Single omic data clustering has proved its impact in the biomedical and biological research fields. Multi-omic data clustering and multi-omic data integration techniques have shown improved clustering performance and biological insight. Cancer subtype clustering is an important task in the medical field to be able to identify a suitable treatment procedure and prognosis for cancer patients. State of the art multi-view clustering methods are based on non-convex objectives which only guarantee non-global solutions that are high in computational complexity. Only a few convex multi-view methods are present. However, their models do not take into account the intrinsic manifold structure of the data. In this paper, we introduce a convex graph regularized multi-view clustering method that is robust to outliers. We compare our algorithm to state of the art convex and non-convex multi-view and single view clustering methods, and show its superiority in clustering cancer subtypes on publicly available cancer genomic datasets from the TCGA repository. We also show our method's better ability to potentially discover cancer subtypes compared to other state of the art multi-view methods.
Omar Shetta, Mahesan Niranjan, Srinandan Dasmahapatra
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 A Numerical Measure of the Instability of Mapper-Type Algorithms
abstract
Mapper is an unsupervised machine learning algorithm generalising the notion of clustering to obtain a geometric description of a dataset. The procedure splits the data into possibly overlapping bins which are then clustered. The output of the algorithm is a graph where nodes represent clusters and edges represent the sharing of data points between two clusters. However, several parameters must be selected before applying Mapper and the resulting graph may vary dramatically with the choice of parameters. We define an intrinsic notion of Mapper instability that measures the variability of the output as a function of the choice of parameters required to construct a Mapper output. Our results and discussion are general and apply to all Mapper-type algorithms. We derive theoretical results that provide estimates for the instability and suggest practical ways to control it. We provide also experiments to illustrate our results and in particular we demonstrate that a reliable candidate Mapper output can be identified as a local minimum of instability regarded as a function of Mapper input parameters.
Francisco Belchí Guillamón, Jacek Brodzki, Matthew Burfitt, Mahesan Niranjan
J. Mach. Learn. Res.4
2020 Estimation of Gaussian mixture models via tensor moments with application to online learning
abstract
In this paper, we present an alternating gradient descent algorithm for estimating parameters of a spherical Gaussian mixture model by the method of moments (AGD-MoM). We formulate the problem as a constrained optimisation problem which simultaneously matches the third order moments from the data, represented as a tensor, and the second order moment, which is the empirical covariance matrix . We derive the necessary gradients (and second derivatives), and use them to implement alternating gradient search to estimate the parameters of the model. We show that the proposed method is applicable in both a batch as well as in a streaming (online) setting. Using synthetic and benchmark datasets, we demonstrate empirically that the proposed algorithm outperforms the more classical algorithms like Expectation Maximisation and variational Bayes.
Donya Rahmani, Mahesan Niranjan, Damien Fay, Akiko Takeda, Jacek Brodzki
Pattern Recognit. Lett.2
2019 Transfer learning across human activities using a cascade neural network architecture
abstract
Cascade Learning (CL) [20] is a new adaptive approach to train deep neural networks. It is particularly suited to transfer learning, as learning is achieved in a layerwise fashion, enabling the transfer of selected layers to optimize the quality of transferred features. In the domain of Human Activity Recognition (HAR), where the consideration of resource consumption is critical, CL is of particular interest as it has demonstrated the ability to achieve significant reductions in computational and memory costs with negligible performance loss. In this paper, we evaluate the use of CL and compare it to end to end (E2E) learning in various transfer learning experiments, all applied to HAR. We consider transfer learning across objectives, for example opening the door features transferred to opening the dishwasher. We additionally consider transfer across sensor locations on the body, as well as across datasets. Over all of our experiments, we find that CL achieves state of the art performance for transfer learning in comparison to previously published work, improving F1 scores by over 15%. In comparison to E2E learning, CL performs similarly considering F1 scores, with the additional advantage of requiring fewer parameters. Finally, the overall results considering HAR classification performance and memory requirements demonstrate that CL is a good approach for transfer learning.
Katayoun Farrahi, Mahesan Niranjan
UbiComp3
2019 Saliency Map on Cnns for Protein Secondary Structure Prediction
abstract
Deep learning, a powerful methodology for data-driven modelling, has been shown to be useful in tackling several problems in the biomedical domain. However, deep neural architectures lack interpretability of how predictions from them are made on any test input. While several approaches to "opening the black box" are being developed, their application to biological and medical data is very much as its infancy. Here, we consider the specific problem of protein secondary structure prediction using the techniques of saliency maps to explain decisions of a deep neural network. The analysis leads to two important observations: (a) one-hot-encoded amino-acids are irrelevant in the presence of PSSM values as extra features; and (b) in predicting α-helices at any position, amino-acids to the right are far more important than those to the left. The latter observation may have a biological basis relating to the synthesis of proteins by ribosome movement from left to right, sequentially adding amino-acids.
Guillermo Romero Moreno, Mahesan Niranjan, Adam Prügel-Bennett
ICASSP2
2019 Representation-dimensionality Trade-off in Biological Sequence-based Inference
abstract
Statistical inference from the analysis of biological sequences is widely used in the prediction of structure and biochemical functions of newly found macromolecules. For the application of machine learning methodologies such as kernel methods and artificial neural networks for such inference, variable length sequence data is often embedded in a finite dimensional real-valued space. The corresponding embedding dimensions are often high, leading to technical difficulties centred around the statistical concept of the curse of dimensionality. We demonstrate a trade-off between fidelity of representation of amino acids of proteins and the resulting dimensionality of the embedding space. Clustering chemically similar amino acids, thereby reducing the alphabet size, reduces the accuracy in their variation, but achieves a reduction in the corresponding feature space. We show this trade-off in three different problems of statistical inference, namely, protein-protein interaction, remote homology and secondary structure prediction. We show that in the reduced space performance often improves similar to what is seen in "diminishing returns" type reward-effort curves. We find alphabet reduction schemes taken from the literature, which are based on some biochemical rationale, perform significantly better than arbitrary random clustering of the alphabets. Statistical feature selection from the full 20 amino acid representation is not competitive with any of these. Dimensionality of representation has an important role when mapping sequence data onto fixed dimensions of an Euclidean space. This work shows that dimensionality reduction based on compressing the amino acid alphabet improves inference performance in two widely studied problems and degrades gracefully in the third. Alphabet reduction, which has a principled biochemical basis, is shown to be superior to feature selection which is purely a statistical exercise.
Bahman Asadi, Mahesan Niranjan
IJCNN2
2019 Classification and Regression Analysis of Lung Tumors from Multi-level Gene Expression Data
abstract
We study classification and regression problems in lung tumors where high throughput gene expression is measured at multiple levels: epi-genetics, transcription and protein. We uncover the correlates of smoking and gender-specificity in lung tumors. Different genes are indicative of smoking levels, gender and survival rates at these different levels. We also carry out an integrative anaysis, by feature selection from the pool of all three levels of features. Our results show that the epigenetic information in DNA methylation is a better marker for smoking status than gene expression either at the transcript or protein levels. Further, surprisingly, integrative anlysis using multi-level gene expression offers no significant advantage over the individual levels in the classification and survival prediction problems considered.
Pratheeba Jeyananthan, Mahesan Niranjan
IJCNN2
2019 Uncovering extensive post-translation regulation during human cell cycle progression by integrative multi-'omics analysis
abstract
BACKGROUND: Analysis of high-throughput multi-'omics interactions across the hierarchy of expression has wide interest in making inferences with regard to biological function and biomarker discovery. Expression levels across different scales are determined by robust synthesis, regulation and degradation processes, and hence transcript (mRNA) measurements made by microarray/RNA-Seq only show modest correlation with corresponding protein levels. RESULTS: In this work we are interested in quantitative modelling of correlation across such gene products. Building on recent work, we develop computational models spanning transcript, translation and protein levels at different stages of the H. sapiens cell cycle. We enhance this analysis by incorporating 25+ sequence-derived features which are likely determinants of cellular protein concentration and quantitatively select for relevant features, producing a vast dataset with thousands of genes. We reveal insights into the complex interplay between expression levels across time, using machine learning methods to highlight outliers with respect to such models as proteins associated with post-translationally regulated modes of action. CONCLUSIONS: We uncover quantitative separation between modified and degraded proteins that have roles in cell cycle regulation, chromatin remodelling and protein catabolism according to Gene Ontology; and highlight the opportunities for providing biological insights in future model systems.
Gregory M. Parkes, Mahesan Niranjan
BMC Bioinform.2
2019 A comparison of multitask and single task learning with artificial neural networks for yield curve forecasting
Manuel Nunes, Enrico H. Gerding, Frank McGroarty, Mahesan Niranjan
Expert Syst. Appl.4
2018 Deep Cascade Learning
abstract
In this paper, we propose a novel approach for efficient training of deep neural networks in a bottom-up fashion using a layered structure. Our algorithm, which we refer to as deep cascade learning, is motivated by the cascade correlation approach of Fahlman and Lebiere, who introduced it in the context of perceptrons. We demonstrate our algorithm on networks of convolutional layers, though its applicability is more general. Such training of deep networks in a cascade directly circumvents the well-known vanishing gradient problem by ensuring that the output is always adjacent to the layer being trained. We present empirical evaluations comparing our deep cascade training with standard end-end training using back propagation of two convolutional neural network architectures on benchmark image classification tasks (CIFAR-10 and CIFAR-100). We then investigate the features learned by the approach and find that better, domain-specific, representations are learned in early layers when compared to what is learned in end-end training. This is partially attributable to the vanishing gradient problem that inhibits early layer filters to change significantly from their initial settings. While both networks perform similarly overall, recognition accuracy increases progressively with each added layer, with discriminative features learned in every stage of the network, whereas in end-end training, no such systematic feature representation was observed. We also show that such cascade training has significant computational and memory advantages over end-end training, and can be used as a pretraining algorithm to obtain a better performance.
Enrique S. Marquez, Jonathon S. Hare, Mahesan Niranjan
IEEE Trans. Neural Networks Learn. Syst.3
2017 Robust Portfolio Risk Minimization Using the Graphical Lasso
Tristan Millington, Mahesan Niranjan
ICONIP (2)2
2017 A Method of Integrating Spatial Proteomics and Protein-Protein Interaction Network Data
Steven Squires, Rob M. Ewing, Adam Prügel-Bennett, Mahesan Niranjan
ICONIP (5)4
2017 Non-Negative Matrix Factorization with Exogenous Inputs for Modeling Financial Data
Steven Squires, Luis Montesdeoca, Adam Prügel-Bennett, Mahesan Niranjan
ICONIP (2)4
2017 Rank Selection in Nonnegative Matrix Factorization using Minimum Description Length
abstract
Nonnegative matrix factorization (NMF) is primarily a linear dimensionality reduction technique that factorizes a nonnegative data matrix into two smaller nonnegative matrices: one that represents the basis of the new subspace and the second that holds the coefficients of all the data points in that new space. In principle, the nonnegativity constraint forces the representation to be sparse and parts based. Instead of extracting holistic features from the data, real parts are extracted that should be significantly easier to interpret and analyze. The size of the new subspace selects how many features will be extracted from the data. An effective choice should minimize the noise while extracting the key features. We propose a mechanism for selecting the subspace size by using a minimum description length technique. We demonstrate that our technique provides plausible estimates for real data as well as accurately predicting the known size of synthetic data. We provide an implementation of our code in a Matlab format.
Steven Squires, Adam Prügel-Bennett, Mahesan Niranjan
Neural Comput.3
2016 Rotation invariant texture descriptors based on Gaussian Markov random fields for classification
Chathurika Dharmagunawardhana, Sasan Mahmoodi, Michael J. Bennett, Mahesan Niranjan
Pattern Recognit. Lett.4
2015 Statistical Machine Translation from and into Morphologically Rich and Low Resourced Languages
Randil Pushpananda, Ruvan Weerasinghe, Mahesan Niranjan
CICLing (1)3
2015 MIAT: A novel attribute selection approach to better predict upper gastrointestinal cancer
abstract
The use of data mining has led to many significant medical discoveries. However, many challenges still exist in using these methods for knowledge discovery within this field given that the large amounts of data medical practitioners collect often creates a curse of dimensionality. To address this challenge, attribute selection approaches have been developed. However, current approaches typically put equal weight on all values within that attribute. At times, and especially within medical domains, we claim that these approaches might miss attributes where only a small subset of attribute values contain a strong indication for one of the target values and thus should still be selected. To quantify this approach, we present MIAT, an algorithm that defines Minority Interesting Attribute Thresholds to find these important attribute values. As we developed MIAT to help better diagnose upper gastrointestinal cancer, we present how we use the attributes selected through this approach to build a predictive model for this cancer. To demonstrate MIAT's generality, we also applied it to a canonical Hungarian Heart Disease Dataset. In both datasets we found that MIAT yields significantly better accuracy and sensitivity over traditional attribute selection approaches.
Avi Rosenfeld, David G. Graham, Rifat Hamoudi, Rommel Butawan, Victor Eneh, Saif Khan, Haroon Miah, Mahesan Niranjan, Laurence B. Lovat
DSAA8
2015 Single-cell transcriptional analysis to uncover regulatory circuits driving cell fate decisions in early mouse development
abstract
MOTIVATION: Transcriptional regulatory networks controlling cell fate decisions in mammalian embryonic development remain elusive despite a long time of research. The recent emergence of single-cell RNA profiling technology raises hope for new discovery. Although experimental works have obtained intriguing insights into the mouse early development, a holistic and systematic view is still missing. Mathematical models of cell fates tend to be concept-based, not designed to learn from real data. To elucidate the regulatory mechanisms behind cell fate decisions, it is highly desirable to synthesize the data-driven and knowledge-driven modeling approaches. RESULTS: We propose a novel method that integrates the structure of a cell lineage tree with transcriptional patterns from single-cell data. This method adopts probabilistic Boolean network (PBN) for network modeling, and genetic algorithm as search strategy. Guided by the 'directionality' of cell development along branches of the cell lineage tree, our method is able to accurately infer the regulatory circuits from single-cell gene expression data, in a holistic way. Applied on the single-cell transcriptional data of mouse preimplantation development, our algorithm outperforms conventional methods of network inference. Given the network topology, our method can also identify the operational interactions in the gene regulatory network (GRN), corresponding to specific cell fate determination. This is one of the first attempts to infer GRNs from single-cell transcriptional data, incorporating dynamics of cell development along a cell lineage tree. AVAILABILITY AND IMPLEMENTATION: Implementation of our algorithm is available from the authors upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haifen Chen, Shital K. Mishra, Paul Robson 0001, Mahesan Niranjan, Jie Zheng 0002
Bioinform.5
2015 Outlier detection at the transcriptome-proteome interface
abstract
BACKGROUND: In high-throughput experimental biology, it is widely acknowledged that while expression levels measured at the levels of transcriptome and the corresponding proteome do not, in general, correlate well, messenger RNA levels are used as convenient proxies for protein levels. Our interest is in developing data-driven computational models that can bridge the gap between these two levels of measurement at which different mechanisms of regulation may act on different molecular species causing any observed lack of correlations. To this end, we build data-driven predictors of protein levels using mRNA levels and known proxies of translation efficiencies as covariates. Previous work showed that in such a setting, outliers with respect to the model are reliable candidates for post-translational regulation. RESULTS: Here, we present and compare two novel formulations of deriving a protein concentration predictor from which outliers may be extracted in a systematic manner. The first approach, outlier rejecting regression, allows explicit specification of a certain fraction of the data as outliers. In a regression setting, this is a non-convex optimization problem which we solve by deriving a difference of convex functions algorithm (DCA). With post-translationally regulated proteins, one expects their concentrations to be affected primarily by disruption of protein stability. Our second algorithm exploits this observation by minimizing an asymmetric loss using quantile regression and extracts outlier proteins whose measured concentrations are lower than what a genome-wide regression would predict. We validate the two approaches on a dataset of yeast transcriptome and proteome. Functional annotation check on detected outliers demonstrate that the methods are able to identify post-translationally regulated genes with high statistical confidence.
Yawwani Gunawardana, Shuhei Fujiwara, Akiko Takeda, Jeongmin Woo, Christopher H. Woelk, Mahesan Niranjan
Bioinform.6
2014 Large-scale Reordering Model for Statistical Machine Translation using Dual Multinomial Logistic Regression
abstract
Phrase reordering is a challenge for statis-tical machine translation systems. Posing phrase movements as a prediction prob-lem using contextual features modeled by maximum entropy-based classifier is su-perior to the commonly used lexicalized reordering model. However, Training this discriminative model using large-scale parallel corpus might be computationally expensive. In this paper, we explore recent advancements in solving large-scale clas-sification problems. Using the dual prob-lem to multinomial logistic regression, we managed to shrink the training data while iterating and produce significant saving in computation and memory while preserv-ing the accuracy. 1
Abdullah Alrajeh, Mahesan Niranjan
EMNLP2
2014 Memory-efficient large-scale linear support vector machine
abstract
Stochastic gradient descent has been advanced as a computationally efficient method for large-scale problems. In classification problems, many proposed linear support vector machines are very effective. However, they assume that the data is already in memory which might be not always the case. Recent work suggests a classical method that divides such a problem into smaller blocks then solves the sub-problems iteratively. We show that a simple modification of shrinking the dataset early will produce significant saving in computation and memory. We further find that on problems larger than previously considered, our approach is able to reach solutions on top-end desktop machines while competing methods cannot.
Abdullah Alrajeh, Akiko Takeda, Mahesan Niranjan
ICMV3
2014 An Inhomogeneous Bayesian Texture Model for Spatially Varying Parameter Estimation
abstract
In statistical model based texture feature extraction, features based on spatially varying parameters achieve higher discriminative performances compared to spatially constant parameters. In this paper we formulate a novel Bayesian framework which achieves texture characterization by spatially varying parameters based on Gaussian Markov random fields. The parameter estimation is carried out by Metropolis-Hastings algorithm. The distributions of estimated spatially varying parameters are then used as successful discriminant texture features in classification and segmentation. Results show that novel features outperform traditional Gaussian Markov random field texture features which use spatially constant parameters. These features capture both pixel spatial dependencies and structural properties of a texture giving improved texture features for effective texture classification and segmentation.
Chathurika Dharmagunawardhana, Sasan Mahmoodi, Michael J. Bennett, Mahesan Niranjan
ICPRAM4
2014 Biomedical visual data analysis to build an intelligent diagnostic decision support system in medical genetics
Kaya Kuru, Mahesan Niranjan, Yusuf Tunca, Erhan Osvank, Tayyaba Azim
Artif. Intell. Medicine2
2014 Gaussian Markov random field based improved texture descriptor for image segmentation
Chathurika Dharmagunawardhana, Sasan Mahmoodi, Michael J. Bennett, Mahesan Niranjan
Image Vis. Comput.4
2013 Enriching Texture Analysis with Semantic Data
abstract
We argue for the importance of explicit semantic modelling in human-centred texture analysis tasks such as retrieval, annotation, synthesis, and zero-shot learning. To this end, low-level attributes are selected and used to define a semantic space for texture. 319 texture classes varying in illumination and rotation are positioned within this semantic space using a pair wise relative comparison procedure. Low-level visual features used by existing texture descriptors are then assessed in terms of their correspondence to the semantic space. Textures with strong presence of attributes connoting randomness and complexity are shown to be poorly modelled by existing descriptors. In a retrieval experiment semantic descriptors are shown to outperform visual descriptors. Semantic modelling of texture is thus shown to provide considerable value in both feature selection and in analysis tasks.
Tim Matthews, Mark S. Nixon, Mahesan Niranjan
CVPR3
2013 Inferring Time-Delayed Gene Regulatory Networks Using Cross-Correlation and Sparse Regression
Piyushkumar A. Mundra, Jie Zheng 0002, Mahesan Niranjan, Roy E. Welsch, Jagath C. Rajapakse
ISBRA3
2013 Bridging the gap between transcriptome and proteome measurements identifies post-translationally regulated genes
abstract
MOTIVATION: Despite much dynamical cellular behaviour being achieved by accurate regulation of protein concentrations, messenger RNA abundances, measured by microarray technology, and more recently by deep sequencing techniques, are widely used as proxies for protein measurements. Although for some species and under some conditions, there is good correlation between transcriptome and proteome level measurements, such correlation is by no means universal due to post-transcriptional and post-translational regulation, both of which are highly prevalent in cells. Here, we seek to develop a data-driven machine learning approach to bridging the gap between these two levels of high-throughput omic measurements on Saccharomyces cerevisiae and deploy the model in a novel way to uncover mRNA-protein pairs that are candidates for post-translational regulation. RESULTS: The application of feature selection by sparsity inducing regression (l₁ norm regularization) leads to a stable set of features: i.e. mRNA, ribosomal occupancy, ribosome density, tRNA adaptation index and codon bias while achieving a feature reduction from 37 to 5. A linear predictor used with these features is capable of predicting protein concentrations fairly accurately (R² = 0.86). Proteins whose concentration cannot be predicted accurately, taken as outliers with respect to the predictor, are shown to have annotation evidence of post-translational modification, significantly more than random subsets of similar size P < 0.02. In a data mining sense, this work also shows a wider point that outliers with respect to a learning method can carry meaningful information about a problem domain.
Yawwani Gunawardana, Mahesan Niranjan
Bioinform.2
2013 On Acoustic Emotion Recognition: Compensating for Covariate Shift
abstract
Pattern recognition tasks often face the situation that training data are not fully representative of test data. This problem is well-recognized in speech recognition, where methods like cepstral mean normalization (CMN), vocal tract length normalization (VTLN) and maximum likelihood linear regression (MLLR) are used to compensate for channel and speaker differences. Speech emotion recognition (SER) is an important emerging field in human-computer interaction and faces the same data shift problems, a fact which has been generally overlooked in this domain. In this paper, we show that compensating for channel and speaker differences can give significant improvements in SER by modelling these differences as a covariate shift. We employ three algorithms from the domain of transfer learning that apply importance weights (IWs) within a support vector machine classifier to reduce the effects of covariate shift. We test these methods on the FAU Aibo Emotion Corpus, which was used in the Interspeech 2009 Emotion Challenge. It consists of two separate parts recorded independently at different schools; hence the two parts exhibit covariate shift. Results show that the IW methods outperform combined CMN and VTLN and significantly improve on the baseline performance of the Challenge. The best of the three methods also improves significantly on the winning contribution to the Challenge.
Ali Hassan 0001, Robert I. Damper, Mahesan Niranjan
IEEE Trans. Speech Audio Process.3
2012 Unsupervised Texture Segmentation using Active Contours and Local Distributions of Gaussian Markov Random Field Parameters
abstract
In this paper, local distributions of low order Gaussian Markov Random Field (GMRF) model parameters are proposed as texture features for unsupervised texture segmentation. Instead of using model parameters as texture features, we exploit the variations in parameter estimates found by model fitting in local region around the given pixel. The spatially localized estimation process is carried out by maximum likelihood method employing a moderately small estimation window which leads to modeling of partial texture characteristics belonging to the local region. Hence significant fluctuations occur in the estimates which can be related to texture pattern complexity. The variations occurred in estimates are quantified by normalized local histograms. Selection of an accurate window size for histogram calculation is crucial and is achieved by a technique based on the entropy of textures. These texture features expand the possibility of using relatively low order GMRF model parameters for segmenting fine to very large texture patterns and offer lower computational cost. Small estimation windows result in better boundary localization. Unsupervised segmentation is performed by integrated active contours, combining the region and boundary information. Experimental results on statistical and structural component textures show improved discriminative ability of the features compared to some recent algorithms in the literature.
Chathurika Dharmagunawardhana, Sasan Mahmoodi, Michael J. Bennett, Mahesan Niranjan
BMVC4
2012 Establishment of a Diagnostic Decision Support System in Genetic Dysmorphology
abstract
In the clinical diagnosis of facial dysmorphology, geneticists attempt to identify the underlying syndromes by associating facial features before cyto or molecular techniques are explored. Specifying genotype-phenotype correlations correctly among many syndromes is labor intensive especially for very rare diseases. The use of a computer based prediagnosis system can offer effective decision support particularly when only very few previous examples exist or in a remote environment where expert knowledge is not readily accessible. In this work we develop and demonstrate that accurate classification of dysmorphic faces is feasible by image processing of two dimensional face images. We test the proposed system on real patient image data by constructing a dataset of dysmorphic faces published in scholarly journals, hence having accurate diagnostic information about the syndrome. Our statistical methodology represents facial image data in terms of principal component analysis (PCA) and a leave one out evaluation scheme to quantify accuracy. The methodology has been tested with 15 syndromes including 75 cases, 5 examples per syndrome. A diagnosis success rate of 79% has been established. It can be concluded that a great number of syndromes indicating a characteristic pattern of facial anomalies can be typically diagnosed by employing computer-assisted machine learning algorithms since a face develops under the influence of many genes, particularly the genes causing syndromes.
Kaya Kuru, Mahesan Niranjan, Yusuf Tunca
ICMLA (2)2
2012 Gaussian process modelling for bicoid mRNA regulation in spatio-temporal Bicoid profile
abstract
MOTIVATION: Bicoid protein molecules, translated from maternally provided bicoid mRNA, establish a concentration gradient in Drosophila early embryonic development. There is experimental evidence that the synthesis and subsequent destruction of this protein is regulated at source by precise control of the stability of the maternal mRNA. Can we infer the driving function at the source from noisy observations of the spatio-temporal protein profile? We use non-parametric Gaussian process regression for modelling the propagation of Bicoid in the embryo and infer aspects of source regulation as a posterior function. RESULTS: With synthetic data from a 1D diffusion model with a source simulated to model mRNA stability regulation, our results establish that the Gaussian process method can accurately infer the driving function and capture the spatio-temporal dynamics of embryonic Bicoid propagation. On real data from the FlyEx database, too, the reconstructed source function is indicative of stability regulation, but is temporally smoother than what we expected, partly due to the fact that the dataset is only partially observed. To be in line with recent thinking on the subject, we also analyse this model with a spatial gradient of maternal mRNA, rather than being fixed at only the anterior pole. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mahesan Niranjan
Bioinform.2
2012 State and parameter estimation of the heat shock response system using Kalman and particle filters
abstract
MOTIVATION: Traditional models of systems biology describe dynamic biological phenomena as solutions to ordinary differential equations, which, when parameters in them are set to correct values, faithfully mimic observations. Often parameter values are tweaked by hand until desired results are achieved, or computed from biochemical experiments carried out in vitro. Of interest in this article, is the use of probabilistic modelling tools with which parameters and unobserved variables, modelled as hidden states, can be estimated from limited noisy observations of parts of a dynamical system. RESULTS: Here we focus on sequential filtering methods and take a detailed look at the capabilities of three members of this family: (i) extended Kalman filter (EKF), (ii) unscented Kalman filter (UKF) and (iii) the particle filter, in estimating parameters and unobserved states of cellular response to sudden temperature elevation of the bacterium Escherichia coli. While previous literature has studied this system with the EKF, we show that parameter estimation is only possible with this method when the initial guesses are sufficiently close to the true values. The same turns out to be true for the UKF. In this thorough empirical exploration, we show that the non-parametric method of particle filtering is able to reliably estimate parameters and states, converging from initial distributions relatively far away from the underlying true values. AVAILABILITY AND IMPLEMENTATION: Software implementation of the three filters on this problem can be freely downloaded from http://users.ecs.soton.ac.uk/mn/HeatShock
Mahesan Niranjan
Bioinform.2
2012 Markov Chain Monte Carlo Methods for State-Space Models with Point Process Observations
abstract
This letter considers how a number of modern Markov chain Monte Carlo (MCMC) methods can be applied for parameter estimation and inference in state-space models with point process observations. We quantified the efficiencies of these MCMC methods on synthetic data, and our results suggest that the Reimannian manifold Hamiltonian Monte Carlo method offers the best performance. We further compared such a method with a previously tested variational Bayes method on two experimental data sets. Results indicate similar performance on the large data sets and superior performance on small ones. The work offers an extensive suite of MCMC algorithms evaluated on an important class of models for physiological signal analysis.
Mark A. Girolami, Mahesan Niranjan
Neural Comput.3
2011 Exploitation of Machine Learning Techniques in Modelling Phrase Movements for Machine Translation
Yizhao Ni, Craig Saunders, Sándor Szedmák, Mahesan Niranjan
J. Mach. Learn. Res.4
2011 Online Variational Inference for State-Space Models with Point-Process Observations
abstract
We present a variational Bayesian (VB) approach for the state and parameter inference of a state-space model with point-process observations, a physiologically plausible model for signal processing of spike data. We also give the derivation of a variational smoother, as well as an efficient online filtering algorithm, which can also be used to track changes in physiological parameters. The methods are assessed on simulated data, and results are compared to expectation-maximization, as well as Monte Carlo estimation techniques, in order to evaluate the accuracy of the proposed approach. The VB filter is further assessed on a data set of taste-response neural cells, showing that the proposed approach can effectively capture dynamical changes in neural responses in real time.
Andrew Zammit-Mangion, Visakan Kadirkamanathan, Mahesan Niranjan, Guido Sanguinetti
Neural Comput.4
2010 Reducing the algorithmic variability in transcriptome-based inference
abstract
MOTIVATION: High-throughput measurements of mRNA abundances from microarrays involve several stages of preprocessing. At each stage, a user has access to a large number of algorithms with no universally agreed guidance on which of these to use. We show that binary representations of gene expressions, retaining only information on whether a gene is expressed or not, reduces the variability in results caused by algorithmic choice, while also improving the quality of inference drawn from microarray studies. RESULTS: Binary representation of transcriptome data has the desirable property of reducing the variability introduced at the preprocessing stages due to algorithmic choice. We compare the effect of the choice of algorithms on different problems and suggest that using binary representation of microarray data with Tanimoto kernel for support vector machine reduces the effect of the choice of algorithm and simultaneously improves the performance of classification of phenotypes.
Salih Tuna, Mahesan Niranjan
Bioinform.2
2010 The application of structured learning in natural language processing
Yizhao Ni, Craig Saunders, Sándor Szedmák, Mahesan Niranjan
Mach. Transl.4
2010 Estimating a State-Space Model from Point Process Observations: A Note on Convergence
abstract
Physiological signals such as neural spikes and heartbeats are discrete events in time, driven by continuous underlying systems. A recently introduced data-driven model to analyze such a system is a state-space model with point process observations, parameters of which and the underlying state sequence are simultaneously identified in a maximum likelihood setting using the expectation-maximization (EM) algorithm. In this note, we observe some simple convergence properties of such a setting, previously un-noticed. Simulations show that the likelihood is unimodal in the unknown parameters, and hence the EM iterations are always able to find the globally optimal solution.
Mahesan Niranjan
Neural Comput.2
2009 Modelling uncertainty in transcriptome measurements enhances network component analysis of yeast metabolic cycle
abstract
Using high throughput DNA binding data for transcription factors and DNA microarray time course data, we constructed four transcription regulatory networks and analysed them using a novel extension to the network component analysis (NCA) approach. We incorporated probe level uncertainties in gene expression measurements into the NCA analysis by the application of probabilistic principal component analysis (PPCA), and applied the method to data from yeast metabolic cycle. Analysis shows statistically significant enhancement to periodicity in a large fraction of the transcription factor activities inferred from the model. For several of these we found literature evidence of post-transcriptional regulation. Accounting for probe level uncertainty of microarray measurements leads to improved network component analysis. Transcription factor profiles showing greater periodicity at their activity levels, rather than at the corresponding mRNA levels, for over half the regulators in the networks points to extensive post-transcriptional regulations.
C. Q. Chang, Yeung Sam Hung, Mahesan Niranjan
ICASSP3
2008 Applying Cost-Sensitive Multiobjective Genetic Programming to Feature Extraction for Spam E-mail Filtering
Yang Zhang 0007, Mahesan Niranjan, Peter I. Rockett
EuroGP3
2007 Prediction of Gene Expression in Embryonic Structures of Drosophila melanogaster
abstract
Understanding how sets of genes are coordinately regulated in space and time to generate the diversity of cell types that characterise complex metazoans is a major challenge in modern biology. The use of high-throughput approaches, such as large-scale in situ hybridisation and genome-wide expression profiling via DNA microarrays, is beginning to provide insights into the complexities of development. However, in many organisms the collection and annotation of comprehensive in situ localisation data is a difficult and time-consuming task. Here, we present a widely applicable computational approach, integrating developmental time-course microarray data with annotated in situ hybridisation studies, that facilitates the de novo prediction of tissue-specific expression for genes that have no in vivo gene expression localisation data available. Using a classification approach, trained with data from microarray and in situ hybridisation studies of gene expression during Drosophila embryonic development, we made a set of predictions on the tissue-specific expression of Drosophila genes that have not been systematically characterised by in situ hybridisation experiments. The reliability of our predictions is confirmed by literature-derived annotations in FlyBase, by overrepresentation of Gene Ontology biological process annotations, and, in a selected set, by detailed gene-specific studies from the literature. Our novel organism-independent method will be of considerable utility in enriching the annotation of gene function and expression in complex multicellular organisms.
Anastassia Samsonova, Mahesan Niranjan, Steven Russell, Alvis Brazma
PLoS Comput. Biol.2
2006 Outlier Detection in Benchmark Classification Tasks
abstract
We present a new outlier detection method which is appropriate for classification problems. It combines estimating the overall probability density and sequential ranking of the data according to observed changes in performance on validation sets. The method has been implemented on ten widely used benchmark datasets and a spam email filtering application. Evaluated by six popular machine learning methods, classification performances are shown to improve after removing outliers in comparison to removing the same number of examples at random from the datasets
Mahesan Niranjan
ICASSP (5)2
2006 Automatic Face Recognition Using Stereo Images
abstract
Face recognition is an important pattern recognition problem in the study of natural and artificial learning systems. In typical optical image based face recognition systems, the systematic variability that arises from representing the three dimensional (3D) shape of a face by a two dimensional (2D) illumination intensity matrix is treated as a random variable, and it is obtained by collecting examples of faces in different poses with respect to the camera. More sophisticated 3D recognition systems employ specialist equipment (e.g. laser scanners) to measure the shape of the face, and they perform either pattern matching in three dimensions or they use projections from 3D models to match against 2D images. It is shown here that optical images obtained with a pair of stereo cameras may be used to extract depth information in the form of disparity values, and thereby significantly enhance the performance of a face recognition system
Anjali Bharatkumar Samani, Joab R. Winkler, Mahesan Niranjan
ICASSP (5)3
2004 Reducing the variability in cDNA microarray image processing by Bayesian inference
abstract
Abstract Motivation: Gene expression levels are obtained from microarray experiments through the extraction of pixel intensities from a scanned image of the slide. It is widely acknowledged that variabilities can occur in expression levels extracted from the same images by different users with the same software packages. These inconsistencies arise due to differences in the refinement of the placement of the microarray ‘grids’. We introduce a novel automated approach to the refinement of grid placements that is based upon the use of Bayesian inference for determining the size, shape and positioning of the microarray ‘spots’, capturing uncertainty that can be passed to downstream analysis. Results: Our experiments demonstrate that variability between users can be significantly reduced using the approach. The automated nature of the approach also saves hours of researchers’ time normally spent in refining the grid placement. Availability: A MATLAB implementation of the algorithm and tiff images of the slides used in our experiments, as well as the code necessary to recreate them are available for non-commercial use from http://www.dcs.shef.ac.uk/~neil/VIS
Neil D. Lawrence, Marta Milo, Mahesan Niranjan, Penny Rashbass, Stephan Soullier
Bioinform.3
2003 Sequential Bayesian Decoding with a Population of Neurons
abstract
Population coding is a simplified model of distributed information processing in the brain. This study investigates the performance and implementation of a sequential Bayesian decoding (SBD) paradigm in the framework of population coding. In the first step of decoding, when no prior knowledge is available, maximum likelihood inference is used; the result forms the prior knowledge of stimulus for the second step of decoding. Estimates are propagated sequentially to apply maximum a posteriori (MAP) decoding in which prior knowledge for any step is taken from estimates from the previous step. Not only do we analyze the performance of SBD, obtaining the optimal form of prior knowledge that achieves the best estimation result, but we also investigate its possible biological realization, in the sense that all operations are performed by the dynamics of a recurrent network. In order to achieve MAP, a crucial point is to identify a mechanism that propagates prior knowledge. We find that this could be achieved by short-term adaptation of network weights according to the Hebbian learning rule. Simulation results on both constant and time-varying stimulus support the analysis.
Si Wu 0001, Danmei Chen, Mahesan Niranjan, Shun-ichi Amari
Neural Comput.3
2002 Diphone subspace mixture trajectory models for HMM complementation
Klaus Reinhard, Mahesan Niranjan
Speech Commun.2
2001 Extractive summarization of voicemail using lexical and prosodic feature subset selection
abstract
This paper presents a novel data-driven approach to summarizing spoken audio transcripts utilizing lexical and prosodic features. The former are obtained from a speech recognizer and the latter are extracted automatically from speech waveforms. We employ a feature subset selection algorithm, based on ROC curves, which examines different combinations of features at different target operating conditions. The approach is evaluated on the IBM Voicemail corpus, demonstrating that it is possible and desirable to avoid complete commitment to a single best classifier or feature set. 1.
Konstantinos Koumpis, Steve Renals, Mahesan Niranjan
INTERSPEECH3
2001 Speech enhancement using a Bayesian evidence approach
Gaafar M. K. Saleh, Mahesan Niranjan
Comput. Speech Lang.2
2000 Matched filter design for diphone subspace models
abstract
Considering the perceptual importance of phonetic transitions as minimal contextual variant units, this paper addresses the problem by modelling explicitly interphone dynamics covered in diphones. Subspace projections based on a time-constrained PCA (TC-PCA) are developed which focus on the temporal evolution. They reveal characteristic trajectories present in a low-dimensional spectral representation facilitating robust parameter estimation and simultaneously optimise the discriminant information. A matched filter design is applied to a multiple hypotheses rescoring scheme which enables operating in very low-dimensional parameter space. Using such multiple hypotheses paradigm the complementary information effectiveness of modelling explicitly inter-phone dynamics covered in diphones can be shown using the TIMIT database, resulting in improved phone error rates.
Klaus Reinhard, Mahesan Niranjan
ICASSP2
2000 Data-dependent kernels in svm classification of speech patterns
abstract
Support Vector Machines (SVMs) have recently proved to be powerful pattern classification tools with a strong connection to statistical learning theory. One of the hurdles to using SVMs in speech recognition, and a crucial aspect of SVM design in general, is the choice of the kernel function for non-separable data, and the setting of its parameters. This is often based on experience or a potentially costly search. This paper gives some experimental justification for the Fisher kernels proposed in [4]; kernels are obtained and their extra regularisation and use of labelled and un- labelled data discussed. Fisher kernels are derived from generarive probability models of the data, and are a firststep to implementing kernels for variable length sequences. 1.
Nathan Smith, Mahesan Niranjan
INTERSPEECH2
2000 Hierarchical Bayesian Models for Regularization in Sequential Learning
abstract
We show that a hierarchical Bayesian modeling approach allows us to perform regularization in sequential learning. We identify three inference levels within this hierarchy: model selection, parameter estimation, and noise estimation. In environments where data arrive sequentially, techniques such as cross validation to achieve regularization or model selection are not possible. The Bayesian approach, with extended Kalman filtering at the parameter estimation level, allows for regularization within a minimum variance framework. A multilayer perceptron is used to generate the extended Kalman filter nonlinear measurements mapping. We describe several algorithms at the noise estimation level that allow us to implement on-line regularization. We also show the theoretical links between adaptive noise estimation in extended Kalman filtering, multiple adaptive learning rates, and multiple smoothing regularization coefficients.
João F. G. de Freitas, Mahesan Niranjan, Andrew H. Gee
Neural Comput.2
2000 Sequential Monte Carlo Methods to Train Neural Network Models
abstract
We discuss a novel strategy for training neural networks using sequential Monte Carlo algorithms and propose a new hybrid gradient descent sampling importance resampling algorithm (HySIR). In terms of computational time and accuracy, the hybrid SIR is a clear improvement over conventional sequential Monte Carlo techniques. The new algorithm may be viewed as a global optimization strategy that allows us to learn the probability distributions of the network weights and outputs in a sequential framework. It is well suited to applications involving on-line, nonlinear, and nongaussian signal processing. We show how the new algorithm outperforms extended Kalman filter training on several problems. In particular, we address the problem of pricing option contracts, traded in financial markets. In this context, we are able to estimate the one-step-ahead probability density functions of the options prices.
João F. G. de Freitas, Mahesan Niranjan, Andrew H. Gee, Arnaud Doucet
Neural Comput.2
1999 Hybrid sequential Monte Carlo/Kalman methods to train neural networks in non-stationary environments
abstract
We propose a novel sequential algorithm for training neural networks in non-stationary environments. The approach is based on a Monte Carlo method known as the sampling-importance resampling simulation algorithm. We derive our algorithm using a Bayesian framework, which allows us to learn the probability density functions of the network weights and outputs. Consequently, it is possible to compute various statistical estimates including centroids, modes, confidence intervals and kurtosis. The algorithm performs a global search for minima in parameter space by monitoring the errors and gradients at several points in the error surface. This global optimisation strategy is shown to perform better than local optimisation paradigms such as the extended Kalman filter.
João F. G. de Freitas, Mahesan Niranjan, Andrew H. Gee
ICASSP2
1999 Sequential Bayesian computation of logistic regression models
abstract
The extended Kalman filter (EKF) algorithm for identification of a state space model is shown to be a sensible tool in estimating a logistic regression model sequentially. A Gaussian probability density over the parameters of the logistic model is propagated on a sample by sample basis. Two other approaches, the Laplace approximation and the variational approximation are compared with the state space formulation. Features of the latter approach, such as the possibility of inferring noise levels by maximising the "innovation probability" are indicated. Experimental illustrations of these ideas on a synthetic problem and two real world problems are discussed.
Mahesan Niranjan
ICASSP1
1999 Diphone multi-trajectory subspace models
abstract
We report on the extension of capturing speech transitions embedded in diphones using trajectory models. The slowly varying dynamics of spectral trajectories carry much discriminant information that is very crudely modelled by traditional approaches such as HMMs. We improved our methodology of explicitly capturing the trajectory of short time spectral parameter vectors introducing multi-trajectory concepts in a probabilistic framework. Optimal subspace selection is presented which finds the most discriminant plane for classification. Using the E-set from the TIMIT database results suggest that discriminant information is preserved in the subspace.
Klaus Reinhard, Mahesan Niranjan
ICASSP2
1999 Diphone subspace models for phone-based HMM complementation
abstract
Considering the perceptual importance of phonetic transitions as minimal contextual variant units, this paper addresses the problem by modelling explicitly interphone dynamics covered in diphones.Subspace projections based on a time-constrained PCA (TC-PCA) are developed which focus on the temporal evolution.They reveal characteristic trajectories present i n a l o wdimensional spectral representation facilitating robust parameter estimation and simultaneously optimise the discriminant information.The applied multiple hypotheses rescoring scheme enables operating in very low-dimensional parameter space.Using such multiple hypotheses paradigm the complementary information eectiveness of modelling explicitly inter-phone dynamics covered in diphones can be shown using the TIMIT database, resulting in improved phone error rates.
Klaus Reinhard, Mahesan Niranjan
EUROSPEECH2
1999 Speech Modelling Using Subspace and EM Techniques
Gavin Smith, João F. G. de Freitas, Tony Robinson, Mahesan Niranjan
NIPS4
1999 Parametric subspace modeling of speech transitions
Klaus Reinhard, Mahesan Niranjan
Speech Commun.2
1998 Realisable Classifiers: Improving Operating Performance on Variable Cost Problems
abstract
A novel method is described for obtaining superior classification performance over a variable range of classification costs. By analysis of a set of existing classifiers using a receiver operating characteristic (###)curve,a set of new realisable classifiers may be obtained by a random combination of two of the existing classifiers. These classifiers lie on the convex hull that contains the original ### points for the existing classifiers. This hull is the maximum realisable ### (#####).
Martin J. J. Scott, Mahesan Niranjan, Richard W. Prager
BMVC2
1998 Parametric subspace modelling of speech transitions
abstract
We report on attempting to capture segmental transition information for speech recognition tasks. The slowly varying dynamics of spectral trajectories carries much discriminant information that is very crudely modelled by traditional approaches such as HMMs. In attempts such as recurrent neural networks there is the hope, but not convincing demonstration, that such transitional information could be captured. We start from the very different position of explicitly capturing the trajectory of short time spectral parameter vectors on a subspace in which the temporal sequence information is preserved (time constrained principal component analysis). On this subspace, we attempt a parametric modelling of the trajectory, and compute a distance metric to perform classification of diphones. Much of the discriminant information is still retained in this subspace. This is illustrated on the isolated transitions /bee/,/dee/ and /gee/.
Klaus Reinhard, Mahesan Niranjan
ICASSP2
1998 Speech enhancement in a Bayesian framework
abstract
We present an approach for the enhancement of speech signals corrupted by additive white noise of Gaussian statistics. The speech enhancement problem is treated as a signal estimation problem within a Bayesian framework. The conventional all-pole speech production model is assumed to govern the behaviour of the clean speech signal. The additive noise level and all-pole model gain are automatically inferred during the speech enhancement process. The strength of the Bayesian approach developed in this paper lies in its ability to perform speech enhancement without the usual requirement of estimating the level of the corrupting noise from "silence" segments of the corrupted signal. The performance of the Bayesian approach is compared to that of the Lim & Oppenheim (1978) framework, to which it follows a similar iterative nature. A significant quality improvement is obtained over the Lim & Oppenheim framework.
Gaafar M. K. Saleh, Mahesan Niranjan
ICASSP2
1998 Markov chain Monte Carlo methods for speech enhancement
abstract
This paper investigates a Bayesian approach to the enhancement of speech signals corrupted by additive white Gaussian noise. Parametric models for the speech and noise processes are constructed, leading to a posterior distribution for the model parameters and uncorrupted speech samples given the observed noisy speech samples. Being analytically intractable, inferences concerning these variables are performed using Markov chain Monte Carlo (MCMC) methods. The efficiency of the sampling scheme within this framework is further improved by employing state-space techniques based on the Kalman filter.
Jaco Vermaak, Mahesan Niranjan
ICASSP2
1998 Global optimisation of neural network models via sequential sampling-importance resampling
abstract
We propose a novel strategy for training neural networks using sequential Monte Carlo algorithms. This global optimisation strategy allows us to learn the probability distribution of the network weights in a sequential framework. It is well suited to applications involving on-line, nonlinear or non-stationary signal processing. We show how the new algorithms can outperform extended Kalman filter (EKF) training.
João F. G. de Freitas, Sue Tranter, Mahesan Niranjan, Andrew H. Gee
ICSLP3
1998 Global Optimisation of Neural Network Models via Sequential Sampling
João F. G. de Freitas, Mahesan Niranjan, Arnaud Doucet, Andrew H. Gee
NIPS2
1998 Feature selection using expected attainable discrimination
David R. Lovell, Christopher R. Dance, Mahesan Niranjan, Richard W. Prager, Kevin J. Dalton, R. Derom
Pattern Recognit. Lett.3
1997 Regularisation in Sequential Learning Algorithms
João F. G. de Freitas, Mahesan Niranjan, Andrew H. Gee
NIPS2
1997 Average-Case Learning Curves for Radial Basis Function Networks
abstract
The application of statistical physics to the study of the learning curves of feedforward connectionist networks has to date been concerned mostly with perceptron-like networks. Recent work has extended the theory to networks such as committee machines and parity machines, and an important direction for current and future research is the extension of this body of theory to further connectionist networks. In this article, we use this formalism to investigate the learning curves of gaussian radial basis function networks (RBFNs) having fixed basis functions. (These networks have also been called generalized linear regression models.) We address the problem of learning linear and nonlinear, realizable and unrealizable, target rules from noise-free training examples using a stochastic training algorithm. Expressions for the generalization error, defined as the expected error for a network with a given set of parameters, are derived for general gaussian RBFNs, for which all parameters, including centers and spread parameters, are adaptable. Specializing to the case of RBFNs with fixed basis functions (basis functions having parameters chosen without reference to the training examples), we then study the learning curves for these networks in the limit of high temperature.
Sean B. Holden, Mahesan Niranjan
Neural Comput.2
1996 Sequential Tracking in Pricing Financial Options using Model Based and Neural Network Approaches
Mahesan Niranjan
NIPS1
1996 Pruning with Replacement on Limited Resource Allocating Networks by F-Projections
abstract
The principle of F-projection, in sequential function estimation, provides a theoretical foundation for a class of gaussian radial basis function networks known as the resource allocating networks (RAN). The ad hoc rules for adaptively changing the size of RAN architectures can be justified from a geometric growth criterion defined in the function space. In this paper, we show that the same arguments can be used to arrive at a pruning with replacement rule for RAN architectures with a limited number of units. We illustrate the algorithm on the laser time series prediction problem of the Santa Fe competition and show that results similar to those of the winners of the competition can be obtained with pruning and replacement.
Christophe Molina, Mahesan Niranjan
Neural Comput.2
1995 Vocal tract modelling with recurrent neural networks
abstract
The speech production system is modelled using true glottal excitation as the source and a recurrent neural network to represent the vocal tract. The hidden nodes have multiple delays of one and two samples, making the network equivalent to a parallel formant synthesiser in the linear regions of the hidden node sigmoids. An ARX model identification is carried out to initialise the neural network parameters. These parameters are re-estimated in an analysis-by-synthesis framework to minimise the synthesis (output) error. Unlike other analysis-by-synthesis speech production models such as CELP, the source and filter in this approach are decoupled, enabling manipulation of the source time-scale to achieve high quality pitch changes.
T. L. Burrows, Mahesan Niranjan
ICASSP2
1995 The use of maximum a posteriori parameters in linear prediction of speech
Gaafar M. K. Saleh, Mahesan Niranjan, William J. Fitzgerald 0001
EUROSPEECH2
1995 On the practical applicability of VC dimension bounds
abstract
This article addresses the question of whether some recent Vapnik-Chervonenkis (VC) dimension-based bounds on sample complexity can be regarded as a practical design tool. Specifically, we are interested in bounds on the sample complexity for the problem of training a pattern classifier such that we can expect it to perform valid generalization. Early results using the VC dimension, while being extremely powerful, suffered from the fact that their sample complexity predictions were rather impractical. More recent results have begun to improve the situation by attempting to take specific account of the precise algorithm used to train the classifier. We perform a series of experiments based on a task involving the classification of sets of vowel formant frequencies. The results of these experiments indicate that the more recent theories provide sample complexity predictions that are significantly more applicable in practice than those provided by earlier theories; however, we also find that the recent theories still have significant shortcomings.
Sean B. Holden, Mahesan Niranjan
Neural Comput.2
1995 On the statistical physics of radial basis function networks
Sean B. Holden, Mahesan Niranjan
Neural Process. Lett.2
1994 Recursive tracking of formants in speech signals
abstract
We report on an approach to recursively track parameters of a cascade formant model. The work follows from that of Rigoll (1986) who showed how an extended Kalman filter (EKF) may be used for recursive estimation of formants. The success of this approach depends on our ability to tune the model noise variances properly. The approach also fails when there is a mismatch between the complexity of the data and that of the model (i.e. wrong number of formants). We show how a multiple model (MM) approach may be used to overcome these problems. We run several models in parallel and use the innovation probabilities of the EKF to recursively evaluate the likelihoods of each of the models. Experimental results demonstrate the feasibility of the approach; accurate switching between models and good tracking of the formants is achieved.>
Mahesan Niranjan, Ingemar J. Cox, Sunita L. Hingorani
ICASSP (2)1
1994 On the design of nonlinear speech predictors with recurrent nets
abstract
A dynamic, nonlinear speech predictor trained with real-time recurrent learning (RTRL) can achieve 2-2.5 dB better predictive gain than a conventional linear predictor. The drawback of the RTRL is that it requires a great deal of computation. For a predictor consisting of N recurrent units, the computational complexity is about O(N/sup 4/). We propose a simplified RTRL by investigating the evolution process of the gradient in a recurrent net and reduce the computational complexity to O(N/sup 3/). On a number of prediction tasks with speech signals, we show that that the simplified RTRL obtains the same prediction accuracy as the RTRL algorithm.>
Lizhong Wu, Mahesan Niranjan
ICASSP (2)2
1994 Fully vector-quantized neural network-based code-excited nonlinear predictive speech coding
abstract
Recent studies have shown that nonlinear predictors can achieve about 2-3 dB improvement in speech prediction over conventional linear predictors. In this paper, we exploit the advantage of the nonlinear prediction capability of neural networks and apply it to the design of improved predictive speech coders. Our studies concentrate on the following three aspects: (a) the development of short-term (formant) and long-term (pitch) nonlinear predictive vector quantizers (b) the analysis of the output variance of the nonlinear predictive filter with respect to the input disturbance (c) the design of nonlinear predictive speech coders. The above studies have resulted in a fully vector-quantized, code-excited, nonlinear predictive speech coder. Performance evaluations and comparisons with linear predictive speech coding are presented. These tests have shown the applicability of nonlinear prediction in speech coding and the improvement in coding performance.>
Lizhong Wu, Mahesan Niranjan, Frank Fallside
IEEE Trans. Speech Audio Process.2
1993 A Function Estimation Approach to Sequential Learning with Neural Networks
abstract
In this paper, we investigate the problem of optimal sequential learning, viewed as a problem of estimating an underlying function sequentially rather than estimating a set of parameters of the neural network. First, we arrive at a suboptimal solution to the sequential estimate that can be mapped by a growing gaussian radial basis function (GaRBF) network. This network adds hidden units for each observation. The function space approach in which the estimates are represented as vectors in a function space is used in developing a growth criterion to limit its growth. A simplification of the criterion leads to two joint criteria on the distance of the present pattern and the existing unit centers in the input space and on the approximation error of the network for the given observation to be satisfied together. This network is similar to the resource allocating network (RAN) (Platt 1991a) and hence RAN can be interpreted from a function space approach to sequential learning. Second, we present an enhancement to the RAN. The RAN either allocates a new unit based on the novelty of an observation or adapts the network parameters by the LMS algorithm. The function space interpretation of the RAN lends itself to an enhancement of the RAN in which the extended Kalman filter (EKF) algorithm is used in place of the LMS algorithm. The performance of the RAN and the enhanced network are compared in the experimental tasks of function approximation and time-series prediction demonstrating the superior performance of the enhanced network with fewer number of hidden units. The approach adopted here has led us toward the minimal network required for a sequential learning problem.
Visakan Kadirkamanathan, Mahesan Niranjan
Neural Comput.2
1992 Models of dynamic complexity for time-series prediction (neural networks)
abstract
A model of dynamic complexity, a growing Gaussian radial basis function (GRBF) network, is developed by analyzing sequential learning in the function space. The criteria for adding a new basis function to the model are based on the angle formed between a new basis function and the existing basis functions and also on the prediction error. When a new basis function is not added the model parameters are adapted by the extended Kalman filter (EKF) algorithm. This model is similar to the resource allocating network (RAN) and hence this work provides an alternative interpretation to the RAN. An enhancement to the RAN is suggested where RAN is combined with EKF. The RAN and its variants are applied to the task of predicting the logistic map and the Mackey-Glass chaotic time-series, and the advantages of the enhanced model are demonstrated.>
Visakan Kadirkamanathan, Mahesan Niranjan, Frank Fallside
ICASSP2
1991 Nonlinear adaptive filtering in nonstationary environments
abstract
The relationship of the F-projections adaptive algorithm to the LMS (least mean square), RLS (recursive least squares), and Kalman algorithms is investigated. A recursive form of nonlinear least squares is developed, and the conditions under which the F-projections algorithm becomes equivalent to it are established. A radial basis function neural network is used as a nonlinear model in analyzing time series under nonstationary environments. The performances of the F-projections and the extended Kalman algorithms for this nonlinear model in predicting a chaotic series and in tracking a time-varying system are compared.>
Visakan Kadirkamanathan, Mahesan Niranjan
ICASSP2
1991 A nonlinear model for time series prediction and signal interpolation
abstract
The approach is an extension of the method of radial basis functions. Parameter estimation for the nonlinear predictor is performed by a gradient descent over a mean squared error measure, starting from a random initialization of the parameters. Results on predicting segments of speech data and the sunspot series are presented and compared to a linear predictor. An approach to adaptive estimation of the model by means of an extended Kalman filter is presented. In terms of prediction residual, the nonlinear predictor is found to perform significantly better than a linear model with the same number of parameters. Difficulties in applying this model in speech processing are discussed.>
Mahesan Niranjan, Visakan Kadirkamanathan
ICASSP1
1990 CELP coding with adaptive output-error model identification
abstract
An approach to enhancing the code excited linear prediction (CELP) model for speech coding is presented. Adaptive infinite impulse response (IIR) filtering is used to reestimate synthesis filter parameters to minimize the output-error. Thus, not only the excitation sequence, but part of the filter is also computed in an analysis by synthesis framework. The method is extended to modeling voiced speech with a stylized glottal excitation. This approach to modeling voiced speech gives the potential of manipulating pitch while retaining a high quality.>
Mahesan Niranjan
ICASSP1
1990 Sequential Adaptation of Radial Basis Function Networks
Visakan Kadirkamanathan, Mahesan Niranjan, Frank Fallside
NIPS2
1990 A theoretical investigation into the performance of the Hopfield model
abstract
An analysis is made of the behavior of the Hopfield model as a content-addressable memory (CAM) and as a method of solving the traveling salesman problem (TSP). The analysis is based on the geometry of the subspace set up by the degenerate eigenvalues of the connection matrix. The dynamic equation is shown to be equivalent to a projection of the input vector onto this subspace. In the case of content-addressable memory, it is shown that spurious fixed points can occur at any corner of the hypercube that is on or near the subspace spanned by the memory vectors. Analysed is why the network can frequently converge to an invalid solution when applied to the traveling salesman problem energy function. With these expressions, the network can be made robust and can reliably solve the traveling salesman problem with tour sizes of 50 cities or more.
Sreeram V. B. Aiyer, Mahesan Niranjan, Frank Fallside
IEEE Trans. Neural Networks2
1989 Temporal decomposition: a framework for enhanced speech recognition
abstract
Short term spectral analysis of source-filter modeling gives a parameterized description of the acoustic signal in terms of a sequence of vectors. These parameter vectors change slowly with time corresponding to a slowly moving vocal tract. The authors consider a model (temporal decomposition) that approximates the time variation by a set of target vectors and interpolation functions that overlap in time. They present a geometric interpretation of the approach, describe an algorithm for decomposing a given utterance into parameters of such a model, and discuss how such modeling can be used in speech recognition systems.>
Mahesan Niranjan, Frank Fallside
ICASSP1