EDBT 2026 Demo / reviewers in the wild / expert
Diego H. Milone
dblp:08/4690
· DBLP profile ↗
61ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0003-2182-4351ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 30 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ET-Pfam: ensemble transfer learning for protein family predictionabstractMOTIVATION: Due to the rapid growth of sequence generation, which has surpassed the expert curators ability to manually review and annotate them, the computational annotation of proteins remains a significant challenge in bioinformatics nowadays. The Pfam database contains a large collection of proteins that are annotated with domain families through profile Hidden Markov models (pHMMs). Using the aligned sequences of a curated family, one HMM is trained independently for each family, missing the opportunity of learning patterns across families, i.e. from a complete view of all the dataset. As an alternative, some deep learning (DL) models have been recently proposed, nevertheless with simple representations of the inputs and moderate improvements in performance. RESULTS: In this work, we present ET-Pfam, a novel approach based on transfer learning and ensembles of multiple DL classifiers to predict functional families in the Pfam database. Several base DL models are first trained using learned representations from protein large language models. Then, the base models are integrated using classical ensemble strategies and novel voting approaches by learning weights for each model and for each Pfam family. Results demonstrate that the proposed ET-Pfam method can consistently diminish error rates compared to individual DL models, boosting prediction performance. Among the novel ensemble strategies presented here, the learned weights by family voting achieved the best performance, with the lowest error rate (7.00%), significantly surpassing the best individual base model error (12.91%) and competitors of the state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Data and source code are available at https://github.com/sinc-lab/ET-Pfam. Sofia A. Duarte, Rosario Vitale, Sofia Escudero, Emilio Fenoy, Leandro A. Bugnon, Diego H. Milone, Georgina Stegmayer |
Bioinform. | 6 |
| 2025 | Comprehensive benchmarking of large language models for RNA secondary structure predictionabstractIn recent years, inspired by the success of large language models (LLMs) for DNA and proteins, several LLMs for RNA have also been developed. These models take massive RNA datasets as inputs and learn, in a self-supervised way, how to represent each RNA base with a semantically rich numerical vector. This is done under the hypothesis that obtaining high-quality RNA representations can enhance data-costly downstream tasks, such as the fundamental RNA secondary structure prediction problem. However, existing RNA-LLM have not been evaluated for this task in a unified experimental setup. Since they are pretrained models, assessment of their generalization capabilities on new structures is a crucial aspect. Nonetheless, this has been just partially addressed in literature. In this work we present a comprehensive experimental and comparative analysis of pretrained RNA-LLM that have been recently proposed. We evaluate the use of these representations for the secondary structure prediction task with a common deep learning architecture. The RNA-LLM were assessed with increasing generalization difficulty on benchmark datasets. Results showed that two LLMs clearly outperform the other models, and revealed significant challenges for generalization in low-homology scenarios. Moreover, in this study we provide curated benchmark datasets of increasing complexity and a unified experimental setup for this scientific endeavor. Source code and curated benchmark datasets with increasing complexity are available in the repository: https://github.com/sinc-lab/rna-llm-folding/. Luciano I Zablocki, Leandro A. Bugnon, Matias Gerard, Leandro E. Di Persia, Georgina Stegmayer, Diego H. Milone |
Briefings Bioinform. | 6 |
| 2025 | Multi-view hybrid graph convolutional network for volume-to-mesh reconstruction in cardiovascular MRI
Nicolás Gaggion, Benjamin A. Matheson, Yan Xia 0002, Rodrigo Bonazzola, Nishant Ravikumar, Zeike A. Taylor, Diego H. Milone, Alejandro F. Frangi, Enzo Ferrante |
Medical Image Anal. | 7 |
| 2024 | sincFold: end-to-end learning of short- and long-range interactions in RNA secondary structureabstractMOTIVATION: Coding and noncoding RNA molecules participate in many important biological processes. Noncoding RNAs fold into well-defined secondary structures to exert their functions. However, the computational prediction of the secondary structure from a raw RNA sequence is a long-standing unsolved problem, which after decades of almost unchanged performance has now re-emerged due to deep learning. Traditional RNA secondary structure prediction algorithms have been mostly based on thermodynamic models and dynamic programming for free energy minimization. More recently deep learning methods have shown competitive performance compared with the classical ones, but there is still a wide margin for improvement. RESULTS: In this work we present sincFold, an end-to-end deep learning approach, that predicts the nucleotides contact matrix using only the RNA sequence as input. The model is based on 1D and 2D residual neural networks that can learn short- and long-range interaction patterns. We show that structures can be accurately predicted with minimal physical assumptions. Extensive experiments were conducted on several benchmark datasets, considering sequence homology and cross-family validation. sincFold was compared with classical methods and recent deep learning models, showing that it can outperform the state-of-the-art methods. Leandro A. Bugnon, Leandro E. Di Persia, Matias Gerard, Jonathan Raad, Santiago Prochetto, Emilio Fenoy, Uciel Chorostecki, Federico Ariel, Georgina Stegmayer, Diego H. Milone |
Briefings Bioinform. | 10 |
| 2024 | Evaluating large language models for annotating proteinsabstractIn UniProtKB, up to date, there are more than 251 million proteins deposited. However, only 0.25% have been annotated with one of the more than 15000 possible Pfam family domains. The current annotation protocol integrates knowledge from manually curated family domains, obtained using sequence alignments and hidden Markov models. This approach has been successful for automatically growing the Pfam annotations, however at a low rate in comparison to protein discovery. Just a few years ago, deep learning models were proposed for automatic Pfam annotation. However, these models demand a considerable amount of training data, which can be a challenge with poorly populated families. To address this issue, we propose and evaluate here a novel protocol based on transfer learningṪhis requires the use of protein large language models (LLMs), trained with self-supervision on big unnanotated datasets in order to obtain sequence embeddings. Then, the embeddings can be used with supervised learning on a small and annotated dataset for a specialized task. In this protocol we have evaluated several cutting-edge protein LLMs together with machine learning architectures to improve the actual prediction of protein domain annotations. Results are significatively better than state-of-the-art for protein families classification, reducing the prediction error by an impressive 60% compared to standard methods. We explain how LLMs embeddings can be used for protein annotation in a concrete and easy way, and provide the pipeline in a github repo. Full source code and data are available at https://github.com/sinc-lab/llm4pfam. Rosario Vitale, Leandro A. Bugnon, Emilio Fenoy, Diego H. Milone, Georgina Stegmayer |
Briefings Bioinform. | 4 |
| 2023 | exp2GO: Improving Prediction of Functions in the Gene Ontology With Expression DataabstractThe computational methods for the prediction of gene function annotations aim to automatically find associations between a gene and a set of Gene Ontology (GO) terms describing its functions. Since the hand-made curation process of novel annotations and the corresponding wet experiments validations are very time-consuming and costly procedures, there is a need for computational tools that can reliably predict likely annotations and boost the discovery of new gene functions. This work proposes a novel method for predicting annotations based on the inference of GO similarities from expression similarities. The novel method was benchmarked against other methods on several public biological datasets, obtaining the best comparative results. exp2GO effectively improved the prediction of GO annotations in comparison to state-of-the-art methods. Furthermore, the proposal was validated with a full genome case where it was capable of predicting relevant and accurate biological functions. The repository of this project withh full data and code is available at https://github.com/sinc-lab/exp2GO. Leandro E. Di Persia, Tiago Lopez, Agustin Arce, Diego H. Milone, Georgina Stegmayer |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Improving Anatomical Plausibility in Medical Image Segmentation via Hybrid Graph Neural Networks: Applications to Chest X-Ray AnalysisabstractAnatomical segmentation is a fundamental task in medical image computing, generally tackled with fully convolutional neural networks which produce dense segmentation masks. These models are often trained with loss functions such as cross-entropy or Dice, which assume pixels to be independent of each other, thus ignoring topological errors and anatomical inconsistencies. We address this limitation by moving from pixel-level to graph representations, which allow to naturally incorporate anatomical constraints by construction. To this end, we introduce HybridGNet, an encoder-decoder neural architecture that leverages standard convolutions for image feature encoding and graph convolutional neural networks (GCNNs) to decode plausible representations of anatomical structures. We also propose a novel image-to-graph skip connection layer which allows localized features to flow from standard convolutional blocks to GCNN blocks, and show that it improves segmentation accuracy. The proposed architecture is extensively evaluated in a variety of domain shift and image occlusion scenarios, and audited considering different types of demographic domain shift. Our comprehensive experimental setup compares HybridGNet with other landmark and pixel-based models for anatomical segmentation in chest x-ray images, and shows that it produces anatomically plausible results in challenging scenarios where other models tend to fail. Nicolás Gaggion, Lucas Mansilla, Candelaria Mosquera, Diego H. Milone, Enzo Ferrante |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Secondary structure prediction of long noncoding RNA: review and experimental comparison of existing approachesabstractMOTIVATION: In contrast to messenger RNAs, the function of the wide range of existing long noncoding RNAs (lncRNAs) largely depends on their structure, which determines interactions with partner molecules. Thus, the determination or prediction of the secondary structure of lncRNAs is critical to uncover their function. Classical approaches for predicting RNA secondary structure have been based on dynamic programming and thermodynamic calculations. In the last 4 years, a growing number of machine learning (ML)-based models, including deep learning (DL), have achieved breakthrough performance in structure prediction of biomolecules such as proteins and have outperformed classical methods in short transcripts folding. Nevertheless, the accurate prediction for lncRNA still remains far from being effectively solved. Notably, the myriad of new proposals has not been systematically and experimentally evaluated. RESULTS: In this work, we compare the performance of the classical methods as well as the most recently proposed approaches for secondary structure prediction of RNA sequences using a unified and consistent experimental setup. We use the publicly available structural profiles for 3023 yeast RNA sequences, and a novel benchmark of well-characterized lncRNA structures from different species. Moreover, we propose a novel metric to assess the predictive performance of methods, exclusively based on the chemical probing data commonly used for profiling RNA structures, avoiding any potential bias incorporated by computational predictions when using dot-bracket references. Our results provide a comprehensive comparative assessment of existing methodologies, and a novel and public benchmark resource to aid in the development and comparison of future approaches. AVAILABILITY: Full source code and benchmark datasets are available at: https://github.com/sinc-lab/lncRNA-folding. CONTACT: [email protected]. Leandro A. Bugnon, Alejandro Edera, Santiago Prochetto, Matias Gerard, Jonathan Raad, Emilio Fenoy, Mariano Rubiolo, Uciel Chorostecki, Toni Gabaldón, Federico Ariel, Leandro E. Di Persia, Diego H. Milone, Georgina Stegmayer |
Briefings Bioinform. | 12 |
| 2022 | Anc2vec: embedding gene ontology terms by preserving ancestors relationshipsabstractThe gene ontology (GO) provides a hierarchical structure with a controlled vocabulary composed of terms describing functions and localization of gene products. Recent works propose vector representations, also known as embeddings, of GO terms that capture meaningful information about them. Significant performance improvements have been observed when these representations are used on diverse downstream tasks, such as the measurement of semantic similarity between GO terms and functional similarity between proteins. Despite the success shown by these approaches, existing embeddings of GO terms still fail to capture crucial structural features of the GO. Here, we present anc2vec, a novel protocol based on neural networks for constructing vector representations of GO terms by preserving three important ontological features: its ontological uniqueness, ancestors hierarchy and sub-ontology membership. The advantages of using anc2vec are demonstrated by systematic experiments on diverse tasks: visualization, sub-ontology prediction, inference of structurally related terms, retrieval of terms from aggregated embeddings, and prediction of protein-protein interactions. In these tasks, experimental results show that the performance of anc2vec representations is better than those of recent approaches. This demonstrates that higher performances on diverse tasks can be achieved by embeddings when the structure of the GO is better represented. Full source code and data are available at https://github.com/sinc-lab/anc2vec. Alejandro Edera, Diego H. Milone, Georgina Stegmayer |
Briefings Bioinform. | 2 |
| 2022 | Hierarchical deep learning for predicting GO annotations by integrating protein knowledgeabstractMOTIVATION: Experimental testing and manual curation are the most precise ways for assigning Gene Ontology (GO) terms describing protein functions. However, they are expensive, time-consuming and cannot cope with the exponential growth of data generated by high-throughput sequencing methods. Hence, researchers need reliable computational systems to help fill the gap with automatic function prediction. The results of the last Critical Assessment of Function Annotation challenge revealed that GO-terms prediction remains a very challenging task. Recent developments on deep learning are significantly breaking out the frontiers leading to new knowledge in protein research thanks to the integration of data from multiple sources. However, deep models hitherto developed for functional prediction are mainly focused on sequence data and have not achieved breakthrough performances yet. RESULTS: We propose DeeProtGO, a novel deep-learning model for predicting GO annotations by integrating protein knowledge. DeeProtGO was trained for solving 18 different prediction problems, defined by the three GO sub-ontologies, the type of proteins, and the taxonomic kingdom. Our experiments reported higher prediction quality when more protein knowledge is integrated. We also benchmarked DeeProtGO against state-of-the-art methods on public datasets, and showed it can effectively improve the prediction of GO annotations. AVAILABILITY AND IMPLEMENTATION: DeeProtGO and a case of use are available at https://github.com/gamerino/DeeProtGO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriela Alejandra Merino, Rabie Saidi, Diego H. Milone, Georgina Stegmayer, Maria Jesus Martin |
Bioinform. | 3 |
| 2022 | miRe2e: a full end-to-end deep model based on transformers for prediction of pre-miRNAsabstractMOTIVATION: MicroRNAs (miRNAs) are small RNA sequences with key roles in the regulation of gene expression at post-transcriptional level in different species. Accurate prediction of novel miRNAs is needed due to their importance in many biological processes and their associations with complicated diseases in humans. Many machine learning approaches were proposed in the last decade for this purpose, but requiring handcrafted features extraction to identify possible de novo miRNAs. More recently, the emergence of deep learning (DL) has allowed the automatic feature extraction, learning relevant representations by themselves. However, the state-of-art deep models require complex pre-processing of the input sequences and prediction of their secondary structure to reach an acceptable performance. RESULTS: In this work, we present miRe2e, the first full end-to-end DL model for pre-miRNA prediction. This model is based on Transformers, a neural architecture that uses attention mechanisms to infer global dependencies between inputs and outputs. It is capable of receiving the raw genome-wide data as input, without any pre-processing nor feature engineering. After a training stage with known pre-miRNAs, hairpin and non-harpin sequences, it can identify all the pre-miRNA sequences within a genome. The model has been validated through several experimental setups using the human genome, and it was compared with state-of-the-art algorithms obtaining 10 times better performance. AVAILABILITY AND IMPLEMENTATION: Webdemo available at https://sinc.unl.edu.ar/web-demo/miRe2e/ and source code available for download at https://github.com/sinc-lab/miRe2e. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jonathan Raad, Leandro A. Bugnon, Diego H. Milone, Georgina Stegmayer |
Bioinform. | 3 |
| 2021 | Domain Generalization via Gradient SurgeryabstractIn real-life applications, machine learning models often face scenarios where there is a change in data distribution between training and test domains. When the aim is to make predictions on distributions different from those seen at training, we incur in a domain generalization problem. Methods to address this issue learn a model using data from multiple source domains, and then apply this model to the unseen target domain. Our hypothesis is that when training with multiple domains, conflicting gradients within each mini-batch contain information specific to the individual domains which is irrelevant to the others, including the test domain. If left untouched, such disagreement may degrade generalization performance. In this work, we characterize the conflicting gradients emerging in domain shift scenarios and devise novel gradient agreement strategies based on gradient surgery to alleviate their effect. We validate our approach in image classification tasks with three multi-domain datasets, showing the value of the proposed agreement strategy in enhancing the generalization capability of deep learning models in domain shift scenarios. Lucas Mansilla, Rodrigo Echeveste, Diego H. Milone, Enzo Ferrante |
ICCV | 3 |
| 2021 | Hybrid Graph Convolutional Neural Networks for Landmark-Based Anatomical SegmentationabstractIn this work we address the problem of landmark-based segmentation for anatomical structures. We propose HybridGNet, an encoder-decoder neural architecture which combines standard convolutions for image feature encoding, with graph convolutional neural networks to decode plausible representations of anatomical structures. We benchmark the proposed architecture considering other standard landmark and pixel-based models for anatomical segmentation in chest x-ray images, and found that HybridGNet is more robust to image occlusions. We also show that it can be used to construct landmark-based segmentations from pixel level annotations. Our experimental results suggest that HybridGNet produces accurate and anatomically plausible landmark-based segmentations, by naturally incorporating shape constraints within the decoding process via spectral convolutions. Nicolás Gaggion, Lucas Mansilla, Diego H. Milone, Enzo Ferrante |
MICCAI (1) | 3 |
| 2021 | Genome-wide discovery of pre-miRNAs: comparison of recent approaches based on machine learningabstractMOTIVATION: The genome-wide discovery of microRNAs (miRNAs) involves identifying sequences having the highest chance of being a novel miRNA precursor (pre-miRNA), within all the possible sequences in a complete genome. The known pre-miRNAs are usually just a few in comparison to the millions of candidates that have to be analyzed. This is of particular interest in non-model species and recently sequenced genomes, where the challenge is to find potential pre-miRNAs only from the sequenced genome. The task is unfeasible without the help of computational methods, such as deep learning. However, it is still very difficult to find an accurate predictor, with a low false positive rate in this genome-wide context. Although there are many available tools, these have not been tested in realistic conditions, with sequences from whole genomes and the high class imbalance inherent to such data. RESULTS: In this work, we review six recent methods for tackling this problem with machine learning. We compare the models in five genome-wide datasets: Arabidopsis thaliana, Caenorhabditis elegans, Anopheles gambiae, Drosophila melanogaster, Homo sapiens. The models have been designed for the pre-miRNAs prediction task, where there is a class of interest that is significantly underrepresented (the known pre-miRNAs) with respect to a very large number of unlabeled samples. It was found that for the smaller genomes and smaller imbalances, all methods perform in a similar way. However, for larger datasets such as the H. sapiens genome, it was found that deep learning approaches using raw information from the sequences reached the best scores, achieving low numbers of false positives. AVAILABILITY: The source code to reproduce these results is in: http://sourceforge.net/projects/sourcesinc/files/gwmirna Additionally, the datasets are freely available in: https://sourceforge.net/projects/sourcesinc/files/mirdata. Leandro A. Bugnon, Cristian A. Yones, Diego H. Milone, Georgina Stegmayer |
Briefings Bioinform. | 3 |
| 2021 | Novel SARS-CoV-2 encoded small RNAs in the passage to humansabstractMOTIVATION: The Severe Acute Respiratory Syndrome-Coronavirus 2 (SARS-CoV-2) has recently emerged as the responsible for the pandemic outbreak of the coronavirus disease 2019. This virus is closely related to coronaviruses infecting bats and Malayan pangolins, species suspected to be an intermediate host in the passage to humans. Several genomic mutations affecting viral proteins have been identified, contributing to the understanding of the recent animal-to-human transmission. However, the capacity of SARS-CoV-2 to encode functional putative microRNAs (miRNAs) remains largely unexplored. RESULTS: We have used deep learning to discover 12 candidate stem-loop structures hidden in the viral protein-coding genome. Among the precursors, the expression of eight mature miRNAs-like sequences was confirmed in small RNA-seq data from SARS-CoV-2 infected human cells. Predicted miRNAs are likely to target a subset of human genes of which 109 are transcriptionally deregulated upon infection. Remarkably, 28 of those genes potentially targeted by SARS-CoV-2 miRNAs are down-regulated in infected human cells. Interestingly, most of them have been related to respiratory diseases and viral infection, including several afflictions previously associated with SARS-CoV-1 and SARS-CoV-2. The comparison of SARS-CoV-2 pre-miRNA sequences with those from bat and pangolin coronaviruses suggests that single nucleotide mutations could have helped its progenitors jumping inter-species boundaries, allowing the gain of novel mature miRNAs targeting human mRNAs. Our results suggest that the recent acquisition of novel miRNAs-like sequences in the SARS-CoV-2 genome may have contributed to modulate the transcriptional reprograming of the new host upon infection. AVAILABILITY AND IMPLEMENTATION: https://github.com/sinc-lab/sarscov2-mirna-discovery. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriela Alejandra Merino, Jonathan Raad, Leandro A. Bugnon, Cristian A. Yones, Laura Kamenetzky, Juan Claus, Federico Ariel, Diego H. Milone, Georgina Stegmayer |
Bioinform. | 8 |
| 2020 | DL4papers: a deep learning approach for the automatic interpretation of scientific articlesabstractMOTIVATION: In precision medicine, next-generation sequencing and novel preclinical reports have led to an increasingly large amount of results, published in the scientific literature. However, identifying novel treatments or predicting a drug response in, for example, cancer patients, from the huge amount of papers available remains a laborious and challenging work. This task can be considered a text mining problem that requires reading a lot of academic documents for identifying a small set of papers describing specific relations between key terms. Due to the infeasibility of the manual curation of these relations, computational methods that can automatically identify them from the available literature are urgently needed. RESULTS: We present DL4papers, a new method based on deep learning that is capable of analyzing and interpreting papers in order to automatically extract relevant relations between specific keywords. DL4papers receives as input a query with the desired keywords, and it returns a ranked list of papers that contain meaningful associations between the keywords. The comparison against related methods showed that our proposal outperformed them in a cancer corpus. The reliability of the DL4papers output list was also measured, revealing that 100% of the first two documents retrieved for a particular search have relevant relations, in average. This shows that our model can guarantee that in the top-2 papers of the ranked list, the relation can be effectively found. Furthermore, the model is capable of highlighting, within each document, the specific fragments that have the associations of the input keywords. This can be very useful in order to pay attention only to the highlighted text, instead of reading the full paper. We believe that our proposal could be used as an accurate tool for rapidly identifying relationships between genes and their mutations, drug responses and treatments in the context of a certain disease. This new approach can certainly be a very useful and valuable resource for the advancement of the precision medicine field. AVAILABILITY AND IMPLEMENTATION: A web-demo is available at: http://sinc.unl.edu.ar/web-demo/dl4papers/. Full source code and data are available at: https://sourceforge.net/projects/sourcesinc/files/dl4papers/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leandro A. Bugnon, Cristian A. Yones, Jonathan Raad, Matias Gerard, Mariano Rubiolo, Gabriela Alejandra Merino, Milton Pividori, Leandro E. Di Persia, Diego H. Milone, Georgina Stegmayer |
Bioinform. | 9 |
| 2020 | Complexity measures of the mature miRNA for improving pre-miRNAs predictionabstractMOTIVATION: The discovery of microRNA (miRNA) in the last decade has certainly changed the understanding of gene regulation in the cell. Although a large number of algorithms with different features have been proposed, they still predict an impractical amount of false positives. Most of the proposed features are based on the structure of precursors of the miRNA only, not considering the important and relevant information contained in the mature miRNA. Such new kind of features could certainly improve the performance of the predictors of new miRNAs. RESULTS: This paper presents three new features that are based on the sequence information contained in the mature miRNA. We will show how these new features, when used by a classical supervised machine learning approach as well as by more recent proposals based on deep learning, improve the prediction performance in a significant way. Moreover, several experimental conditions were defined and tested to evaluate the novel features impact in situations close to genome-wide analysis. The results show that the incorporation of new features based on the mature miRNA allows to improve the detection of new miRNAs independently of the classifier used. AVAILABILITY AND IMPLEMENTATION: https://sourceforge.net/projects/sourcesinc/files/cplxmirna/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jonathan Raad, Georgina Stegmayer, Diego H. Milone |
Bioinform. | 3 |
| 2020 | Learning deformable registration of medical images with anatomical constraints
Lucas Mansilla, Diego H. Milone, Enzo Ferrante |
Neural Networks | 2 |
| 2020 | Dimensional Affect Recognition from HRV: An Approach Based on Supervised SOM and ELMabstractDimensional affect recognition is a challenging topic and current techniques do not yet provide the accuracy necessary for HCI applications. In this work we propose two new methods. The first is a novel self-organizing model that learns from similarity between features and affects. This method produces a graphical representation of the multidimensional data which may assist the expert analysis. The second method uses extreme learning machines, an emerging artificial neural network model. Aiming for minimum intrusiveness, we use only the heart rate variability, which can be recorded using a small set of sensors. The methods were validated with two datasets. The first is composed of 16 sessions with different participants and was used to evaluate the models in a classification task. The second one was the publicly available Remote Collaborative and Affective Interaction (RECOLA) dataset, which was used for dimensional affect estimation. The performance evaluation used the kappa score, unweighted average recall and the concordance correlation coefficient. The concordance coefficient on the RECOLA test partition was 0.421 in arousal and 0.321 in valence. Results show that our models outperform state-of-the-art models on the same data and provides new ways to analyze affective states. Leandro A. Bugnon, Rafael A. Calvo, Diego H. Milone |
IEEE Trans. Affect. Comput. | 3 |
| 2020 | Deep Neural Architectures for Highly Imbalanced Data in BioinformaticsabstractIn the postgenome era, many problems in bioinformatics have arisen due to the generation of large amounts of imbalanced data. In particular, the computational classification of precursor microRNA (pre-miRNA) involves a high imbalance in the classes. For this task, a classifier is trained to identify RNA sequences having the highest chance of being miRNA precursors. The big issue is that well-known pre-miRNAs are usually just a few in comparison to the hundreds of thousands of candidate sequences in a genome, which results in highly imbalanced data. This imbalance has a strong influence on most standard classifiers and, if not properly addressed, the classifier is not able to work properly in a real-life scenario. This work provides a comparative assessment of recent deep neural architectures for dealing with the large imbalanced data issue in the classification of pre-miRNAs. We present and analyze recent architectures in a benchmark framework with genomes of animals and plants, with increasing imbalance ratios up to 1:2000. We also propose a new graphical way for comparing classifiers performance in the context of high-class imbalance. The comparative results obtained show that, at a very high imbalance, deep belief neural networks can provide the best performance. Leandro A. Bugnon, Cristian A. Yones, Diego H. Milone, Georgina Stegmayer |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Predicting novel microRNA: a comprehensive comparison of machine learning approachesabstractMOTIVATION: The importance of microRNAs (miRNAs) is widely recognized in the community nowadays because these short segments of RNA can play several roles in almost all biological processes. The computational prediction of novel miRNAs involves training a classifier for identifying sequences having the highest chance of being precursors of miRNAs (pre-miRNAs). The big issue with this task is that well-known pre-miRNAs are usually few in comparison with the hundreds of thousands of candidate sequences in a genome, which results in high class imbalance. This imbalance has a strong influence on most standard classifiers, and if not properly addressed in the model and the experiments, not only performance reported can be completely unrealistic but also the classifier will not be able to work properly for pre-miRNA prediction. Besides, another important issue is that for most of the machine learning (ML) approaches already used (supervised methods), it is necessary to have both positive and negative examples. The selection of positive examples is straightforward (well-known pre-miRNAs). However, it is difficult to build a representative set of negative examples because they should be sequences with hairpin structure that do not contain a pre-miRNA. RESULTS: This review provides a comprehensive study and comparative assessment of methods from these two ML approaches for dealing with the prediction of novel pre-miRNAs: supervised and unsupervised training. We present and analyze the ML proposals that have appeared during the past 10 years in literature. They have been compared in several prediction tasks involving two model genomes and increasing imbalance levels. This work provides a review of existing ML approaches for pre-miRNA prediction and fair comparisons of the classifiers with same features and data sets, instead of just a revision of published software tools. The results and the discussion can help the community to select the most adequate bioinformatics approach according to the prediction task at hand. The comparative results obtained suggest that from low to mid-imbalance levels between classes, supervised methods can be the best. However, at very high imbalance levels, closer to real case scenarios, models including unsupervised and deep learning can provide better performance. Georgina Stegmayer, Leandro E. Di Persia, Mariano Rubiolo, Matias Gerard, Milton Pividori, Cristian A. Yones, Leandro A. Bugnon, Tadeo Rodriguez, Jonathan Raad, Diego H. Milone |
Briefings Bioinform. | 10 |
| 2019 | Clustermatch: discovering hidden relations in highly diverse kinds of qualitative and quantitative data without standardizationabstractMOTIVATION: Heterogeneous and voluminous data sources are common in modern datasets, particularly in systems biology studies. For instance, in multi-holistic approaches in the fruit biology field, data sources can include a mix of measurements such as morpho-agronomic traits, different kinds of molecules (nucleic acids and metabolites) and consumer preferences. These sources not only have different types of data (quantitative and qualitative), but also large amounts of variables with possibly non-linear relationships among them. An integrative analysis is usually hard to conduct, since it requires several manual standardization steps, with a direct and critical impact on the results obtained. These are important issues in clustering applications, which highlight the need of new methods for uncovering complex relationships in such diverse repositories. RESULTS: We designed a new method named Clustermatch to easily and efficiently perform data-mining tasks on large and highly heterogeneous datasets. Our approach can derive a similarity measure between any quantitative or qualitative variables by looking on how they influence on the clustering of the biological materials under study. Comparisons with other methods in both simulated and real datasets show that Clustermatch is better suited for finding meaningful relationships in complex datasets. AVAILABILITY AND IMPLEMENTATION: Files can be downloaded from https://sourceforge.net/projects/sourcesinc/files/clustermatch/ and https://bitbucket.org/sinc-lab/clustermatch/. In addition, a web-demo is available at http://sinc.unl.edu.ar/web-demo/clustermatch/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Milton Pividori, Andres Cernadas, Luis A. de Haro, Fernando Carrari, Georgina Stegmayer, Diego H. Milone |
Bioinform. | 6 |
| 2018 | Extreme learning machines for reverse engineering of gene regulatory networks from expression time seriesabstractMotivation: The reconstruction of gene regulatory networks (GRNs) from genes profiles has a growing interest in bioinformatics for understanding the complex regulatory mechanisms in cellular systems. GRNs explicitly represent the cause-effect of regulation among a group of genes and its reconstruction is today a challenging computational problem. Several methods were proposed, but most of them require different input sources to provide an acceptable prediction. Thus, it is a great challenge to reconstruct a GRN only from temporal gene expression data. Results: Extreme Learning Machine (ELM) is a new supervised neural model that has gained interest in the last years because of its higher learning rate and better performance than existing supervised models in terms of predictive power. This work proposes a novel approach for GRNs reconstruction in which ELMs are used for modeling the relationships between gene expression time series. Artificial datasets generated with the well-known benchmark tool used in DREAM competitions were used. Real datasets were used for validation of this novel proposal with well-known GRNs underlying the time series. The impact of increasing the size of GRNs was analyzed in detail for the compared methods. The results obtained confirm the superiority of the ELM approach against very recent state-of-the-art methods in the same experimental conditions. Availability and implementation: The web demo can be found at http://sinc.unl.edu.ar/web-demo/elm-grnnminer/. The source code is available at https://sourceforge.net/projects/sourcesinc/files/elm-grnnminer. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Mariano Rubiolo, Diego H. Milone, Georgina Stegmayer |
Bioinform. | 2 |
| 2018 | Genome-wide pre-miRNA discovery from few labeled examplesabstractMotivation: Although many machine learning techniques have been proposed for distinguishing miRNA hairpins from other stem-loop sequences, most of the current methods use supervised learning, which requires a very good set of positive and negative examples. Those methods have important practical limitations when they have to be applied to a real prediction task. First, there is the challenge of dealing with a scarce number of positive (well-known) pre-miRNA examples. Secondly, it is very difficult to build a good set of negative examples for representing the full spectrum of non-miRNA sequences. Thirdly, in any genome, there is a huge class imbalance (1: 10 000) that is well-known for particularly affecting supervised classifiers. Results: To enable efficient and speedy genome-wide predictions of novel miRNAs, we present miRNAss, which is a novel method based on semi-supervised learning. It takes advantage of the information provided by the unlabeled stem-loops, thereby improving the prediction rates, even when the number of labeled examples is low and not representative of the classes. An automatic method for searching negative examples to initialize the algorithm is also proposed so as to spare the user this difficult task. MiRNAss obtained better prediction rates and shorter execution times than state-of-the-art supervised methods. It was validated with genome-wide data from three model species, with more than one million of hairpin sequences each, thereby demonstrating its applicability to a real prediction task. Availability and implementation: An R package can be downloaded from https://cran.r-project.org/package=miRNAss. In addition, a web-demo for testing the algorithm is available at http://fich.unl.edu.ar/sinc/web-demo/mirnass. All the datasets that were used in this study and the sets of predicted pre-miRNA are available on http://sourceforge.net/projects/sourcesinc/files/mirnass. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Cristian A. Yones, Georgina Stegmayer, Diego H. Milone |
Bioinform. | 3 |
| 2018 | Blankets Joint Posterior score for learning Markov network structures
Federico Schlüter, Jan Strappa, Diego H. Milone, Facundo Bromberg |
Int. J. Approx. Reason. | 3 |
| 2018 | Inferring Unknown Biological Function by Integration of GO Annotations and Gene Expression DataabstractCharacterizing genes with semantic information is an important process regarding the description of gene products. In spite that complete genomes of many organisms have been already sequenced, the biological functions of all of their genes are still unknown. Since experimentally studying the functions of those genes, one by one, would be unfeasible, new computational methods for gene functions inference are needed. We present here a novel computational approach for inferring biological function for a set of genes with previously unknown function, given a set of genes with well-known information. This approach is based on the premise that genes with similar behaviour should be grouped together. This is known as the guilt-by-association principle. Thus, it is possible to take advantage of clustering techniques to obtain groups of unknown genes that are co-clustered with genes that have well-known semantic information (GO annotations). Meaningful knowledge to infer unknown semantic information can therefore be provided by these well-known genes. We provide a method to explore the potential function of new genes according to those currently annotated. The results obtained indicate that the proposed approach could be a useful and effective tool when used by biologists to guide the inference of biological functions for recently discovered genes. Our work sets an important landmark in the field of identifying unknown gene functions through clustering, using an external source of biological input. A simple web interface to this proposal can be found at http://fich.unl.edu.ar/sinc/webdemo/gamma-am/. Guillermo Leale, Ariel E. Bayá, Diego H. Milone, Pablo M. Granitto, Georgina Stegmayer |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Assessment of Homomorphic Analysis for Human Activity Recognition From Acceleration SignalsabstractUnobtrusive activity monitoring can provide valuable information for medical and sports applications. In recent years, human activity recognition has moved to wearable sensors to deal with unconstrained scenarios. Accelerometers are the preferred sensors due to their simplicity and availability. Previous studies have examined several classic techniques for extracting features from acceleration signals, including time-domain, time-frequency, frequency-domain, and other heuristic features. Spectral and temporal features are the preferred ones and they are generally computed from acceleration components, leaving the acceleration magnitude potential unexplored. In this study, a new type of feature extraction stage, based on homomorphic analysis, is proposed in order to exploit discriminative activity information present in acceleration signals. Homomorphic analysis can isolate the information about whole body dynamics and translate it into a compact representation, called cepstral coefficients. Experiments have explored several configurations of the proposed features, including size of representation, signals to be used, and fusion with other features. Cepstral features computed from acceleration magnitude obtained one of the highest recognition rates. In addition, a beneficial contribution was found when time-domain and moving pace information was included in the feature vector. Overall, the proposed system achieved a recognition rate of 91.21% on the publicly available SCUT-NAA dataset. To the best of our knowledge, this is the highest recognition rate on this dataset. Sebastián R. Vanrell, Diego H. Milone, Hugo Leonardo Rufiner |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Extreme learning machine prediction under high class imbalance in bioinformaticsabstractClass imbalance in machine learning is when there are significantly fewer training instances of one class in comparison to another one. In bioinformatics, there is such a problem in the computational prediction of novel microRNA (miRNAs) within a full genome. The well-known precursors miRNA (pre-miRNA) are usually only a few in comparison to the hundreds of thousands of potential candidates, which makes this task a high class imbalance classification problem. It is well-known that high class imbalance usually affects any classical supervised machine learning classifier. Thus the imbalance must be explicitly considered. Extreme Learning Machine (ELM) is a supervised artificial neural network model that has gained interest in the last years because of its high learning rate and performance. In this work, we propose a novel approach to overcome the high class imbalance in pre-miRNAs prediction data in which ELMs are used for predicting good candidates to pre-miRNA, without needing balanced data sets. Real datasets were used for validation of the proposal with several class imbalance levels. The results obtained showed the superiority of the ELM approach against very recent state-of-the-art methods in the same experimental conditions. Tadeo Rodriguez, Leandro E. Di Persia, Diego H. Milone, Georgina Stegmayer |
CLEI | 3 |
| 2017 | Feature extraction based on bio-inspired model for robust emotion recognition
Enrique Marcelo Albornoz, Diego H. Milone, Hugo Leonardo Rufiner |
Soft Comput. | 2 |
| 2017 | Emotion Recognition in Never-Seen Languages Using a Novel Ensemble Method with Emotion ProfilesabstractOver the last years, researchers have addressed emotional state identification because it is an important issue to achieve more natural speech interactive systems. There are several theories that explain emotional expressiveness as a result of natural evolution, as a social construction, or a combination of both. In this work, we propose a novel system to model each language independently, preserving the cultural properties. In a second stage, we use the concept of universality of emotions to map and predict emotions in never-seen languages. Features and classifiers widely tested for similar tasks were used to set the baselines. We developed a novel ensemble classifier to deal with multiple languages and tested it on never-seen languages. Furthermore, this ensemble uses the Emotion Profiles technique in order to map features from diverse languages in a more tractable space. The experiments were performed in a language-independent scheme. Results show that the proposed model improves the baseline accuracy, whereas its modular design allows the incorporation of a new language without having to train the whole system. Enrique Marcelo Albornoz, Diego H. Milone |
IEEE Trans. Affect. Comput. | 2 |
| 2017 | High Class-Imbalance in pre-miRNA Prediction: A Novel Approach Based on deepSOMabstractThe computational prediction of novel microRNA within a full genome involves identifying sequences having the highest chance of being a miRNA precursor (pre-miRNA). These sequences are usually named candidates to miRNA. The well-known pre-miRNAs are usually only a few in comparison to the hundreds of thousands of potential candidates to miRNA that have to be analyzed, which makes this task a high class-imbalance classification problem. The classical way of approaching it has been training a binary classifier in a supervised manner, using well-known pre-miRNAs as positive class and artificially defining the negative class. However, although the selection of positive labeled examples is straightforward, it is very difficult to build a set of negative examples in order to obtain a good set of training samples for a supervised method. In this work, we propose a novel and effective way of approaching this problem using machine learning, without the definition of negative examples. The proposal is based on clustering unlabeled sequences of a genome together with well-known miRNA precursors for the organism under study, which allows for the quick identification of the best candidates to miRNA as those sequences clustered with known precursors. Furthermore, we propose a deep model to overcome the problem of having very few positive class labels. They are always maintained in the deep levels as positive class while less likely pre-miRNA sequences are filtered level after level. Our approach has been compared with other methods for pre-miRNAs prediction in several species, showing effective predictivity of novel miRNAs. Additionally, we will show that our approach has a lower training time and allows for a better graphical navegability and interpretation of the results. A web-demo interface to try deepSOM is available at http://fich.unl.edu.ar/sinc/web-demo/deepsom/. Georgina Stegmayer, Cristian A. Yones, Laura Kamenetzky, Diego H. Milone |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | A very simple and fast way to access and validate algorithms in reproducible researchabstractThe reproducibility of research in bioinformatics refers to the notion that new methodologies/algorithms and scientific claims have to be published together with their data and source code, in a way that other researchers may verify the findings to further build more knowledge on them. The replication and corroboration of research results are key to the scientific process, and many journals are discussing the matter nowadays, taking concrete steps in this direction. In this journal itself, a recent opinion note has appeared highlighting the increasing importance of this topic in bioinformatics and computational biology, inviting the community to further discuss the matter. In agreement with that article, we would like to propose here another step into that direction with a tool that allows the automatic generation of a web interface, named web-demo, directly from source code in a simple and straightforward way. We believe this contribution can help make research not only reproducible but also more easily accessible. A web-demo associated to a published paper can accelerate an algorithm validation with real data, wide-spreading its use with just a few clicks. Georgina Stegmayer, Milton Pividori, Diego H. Milone |
Briefings Bioinform. | 3 |
| 2016 | A new index for clustering validation with overlapped clusters
David Nazareno Campo, Georgina Stegmayer, Diego H. Milone |
Expert Syst. Appl. | 3 |
| 2016 | Multi-objective optimisation of wavelet features for phoneme recognitionabstractState‐of‐the‐art speech representations provide acceptable recognition results under optimal conditions, though their performance in adverse conditions still needs to be improved. In this direction, many advances involving wavelet processing have been reported, showing significant improvements in classification performance for different kinds of signals. However, for speech signals, the problem of finding a convenient wavelet‐based representation is still an open challenge. This study proposes the use of a multi‐objective genetic algorithm for the optimisation of a wavelet‐based representation of speech. The most relevant features are selected from a complete wavelet packet decomposition in order to maximise phoneme classification performance. Classification results for English phonemes, in different noise conditions, show significant improvements compared with well‐known speech representations. Leandro Daniel Vignolo, Hugo Leonardo Rufiner, Diego H. Milone |
IET Signal Process. | 3 |
| 2016 | Diversity control for improving the analysis of consensus clustering
Milton Pividori, Georgina Stegmayer, Diego H. Milone |
Inf. Sci. | 3 |
| 2016 | Feature optimisation for stress recognition in speech
Leandro Daniel Vignolo, S. R. Mahadeva Prasanna, Samarendra Dandapat, Hugo Leonardo Rufiner, Diego H. Milone |
Pattern Recognit. Lett. | 5 |
| 2016 | Using multiple frequency bins for stabilization of FD-ICA algorithms
Leandro E. Di Persia, Diego H. Milone |
Signal Process. | 2 |
| 2015 | Wavelet shrinkage using adaptive structured sparsity constraintsabstractStructured sparsity approaches have recently received much attention in the statistics, machine learning, and signal processing communities. A common strategy is to exploit or assume prior information about structural dependencies inherent in the data; the solution is encouraged to behave as such by the inclusion of an appropriate regularisation term which enforces structured sparsity constraints over sub-groups of data. An important variant of this idea considers the tree-like dependency structures often apparent in wavelet decompositions. However, both the constituent groups and their associated weights in the regularisation term are typically defined a priori. We here introduce an adaptive wavelet denoising framework whereby a sparsity-inducing regulariser is modified based on information extracted from the signal itself. In particular, we use the same wavelet decomposition to detect the location of salient features in the signal, such as jumps or sharp bumps. Given these locations, the weights in the regulariser associated to the groups of coefficients that cover these time locations are modified in order to favour retention of those coefficients. Denoising experiments show that, not only does the adaptive method preserve the salient features better than the non-adaptive constraints, but it also delivers significantly better shrinkage over the signal as a whole. Diego Tomassi, Diego H. Milone, James D. B. Nelson |
Signal Process. | 2 |
| 2015 | Mining Gene Regulatory Networks by Neural Modeling of Expression Time-SeriesabstractDiscovering gene regulatory networks from data is one of the most studied topics in recent years. Neural networks can be successfully used to infer an underlying gene network by modeling expression profiles as times series. This work proposes a novel method based on a pool of neural networks for obtaining a gene regulatory network from a gene expression dataset. They are used for modeling each possible interaction between pairs of genes in the dataset, and a set of mining rules is applied to accurately detect the subjacent relations among genes. The results obtained on artificial and real datasets confirm the method effectiveness for discovering regulatory networks from a proper modeling of the temporal dynamics of gene expression profiles. Mariano Rubiolo, Diego H. Milone, Georgina Stegmayer |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2014 | Improving clustering with metabolic pathway dataabstractBACKGROUND: It is a common practice in bioinformatics to validate each group returned by a clustering algorithm through manual analysis, according to a-priori biological knowledge. This procedure helps finding functionally related patterns to propose hypotheses for their behavior and the biological processes involved. Therefore, this knowledge is used only as a second step, after data are just clustered according to their expression patterns. Thus, it could be very useful to be able to improve the clustering of biological data by incorporating prior knowledge into the cluster formation itself, in order to enhance the biological value of the clusters. RESULTS: A novel training algorithm for clustering is presented, which evaluates the biological internal connections of the data points while the clusters are being formed. Within this training algorithm, the calculation of distances among data points and neurons centroids includes a new term based on information from well-known metabolic pathways. The standard self-organizing map (SOM) training versus the biologically-inspired SOM (bSOM) training were tested with two real data sets of transcripts and metabolites from Solanum lycopersicum and Arabidopsis thaliana species. Classical data mining validation measures were used to evaluate the clustering solutions obtained by both algorithms. Moreover, a new measure that takes into account the biological connectivity of the clusters was applied. The results of bSOM show important improvements in the convergence and performance for the proposed clustering method in comparison to standard SOM training, in particular, from the application point of view. CONCLUSIONS: Analyses of the clusters obtained with bSOM indicate that including biological information during training can certainly increase the biological value of the clusters found with the proposed method. It is worth to highlight that this fact has effectively improved the results, which can simplify their further analysis.The algorithm is available as a web-demo at http://fich.unl.edu.ar/sinc/web-demo/bsom-lite/. The source code and the data sets supporting the results of this article are available at http://sourceforge.net/projects/sourcesinc/files/bsom. Diego H. Milone, Georgina Stegmayer, Mariana G. Lopez, Laura Kamenetzky, Fernando Carrari |
BMC Bioinform. | 1 |
| 2013 | Clustering biological data with SOMs: On topology preservation in non-linear dimensional reduction
Diego H. Milone, Georgina Stegmayer, Laura Kamenetzky, Mariana G. Lopez, Fernando Carrari |
Expert Syst. Appl. | 1 |
| 2013 | Automatic recognition of quarantine citrus diseases
Georgina Stegmayer, Diego H. Milone, Sergio Garran, Lourdes Burdyn |
Expert Syst. Appl. | 2 |
| 2013 | Genetic wavelet packets for speech recognition
Leandro Daniel Vignolo, Diego H. Milone, Hugo Leonardo Rufiner |
Expert Syst. Appl. | 2 |
| 2013 | Feature selection for face recognition based on multi-objective evolutionary wrappers
Leandro Daniel Vignolo, Diego H. Milone, Jacob Scharcanski |
Expert Syst. Appl. | 2 |
| 2013 | Compressing arrays of classifiers using Volterra-neural network: application to face recognition
Mariano Rubiolo, Georgina Stegmayer, Diego H. Milone |
Neural Comput. Appl. | 3 |
| 2012 | An evolutionary wrapper for feature selection in face recognition applicationsabstractActive shape models is an adaptive shape-matching technique that has been used for locating facial features in images. However, when a number of features is extracted for each landmark point, distortions caused by noise or illumination, and the dimensionality of the final representation, have a negative impact in the performance of a classifier. In this paper, an evolutionary wrapper for selection of the most relevant set of features for face recognition is presented. The proposed strategy explores the space of multiple feasible selections using genetic algorithms. Experimental results show that the proposed approach allows to improve the classification performance in comparison with another enhanced method and a state of the art face recognition approach. Leandro Daniel Vignolo, Diego H. Milone, Carlos A. R. Behaine, Jacob Scharcanski |
SMC | 2 |
| 2012 | Bioinspired sparse spectro-temporal representation of speech for robust classification
César Ernesto Martínez, John Goddard Close, Diego H. Milone, Hugo Leonardo Rufiner |
Comput. Speech Lang. | 3 |
| 2012 | A Biologically Inspired Validity Measure for Comparison of Clustering Methods over Metabolic Data SetsabstractIn the biological domain, clustering is based on the assumption that genes or metabolites involved in a common biological process are coexpressed/coaccumulated under the control of the same regulatory network. Thus, a detailed inspection of the grouped patterns to verify their memberships to well-known metabolic pathways could be very useful for the evaluation of clusters from a biological perspective. The aim of this work is to propose a novel approach for the comparison of clustering methods over metabolic data sets, including prior biological knowledge about the relation among elements that constitute the clusters. A way of measuring the biological significance of clustering solutions is proposed. This is addressed from the perspective of the usefulness of the clusters to identify those patterns that change in coordination and belong to common pathways of metabolic regulation. The measure summarizes in a compact way the objective analysis of clustering methods, which respects coherence and clusters distribution. It also evaluates the biological internal connections of such clusters considering common pathways. The proposed measure was tested in two biological databases using three clustering methods. Georgina Stegmayer, Diego H. Milone, Laura Kamenetzky, Mariana G. Lopez, Fernando Carrari |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | Spoken emotion recognition using hierarchical classifiers
Enrique Marcelo Albornoz, Diego H. Milone, Hugo Leonardo Rufiner |
Comput. Speech Lang. | 2 |
| 2010 | Compressing a neural network classifier using a Volterra-Neural Network modelabstractModel compression is a required task when slow and large models are used, for example, for classification, but there are transmissions, space, time or computing capabilities constraints that have to be fulfilled. Multilayer Perceptron (MLP) models have been traditionally used as classifiers. Depending on the problem, they may need a large number of parameters (neuron functions, weights and bias) to obtain an acceptable performance. This work proposes a technique to compress an MLP model preserving, at the same time, its classification performance, through the kernels of a Volterra series model. The Volterra kernels can be used to represent the information that a Neural Network (NN) model has learnt with almost the same accuracy but compressed into less parameters. The Volterra-NN approach proposed in this work has two parts. First of all, it allows extracting the Volterra kernels from the NN parameters after training, which will contain the classifier knowledge. Second, it allows building different orders Volterra series model for the original problem using the Volterra kernels, significantly reducing the number of neural parameters involved to a very few Volterra-NN parameters (kernels). Experimental results are presented over the standard Iris classification problem, showing the good Volterra-NN model compression capabilities. Mariano Rubiolo, Georgina Stegmayer, Diego H. Milone |
IJCNN | 3 |
| 2010 | *omeSOM: a software for clustering and visualization of transcriptional and metabolite data mined from interspecific crosses of crop plantsabstractBACKGROUND: modern biology uses experimental systems that involve the exploration of phenotypic variation as a result of the recombination of several genomes. Such systems are useful to investigate the functional evolution of metabolic networks. One such approach is the analysis of transcript and metabolite profiles. These kinds of studies generate a large amount of data, which require dedicated computational tools for their analysis. RESULTS: this paper presents a novel software named *omeSOM (transcript/metabol-ome Self Organizing Map) that implements a neural model for biological data clustering and visualization. It allows the discovery of relationships between changes in transcripts and metabolites of crop plants harboring introgressed exotic alleles and furthermore, its use can be extended to other type of omics data. The software is focused on the easy identification of groups including different molecular entities, independently of the number of clusters formed. The *omeSOM software provides easy-to-visualize interfaces for the identification of coordinated variations in the co-expressed genes and co-accumulated metabolites. Additionally, this information is linked to the most widely used gene annotation and metabolic pathway databases. CONCLUSIONS: *omeSOM is a software designed to give support to the data mining task of metabolic and transcriptional datasets derived from different databases. It provides a user-friendly interface and offers several visualization features, easy to understand by non-expert users. Therefore, *omeSOM provides support for data mining tasks and it is applicable to basic research as well as applied breeding programs. The software and a sample dataset are available free of charge at http://sourcesinc.sourceforge.net/omesom/. Diego H. Milone, Georgina Stegmayer, Laura Kamenetzky, Mariana G. Lopez, Je Min Lee, James J. Giovannoni, Fernando Carrari |
BMC Bioinform. | 1 |
| 2010 | Denoising and recognition using hidden Markov models with observation distributions modeled by hidden Markov trees
Diego H. Milone, Leandro E. Di Persia, María Eugenia Torres |
Pattern Recognit. | 1 |
| 2010 | Minimum classification error learning for sequential data in the wavelet domain
Diego Tomassi, Diego H. Milone, Liliana Forzani |
Pattern Recognit. | 2 |
| 2009 | Artificial Life Contest - A Tool for Informal Teaching of Artificial Intelligence
Diego H. Milone, Georgina Stegmayer, Daniel Beber |
CSEDU (1) | 1 |
| 2009 | Neural network model for integration and visualization of introgressed genome and metabolite dataabstractThe volume of information derived from post-genomic technologies is rapidly increasing. Due to the amount of data involved, novel computational models are needed for introducing order into the massive data sets produced by these new technologies. Data integration is also gaining increasing attention for merging signals in order to discover unknown pathways. These topics require the development of adequate soft computing tools. This work proposes a neural network model for discovering relationships between gene expression and metabolite profiles of introgressed lines. It also provides a simple visualization interface for identification of coordinated variations in mRNA and metabolites. This may be useful when the focus is on the easily identification of groups of different patterns, independently of the number of formed clusters. This kind of analysis may help for the inference of a-priori unknown metabolic pathways involving the grouped data. The model has been used on a case study involving data from tomato fruits. Georgina Stegmayer, Diego H. Milone, Laura Kamenetzky, Mariana G. Lopez, Fernando Carrari |
IJCNN | 2 |
| 2009 | Indeterminacy Free Frequency-Domain Blind Separation of Reverberant Audio SourcesabstractBlind separation of convolutive mixtures is a very complicated task that has applications in many fields of speech and audio processing, such as hearing aids and man-machine interfaces. One of the proposed solutions is the frequency-domain independent component analysis. The main disadvantage of this method is the presence of permutation ambiguities among consecutive frequency bins. Moreover, this problem is worst when reverberation time increases. Presented in this paper is a new frequency-domain method, that uses a simplified mixing model, where the impulse responses from one source to each microphone are expressed as scaled and delayed versions of one of these impulse responses. This assumption, based on the similitude among waveforms of the impulse responses, is valid for a small spacing of the microphones. Under this model, separation is performed without any permutation or amplitude ambiguity among consecutive frequency bins. This new method is aimed mainly to obtain separation, with a small reduction of reverberation. Nevertheless, as the reverberation is included in the model, the new method is capable of performing separation for a wide range of reverberant conditions, with very high speed. The separation quality is evaluated using a perceptually designed objective measure. Also, an automatic speech recognition system is used to test the advantages of the algorithm in a real application. Very good results are obtained for both, artificial and real mixtures. The results are significantly better than those by other standard blind source separation algorithms. Leandro E. Di Persia, Diego H. Milone, Masuzo Yanagida |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Perceptual evaluation of blind source separation for robust speech recognition
Leandro E. Di Persia, Diego H. Milone, Hugo Leonardo Rufiner, Masuzo Yanagida |
Signal Process. | 2 |
| 2007 | Objective quality evaluation in blind source separation for speech recognition in a real room
Leandro E. Di Persia, Masuzo Yanagida, Hugo Leonardo Rufiner, Diego H. Milone |
Signal Process. | 4 |
| 2003 | Prosodic and accentual information for automatic speech recognitionabstractVarious aspects relating to the human production and perception of speech have gradually been incorporated into automatic speech recognition systems. Nevertheless, the set of speech prosodic features has not yet been used in an explicit way in the recognition process itself. This study presents an analysis of prosody's three most important parameters, namely energy, fundamental frequency and duration, together with a method for incorporating this information into automatic speech recognition. On the basis of a preliminary analysis, a design is proposed for a prosodic feature classifier in which these parameters are associated with orthographic accentuation. Prosodic-accentual features are incorporated in a hidden Markov model recognizer; their theoretical formulation and experimental setup are then presented. Several experiments were conducted to show how the method performs with a Spanish continuous-speech database. Using this approach to process other database subsets, we obtained a word recognition error reduction rate of 28.91%. Diego H. Milone, Antonio J. Rubio |
IEEE Trans. Speech Audio Process. | 1 |
| 2002 | Evolutionary algorithm for speech segmentationabstractSpeech segmentation is one of the problems in the speech processing area. The main techniques that attempt to solve it are manual segmentation and hidden Markov model alignment. In this work a new technique based on an evolutionary algorithm that permits to segment the speech without a previous training process is presented. Diego H. Milone, Juan Julián Merelo Guervós, Hugo Leonardo Rufiner |
IEEE Congress on Evolutionary Computation | 1 |
| 2001 | A new technique based on augmented language models to improve the performance of spoken dialogue systemsabstractThis paper presents a new technique that aims to improve the performance of spoken dialogue systems by using the so-called augmented language models. We define an augmented language model as a compound of a language model and a set of values concerning parameters that can influence the speech recognition when the language model is used. The diverse language models used by a dialogue system can be very different, in terms of perplexity for example. Then, the aim of the technique is to find and use the combination of values concerning the different parameters that leads to the best recognition results when the different language models are used by a dialogue system. The technique has been applied to a dialogue system for the fast food domain. The results show that when the augmented language models are used the system’s performance is enhanced. In the experiments we have achieved a reduction of 9,33% in the word error rate and an increment of 11,26% in the sentence understanding. Ramón López-Cózar, Diego H. Milone |
INTERSPEECH | 2 |